MMIP Policy Tracker

Legislative intelligence for the fight against MMIP

Role. Design Engineer & Project Lead — product definition through implementation across four iterations: automating the analysis, then automating search and discovery across all 50 states and federal sources, then the UI engineering to turn 275 tracked policies into a public dashboard, with community UX research defining what comes after. Built in service of UIC's MMIP advocacy mission.

Design Engineer · Project Lead · 2024–present
Built for Urban Indigenous Collective's MMIP Policy Tracker.

AI Policy Analyzer on MacBook — MMIP legislative data tool for Urban Indigenous Collective

MMIP data is life data

In 2021, Urban Indigenous Collective initiated a policy tracker to monitor all state-wide and national MMIP-related legislation in the United States. Its goals were concrete: demonstrate the current state of the crisis, track progress toward supporting survivors, and expose which jurisdictions are neglecting the issue entirely.

Behind that mission is a deeper injustice. Indigenous communities face a legacy of data genocide — an intentional erasure through lack of data collection and insufficient funding for new research. This erasure has long suppressed visibility for Indigenous issues, and nowhere is that more damaging than in the MMIP crisis, where absence of data translates directly into absence of policy response.

Addressing that gap wasn't just a technical problem. It was a strategic one: how do you build a system that generates culturally relevant data at scale, fast enough to keep pace with a legislative calendar?

A tangle of conflicting data

The first obstacle was the data itself. UIC's existing MMIP tracker contained conflicting datasets — years of manual entry with no standardized format. Fields like "Sponsors" and "Indigenous Sponsors" had no systematic way to identify which sponsors were Indigenous, leading to underreporting and inaccurate records. Team members had to rely on name recognition or manual research to fill the gap.

Mixing qualitative and quantitative data in the same spreadsheet blocked any kind of automated reporting. Over 25 columns required manual updates, and without a single source of truth, each dataset contradicted the others. The tracker was technically functional — but practically unusable for the kind of rapid, scalable analysis UIC needed.

By leveraging ChatGPT and automating legislative data analysis, we are addressing this data genocide head-on — giving Indigenous communities a way to reclaim and generate culturally relevant data at scale.

Iteration one: automating the analysis

One unified dataset

The first move was merging the conflicting datasets into a single Excel document — reconciling years of inconsistency without losing any critical information. From there, the tracker migrated to AirTable, which provided the standardized data formats, automated visualizations, and interface builder UIC needed to turn raw legislative data into actionable intelligence.

Knowing who is Indigenous

To solve the Indigenous sponsor identification problem, I built a standalone database of Indigenous politicians and legislators — sourced from Wikipedia, the only public data available — and integrated it directly into the policy analyzer. What was once a manual, error-prone research task is now automated: the system cross-references each bill's sponsor list against the database and flags Indigenous legislators without any manual lookup.

This tool exists as a standalone resource as well as an integrated component of the analyzer — available to any researcher who needs to identify Indigenous legislative representation.

An LLM as legislative analyst

The analyzer, built in Python, pulls legislative details directly from the Legiscan API and — in this first iteration — used GPT-4 to parse bill text through culturally tailored questions — generating new structured data for analysis. What once took over an hour of manual reading and annotation now takes minutes.

The AI-generated analysis is reviewed by volunteers and the MMIP Program Associate before upload, balancing speed with the human oversight that culturally sensitive data demands. Automation handles the volume; people ensure the integrity.

One entry: 130 minutes to 10

Automating the analysis changed UIC’s data work at the unit level. A single legislative entry took two hours and ten minutes of manual reading, annotation, and data entry. It now takes ten minutes — a 92% reduction. Conflicting datasets are consolidated, Indigenous sponsors are identified systematically rather than by name recognition, and what reaches AirTable is consistent and structured.

That fixed the cost of handling a bill. It did not change how many bills UIC could find.

130 min → 10 min per entry 92% reduction 25+ columns automated Legiscan API Human-reviewed outputs

Iteration two: automating discovery

Automating the analysis exposed the real bottleneck: a person was still the discovery mechanism. Someone had to know a bill existed, find it, and paste in a link. That works when you are following a handful of bills you already know about. It does not work as a public record, because you cannot publish what nobody thought to look for.

Over the first weekend of August 2026 I rebuilt the front half of the product. It now monitors all fifty state legislatures, LegiScan, and federal sources continuously — and it captures more than bills. Executive orders, proclamations, and Department of Justice initiatives all shape the MMIP policy landscape, and none of them were being tracked at all.

This is the phase where the data was actually reclaimed at scale — not by analyzing faster, but by no longer depending on a person to notice. There are now 275 policies in the dataset, ready to publish as soon as the interface can carry them.

50 state legislatures LegiScan + federal sources Bills, EOs, proclamations, DOJ 275 policies surfaced Continuous monitoring

In use

Use has spread beyond UIC. A tribal liaison inside a community outreach program at the Massachusetts Department of Health has been working from the tracker — a role that covers every tribe in Massachusetts and pulls in Connecticut and Rhode Island tribes as well.

Feedback from UIC’s MMIP Taskforce has centered on the same point: the value is in the dataset. Which is precisely what makes the next phase an interface problem rather than a data one.

Sovereign by design

A project whose premise is that Indigenous data has been erased by others cannot credibly hand that data to a third party to process. The first iteration ran its analysis through OpenAI’s API, which was the fastest way to prove the idea worked — and the wrong place for it to stay.

The analysis now runs on a Qwen model hosted locally on a 48 GB VRAM machine. Nothing about a bill, a sponsor, or a jurisdiction leaves infrastructure the community controls. AirTable, where the tracker is published, is the one remaining third-party dependency; everything upstream of it is self-hosted.

The practical benefits are real — no per-token cost drawn against grant funding, no rate limits during a legislative session, no exposure to a vendor deprecating a model mid-project. But data sovereignty was the reason for the decision, and the rest is consequence.

Self-hosted Qwen 48 GB VRAM, local No third-party inference AirTable the only external dependency

Learnings

Solving a problem well moves the bottleneck rather than removing it. Automating the analysis made each entry cheap, which immediately exposed discovery as the real constraint. Automating discovery produced 275 policies, which immediately exposed the interface as the next one. Each phase earned its successor — and planning the whole roadmap up front would have produced the wrong second phase, because the second problem did not exist until the first one was solved.

A values commitment has to be structural to count for anything. Running the first iteration through a third-party API was the fastest way to prove the idea worked, and the wrong place for it to live permanently — a project about Indigenous data sovereignty cannot rent its inference from someone else indefinitely. Moving to self-hosted inference turned the principle into a property of the system instead of a claim in a document.

A prototype that proves an idea is not a foundation. The hackathon front end did its job: it demonstrated the concept convincingly enough to justify everything built after it. Recognising that it needed replacing rather than extending — and that its real limitation was an assumption about scale, not a shortage of features — mattered more than any individual design decision that follows.

A key lesson from this project has been understanding the balance between automation and human oversight when dealing with culturally sensitive data. The tool is designed to amplify human capacity, not replace human judgment — and that distinction matters enormously when the data directly concerns communities facing a public safety crisis.

Clear data governance protocols were critical as the team expanded. As volunteers began contributing to the review process, establishing workflows that enabled participation while ensuring data accuracy and cultural integrity became as important as the technical build itself.

Built with

Built in Python, pulling legislative data from Legiscan and running culturally tailored analysis through a Qwen model hosted locally — no inference leaves community-controlled infrastructure. The Indigenous sponsor identification database draws from Wikipedia. Structured outputs flow into AirTable for visualization and team review.

Python
Qwen (self-hosted)
AirTable
Legiscan API
Wikipedia (source)

What's next

The first two phases made the data exist. The next two make it usable and accessible — and solving discovery created that problem. There are now 275 policies ready to be made public, and the front end cannot hold them. It was built by a team at a hackathon as a lightweight proof of concept, shaped around a person who already knew which bill they cared about — it was never meant to carry a browsable public record.

Iteration three is a full interface overhaul: a real dashboard, with embeddable elements so partner organizations can surface MMIP legislative data on their own sites instead of sending people here. This is the part of the work I care most about — UI engineering is the craft I keep coming back to.

Iteration four is research-led. Over the next month I am running user experience research with MMIP advocates to learn how they have actually used the tracker so far, and what would support their advocacy going forward. The candidate features are hypotheses, not decisions: a monthly digest of everything that changed, an LLM interface for asking questions of the corpus, per-state subscriptions. Which one gets built depends on what the research says.

The Indigenous sponsor identification database remains a standalone tool available to any organization researching Indigenous legislative representation.