Recent Projects
Ecosystem Mapping: Classifying Endangered Plant Communities to Quantify the Ecological Cost of Data Centers
Independent Ecological Data Science Research
Ecosystem classification at the association level—the most precise and ecologically meaningful scale—is often costly, time-consuming, and incomplete, which means conservation resources may be misallocated and protective measures can lag behind ecological changes. This has become urgent as data center expansion creates substantial, localized demand for water and land, often diverted from ecosystems that depend on it. This project evaluates whether Earth Observation data, combined with California's vegetation datasets, can be used to train models that map ecosystems at the association level, producing a baseline for comparing infrastructure siting decisions against ecological risk.
Informed Seattle: Collective Sensemaking Infrastructure with AI Supported Legislative Plain Text Summaries
Civic Information Infrastructure for Local Democracy
In collaboration with UW eSciences Institute
Democratic participation is unequally distributed because civic information is unequally accessible. Every week, local governments make decisions that shape housing affordability, transportation, public safety, climate resilience, disability services, childcare, and public health. While these decisions are technically public, they are rarely understandable to the people most affected by them. Residents with fewer resources—working parents, renters, immigrants, people with disabilities, young adults, and those unfamiliar with legal or bureaucratic language—often cannot afford the time or expertise required to interpret legislation before decisions are made. As a result, public participation disproportionately reflects organizations and individuals who already possess the time, knowledge, and relationships to navigate government. This creates a feedback loop where...
Counting the Uncounted: Privacy-Preserving Linkage for Public Health
U.S. Census Bureau Emerging Technology (xD) Fellowship / eHealth and Northwestern University CAPriCORN network
Public health research depends on linking people across datasets—connecting someone's record to their Census responses to see how neighborhood conditions shape chronic disease—but the infrastructure to do this safely doesn't exist at scale, and the identifiers most linkage methods rely on are least available for the unhoused, low-income, and undocumented people the research is meant to serve. This project piloted VaultDB, a secure multiparty computation tool that lets the Census Bureau and Northwestern's ORION network compute across ACS and EHR data without either side exposing its records, and recommended moving from probabilistic matching with human review to a deterministic or hybrid method that avoids disclosing patient identity. The technical path...
Remote Sensing–ML Approach for Household Wealth Index Estimation
U.S. Census Bureau Emerging Technology (xD) Fellowship / SEHSD
In collaboration with the AI & Global Development Lab
Existing tools for understanding economic wellbeing at fine spatial resolution are constrained by fundamental survey design limits. Traditional small area estimation methods become statistically unreliable below the block group level, carry high compliance overhead, and can take years to reflect ground conditions. This project developed a geo-temporal Earth observation and machine learning pipeline to predict average material wealth at the 1km grid level—framed as a gap measure between absolute income and relative cost of living—enabling program targeting, infrastructure planning, disaster preparedness, and causal analysis of policy treatment effects at the neighborhood level. The pipeline proceeded through complete data preparation phases in R and Python, with Google Earth Engine satellite image...
Responsible AI in National Statistical Agencies
Challenges in applying responsible AI principles in daily practice and what practitioners require to address these gaps.
U.S. Census Bureau Emerging Technology (xD) Fellowship
Responsible AI is guided by principles in policy documents, NIST frameworks, and Executive Orders, but a persistent gap separates those principles from daily practice—visible in decisions like disclosure avoidance under deadline pressure, imputation choices for small datasets, and feature engineering that doesn't fully account for bias. This project conducted qualitative research with practitioners at eight federal statistical agencies, tracing where responsible AI principles break down at each stage of the data lifecycle—from raw data through deployment—and translating those findings into practitioner-focused recommendations: a policy navigator, shared evaluation rubrics with bias thresholds, community and domain-expert review, tiered model governance, and post-deployment feedback loops.
Responsible Feature Engineering: Advancing Fairness from Post-Hoc Audit to a Design Discipline
U.S. Census Bureau Emerging Technology (xD) Fellowship
With Atul Rawal
Most AI fairness interventions occur after model development, through output audits that address harms only after they've affected people—but by the time a model generates predictions, the decisions about data inclusion and feature selection that embedded the bias have already been finalized. This project, with Atul Rawal, developed Responsible Feature Engineering (RFE), a model-agnostic framework that evaluates whether a feature's influence varies across the demographic groups a model affects, surfacing proxy variables like latitude and longitude (which stand in for race via residential segregation) before they reach training. RFE doesn't automate the resulting judgment call—it makes an implicit, often invisible decision explicit and puts it in front of domain experts...
Measuring the Societal Impacts of AI: The MIDAS Initiative
U.S. Census Bureau Emerging Technology (xD) Fellowship
Algorithmic systems make consequential decisions about people at scale—screening job applicants, setting credit eligibility, allocating housing and benefits—but the government has no systematic, nationally representative picture of where AI is deployed, how people experience it, or whether those affected even know it's happening. This project helped design measurement infrastructure from two directions: Household Pulse Survey questions that ask about behaviors first rather than gating respondents on AI literacy, and an analysis framework that reframes the driving question from how many people use AI to where AI compounds privacy violations, who is screened without being told, and how exposure differs by race, income, and geography.
Environmental Data Landscape Scan for Federal Climate Policy
U.S. Census Bureau Emerging Technology (xD) Fellowship / SEHSD
The federal government lacks a coherent environmental data framework—a shared understanding of what climate and ecological data exists, who needs it, and how it connects environmental conditions to the people and places affected. This project ran a structured landscape scan and stakeholder consultation across Census's Economic and Housing Statistics Division, the Environmental Impacts Frame initiative, and the interagency natural capital accounting effort, producing a stakeholder map and research questions designed to be picked up by whoever secures funding or an institutional home for the next phase.
Recurve Market Access Reporting
Data product for CPUC oversight of $150M in clean-energy program funding
The California Public Utilities Commission oversees $150 million in Market Access program funding meant to accelerate the shift to a clean-energy economy through utility-run, demand-side programs—but funding was allocated using estimated savings rather than verified results, risking both ratepayer accountability and the low-income and hard-to-reach households the programs were meant to serve. This project built a configurable Market Access reporting pipeline that measures outcomes directly from meter data, developed through weekly collaboration with business stakeholders as CPUC requirements evolved. A dbt/LaTeX/Looker abstraction layer made the pipeline portable across utility clients with different datasets, cutting new-client deployment effort 30x, while anonymized building attribute data let energy efficiency companies evaluate performance across...