Build healthcare data infrastructure that supports clinicians, clinic managers, and a team of 20+ engineers, data scientists, and analysts.
- Built a Databricks platform with configuration-driven source onboarding and centralized data-quality management
- Led delivery of incremental sync and major GraphQL API components over PostgreSQL, bringing claims and quality scores to clinicians and outreach priorities to clinic managers
- Built an enterprise master patient index for 8M patients as sole engineer, using Splink, deterministic matching constraints, match review, and replayable match history
- Mentor six engineers on CI/CD and software design, and coach analysts and data scientists daily on platform usage and software engineering practices
- Lead incident response and root-cause analysis, turning findings into regression tests and platform fixes
- Co-developed a clinical NLP pipeline with the lead data scientist, using OCR and third-party models to extract diagnoses, procedures, and labs from encounter PDFs for clinician review
Data Engineer at Medisolv
Jul 2023 - Jul 2025
Owned patient-data ingestion and transformation for hundreds of hospital clients, processing 10B+ records weekly into a 150TB+ data lake.
- Migrated 1,500+ Azure Data Factory pipelines from the portal to source control, adding code review, rollback, one-button deployments, and Pydantic/pytest validation
- Automated client onboarding with Python and the Azure SDK, reducing setup from over a day of manual work to a five-minute deployment
- Open sourced sparkparse to expose exploding joins and memory pressure; guided rewrites cut tasks from hours to 10–20 minutes on the same compute
- Built monitoring with Azure Event Hubs and Power BI that correlated task start/stop events to detect stalled tasks during execution rather than hours later
Built analytics and quality-reporting workflows for NewYork Quality Care, the accountable care organization of NewYork-Presbyterian, Weill Cornell, and Columbia.
- Cut Tableau load times from 2–3 minutes to under 10 seconds with pre-aggregated summary tables and a restructured star schema
- Built the ELT pipeline behind dashboards and measure reporting, processing 100M+ records a week for 35,000 Medicare patients
- Built address-cleaning and geocoding workflows that prioritized free Census matches and routed low-quality results to paid Esri ArcGIS matching for patient-population mapping
- Modeled CMS quality scores against sample size to determine reporting depth within the ranked patient sample for a Medicare shared-savings program
Data Engineer (contract) at UTHealth
May 2020 - Sep 2023
Operated the daily data pipeline for UTHealth’s Texas COVID-19
dashboard, integrating state and third-party sources with unit tests and Slack monitoring for data-quality issues.
- Authored thesis (100+ citations) on biomarkers of traumatic brain injury (TBI) and provided data support for other research efforts
- Applied variable selection on hundreds of biomarker combinations to identify TBI predictors
- Built analysis pipelines and created publication-ready data visualization in R