I’m a Senior Clinical Context Engineer at Doctronic, where I build clinical AI and the systems that measure it.

Day to day that means two things. On architecture, I design capabilities inside a multi-agent medical harness and the deterministic gates that keep clinical claims grounded. On evaluation, I build the infrastructure that decides whether any of it is safe to ship.

Producing a plausible clinical agent is cheap now. Knowing whether it works is not, and in medicine that gap is the whole problem. Most of what I’ve learned this year came from finding defects in the instrument rather than the model: a simulated patient that told the AI doctor it was a test case and inflated our benchmark scores, a routing signal with no schema behind it, per-case clinical data that never reached the prompt it was meant to populate. The number you’re about to trust is usually the believable one nobody predicted.

Before Doctronic I was a Med-Peds Infectious Diseases fellow at NIAID/CNH, doing genomics and EHR informatics research at the NHGRI Center for Precision Health Research under Josh Denny’s mentorship. I studied how genetics and clinical factors shape susceptibility to infectious disease using the NIH All of Us Research Program, including the first multi-component computable phenotype for respiratory viruses.

I’m board-certified in Internal Medicine, Pediatrics, and Adult Infectious Diseases. The clinical training is what lets me write the criteria, implement the agent, and then judge it.

This blog documents real computational work with frontier models — actual workflows, actual code, and the failures worth reporting.

Outside work: family adventures with my wife and girls, cycling, yoga, and a lot of live music.

Résumé (PDF) Full CV