Retrieval, ranking, and the agentic layers that act on them — productionized and A/B-tested, not just prototyped. Plus the post-training and evaluation behind them.
Every project below links to running code.
I build, train, and evaluate LLM and recommendation systems — post-training, agentic pipelines, retrieval, and the infrastructure behind them.
The keystone build runs live in your browser right now — plus a full post-training lab and a self-improving training loop with evaluation built in.
Production-shaped LLM systems — retrieval, routing, agents, observability, and inference.
Hypothesis → experiment → measurement. Each project asks a concrete question and reports what the data actually says.
Earlier interactive builds and data-visualization work — plus the full portfolio archive.
Tools others can install and run — plus earlier research software licensed by UC / LBNL.
Twenty interactive essays, organized as five short series — read any one alone, or follow a series top to bottom. Earlier long-form work lives on Medium.
Why long tasks fail and the machinery that fixes it — pacing, memory, forgetting, rollback, and the loop itself.
Evaluation and trust — from grading traces to catching gamed metrics to deciding when an agent may act alone.
The surfaces agents touch and produce — tool contracts, visualization as infrastructure, and answers that render.
Where the durable value sits — the harness over the model, domain depth over generality, and the real cost of more agents.
The modeling layer underneath — retrieval systems and explaining what the ranker did.
Where the models meet production — experience across industry research and the infrastructure that scales AI systems.
The systems work behind the models. At Intel / Intuit I help build DeepInsight, a real-time ML analytics engine — streaming pipelines on Kafka and Apache Flink, services in C++ / GoLang / Python on Kubernetes, with ML anomaly detection. The same instinct runs through my open-source work: distributed training (DeepSpeed / FSDP / DDP), inference serving, and agent observability — infrastructure in service of shipping and scaling AI systems.
Peer-reviewed research — modelling data through mathematical models and visualization, with a Best Paper award and a patent.
Organized for AI research engineering. Highlighted skills are where I do my deepest work.
The fastest ways to reach me and see the work.
I'm always happy to talk about LLM post-training, agent evaluation, RAG, and the infrastructure that scales AI systems. Email is best — or browse the code on GitHub.