Applied ML engineer working on retrieval, evaluation and low-latency inference. I take ideas from a paper to production traffic without losing the science.
180ms
p95 inference latency in production
12M
daily requests served by my systems
3
first-author workshop papers
Designed a hybrid retrieval layer (BM25 plus dense) with a reranker, feeding a grounded answer model for a customer support product.
Cut hallucinated answers by 41% and raised deflection rate from 28% to 52%.
Built a reproducible eval harness scoring every model candidate on a curated suite before it reaches traffic, with regression gates in CI.
Blocked 9 quality regressions pre-release and made model swaps a one-day decision.
Rewrote a batched inference service with continuous batching and quantized weights to hold latency under load.
Held p95 under 200ms at 12M requests per day while cutting GPU cost 34%.
Aperture AI
Own the retrieval and evaluation stack for a support assistant used by millions; lead a squad of four.
Loop Research
Shipped the first production recommendation model and the team's model serving platform.
University Vision Lab
Published on efficient inference and maintained the lab's training infrastructure.
MLSys Workshop, talk
2024Amazon Web Services
2023DeepLearning.AI
2021Open to senior and staff ML roles, and to research collaborations on retrieval and evaluation.