Machine Learning Engineer

I turn research into models that ship.

Applied ML engineer working on retrieval, evaluation and low-latency inference. I take ideas from a paper to production traffic without losing the science.

LLMsRetrievalPyTorchEvaluation
Neural network visualization

180ms

p95 inference latency in production

12M

daily requests served by my systems

3

first-author workshop papers

Selected projects

Retrieval gateway for a support assistant

Retrieval gateway for a support assistant

Designed a hybrid retrieval layer (BM25 plus dense) with a reranker, feeding a grounded answer model for a customer support product.

Cut hallucinated answers by 41% and raised deflection rate from 28% to 52%.

PythonPyTorchFAISSvLLMRay
Offline evaluation harness for LLM releases

Offline evaluation harness for LLM releases

Built a reproducible eval harness scoring every model candidate on a curated suite before it reaches traffic, with regression gates in CI.

Blocked 9 quality regressions pre-release and made model swaps a one-day decision.

PythonLangChainWeights & BiasesGitHub Actions
Low-latency inference service

Low-latency inference service

Rewrote a batched inference service with continuous batching and quantized weights to hold latency under load.

Held p95 under 200ms at 12M requests per day while cutting GPU cost 34%.

RustTritonCUDAKubernetes

Skills

Modeling

PyTorchTransformersRAGFine-tuningRLHF

Serving

vLLMTritonTensorRTQuantizationRay

Data & Eval

PandasSparkLangChainRagasW&B

Platform

PythonRustDockerKubernetesAWS

Experience

  1. Since 2022

    Senior ML Engineer

    Aperture AI

    Own the retrieval and evaluation stack for a support assistant used by millions; lead a squad of four.

  2. 2020 to 2022

    Machine Learning Engineer

    Loop Research

    Shipped the first production recommendation model and the team's model serving platform.

  3. 2018 to 2020

    Research Engineer

    University Vision Lab

    Published on efficient inference and maintained the lab's training infrastructure.

Publications

  • Continuous Batching for Low-Latency LLM ServingMLSys Workshop, 2024
  • A Reproducible Evaluation Harness for RAG SystemsEMNLP Industry Track, 2023
  • Distilling Rerankers for Edge InferencearXiv preprint, 2022

Certifications & talks

Efficient LLM Serving at Scale

MLSys Workshop, talk

2024

AWS Machine Learning, Specialty

Amazon Web Services

2023

Deep Learning Specialization

DeepLearning.AI

2021

Let's build something that ships

Open to senior and staff ML roles, and to research collaborations on retrieval and evaluation.