LLM Evaluation Infrastructure
Lead engineer
Solo-built evaluation infrastructure for production LLM features — a labeled corpus, human-graded ground truth, LLM-as-judge scoring, per-call cost tracing and version-controlled prompts.
- typescript
- llms
- ai-agents
- prompt-engineering
- observability
The problem
Once an LLM feature ships, "it looked good in a demo" stops being a useful signal. Prompts drift, models get swapped, and without a way to measure quality on a fixed corpus, every change becomes an argument instead of a number.
What I built
I built the evaluation infrastructure that turns "does this still work" into something measurable, on my own, end to end:
- A labeled corpus of real cases with human-graded ground truth, so changes have something fixed to compare against.
- LLM-as-judge scoring for the cases where a strict string match does not capture correctness, calibrated against the human-graded set rather than trusted blindly.
- Per-call cost tracing, so a quality improvement is judged against what it costs to run, not in isolation.
- Version-controlled prompts, so a regression can be bisected the same way a code regression can.
- A judge model deliberately resolved from a different provider than the agent under test — a provider grading its own output tends to be biased, so I moved model selection out of environment variables to make an opposing-provider judge the default.
- A file-naming convention that isolates model-dependent code into a separate CI lane, so real-model evals only run when a prompt or agent actually changes instead of on every pull request.
Why it mattered
This is the piece that made every other LLM feature — the document extraction pipeline, message classification — accountable to a number instead of a feeling. It is also the project I point to for "owns a large-scope effort end to end": nobody assigned the scope, I decided what mattered and built it.