← Rishi Sangare
Across Tamago and RefineCV

Evals that tell the truth

My first result said +30%. I didn't trust a number that good, found the artifact, and reported +12.9%.

Designed and ran the evaluations2025 – 2026
+12.9%honest holdout recall@20
$2.31for about 48,000 LLM judgements
3,162live turns in a 24-model bake-off
0.982 vs 0.907winner vs incumbent

The reranker trial

I tested an LLM reranker on the Tamago matcher: about 48,000 judgements for $2.31. My first read said +30%. A number that good made me suspicious; it turned out to be an artifact of how I'd split the data. The honest holdout result was +12.9% recall@20. A bias audit over 540 probes found a maximum difference of 0.09. Cost: about $0.008 per search.

Choosing a model by measurement

For RefineCV's CV-editing assistant I ran a 24-model, 61-scenario bake-off over 3,162 live turns. The winner scored 0.982 against the incumbent's 0.907.

Why this matters

Most LLM features fail quietly. The work I'm proudest of is building the golden sets, judges and holdouts that make a number trustworthy, and reporting the smaller number when it's the true one.

Stack

Golden setsrecall@kLLM-as-judgeholdout splitsbias probesPython
nextRefineCV →