Evals that tell the truth
My first result said +30%. I didn't trust a number that good, found the artifact, and reported +12.9%.
The reranker trial
I tested an LLM reranker on the Tamago matcher: about 48,000 judgements for $2.31. My first read said +30%. A number that good made me suspicious; it turned out to be an artifact of how I'd split the data. The honest holdout result was +12.9% recall@20. A bias audit over 540 probes found a maximum difference of 0.09. Cost: about $0.008 per search.
Choosing a model by measurement
For RefineCV's CV-editing assistant I ran a 24-model, 61-scenario bake-off over 3,162 live turns. The winner scored 0.982 against the incumbent's 0.907.
Why this matters
Most LLM features fail quietly. The work I'm proudest of is building the golden sets, judges and holdouts that make a number trustworthy, and reporting the smaller number when it's the true one.