← Rishi Sangare
LD Technologies · for Tamago

LLM matching for a Japanese recruiting platform

Recruiters type a job description and get a ranked shortlist with reasons, in English and Japanese, fast enough to use live.

Backend / LLM engineer on a team of 3Oct 2025 – now
0.19 → 0.57recall@10, same model, honest golden set
85% → 0%broken clarifying questions
16/16 → 0/16English leaking into Japanese sessions
253commits · 128 PRs

The problem

A large Japanese recruiting database wanted recruiters to type a job description, or pick a candidate, and get a ranked shortlist with reasons. It had to work in English and Japanese and be fast enough to use live.

What I built

  • The service foundation: FastAPI, Docker, Caddy/TLS, GitHub Actions to GHCR to staging and production, with Slack alerts.
  • Retrieval: Elasticsearch (kuromoji for Japanese) with wage/currency, language, age, experience and excluded-company filters, routed per tenant.
  • Interactive pre-screening, in two flows: the LLM asks clarifying questions and a deal-breaker question, classifies the answers in parallel, then runs retrieval and parallel LLM evaluation in the background and calls back with an HMAC-signed webhook. Sessions are idempotent, with status guards and TTLs.
  • Reliability: a primary and fallback LLM provider chain hardened against malformed JSON, which removed a whole class of "all evaluations failed" 500/503 errors. A 4-agent robustness audit found 8 issues; all were fixed.

The measurement work

  • Request filters were silently deleting 68% of the "ideal" answers in the golden set. I rebuilt it to be filter-aware. Same model outputs: recall@10 went from 0.19 to 0.57 in English, hits@10 from 2.2 to 6.6 in Japanese.
  • Clarifying questions were often degenerate: 85% → 0%. Visa over-triggering 90% → 0%. Options per question 3.3 → 5.9. Verified on 72 live scenarios.
  • Elasticsearch boost tuning gained 6.8 points on train and under 1 on holdout. It didn't generalize, so I said so: retrieval depth was the real lever.
  • An open-source reranker scored recall@20 0.153 against the paid one's 0.536, so I recommended not switching yet.

Stack

PythonFastAPIPydanticElasticsearch 8/9SupabaseCerebrasOpenRouterAWS BedrockDockerCaddyGitHub Actionspytest
nextEvals that tell the truth →