Multilingual preference & critique for post-training
The client needed consistent human preference and critique data across several languages to improve a general-purpose assistant — and their internal team couldn’t hold quality steady as volume grew.
Recruiting, qualification, rubric training, calibration, production, layered QA and adjudication — run as a managed program with dedicated leads.
Pairwise preference judgments with documented rationale, structured response critiques, and rewrite tasks, across multiple languages and domains.
Gold tasks and calibration sets per language; inter-annotator agreement tracked per batch; disagreements routed to senior adjudication.
Access-controlled environment, workforce NDAs, and data handling per the client’s requirements.
A stable, repeatable calibration loop the client’s research team could iterate their rubric against — with agreement held steady as the program scaled.