Measure what benchmarks can’t.
Automated metrics tell you a model changed. Human evaluation tells you whether it got better — for real users, on the dimensions you care about. OmniTech runs rigorous, calibrated human evaluation across models, prompts and agent trajectories.
Rigorous human judgment, applied consistently.
Rubric-based scoring
Multi-dimensional scoring against rubrics you define — helpfulness, correctness, tone, safety, format adherence and more.
Pairwise comparison
Side-by-side A/B judgments to rank models or checkpoints, with forced rationale and tie-handling policies.
Factuality & groundedness
Claim-level verification against sources: is it true, is it supported, and does it cite what it should?
Hallucination analysis
Categorize and rate fabrications, unsupported claims and overconfidence, with taxonomies your team can act on.
Benchmark & release support
Human evaluation tied to your eval sets and release cadence, so every version gets a comparable read.
Model comparison
Head-to-head evaluation across your model and alternatives, with statistically legible reporting.
Judgment for systems that take actions.
Agentic systems don’t just answer — they plan, call tools and take multi-step actions. Evaluating them needs humans who can follow a trajectory and judge whether each step, and the whole, was right.
Trajectory assessment
Step-by-step review of an agent’s plan and actions — was each decision reasonable given what it knew at the time?
Tool-call correctness
Did the agent call the right tool, with the right arguments, and interpret the result correctly?
Task completion
Judgment on whether the end-to-end task actually succeeded — not just whether it looked done.
Long-horizon & multi-turn
Evaluation of coherence, recovery from errors and goal-tracking across long, multi-turn interactions.
Evaluation you can defend in a research review.
A quality signal is only useful if it’s reliable. We calibrate evaluators, embed gold checks, track inter-rater agreement, and route disagreement to senior adjudication — so your scores reflect the model, not the mood of the crowd.
How quality is measuredSignals we report
We report methodology and the signals above. We don’t publish invented accuracy figures — your numbers come from your program.
Run a calibrated human evaluation.
Send us your rubric or let us help build one. We’ll qualify evaluators, calibrate them, and return scores you can trust.