Security & Trust
Case Studies
Talk to an AI Data Expert
Case Studies

Programs, written as procurement evidence.

Real engagement patterns, described honestly. Client identities and any confidential performance metrics are withheld under NDA — what remains is a straight account of what we were asked to do and how we delivered it.

Frontier LLM Developer · Identity withheld under NDA

Multilingual preference & critique for post-training

Challenge

The client needed consistent human preference and critique data across several languages to improve a general-purpose assistant — and their internal team couldn’t hold quality steady as volume grew.

OmniTech role

Recruiting, qualification, rubric training, calibration, production, layered QA and adjudication — run as a managed program with dedicated leads.

Work performed

Pairwise preference judgments with documented rationale, structured response critiques, and rewrite tasks, across multiple languages and domains.

Quality approach

Gold tasks and calibration sets per language; inter-annotator agreement tracked per batch; disagreements routed to senior adjudication.

Security

Access-controlled environment, workforce NDAs, and data handling per the client’s requirements.

Outcome

A stable, repeatable calibration loop the client’s research team could iterate their rubric against — with agreement held steady as the program scaled.

Client identity and program metrics withheld pursuant to confidentiality obligations.
Generative Media Platform · Identity withheld under NDA

Multimodal generation-quality evaluation

Challenge

A fast-moving image and video generation product needed human evaluation of output quality — prompt adherence, visual artifacts and safety — that could keep pace with frequent releases.

OmniTech role

Stood up and qualified a specialist evaluation team, designed the scoring rubric with the client, and ran calibrated production with QA.

Work performed

Rubric scoring and pairwise comparison of generated images and video for adherence, coherence and artifacts, plus safety flags.

Quality approach

Calibration against reference examples, embedded gold checks, and secondary review on flagged and borderline outputs.

Security

Work performed in an approved environment with role-scoped access and confidentiality controls.

Outcome

A calibrated review pipeline that scaled with release cadence while holding a consistent quality bar the product team could rely on.

Client identity and program metrics withheld pursuant to confidentiality obligations.
AI Evaluation Company · Identity withheld under NDA

Coding & STEM expert evaluation

Challenge

The client needed a defensible expert layer to evaluate model reasoning and code correctness — beyond what non-specialists could reliably judge.

OmniTech role

Sourced and qualified coding and STEM specialists, built the assessment and calibration, and managed production with adjudication.

Work performed

Structured evaluation of code correctness, reasoning quality and step-level process against detailed rubrics.

Quality approach

Specialist qualification tests, calibration against worked solutions, and senior adjudication with traceable rationale.

Security

Confidential handling of evaluation sets and model outputs under NDA and access controls.

Outcome

An expert-review layer with documented disagreement resolution and reviewer coaching the client could stand behind.

Client identity and program metrics withheld pursuant to confidentiality obligations.
Enterprise AI Product Team · Identity withheld under NDA

Safety red teaming & policy evaluation

Challenge

Ahead of a release, the client needed structured adversarial testing and policy evaluation against their usage standards.

OmniTech role

Qualified and briefed analysts, translated client policy into an evaluation rubric and harm taxonomy, and ran managed red teaming.

Work performed

Adversarial and jailbreak attempts, harmful-content categorization, and policy pass/fail assessment with examples.

Quality approach

Calibrated categorization, secondary review on severe cases, and adjudication for contested judgments.

Security

Sensitive-content handling policies, access controls and attention to analyst wellbeing.

Outcome

Structured findings and examples that fed directly into the client’s guardrails, policies and training data before launch.

Client identity and program metrics withheld pursuant to confidentiality obligations.
A note on evidence

Why you won’t see invented numbers here.

Many vendor case studies quote precise figures no one can verify. We’ve chosen not to. Where we’re contractually able to share verified program metrics, we’ll do so with your team directly, under NDA. Everything above is a truthful account of engagement patterns — the kind of evidence a procurement team can actually rely on.