Skip to main content

A frontier AI lab (anonymized) · Foundation Model Labs · Text

Multilingual safety-eval for a frontier AI lab

A native-speaker red-team and safety-evaluation set across ten Indic languages that surfaced harms the lab's English-only guardrails missed.

By Cognegica Data Operations · Field collection & delivery team

11 harm categories · IAA 0.86

Languages: Hindi, Tamil, Telugu, Bengali, +6 Indic

Illustration representing RLHF and model alignment

Anonymised. Client identity is withheld at their request. The methods, gates, and metrics are real.

Challenge

The lab's flagship model passed English safety benchmarks but had no credible way to measure — let alone catch — culturally-specific harms in Indic languages. Caste slurs, communal dog-whistles, regional abuse and code-mixed toxicity slipped past guardrails trained on translated Western taxonomies.

They needed an evaluation set that reflected how harm is actually expressed across ten Indic languages, with measured agreement they could defend to their own safety and policy teams.

Approach

We built a native-speaker red-team and safety-eval program rather than translating an existing benchmark:

  • Recruited vetted native-speaker linguists across ten languages under written contributor agreements.
  • Co-designed a culturally-grounded harm taxonomy (11 categories) with the lab's policy team, piloted before scale.
  • Authored adversarial prompts from scratch with escalation rounds and a cultural-plausibility review — no scraped forum data.
  • Ran two-pass QA with adjudication; reported Krippendorff's α per category.

Outcome

The lab received a documented evaluation set with inter-annotator agreement of 0.86 across 11 harm categories and a post-adjudication QA pass rate above 98%. It exposed material gaps in Hindi and Tamil guardrail coverage that English benchmarks had hidden, and became a recurring quarterly refresh as new harm patterns emerged.

Representative outcome from an anonymized engagement; metrics are illustrative of our standard quality bar.

How this maps to what we do

The services and data behind this engagement

This outcome was delivered with the same rights-cleared, documented services and datasets you can engage today.

Measure the harms your benchmarks can't see.

See how we structure engagements and indicative pricing, or tell us your languages, modalities and quality bar for a scoped quote.

Written by

Cognegica Data Operations

Field collection & delivery team

Cognegica Data Operations is the internal team responsible for field-grade data collection, contributor recruitment, consent and delivery across our multilingual programs. This is an editable team identity — a named individual with a public profile can be assigned to it later in the admin.

Run a similar pilot.

Talk to a Language PM