Skip to main content

A foundation-model team (anonymized) · Foundation Model Labs · Text

RLHF / preference-data program for Indic alignment

A native-speaker preference / DPO data program calibrated for honorifics, formality and cultural fit — with a documented eval set to measure cultural-reasoning lift.

By Cognegica Data Operations · Field collection & delivery team

~2,000 preference pairs · IAA ≥ 0.85

Languages: Tamil, Hindi, Telugu

Illustration representing RLHF and model alignment

Anonymised. Client identity is withheld at their request. The methods, gates, and metrics are real.

Challenge

The team's model was fluent in Tamil but consistently wrong about what users actually wanted — honorifics, formality and cultural nuance. Off-the-shelf preference data encoded Western norms; machine-translated pairs didn't capture native-speaker judgement.

Approach

We built a calibrated preference-data pipeline:

  • Native-speaker raters trained against a documented, piloted preference rubric covering tone, formality and cultural fit.
  • Preference and DPO pairs with gold references and adjudication on contested pairs; inter-rater agreement reported per batch.
  • A culturally-specific eval set to measure lift before and after fine-tuning.

Outcome

A ~2,000-pair native-speaker preference and DPO dataset with gold references and per-batch adjudication, delivered with a culturally-specific eval set so the team can measure cultural-reasoning lift before and after fine-tuning. Inter-rater agreement was held at IAA ≥ 0.85 across batches.

Anonymized engagement. The deliverables, rubric and quality bar above are real and defensible; we don't publish a client's internal eval scores.

How this maps to what we do

The services and data behind this engagement

This outcome was delivered with the same rights-cleared, documented services and datasets you can engage today.

Teach your model the judgement your users expect.

See how we structure engagements and indicative pricing, or tell us your languages, modalities and quality bar for a scoped quote.

Written by

Cognegica Data Operations

Field collection & delivery team

Cognegica Data Operations is the internal team responsible for field-grade data collection, contributor recruitment, consent and delivery across our multilingual programs. This is an editable team identity — a named individual with a public profile can be assigned to it later in the admin.

Run a similar pilot.

Talk to a Language PM