Skip to main content

A legal-AI company (anonymized) · Legal · Text

Indic legal annotation for a legal-AI product

Expert clause-typing and entity annotation across multilingual Indian judgments and contracts for a legal-AI product that couldn't hallucinate.

By Cognegica Quality & Standards · QA & annotation-standards team

Clause & entity IAA 0.89

Languages: Hindi, Marathi, Tamil, English

Illustration representing data annotation

Anonymised. Client identity is withheld at their request. The methods, gates, and metrics are real.

Challenge

The product's contract-review and case-research copilots failed on Indian legal text: multilingual judgments, code-mixed affidavits and scanned cause-lists. A wrong citation or misread clause was malpractice risk, so the model needed expert-grade ground truth — not crowd labels.

Approach

We assembled legally-trained annotators and a calibrated rubric:

  • Clause typing, entity and obligation extraction, and structure labeling across contracts, judgments and pleadings.
  • Guidelines co-written with the client's legal SMEs and piloted before scale; two-pass QA with adjudication on contested items.
  • Access-controlled handling, de-identification and documented chain of custody for confidential matters.

Outcome

Annotation reached inter-annotator agreement of 0.89 on clause and entity labels across four languages, with a fully documented, rights-cleared and de-identified delivery the client could defend in a procurement and compliance review.

Representative outcome from an anonymized engagement.

How this maps to what we do

The services and data behind this engagement

This outcome was delivered with the same rights-cleared, documented services and datasets you can engage today.

Build legal AI that holds up.

See how we structure engagements and indicative pricing, or tell us your languages, modalities and quality bar for a scoped quote.

Written by

Cognegica Quality & Standards

QA & annotation-standards team

Cognegica Quality & Standards is the internal team that defines and enforces our annotation guidelines, multi-layer QA, native-linguist review and inter-annotator agreement reporting. This is an editable team identity — a named reviewer with a public profile can be assigned to it later in the admin.

Run a similar pilot.

Talk to a Language PM