Skip to main content

Guide

Transcription Guidelines — the seven core standards

A reference doc covering Cognegica's seven transcription standards: native-linguist transcribers, orthographic accuracy, timestamping, diarization, non-speech events, dialect handling and formatting.

Illustration representing speech transcription

This reference documents the seven core standards behind every Cognegica transcription delivery. They exist so that transcripts are accurate, consistent and machine-readable across languages and scripts.

  1. Native-linguist transcribers for each language and script.
  2. Orthographic accuracy to each script's conventions.
  3. Timestamping and segmentation at a defined granularity.
  4. Speaker identification and diarization across multi-speaker recordings.
  5. Non-speech event annotation: [noise], [laughter], [music], [overlapping speech].
  6. Dialect and accent considerations handled by native linguists.
  7. Data formatting and consistency: UTF-8, standardized punctuation, consistent spacing and project annotation tags.

Together these feed the multi-layer QA framework and downstream training and evaluation.

Using these standards

Questions about the transcription standards

Why native-linguist transcribers?

Orthography, dialect and accent edge cases need native judgement that generic typists can't provide. It's the first of the seven standards for that reason.

How are non-speech events handled?

With a fixed vocabulary — [noise], [laughter], [music], [overlapping speech] — applied consistently so downstream models can account for them.

What formatting do deliveries use?

UTF-8 with standardized punctuation, consistent spacing and project annotation tags, applied uniformly across the delivery.

Need data like this?

License a proprietary dataset, or commission a collection in the languages you need.