Skip to main content
Trust & compliance

PII Scrubbing & Data Minimisation

Automated detection and redaction of faces, license plates and names, with manual audit and QA sign-off — data minimisation before delivery.

Data minimisation means a dataset carries only what the use case needs. Cognegica strips personal information — faces, license plates, names and other identifiers — through an automated detection-and-redaction pipeline backed by a manual audit and QA sign-off, so personal data is removed before delivery rather than shipped and regretted. This page documents that pipeline stage by stage.

Why minimise

How is personal information removed from a dataset?

The safest personal data is the data you no longer hold. Before a dataset is delivered, identifiers that are not needed for the task are detected and redacted — and the redaction is checked by a human, because automated detection alone misses edge cases.

Automated detection and redaction

Standard pipeline stages detect and redact faces, license plates, names and other PII across the relevant modalities. Detection thresholds and the exact identifier set are configured per engagement.

Manual audit and QA sign-off

A manual audit reviews a sample for missed or over-redacted content, feeding corrections back before QA sign-off. This is the same multi-layer QA discipline used across our annotation quality frameworks. De-identification is one stage of the wider provenance framework, and keeps sovereign datasets responsible to build.

The scrubbing pipeline

From detection to QA sign-off

An automated pass followed by human audit. The identifier set and thresholds are configured per engagement.

  1. 1

    Automated detection

    Standard pipeline stages detect PII candidates — faces, license plates, names and other identifiers — across the relevant modalities.

    Detection pass complete

  2. 2

    Redaction (faces / plates / names / PII)

    Detected identifiers are redacted or masked so the delivered data carries only what the use case needs.

    Identifiers redacted

  3. 3

    Manual audit

    A human audit reviews a sample for missed or over-redacted content and routes corrections back.

    Manual audit sample passed

  4. 4

    QA sign-off

    Multi-layer QA confirms minimisation against the agreed bar before delivery.

    QA sign-off recorded

Data type, method, check

What gets scrubbed, and how

Each personal-data type, the automated method that detects and redacts it, and the manual check that verifies it.

Data typeAutomated methodManual check
FacesAutomated face detection and blurring / masking in images and videoAudit sample for missed or partial faces
License platesAutomated plate detection and redaction in images and videoAudit sample for missed or partial plates
Names & named entitiesAutomated named-entity detection and masking in text and transcriptsNative-linguist review for context-dependent names
Other PII (IDs, contacts, locations)Pattern- and entity-based detection and redactionManual audit against the engagement's identifier set

Standard pipeline stages — no vendor names or accuracy figures are implied. The identifier set, detection thresholds and any reported metric are editable placeholders scoped and confirmed per engagement.

Related work

Minimisation in context

  • Data provenance framework

    Where de-identification sits in the end-to-end provenance lifecycle.

    See the framework
  • Sovereign data

    Residency, consent and provenance for regulated programs.

    Learn more
  • Annotation quality frameworks

    The multi-layer QA discipline behind the manual audit.

    Read more

About PII scrubbing

Questions about PII and minimisation

What personal data do you remove?

Faces, license plates, names and named entities, and other identifiers such as IDs, contacts and locations — scoped to the engagement's identifier set.

Is removal automated or manual?

Both. Automated detection and redaction run first, then a manual audit reviews a sample for missed or over-redacted content before QA sign-off.

Do you publish accuracy figures?

We don't imply fixed accuracy figures here. Detection thresholds and any reported metric are scoped and confirmed per engagement rather than asserted as a blanket number.

Ship datasets that carry only what they need.

Tell us the modalities and identifier set — we'll scope automated scrubbing with a manual audit.