BUYER_DOCS / ANNOTATION_METHODOLOGY

Annotation methodology and quality standard

Goldset — vetted human data for AI Version 1.0 · July 2026 · Review cycle: every 6 months or on material change

A one-page answer to the question a data-ops lead asks first: how do you know the labels are good? Everything below is what we already do. Nothing here is aspirational; where something is not yet in place it is marked Roadmap.


1. The short version

We do not scale with whoever signs up. Every contributor is screened on your rubric before touching production data, gold items are seeded into every batch, and you receive the scorecard with the delivery — so quality is measured against your task rather than asserted against ours.

Screening pass mark80% across classification, preference, reasoning and instruction-following
Gold seedingEvery batch, continuously — not a one-off calibration round
ReviewMulti-pass, with tracked inter-annotator agreement
Delivered with every batchLabelled data · IAA report · QA scorecard
If we miss the agreed thresholdWe redo it free
SubcontractingNone

2. Contributor screening

Screening is task-specific. A generic platform quiz tells you a person can use a platform; it does not tell you they can do your task.

Stage 1 — Baseline screening. A live, scored seven-task assessment covering classification, RLHF preference judgement, code reasoning, error-spotting and rating, graded against a gold standard. Pass mark 80%. It runs in the browser and you can take it yourself — this is the single artefact we most want buyers to inspect before signing anything.

Stage 2 — Rubric calibration. Before production, contributors are calibrated on your rubric against a gold set we build with you (§3). Contributors who do not reach the agreed threshold on your task do not work on your task, regardless of Stage 1.

Stage 3 — Domain verification. For code and STEM work, contributors are screened engineers and PhDs; domain claims are verified before assignment, not taken on trust.

Standing rule. Screening is a gate, not a certificate. A contributor who drifts below threshold in production (§4) is removed from the queue and their recent work is re-reviewed.


3. Building the gold set

Gold items are the measurement instrument. If they are wrong, every number downstream is decorative.

  1. You supply a rubric and a small set of adjudicated examples — or we draft the

rubric from your task description and you adjudicate it. Either way, you own the final call on every gold item.

  1. We stress the rubric before production: we deliberately surface the edge cases

where two reasonable annotators would disagree, and force a decision. Most rubric failure is not annotator error — it is an ambiguity nobody resolved in advance.

  1. Gold is versioned. When a rubric changes mid-project, the gold set is versioned

with it and the change is recorded in the QA scorecard, so a shift in agreement can be attributed to the rubric rather than to the people.

  1. Gold is never exhausted. We hold back a reserve so that late-batch items are

scored against items no contributor has seen.


4. Production quality control

Continuous gold seeding. Gold items are distributed through every batch at an agreed rate, indistinguishable from live items. Per-contributor accuracy against gold is tracked over time, so drift is visible while it is still cheap to fix.

Inter-annotator agreement. Multi-pass review on an agreed proportion of items, with agreement computed and reported per batch. We report the metric appropriate to the task — percentage agreement for simple categorical work, and a chance-corrected statistic (Cohen's or Fleiss' kappa, or Krippendorff's alpha for ordinal and preference data) where the label space makes raw agreement misleading. The metric and the target are agreed with you in writing before production starts, and both appear in the scorecard.

Adjudication. Disagreements above the agreed threshold are escalated to a senior reviewer, adjudicated against the gold standard, and — where the disagreement reveals a rubric gap rather than an annotator error — fed back into the rubric with your approval.

Threshold and remedy. The acceptance threshold is agreed in writing before work begins. If a delivered batch misses it, we redo the work at our cost. That is a commercial commitment, not a best-effort statement.


5. What you receive

Every delivery includes three artefacts:

  1. The labelled data, in the schema agreed at scoping.
  2. An inter-annotator agreement report — the metric, the value, the sample it was

computed on, and the comparison against the agreed target.

  1. A QA scorecard — gold accuracy per batch, adjudication volume, rubric versions in

force, and any items flagged as ambiguous for your review.

A worked, anonymised example of all three is available as the Sample Deliverable Pack (see 02-sample-deliverable-pack.md).


6. Pilot terms

Duration5 business days, fixed scope
Code / STEMUSD 2,500 – 4,500
Indic / multilingualUSD 1,500 – 2,500
Lock-inNone
You sendOne task and a rubric
You receiveLabelled sample, IAA report, QA scorecard

The pilot exists to be judged. If the scorecard does not convince you, there is nothing further to discuss and no commitment either way.


7. Scope and honest limits

(see 03-data-handling-policy.md). There is no subcontracting to third-party vendors or downstream crowd platforms.

is an advantage on accountability and a constraint on raw throughput — for programmes requiring sustained thousand-annotator scale, say so at scoping and we will tell you honestly whether we are the right supplier.


Goldset is the AI-data venture of PMC DXB — the same principal, a separate company and site. Entity: PMCDXB Corporate Services Provider (CSP) L.L.C S.O.C, Dubai commercial licence 1638875, Dubai Department of Economy & Tourism.

← All documents Start a pilot