BUYER_DOCS / ANNOTATION_METHODOLOGY
Annotation methodology and quality standard
Goldset — vetted human data for AI Version 1.0 · July 2026 · Review cycle: every 6 months or on material change
A one-page answer to the question a data-ops lead asks first: how do you know the labels are good? Everything below is what we already do. Nothing here is aspirational; where something is not yet in place it is marked Roadmap.
1. The short version
We do not scale with whoever signs up. Every contributor is screened on your rubric before touching production data, gold items are seeded into every batch, and you receive the scorecard with the delivery — so quality is measured against your task rather than asserted against ours.
| Screening pass mark | 80% across classification, preference, reasoning and instruction-following |
| Gold seeding | Every batch, continuously — not a one-off calibration round |
| Review | Multi-pass, with tracked inter-annotator agreement |
| Delivered with every batch | Labelled data · IAA report · QA scorecard |
| If we miss the agreed threshold | We redo it free |
| Subcontracting | None |
2. Contributor screening
Screening is task-specific. A generic platform quiz tells you a person can use a platform; it does not tell you they can do your task.
Stage 1 — Baseline screening. A live, scored seven-task assessment covering classification, RLHF preference judgement, code reasoning, error-spotting and rating, graded against a gold standard. Pass mark 80%. It runs in the browser and you can take it yourself — this is the single artefact we most want buyers to inspect before signing anything.
Stage 2 — Rubric calibration. Before production, contributors are calibrated on your rubric against a gold set we build with you (§3). Contributors who do not reach the agreed threshold on your task do not work on your task, regardless of Stage 1.
Stage 3 — Domain verification. For code and STEM work, contributors are screened engineers and PhDs; domain claims are verified before assignment, not taken on trust.
Standing rule. Screening is a gate, not a certificate. A contributor who drifts below threshold in production (§4) is removed from the queue and their recent work is re-reviewed.
3. Building the gold set
Gold items are the measurement instrument. If they are wrong, every number downstream is decorative.
- You supply a rubric and a small set of adjudicated examples — or we draft the
rubric from your task description and you adjudicate it. Either way, you own the final call on every gold item.
- We stress the rubric before production: we deliberately surface the edge cases
where two reasonable annotators would disagree, and force a decision. Most rubric failure is not annotator error — it is an ambiguity nobody resolved in advance.
- Gold is versioned. When a rubric changes mid-project, the gold set is versioned
with it and the change is recorded in the QA scorecard, so a shift in agreement can be attributed to the rubric rather than to the people.
- Gold is never exhausted. We hold back a reserve so that late-batch items are
scored against items no contributor has seen.
4. Production quality control
Continuous gold seeding. Gold items are distributed through every batch at an agreed rate, indistinguishable from live items. Per-contributor accuracy against gold is tracked over time, so drift is visible while it is still cheap to fix.
Inter-annotator agreement. Multi-pass review on an agreed proportion of items, with agreement computed and reported per batch. We report the metric appropriate to the task — percentage agreement for simple categorical work, and a chance-corrected statistic (Cohen's or Fleiss' kappa, or Krippendorff's alpha for ordinal and preference data) where the label space makes raw agreement misleading. The metric and the target are agreed with you in writing before production starts, and both appear in the scorecard.
Adjudication. Disagreements above the agreed threshold are escalated to a senior reviewer, adjudicated against the gold standard, and — where the disagreement reveals a rubric gap rather than an annotator error — fed back into the rubric with your approval.
Threshold and remedy. The acceptance threshold is agreed in writing before work begins. If a delivered batch misses it, we redo the work at our cost. That is a commercial commitment, not a best-effort statement.
5. What you receive
Every delivery includes three artefacts:
- The labelled data, in the schema agreed at scoping.
- An inter-annotator agreement report — the metric, the value, the sample it was
computed on, and the comparison against the agreed target.
- A QA scorecard — gold accuracy per batch, adjudication volume, rubric versions in
force, and any items flagged as ambiguous for your review.
A worked, anonymised example of all three is available as the Sample Deliverable Pack (see 02-sample-deliverable-pack.md).
6. Pilot terms
| Duration | 5 business days, fixed scope |
| Code / STEM | USD 2,500 – 4,500 |
| Indic / multilingual | USD 1,500 – 2,500 |
| Lock-in | None |
| You send | One task and a rubric |
| You receive | Labelled sample, IAA report, QA scorecard |
The pilot exists to be judged. If the scorecard does not convince you, there is nothing further to discuss and no commitment either way.
7. Scope and honest limits
- Contributors are screened and based in India; client data is hosted in the UAE
(see 03-data-handling-policy.md). There is no subcontracting to third-party vendors or downstream crowd platforms.
- We are founder-led. You work directly with the principal who owns delivery. That
is an advantage on accountability and a constraint on raw throughput — for programmes requiring sustained thousand-annotator scale, say so at scoping and we will tell you honestly whether we are the right supplier.
- SOC 2 and ISO 27001 are on the roadmap, not in place. We will not imply otherwise.
Goldset is the AI-data venture of PMC DXB — the same principal, a separate company and site. Entity: PMCDXB Corporate Services Provider (CSP) L.L.C S.O.C, Dubai commercial licence 1638875, Dubai Department of Economy & Tourism.