BUYER_DOCS / SAMPLE_DELIVERABLE_PACK
Sample deliverable pack — what a Goldset delivery actually contains
Goldset — vetted human data for AI Version 1.0 · July 2026
This document exists so that an evaluator can see the shape of a delivery before committing to a pilot, and so that whoever they forward it to internally can judge it without us in the room.
Every batch ships three artefacts. Below is what each contains, with a worked anonymised example.
Artefact 1 — The labelled data
Delivered in the schema agreed at scoping (JSONL by default; CSV, Parquet or a platform-native export on request). Every record carries provenance.
{
"item_id": "GS-4417-00212",
"task": "rlhf_preference",
"rubric_version": "v1.2",
"prompt_id": "P-0881",
"label": {"preferred": "B", "strength": 2, "rationale": "B follows the stated constraint on units; A silently converts to metric."},
"annotator_id": "A-113",
"annotator_tier": "expert",
"reviewed_by": "R-004",
"review_outcome": "confirmed",
"gold": false,
"time_on_task_s": 96,
"labelled_at": "2026-07-14T09:22:41Z"
}
Why each field is there. rubric_version lets you attribute a change in agreement to a rubric edit rather than to people. annotator_id is pseudonymous and stable, so you can audit consistency per contributor without receiving personal data. gold marks measurement items. review_outcome shows whether a second pass changed the label. time_on_task_s is included because implausibly fast work is the earliest signal of a quality problem — we would rather you could see it than take our word.
Artefact 2 — Inter-annotator agreement report
GOLDSET · IAA REPORT
Project Preference data, technical domain
Batch 4417 Items 2,000 Multi-pass sample 400 (20%)
Rubric v1.2 (in force from item 00961; v1.1 before)
Metric Krippendorff's alpha (ordinal) Agreed target ≥ 0.75
Result 0.81 PASS
Rubric v1.1 segment (n=192) alpha 0.74
Rubric v1.2 segment (n=208) alpha 0.87
Note: the v1.2 revision resolved the unit-conversion ambiguity flagged in
batch 4416. The lift is attributable to the rubric, not to staffing.
Disagreements escalated 37 (9.3% of multi-pass sample)
Adjudicated by R-004, R-009
Rubric changes proposed 1 (accepted by client 2026-07-13)
Why the segmentation matters. A single headline number hides the most useful information in the report. Splitting by rubric version is how you tell a people problem from a specification problem — and in our experience most agreement failures are the latter.
Artefact 3 — QA scorecard
GOLDSET · QA SCORECARD
Batch 4417 Delivered 2026-07-15
Gold items seeded 200 of 2,000 (10%)
Gold accuracy, batch 94.5% Agreed threshold ≥ 90% PASS
Per contributor (pseudonymous)
A-108 97.0% n=41 A-113 96.1% n=52 A-121 95.2% n=42
A-117 92.8% n=39 A-126 86.4% n=26 ← below threshold
A-126 fell below threshold at item ~00140. Removed from the queue; their
preceding 26 items were re-reviewed in full and 3 labels corrected.
Corrections are included in the delivered data and listed in the appendix.
Items flagged ambiguous for client review 14
Redo triggered under the QA guarantee No
Why we publish the failure. A scorecard with no bad rows is a scorecard nobody is really running. The value of this artefact is that it shows the control working — a contributor drifted, the gold caught it, the work was re-reviewed, and the corrections are traceable. That is the mechanism you are actually buying.
How to request the real pack
The anonymised pack — a genuine labelled batch with its IAA report and scorecard, with all client-identifying content removed — is available on request.
- Under a mutual NDA (
04a-mutual-nda.md) if you would like it before any
commercial discussion.
- No NDA required for this document and the methodology one-pager.
Email [email protected] or start a pilot directly.
Note on scoping
Two things determine whether a pilot is worth running, and both are decided before any data moves:
- The rubric. Send the one you have, however rough. If you do not have one, send
the task description and we will draft it for your adjudication.
- The acceptance threshold and the metric. Agreed in writing first. A delivery that
cannot be judged against a number agreed in advance is not a delivery, it is an opinion.
Goldset is the AI-data venture of PMC DXB. Entity: PMCDXB Corporate Services Provider (CSP) L.L.C S.O.C, Dubai commercial licence 1638875.