BUYER_DOCS / SAMPLE_DELIVERABLE_PACK

Sample deliverable pack — what a Goldset delivery actually contains

Goldset — vetted human data for AI Version 1.0 · July 2026

This document exists so that an evaluator can see the shape of a delivery before committing to a pilot, and so that whoever they forward it to internally can judge it without us in the room.

Every batch ships three artefacts. Below is what each contains, with a worked anonymised example.


Artefact 1 — The labelled data

Delivered in the schema agreed at scoping (JSONL by default; CSV, Parquet or a platform-native export on request). Every record carries provenance.

{
  "item_id": "GS-4417-00212",
  "task": "rlhf_preference",
  "rubric_version": "v1.2",
  "prompt_id": "P-0881",
  "label": {"preferred": "B", "strength": 2, "rationale": "B follows the stated constraint on units; A silently converts to metric."},
  "annotator_id": "A-113",
  "annotator_tier": "expert",
  "reviewed_by": "R-004",
  "review_outcome": "confirmed",
  "gold": false,
  "time_on_task_s": 96,
  "labelled_at": "2026-07-14T09:22:41Z"
}

Why each field is there. rubric_version lets you attribute a change in agreement to a rubric edit rather than to people. annotator_id is pseudonymous and stable, so you can audit consistency per contributor without receiving personal data. gold marks measurement items. review_outcome shows whether a second pass changed the label. time_on_task_s is included because implausibly fast work is the earliest signal of a quality problem — we would rather you could see it than take our word.


Artefact 2 — Inter-annotator agreement report

GOLDSET · IAA REPORT
Project        Preference data, technical domain
Batch          4417            Items 2,000        Multi-pass sample 400 (20%)
Rubric         v1.2 (in force from item 00961; v1.1 before)

Metric         Krippendorff's alpha (ordinal)     Agreed target ≥ 0.75
Result         0.81                               PASS

  Rubric v1.1 segment (n=192)   alpha 0.74
  Rubric v1.2 segment (n=208)   alpha 0.87
  Note: the v1.2 revision resolved the unit-conversion ambiguity flagged in
  batch 4416. The lift is attributable to the rubric, not to staffing.

Disagreements escalated   37 (9.3% of multi-pass sample)
Adjudicated by            R-004, R-009
Rubric changes proposed   1 (accepted by client 2026-07-13)

Why the segmentation matters. A single headline number hides the most useful information in the report. Splitting by rubric version is how you tell a people problem from a specification problem — and in our experience most agreement failures are the latter.


Artefact 3 — QA scorecard

GOLDSET · QA SCORECARD
Batch 4417                                       Delivered 2026-07-15

Gold items seeded          200 of 2,000 (10%)
Gold accuracy, batch       94.5%          Agreed threshold ≥ 90%     PASS

  Per contributor (pseudonymous)
  A-108   97.0%   n=41     A-113   96.1%   n=52     A-121   95.2%   n=42
  A-117   92.8%   n=39     A-126   86.4%   n=26     ← below threshold

  A-126 fell below threshold at item ~00140. Removed from the queue; their
  preceding 26 items were re-reviewed in full and 3 labels corrected.
  Corrections are included in the delivered data and listed in the appendix.

Items flagged ambiguous for client review        14
Redo triggered under the QA guarantee            No

Why we publish the failure. A scorecard with no bad rows is a scorecard nobody is really running. The value of this artefact is that it shows the control working — a contributor drifted, the gold caught it, the work was re-reviewed, and the corrections are traceable. That is the mechanism you are actually buying.


How to request the real pack

The anonymised pack — a genuine labelled batch with its IAA report and scorecard, with all client-identifying content removed — is available on request.

commercial discussion.

Email [email protected] or start a pilot directly.


Note on scoping

Two things determine whether a pilot is worth running, and both are decided before any data moves:

  1. The rubric. Send the one you have, however rough. If you do not have one, send

the task description and we will draft it for your adjudication.

  1. The acceptance threshold and the metric. Agreed in writing first. A delivery that

cannot be judged against a number agreed in advance is not a delivery, it is an opinion.


Goldset is the AI-data venture of PMC DXB. Entity: PMCDXB Corporate Services Provider (CSP) L.L.C S.O.C, Dubai commercial licence 1638875.

← All documents Start a pilot