WHO WE ARE · VETTED HUMAN LAYER · ONLINE

Certified expert judgment for frontier AI.

We build vetted contributor pools in India for AI labs — RLHF, expert annotation, and evals screened on your rubric, measured in every batch, and accountable to the founders. Built for AI teams in the US and Europe.

SCREENED ON YOUR RUBRIC → MEASURED IN EVERY BATCH → FOUNDER-ACCOUNTABLE

RLHF preference Code & STEM annotation Evals & red-team Multilingual · Indic Physical-AI · video
Our operating thesis

Models are only as good as the human judgment they learn from.

Most AI data comes from anonymous crowds graded on generic quizzes — hidden variance, unaccountable labels, acceptance tests that pass while behaviour degrades.

We run a different model.

Every Goldset contributor is screened on your rubric, scored against seeded gold in every batch, and traceable to a vetted expert under NDA.

We give you the scorecard — not just the labels. If a batch misses the agreed threshold, we redo it at our cost.

We're not a labelling vendor — we're the vetted human layer that turns expert judgment into trainable data, for language, code, and now physical & temporal AI.

What we do

Human data for the hard parts of AI.

Four workstreams, one standard of quality — scoped to your rubric, staffed by vetted specialists.

RL

RLHF & preference data

Pairwise comparisons, rankings and rationales to train and stress-test reward models.

EX

Expert annotation

Code, maths and STEM judgment from screened engineers and PhDs — not a generic crowd.

EV

Evals & red-teaming

We build rubrics, run them at scale, and surface the failure modes benchmarks miss.

ML

Multilingual & Indic

Native-tested data across India's major languages, for models going beyond English.

Physical-skill data for embodied AI · Emerging · in pilot

Vetted domain experts annotate professional-work video for embodied-AI and VLA training — causal event graphs, action-language captions, and 2D contact/keypoint labels. Backed by a proprietary archive of annotated multi-profession video in India.

How we work

Four principles we don't bend.

01

Screen on your rubric, not a generic quiz

A task-specific certification, built from your spec, that contributors must pass (80%) before they touch production data.

02

Seed gold into every batch

Continuous QA with tracked inter-annotator agreement, so drift is caught in the batch — not by you, weeks later.

03

Named, traceable accountability

Every label traces to a vetted, NDA-bound contributor and a batch reviewer. No anonymous crowd.

04

Founder-owned delivery

You work directly with the people accountable for the outcome — and we redo off-threshold work at our cost.

Security & contracts

Built for how AI labs buy.

We start behind an NDA, before any of your data changes hands.

Contracts

Per-contributor NDA + IP assignment, a DPA, and non-circumvention on every engagement.

Data handling

Access-controlled workspace, no subcontracting, activity logging, and a retention/deletion policy. GDPR & India DPDP-aware.

On the roadmap

SOC 2 / ISO 27001. We share our current security posture on request.

The team

Founder-led, from the UAE.

Two founders on the ground, an expert pool sourced across India, and a focus on serving AI labs in the US and Europe. You work with the people who own the outcome — not an account queue.

01
Delivery & OperationsFounder · vetting, QA and deliveryUAE-based
02
Partnerships & ComplianceFounder · US & EU go-to-marketUAE-based
Base · United Arab Emirates Talent · sourced across India Market · US & EU labs Model · founder-led, bootstrapped

Build on judgment you can trust.

Send us one hard task and a rubric. We'll vet a pool, label a sample, and show you measured quality — within a week.