THE VETTED HUMAN LAYER FOR AI · ONLINE

AI data from a pool you can actually screen.

RLHF preference data, code & STEM annotation, and model evals for AI labs — delivered by vetted Indian domain experts who pass your rubric before they touch production data.

SEND ONE TASK + RUBRIC → LABELLED SAMPLE + QUALITY SCORECARD · NO LOCK-IN

RLHF preference Code & STEM annotation Model evals Multilingual · Indic Physical-AI · video
// PROOF_NOT_PROMISES

See the exact test our annotators pass.

A live, scored 7-task screening — classification, RLHF preference, code reasoning, error-spotting and rating — graded against a gold standard. Run it yourself in the browser.

Open vetting workspace
[ DELIVERABLES ]

The human data behind better models.

Four workstreams, one standard of quality — scoped to your rubric, staffed by vetted specialists.

RLHF

RLHF & preference data

Pairwise comparisons, rankings, and written rationales to train and stress-test reward models.

TIER · EXPERTSTATUS · LIVE
EXPERT

Code & STEM annotation

Judgment from screened engineers and PhDs — not a generic crowd.

TIER · SWE/PHDSTATUS · LIVE
EVALS

Evals & red-teaming

We build rubrics, run them at scale, and surface the failure modes benchmarks miss.

TIER · EXPERTSTATUS · LIVE
INDIC

Multilingual data

Native testers across Hindi, Tamil, Telugu, Bengali, Marathi, Kannada, Malayalam, Gujarati, Punjabi, Odia and more.

LANGS · 10+STATUS · LIVE
// EMERGING · IN_PILOT

Physical-skill data for embodied AI

Robotics' bottleneck isn't compute — it's data on how humans physically do skilled work. Our vetted domain experts turn professional-work video into training-ready labels for VLA and embodied-AI models: causal event graphs, action-language captions, and 2D contact/keypoint labels — backed by a proprietary archive of annotated multi-profession video.

Talk about a video pilot
[ WHY_GOLDSET ]

We build your pool on purpose.

Most vendors scale with whoever shows up and test them on a generic platform quiz. We screen every contributor on your rubric, seed gold into every batch, and hand you the scorecard — so quality is measured against your task, not assumed.

  • Screened on your gold standardTask-specific vetting before anyone touches production data — an 80% pass mark across classification, preference, reasoning and instruction-following.
  • Gold seeded into every batchContinuous QA with multi-pass review and tracked inter-annotator agreement — drift shows up immediately.
  • Founder-led and accountableYou work directly with the founders who own delivery — not an account manager and a ticket queue.
[ THE_POOL ]

Specialists, not a generic crowd.

We recruit for depth in the domains AI teams actually pay for.

01 · ENGINEERS

Engineers & competitive coders

IIT/NIT-grade software engineers for code generation, review, and coding-agent evaluations.

02 · STEM

STEM & PhD experts

Maths, science and domain specialists for reasoning data and hard evals.

03 · LINGUISTS

Native Indic linguists

Fluency-tested native speakers for multilingual data across India's major languages.

[ PROCESS ]

From rubric to reliable data in four steps.

01 · SCOPE

Share the task

You bring the task, rubric, and volume. We pressure-test the spec and set the gold standard.

02 · STAFF

Assemble the pool

We screen and NDA a scored expert pool matched to your domain and languages.

03 · LABEL

Produce with QA

Production runs with seeded gold and multi-pass review — quality is measured, not assumed.

04 · DELIVER

Ship & iterate

You get clean data plus agreement reports; we tighten the rubric and scale what works.

RETURNED_WITH_EVERY_BATCH → labelled_data.json agreement_report.pdf gold_scorecard
[ SECURITY_&_CONTRACTS ]

Built for how AI labs buy.

We start behind an NDA, before any of your data changes hands.

[✓]Per-contributor NDA + IP assignmentstandard
[✓]Data hosted in the UAE · annotators screened & based in Indiain place
[✓]Encrypted, access-controlled workspace · no subcontractingin place
[✓]Your data deleted after delivery · retention on requestpolicy
[✓]DPA available · GDPR & India DPDP-awareavailable
[ ]SOC 2 / ISO 27001 ROADMAPon the roadmap
[ FAQ ]

What buyers ask first.

How is this different from Scale, Surge or Mercor?

We screen every contributor on your rubric — not a generic platform test — seed gold items into every batch, and you deal directly with the founders who own delivery. Smaller, more specialised, accountable per engagement.

You're new — how do I know the quality is real?

Try the screening our annotators pass, live in your browser, and start with a small paid pilot judged on your own acceptance criteria. You see the scorecard and agreement report, not just labels.

How do you handle security, IP and data privacy?

Per-contributor NDAs and IP assignment, a DPA, an access-controlled workspace with no subcontracting, activity logging, and a retention/deletion policy. GDPR- and India-DPDP-aware; SOC 2 / ISO 27001 on the roadmap. NDA before any data exchange.

Which languages and domains do you cover?

Software engineering and competitive coding, maths/STEM and PhD-level domains, native Indic languages, and — in pilot — physical-skill video for embodied AI.

How fast, and how much?

Pilots run in about five working days and are paid; pricing is per-task, scoped to your rubric and volume, with no lock-in.

Run a pilot on your hardest task.

Send us one task and a rubric. We'll vet a pool, label a sample, and return measured quality — data, agreement report and scorecard — within a week.

THE PILOT · WHAT YOU GET
  • 250–500 expert-labelled examples on your exact task + rubric
  • 5–10 business days · fixed scope
  • Deliverables: labelled data + inter-annotator agreement report + QA scorecard
  • Code/STEM $2,500–$4,500 · Indic/multilingual $1,500–$2,500 · no lock-in
  • Miss the agreed QA threshold? We redo it free.
We reply within one business day. NDA before any data changes hands.
✓ FOUNDER GUARANTEE If the first batch misses your agreed threshold, we redo it at our cost — and you deal with the founders, not a sales queue.

SMALL · PAID · NO LOCK-IN

Expert in code, STEM, or an Indian language?

Join the vetted pool. Take the 7-task screening and, if you pass, get matched to paid AI-data work.

Take the screening