RLHF & preference data
Pairwise comparisons, rankings, and written rationales to train and stress-test reward models.
RLHF preference data, code & STEM annotation, and model evals for AI labs — delivered by vetted Indian domain experts who pass your rubric before they touch production data.
SEND ONE TASK + RUBRIC → LABELLED SAMPLE + QUALITY SCORECARD · NO LOCK-IN
A live, scored 7-task screening — classification, RLHF preference, code reasoning, error-spotting and rating — graded against a gold standard. Run it yourself in the browser.
Four workstreams, one standard of quality — scoped to your rubric, staffed by vetted specialists.
Pairwise comparisons, rankings, and written rationales to train and stress-test reward models.
Judgment from screened engineers and PhDs — not a generic crowd.
We build rubrics, run them at scale, and surface the failure modes benchmarks miss.
Native testers across Hindi, Tamil, Telugu, Bengali, Marathi, Kannada, Malayalam, Gujarati, Punjabi, Odia and more.
Robotics' bottleneck isn't compute — it's data on how humans physically do skilled work. Our vetted domain experts turn professional-work video into training-ready labels for VLA and embodied-AI models: causal event graphs, action-language captions, and 2D contact/keypoint labels — backed by a proprietary archive of annotated multi-profession video.
Most vendors scale with whoever shows up and test them on a generic platform quiz. We screen every contributor on your rubric, seed gold into every batch, and hand you the scorecard — so quality is measured against your task, not assumed.
We recruit for depth in the domains AI teams actually pay for.
IIT/NIT-grade software engineers for code generation, review, and coding-agent evaluations.
Maths, science and domain specialists for reasoning data and hard evals.
Fluency-tested native speakers for multilingual data across India's major languages.
You bring the task, rubric, and volume. We pressure-test the spec and set the gold standard.
We screen and NDA a scored expert pool matched to your domain and languages.
Production runs with seeded gold and multi-pass review — quality is measured, not assumed.
You get clean data plus agreement reports; we tighten the rubric and scale what works.
We start behind an NDA, before any of your data changes hands.
We screen every contributor on your rubric — not a generic platform test — seed gold items into every batch, and you deal directly with the founders who own delivery. Smaller, more specialised, accountable per engagement.
Try the screening our annotators pass, live in your browser, and start with a small paid pilot judged on your own acceptance criteria. You see the scorecard and agreement report, not just labels.
Per-contributor NDAs and IP assignment, a DPA, an access-controlled workspace with no subcontracting, activity logging, and a retention/deletion policy. GDPR- and India-DPDP-aware; SOC 2 / ISO 27001 on the roadmap. NDA before any data exchange.
Software engineering and competitive coding, maths/STEM and PhD-level domains, native Indic languages, and — in pilot — physical-skill video for embodied AI.
Pilots run in about five working days and are paid; pricing is per-task, scoped to your rubric and volume, with no lock-in.
Send us one task and a rubric. We'll vet a pool, label a sample, and return measured quality — data, agreement report and scorecard — within a week.
SMALL · PAID · NO LOCK-IN
Join the vetted pool. Take the 7-task screening and, if you pass, get matched to paid AI-data work.