Mumbai Founded 2024 Expert human data for code & agentic AI

The people reviewing your model's code have shipped production code themselves.

We build software-engineering tasks, evaluate agent trajectories, and produce code-preference and RLHF data — reviewed by practising engineers, not annotators.

The thesis

Most annotation scales with headcount. Code and agentic data doesn't.

The only constraint that matters is whether the person reviewing the output can tell correct from plausible.

Where non-engineers miss it

01

Patches that compile and pass the obvious tests but are still wrong.

02

Agent rollouts where the failure happened twenty steps before the visible error.

03

Preference pairs with no clean winner, where the written rationale is the actual signal.

What we do

Four kinds of work. All of it code or agent.

01

SWE task authoring

Difficulty-graded engineering tasks with executable test harnesses and reference solutions, written from real repository contexts.

02

Agentic trajectory evaluation

Step-level annotation of multi-turn tool-use rollouts: where the plan broke, whether recovery was correct, whether the end state met the goal.

03

Code preference & RLHF data

Ranked pairs with written engineering rationale. Disagreements go to adjudication before delivery.

04

Evaluation & red-team sets

Held-out benchmark construction, contamination screening, and adversarial probing of code and agent behaviour.

Everything we do is code or agent work. We don't take general-purpose annotation projects.

The bench

Vetted engineers, not a general annotation pool.

Two thirds of the people reviewing your model's code have shipped production code themselves. Admission is selective and agreement is monitored from the first live task.

0

engineers on bench

0

software engineers

0

scale ceiling

Sourcing funneln = 320
01Applicants sourced320

Referral-led from engineering networks at top Indian engineering institutes.

02Passed technical screen46.87%

Live coding and reasoning assessment.

03Cleared paid trialstage

Real tasks, blind-scored against gold.

04Admitted to benchstage

Agreement monitored from day one.

How we run the work

Quality is a system, not a checkpoint.

Stage 01

Calibration

Guidelines written jointly with your team. Every engineer clears a gold set before live work. Weekly recalibration as task types drift.

Stage 02

Measurement

Inter-annotator agreement per task family. Blind re-annotation on a rolling sample. Per-person quality scores, visible to you.

Stage 03

Adjudication

Disagreements escalate to a senior reviewer. A written resolution is attached to the item. Patterns feed back into the guidelines.

Stage 04

Tooling

We work inside your environment — Feather, or any internal platform. Nothing leaves your systems, and we onboard to a new tool in two days. Around it we add sandboxed execution, an agreement dashboard, and an adjudication queue.

What an audit looks like

Rather than tell you we're rigorous, here is one item mid-flight.

Every disagreement in a preference batch lands in a queue like this before anything ships. The written resolution stays attached to the item, and the pattern goes back into the guidelines.

Adjudication queue item 0447 · py-concurrency

Task. Rank two candidate patches for a rate limiter under concurrent access. Both compile. Both pass the supplied suite.

Rationale · R-014 A holds no lock across the await, so two callers can both pass the budget check. B is slower by one allocation and correct under load. Correctness outranks the allocation.

Senior resolution

Upheld B > A. R-027 scored A higher on readability and did not model the interleaving — a real gap, not a mistake. Guideline updated: for any concurrency task family, correctness under interleaving is scored before style, and reviewers must state the interleaving they tested.

closed in 3h 12m · rationale attached to item · pattern logged

Agreement rolling 30d
SWE task authoring0.91
Trajectory evaluation0.84
Code preference0.72
Preference agreement · after recalibration0.72
day 1 · 0.61day 30 · 0.72

Preference tasks sit lowest by design — the ones with no clean winner are the ones worth adjudicating. We report the number we'd rather not.

Security & compliance

We would rather raise this ourselves than have procurement find it.

Two of these are not done. They are listed at the top rather than the bottom.

ControlDetailStatus
SSPA enrolment Ready to initiate with a Microsoft business sponsor. Gating item
ISO 27001 / 27701 Underway. Target September 2026. In progress
Individual NDAs Signed by every engineer before bench admission. In place
Access control Named users, least privilege, full audit logging. In place
No model assistance Contractual, monitored, and tested with seeded probe tasks. In place

Who we are

We're small and new. We're also specialised on purpose.

We have no client logos to show you. What we have is a bench where two thirds of the people reviewing your model's code have shipped production code themselves, and a quality system we will let you audit.

What we'd rather be judged on

Ask us for the gold set, the agreement numbers per task family, and the adjudication log for a live batch. Read the written rationales. If the reasoning doesn't hold up, don't sign anything.

That offer is the whole pitch. Everything above is just the detail.

Contact

Bring us one task family and your guidelines. We'll show you the rest.

Send a single batch. We'll return it calibrated, with agreement numbers and the written rationales attached, and you can judge from there. We reply within one business day.

Your mail client is opening with this message addressed to ceo@hirepowerai.tech. If nothing opened, copy the address and send it directly — we'll reply within one business day.