01
SWE task authoring
Difficulty-graded engineering tasks with executable test harnesses and reference solutions, written from real repository contexts.
Mumbai Founded 2024 Expert human data for code & agentic AI
We build software-engineering tasks, evaluate agent trajectories, and produce code-preference and RLHF data — reviewed by practising engineers, not annotators.
await on line 9 yields control between the limit check and the increment. Parallel tool calls all read self.used pre-update and every one of them passes. Single-threaded tests never catch it; agent loops overrun the budget silently.
The thesis
The only constraint that matters is whether the person reviewing the output can tell correct from plausible.
Where non-engineers miss it
01
Patches that compile and pass the obvious tests but are still wrong.
02
Agent rollouts where the failure happened twenty steps before the visible error.
03
Preference pairs with no clean winner, where the written rationale is the actual signal.
What we do
01
Difficulty-graded engineering tasks with executable test harnesses and reference solutions, written from real repository contexts.
02
Step-level annotation of multi-turn tool-use rollouts: where the plan broke, whether recovery was correct, whether the end state met the goal.
03
Ranked pairs with written engineering rationale. Disagreements go to adjudication before delivery.
04
Held-out benchmark construction, contamination screening, and adversarial probing of code and agent behaviour.
Everything we do is code or agent work. We don't take general-purpose annotation projects.
The bench
Two thirds of the people reviewing your model's code have shipped production code themselves. Admission is selective and agreement is monitored from the first live task.
0
engineers on bench
0
software engineers
0
scale ceiling
Referral-led from engineering networks at top Indian engineering institutes.
Live coding and reasoning assessment.
Real tasks, blind-scored against gold.
Agreement monitored from day one.
How we run the work
Stage 01
Guidelines written jointly with your team. Every engineer clears a gold set before live work. Weekly recalibration as task types drift.
Stage 02
Inter-annotator agreement per task family. Blind re-annotation on a rolling sample. Per-person quality scores, visible to you.
Stage 03
Disagreements escalate to a senior reviewer. A written resolution is attached to the item. Patterns feed back into the guidelines.
Stage 04
We work inside your environment — Feather, or any internal platform. Nothing leaves your systems, and we onboard to a new tool in two days. Around it we add sandboxed execution, an agreement dashboard, and an adjudication queue.
What an audit looks like
Every disagreement in a preference batch lands in a queue like this before anything ships. The written resolution stays attached to the item, and the pattern goes back into the guidelines.
Task. Rank two candidate patches for a rate limiter under concurrent access. Both compile. Both pass the supplied suite.
Senior resolution
Upheld B > A. R-027 scored A higher on readability and did not model the interleaving — a real gap, not a mistake. Guideline updated: for any concurrency task family, correctness under interleaving is scored before style, and reviewers must state the interleaving they tested.
closed in 3h 12m · rationale attached to item · pattern logged
Preference tasks sit lowest by design — the ones with no clean winner are the ones worth adjudicating. We report the number we'd rather not.
Security & compliance
Two of these are not done. They are listed at the top rather than the bottom.
| Control | Detail | Status |
|---|---|---|
| SSPA enrolment | Ready to initiate with a Microsoft business sponsor. | Gating item |
| ISO 27001 / 27701 | Underway. Target September 2026. | In progress |
| Individual NDAs | Signed by every engineer before bench admission. | In place |
| Access control | Named users, least privilege, full audit logging. | In place |
| No model assistance | Contractual, monitored, and tested with seeded probe tasks. | In place |
Who we are
We have no client logos to show you. What we have is a bench where two thirds of the people reviewing your model's code have shipped production code themselves, and a quality system we will let you audit.
What we'd rather be judged on
Ask us for the gold set, the agreement numbers per task family, and the adjudication log for a live batch. Read the written rationales. If the reasoning doesn't hold up, don't sign anything.
That offer is the whole pitch. Everything above is just the detail.
Contact
Send a single batch. We'll return it calibrated, with agreement numbers and the written rationales attached, and you can judge from there. We reply within one business day.