The coding data layer
for frontier models.
We design, build and verify agentic coding data — executable tasks in reproducible environments, with difficulty measured on our own rollouts before anything ships.
Three principles.
Frontier benchmark alignment
We track frontier benchmarks continuously and evolve our task designs as evaluation paradigms move. When the field shifts from single-file patches to long-horizon repository work — or from static evals to live agentic environments — our production lines move first, so the data trains toward where the frontier is going, not where it was.
Expert in the loop
Automated generation and grading go far — and we use them wherever they're honest. Where they aren't, domain experts take over: authoring research-grade problems, arbitrating ambiguous specs, and reviewing the work no verifier can judge. The boundary between automation and expertise is a design decision we make per task family, not a cost decision.
Scenario diversity
Real software work is long-tail: evolving requirements, messy toolchains, codebases with history. We prioritize diverse, realistic scenarios over benchmark-shaped repetition — because a model trained on a thousand variations of the same task learns the shape of the benchmark, not the shape of the work.
How a task is built.
Environment first
Each task starts as a reproducible container — pinned dependencies, seeded state, and a spec an agent can actually act on. If it doesn't run identically twice, it doesn't ship.
Hidden tests + reference solution
Fail-to-pass hidden suites define success. A reference solution and a programmatic verifier prove the task is solvable and the gate is honest — in both directions.
Rollouts before delivery
We run frontier models against every batch on our own rollout infrastructure and put per-task pass rates on the datasheet. Difficulty is measured, not asserted — and saturated tasks are flagged, not buried.
Capability areas.
Issue-to-patch engineering on real codebases — from focused bug fixes to long-horizon problems that demand exploration, reproduction, multi-file changes and disciplined verification across hundreds of agent steps. This is our deepest line, spanning the full difficulty spectrum up to tasks current frontier agents still fail.
Task designs stay aligned with SWE-bench-class and successor evaluation paradigms, while the underlying repositories are freshly sourced and screened against public benchmarks.
- agent SFT trajectories
- RL environments · fail-to-pass gates
- contamination-controlled evals
Agents earn their keep operating real toolchains: shells, filesystems, APIs and MCP servers. We build shell-native tasks with programmatic checkers, multi-tool workflows captured end-to-end with outcomes, and multi-turn sessions where requirements arrive progressively, evolve, and conflict — the way real users actually drive coding agents.
Long-horizon project work belongs here too: milestone-chained engineering where an agent carries context across stages with regression gates between them.
- terminal-agent RL tasks
- tool-call & MCP trajectories
- multi-turn session data
Frontend and full-stack tasks where the output is judged the way users judge it: does it render, respond, and feel right. Design-to-code implementation, interactive component builds, responsive layouts and accessibility work — with outcomes verified both programmatically and visually.
This site is built by the same team, on the same standards we hold the data to.
- build tasks · visual verification
- UI/UX preference data
- web-arena-style evals
Past the ceiling of crowd annotation: problems where correctness requires knowing the science, not just the syntax. Numerical methods, algorithms, domain modeling and skill-grounded engineering tasks — authored and reviewed by graduate-level domain experts, with difficulty banded and every claim backed by an automated check.
- expert-authored task sets
- skill-taxonomy coverage
- hard capability evals
Where coding meets geometry: programs whose output is a part, a scene, or a simulation. Spec-to-model tasks in code-first CAD — the agent writes the program, the geometry is rendered and checked against dimensions, constraints and topology — extending the same verify-by-execution recipe into Blender scripting, asset pipelines and physics simulation.
A natural bridge to our world-model and robotics programs, which share the same measured-ground-truth discipline.
- spec → code → geometry tasks
- automated geometric checks
- 3D workflow trajectories
The long tail is where scenario diversity is won: niche languages, legacy stacks, infrastructure and build systems, data engineering, embedded targets. We scope task mixes, difficulty bands, domains and formats to your training stage on standing production lines — if the environment can be made reproducible and the outcome verifiable, we can build the line.
- custom production lines
- scoped task mixes & bands
- SFT / RL / eval formats
From overview to catalog.
This page — what we cover, how we build, and why it's differentiated.
Representative task packages and datasheets with measured pass rates, on request.
A scoped batch against your acceptance criteria, evaluated before anything scales.
Complete task inventories, per-family scale and difficulty data — customer access.
The detailed catalog — task inventories, per-family scale, and measured difficulty data — is shared under NDA. Request access and we'll respond within one business day.
See the data behind this page.
We'll share representative task packages, datasheets with measured pass rates, and scoping for custom production lines.