Blog & insights
Notes from
Notes from
the data floor.
How we think about coding agents, benchmarks, data quality and evaluation — written by the team that builds and verifies the tasks.
All posts
research notes · industry observations · company news
Evaluation2026 · 08
Pass rates belong on the datasheet
Why we run frontier-model rollouts on every coding task we sell — and what "hard" should actually mean.
Coding Datacoming soon
What long-horizon really demands of a task
Anatomy of repository-engineering work that still resists frontier agents after hundreds of steps.
Benchmarkscoming soon
Long-horizon is not just more turns
Evolving requirements, milestone chains, and why multi-turn data needs structure — not just length.
Data Qualitycoming soon
Contamination is a supply-chain problem
Screening against public benchmarks is table stakes; the real fix is sourcing material that was never public at all.
Write with us.
We publish guest posts from researchers working on coding agents, evaluation, robotics and world models.