Tidel.ai
Capability · Coding & agents

The coding data layer
for frontier models.

We design, build and verify agentic coding data — executable tasks in reproducible environments, with difficulty measured on our own rollouts before anything ships.

SFT trajectories·RL environments·contamination-controlled evals
echo tab ↹ fill · ↵ run
How we approach coding data

Three principles.

philosophy before catalog
01

Frontier benchmark alignment

We track frontier benchmarks continuously and evolve our task designs as evaluation paradigms move. When the field shifts from single-file patches to long-horizon repository work — or from static evals to live agentic environments — our production lines move first, so the data trains toward where the frontier is going, not where it was.

02

Expert in the loop

Automated generation and grading go far — and we use them wherever they're honest. Where they aren't, domain experts take over: authoring research-grade problems, arbitrating ambiguous specs, and reviewing the work no verifier can judge. The boundary between automation and expertise is a design decision we make per task family, not a cost decision.

03

Scenario diversity

Real software work is long-tail: evolving requirements, messy toolchains, codebases with history. We prioritize diverse, realistic scenarios over benchmark-shaped repetition — because a model trained on a thousand variations of the same task learns the shape of the benchmark, not the shape of the work.

Methodology

How a task is built.

the same recipe across every family
Step 1 · Build

Environment first

Each task starts as a reproducible container — pinned dependencies, seeded state, and a spec an agent can actually act on. If it doesn't run identically twice, it doesn't ship.

Step 2 · Gate

Hidden tests + reference solution

Fail-to-pass hidden suites define success. A reference solution and a programmatic verifier prove the task is solvable and the gate is honest — in both directions.

Step 3 · Measure

Rollouts before delivery

We run frontier models against every batch on our own rollout infrastructure and put per-task pass rates on the datasheet. Difficulty is measured, not asserted — and saturated tasks are flagged, not buried.

sample datasheet · pass ratesillustrative
pass@8 = 0%frontier band
0 – 25%hard training band
25 – 75%core training band
> 75%flagged · saturation
every delivered batch carries per-task pass rates from our own frontier rollouts
What we cover

Capability areas.

six areas · scoped to your training stage

Issue-to-patch engineering on real codebases — from focused bug fixes to long-horizon problems that demand exploration, reproduction, multi-file changes and disciplined verification across hundreds of agent steps. This is our deepest line, spanning the full difficulty spectrum up to tasks current frontier agents still fail.

Task designs stay aligned with SWE-bench-class and successor evaluation paradigms, while the underlying repositories are freshly sourced and screened against public benchmarks.

Ships as
  • agent SFT trajectories
  • RL environments · fail-to-pass gates
  • contamination-controlled evals

Agents earn their keep operating real toolchains: shells, filesystems, APIs and MCP servers. We build shell-native tasks with programmatic checkers, multi-tool workflows captured end-to-end with outcomes, and multi-turn sessions where requirements arrive progressively, evolve, and conflict — the way real users actually drive coding agents.

Long-horizon project work belongs here too: milestone-chained engineering where an agent carries context across stages with regression gates between them.

Ships as
  • terminal-agent RL tasks
  • tool-call & MCP trajectories
  • multi-turn session data

Frontend and full-stack tasks where the output is judged the way users judge it: does it render, respond, and feel right. Design-to-code implementation, interactive component builds, responsive layouts and accessibility work — with outcomes verified both programmatically and visually.

This site is built by the same team, on the same standards we hold the data to.

Ships as
  • build tasks · visual verification
  • UI/UX preference data
  • web-arena-style evals

Past the ceiling of crowd annotation: problems where correctness requires knowing the science, not just the syntax. Numerical methods, algorithms, domain modeling and skill-grounded engineering tasks — authored and reviewed by graduate-level domain experts, with difficulty banded and every claim backed by an automated check.

Ships as
  • expert-authored task sets
  • skill-taxonomy coverage
  • hard capability evals

Where coding meets geometry: programs whose output is a part, a scene, or a simulation. Spec-to-model tasks in code-first CAD — the agent writes the program, the geometry is rendered and checked against dimensions, constraints and topology — extending the same verify-by-execution recipe into Blender scripting, asset pipelines and physics simulation.

A natural bridge to our world-model and robotics programs, which share the same measured-ground-truth discipline.

Ships as
  • spec → code → geometry tasks
  • automated geometric checks
  • 3D workflow trajectories

The long tail is where scenario diversity is won: niche languages, legacy stacks, infrastructure and build systems, data engineering, embedded targets. We scope task mixes, difficulty bands, domains and formats to your training stage on standing production lines — if the environment can be made reproducible and the outcome verifiable, we can build the line.

Ships as
  • custom production lines
  • scoped task mixes & bands
  • SFT / RL / eval formats
Going deeper

From overview to catalog.

detail is disclosed progressively — by design
Layer 01 Overview

This page — what we cover, how we build, and why it's differentiated.

Layer 02 Samples

Representative task packages and datasheets with measured pass rates, on request.

Layer 03 Pilot

A scoped batch against your acceptance criteria, evaluated before anything scales.

Layer 04 Full catalog

Complete task inventories, per-family scale and difficulty data — customer access.

The detailed catalog — task inventories, per-family scale, and measured difficulty data — is shared under NDA. Request access and we'll respond within one business day.

See the data behind this page.

We'll share representative task packages, datasheets with measured pass rates, and scoping for custom production lines.