TermGrade: 1k graded terminal environments, 36k trajectories, and the RL run they trained
Graded six times over; all 36,144 attempts released, the 14,234 failures included. Every task carries a measured pass rate per model, not a difficulty label somebody chose, and that measurement picked the RL training data: +3.1 on Terminal-Bench 2.1, and +2.1 across five runs of the same recipe.
Yagiz Calik, Founding Member of Technical Staff (Post-Training)
Jianbo Wu, Member of Technical Staff (Post-Training)
Shimpei Hara, President & Co-Founder
Graded by Kimi-K3, DeepSeek-V4-Pro, GLM-5.2, Qwen3.6-27B and gemma-4-31B-it (with and without reasoning), released per trial, not as summary rows.
TL;DR
Today we’re releasing TermGrade: 1,004 terminal environments for training and testing AI agents. Each one is a real task on a Linux machine, with automated tests that check the result, and we kept only tasks whose own reference solution passes those tests.
We had six models attempt every task, and we’re releasing every attempt, the failures included, each with the exact test it broke.
A model learns the most from tasks it solves about half the time. So we graded gemma-4-31B-it separately, in the setup we train it in, and trained it on the tasks it solved half the time. It gained 3.1 points on Terminal-Bench 2.1, and five repeat runs of the training all beat the base model.
We also found that difficulty depends on the model: the tasks one model solves half the time are mostly different from the next one’s. So you grade against the model you’re training.
Reinforcement learning rewards the attempts that succeed. But when a model always solves a task, or never solves it, every attempt receives the same reward, leaving no difference for the training update to learn from. Tasks it gets right about half the time provide both successes to reinforce and failures to compare them against. So the question before training is which tasks fall in that range for the model you plan to train.
Selecting tasks by measured difficulty is already an established approach (see the references), yet task selection in agentic RL often relies on intuition. Measuring difficulty takes more work because the results depend on the model and need to be checked again as it changes.
Those measurements are specific to the model. A task can be easy for one model and difficult for another, so another model’s pass rates cannot reliably tell you which tasks to train on. A new model version or tool setup can change the results again. Tasks that were useful last quarter may no longer be the right choice today.
We evaluated six model configurations on the same 1,004 tasks. The tasks each solved about half the time were largely different. For our training run, we therefore tested candidates separately with the model and tool setup we planned to use, then selected tasks with pass rates near 50%.
We built the evaluation and selection pipeline to be inexpensive to rerun, so we can repeat these measurements and update the training selection as models change.
We’re releasing the task set and its measured pass rates, every recorded attempt behind those measurements, the exact training selection, and the trained checkpoint:
messages column feeds a chat template unmodified). The 14,234 failures ship too.The pipeline end to end
This grading is the release artifact. It never selects training data.
Graded in the scaffold we train in, because difficulty shifts with the scaffold. Not interchangeable with the Terminus-2 rates above.
01 Building and grading the corpus
A terminal task is worth grading only if its environment runs and its verifier can tell a correct solution from a broken one, so before grading any model we had to build tasks we could trust.
DeepSeek-V4-Pro† generated the candidate tasks, and we ran every one of them to see which held up. A task made it into the corpus only if its own reference solution passed its own verifier, offline, inside the built image. Of 66,165 candidate task directories, 1,004 did (1.5%).
Verifiers average 16.2 named assertions per task, so a failure names the behaviour that broke.
We then graded all of it for the release: six models under the Terminus-2 scaffold through Harbor, k=8 for most and k=2 for the two costliest, DeepSeek-V4-Pro and GLM-5.2, under a 12-hour scoring budget. That gives 6,024 task-by-model cells, each with a measured pass rate. pass@1 is the fraction of tasks solved on a single attempt; pass@8, the fraction solved at least once in eight.
Measured pass rates across six models
pass@1 (all six) and pass@8 (k=8 models only). The top three do not resolve; Qwen3.6-27B separates below them.
View as table
| Model | k | pass@1 | pass@8 |
|---|---|---|---|
| gemma-4-31B-it | 8 | 0.4986 | 0.7211 |
| gemma-4-31B-it (thinking) | 8 | 0.5774 | 0.7769 |
| Qwen3.6-27B | 8 | 0.6320 | 0.8257 |
| Kimi-K3 | 8 | 0.6770 | 0.8367 |
| DeepSeek-V4-Pro † | 2 | 0.6853 | — |
| GLM-5.2 | 2 | 0.6858 | — |
The grading does not separate the top three models. Paired per task over all 1,004 tasks, the differences between Kimi-K3, DeepSeek-V4-Pro and GLM-5.2 are +0.05pp (p=0.97), −0.87pp (p=0.42) and −0.82pp (p=0.51), none of them significant. Qwen3.6-27B, below them, is separated from all three at p < 0.001.
started_at, so the window is auditable per row.A tie on pass rate hides a large gap in cost. Kimi-K3 reaches the same score with a median winning trial of 5.6 minutes against 92.9 and 106.9, and roughly 5,600 output tokens per win against 100,000 and 73,000. That is 11 to 13 times less wall clock per attempt, which is why it carries k=8 here while the other two got k=2. Eight attempts on all 1,004 tasks cost 1,929 trial-hours for Kimi-K3; two attempts each for GLM-5.2 and DeepSeek-V4-Pro cost 11,648.
Time per win among the three tied models
Median wall-clock of a winning trial, for three models that finish in a statistical tie on pass@1.
02 The selection rule
The useful tasks are the ones a model solves about half the time. That falls out of how the training update works.
Graded twice, for two different purposes
Six models
Terminus-2 scaffold, via Harbor
The base model alone
bash tool-calling scaffold
For GRPO, the gradient contribution of a prompt scales with the variance of its rollout rewards. With binary rewards that variance is p(1−p), which is maximal at p = 0.50, so a task the model always solves and one it never solves both contribute nothing.
Why p = 0.50: gradient weight is p(1−p)
GRPO gradient scales with reward variance: zero at the ends, maximal in the middle.
The rule: select the tasks whose measured base pass rate is closest to 0.50. No hand-curation and no heuristic on top, so anyone with a base model and a grader can reproduce it. The idea is not ours: Kimi k1.5 samples prompts by measured pass rate, and recent work formalises reward variance as what speeds group-relative RL. What we add is the measurement itself: six models, every task, published. The rule yielded 151 tasks (Figure 6): 118 from the corpus above plus 33 from SETA, a synthetic terminal-environment set from CAMEL-AI, graded identically and held to the same 0.50 bar. Of the 151, 137 carried gradient once training began. (All 151 are published, SETA included, in termgrade-bandcentral-gemma4-31b-bash. The 1,004-task corpus flags 113 of its own tasks as train151 for decontamination: the 118 minus five dropped during release QA.)
From 66,165 candidates to 118 band-central tasks
Every stage is passed by running the task. Log-scaled. First percentage is of the stage above, second of everything generated.
termgrade-bandcentral-gemma4-31b-bash; 137 carried gradient once training began.Does volume matter? A 707-task corpus of loosely in-band tasks, same base, validation set and hyperparameters, scores 46.6 against the 151-task selection’s 46.1: not distinguishable from zero (t = 0.80, 95% CI −0.8 to +1.8; means of three and seven evaluations). Nearly five times the training data bought no measurable advantage.
Figure 5 explains why. The curve p(1−p) is flat across the middle: anything from 0.3 to 0.7 delivers 84% or more of the peak gradient weight, and 0.45 delivers 99%. So what the rule really does is exclude the tasks at the ends. It does not need to be precise inside the band. That is also why it cost us so little when our tasks drifted to 0.36 by training time: 0.36 still delivers 92% of the peak.
The saving is in the RL run, not in the grading. You cannot build just 151 environments and skip the rest, because finding the 151 took the rest: several thousand task-gradings against the base model, narrowed to 555 tasks in the 0.25–0.75 range, then to the 151 closest to 0.50.
03 Training
Choosing the data was the novel part. The training itself is a recipe suited to long-horizon, sparse-reward, many-turn terminal tasks. The full configuration follows, and two of its choices exist because of how the data was selected.
The recipe is GRPO with DAPO’s modifications for long-horizon agentic RL plus Dr.GRPO’s advantage handling (GRPO++). The changes from default GRPO:
No KL term at all, neither in the reward nor as a loss. Long agentic rollouts drift far from the reference policy by design; penalising that fights the thing you want.
Clip-Higher: asymmetric PPO clipping at 0.2 / 0.28. The looser upper bound stops low-probability-but-correct tokens from being clipped away, which is the main entropy-collapse path in long-CoT RL.
Token-mean loss aggregation, so a 40-turn trajectory contributes proportionally to its length rather than being averaged down to the weight of a 3-turn one.
No standard-deviation normalisation of advantages (Dr.GRPO): mean-centre only. Dividing by group std silently up-weights easy and hopeless prompts, which is precisely the bias the selection rule exists to remove.
Dynamic sampling:
filter_groupsdrops any prompt whose 16 rollouts come out all-pass or all-fail, then resamples, up to 10 generation batches per step.Fully on-policy: batch 16 prompts, mini-batch 16, one optimizer update per step, no stale-gradient drift within a step. 16 rollouts per prompt at temperature 0.6.
Base: gemma-4-31B-it · LoRA rank 32 · lr 1e-5 · binary reward · 5 updates ≈ one epoch.
Trained in a bash tool-calling scaffold, evaluated under Terminus-2. The cross-scaffold transfer is part of the result.
Held-out validation: 120 disjoint tasks, deliberately bimodal (60 at p≈0.25, 60 at p≈0.75), so it measures generalisation across a difficulty shift rather than in-band recall.
A run takes about seven hours on one 8×H200 node: four of training and three of validation, since we evaluated the held-out set at every step.
Dynamic sampling and band-central selection are the same idea at two points in the pipeline. filter_groups throws away a group with no spread in reward, but only after paying for its 16 rollouts. Selecting on measured pass rate avoids generating those groups in the first place. Some still got through: 14 selected tasks never produced a passing rollout, and filter_groups caught them, which left 137 that carried gradient.
Even those 137 do not all train in a given run. Each update draws 16 prompts with 16 rollouts each, and a batch fills in one go only if none of its groups is all-pass or all-fail. In practice it always drew twice, 32 prompts per step, and kept 16. About 78 distinct prompts contribute gradient per run, and which 78 depends on how the rollouts landed. That is much of why two runs diverge, and the clearest argument we have that batch size should scale with corpus size.
One update, and why only part of it trains anything
no spread
discardedno spread
discardedspread is maximal near 0.50
trainsAbout 78 distinct prompts carry gradient in a whole run; which 78 depends on how the rollouts land.
The discarded groups were those 14 tasks, not band-central tasks misbehaving. A task genuinely at p = 0.50 produces sixteen identical rollouts only about three times in a hundred thousand. But across the prompts that did carry gradient, the mean pass rate was 0.36, not 0.50. The band had moved by the time we trained on it. A difficulty label is a measurement of one model at one moment, and that includes your own model a few updates ago.
Dynamic sampling also gives a cheap check that the band still holds. num_gen_batches, the number of draws it took to fill a batch, sat at 2 of a possible 10 for all five updates, so the corpus stayed band-central for the updated policy. If it climbs toward the cap, that is the signal to re-band rather than keep training.
04 Results across five runs
The released checkpoint scores 46.1 on Terminal-Bench 2.1, +3.1 over the base model: the mean of seven evaluation runs on the exact checkpoint we ship, rerunnable against the released adapter.
Band-central RL on gemma-4-31B-it
Terminal-Bench 2.1, pass@1.
The more useful question is what the method delivers, not the one checkpoint we chose to release, so we trained the identical recipe five times and evaluated every checkpoint. The five average 45.1, or +2.1 over the base model (SE 0.6), and all five land above it: the weakest by 0.4 points, the strongest by 3.3. The gain is distinguishable from zero (t = 3.8, p ≈ 0.025). The released checkpoint scored +3.1; that is one run, and +2.1 is the mean of five drawn from the same recipe.
We follow Artificial Analysis’ Terminal-Bench 2.1 methodology: 89 tasks, Terminus-2, pass@1 over three repeats per task, two-hour per-task timeout; the one difference is the sandbox (Daytona through Harbor, not E2B). They report 43.0 for gemma-4-31B-it, our reproduction came out close, and we quote deltas against their figure.
| Terminal-Bench 2.1 | vs 43.0 | |
|---|---|---|
| released checkpoint | 46.1 | +3.1 |
| five runs, same recipe | 45.1 | +2.1 ± 0.6 |
We also track a held-out validation set of 120 tasks in a different scaffold and difficulty distribution. It moved with the benchmark in four of five runs and against it in one, so Terminal-Bench is the measurement of record, and we did not select checkpoints or runs on validation.
Averaged over the five runs, validation rises, dips at the fourth update to slightly below its start, then jumps at the fifth, in every run without exception; it is the most reproducible thing we measured. Whether five updates is the optimum or one phase of a longer oscillation is the next experiment.
How the gain happens is what we would most like others to check. Both selections built from measured in-band tasks, the 151 and the 707, raised task completion and answer quality together. Other setups we tried raised completion by other means and paid for it in correctness. The gains concentrate on tasks where the base model ran out of turns, and are flat where it already finished. Like everything else here, this rests on a small number of runs.
05 What the trajectory data contains
Most released trajectory sets contain only the successes. This one contains every attempt, and the failures carry much of the information. Take one of them (Figure 9). Asked to build /app/fixcalc, a fixed-point calculator, gemma-4-31B-it writes it in twelve turns, passes 44 of 45 assertions, and fails test_empty_input. Its reward is 0.0, the same as a model that never started.
44 of the task’s 45 assertions
gemma-4-31B-it writes it in twelve turns.
test_empty_inputThis case is common. 6,544 released trials failed exactly one assertion while passing eight or more, 18% of everything we ran. A score records only pass or fail, which is enough to compute a reward but says nothing about how the model got there. The per-test data keeps that: which assertion broke, and what the model typed. We tried using it directly as per-test partial credit, and on this corpus it did worse than a binary reward, so we went back to binary. It remains the most interesting unexplored direction here.
The scale would be hard for anyone else to reproduce, but the schema is what makes it usable. Every trial ships as two views of the same record: a chat-formatted messages column that feeds a template directly, and the complete Terminus-2 episode with everything the projection throws away.
| Field | What it gives you |
|---|---|
messages | {role, content, reasoning_content, parsed}, SFT-ready, strict alternation verified on every row |
terminus_episode | Every keystroke sent, every raw terminal frame returned, the scaffold’s injected warnings |
steps[].metrics | Per-turn prompt_tokens, completion_tokens, cached_tokens |
steps[].timestamp | Per-turn wall clock (over completion tokens: that turn’s generation rate) |
agent.extra | Serving config: parser, temperature, max_tokens |
tests | 36,074 rows of named assertion verdicts |
verifier_output | 72,148 rows: pytest output plus program stdout, two per trial |
release_meta | Scoring record inline: reward, reward_uncapped, cap flags, status |
In total: 36,144 counting trials (62,481 including the retries beyond k), 3.68 billion input tokens against 0.80 billion output, and 19,875 hours of wall clock. The input-to-output ratio is 4.6:1, because a turn resends the whole transcript. DeepSeek-V4-Pro and GLM-5.2 account for 58% of that compute and 11% of the trials.
pass@1 and total wall clock across six models
One point per model; wall clock is on a log scale. pass@1 spans 1.4× across six models; wall clock spans 10.6×.
A few things fall out of the data immediately:
gemma-4-31B-it with thinking on ends a third of its trials in two turns or fewer. It writes a script, declares completion, and never runs it.
Score and cost move independently. pass@1 spans 1.4× across these models, while the median tokens per win spans 34×.
GLM emits turns containing no valid tool call, and those turns are longer than its successful ones: analysis that never resolves into an action.
More turns is not progress. Paired within a task, winning trials use fewer turns than failing ones for Qwen and GLM; the other models are flat.
06 Difficulty labels don’t transfer between models
We measured the half-solved band for all six models, and the bands barely overlap: Jaccard similarity between any two distinct models’ in-band sets runs 0.11 to 0.21, and not one of the 1,004 tasks is band-central for all six.
Pairwise Jaccard overlap of the six in-band task sets
Jaccard overlap between each pair’s “solved about half the time” sets (base pass rate 0.25 to 0.75).
Difficulty belongs to the combination of task, model and scaffold. If you select training data by difficulty you cannot inherit anyone else’s labels, ours included; you grade against the model you intend to train. That is why we release every trajectory and not only the pass rates: so the measurement can be redone for a different model.
07 What’s next
A few directions are already clear.
Breadth, without giving up measurement. 1,004 tasks is small because each carries a proven reference solution, six difficulty measurements and per-test evidence. Next: more tasks and domains, each still measured rather than assumed.
A fast regrader. If difficulty has to be measured per model, the tool that matters is the one that lets anyone re-band the corpus against their own model in an afternoon rather than a week.
Less grading overhead. 20 to 30% of this release’s wall clock is client-retry overhead from a per-attempt timeout, not model compute: real headroom in the grading pipeline.
Other objectives. The rule rests on one fact: for group-relative RL with binary rewards, a prompt’s gradient weight is its reward variance p(1−p), largest at 0.50. That is not specific to GRPO and should carry to other group-relative objectives; DPO would want the analogue, pairs the model finds genuinely close. We have measured neither, and the released corpus is meant to make both answerable.
We hope to follow this with a more detailed paper.
The released repos carry per-trial metadata rather than summary rows, so every pass rate can be verified. One note if you benchmark on the corpus: 202 of the 1,004 tasks are in our train or validation split, so score on the 802-task unseen split. All 271 train and validation tasks ship in the bandcentral repo.
Everything is on Hugging Face
The corpus, the graded trajectories, the band-central splits, and the checkpoint all ship under ai-and. Regrade against your own model and tell us what you find.
Contributions
Y.C. designed and led the work: corpus generation and verification, the grading pipeline, band-central selection, the reinforcement learning runs and their evaluation, and this write-up. J.W. ran a broad set of experiments and evaluations across the data, environments, models and training setup, which helped shape how the work was tested and refined. S.H. provided strategic direction and contributed to this write-up.
Acknowledgements
We would like to thank Adithya Kolavi (Hugging Face) and Maxime Labonne (Liquid AI) for reviewing this work and for their feedback on the write-up.
References
Method.
DeepSeekMath (Shao et al., 2024): introduces GRPO, the group-relative objective whose per-prompt gradient weight is the reward variance our selection rule targets.
DAPO (Yu et al., 2025): dynamic sampling and Clip-Higher for long-horizon agentic RL.
Dr.GRPO (Liu et al., 2025): advantage handling without standard-deviation normalisation.
LoRA (Hu et al., 2021): the adapter we train.
verl / HybridFlow (Sheng et al., 2024): the RL framework the runs execute in.
Benchmark and harness.
Terminal-Bench (Merrill, Shaw et al., 2026): the 89-task benchmark; we score on the 2.1 revision, which fixes 28 of those tasks.
Terminus: the agent scaffold. We run Terminus 2, which has no separate writeup of its own.
Harbor, the run harness, and Daytona, the container sandbox underneath it.
Artificial Analysis: the published methodology we follow, and the source of the 43.0 base figure we quote against.
Task corpora we build on or sit beside.
SETA (Shen et al., CAMEL-AI, 2026): 4,567 terminal environments; 33 of our 151 training tasks come from it, graded under the same rule as our own.
Tmax (UW / AI2, 2026): an open RL recipe for terminal agents, released with data, models and code.
Prior work on selecting training data by measured difficulty. The rule in section 02 is not new, and the literature around it is active.
Kimi k1.5 (2025): computes per-prompt pass rate as a difficulty proxy and samples on it, predating DAPO. It weights toward harder prompts rather than toward peak variance, which is a real divergence from our rule.
Accelerating RLHF Training with Reward Variance Increase (Yang et al., 2025): formalises the mechanism: higher initial-policy reward variance provably speeds GRPO-based training.
Spend Your Rollouts Where It Counts (Kim et al., 2026): the same p(1−p) argument run online as a rollout scheduler rather than an offline filter.
CLPO (Zhang et al., 2025): buckets prompts by on-policy accuracy and trains around the medium band.
Mechanistically Interpreting Sample Difficulty in RLVR (Cheng et al., 2026): finds the difficulty–gain relationship non-monotonic, with easy-to-medium prompts giving the most stable gains. That cuts against a hard claim that the optimum sits exactly at 0.50.
Citation
@misc{calik2026termgrade, title = {{TermGrade}: 1k Graded Terminal Environments, 36k Trajectories, and the {RL} Run They Trained}, author = {Calik, Yagiz and Wu, Jianbo and Hara, Shimpei}, year = {2026}, month = oct, howpublished = {\url{https://www.aiand.com/newsroom/termgrade}}, note = {ai\& Research blog post}}