TermGrade: 1k graded terminal environments, 36k trajectories, and the RL run they trained

Graded six times over; all 36,144 attempts released, the 14,234 failures included. Every task carries a measured pass rate per model, not a difficulty label somebody chose, and that measurement picked the RL training data: +3.1 on Terminal-Bench 2.1, and +2.1 across five runs of the same recipe.

ai& Research · 17 min read

ai& Research · 17 min read

October 8th, 2026
October 8th, 2026

Yagiz Calik, Founding Member of Technical Staff (Post-Training)

Jianbo Wu, Member of Technical Staff (Post-Training)

Shimpei Hara, President & Co-Founder

Executable environments
1,004
container + instruction + pytest verifier
Graded trajectories
36k
14k of them failures, kept in full
Models graded
6
across 6,024 task-by-model cells
Compute behind it
20k h
4.5B tokens of agent rollout

Graded by Kimi-K3, DeepSeek-V4-Pro, GLM-5.2, Qwen3.6-27B and gemma-4-31B-it (with and without reasoning), released per trial, not as summary rows.

TL;DR

  • Today we’re releasing TermGrade: 1,004 terminal environments for training and testing AI agents. Each one is a real task on a Linux machine, with automated tests that check the result, and we kept only tasks whose own reference solution passes those tests.

  • We had six models attempt every task, and we’re releasing every attempt, the failures included, each with the exact test it broke.

  • A model learns the most from tasks it solves about half the time. So we graded gemma-4-31B-it separately, in the setup we train it in, and trained it on the tasks it solved half the time. It gained 3.1 points on Terminal-Bench 2.1, and five repeat runs of the training all beat the base model.

  • We also found that difficulty depends on the model: the tasks one model solves half the time are mostly different from the next one’s. So you grade against the model you’re training.

Reinforcement learning rewards the attempts that succeed. But when a model always solves a task, or never solves it, every attempt receives the same reward, leaving no difference for the training update to learn from. Tasks it gets right about half the time provide both successes to reinforce and failures to compare them against. So the question before training is which tasks fall in that range for the model you plan to train.

Selecting tasks by measured difficulty is already an established approach (see the references), yet task selection in agentic RL often relies on intuition. Measuring difficulty takes more work because the results depend on the model and need to be checked again as it changes.

Those measurements are specific to the model. A task can be easy for one model and difficult for another, so another model’s pass rates cannot reliably tell you which tasks to train on. A new model version or tool setup can change the results again. Tasks that were useful last quarter may no longer be the right choice today.

We evaluated six model configurations on the same 1,004 tasks. The tasks each solved about half the time were largely different. For our training run, we therefore tested candidates separately with the model and tool setup we planned to use, then selected tasks with pass rates near 50%.

We built the evaluation and selection pipeline to be inexpensive to rerun, so we can repeat these measurements and update the training selection as models change.

We’re releasing the task set and its measured pass rates, every recorded attempt behind those measurements, the exact training selection, and the trained checkpoint:

1,004 executable Linux tasks: container + instruction + pytest verifier, with a measured pass rate for six models.
36,144 graded attempts: every command, every raw terminal frame, per-test verdicts, reasoning where exposed. 21,910 pass: an SFT set usable today (Apache-2.0; the messages column feeds a chat template unmodified). The 14,234 failures ship too.
The selection we trained on: 151 training and 120 validation environments, exact parquet splits.
The resulting adapter and, for convenience, the merged weights.
Figure 1

The pipeline end to end

Stage 1Task generation
Domain + difficulty speca prompt, with a repair loop
A model writes the taskthe whole thing, not a prompt
Four files per task
environment/tests/instruction.mdsolution/
66,165 candidate tasks
Stage 2Filtering by execution, not by review
Build the imageoffline: no network at grade time
Gold-passthe task’s own solution must pass
Unsatisfiable-test auditno implementation could pass
Discarded
fail
1,004 released corpus
Stage 3Release grading — what the corpus carries
Six modelsTerminus-2 scaffold, via Harbor
36,144
A pass rate per modelplus every trajectory behind it
6,024

This grading is the release artifact. It never selects training data.

36,144 graded trajectories
Stage 4Training-band grading — a separate measurement
The base model alonebash tool-calling scaffold
Pass rate near 0.50maximises p(1−p)
151
GRPO++, 5 updates16 rollouts, LoRA r32
+3.1

Graded in the scaffold we train in, because difficulty shifts with the scaffold. Not interchangeable with the Terminus-2 rates above.

+3.1 Terminal-Bench 2.1
Open releaseenvironments/trajectoriesper-model labelsadapter + merged
Stages 1 and 2 build the corpus by running it. It is then graded twice: stage 3, six models under Terminus-2, produces the released labels; stage 4, the base model in our own scaffold, is what selects training data.
Jump target #corpus · invisible on the published site

01 Building and grading the corpus

A terminal task is worth grading only if its environment runs and its verifier can tell a correct solution from a broken one, so before grading any model we had to build tasks we could trust.

DeepSeek-V4-Pro† generated the candidate tasks, and we ran every one of them to see which held up. A task made it into the corpus only if its own reference solution passed its own verifier, offline, inside the built image. Of 66,165 candidate task directories, 1,004 did (1.5%).

Verifiers average 16.2 named assertions per task, so a failure names the behaviour that broke.

We then graded all of it for the release: six models under the Terminus-2 scaffold through Harbor, k=8 for most and k=2 for the two costliest, DeepSeek-V4-Pro and GLM-5.2, under a 12-hour scoring budget. That gives 6,024 task-by-model cells, each with a measured pass rate. pass@1 is the fraction of tasks solved on a single attempt; pass@8, the fraction solved at least once in eight.

Figure 2

Measured pass rates across six models

pass@1 (all six) and pass@8 (k=8 models only). The top three do not resolve; Qwen3.6-27B separates below them.

pass@1pass@8
gemma-4-31B-it
0.50
0.72
gemma-4-31B-it (thinking)
0.58
0.78
Qwen3.6-27B
0.63
0.83
statistically tied
Kimi-K3
0.68
0.84
DeepSeek-V4-Pro
0.69
pass@8 not run — k=2
GLM-5.2
0.69
pass@8 not run — k=2
View as table
Modelkpass@1pass@8
gemma-4-31B-it80.49860.7211
gemma-4-31B-it (thinking)80.57740.7769
Qwen3.6-27B80.63200.8257
Kimi-K380.67700.8367
DeepSeek-V4-Pro †20.6853—
GLM-5.220.6858—
Shaded rows are statistically tied. Chart values are rounded to two decimals; the table carries full precision.

The grading does not separate the top three models. Paired per task over all 1,004 tasks, the differences between Kimi-K3, DeepSeek-V4-Pro and GLM-5.2 are +0.05pp (p=0.97), −0.87pp (p=0.42) and −0.82pp (p=0.51), none of them significant. Qwen3.6-27B, below them, is separated from all three at p < 0.001.

† On the DeepSeek-V4-Pro build. All DeepSeek-V4-Pro numbers refer to the build we graded, not the 0813 release that shipped inside our grading window. Every trial records started_at, so the window is auditable per row.

A tie on pass rate hides a large gap in cost. Kimi-K3 reaches the same score with a median winning trial of 5.6 minutes against 92.9 and 106.9, and roughly 5,600 output tokens per win against 100,000 and 73,000. That is 11 to 13 times less wall clock per attempt, which is why it carries k=8 here while the other two got k=2. Eight attempts on all 1,004 tasks cost 1,929 trial-hours for Kimi-K3; two attempts each for GLM-5.2 and DeepSeek-V4-Pro cost 11,648.

Figure 3

Time per win among the three tied models

Median wall-clock of a winning trial, for three models that finish in a statistical tie on pass@1.

Kimi-K3
5.6 min
Tied peer
92.9 min
Tied peer
106.9 min
pass@1 0.677, 0.685 and 0.686: a tie with DeepSeek-V4-Pro and GLM-5.2 that the run does not resolve. Kimi-K3 wins in a seventeenth to a nineteenth of the wall clock.
Jump target #selection · invisible on the published site

02 The selection rule

The useful tasks are the ones a model solves about half the time. That falls out of how the training update works.

Figure 4

Graded twice, for two different purposes

The released corpus1,004 tasks
Release grading — what the corpus carries
Six models

Terminus-2 scaffold, via Harbor

36,144
It never selects training data
Training-band grading — a separate measurement
The base model alone

bash tool-calling scaffold

151
selects training data
The two pass rates are not interchangeable.


For GRPO, the gradient contribution of a prompt scales with the variance of its rollout rewards. With binary rewards that variance is p(1−p), which is maximal at p = 0.50, so a task the model always solves and one it never solves both contribute nothing.

Figure 5

Why p = 0.50: gradient weight is p(1−p)

GRPO gradient scales with reward variance: zero at the ends, maximal in the middle.

Prompts at p=0 and p=1 contribute nothing. Selection targets the peak for the model about to be trained.

The rule: select the tasks whose measured base pass rate is closest to 0.50. No hand-curation and no heuristic on top, so anyone with a base model and a grader can reproduce it. The idea is not ours: Kimi k1.5 samples prompts by measured pass rate, and recent work formalises reward variance as what speeds group-relative RL. What we add is the measurement itself: six models, every task, published. The rule yielded 151 tasks (Figure 6): 118 from the corpus above plus 33 from SETA, a synthetic terminal-environment set from CAMEL-AI, graded identically and held to the same 0.50 bar. Of the 151, 137 carried gradient once training began. (All 151 are published, SETA included, in termgrade-bandcentral-gemma4-31b-bash. The 1,004-task corpus flags 113 of its own tasks as train151 for decontamination: the 118 minus five dropped during release QA.)

Figure 6

From 66,165 candidates to 118 band-central tasks

Every stage is passed by running the task. Log-scaled. First percentage is of the stage above, second of everything generated.

Candidate task directories written
66,165
100%
Survived execution filtering: the released corpus
1,004
1.5%1.5% of all
Band-central for the base model
118
11.8%0.18% of all
The training set was 151 tasks, not 118: these 118 plus 33 from SETA, which is not part of the 1,004 and so is not shown above. All 151 ship in termgrade-bandcentral-gemma4-31b-bash; 137 carried gradient once training began.
Execution filtering first; band-central selection on the survivors.
Two separate measurements. The six-model grading runs under Terminus-2 and is what the released corpus carries. The selection we trained on was graded separately, in the bash tool-calling scaffold we train in, because difficulty shifts with the scaffold as it does with the model. A task’s Terminus-2 pass rate is not the number that put it in our training set.

Does volume matter? A 707-task corpus of loosely in-band tasks, same base, validation set and hyperparameters, scores 46.6 against the 151-task selection’s 46.1: not distinguishable from zero (t = 0.80, 95% CI −0.8 to +1.8; means of three and seven evaluations). Nearly five times the training data bought no measurable advantage.

Figure 5 explains why. The curve p(1−p) is flat across the middle: anything from 0.3 to 0.7 delivers 84% or more of the peak gradient weight, and 0.45 delivers 99%. So what the rule really does is exclude the tasks at the ends. It does not need to be precise inside the band. That is also why it cost us so little when our tasks drifted to 0.36 by training time: 0.36 still delivers 92% of the peak.

The saving is in the RL run, not in the grading. You cannot build just 151 environments and skip the rest, because finding the 151 took the rest: several thousand task-gradings against the base model, narrowed to 555 tasks in the 0.25–0.75 range, then to the 151 closest to 0.50.

Jump target #training · invisible on the published site

03 Training

Choosing the data was the novel part. The training itself is a recipe suited to long-horizon, sparse-reward, many-turn terminal tasks. The full configuration follows, and two of its choices exist because of how the data was selected.

The recipe is GRPO with DAPO’s modifications for long-horizon agentic RL plus Dr.GRPO’s advantage handling (GRPO++). The changes from default GRPO:

  • No KL term at all, neither in the reward nor as a loss. Long agentic rollouts drift far from the reference policy by design; penalising that fights the thing you want.

  • Clip-Higher: asymmetric PPO clipping at 0.2 / 0.28. The looser upper bound stops low-probability-but-correct tokens from being clipped away, which is the main entropy-collapse path in long-CoT RL.

  • Token-mean loss aggregation, so a 40-turn trajectory contributes proportionally to its length rather than being averaged down to the weight of a 3-turn one.

  • No standard-deviation normalisation of advantages (Dr.GRPO): mean-centre only. Dividing by group std silently up-weights easy and hopeless prompts, which is precisely the bias the selection rule exists to remove.

  • Dynamic sampling: filter_groups drops any prompt whose 16 rollouts come out all-pass or all-fail, then resamples, up to 10 generation batches per step.

  • Fully on-policy: batch 16 prompts, mini-batch 16, one optimizer update per step, no stale-gradient drift within a step. 16 rollouts per prompt at temperature 0.6.

  • Base: gemma-4-31B-it · LoRA rank 32 · lr 1e-5 · binary reward · 5 updates ≈ one epoch.

  • Trained in a bash tool-calling scaffold, evaluated under Terminus-2. The cross-scaffold transfer is part of the result.

  • Held-out validation: 120 disjoint tasks, deliberately bimodal (60 at p≈0.25, 60 at p≈0.75), so it measures generalisation across a difficulty shift rather than in-band recall.

A run takes about seven hours on one 8×H200 node: four of training and three of validation, since we evaluated the held-out set at every step.

Dynamic sampling and band-central selection are the same idea at two points in the pipeline. filter_groups throws away a group with no spread in reward, but only after paying for its 16 rollouts. Selecting on measured pass rate avoids generating those groups in the first place. Some still got through: 14 selected tasks never produced a passing rollout, and filter_groups caught them, which left 137 that carried gradient.

Even those 137 do not all train in a given run. Each update draws 16 prompts with 16 rollouts each, and a batch fills in one go only if none of its groups is all-pass or all-fail. In practice it always drew twice, 32 prompts per step, and kept 16. About 78 distinct prompts contribute gradient per run, and which 78 depends on how the rollouts landed. That is much of why two runs diverge, and the clearest argument we have that batch size should scale with corpus size.

Figure 7

One update, and why only part of it trains anything

151-task bandall at p near 0.50
Draw 16 promptsa batch
16 rollouts eachtemperature 0.6, in containers
Each group of 16 is then judged only on its spread of binary rewards
every rollout passes

no spread

discarded
every rollout fails

no spread

discarded
some pass, some fail

spread is maximal near 0.50

trains
rollout solved the taskrollout did not
A discard forces a redrawso 32 prompts are generated
The batch fills at 16the surplus is thrown away
One optimizer stepLoRA r32, no KL term

About 78 distinct prompts carry gradient in a whole run; which 78 depends on how the rollouts land.

A group is judged on reward spread, so always-solved and never-solved prompts are discarded alike. Band-central selection applies that test before paying for the rollouts.

The discarded groups were those 14 tasks, not band-central tasks misbehaving. A task genuinely at p = 0.50 produces sixteen identical rollouts only about three times in a hundred thousand. But across the prompts that did carry gradient, the mean pass rate was 0.36, not 0.50. The band had moved by the time we trained on it. A difficulty label is a measurement of one model at one moment, and that includes your own model a few updates ago.

Dynamic sampling also gives a cheap check that the band still holds. num_gen_batches, the number of draws it took to fill a batch, sat at 2 of a possible 10 for all five updates, so the corpus stayed band-central for the updated policy. If it climbs toward the cap, that is the signal to re-band rather than keep training.

Jump target #results · invisible on the published site

04 Results across five runs

The released checkpoint scores 46.1 on Terminal-Bench 2.1, +3.1 over the base model: the mean of seven evaluation runs on the exact checkpoint we ship, rerunnable against the released adapter.

Figure 8

Band-central RL on gemma-4-31B-it

Terminal-Bench 2.1, pass@1.

43.0
46.1
+3.1
gemma-4-31B-itbase model
+ band-central RLreleased checkpoint
Base 43.0 from Artificial Analysis. 46.1 is the mean of seven runs, three attempts per task.

The more useful question is what the method delivers, not the one checkpoint we chose to release, so we trained the identical recipe five times and evaluated every checkpoint. The five average 45.1, or +2.1 over the base model (SE 0.6), and all five land above it: the weakest by 0.4 points, the strongest by 3.3. The gain is distinguishable from zero (t = 3.8, p ≈ 0.025). The released checkpoint scored +3.1; that is one run, and +2.1 is the mean of five drawn from the same recipe.

We follow Artificial Analysis’ Terminal-Bench 2.1 methodology: 89 tasks, Terminus-2, pass@1 over three repeats per task, two-hour per-task timeout; the one difference is the sandbox (Daytona through Harbor, not E2B). They report 43.0 for gemma-4-31B-it, our reproduction came out close, and we quote deltas against their figure.

Every run is the identical recipe with no fixed random seed. Run 1 is the checkpoint we release. All five landed above the base model.
Terminal-Bench 2.1vs 43.0
released checkpoint46.1+3.1
five runs, same recipe45.1+2.1 ± 0.6


We also track a held-out validation set of 120 tasks in a different scaffold and difficulty distribution. It moved with the benchmark in four of five runs and against it in one, so Terminal-Bench is the measurement of record, and we did not select checkpoints or runs on validation.

Averaged over the five runs, validation rises, dips at the fourth update to slightly below its start, then jumps at the fifth, in every run without exception; it is the most reproducible thing we measured. Whether five updates is the optimum or one phase of a longer oscillation is the next experiment.

How the gain happens is what we would most like others to check. Both selections built from measured in-band tasks, the 151 and the 707, raised task completion and answer quality together. Other setups we tried raised completion by other means and paid for it in correctness. The gains concentrate on tasks where the base model ran out of turns, and are flat where it already finished. Like everything else here, this rests on a small number of runs.

Jump target #trajectories · invisible on the published site

05 What the trajectory data contains

Most released trajectory sets contain only the successes. This one contains every attempt, and the failures carry much of the information. Take one of them (Figure 9). Asked to build /app/fixcalc, a fixed-point calculator, gemma-4-31B-it writes it in twelve turns, passes 44 of 45 assertions, and fails test_empty_input. Its reward is 0.0, the same as a model that never started.

Figure 9

44 of the task’s 45 assertions

gemma-4-31B-it writes it in twelve turns.

task fixcalcmodel gemma-4-31B-itturns 12assertions 45
44 passed 1 failed — test_empty_input
reward0.00
Reward 0.0: in a score column, indistinguishable from a model that never started.

This case is common. 6,544 released trials failed exactly one assertion while passing eight or more, 18% of everything we ran. A score records only pass or fail, which is enough to compute a reward but says nothing about how the model got there. The per-test data keeps that: which assertion broke, and what the model typed. We tried using it directly as per-test partial credit, and on this corpus it did worse than a binary reward, so we went back to binary. It remains the most interesting unexplored direction here.

The scale would be hard for anyone else to reproduce, but the schema is what makes it usable. Every trial ships as two views of the same record: a chat-formatted messages column that feeds a template directly, and the complete Terminus-2 episode with everything the projection throws away.

Per-trial fields. Two configs carry the same 36,144 trials with one heavy column each; both join on task + model + trial.
FieldWhat it gives you
messages{role, content, reasoning_content, parsed}, SFT-ready, strict alternation verified on every row
terminus_episodeEvery keystroke sent, every raw terminal frame returned, the scaffold’s injected warnings
steps[].metricsPer-turn prompt_tokens, completion_tokens, cached_tokens
steps[].timestampPer-turn wall clock (over completion tokens: that turn’s generation rate)
agent.extraServing config: parser, temperature, max_tokens
tests36,074 rows of named assertion verdicts
verifier_output72,148 rows: pytest output plus program stdout, two per trial
release_metaScoring record inline: reward, reward_uncapped, cap flags, status


In total: 36,144 counting trials (62,481 including the retries beyond k), 3.68 billion input tokens against 0.80 billion output, and 19,875 hours of wall clock. The input-to-output ratio is 4.6:1, because a turn resends the whole transcript. DeepSeek-V4-Pro and GLM-5.2 account for 58% of that compute and 11% of the trials.

Figure 10

pass@1 and total wall clock across six models

One point per model; wall clock is on a log scale. pass@1 spans 1.4× across six models; wall clock spans 10.6×.

Kimi-K3 ties DeepSeek-V4-Pro and GLM-5.2 on pass@1, in roughly a third of the wall clock.

A few things fall out of the data immediately:

  • gemma-4-31B-it with thinking on ends a third of its trials in two turns or fewer. It writes a script, declares completion, and never runs it.

  • Score and cost move independently. pass@1 spans 1.4× across these models, while the median tokens per win spans 34×.

  • GLM emits turns containing no valid tool call, and those turns are longer than its successful ones: analysis that never resolves into an action.

  • More turns is not progress. Paired within a task, winning trials use fewer turns than failing ones for Qwen and GLM; the other models are flat.

Jump target #labels · invisible on the published site

06 Difficulty labels don’t transfer between models

We measured the half-solved band for all six models, and the bands barely overlap: Jaccard similarity between any two distinct models’ in-band sets runs 0.11 to 0.21, and not one of the 1,004 tasks is band-central for all six.

Figure 11

Pairwise Jaccard overlap of the six in-band task sets

Jaccard overlap between each pair’s “solved about half the time” sets (base pass rate 0.25 to 0.75).

0.11–0.21
1.0
0  disjointidentical  1.0
Every distinct pair falls in the dark band, 0.11 to 0.21; the highest, 0.29, is gemma-4-31B-it against itself with thinking on. Task-level difficulty would put these near 1.0.

Difficulty belongs to the combination of task, model and scaffold. If you select training data by difficulty you cannot inherit anyone else’s labels, ours included; you grade against the model you intend to train. That is why we release every trajectory and not only the pass rates: so the measurement can be redone for a different model.

07 What’s next

A few directions are already clear.

  • Breadth, without giving up measurement. 1,004 tasks is small because each carries a proven reference solution, six difficulty measurements and per-test evidence. Next: more tasks and domains, each still measured rather than assumed.

  • A fast regrader. If difficulty has to be measured per model, the tool that matters is the one that lets anyone re-band the corpus against their own model in an afternoon rather than a week.

  • Less grading overhead. 20 to 30% of this release’s wall clock is client-retry overhead from a per-attempt timeout, not model compute: real headroom in the grading pipeline.

  • Other objectives. The rule rests on one fact: for group-relative RL with binary rewards, a prompt’s gradient weight is its reward variance p(1−p), largest at 0.50. That is not specific to GRPO and should carry to other group-relative objectives; DPO would want the analogue, pairs the model finds genuinely close. We have measured neither, and the released corpus is meant to make both answerable.

We hope to follow this with a more detailed paper.

The released repos carry per-trial metadata rather than summary rows, so every pass rate can be verified. One note if you benchmark on the corpus: 202 of the 1,004 tasks are in our train or validation split, so score on the 802-task unseen split. All 271 train and validation tasks ship in the bandcentral repo.

Everything is on Hugging Face

The corpus, the graded trajectories, the band-central splits, and the checkpoint all ship under ai-and. Regrade against your own model and tell us what you find.

Browse the release Talk to us

Contributions

Y.C. designed and led the work: corpus generation and verification, the grading pipeline, band-central selection, the reinforcement learning runs and their evaluation, and this write-up. J.W. ran a broad set of experiments and evaluations across the data, environments, models and training setup, which helped shape how the work was tested and refined. S.H. provided strategic direction and contributed to this write-up.

Acknowledgements

We would like to thank Adithya Kolavi (Hugging Face) and Maxime Labonne (Liquid AI) for reviewing this work and for their feedback on the write-up.

References

Method.

  • DeepSeekMath (Shao et al., 2024): introduces GRPO, the group-relative objective whose per-prompt gradient weight is the reward variance our selection rule targets.

  • DAPO (Yu et al., 2025): dynamic sampling and Clip-Higher for long-horizon agentic RL.

  • Dr.GRPO (Liu et al., 2025): advantage handling without standard-deviation normalisation.

  • LoRA (Hu et al., 2021): the adapter we train.

  • verl / HybridFlow (Sheng et al., 2024): the RL framework the runs execute in.

Benchmark and harness.

  • Terminal-Bench (Merrill, Shaw et al., 2026): the 89-task benchmark; we score on the 2.1 revision, which fixes 28 of those tasks.

  • Terminus: the agent scaffold. We run Terminus 2, which has no separate writeup of its own.

  • Harbor, the run harness, and Daytona, the container sandbox underneath it.

  • Artificial Analysis: the published methodology we follow, and the source of the 43.0 base figure we quote against.

Task corpora we build on or sit beside.

  • SETA (Shen et al., CAMEL-AI, 2026): 4,567 terminal environments; 33 of our 151 training tasks come from it, graded under the same rule as our own.

  • Tmax (UW / AI2, 2026): an open RL recipe for terminal agents, released with data, models and code.

Prior work on selecting training data by measured difficulty. The rule in section 02 is not new, and the literature around it is active.

  • Kimi k1.5 (2025): computes per-prompt pass rate as a difficulty proxy and samples on it, predating DAPO. It weights toward harder prompts rather than toward peak variance, which is a real divergence from our rule.

  • Accelerating RLHF Training with Reward Variance Increase (Yang et al., 2025): formalises the mechanism: higher initial-policy reward variance provably speeds GRPO-based training.

  • Spend Your Rollouts Where It Counts (Kim et al., 2026): the same p(1−p) argument run online as a rollout scheduler rather than an offline filter.

  • CLPO (Zhang et al., 2025): buckets prompts by on-policy accuracy and trains around the medium band.

  • Mechanistically Interpreting Sample Difficulty in RLVR (Cheng et al., 2026): finds the difficulty–gain relationship non-monotonic, with easy-to-medium prompts giving the most stable gains. That cuts against a hard claim that the optimum sits exactly at 0.50.

Citation

@misc{calik2026termgrade,  title        = {{TermGrade}: 1k Graded Terminal Environments, 36k Trajectories, and the {RL} Run They Trained},  author       = {Calik, Yagiz and Wu, Jianbo and Hara, Shimpei},  year         = {2026},  month        = oct,  howpublished = {\url{https://www.aiand.com/newsroom/termgrade}},  note         = {ai\& Research blog post}}