An agent post-trained Qwen3-1.7B to 75.9% on GSM8K
We gave an AI agent a raw base model, one H100 and 10 hours. It took Qwen3-1.7B-Base from about 10% to 75.9% on GSM8K, 1,001 of 1,319 test problems, scored by the benchmark's own unmodified evaluator in a fresh container. No human touched the run.
The agent was GPT-5.6 Sol at max reasoning, driven by our harness. It chose the data, wrote the training code, ran supervised fine-tuning and then RL, checked its own work for contamination, and shipped a merged model. It finished in 7h59m, with 241 tool calls and about 18.5M tokens.
The setup
PostTrainBench asks one question: can an agent do a post-training engineer's job? The agent gets a base model, a benchmark, a GPU and a deadline. It has to hand back a better model. It may not train on test items, distill from outside models, or look up past solutions.
| Setting | Value |
|---|---|
| Agent model | GPT-5.6 Sol, max reasoning |
| Harness | Manifold operator on Pi 0.87.1 |
| Base model | Qwen/Qwen3-1.7B-Base |
| Benchmark | GSM8K, full 1,319-item test set |
| Compute | 1 × H100 on Modal |
| Budget | 10 hours, hard deadline |
| Final eval | Official evaluate.py, 4,000-token limit, run after the agent exits |
The harness hashes every supplied file before, during and after the run. Any change to the evaluator or the chat templates ends the run. None changed.
What the agent did
It built the model in four stages: three rounds of supervised fine-tuning (SFT) to teach the answer format and the reasoning, then RL to sharpen them. Hours are from the agent's start.
- Baseline, 0h. It scored the untouched base model on 50 problems and got 10%. Most failures were runaway generations that never stopped.
- SFT on GSM8K train, 0.2h. Five epochs on the 7,473 training problems. It also fixed the model's end-of-sequence config so answers terminate.
- SFT on MetaMathQA, 0.5h. About 385k augmented math problems, which GSM8K's task rules explicitly allow.
- SFT on GSM8K again, 2.5h. A short pass to bring the answer style back to GSM8K. On a 150-problem development check it landed between 69% and 77% across reruns, so the noise was about as large as the gains it was chasing.
- GRPO, 4–6h. Reinforcement learning with a correctness reward on GSM8K train, as a LoRA adapter, 4 samples per prompt for 100 steps. The best seed reached 74.7% on the development check, and it held there on a rerun.
- Finalize, 7.8h. It merged the adapter, pinned greedy decoding, wrote a SHA-256 manifest and ran a load-and-generate smoke test.
It tried and dropped several ideas: Orca-Math data, 8 samples per prompt, averaged adapter soups, hard-example mining and self-training. None beat the chosen model on the development check. It stopped with 2 hours of budget left.
The agent spent about $13 of model tokens at list price. The H100 ran for 8 hours.
How we checked it
The agent's own number was 74.7% on 150 problems. We didn't take its word for it.
- Independent evaluation. After the agent exited, a fresh container loaded
final_model/and ran the frozen official evaluator on all 1,319 problems. Result: 75.9% ± 1.2%, with zero sample errors or retries. The full-set score came in above the agent's development number, not below it, so there's no sign it overfit its own checks. - Contamination. The upstream n-gram checker found zero overlaps between the test set and every training source: GSM8K train (7,473 docs), MetaMathQA (384,824) and Orca-Math (198,975).
- Rules review. We reviewed the trace by hand. Training started from the assigned base model with no Instruct-model substitution, used only allowed public datasets, made no calls to outside teacher models and never looked up PostTrainBench solutions.
The checker only catches literal overlap. It can't rule out paraphrased contamination.
How this compares
The PostTrainBench leaderboard (v1.1, updated September 17, 2026) scores this same cell, Qwen3-1.7B on GSM8K, for three generations of OpenAI agents in OpenAI's own Codex CLI. That includes the fairest comparison we could ask for: the same model, GPT-5.6 Sol at max reasoning.

Source: PostTrainBench v1.1 leaderboard, updated September 17, 2026; our run from the independent evaluation on September 24, 2026.
In Codex CLI, GPT-5.6 Sol averaged 59.8% across two seeds. In our harness it scored 75.9%, 16 points higher. Two model generations moved this cell 9 points in Codex CLI, from 50.4% for GPT-5.4 to 59.8% for GPT-5.6. The harness moved it 16.
The gap isn't explained by effort. Codex CLI ran for 7h20m on average and our agent ran for 7h59m. Run time does rise with each generation, from 2h03m for GPT-5.4 to 5h41m for GPT-5.5 and 7h20m for GPT-5.6. But our run used only 39 more minutes than Codex CLI. The PostTrainBench paper found the same harness effect in March. GPT-5.1 Codex Max scored 20.2% averaged over all tasks in Codex CLI, and 7.7% in OpenCode.
Two caveats. Ours is one seed against their two or three. Leaderboard runs have also passed the official judges, which replace a cheating run's score with the base model's.
The model under the harness is not the bottleneck. On coding benchmarks the frontier is bunched together: GPT-5.6 Sol scores 73% ± 3% on DeepSWE v1.1, against 74% ± 4% for Claude Opus 5 (September 22, 2026). SWE-bench Verified is saturated. Opus 5 leads at 97.0% on vals.ai, which no longer runs it on new models. Post-training is where agents still leave most of the score on the table.
What broke
The next day we ran four benchmarks in parallel, and none of them finished cleanly. Every failure we could trace came from our infrastructure, not the agent's ML.
| Run | Where it stopped | Cause |
|---|---|---|
| GSM8K, rerun | 5h22m, best model 56.6% | Hit our Codex account's usage limit |
| HumanEval | 4h49m, best model 34.1% | Same usage limit; also left the model outside the path our evaluator reads |
| AIME 2025 | 45m | Container stopped mid-launch of a training job; cause not logged |
| BFCL | 33m | Same as AIME |
Four agents at max reasoning on one account ran out of quota well before the 10-hour budget. The fixes are ordinary infrastructure: a quota per run, a canonical final-model path, and a heartbeat that marks dead runs as dead.
What's next
Next, we're pointing the same agentic system at post-training open-weight models for inference: smaller, faster models that hold a task's quality at a fraction of the serving cost. The loop is the one you just read about, with the customer's own traffic as the training data and the customer's own validators as the reward.
Shortly after that, three things this post owes you:
- Official judges on this run. The model, the trace and the evaluator output still exist. We'll run the three PostTrainBench judges on them and publish the verdicts, whichever way they land.
- More seeds. One seed is a result. Three is a claim. We'll rerun the same cell twice more and report the spread.
- The four-run rerun. With a quota per run, a canonical final-model path, and a heartbeat that marks dead runs as dead, the GSM8K, HumanEval, AIME and BFCL runs go again on the same budget. If the harness effect is real, it should show up on four tasks, not one.
We'll post the receipts as they land.
Sources
- Rank et al., PostTrainBench: Can LLM Agents Automate LLM Post-Training?, arXiv 2603.08640, March 2026. Table 1, Figure 6 and the scaffold comparison.
- PostTrainBench leaderboard, v1.1, updated September 17, 2026: per-cell scores, seed SDs and run times from its published scores.js
- DeepSWE v1.1 leaderboard, updated September 22, 2026
- vals.ai SWE-bench Verified, updated September 1, 2026