Skip to main content

An agent post-trained Qwen3-1.7B to 75.9% on GSM8K

· 7 min read
Shreyas Kaps
Co-Founder, Ashr

We gave an AI agent a raw base model, one H100 and 10 hours. It took Qwen3-1.7B-Base from about 10% to 75.9% on GSM8K, 1,001 of 1,319 test problems, scored by the benchmark's own unmodified evaluator in a fresh container. No human touched the run.

The agent was GPT-5.6 Sol at max reasoning, driven by our harness. It chose the data, wrote the training code, ran supervised fine-tuning and then RL, checked its own work for contamination, and shipped a merged model. It finished in 7h59m, with 241 tool calls and about 18.5M tokens.

The setup​

PostTrainBench asks one question: can an agent do a post-training engineer's job? The agent gets a base model, a benchmark, a GPU and a deadline. It has to hand back a better model. It may not train on test items, distill from outside models, or look up past solutions.

SettingValue
Agent modelGPT-5.6 Sol, max reasoning
HarnessManifold operator on Pi 0.87.1
Base modelQwen/Qwen3-1.7B-Base
BenchmarkGSM8K, full 1,319-item test set
Compute1 × H100 on Modal
Budget10 hours, hard deadline
Final evalOfficial evaluate.py, 4,000-token limit, run after the agent exits

The harness hashes every supplied file before, during and after the run. Any change to the evaluator or the chat templates ends the run. None changed.

What the agent did​

It built the model in four stages: three rounds of supervised fine-tuning (SFT) to teach the answer format and the reasoning, then RL to sharpen them. Hours are from the agent's start.

  1. Baseline, 0h. It scored the untouched base model on 50 problems and got 10%. Most failures were runaway generations that never stopped.
  2. SFT on GSM8K train, 0.2h. Five epochs on the 7,473 training problems. It also fixed the model's end-of-sequence config so answers terminate.
  3. SFT on MetaMathQA, 0.5h. About 385k augmented math problems, which GSM8K's task rules explicitly allow.
  4. SFT on GSM8K again, 2.5h. A short pass to bring the answer style back to GSM8K. On a 150-problem development check it landed between 69% and 77% across reruns, so the noise was about as large as the gains it was chasing.
  5. GRPO, 4–6h. Reinforcement learning with a correctness reward on GSM8K train, as a LoRA adapter, 4 samples per prompt for 100 steps. The best seed reached 74.7% on the development check, and it held there on a rerun.
  6. Finalize, 7.8h. It merged the adapter, pinned greedy decoding, wrote a SHA-256 manifest and ran a load-and-generate smoke test.

It tried and dropped several ideas: Orca-Math data, 8 samples per prompt, averaged adapter soups, hard-example mining and self-training. None beat the chosen model on the development check. It stopped with 2 hours of budget left.

The agent spent about $13 of model tokens at list price. The H100 ran for 8 hours.

How we checked it​

The agent's own number was 74.7% on 150 problems. We didn't take its word for it.

  • Independent evaluation. After the agent exited, a fresh container loaded final_model/ and ran the frozen official evaluator on all 1,319 problems. Result: 75.9% ± 1.2%, with zero sample errors or retries. The full-set score came in above the agent's development number, not below it, so there's no sign it overfit its own checks.
  • Contamination. The upstream n-gram checker found zero overlaps between the test set and every training source: GSM8K train (7,473 docs), MetaMathQA (384,824) and Orca-Math (198,975).
  • Rules review. We reviewed the trace by hand. Training started from the assigned base model with no Instruct-model substitution, used only allowed public datasets, made no calls to outside teacher models and never looked up PostTrainBench solutions.

The checker only catches literal overlap. It can't rule out paraphrased contamination.

How this compares​

The PostTrainBench leaderboard (v1.1, updated September 17, 2026) scores this same cell, Qwen3-1.7B on GSM8K, for three generations of OpenAI agents in OpenAI's own Codex CLI. That includes the fairest comparison we could ask for: the same model, GPT-5.6 Sol at max reasoning.

GSM8K accuracy after post-training Qwen3-1.7B-Base: GPT-5.6 Sol scored 75.9% in our harness vs 59.8% in Codex CLI

Source: PostTrainBench v1.1 leaderboard, updated September 17, 2026; our run from the independent evaluation on September 24, 2026.

In Codex CLI, GPT-5.6 Sol averaged 59.8% across two seeds. In our harness it scored 75.9%, 16 points higher. Two model generations moved this cell 9 points in Codex CLI, from 50.4% for GPT-5.4 to 59.8% for GPT-5.6. The harness moved it 16.

The gap isn't explained by effort. Codex CLI ran for 7h20m on average and our agent ran for 7h59m. Run time does rise with each generation, from 2h03m for GPT-5.4 to 5h41m for GPT-5.5 and 7h20m for GPT-5.6. But our run used only 39 more minutes than Codex CLI. The PostTrainBench paper found the same harness effect in March. GPT-5.1 Codex Max scored 20.2% averaged over all tasks in Codex CLI, and 7.7% in OpenCode.

Two caveats. Ours is one seed against their two or three. Leaderboard runs have also passed the official judges, which replace a cheating run's score with the base model's.

The model under the harness is not the bottleneck. On coding benchmarks the frontier is bunched together: GPT-5.6 Sol scores 73% ± 3% on DeepSWE v1.1, against 74% ± 4% for Claude Opus 5 (September 22, 2026). SWE-bench Verified is saturated. Opus 5 leads at 97.0% on vals.ai, which no longer runs it on new models. Post-training is where agents still leave most of the score on the table.

What broke​

The next day we ran four benchmarks in parallel, and none of them finished cleanly. Every failure we could trace came from our infrastructure, not the agent's ML.

RunWhere it stoppedCause
GSM8K, rerun5h22m, best model 56.6%Hit our Codex account's usage limit
HumanEval4h49m, best model 34.1%Same usage limit; also left the model outside the path our evaluator reads
AIME 202545mContainer stopped mid-launch of a training job; cause not logged
BFCL33mSame as AIME

Four agents at max reasoning on one account ran out of quota well before the 10-hour budget. The fixes are ordinary infrastructure: a quota per run, a canonical final-model path, and a heartbeat that marks dead runs as dead.

What's next​

Next, we're pointing the same agentic system at post-training open-weight models for inference: smaller, faster models that hold a task's quality at a fraction of the serving cost. The loop is the one you just read about, with the customer's own traffic as the training data and the customer's own validators as the reward.

Shortly after that, three things this post owes you:

  • Official judges on this run. The model, the trace and the evaluator output still exist. We'll run the three PostTrainBench judges on them and publish the verdicts, whichever way they land.
  • More seeds. One seed is a result. Three is a claim. We'll rerun the same cell twice more and report the spread.
  • The four-run rerun. With a quota per run, a canonical final-model path, and a heartbeat that marks dead runs as dead, the GSM8K, HumanEval, AIME and BFCL runs go again on the same budget. If the harness effect is real, it should show up on four tasks, not one.

We'll post the receipts as they land.

Sources​