Manifold levels up · Research teams

A Specialized Agent Harness for Model Training

A general coding agent can write a training script that looks like it works, until it blows up a large model's training budget on GPUs and pathological runs. The solution to this is a specialized harness with a focus on training AI models, with a repository of individually tested and vetted training recipes for every model and use case, developed by engineers who lived the problem in post-training.

Mistakes are expensive and slow to surface

The reality is that general purpose coding agents work best in environments where errors are quick to surface and easy to check. In a full stack project, agents can quickly work with unit tests, language-specific linters, and even computer-use, to quickly iterate beyond their errors into a functioning, production-ready solution. That is not the case with model training. Errors are subtle: data curation can be slightly incorrect, resulting in an incorrect loss function, loss can be assigned to the wrong tokens, resulting in a model optimized for the wrong things, and the newest models and architectures can be served or trained in more efficient ways than a general-purpose coding agent's knowledge cutoff has access to.

The recipe has to match your stack, and it keeps changing

Post-training methods and the libraries under them change monthly. GRPO variants, data filters, trainer APIs, quantization paths. Any agent can read the latest docs. The question is whether the recipe it writes from them has ever been run against your trainer, your pinned dependencies, and your model architecture.

Our specialized harness keeps a library of recipes for every major operation, each one executed against real stacks, each one versioned, and the version recorded in the receipt of every run that used it. When a method changes, the recipe is re-tested and the library moves. The agent selects from it; it does not reconstruct a trainer from memory and a README.

The small operations are the ones that go wrong

Post-training has a long tail of steps that look minor and are not. Generating synthetic data to fill the gaps in your traces without inventing behavior your agent never showed. Converting traces into each model's chat and tool-call schema, which differ between Qwen, Llama, Gemma and Mistral and change between versions. Knowing what a good training example looks like for a support transcript, a robot telemetry log, or a code review, and what a plausible-looking wrong one looks like.

The harness ships these as built-in functions, written and tested once and used on every run: synthetic generation grounded in the real traces, schema conversion per model family, qualification that knows the shape of the task. The agent calls them. It does not reinvent them.

A coding agent rebuilds each of these from general knowledge, per run. It does not know Qwen3's tool-call format from Llama's, and a tool call serialized in the wrong format trains a model that cannot call tools. The run finishes, the loss looks fine, and the model is broken in a way nothing measured.

Try it on one run

Bring one experiment your team is about to run anyway. We run it in the harness alongside your own attempt and compare the receipts.

Book a demo

Back to the overview