Manifold levels up · Growth-stage startups

From traces to a model of your own

Your traces in, a model of your own out, and a platform that keeps it improving. Here is what that includes.

Connect your traces

Point us at where they live: a read-only credential to your trace warehouse (ClickHouse, BigQuery, Snowflake, Postgres, S3), or a read-only key to your tracing tool (Arize, Braintrust, LangSmith, Langfuse). The agents write the import for whatever shape they find. Every training example stays traceable to the run it came from.

Traces become environments

The agents turn your traces into environments the model can be run in again and graded. From the traces they derive what a good run looks like, what a failed one looks like, and the exceptions that should stop a run. You review those criteria once. They become the grader for benchmarking, training and production.

Benchmark and choose a base

Candidate open-weight models run through your environments untrained. Each gets a task score, a latency and a cost per run, alongside the frontier model you call today. The agent recommends. You pick.

Fine-tune within a cap

The base is trained on your traces inside a spend cap you set and the agent cannot raise, in your cloud or on ours with zero data retention. A candidate ships only if it beats the base on traces it never saw. Every run leaves a receipt.

A drop-in endpoint

The model is served behind an endpoint that speaks the OpenAI or Anthropic model spec. Switching in is a model name and a base URL in your config; switching back is the previous value. We need write access to that config to promote and roll back. Or take the weights and host them yourself.

Serving watched by agents

A second set of agents, built for serving, takes over once the model is live. They tune batching, caching, quantization and placement against your real traffic so cost per task keeps falling. They catch runs that loop, stall or burn tokens and pull them into the environments as new cases. When the fix is the prompt, the tools or the stop conditions rather than the weights, they propose the harness change with a before and after.

Continual learning

You decide when the loop runs, through hooks you set in the platform:

  • Retrain when a number of new traces has accumulated, a score in your environments drops below a line, a schedule fires, or you press the button.
  • Start a new model when a new open-weight base is released, a cost-per-task target is missed, or a task changes enough that the environments say so.

Each hook starts the same loop: new cases in, benchmark, train, gate, promote. Every run, receipt and promotion is visible in the platform, and nothing goes to production without clearing the gate you set.

What we need from you

AccessFor
Read-only credential or API key to your tracesImport and ongoing monitoring
An hour with someone who knows the taskReviewing the pass and exception criteria
Write access to your model configPromoting and rolling back the endpoint
A cloud account for training, or none if you train on oursThe spend cap is yours either way

Pricing

A flat monthly fee plus usage on the model. A generous usage allowance is included each month; above it you pay per call, at a fraction of the frontier API rate. Training, environments, benchmarks and the serving agents are covered by the fee.

Start

Thirty minutes. Bring a pointer to your traces and the one task you most want off the frontier API.

Book a demo

Back to the overview