From traces to a model of your own
Your traces in, a model of your own out, and a platform that keeps it improving. Here is what that includes.
Connect your traces
Point us at where they live: a read-only credential to your trace warehouse (ClickHouse, BigQuery, Snowflake, Postgres, S3), or a read-only key to your tracing tool (Arize, Braintrust, LangSmith, Langfuse). The agents write the import for whatever shape they find. Every training example stays traceable to the run it came from.
Traces become environments
The agents turn your traces into environments the model can be run in again and graded. From the traces they derive what a good run looks like, what a failed one looks like, and the exceptions that should stop a run. You review those criteria once. They become the grader for benchmarking, training and production.
Benchmark and choose a base
Candidate open-weight models run through your environments untrained. Each gets a task score, a latency and a cost per run, alongside the frontier model you call today. The agent recommends. You pick.
Fine-tune within a cap
The base is trained on your traces inside a spend cap you set and the agent cannot raise, in your cloud or on ours with zero data retention. A candidate ships only if it beats the base on traces it never saw. Every run leaves a receipt.
A drop-in endpoint
The model is served behind an endpoint that speaks the OpenAI or Anthropic model spec. Switching in is a model name and a base URL in your config; switching back is the previous value. We need write access to that config to promote and roll back. Or take the weights and host them yourself.
Serving watched by agents
A second set of agents, built for serving, takes over once the model is live. They tune batching, caching, quantization and placement against your real traffic so cost per task keeps falling. They catch runs that loop, stall or burn tokens and pull them into the environments as new cases. When the fix is the prompt, the tools or the stop conditions rather than the weights, they propose the harness change with a before and after.
Continual learning
You decide when the loop runs, through hooks you set in the platform:
- Retrain when a number of new traces has accumulated, a score in your environments drops below a line, a schedule fires, or you press the button.
- Start a new model when a new open-weight base is released, a cost-per-task target is missed, or a task changes enough that the environments say so.
Each hook starts the same loop: new cases in, benchmark, train, gate, promote. Every run, receipt and promotion is visible in the platform, and nothing goes to production without clearing the gate you set.
What we need from you
| Access | For |
|---|---|
| Read-only credential or API key to your traces | Import and ongoing monitoring |
| An hour with someone who knows the task | Reviewing the pass and exception criteria |
| Write access to your model config | Promoting and rolling back the endpoint |
| A cloud account for training, or none if you train on ours | The spend cap is yours either way |
Pricing
A flat monthly fee plus usage on the model. A generous usage allowance is included each month; above it you pay per call, at a fraction of the frontier API rate. Training, environments, benchmarks and the serving agents are covered by the fee.
Start
Thirty minutes. Bring a pointer to your traces and the one task you most want off the frontier API.
Book a demo