Production traces → curated data & environments → evals → post-training → cheaper specialized model -> deployment
Our hypothesis: a post-trained Qwen3.8-27B can handle ~90% of production tasks at frontier-level performance, while cutting inference costs by up to 10×.
We have proved this with one of our pilots where we decreased their costs from $50k/month to $4.6k/month, we believe this is the closest we have gotten to self evolving agents.
We’re looking for teams running agents in production to test this with us.