5 pointsby thymikee3 hours ago1 comment
  • thymikee3 hours ago
    Hi HN, I’m Michal. Together with Piotr, Artur, Lech, and Mike, for the last 8 months we've been building Apex, a coding model for mobile and web development. A model which we trained on React Native expertise and turned out to be good at React web too.

    We're a small R&D team at Callstack, where we spend a lot of time working on React Native apps and create open-source libraries. We wanted to see how far that experience could take an open-weight model through domain-specific training. Could we achieve frontier-level performance for specific tasks at a much lower cost? After a few base model iterations, we sticked to Qwen as a foundation, trained on curated framework documentation, source code, APIs, and repository-grounded conversations, evaluated by our engineers.

    On our React Native Evals (rn-evals.vercel.app), Apex scores 87.8%, compared with 88.6% for Claude Opus 5.5. The reported evaluation cost is $7.72 versus $28.75. These are our benchmarks, so we’d encourage you to try tasks from your own repository too.

    Over the weekend we released a stealth version of Apex, Pixel Canary on Vercel AI Gateway. In Vercel’s Next.js evaluation, that preview passed 28 of 31 tasks, or 30 with documentation in context, matching GPT-6 Astra. Opening free access to the model on AI Gateway smashed our infrastructure and in these first days Pixel Canary was slow. We're not DevOps by train, but over these 5 days we had a quick crash course on scaling the model performance on B200s with custom prefill and decode setups to optimize load accordingly to the traffic we made, allowing us to achieve ~100 TPS at 4s p50 latency on the last day, with up to 80B tokens / day throughput.

    The API is now available for business users (we need few more weeks to sort out EU consumer rights). It works with OpenAI- and Anthropic-compatible coding tools. We charge $0.50/M input, $0.20/M cache, and $3/M for output.

    We’re interested in where specialization pays off and where it falls short. We're planning more specialized models for native iOS and Android development, and more.

    I'm using this model for other languages and it's just as good as running Sonnet or Sol on medium. It's still Qwen, but tuning seems to make it better than barebones version on coding tasks.

    If you try it, I’d love to hear which tasks it handles well and where you still reach for another model.