DS 4.1 flash is my main powerhouse and Opus/Astra my auditors (when they're not out of tokens) otherwise K3 or DS4 pro
1. who hosts the inference
2. which harness are you using with it, still CC?
1. I go direct to source, i.e. DS platform, I find it cheaper than paying the openrouter tax -- I also switch it up a bit
2. I built a local LLM router, that I update with new profiles that have my preferred provider of the week (lowest token costs/speed) with fallbacks, like mimo --> DS4 etc.. if there is overloading,
3. I use 3 diff harnesses, CC/Codex + Opencode -- they all talk to each other through a custom rig system that routes messages between llms using a Rust backed structured JSON system
Not saying this is the best, it's just what I like and works for me^.
I can flow quite naturally between Opus/Astra/K3/GLM/MiMo/DS/etc.. this way and often do...more so these days with subs no longer great as they used to be.
I’ll be trying these models out and may end up switching my subscriptions if this craziness continues
If outages like this happened on every major deploy happened at any other company; this would be viewed as unacceptable, especially if it was something like Google Search going down on every update.
Proof that vibe coded or not, people pay someone else mostly for liability.
Not sure how can claude even compete at this point unless openAI seriously downgrade the models to save money.