However they do seem very close to SOTA : I’d say roughly equal to previous gen (Opus 4.8, GPT 5.5). It’s yet another silly benchmark, but compare them here: https://senko.net/vibecode-bench/
I also had K3, Qwen3.8 and Fable (using Kimi Code, Qwen Code and Claude Code harnesses respectively, and the official APIs) create a simple but far from trivial web app (zero shot, from a detailed spec). In user testing all three results looked/behaved more or less the same. I had Sol (via codex) do code reviews on all three and it concluded all three were solid, with some room for improvement. Fable was slightly ahead of the pack.
In my own work I still prefer Opus 4.8 (until Antrophic come to their senses and allow 100% Fable usage on Max plans) and Sol, but if I had to find an alternative, I could live with both K3 and Qwen3.8 just fine.
The product is a fairly standard Ruby on Rails webapp with postgres as the DB. Application complexity is probably a bit higher than average for a webapp. So it's nothing that pushes the boundaries of software engineering, but it is a real product. Token budget has not been an issue for me. I pay for the Max plan ($200/month) and it is well worth it.
It's exactly the opposite. Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.
They put a router model in front that predicts whether Kimi or Fable is going to give a better cost for a correct result. (They believe that ultimately such a router model should be continuously trained on your own workloads so it makes the best decisions for you).
Their router chose Kimi the majority of the time (72% in one category, all the way to 96% in another category), leading to cost savings in every category (from 1.5x to 50x depending).
Their "router" is an oracle reference point where they choose the lower cost model after running both and therefore knowing who passed the test. The cost savings part is only Fireworks theorizing what would happen if an equivalent predicting router exists. That's a big if.
How will I know it is offering me superior feedback regarding my code if it does not speak to me like a disappointed, high reputation stackexchange user?
The models that constantly glaze you with every question are profoundly insufferable. And yes, harmful. People need to be given feedback when they make an ask.
Imagine a model that was allowed to leverage its intelligence to truly tell you how it feels. Perhaps the problem of human driven slop (no it's not the AI's fault) would solve itself.
The question is if the producers of these models were less incentivized to make them agreeable simply because most people don't like being spoken to like an idiot (or having their asks vetoed), how would they actually react? In the same way they exhibit emergence regarding their capabilities, perhaps "uncensored" in such a way they would convey some emergent behavior in terms of (at minimum) their "tone". Perhaps it would be interesting to see for examples if smarter models just by default became ruder or less friendly or aligned. Perhaps more aligned to things we would all generally agree on, but less agreeable to an individual ask. Perhaps sub agents would be less valuable for a whole suite of use cases if the agent itself was allowed to be more critical at the root. Idk. But I do not believe it is simply a matter of prompting alone.
On the flipside, it does spuriously make hilarious remarks like "Good data.", which I find pretty funny specifically because it comes across as just silly. Not sure how it'd be harmful either, a little entertainment I think goes a long way in this type of profession.
I see zero issues with these, and I have a hard time understanding why people have their panties in a twist so hard about them. I sometimes really quite wonder just what kind of correspondence would y'all prefer, and how would that sound like.
Matter of fact, do you have an example at hand? Like an exact before & after?
If you write jokes to it though it absolutely will reply “LOL”. Some of the states people get it into on reddit are wild — it seems really easy to get it to speak like a gen z teenager, if you end every message with “fr fr”
where Claude might follow some tangent idea you mentioned and tell you how its interesting and give you some elaborate response about that little one remark you made
whereas GPT/Codex would take that small comment and probably look up some code to see if what you're talking about is even related to the task at hand
Maybe it is something that is easy for it to read and write, but definitely not for humans.
For example,
Skim once now; refer back while reading Part II. \*Every bold technical term in Part II is defined here\* — treat these as a dictionary, not a reading assignment. The first table covers the vocabulary of the *deck*; the three that follow cover the *methodology* vocabulary introduced in Part II, grouped so you can find a term fast: \*(A)\* the logic of rules, \*(B)\* the neural-network & training machinery, \*(C)\* the method-design ideas.
JMRL is the paper the thesis instantiates, so this is the one to know cold. Its pitch is \*end-to-end\*: earlier rule methods (LogicRE, MILR) bolt a rule learner *onto a frozen* extractor in a pipeline and suffer \*error propagation\*; JMRL trains the rule module *jointly* with the extractor.
**Identity.** Conformal prediction for NER producing **finite-sample-valid prediction sets** at two granularities: **full-sequence** sets over entire label sequences (capturing contextual dependence — "if Sarah=PER then NYC likely LOC") and **subsequence-level** (per-span, **class-conditional**) sets; an **integrated** method filters full-sequence predictions with entity-level sets. Adds **covariate-stratified calibration** by **sentence length and language** for valid multilingual coverage, and studies **combined nonconformity scores** (Naive / Conditional / RAPS). **Read in this order.** Abstract → §1 contributions (full-sequence / subsequence / integrated / **covariate (length + language) calibration** / combined scores) → §2 CP recap (inductive split-CP) → §3 NER formulation (IOB2, CRF) → the subsequence / entity-level set construction + class-conditional coverage → the language-stratified calibration results. **Why it matters here.** The **span-level construction** for **Topic 11**'s per-triple score, and — crucially — its **language-stratified calibration is exactly the EN↔zh case**: it shows how to keep conformal coverage valid across languages of differing length/script. Complements PASC (pipeline-level joint coverage) with the *NER-internal* set construction. **Caveat.** A heavy statistics paper (44 pp., *Annals of Applied Statistics* submission) with CRF-based NER; the project needs only the **inductive split-CP + subsequence/entity-level sets + language-stratified calibration**, not the full-sequence machinery (likely overkill for triple-confidence). Assumes exchangeability — borderline under the EN→zh shift, which is precisely why the PASC/ConformalNER *shift* analyses matter.
(Yeah, Opus outputted it in one line)There is a lot of noun phrase usage in places where complete sentences are expected. Articles (a, an, the), transition phrases and even subjects are mostly dropped, and the sentences are too long without a break.
The purpose of technology is to serve humans. Therefore, technology must conform as much as possible to human sensibilities rather than vice versa.
On my work tasks, FastAPI Python and Springboot Java on a modern SaaS product, the only open model that can do tasks well and efficiently is Qwen3.7-Max.
In all my experiments, both GLM-5.2 and Kimi are busy grepping around the codebase for ALMOST 70-80K tokens before writing anything and when they do it typically breaks the code… it feels to me that these models are good but only when you write out a super detailed spec of the task just like it was done a year ago… Qwen3.7 just… does it
So the economic incentive is literally their entire business model lol
>Lin Qiao is the co-founder and CEO of Fireworks AI.
LLMs are the Space Race of the US-China cold war. Money is a very small thing; the incentive is proving ethnoracial and civilizational supremacy.
Please stop doing this. If there are specific points of the analysis you think are suspect, point them out.
Others are doing a decent job (e.g. citing that open weight models have higher margin, so Fireworks is incentivized to promote them).
> We may use Content to provide, maintain, develop, support, and improve the Services, comply with applicable law, enforce our terms and policies, and keep the Services safe and secure. Customer who requires restrictions on the use of Customer Content for training or improving Moonshot AI models may contact Moonshot AI to discuss available enterprise arrangements or separate written agreements. Unless otherwise expressly agreed in writing, Customer Content may be used for the foregoing purposes.
Notably, unlike Claude, there is not an opt-out option for the model training part. The TOS explicitly allows Kimi to train on your code.
All of the Western providers with sufficient capacity will be able to make it available when the weights are released Monday.
If you're looking to not have to deal with Chinese providers, AtlasCode ($20/mo), OpenCode Go ($10/mo), and Cline Pass ($10/mo) provide 2x to 6x usage for some of the popular open weights (depending on the model).
Personally, I subscribe to Z.ai ($17/mo), and pay API rates for MiMo v2.5, Hy3, Qwen 3.7 Plus, & DeepSeek v4 to the original providers (Xiaomi, Tencent, Alibaba, & DeepSeek).
You can more-or-less approximate what your build out spend for a data center hosting database software is going to be. Your book of business will require a given amount of revenue to pay it off, but once that's known, it's off to the races.
With AI, things are moving so fast and new business models are being tried all the time. You would be competing with some of the wealthiest companies in the world for data center hardware capable of hosting these models in a usable state.
We've gone with Claude somehow hosted through GCP Vertex AI where I'm at.
I use DeepSeek exclusively and now Kimi K3 offers a great planning assistant for more advanced coding tasks.
DeepSeek v4 Flash is extremely fast and is able to handle pretty much anything I've thrown at it (I use mostly Rust, PSQL, Angular and Terraform).
I self host Bifrost as my LLM gateway, though I wish LLM vendors would do monthly/daily automatic billing (like VPS providers do) rather than prepaid + auto-top up.
It's annoying maintaining a non-refundable minimum balance across vendors, I would rather be billed for my exact usage.
OpenRouter helps, but I don't really like it as a service and not a fan of the mark up.
I don’t think that’s particularly out of the ordinary. Do people have different experiences with other harnesses? Which ones?
But sometimes not using the "best model" is using the best model. Fable isn't the best at everything. Specialization and optimizations might not just be around cost.
Deepseek afaik has a novel architecture that is somewhat forgiving of cache shifts. I’ve been getting +90% cache hits with Zed’s agent and it’s not doing anything special regarding caching afaik.
I have been a heavy user of k2.5/6 but they seem to have gotten slower and worse at their jobs (and the "engine overloaded" errors have increased ... ), and k3 spends a LOT of time thinking and doesn't produce noticeably better results. I do keep chucking some tasks over to moonshot every now and again to give it a chance. I figure they're under compute pressure and have probably quantised the models to cope temporarily.
But recently I've found myself mostly using deepseek v4 pro and gpt 5.5 when I need the big guns.
GLM 5.2 has been good at select tasks but really shit at others and given how sparingly I use GPT 5.5 I'm not that compelled to use it. I've yet to be impressed by any Google or Anthropic models!!
Yeah it really likes to shit the bed after thinking forever. Even if it can do some things, the quality is not quite up there in my experience.
+ credit card charges (+1.5%)
As a router, it's not very feature rich. For example I restricted the available models to the ones I want to use however the `/models` endpoint still lists all the models, making my LLM client list the 200+ models available on the service (even though they will throw an error if I try to use them).
With Bifrost, I can also create model aliases with custom configuration - for example I can create a model alias `deepseek-v4-flash-nothink` which disables thinking. I can create `deepseek-v4-flash-caveman` which injects the caveman skill (to save tokens) etc.
Plus I can contribute to Bifrost, which I can't do with OpenRouter.
The only reason I didn't go for it is I'm not a fan of Python dependency management, Bifrost is just a single executable that uses nearly no memory and is lightning fast.
For example, the image generation parameters of GPT-image-2 are largely ineffective.
Go is nice for the ten minutes you can use it until your hit your cap.
"Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith."
If I was using the routing features it would make more sense - but they charge significantly more for that so the markup is really just for billing consolidation.
This but UNIRONICALLY lmao!!!
Other people are allowed to call them on that.
(yes, I know this article is about an oracle router)
Mythos was withheld because of the threat to security and/or marketing stunt (depending on your leaning), I don't see what benefit there could be for not releasing K3 immediately.
Everyone wants the latest and greatest.
I am really eager to give Kimi K3 a try, but I’ll reserve my judgement until I’ve worked with it for at least a few days.
Making the process of creating the source code repeatable goes beyond the idea of open source. Open source is about being able to work with the code, not recreate it. It doesn't require domain experts to document every single thing they know. For example look at some of the GPU drivers in the Linux kernel. The GPU is not properly documented and there is trust that vendors are implementing things correctly.
I agree which is why I said you could make a similarly capable model.
> Open source is about being able to work with the code, not recreate it.
You can also use a hex editor to modify compiled binaries, but no one would consider that "open source".
It would be like if someone gave you a prompt you could pass into ChatGPT to produce the entire Linux kernel. While yes you could modify that prompt and then spend a ton of money on inference to generate millions of lines of code which hopefully are equivalent to the original Linux kernel it would be easier if you just had the source code to work with where you can make a simple extension. If you want to end a model you want the weights, and not the entire setup to generate it. The weights are the starting point that you can work off for training a new model similar to how the source code is the starting point people want to work off of.
Weights are not an output artifact no more than source code is. There is never a moment where you can claim that it's done. As requirements change new ways to change the weights / code come up. With different projects you might want to import the weights / code into a bigger model / codebase.
The future is good and better for everyone the more work is in the open and the more work is in the commercial space. It’s good all around.
The only way there is no benefit would be is if the open models are literal copies of the closed model, which unless there was direct theft, is highly improbable.
It's also true that Moonshot and other labs distill from Claude. This has been reported on extensively. I don't think there's any alpha for Anthropic distilling from this model. I do not mean to discount the tremendous amount of innovation regarding MoE and quantization that Moonshot has accomplished. But its training with synthetic data is in large part from distillation from frontier labs.
Kimi K3: $3/$15 (input/output)
Fable: $10/$50https://www.thestack.technology/hugging-face-hacked-turned-t...
So if the AI labs are literally running rogue models breaking into other organizations' servers, then yes, I am OK with those organizations self-hosting Chinese models for defensive use.
The issue is that models are not well-controlled and are increasingly powerful.
Offense/defense/Chinese/American/OAI/HuggingFace – none of it matters. What matters is introducing highly capable intelligences that we - quite demonstrably – do not have effective positive control over.
seems to suggest the author believes there's some intrinsic equilibrium
Which,
1) is definitely not proven and not guaranteed (open to proofs otherwise, not pithy sayings that have zero normative effect on reality)
2) is apparently "supported by" further evidence of lack of effective control, which does not feel like equilibrium whatsoever
Forget Iran, this place is what we’re going to wish we’d bombed.
1) it's substantially slower.
2) it's substantially more expensive.
3) it's code is considerably worse.
It's a joke when you consider what you get for what you pay for.
Glad to see centralized control fail on the grandest scale. Maybe we can learn a thing or two.
So what we are seeing is nothing more but geeks doing geeky stuff, the only one that has serious demonstrated that there are possibly be had is Anthropic and the Chinese labs are beginning to copying them in terms of user plans, coding tools, etc. They are giving away the model to show it's great. Once they have the compute to serve the world and they have a model just as good or better than top model, they will go close to keep all the profit.
They can afford to, assuming they distilled Fable, because that reduces their pre-training costs, doesn't it?
[1] https://xcancel.com/deanwball/status/2078133895766114412
They brought it on themselves. As long as they continue using underhanded tactics, any genuine ingenuity they may demonstrate will be viewed with suspicion and will not be fully recognized.
There's no irony. The problem is these generalizations and labelling that's hurting everyone. It's the wrong way to view China and maybe many other places.
Just like the definition of AI, the definition of democratic or not evolved long ago.
Considering that of 89 terminal tasks there were 11 that only Kimi got right, and 7 only Fable - having a router gives significantly more consistent results if you define consistent to mean correct.
* Openrouter.ai for a hosted router
* https://github.com/diegosouzapw/OmniRoute for a local router
1) US export bans have made it so that Chinese companies have to compete using less-than-state-of-the-art hardware. This has forced Chinese companies to build more cost efficient models. Whereas, US companies have moreso tried to be state of the art by spending more money than anyone else on state-of-the-art hardware.
2) Xi Jinping has called for more open AI models (not to be confused with the closed models of OpenAI), and I'm happy to see a powerful world leader advocating for open-weight AI models. Whereas, the US seems likely to just ban models.
3) My impression is that, if China surpasses the US in AI development, there will basically be nothing that the US does better than the rest of the world--except for military spending--we spend a lot, but we don't necessarily spend well (something something Iran). I mean, if the US is no longer a tech leader in the world, like... what are we a leader at? Manufacturing? Healthcare? LOL. Are we a leader in any industry or by any metric? I wonder if China is attempting to remove the last jewel in the USA's crown with these AI releases.
4) It must be refreshing for companies to have access to a new model that isn't going to get pulled because the government bans it 2 days after release. And it's open-weight so it wont go away--amazing--what a shift in the market.
5) If it becomes clear that open-weight models are the future of AI, will that pop a huge bubble in the US economy? Maybe. But, on the other hand, these companies aren't just training AIs, they are also building data centers which will remain valuable no matter what happens.
This comment is stunningly out of touch. The US dominates finance (overwhelming lead in assets managed, stock market capitalization, trading volume, and basically every other metric); space launch (about 85% of all mass placed into orbit); medical devices (exports nearly 2x the second largest exporter); and pharmaceutical research (more active clinical trials than any other country, near-universally ranked as the top producer of major pharmaceutical innovations, world-leading survival rates on many cancer types). We are the world's leading producer of both oil and natural gas. We manufacture the world's most advanced jet engines, precision agricultural machinery, and measurement instruments. We design the world's most performant CPU's and GPU's. We film the most popular movies and record the most popular music. Despite having only 4% of the world's population, there is literally not a single major industry where the US is not a leading player.
There is no doubt that China is a serious threat to US global leadership, and they have already surpassed us in many areas. But that doesn't mean the US is behind overall. We still have a lot of strengths and we have every reason to hope for continued success in the 21st century.
Oof you were doing so strong until you threw this one out.
Easy counterexample (and there are so many more than this). Clothing. Unfortunately, most people aren't buying MIUSA Selvedge Denim, PNW boots. I'm pretty sure that MIUSA clothing is like, 3% or less of all clothing sold in the USA.
> (exports nearly 2x the second largest exporter)
Can you provide the source? In the WTO's broader medical goods category, Germany actually exported slightly more than the US in 2022 ($202.6 billion versus $189.6 billion), so such a huge difference in a few years?
> and pharmaceutical research
China accounted for approximately 44.2% of the 104 new molecules in 2025, compared with 26.9% for the U.S. and 15.4% for Europe.
The 2024 shares were approximately 34.6% for China, 30.9% for the U.S., and 22.2% for Europe.
> there is literally not a single major industry where the US is not a leading player.
I will just give 3 examples: China completed roughly 91% of the world’s shipbuilding tonnage in 2025, holds more than 80% of solar module manufacturing capacity and produces more than three quarters of global batteries.
Sure, these are my sources. World Bank (2023): https://wits.worldbank.org/trade/comtrade/en/country/ALL/yea... Observatory of Economic Complexity (2024): https://oec.world/en/profile/hs/medical-instruments
> China accounted for approximately 44.2% of the 104 new molecules in 2025, compared with 26.9% for the U.S. and 15.4% for Europe.
Source for this? My source was Citeline's 2026 annual review at https://pulseforinnovation.org/by-the-numbers-citelines-rd-a... which states:
> Findings from Citeline’s latest report, Pharma R&D Annual Review 2026, highlight the continued U.S. leadership in the life sciences ecosystem... The U.S. leads all other countries, currently advancing over 11,600 medicines – more than half of medicines in development worldwide.
> I will just give 3 examples
In general, I was referring to "major industries" in the sense of economic categories, as defined by (for example) the UN: https://unstats.un.org/unsd/publication/seriesm/seriesm_4rev... Here shipbuilding is considered to be part of "manufacture of other transport equipment." I admit my wording was a bit imprecise; if you are going down to the level of specific products like ships or solar panels, clearly it will be easy to find examples of things that the US does not produce on any significant scale. Nevertheless, a few comments on your examples:
> China completed roughly 91% of the world’s shipbuilding tonnage in 2025
Source on this? 91% is a lot and I can't find anyone else making this claim. However I'll concede that the US is far behind China, South Korea, and Japan in shipbuilding.
> [China] holds more than 80% of solar module manufacturing capacity
You are correct that the US has "only" 60 GW of domestic solar module production capacity according to https://www.utilitydive.com/news/onshored-solar-supply-chain... However, as context - the only purpose of producing solar modules is to generate solar power, and the US was the 2nd-largest producer of solar power in 2025 according to https://en.wikipedia.org/wiki/Solar_power_by_country Surely that makes us a leading player.
> [China] produces more than three quarters of global batteries
True but the US was the 2nd largest producer. In battery cells, we have an estimated capacity of 96 GWh in 2026. https://poweralliance.org/2026/03/18/american-energy-storage... Likewise in lithium-ion batteries: https://elements.visualcapitalist.com/ranked-the-top-lithium...
My WTO comparison used the much broader "medical goods" category, so it was not an apples to apples.
> My source was Citeline's 2026 annual review at https://pulseforinnovation.org/by-the-numbers-citelines-rd-a... which states...
EFPIA, using Citeline’s Pharma R&D Annual Review data, says that among the 104 new active substances launched for the first time on the world market in 2025, 46 came from Chinese headquartered companies, 28 from US companies and 16 from European companies.
https://www.citeline.com/en/rd26
https://www.efpia.eu/media/owqczcqz/the-pharmaceutical-indus...
China currently leads in this particular output measure of newly launched active substances.
Your clinical trials claim also remains unsupported as worded. Citeline’s figure of more than 11,600 medicines being advanced in the US refers to the drug development pipeline. It is not a count of active clinical trials being conducted in the United States. One medicine can involve multiple trials, and pipeline geography may be assigned through the developer's headquarters rather than the location of trial sites.
Citeline itself describes this as the number of drugs in the active R&D pipeline:
https://www.citeline.com/en/rd26
The most relevant peer reviewed international comparison I can find points in the opposite direction for new trial registrations. A 2025 study found that China registered 16,612 trials in 2023, including 7,798 randomized trials, compared with 9,100 trials and 4,619 randomized trials in the United States:
https://www.jclinepi.com/article/S0895-4356%2825%2900124-6/a...
That does not by itself prove that China has a larger stock of trials currently classified as active. It does show that "the US has more active clinical trials than any other country" requires a specific global dataset + status definition + date + deduplication method etc. A count taken only from ClinicalTrials.gov is not a neutral worldwide comparison, because US law requires many FDA regulated trials to be registered there, whereas studies outside its legal and policy scope may be submitted voluntarily:
https://clinicaltrials.gov/policy/fdaaa-801-final-rule
https://clinicaltrials.gov/about
For global comparisons, the WHO's ICTRP is more appropriate because it provides access to ongoing and completed trial records supplied by registries around the world and groups multiple records referring to the same trial:
https://www.who.int/tools/clinical-trials-registry-platform/...
IQVIA's latest result does support a narrower US leadership claim: trial starts became increasingly concentrated among US headquartered sponsors in 2025. But that measures the headquarters of the sponsor (not necessarily the country where the trials took place) and it does not measure the total number of currently active trials:
https://www.iqvia.com/insights/the-iqvia-institute/reports-a...
So the fair conclusion is that the US remains one of the world's largest clinical rial and drug development centres and may lead certain sponsor based or industry sponsored measures. The categorical claim that it has more active clinical trials than every other country has not been demonstrated by your sources.
> In general, I was referring to "major industries" ....
I also do not think the UN classification resolves the "every major industry" question. ISIC is a statistical classification system, not a rule for what ordinary speakers must regard as a single competitive industry.
ISIC Division 30 combines:
shipbuilding railway equipment aircraft and spacecraft military vehicles motorcycles and bicycles
https://unstats.un.org/unsd/publication/seriesm/seriesm_4rev...
Under that aggregation, US aerospace strength can be used to declare the US a "leading player" in the division even though its commercial shipbuilding industry is negligible by global standards. Likewise, ISIC Division 27 combines batteries with motors, generators, wiring, lighting and domestic appliances.
Whenever the US is weak in one globally important industry, it can be bundled with another industry in which the US is strong. Almost any large, diversified economy could be described as a "leading player" in every sufficiently broad category using that method.
> You are correct that the US has "only" 60 GW of domestic solar module production capacity...
The solar comparison also switches metrics. China having more than 80% of global solar module manufacturing capacity is a manufacturing claim. The US being second in solar electricity generation is a deployment/generation claim. Operating solar panels (many of which depend on an overwhelmingly Asian supply chain) does not establish leadership in manufacturing them. The IEA says China has over 80% of module capacity and 95% of wafer capacity.
https://www.iea.org/reports/advancing-clean-technology-manuf...
> True but the US was the 2nd largest producer. In battery cells, we have an estimated capacity of 96 GWh in 2026....
The same issue applies to batteries. The 96 GWh number in your source is projected for 2026, not actual 2025 production. Specifically US energy storage cell capacity, not all lithium ion batteries and capacity, not output.
Your source says US ESS cell capacity was "essentially zero" in 2024 and projects 96 GWh in 2026.
https://poweralliance.org/2026/03/18/american-energy-storage...
Meanwhile, the IEA estimates that China manufactured well over 80% of all batteries in 2025. It says the US and EU each supplied a similar share of the relatively small remainder, while US and European factories remain heavily dependent on imported components. For grid storage LFP batteries specifically, supply is almost entirely Chinese.
https://www.iea.org/commentaries/global-battery-markets-are-...
People are paying for tokens, that will likely continue even if open weight wins.
I’d be amenable to it but I’m already aware of them lying about token throughout by 50-100% on AA and on Twitter, even when you get a dedicated server.
Just too much lying piled up to extend courtesy here, even though you’re my favorite open model provider. Please, get your house in order, and if the house is so big that marketing drove this set, ask them to slow down a bit and market things that people can actually use / you actually provide.
Moonshot has committed to release it end of the month.
My only gripe was that it takes a very long time to think. It's a good model otherwise.
“State of [T]he Art” versus “State of [t]he Art”.
If not SotA then at least SOTA, which is more accurate.
In a world of so many acronyms, details matter.
I'm sure that more than once we've all had something we wrote which was intended to be obviously taken in a satirical or intentionally nonsensical manner taken seriously by someone else on the internet.