62 pointsby theanonymousone3 hours ago13 comments
  • redox992 minutes ago
    They probably figured they could increase their margins by releasing "6 Terra" as "6 Sol". After being surprised by the Opus 5.5 launch, they release the true "6 Sol" as "6.1 Sol" and with very aggressive pricing.
  • atkrista2 hours ago
    I would just LOVE to see all the behind-the-scenes shithousery both companies are employing to one-up the other in this, largely, 2-horse AGI race. Someone should make a mockumentary when all is said and done!
    • nbardy2 hours ago
      I think it's weirdly just a choice of deciding to cut releases.

      We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

      • 233mhzan hour ago
        > We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

        I mean, isn't it almost a guarantee that what we get is a gimped version of what they use internally? They probably already serve themselves next gen level models at 1k+ tps from cerebras machines hosted on perm while we get quantized astra/opus at 50tps on a good day

      • iLoveOncall2 hours ago
        > We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

        You're just believing their own bullshit. There's no indication that this is true except from claims from people working at OpenAI.

        If they really had a much more powerful model, it would make absolutely no sense to sit on it.

        • meowfacean hour ago
          With all due respect, you do not have a single clue what you're talking about.
          • owebmasteran hour ago
            With the same respect, you don't either. Simping for openai don't make you part of their in group
            • meowfacean hour ago
              I definitely do not know what I'm talking about, but that individual doesn't know what they're talking about even more than I don't know what I'm talking about.
        • adamzenith2 hours ago
          You don't think having a more intelligent model they can use internally that others can't is an advantage?
          • owebmasteran hour ago
            If that's the case, do Anthropic have an even better one helping them? The Chinese labs too? Where this OpenAI "advantage" is taking them?
          • iLoveOncallan hour ago
            No? The top of human engineers are much better than any model would be, so AI models really aren't a big advantage when you're trying to develop anything that is SOTA.
            • ceejayozan hour ago
              Not every problem is best addressed by a top engineer.

              Plenty of the tasks that keep a company running can benefit from good-enough (and better than the competition).

              • iLoveOncallan hour ago
                Yes, and none of those tasks require even the current SOTA models.
        • 233mhzan hour ago
          > If they really had a much more powerful model, it would make absolutely no sense to sit on it.

          Makes total sense if they don't have the compute and can't serve it in an economically viable way. Also it lets you build things no one else in the world can build as fast as you until it's released

          • iLoveOncall12 minutes ago
            > Also it lets you build things no one else in the world can build as fast as you until it's released

            "We can generate slop faster than anyone in the world" :evil_emoji:

        • simonwan hour ago
          It makes sense for them to sit on it until they've finished testing it. More powerful but also more likely to delete all your email by mistake = you shouldn't release it yet.
        • howdareme9an hour ago
          its not done training, why would they release a model that hasn't finished training?

          besides, we know anthropic are sitting on models too

          • iLoveOncallan hour ago
            > its not done training, why would they release a model that hasn't finished training?

            Because clearly they have no problem with releasing newer versions of models even just a week apart.

            > besides, we know anthropic are sitting on models too

            It is from your crystal ball or from other bullshit you heard from Anthropic employees on Twitter?

            We all know Anthropic had Mythos and Fable, and they turned out to be completely normal models, entirely in line with the capability of their predecessors.

            All they do is lie, and you're believing their lies.

            • meowfacean hour ago
              The poster is not claiming Bel is a secret AGI. Just that it exists and only exists internally at the moment.

              It's rumored to be over 10T parameters. When released it'll probably be very good at certain tasks, albeit slow and expensive and not necessarily "wiser". You don't have to make this a binary.

              Also, Mythos was in fact a significant step-up in several ways. It fits the trend line, but only because the trend line for LLMs is quite steep. Plus Astra is still in many ways less intelligent than Fable/Mythos despite being released much later.

            • azan_an hour ago
              > its not done training, why would they release a model that hasn't finished training?

              > Because clearly they have no problem with releasing newer versions of models even just a week apart.

              You can see how it is pure non-sequitur, right?

              > We all know Anthropic had Mythos and Fable, and they turned out to be completely normal models, entirely in line with the capability of their predecessors.

              Fable was absolutely not in line with capabilities of other models when released. For cybersec work it was much, much better.

      • mFixman2 hours ago
        Any strong enough model with weak enough safeguards can cause an AI Chernobyl event that will make people and governments against AI development and deployment, just like Chernobyl did for nuclear energy.
        • ChromeUltronan hour ago
          tell me "I drank the kool aid" without telling me you drank the kool aid.
          • mFixmanan hour ago
            The US government and most large companies drank the kool aid, and they will be the ones blaming Big AI if things go very wrong.
            • esseph31 minutes ago
              I think there's going to be a brutal backlash against the entire technology sector.

              Gov and Corp will throw their hands up and explain how it's not their fault.

    • michelban hour ago
      I REALLY need new seasons of 'Silicon Valley'
      • tom133730 minutes ago
        Still waiting for the moment where GPT either orders 4,000 pounds of raw beef or just deletes the whole OpenAI repo because it "thought the easiest way to remove all bugs is by deleting the whole repository"
    • petesergeant2 hours ago
      > largely 2-horse

      The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.

      • bayindirh2 hours ago
        Gemini is also pretty nice for researching things. It turns out that having the whole internet indexed and having unlimited access to YouTube is a force multiplier of some kind.

        Since Google has their own TPUs, TPS is also pretty high w.r.t. Claude, for example.

        • thunfischtoastan hour ago
          I think Gemini could be a great product, if they didn't decide to stuff it down my throat at every possible occasion.

          Recent example: on my e-reader, tapping a word I don't now and clicking "Translate" pulls up the possible translations from a dictionary, a local file just a couple of megabytes big, near instantly, on this tiny processor.

          Doing the same on my Android phone starts a Gemini-chat with the prompt "Translate the word x into y". Takes forever, internet access needed, results vary, burns who knows who much energy.

          Why? Just why?

          • bayindirhan hour ago
            That kind of shoving down is pretty bad, I agree.

            I neither use Android devices nor Google Search, so the only Gemini thing I see is the Gemini chat interface.

            I understand the pain, though.

          • selestifyan hour ago
            So that some team at Google can hit their OKRs for user adoption and get promoted.
        • lxgran hour ago
          Gemini is unbelievably bad for research in my experience. It hallucinates like it's 2023, doesn't use its own search, makes up fake rationalizations for why it didn't need to etc.

          It's baffling that OpenAI managed to get better at web search than the company literally synonymous with web search.

          • bayindirhan hour ago
            Pretty interesting. It one-shots correct information with references and a further reading list 99% of the time for me.

            I don't want it to fill in the gaps though, but make it reference anything and everything it brings, hence it doesn't hallucinate much.

            If something feels off, I ask it to back it with concrete data, and if it can't, I don't consider that information correct. That happened once, though, and web doesn't have any information on that thing either. So in that case, not only it had no information on the web, the training data had no information on that thing either short of feeding confidential design documents if they were ever present in the first place.

            The question was about an instrument preamp though, so nothing crucial.

      • Bluestein2 hours ago
        ... and, must be said a plethora of largely unsung, small, unknown "labs", outfits, "researchers" and the like. There is a long tail of smart people having at this. I guess, sheer compute aside, I think much progress - or, at least, important pieces thereof, will come from there.-
        • bayindirh2 hours ago
          There are some niche research areas where bog standard machine learning algorithms make miracles. LLM is just the poster child. AI/ML is a much larger and wider research area.
          • Bluesteinan hour ago
            Your "bog" if unrelated, put me in mind of a "Cambrian explosion" (of intelligence) in this case, ongoing.-

            In a way: We have had "recursive self improvement" (RSI) - that's what genetics is, giving rise to ... us.-

            This time around, I think the truly worrying thing is that, having transcended the biological substrate, the pace is unlike anything previously seen.-

            Just a thought.-

    • TeMPOraL2 hours ago
      AGI will make one, about humanity, after we're all gone - "They were so dumb, they just deserved to die".
      • cindyllm2 hours ago
        [dead]
      • f6v2 hours ago
        The sooner the better, brother.
  • bearjawsan hour ago
    What do people get out of being so reductionist?

    Every time a new model comes out, people come out in droves "oh I don't notice anything different".

    People have been saying this about <currentModel-1> for 2 years now, and the entire state of AI has changed dramatically.

    It cannot be that the next AI model isn't better, but also suddenly what they are capable of is on an entirely different level.

    • ameliusan hour ago
      I suppose it's the reverse of Amara's law:

      "We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run"

    • singpolyma340 minutes ago
      They've been pretty capable for a very long time. I don't think the models are getting more capable so much as people are getting better at using them and more people are getting the opportunity to be impressed.
    • alstonitean hour ago
      I’ve seen overwhelmingly that when a model is good people see it, and when it isn’t, they criticize. I’ve seen nothing but ‘wow this is a huge step up’ from Opus 5.5. I felt this way about Opus 4.5, GPT-5.6 Sol/Luna, and to a lesser extent with Fable and Astra. But Opus 5 was ass, and the entire gpt 6 line feels like OpenAI’s version of that.
    • ulimnan hour ago
      I suspect it's partly because people didn't jump from GPT-3 to GPT-6.1 Sol and partly because SOTA models from the last few(?) months have been able to tackle most of the regular tasks. It means this new model isn't different in that regard from Opus 4.8, if your mental benchmark is that they both are capable of implementing something like a CRUD app.
    • sosukean hour ago
      I like the fast releases. 6.1 coming so fast after 6.0 means they found some improvement solid enough for a new rollout.

      The only time I remember the newest model making an obvious regression was when the first rolled out MoE. Super speed update but each request had less intelligence at hand. We’re way past that now

      • hgoelan hour ago
        It seems strange though.

        Did they find this improvement within the week? If so, given all their whinging about safety, it seems irresponsible to only test the improved model for less than a week.

        Did they find the improvement more than a week ago? If so, why bother releasing GPT 6 if they knew they had a better version essentially ready to go?

        • sznioan hour ago
          Pipelining.

          They probably have 6.2 just about ready, while 7 and 8 are still being cooked up. You can work on multiple releases in parallel.

          We're just getting the latest training checkpoints constantly, just to edge out the other lab, while they are trying to come up with something worthy of a new major version number. It's actually worrying, since there's so much pressure to release now.

    • raincolean hour ago
      I think everyone, I mean everyone, has noticed how different GPT-6 sol is from 5.6. Just not the direction OpenAI hoped for.
    • owebmasteran hour ago
      I don't think a great model would be replaced in a week.

      If the previous one was as good as worth retiring in 7 days, there's a good chance the new one is also not great.

  • moomin2 hours ago
    This has got to be a panic move from OpenAI, right? They’ve had some bad press lately because from their billing changes, and Anthropic have finally released a fast, relatively cheap Opus with improved written English.
    • nba456_2 hours ago
      Didn't they announce the billing changes the same day?
  • Dinuda3 hours ago
    After 5.5, I basically don't notice a jump in model performance, other than my usage ending sooner.
    • baq2 hours ago
      Astra is noticeably smarter than any OpenAI model before it. Sol 6.1 is very noticeably smarter than sol 6 even after half a day of using it (sol 6 was actually terra 6 and opus 5.5 has taken them by a total complete surprise)
      • rimliuan hour ago
        none of them is smart.
    • whazor19 minutes ago
      Meanwhile Fable 5 and Opus 5.5 are in a different league.
    • user439282 hours ago
      I do.

      The results are less buggy, animations are much better.

      It can work autonomously for hours and the result is decent most of the time.

      That wasn't usually the case with 5.5, which needed more feedback and iterations to get things right.

    • tom13372 hours ago
      Kinda same but I miss my 5.3 Codex. Thing lasted forever on my $20 subscription and with detailed prompts was able to pretty much implement everything I requested it to do with a acceptable quality.
    • zero10092 hours ago
      Felt the same until I started using Luna. I feel like I get similar performance, but faster, and my usage lasts so much longer.
    • ModernMech2 hours ago
      Same. 5.5 got work done then 5.6 was also fine then 6 was maybe not quite as good. Now with 6.1 they are cutting usage and raising prices and introducing ultra fast mode, but things were good enough 5 months ago.
    • ndbe2 hours ago
      [dead]
  • egeozcan2 hours ago
    Worries of AI going rogue take so much attention that no governance body seems to care about the shady subscriptions and limits business.
    • entrope2 hours ago
      What about those do you think are "shady"? Price discrimination in favor of small customers at the expense of large customers is somewhat common; businesses want customers to buy more of their products, but customers are not obliged to buy more if they like their current deals. Having limits on using finite resources seems even easier to justify.
      • egeozcanan hour ago
        Shady as in not clearly communicated and change with random tweets being the only indicator, or a blog post if you're lucky.

        x20? 25x? 5x? All in comparison to some other plan that also doesn't have clear limits. "Oh BTW we mean session limits, your weekly is 2x of 5x". "This model eats your limits twice as fast and you can use half of your weekly limits on it". "No this time we mean weekly only".

      • SkyBelowan hour ago
        Humans see differences in prices as unfair. The greatest example of this is price gouging during an emergency, but also look at the level of hate that scalpers get.

        Using AI to do this, if anything, given the common negative sentiment, is seen as even worse. People think of this as AI using information asymetry to squeeze more out of users, not to cut people a deal.

        Taking advantage of information asymmetry is generally looked down upon as well. Look at all the laws we have protecting kids from this. Businesses often don't seem similar protections because those are businesses with big legal teams (and when it is a big legal team vs a small mom and pop store without a single lawyer on payroll, people do start taking issues with it). The power difference between the average company using AI pricing and the average consumer falls pretty solidly in the 'we don't accept this' side of taking advantage of information asymmetry.

        I could keep going, but I think these are already plenty enough reasons to why people look at AI price discrimination as not just a bad business practice they don't like, but an immoral/unethical one.

        Sure, one can make economical counter arguments, but that's arguing on an orthogonal dimension that simply isn't relevant to where these feelings/thoughts come from.

    • semiquaver2 hours ago
      Do you mean the subsidized subscriptions which allow individuals to pay a tenth of the normal API cost for tokens?
      • muyuuan hour ago
        Dumping is shady af.
  • madsgarff2 hours ago
    It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode. I use claude, and I wanna build a feeling for what high, medium, etc. actually gives me. So far their comparisons, and having tried several different models for my work, has given me a feel of what 50 intelligence actually is. And I believe it would be be of even greater value to get a feel inside the single model I actually use, as most people do, because not many, I believe, switch heavily between models when working. I understand that the cost here is greater but the model provivders should obviously give you free access, because of the great work you are doing.
    • nextaccountic2 hours ago
      > It would be very nice if artificialintelligence.ai actually, from a UX perspective, did the models in more than one thinking mode

      But they do. See for example the pareto curve they have, try to locate GPT-6 Luna (max), (xhigh), (high), (medium), (low)

    • rrr_oh_man2 hours ago
      > I wanna build a feeling for what high, medium, etc. actually gives me

      Nothing, really. It's like oversampling your data set. You usually get a much better overall baseline performance if you use the default setting.

    • girvo2 hours ago
      It does!

      …but not for all models, which is pretty annoying.

  • mapontoseventhsan hour ago
    Wow, and it only cost me half my usage!

    Now I can finally build that half a thing I've had my eye on.

  • max9792 hours ago
    Wild pace. Guess they found a critical bug or a quick win to push it out so fast. Astra is a high bar.
  • Pythagon2 hours ago
    Is this a duplicate thread of this? https://news.ycombinator.com/item?id=49896586
  • sscaryterry2 hours ago
    Its nerfed as fuck. Unusable.
    • K0baltan hour ago
      What is your use case?
  • 77rushi772 hours ago
    [flagged]
  • tobin19942 hours ago
    [dead]