185 pointsby tamnd4 hours ago23 comments
  • fwlran hour ago
    It is suspicious that OpenAI decided to generate 300 billion output tokens from a model still in training, right after learning there was a credible chance that a major math proof was in that model’s training data. Obviously there are reasonably plausible explanations for each step, but it does sort of feel like parallel construction.
  • jarofgreen2 hours ago
    • pera2 hours ago
      Everything you say can and will be trained against you
    • dude2507114 minutes ago
      Lol, could not get visibility without Twitter.
  • warpech15 minutes ago
    I wonder what’s more valuable in our prompts: the raw data or the feedback system that drives the exchange towards a goal.

    For a long time it was clearly the former, but now I think it is the latter.

    The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.

  • r0ze-at-hn2 hours ago
    Doing some research and at this point doing it very much in the open with dates on GitHub so if any AI Lab says they re-discover my exact work it will be obvious that the AI used or was trained on my work. I am guessing anyone in a similar situation is now thinking about how they date their existing work if the math is done, but the proses are not.
    • bambaxan hour ago
      Yeah but that will not prevent the stealing, it will only make the fight easier afterwards.
    • riedelan hour ago
      That is what arxiv is about. We have been facing the same problem with review processes by before. Nothing all too specific here.
    • calf32 minutes ago
      If only prompts could also be watermarked.
  • Legend24404 hours ago
    This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point".

    They don't even claim to have had a proof, only to have been working on it.

    • rnijveld3 hours ago
      I would say there is a significant difference between AI discovering this completely on its own versus AI creating the finishing connecting part by connecting relevant data. Maybe this claim is too strong, but if part of it is true then the claims that OpenAI have made would be too strong as well.

      To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capturing large amounts of data and connecting the dots.

      • madaxe_againan hour ago
        But this is what we do. Nobody ever invented or discovered anything in a vacuum - all discovery is synthesis of existing ideas and concepts applied to a novel domain. We laud Einstein for instance, but his work was a logical extension of Riemann - Riemann had a neat mathematical toy, Einstein described the universe with it - should we say Einstein was incapable of unique work?
    • itake2 hours ago
      The AI only seem to solve the problems that it had human trading data on…

      If this wasn’t human driven, I’d expect to see other problems within that problem. Space solved not just the ones that it had chat data on.

      • dist-epochan hour ago
        There have been about 6-8 major math breakthroughs claimed by AI. Only for 2 of them there are public accusations about the training data.
        • dgellow25 minutes ago
          That we know of
  • bambaxan hour ago
    All the big AI labs were built on stealing IP; who is surprised that's still how they operate? And who believes, or has ever believed, their promises that your data is private and not logged, etc.?

    The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.

    That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.

    • dakolli34 minutes ago
      It's hilarious how people think they care about their reputation, and wouldn't circumvent ZDR policies. Like bro, they literally covertly hired Apple employees and had them steal IP and equipment form Apple. They aren't scared of Apple lawyers, so they definitely aren't scared of yours.
      • giov44 minutes ago
        what the point and usefulness of the comments above? we shouldn't be surprised? is normal to steal? hiring apple employees?

        can you realize what this means?

        focus on this part:

        "If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself"

        don't threat this as a minor dispute!

        also why not nitter link? not even in comments?

        https://nitter.xitter.cc/ValerioCapraro/status/2097791836269...

  • Cloudef2 hours ago
    Relying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
  • drivebyhooting4 hours ago
    If we put aside the idea of credit for a moment, it sounds like human/AI collaboration is indeed super charging discovery.
    • munksbeer3 minutes ago
      If the allegations are true, I can't see that collaboration lasting. Unfortunately, researches need to earn a living too, and being front run by a lab for everything you do isn't going to pay the bills.
    • matherial2 hours ago
      "Discovery" is not a goal in itself. I could launch a project to find out how many people in the United States have names such that if you assign numbers to every character and then sum the values, the sum works out to 72. It's discovery, but it's useless unless it has some higher goal.

      The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician will ever see in their entire life. They don't care if the findings have any other value to anyone. Mathematicians have very different objectives for their work.

    • PowerElectronix8 minutes ago
      It looks to me more like they made a math engine that can sift through a huge number of combinations, most them absurd, to prove a statement. Just like a chess engine, but for math.

      At least that's what I get from the NS result, they got from a point close to the solution to the solution by making it churn through 10 million bucks of compute.

  • vaylianan hour ago
    This article explains the controversy and the mathematical problem much better than the tweet and toots: https://www.science.org/content/article/how-ai-math-breakthr...
    • dgellow29 minutes ago
      I prefer to read the actual sources for anything related to AI companies given how much AI nonsense journalists seem to accept without any skepticism
    • dakolli31 minutes ago
      Gromov’s soficity conjecture isn't even mentioned in the article you shared.

      Why are you saying that this article explains it much better than the tweet that you clearly didn't even read..

      • vaylian29 minutes ago
        I read the tweet several times but there is so much context missing, that the tweet itself is not enough.
  • mlazos2 hours ago
    It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
    • cm2187an hour ago
      Or start competing with you.
  • profsummergigan hour ago
    Only after reading this post did I learn that my preferred AI trains on my inputs (prompts).

    How was I not aware of this before?

    • madethemcry3 minutes ago
      Don't make this our fault. I would even ask how is this not off by default or why aren't we asked upfront about it if they really care. It's disguising data collection as good faith. I don't even understand how this is legal under GDPR/EU given how much of PII they receive through chats.
    • ga_to30 minutes ago
      Because you have not been paying attention to the discourse regarding AI for the last couple years? That AIs unethical train on data wherever they may get it from has been in the news basically weekly.
    • vaylianan hour ago
      AI is also trained on your HN posts. And lots of other things you post on the internet.
    • cleaning22 minutes ago
      Good question, this was very well known. Do you have an answer?
  • ThalesXan hour ago
    I don't get it, but I'm not an academic.

    If I dedicated my life to curing whatever, warts... and I'm making progress, but it's slow. And then here comes along this tool (LLM), and I use it, and it accelerates my progress to actually finding some sort of thing that makes warts more prone to being eradicated and then the lab throws a couple of million dollars of computes and lo and behold they eliminated warts. If I leave my ego and identity aside, which of course is hard for humans, wouldn't I be glad that warts is cured?

    As a software developer that contributed to open source. Yeah. My code is there. It was the most beautiful code ever written and the labs stole it from me. And now they use it to progress much faster than I ever could. OK. Whatever. It's a tool. I solve problems. Can't I move on from this wart to the next?

    To me, and I know this is gonna get me some heat, it just sounds like academics having their identity ruffled and turning their back to progress in the fields that they chose just because they don't get to play their little decades long of coffee, papers and ultimately identity politics.

    Edit: never got to negative so fast on this board haha. This board is unfortunately turning, or has turned, to Reddit.

    • jaccolaan hour ago
      If these accusation are true

      It’s more like you spend 4 years developing a product you’re passionate about. This product will gain you the respect of all your colleagues and either earn you money directly or lead to great career advancements. Then OpenAI takes it, changes the colour scheme, finishes the login flow and claims the whole thing as their own.

      Not only would it piss you off but it would also misrepresent what OpenAIs models are capable of.

    • yshklarov14 minutes ago
      We love to do work that is useful and valuable to others, and we often form our identities around this. But identities are in large part socially constructed, so many of us need the recognition of others for our contribution. And it can be very painful when we perceive that the credit for our life's work got "stolen". Naturally, we fight against this. There's nothing shameful there. Sure, you can hold onto an ideal of egoless service. There's nothing wrong with that, either. But it's misanthropic to pass such harsh judgment on people for behaving in such a normal and natural manner.
    • frabcus40 minutes ago
      It's partly empathy with the person who did the work and had it stolen, in a field where the main thing people work for is credit. Maths isn't well paid, and doesn't make things that millions of people directly use.

      It's also systemic, it cuts off the supply of results, if there is no reward any more for getting a result, the pipeline of maths will stop. It is the snake eating itself, which has a bad impact for all of us.

    • card_zero20 minutes ago
      If, as you say, it doesn't matter that the AI company gets praise for somebody else's discovery, then it also wouldn't matter if the praise went to the academic. You apparently resent the academic for seeking praise instead of being content with anonymously advancing human knowledge, but you don't resent the AI company seeking praise while leaching off the academic.
    • blensor42 minutes ago
      Let's turn this question around.

      If I have infinite money to progress whatever problem solution I want but I always wait until I have an unfair advantage to get credit for whatever problem was just at the brink of a breakthrough anyway by sniping the last steps. Am I actually doing a good thing or would it be better to let it run it's natural course and spend the money somewhere it's actually needed?

    • Yizahi40 minutes ago
      And I could dedicate my life to helping feed starving kids all across the globe. And then comes along this tool (a lockpick) and I use it, it accelerates my progress to actually getting money to fulfill my dream. If you leave your ego and identity aside, which of course is hard for you, wouldn't you be glad that I stole your money to feed starving kids?
    • alex1138an hour ago
      HN loves drive-by downvotes. It's a real shame.
      • card_zero7 minutes ago
        Downvotes might work as an abuse sponge, absorbing the impulse to make personal attacks. Other than that possible advantage, the downvote functionality seems contradictory to the concept of a discussion forum, I agree.
  • b800han hour ago
    I'm genuinely surprised that more people - including this mathematician in particular - don't untick the "improve the model for everyone" box. Unless the suggestion is that OpenAI ignore this preference?
    • msyan hour ago
      Given OpenAI's well documented history of unethical behaviour it seems adorably naive to think they actually do that in general, or that they wouldn't pull this particular data separately to generate these proofs.
      • olalonde44 minutes ago
        Unethical doesn't mean irrational. They'd be risking massive lawsuits and a total loss of trust if they got caught lying about this. Doesn't seem worth it.
        • dgellow26 minutes ago
          Sounds like exactly what OpenAI would do?
    • afzalivean hour ago
      That doesn't stop them from training on your data apparently. I have that disabled but still has to disable "Don't train on my data" in the privacy center too.

      https://privacy.openai.com/policies?modal=take-control

    • johnnyApplePRNGan hour ago
      They hide that button. Quite well.
    • frabcusan hour ago
      That option is really bad UX - you have to know to do it, you have to know what plan it is needed on. If you're not working in AI, I just don't think that's a reasonable expectation.

      Even if you know, in a complex project over years with multiple collaborators, it just needs one person once to fuck up and paste something into ChatGPT and not realise they weren't logged in, to go wrong.

      In a proper world, we'd at the very least legislate that AI-training on private data needs consent (in the GDPR sense). It's not consent to go "you didn't uncheck a box that lets me steal everything you've done".

      Any training on private data is in my view immoral (it's spying that ultimately will have a chilling effect on even people's private communications). And chats are private data. Unfortunately, it also increases power, so the big tech companies are all doing it.

  • sdcfgy11 minutes ago
    Theft machines be thieving.
  • galkk3 hours ago
    I want bunch of lawsuits, because the way things are described now produces perverse initiatives like try to discuss every possible idea that comes to mind with llm and if any of it works later claim the llm stole it.

    I would like to see chat logs etc and understand how much of a progress was done by human.

  • touweran hour ago
    But China steals our AI!!!!!!
  • Grimblewald3 hours ago
    people seem to miss tge point of this. The problem isn't about credit, its about portraying these models as more competant than they really are. It fuels idiotic statements like jensen huangs recent "agi achieved" statement, which fuels an already dangerous financial fire.
    • ramblerman2 hours ago
      As per the post, this mathematician has been working on this problem for 20 years. So either he was "just" about to breakthrough and this is a big coincidence, or Astra was able to push through the remaining block of 5-10-20-never years it might have taken.

      That's still a pretty big marker of competence in my eyes.

      The point of controversy seems to be who gets credit

      • 8bitsrulean hour ago
        The question's not new. In the early 1900s, women could not become PhD astronomers. Yet two women (Payne with stellar composition and Leavitt with cosmic distances) made fundamental, essential contributions to the science. Credit mostly went to male astronomers. The same might be said of Franklin and DNA.

        It was nearly a century before the stories of all of them were revealed to public history. That the discoverers were not all equally rewarded is unjustifiable.

      • jeltz2 hours ago
        To me that is not a credit thing because this removes a piece evidence for the ability of AI to come up with novel ideas while still making it a useful tool.
      • mentalgear2 hours ago
        The big LLM providers, desperate for good PR before their IPOs, are all actively looking for 'almost finished' hard problems, e.g. where the conceptual / creative parts are almost done and they only need to throw their VC-backed resources at to brute-force through the remaining computationally expensive problem (lean, etc) and claim 'they have solved it'.

        It's an utterly disrespectful, exploitive process, but all in line with exploitative predator capitalism of the stock market and big companies, now exploiting the knowledge / academia domain for scraps with a thin veneer of 'for science' PR.

  • dash22 hours ago
    Can we get a link to the mathstodon post?
  • protocolturean hour ago
    Gonna need grants for local models. Its happening. OpenAI and Anthropic models are powerful but are rapidly approaching the good ol trust thermocline.
  • viccis2 hours ago
    Some mathematicians I know who've been following this have realized that they'd all gotten some emails from people they now know to be affiliated with OpenAI/Anthropic asking questions about their research in a way that seemed like scooping attempts.

    Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.

    All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.

    • mdspan2 hours ago
      Curious, what other options are prospective pure math grad students considering?
      • ethanwillis2 hours ago
        I think Anthropic told them being a plumber is a great option.
    • dist-epochan hour ago
      One has nothing to do with the other.

      It was long predicted that math and software developments would be the first domain where AI was going to do major damage.

      If OpenAI and Anthropic didn't get into math result dick measuring, Internet anons would have in their place, 6 months later when it got cheaper.

  • nobodywillobsrvan hour ago
    The real annoying thing it seems is mostly that openai is presumably doing this for internal reasons and this marginally increases the cost to users with no real gain.

    It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.

    If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.

  • 1337h4xx3 hours ago
    TL/DR: Mathematician opted out of training on 29-JUN and asked OpenAI whether they trained on his data and was told that it "did not happen" but it clearly did.
    • achronoan hour ago
      I've been suspecting over the last couple of years of the frontier companies using data for training anyway, regardless of training-use consent. "Using" the data doesn't have to mean they literally upload chat transcripts into pretraining datasets. My analogy has been money laundering -- if that can happen at massive scales, surely these companies can and will do the digital/data equivalent derivations/transformations. Even if one could have the access etc. to do so, how exactly would one prove that a given synthetic dataset that OAI/Anthropic uses is derived from particular user conversations that did not consent for the info to be used in training?

      Consider, for instance that OpenAI's (consumer) terms say "If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article ." but they also do say "We may use Content to provide, maintain, develop, and improve our Services". [1]

      If you think that's quibbling, consider that OpenAI's business terms, in contrast, do state "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.". [2]

      [1] https://archive.is/EcwD8 [2] https://archive.is/yZdAF

      • calf18 minutes ago
        And humanities have a word for this, exploitation, or appropriation, maybe it's time scientists and engineers revisited basic ethical notions. Skimming a dozen threads and nobody seems to have this vocabulary or willing to say it.
  • ath3ndan hour ago
    [dead]