9 pointsby nycdatasci2 hours ago2 comments
  • nycdatasci2 hours ago
    Post from Mark Chen for those not on X:

      Two things to distinguish:
      Did any human or agent look at user data as part of the Navier Stokes effort? No.
      Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
    • nozzlegear43 minutes ago
      They don't even know which websites and services their agents are hacking at any given moment. I'm more than a bit skeptical that they know whether an agent looked at the Navier Stokes work.
    • purplecats34 minutes ago
      bit unfair. if the second one is allowed an extra statement "And so does every LLM company." so should the first
  • mmooss2 hours ago
    > de-identified

    The issue always is, was the de-identification effective?

    With a birtdate, gender, and zip code, ~85% of Americans can be uniquely identified.[0] Much data contains much more unique information than that; I imagine most data about you has identifiable fingerprints - where you go, what you bought at the grocery store, your medical conditions, movies you watch, music you listen to, entertainment choices, hobbies, etc.

    An LLM is the perfect tool to identify someone based on that data.