79 pointsby sparticle629 hours ago12 comments
  • JPLeRouzic5 hours ago
    I have read somewhere that Transformer architecture has a quadratic cost [0] (which explains the high costs associated with LLMs and the difficulty for constant improvement without state size pockets).

    For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need to load the full KV in memory to generate a single token:

    [0] https://aclanthology.org/2023.findings-emnlp.936/

    https://arxiv.org/abs/2503.00392

    https://papers.nips.cc/paper_files/paper/2023/hash/6ceefa7b1...

  • prospero_6 hours ago
    Apparently none of the people complaining about rust use read past the title because it's the 2nd section of the readme and impossible to miss.
    • kasumispencer26 hours ago
      That part is newly added after people complained.

      Edit: the claimed reason of "updates" also seems to not exist in code. If the author is going to use LLM for this, the very least they can do is to ask it to check properly before publishing.

    • hashar6 hours ago
      To quote the readme:

      > The implementation language is a detail, and a Python port is welcome.

      I guess the post title could have dropped "in rust"

  • mllev159 hours ago
    You can just tell when the idea itself was generated
  • a194868 hours ago
    Homie just had some tokens to burn at the end of the month and “in Rust” is pure HN clickbait.
  • janalsncm8 hours ago
    OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.

    You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.

    Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?

    In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.

    • yjftsjthsd-h6 hours ago
      I dunno, I could probably be convinced to try a new tool purely on the basis of not having to deal with installing pytorch
      • intoXbox6 hours ago
        I’m curious, what’s the criticism for PyTorch?
        • pseudocomposer17 minutes ago
          The entire Python ecosystem is horrible and shouldn’t have been as falsely boosted by institutions as it was in the 2010s. Yes, it got less bad with 3.8 or whatever version added type annotations. But making so much of ML depend on Python has made it distasteful to a lot of devs who would otherwise have contributed more to it.

          We really need to move all AI/ML research off PyTorch to Candle or… just anything that isn’t Python or another old-gen, broken language like it.

        • jeroenhd5 hours ago
          I don't think it's caused by PyTorch on its own, but every AI-related Python project I try out locally manages to depend on a version of PyTorch that isn't in my disk cache yet. Having to download a gigabyte of dependencies for every project gets tiresome.

          The Rust compile cycle will probably generate a gigabyte of files locally as well, but at least they can be `rm`'d out of `target/` once it's done.

          It should be said that for this project that's entirely irrelevant of course, but seeing PyTorch has made me skip over projects on the HN homepage before and probably will again in the future.

        • tomtom13376 hours ago
          One criticism is that you have to install the same package, torch, but from different Python indexes in order to install the cpu version or gpu version, on Linux. On windows, `pip install torch` gets you the cpu version. On linux, that gets you a ton of Nvidia extras that take a lot of space.

          GPU support should really be a optional extra eg `torch[gpu]` or `torch[nvidia]`.

      • IshKebab6 hours ago
        Yeah likewise. Pytorch needs to die.

        That said I don't know what's wrong with using a Rust AI framework like Candle.

    • alightsoul8 hours ago
      Apparently op cares a lot about speed, which is fine, but ML researchers care about correctness first, speed second. And it makes sense, because they are not as resource constrained as OP.
      • janalsncm7 hours ago
        Most PyTorch tensor operations are cython not python. So imo rewriting in rust is not going to have an enormous speed up. If that really was the concern we should see a throughput comparison vs PyTorch or something.
      • skeledrew4 hours ago
        Are you trying to make an argument here that speed is more important than correctness? I'm finding it difficult to interpret - the purpose of - this comment otherwise, and if you are, I'd consider such an argument pretty wild.
      • lunchbucket5 hours ago
        They aren't using Rust for speed.
      • nicman236 hours ago
        yeah that is why they compute in fp32 lol
  • Marcuss27 hours ago
    How would it compare to current state of the art tested state space layers like Kimi Delta Attention?
  • saglogog6 hours ago
    Nice idea actually, I always wondered why there was no actual programming in ML!
    • skeledrew4 hours ago
      Well, to be a bit pedantic, it's called Machine Learning, not Machine Programming. But also learning is just the other side of the skills transfer coin (that teaching, or programming, is on), so it's more a matter of perspective. And then, actually there has always been - traditional - programming in ML: the data has to be prepped, architecture created, etc.
  • meredithbloom6 hours ago
    > The architecture is the claim here.

    Weird turn of phrase very typical of AI.

  • kasumispencer28 hours ago
    Isn't this just RNN and nearly nothing about this is actually new?
  • anon2919 hours ago
    Seems similar to a neural turing machine.
    • sparticle628 hours ago
      That's the closest prior work, yes, and the memory half sits squarely in that lineage. Three differences worth naming. Reads are content-addressed in the Poincare ball rather than by cosine similarity in R^n, which is what lets a bounded top-4 read cover a hierarchy of contexts instead of a flat neighborhood. There is no controller emitting explicit read and write heads; writes are novelty-gated with a refractory counter that rate-limits overwriting the same slot, so repeated contradictory updates do less damage. And the fast weights do not stay external forever, they get consolidated into the recurrent transition matrix by closed-form ridge regression.

      The backbone is also a selective SSM rather than an LSTM controller, so the sequence half is closer to Mamba than to an NTM. DNC comparisons are fair too, and I should have cited both in the readme.

  • purple-leafy9 hours ago
    “In Rust” man of all the cliche title baits, I hate this one the most.

    Why did you choose Rust? Why does that matter?

    Nothing wrong with Rust. Lots wrong with the bandwagon that “in Rust” somehow adds value.

    • janalsncm8 hours ago
      Rust has its place but python is default for this kind of thing. By doing it in rust, they have now changed two things: the implementation of the transformer, and this new model.
    • echelon8 hours ago
      Counterpoint: I only clicked on it because it said "in Rust".
    • IshKebab6 hours ago
      Rust matters because it means you can actually deploy it without going insane. I wonder if the anti-Rust zealots have actually ever used pip, especially on Windows.

      He could have said "... not written in Python" - would that have been acceptable?

      • skeledrew5 hours ago
        > actually ever used pip, especially on Windows

        There are a few layers of irony here, but see uv.

        - https://docs.astral.sh/uv/

      • purple-leafy5 hours ago
        No he could have left out languages entirely. It’s not relevant for actual discussion.

        A non transformer language model written from scratch is more interesting than

        “I did X in Python/Rust/Flavour-of-the-month-thing”

    • anon2919 hours ago
      Yeah I read the title and did a double-take. It's an uninteresting choice for systems like these. The more pressing concerns, which go undescribed, are the exact mathematical choices behind the actual model. Rust provides almost zero value here because tensor stuff is all just 2d-arrays of floats for the most part.
  • octoberfranklin9 hours ago
    vibe coded vibe coder!