For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need to load the full KV in memory to generate a single token:
[0] https://aclanthology.org/2023.findings-emnlp.936/
https://arxiv.org/abs/2503.00392
https://papers.nips.cc/paper_files/paper/2023/hash/6ceefa7b1...
Edit: the claimed reason of "updates" also seems to not exist in code. If the author is going to use LLM for this, the very least they can do is to ask it to check properly before publishing.
> The implementation language is a detail, and a Python port is welcome.
I guess the post title could have dropped "in rust"
You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.
Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?
In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.
We really need to move all AI/ML research off PyTorch to Candle or… just anything that isn’t Python or another old-gen, broken language like it.
The Rust compile cycle will probably generate a gigabyte of files locally as well, but at least they can be `rm`'d out of `target/` once it's done.
It should be said that for this project that's entirely irrelevant of course, but seeing PyTorch has made me skip over projects on the HN homepage before and probably will again in the future.
GPU support should really be a optional extra eg `torch[gpu]` or `torch[nvidia]`.
That said I don't know what's wrong with using a Rust AI framework like Candle.
Weird turn of phrase very typical of AI.
The backbone is also a selective SSM rather than an LSTM controller, so the sequence half is closer to Mamba than to an NTM. DNC comparisons are fair too, and I should have cited both in the readme.
Why did you choose Rust? Why does that matter?
Nothing wrong with Rust. Lots wrong with the bandwagon that “in Rust” somehow adds value.
He could have said "... not written in Python" - would that have been acceptable?
There are a few layers of irony here, but see uv.
A non transformer language model written from scratch is more interesting than
“I did X in Python/Rust/Flavour-of-the-month-thing”