A custom AI architecture being developed in rust . Contribute to Sparticle62ops/pssa development by creating an account on GitHub.

81 points•sparticle62•1 day ago•35 comments•

35 comments

JPLeRouzicabout 20 hours ago
I have read somewhere that Transformer architecture has a quadratic cost [0] (which explains the high costs associated with LLMs and the difficulty for constant improvement without state size pockets).

For what I understand PSSA belongs to a line of research for LLMs with scalable architecture because you don't need to load the full KV in memory to generate a single token:

[0] https://aclanthology.org/2023.findings-emnlp.936/

https://arxiv.org/abs/2503.00392

https://papers.nips.cc/paper_files/paper/2023/hash/6ceefa7b1...

prospero_about 21 hours ago
Apparently none of the people complaining about rust use read past the title because it's the 2nd section of the readme and impossible to miss.
kasumispencer2about 21 hours ago
That part is newly added after people complained.

Edit: the claimed reason of "updates" also seems to not exist in code. If the author is going to use LLM for this, the very least they can do is to ask it to check properly before publishing.

hasharabout 21 hours ago
To quote the readme:

> The implementation language is a detail, and a Python port is welcome.

I guess the post title could have dropped "in rust"

mllev15about 24 hours ago
You can just tell when the idea itself was generated
a19486about 22 hours ago
Homie just had some tokens to burn at the end of the month and “in Rust” is pure HN clickbait.
janalsncmabout 23 hours ago
OP, you should not have written this in Rust. It should be in PyTorch, which is by far the most popular. We can’t tell if this architecture is good or whether there is a problem in your implementation.

You can test the whole thing for free on a GPU with Google Colab. Test both the transformer and your new architecture on a larger dataset. Something that maxes out the GPU for an hour each run.

Also, the readme mentions keeping the same optimizer schedule which sounds nice at first but they are completely different architectures. The loss is high on the transformer, did you try raising the learning rate on it?

In general I’m interested in parameter efficient architectures. I don’t think transformers are optimal, and indeed many improvements have been made to vanilla transformers. But if you have an idea for something better you need to show it.

yjftsjthsd-habout 21 hours ago
I dunno, I could probably be convinced to try a new tool purely on the basis of not having to deal with installing pytorch
intoXboxabout 21 hours ago
I’m curious, what’s the criticism for PyTorch?
IshKebababout 21 hours ago
Yeah likewise. Pytorch needs to die.

That said I don't know what's wrong with using a Rust AI framework like Candle.

alightsoulabout 23 hours ago
Apparently op cares a lot about speed, which is fine, but ML researchers care about correctness first, speed second. And it makes sense, because they are not as resource constrained as OP.
janalsncmabout 21 hours ago
Most PyTorch tensor operations are cython not python. So imo rewriting in rust is not going to have an enormous speed up. If that really was the concern we should see a throughput comparison vs PyTorch or something.
skeledrewabout 19 hours ago
Are you trying to make an argument here that speed is more important than correctness? I'm finding it difficult to interpret - the purpose of - this comment otherwise, and if you are, I'd consider such an argument pretty wild.
lunchbucketabout 19 hours ago
They aren't using Rust for speed.
nicman23about 21 hours ago
yeah that is why they compute in fp32 lol

Read the full thread on Hacker News →

Related stories