transformer

12 stories and discussions about transformer, aggregated from every source we track.

1.

A custom AI architecture being developed in rust . Contribute to Sparticle62ops/pssa development by creating an account on GitHub.

77 points•sparticle62•about 21 hours ago•32 comments•
2.

Train a tiny GPT in under a minute (CUDA only). Contribute to lostmsu/TurboGPT development by creating an account on GitHub.

53 points•lostmsu•1 day ago•10 comments•
3.

Posted by Jakob Uszkoreit, Software Engineer, Natural Language Understanding Neural networks, in particular recurrent neural networks (RNNs), are n...

4 points•adsouza•about 9 years ago•0 comments
4.

Paper Tape is All You Need. Contribute to dbrll/ATTN-11 development by creating an account on GitHub.

3 points•raymii•6 months ago•0 comments
5.

While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model…

2 points•sbulaev•5 days ago•0 comments•
6.

I've been learning about Vision Transformers, and it's been a ride. It's probably the densest...

2 points•marshateo•8 days ago•0 comments
7.

Performing inference on a Transformer can be very different from training. Partly this is because inference adds a new factor to consider: latency. In this section, we will go all the way from sampling a single new…

2 points•aray07•9 days ago•0 comments•
8.
1 points•ksec•about 7 hours ago•0 comments•
9.
1 points•yarapavan•about 23 hours ago•0 comments•
10.

Nemotron-H is a series of hybrid Mamba-Transformer models which offer either better or on-par accuracy and improved inference speed (up to 3x) compared to other similarly-sized state-of-the-art open-sourced pure…

1 points•Bluestein•8 days ago•0 comments•
11.

Paradigma releases Limite 1B - Violetto, a 1-billion parameter transformer for high-throughput mathematical reasoning, alongside its evaluations, a training-time value model, and a custom vLLM inference plugin.

1 points•E-Reverance•9 days ago•0 comments•
12.

A visual technical deep dive into Jev: transformer inference, typed probabilities, proper scoring rules, policy gradients, and what RLCD changes.

1 points•madmax108•10 days ago•0 comments•

Related topics