transformer
12 stories and discussions about transformer, aggregated from every source we track.
A custom AI architecture being developed in rust . Contribute to Sparticle62ops/pssa development by creating an account on GitHub.
Train a tiny GPT in under a minute (CUDA only). Contribute to lostmsu/TurboGPT development by creating an account on GitHub.
Posted by Jakob Uszkoreit, Software Engineer, Natural Language Understanding Neural networks, in particular recurrent neural networks (RNNs), are n...
Paper Tape is All You Need. Contribute to dbrll/ATTN-11 development by creating an account on GitHub.
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model…
I've been learning about Vision Transformers, and it's been a ride. It's probably the densest...
Performing inference on a Transformer can be very different from training. Partly this is because inference adds a new factor to consider: latency. In this section, we will go all the way from sampling a single new…
Nemotron-H is a series of hybrid Mamba-Transformer models which offer either better or on-par accuracy and improved inference speed (up to 3x) compared to other similarly-sized state-of-the-art open-sourced pure…
Paradigma releases Limite 1B - Violetto, a 1-billion parameter transformer for high-throughput mathematical reasoning, alongside its evaluations, a training-time value model, and a custom vLLM inference plugin.
A visual technical deep dive into Jev: transformer inference, typed probabilities, proper scoring rules, policy gradients, and what RLCD changes.