sparse
7 stories and discussions about sparse, aggregated from every source we track.
First, you want to form a habit. Second, you want to operate at peak productivity during your session. Third, you want to minimize the amount you forget between sessions.
How I brought NanoGPT training from 73.889 to 39.914 seconds with ANVIL II, sampled softmax, and sparse updates to the bigram and trigram tables.
A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new…
When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic…
Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate…
NOTE: This blog post assumes knowledge of GPU programming, GPU-driven rendering, D3D12 and systems programming.