sparse

7 stories and discussions about sparse, aggregated from every source we track.

1.

First, you want to form a habit. Second, you want to operate at peak productivity during your session. Third, you want to minimize the amount you forget between sessions.

4 points•gmays•8 days ago•0 comments•
2.

How I brought NanoGPT training from 73.889 to 39.914 seconds with ANVIL II, sampled softmax, and sparse updates to the bigram and trigram tables.

2 points•Mizza•2 days ago•1 comment•
3.

A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new…

1 points•theanonymousone•1 day ago•0 comments•
4.

When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic…

1 points•vivekkalyanaran•1 day ago•0 comments•
5.
1 points•andsoitis•4 days ago•0 comments•
6.

Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate…

1 points•ksec•7 days ago•0 comments•
7.

NOTE: This blog post assumes knowledge of GPU programming, GPU-driven rendering, D3D12 and systems programming.

1 points•ibobev•10 days ago•0 comments•

Related topics