attention
22 stories and discussions about attention, aggregated from every source we track.
The Tetris effect is one of psychology’s most easy to reproduce experiments.Simply spend a bit of time playing the eponymous game every day for a few weeks.A...
The Tetris effect is one of psychology’s most easy to reproduce experiments.Simply spend a bit of time playing the eponymous game every day for a few weeks.A...
Ignoring licenses can be dangerous. Let's try to understand them.
Cambridge scientists have shown that girls and boys differ in what they pay attention to, even at birth. Since these differences are present so early, it is possible they emerge due to prenatal factors. The findings…
Cambridge scientists have shown that girls and boys show differences in what they pay attention to, even at birth. Since these differences are present so early, it is possible they emerge due to prenatal factors.
The Vatican Dicastery for Communication releases the theme for the 2027 World Communications Day, which is “A New Gaze: Paying Attention in the Age of ...
Replaceme checks the communication already on your Mac and sends one concise iMessage when something appears to need attention.
A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new…
It figures out when a task needs my attention based on how much work it might take. I’ve used Getting Things Done on and off for years. My problem with GTD is simple: it only works as well as I remember to check it.…
Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate…
I found this great blog post: Attention is all you have on the front page of Hacker News. It is well worth your time to read it, but what it is essentially try…
AI coding interfaces present generated code either all at once or token-by-token. These rendering strategies reflect model generation rather than how programmers actually read code: selectively, non-linearly, and…
Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity. Prior work has shown that composing two delta-rule transitions…
The first attention kernel proven minimal BEFORE code was written - womenflyplanes/moa-attention-verified-mullin