Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity. Prior work has shown that composing two delta-rule transitions…
0 comments
No comments yet.
Related stories
- Hacker News · 2 points · 10 days ago
- Hacker News · 1 points · 9 days ago
- Hacker News · 5 points · 7 days ago
- Hacker News · 8 points · 8 days ago
- Hacker News · 3 points · 8 days ago
- Transformer: A Novel Neural Network Architecture for Language Understandingresearch.googleblog.comLobsters · 4 points · about 9 years ago