Declarative Attention: A Model That Says Which Parts to Reread

A new method lets a long-context language model declare which parts of its input it needs, cutting the context it reads by up to 52% for a small accuracy cost, with no retraining.
artificial-intelligence
Author

Kabui, Charles

Published

2026-09-29

Keywords

declarative-attention, long-context, kv-cache, sparse-attention, chain-of-thought