Attention, Part 4: Linear Attention, Gated DeltaNet and Hybrid Models
The newest idea in attention: replace the ever-growing KV cache with a fixed-size memory. Linear attention, the delta rule, forget gates, Gated DeltaNet, Kimi Delta Attention, gated attention, and the hybrid models built from them (Qwen3-Next, Kimi Linear, Nemotron-H). Explained from zero with equations, paper screenshots and real runs, including three small models trained side by side.
In Part 2, each token stored less. In Part 3, each token looked at fewer earlier tokens. Both kept the same basic picture: a list of keys and values that grows by one entry for every token.
This final part changes that picture. The idea: give each layer a memory of fixed size, and update it once per token, the way a person keeps notes on one page instead of re-reading the whole book. A 1,000-token chat and a 1,000,000-token chat then need the same memory in those layers.
That idea is old (2020), but it only became good enough for frontier models in 2025, thanks to two additions: a smarter way to write (the delta rule) and a way to forget (gates). Today Qwen3-Next, Kimi Linear and Nemotron-H all use it for most of their layers, keeping only a few ordinary attention layers.
As in every part: each new word gets a yellow box, each equation gets a list explaining every symbol, every quote is checked against the paper, and every number comes from code I ran.
1. Gated attention: a small fix for softmax attention#
Before replacing attention, here is the one improvement to ordinary attention that the newest hybrid models use.
Recall the attention sink from Part 1 and Part 3: softmax forces every token's weights to add up to 1, so a head that has nothing useful to say still has to put its weight somewhere. Trained models learn to dump it on the first token. In Part 3, cutting that first token off broke a model completely.
A 2025 paper from the Qwen team tested a simple fix: let each head turn its own output down when it has nothing useful.
The gate is applied right after attention, before the output matrix:
Y′=Y⊙σ(XWθ)
where:
Y is the normal attention output (what Part 1 called "the weighted mix of values"), one vector per token and head;
X is the layer's input (the token vectors);
Wθ is a new learned matrix, so the gate is computed from the token itself;
σ is the sigmoid, so every gate value is between 0 and 1;
⊙ means "multiply number by number" (each output number has its own gate);
Y′ is the gated output, which then goes through the usual output matrix Wo.
Here is that equation as it appears in the paper:
Equation (5) of Qiu et al., 2025, "Gated Attention for Large Language Models".Gated attention. The attention output Y is multiplied, number by number, by a gate between 0 and 1 computed from the input x. Only then does it go through the output matrix.The paper on arXiv (2505.06708). They tested 30 variants on models up to 15B parameters trained on 3.5 trillion tokens.
Why does it help with sinks? With a gate, a head can say "nothing here" by setting its gate near 0. It no longer needs the trick of parking its attention on the first token.
It is in a real model. Qwen3-Next's attention layers do exactly this. From its code in the Hugging Face transformers library, the query projection produces both the query and the gate, and the gate is applied right after attention:
We will measure the sink effect ourselves in section 10.
2. Linear attention: remove softmax, get a fixed-size memory#
Now the big idea. Look again at the attention formula without the mask:
output=softmax(QK⊤)V
Softmax sits betweenQK⊤ and V, so we must first build the full T×T grid of scores. What if there were no softmax?
Without softmax, (QK⊤)V=Q(K⊤V). And K⊤V is only d×d, whatever the length T:
Same numbers, different order. Left: softmax attention must build the T × T grid, which grows with the square of the text length. Right: without softmax, Kᵀ V is computed first, and it is a small d × d matrix whose size never changes.
But we cannot just delete softmax: softmax made all weights positive, and that matters. Katharopoulos et al. (2020) replaced it with a feature map.
Linear attention for token t is then:
ot=∑i≤tϕ(qt)⋅ϕ(ki)∑i≤t(ϕ(qt)⋅ϕ(ki))vi
where:
qt is the current token's query, and ki,vi are the key and value of an earlier token i;
ϕ(qt)⋅ϕ(ki) is the (always positive) score, replacing eqt⋅ki;
the bottom line divides by the total score, so the weights still add up to 1, just like softmax did.
Now move the brackets. Because ϕ(qt) is the same for every i, it can come out of the sum:
where N2 comes from comparing every pair of tokens, and DM is the size of the memory matrix that each token updates. When the text is longer than the head size (N>D), linear attention does less work, and the gap grows with N.
For text written left to right, the sums only run over earlier tokens, which the paper writes like this:
And here is the punchline: St can be updated one token at a time:
St=St−1+ϕ(kt)vt⊤,zt=zt−1+ϕ(kt)
That is an RNN: a fixed-size state, updated once per token. Each step costs the same, no matter how long the text is, and there is no KV cache that grows.
Two kinds of memory. Softmax attention keeps every key and value in a list that grows forever. Linear attention (and everything after it in this part) keeps one fixed-size matrix and updates it once per token.The 2020 paper that started linear attention for transformers (arXiv 2006.16236).
I wrote both forms in part4_linear.py: the parallel one (build the masked T×T grid, like attention) and the recurrent one (one token at a time, keeping only S and z):
python
def linear_attention_parallel(q, k, v): """Like causal attention, but exp(q.k) is replaced by phi(q).phi(k), and there is no softmax.""" Q, K = phi(q), phi(k) scores = (Q @ K.T).tril() # (T, T), future set to 0 return (scores @ v) / scores.sum(-1, keepdim=True)def linear_attention_recurrent(q, k, v): """The same thing as an RNN: a fixed-size state S (d_k x d_v) and a normaliser z (d_k).""" S = torch.zeros(k.shape[-1], v.shape[-1], dtype=q.dtype) z = torch.zeros(k.shape[-1], dtype=q.dtype) out = [] for t in range(q.shape[0]): S = S + torch.outer(phi(k[t]), v[t]) # write: add this token's key-value pair z = z + phi(k[t]) out.append((phi(q[t]) @ S) / (phi(q[t]) @ z)) # read: one small matrix-vector product return torch.stack(out)
plain text
1. Linear attention, parallel form vs recurrent form (T = 200): max |diff| = 4.4e-16 the recurrent form keeps only 272 numbers, however long the text
Identical (the difference is rounding noise). And the recurrent form kept just 16×16+16=272 numbers for 200 tokens. For 200,000 tokens it would still be 272.
A memory that never grows sounds too good to be true. It is. Softmax attention keeps every key and value separately, so it can always find an old fact exactly. Linear attention adds everything into one matrix, and things start to blur.
To see why, write a few facts into S (I drop ϕ and the divisor here to keep it simple) and read one back with its own key ki:
Ski=j∑vj(kj⋅ki)=the fact you wantedvi(ki⋅ki)+noise from every other factj=i∑vj(kj⋅ki)
where S=∑jvjkj⊤ is the memory (in the "value × key" layout used by the Gated DeltaNet paper), and kj⋅ki is how much key j overlaps with key i.
There is a second problem: linear attention can only add. It has no way to change a fact.
Test 1: overwriting. I wrote x = 1, then y = 7, then x = 5 (each fact stored as an outer product of a key and a value, with perpendicular keys for x and y), and read x and y back:
plain text
2. Write "x = 1", "y = 7", then "x = 5". Read x and y back: linear attention: x = 6.00, y = 7.00 delta rule: x = 5.00, y = 7.00
Overwriting a fact. Linear attention can only add, so x comes back as 1 + 5 = 6. The delta rule (next section) replaces the old value and returns 5.
Linear attention says x = 6: it added the new value on top of the old one.
Test 2: capacity. I stored n random facts (random unit-length keys, random values) in a 64×64 memory, then read all of them back and measured the error (0 = perfect, 1 = as wrong as the value itself):
plain text
3. Store n random facts in a 64 x 64 memory, then read all of them back (relative error, 0 = perfect): n = 4: linear attention 0.21 delta rule 0.10 n = 8: linear attention 0.32 delta rule 0.20 n = 16: linear attention 0.49 delta rule 0.33 n = 32: linear attention 0.70 delta rule 0.50 n = 48: linear attention 0.86 delta rule 0.63 n = 64: linear attention 1.00 delta rule 0.73 n = 96: linear attention 1.22 delta rule 0.89 n = 128: linear attention 1.41 delta rule 0.99 n = 192: linear attention 1.74 delta rule 1.12 n = 256: linear attention 2.01 delta rule 1.20
Reading back n random facts from a 64 × 64 memory. The error grows as the memory fills up. The delta rule is always better than plain adding, but neither is perfect: a fixed-size memory has a fixed capacity.
Two lessons. Plain adding gets worse with every fact (by 64 facts, the error is as big as the signal). And even the better rule cannot escape the limit: a fixed-size memory has a fixed capacity. Keep this in mind; it is why hybrid models exist (section 9).
A 2023 study found that this is exactly where cheaper models lose to attention on real text:
The fix for overwriting is old: the delta rule (Widrow and Hoff, 1960), brought to transformers as "DeltaNet" (Schlag et al., 2021; Yang et al., 2024). The idea in plain words:
Before writing, look up what the memory currently says for this key: St−1kt.
Compare it with the value you want to store, vt. The difference is the error ("delta").
Change the memory by just that error, scaled by a writing strength βt.
As an equation (memory in "value × key" layout, keys of length 1):
St−1kt is the old value the memory holds for key kt;
vt−St−1kt is the error: what we want minus what is there;
βt, between 0 and 1, is the writing strength (1 = replace completely, 0.5 = move halfway);
I is the identity matrix (the matrix that changes nothing); I−βtktkt⊤ removes part of whatever was stored along the direction kt;
βtvtkt⊤ then writes the new value there.
With βt=1 and a unit-length key, reading right after writing gives exactly Stkt=vt: the old value is gone. That is why the test above returned x = 5.
Here is the beautiful part: the delta rule is one step of gradient descent, taken at every token, on the error "how wrong is my memory for this key?". The Kimi Linear paper writes it this way (in their layout, keys and values swap sides):
The delta rule as one gradient-descent step, from the Kimi Linear paper (Section 2). The loss being reduced is ½‖Sᵀk − v‖².Lt(S)=21Skt−vt2
where Lt is the squared error between what the memory returns for key kt and the value vt it should return, and βt plays the role of the step size. So the memory is literally learning while it reads, one token at a time. This is why this family is sometimes called "test-time training".
In code it is one line:
python
def write_delta(S, k, v, beta=1.0): return S - beta * torch.outer(S @ k - v, k) # S (I - beta k k^T) + beta v k^T
The delta rule fixes one fact at a time. But sometimes the model needs to forget a lot at once: a new paragraph, a new topic, a new document. For that, models add a forget gate.
Mamba-2 (2024) can be written as linear attention with exactly this gate. As written in the Gated DeltaNet paper:
Mamba-2 as a gated linear attention, from the Gated DeltaNet paper (Section 2).St=αtSt−1+vtkt⊤,ot=Stqt
where αt is the forget gate (computed from token t), vtkt⊤ writes the new fact, and ot=Stqt reads the memory with the query.
How fast do facts fade? It depends strongly on α:
plain text
4. Strength of a fact after t more tokens, with forget gate alpha (alpha^t): alpha = 0.9 : t=0: 1.000 t=10: 0.349 t=50: 0.005 t=100: 0.000 t=500: 0.000 alpha = 0.99 : t=0: 1.000 t=10: 0.904 t=50: 0.605 t=100: 0.366 t=500: 0.007 alpha = 0.999: t=0: 1.000 t=10: 0.990 t=50: 0.951 t=100: 0.905 t=500: 0.606
How much of a stored fact is left after t more tokens. With α = 0.9 a fact is almost gone after 50 tokens; with α = 0.999 more than half survives 500 tokens. Because α is computed from each token, the model can choose: remember (α near 1) or wipe (α near 0).
The important word is data-dependent: αt is computed from token t itself. So the model can keep α near 1 inside a paragraph and drop it towards 0 at a topic change.
Gating is good at forgetting everything a bit. The delta rule is good at changing one fact exactly. Gated DeltaNet (Yang, Kautz and Hatamizadeh, ICLR 2025) uses both:
St is the memory, a dv×dk matrix (one per head);
αt∈(0,1) is the forget gate: shrink the whole memory a little (or a lot);
βt∈(0,1) is the writing strength of the delta rule;
I−βtktkt⊤ removes (part of) the old value stored at key kt;
βtvtkt⊤ writes the new value at key kt;
qt is the query; ot=Stqt reads the memory;
if αt=1 it becomes the plain delta rule; if βt=0 it only forgets.
One Gated DeltaNet step, in four moves: forget a little, look up what is stored at the key, correct it towards the new value, and read the answer with the query.
There is a neat way to see all these memories as one family. At each token, the new memory St is the answer to a tiny optimisation problem: stay close to the old memory, but fit the new fact.
Both are computed from the token by small learned layers, so the model decides token by token. In Qwen3-Next's code:
python
beta = b.sigmoid()# If the model is loaded in fp16, without the .float() here, A might be -infg = -self.A_log.float().exp() * F.softplus(a.float() + self.dt_bias)
where:
b and a are produced from the token by a learned matrix;
beta = sigmoid(b) keeps βt between 0 and 1;
g is logαt, the forget gate stored as a logarithm, and αt=eg;
softplus(x)=log(1+ex) is always positive and A_log.exp() is positive, so g is always negative and αt=eg is always between 0 and 1.
A real Gated DeltaNet layer adds a few more parts around this core: queries and keys are normalised to length 1 (so ktkt⊤ behaves well), a short convolution mixes each token with its 3 neighbours before the memory step, and the output passes through a norm and an output gate. The Gated DeltaNet paper describes the query and key path as "linear proj., shortconv., SiLU and L2 norm".
I wrote the gated delta rule from scratch, line by line from the equation:
python
def gated_delta_rule(q, k, v, alpha, beta): """S_t = alpha_t * S_{t-1} (I - beta_t k_t k_t^T) + beta_t v_t k_t^T ; o_t = S_t q_t. q, k: (H, T, d_k) L2-normalised; v: (H, T, d_v); alpha, beta: (H, T) in (0, 1).""" H, T, dk = k.shape S = torch.zeros(H, v.shape[-1], dk, dtype=q.dtype) out = [] for t in range(T): kt, vt = k[:, t], v[:, t] S = alpha[:, t, None, None] * S # 1. forget a little (gate) S = S - beta[:, t, None, None] * torch.einsum('hv,hk->hvk', torch.einsum('hvk,hk->hv', S, kt) - vt, kt) # 2. correct (delta) out.append(torch.einsum('hvk,hk->hv', S, q[:, t])) # 3. read return torch.stack(out, 1)
The transformers library ships Qwen3-Next's reference implementation in two versions: a step-by-step one (used when generating) and a chunked one (used for long inputs: it processes 64 tokens at a time with clever algebra, giving the same result). I compared mine with both, on 4 heads and 256 tokens of random inputs:
plain text
5. Our gated delta rule vs the Qwen3-Next reference code in transformers (float32 inside): vs step-by-step version: max |diff| = 3.0e-08 vs chunked version: max |diff| = 2.3e-07 (outputs up to 0.17)
Differences around 10−7 are the rounding noise of 32-bit numbers, which their code uses internally. So the equation above really is what runs inside Qwen3-Next.
7. Kimi Delta Attention: a forget rate for every channel#
Kimi Linear (Moonshot AI, 2025) makes one more change. In Gated DeltaNet, αt is one number per head: the whole memory of a head fades at the same speed. Kimi Delta Attention (KDA) gives every channel its own forget rate.
Equation (1) of the Kimi Linear paper: Kimi Delta Attention. (Their memory is stored as d_k × d_v, the transpose of the Gated DeltaNet layout, so keys and values swap sides.)St=(I−βtktkt⊤)Diag(αt)St−1+βtktvt⊤,ot=St⊤qt
where everything is as in Gated DeltaNet, except that αt is now a vector (one forget rate per key channel) instead of a single number.
Why would that help? Some channels can hold long-lived facts (α near 1) while others act as scratch space that is quickly cleared (α small), all inside one head.
Written out entry by entry, the forget step of KDA is
(Diag(αt)St−1)ij=αt,i⋅(St−1)ij
where i indexes the key channels (rows of KDA's dk×dv memory) and j the value channels. Every entry in row i fades at its own rate αt,i. In Gated DeltaNet the same step is αt⋅(St−1)ij: one rate for the whole head.
Each new token costs a linear layer one memory update, the same size every time. Softmax attention instead reads the whole cache. I timed one decode step of one layer, using Qwen3-Next's real layer shapes (gated attention: 16 query heads sharing 2 key/value heads of size 256; Gated DeltaNet: 32 heads with a 128×128 memory each), on an Apple M5 Pro GPU:
plain text
7. One decode step on mps (Qwen3-Next layer shapes): 4096 tokens so far: gated attention 0.097 ms Gated DeltaNet 0.180 ms 16384 tokens so far: gated attention 0.284 ms Gated DeltaNet 0.181 ms 65536 tokens so far: gated attention 1.238 ms Gated DeltaNet 0.180 ms 262144 tokens so far: gated attention 5.831 ms Gated DeltaNet 0.180 ms
Time for one decode step of one layer. Attention grows with the conversation; the Gated DeltaNet step stays flat at 0.18 ms. Below about 10,000 tokens attention is actually faster in this test; beyond that, the fixed-size memory wins by more and more.
Note the honest detail: at 4,096 tokens, attention was faster (0.097 vs 0.180 ms). My Gated DeltaNet step is several small unfused operations, and a 32×128×128 memory is not tiny. The advantage appears only once the conversation is long, and then it keeps growing: 32× faster at 262K tokens.
9. Hybrid models: a few attention layers keep recall sharp#
Section 3 showed the weakness: a fixed-size memory has fixed capacity, so it cannot recall every detail of a long text exactly. Softmax attention can. The solution all three of today's big designs use: mostly linear layers, plus a few full attention layers.
Layer patterns from the published configurations. Qwen3-Next: 3 Gated DeltaNet layers, then 1 gated attention layer, repeated (36 + 12). Kimi Linear: 3 KDA layers, then 1 MLA layer (20 + 7). Nemotron-H-8B: 24 Mamba-2 layers, 24 MLP-only layers, and just 4 attention layers. Hover over a layer to see its type.
Qwen3-Next-80B (Qwen, 2025). From its configuration on Hugging Face:
full_attention_interval: 4 means every 4th layer is (gated) full attention: 12 of 48. The other 36 are Gated DeltaNet, each with 32 heads holding a 128×128 memory, and a short convolution of 4 tokens.
Kimi Linear 48B (Moonshot AI, 2025): 27 layers; layers 4, 8, 12, 16, 20, 24 and 27 are MLA (the compressed attention from Part 2), the other 20 are KDA.
Nemotron-H (NVIDIA, 2025) uses Mamba-2 (the gated linear attention from section 5) instead of a delta rule:
For attention layers the memory grows with the text; for linear layers it is fixed. With 2 bytes per number:
memory=attention layers: grows with nLattn×2×Hkv×dh×n×2B+linear layers: fixedLlin×H×dk×dv×2B
where n is the number of tokens, Lattn and Llin count the two kinds of layers, and the other symbols are the head counts and sizes from the configuration.
plain text
6. Memory, 2 bytes per number Qwen3-Next-80B at 262144 tokens: 12 attention layers 6.00 GiB (all 48 as attention: 24.00 GiB), plus a fixed 36 MiB of Gated DeltaNet state Kimi Linear 48B: 8064 B per token with 7 MLA layers vs 31104 B if all 27 were MLA (74% less), plus a fixed 20 MiB of KDA state
Memory for one conversation in Qwen3-Next-80B: if all 48 layers were attention (orange) versus the real hybrid (blue). At 262K tokens: 24 GiB versus about 6 GiB.
Qwen3-Next at 262K tokens: 24 GiB if every layer were attention, 6 GiB as built, plus a fixed 36 MiB for all the Gated DeltaNet memories together.
Kimi Linear: 7 MLA layers out of 27 means 74% less cache per token, matching the paper's "up to 75%".
10. Experiment: three small models, one difference#
Equations and timings are one thing. Do these layers actually learn as well as attention? I trained three small models that are identical except for their token-mixing layers:
Model
Its 4 layers
softmax attention
4 × softmax attention (with RoPE)
gated attention
4 × softmax attention + output gate (section 1)
linear attention
4 × linear attention, ϕ=elu+1 (section 2)
Everything else is the same: 4 layers, width 256, 4 heads, the same feed-forward blocks, the same optimizer, the same random seed, the same data, the same number of steps. Each model has about 3.3 to 3.5 million parameters.
Task 1: predicting text. Each model reads public-domain books one byte at a time and learns to predict the next byte (1,500 steps of 32 × 256 bytes, from six books). It is then tested on a book it never saw, The Adventures of Tom Sawyer. The score is bits per byte: lower is better.
Task 2: recall. The model sees n facts (a key and its value at each position), then n questions (just a key, in a new order), and must answer each with the right value. Keys come from 1,024 possible tokens and values from 256. Each model trains for 3,000 steps on random sets of 32 to 256 facts, then is tested on fresh ones. This is a simplified, one-step version of the "multi-query associative recall" test from the Zoology paper (Arora et al., 2023), which found that recall explains most of the gap between attention and its cheaper rivals.
Bits per byte on a book none of the models saw during training (lower is better). Softmax and gated attention are almost tied; linear attention is clearly behind.
Model
Bits per byte (lower is better)
softmax attention
1.994
gated attention
1.989
linear attention
2.267
Gated attention is slightly better than plain softmax (1.989 vs 1.994). That is the direction the paper reports, but the gap here is tiny, and with one run per model I would not call it a real win at this scale.
Linear attention is clearly worse (2.267, about 14% more bits). Its fixed memory cannot keep exact track of the recent letters and words the way softmax attention can. This is the "fixed memory gets confused" problem from section 3, showing up in real training. It is exactly what the delta rule and gates were invented to fix.
Share of lookups answered correctly as the number of facts grows. Softmax attention stays at 100%. Linear attention starts near 100% but falls to 78% at 256 facts: its memory is full. Gated attention (this run) failed to learn the task in 3,000 steps; see the note below.
Facts to remember
32
64
128
256
softmax attention
100%
100%
100%
100%
linear attention
99.8%
99.3%
95.6%
78.1%
gated attention (seed 0)
15.9%
14.7%
12.2%
7.9%
Linear attention shows its capacity limit. With few facts it is nearly perfect, but as the facts pile up, its fixed memory (4 heads of 64×64 per layer) starts mixing them, exactly like the capacity test in section 3. Softmax attention keeps every fact separately and stays at 100%.
The gated attention paper says the gate removes attention sinks. I measured how much attention the first token receives in the trained softmax and gated models (queries from position 64 on, averaged per layer):
plain text
softmax first-token attention per softmax layer: 0.003 0.001 0.000 0.000gated first-token attention per softmax layer: 0.003 0.000 0.000 0.000
Both are near zero: these tiny models never formed an attention sink at all, so there was nothing for the gate to remove. Sinks are known to appear in larger models trained for much longer, and my byte-level texts start mid-sentence with no special first token. So this experiment cannot confirm or refute the paper's 46.7% → 4.8% result; it only shows that sinks are not automatic.
In 2024, linear attention was mostly a research topic. By late 2025, three major model families shipped it in most of their layers. The results the companies reported are what made the difference:
Use cases in one line each:
Very long documents and codebases (hundreds of thousands to millions of tokens): hybrids keep memory and per-token cost almost flat.
Reasoning models and agents that write very long outputs: each new token stays cheap, as in Kimi Linear's 6× faster decoding at 1M tokens.
High-volume serving: more requests per GPU (Nemotron-H's 2.4× throughput) means lower cost per answer.
Devices with little memory: a fixed-size state means a long conversation does not need a growing cache.
Where full attention still wins: exact lookups of details far back in the text. That is why every model above keeps some attention layers.
Gated attention multiplies each attention output by a sigmoid gate, Y′=Y⊙σ(XWθ). The paper cut first-token attention from 46.7% to 4.8%; Qwen3-Next uses it.
Linear attention replaces softmax with ϕ(q)⋅ϕ(k). Moving the brackets turns attention into an RNN with a fixed memory St=St−1+ϕ(kt)vt⊤. The parallel and recurrent forms matched to 4.4×10−16.
A fixed memory gets confused: plain adding cannot overwrite (x came back as 6 instead of 5), and errors grow as facts pile up.
The delta rule writes by correcting: St=St−1+βt(vt−St−1kt)kt⊤, one step of gradient descent per token. It overwrote correctly and stored random facts with less error.
Forget gatesαt fade old memories (αt), chosen token by token.
Gated DeltaNet combines both: St=St−1αt(I−βtktkt⊤)+βtvtkt⊤. My version matched Qwen3-Next's reference code to 3×10−8. KDA gives every channel its own forget rate.
One Gated DeltaNet decode step stayed at 0.18 ms at any length; attention grew to 5.8 ms at 262K tokens.
Hybrids keep a few full attention layers for exact recall: Qwen3-Next (36 + 12), Kimi Linear (20 + 7, 74% less cache), Nemotron-H (4 attention layers of 52).
In my three small models, linear attention was clearly worse at predicting text (2.267 vs 1.994 bits per byte) and its recall fell to 78% at 256 facts, while softmax attention stayed at 100%. Gated and plain softmax were too close to rank from single runs.
That is the end of the series. Attention started in 2014 as "let the decoder look back at every word"; in 2025 the best models look back with only a handful of layers, and remember everything else in memories that never grow.
Run it yourself
code/attention/part4_linear.py: linear attention forms, overwrite and capacity tests, forget-gate fading, the comparison with Qwen3-Next's reference code, memory and decode timing. Runs in about a minute.
code/attention/part4_train.py: the small models, trained on books and on the recall task. Pass model names to choose: python part4_train.py softmax gated linear (what this article reports). It can also train gdn and hybrid, but those are very slow without the special GPU kernels. SEED=1 changes the random start and RECALL_ONLY=1 skips the book task, as in the extra checks.