Attention, Part 4: Linear Attention, Gated DeltaNet and Hybrid Models

The newest idea in attention: replace the ever-growing KV cache with a fixed-size memory. Linear attention, the delta rule, forget gates, Gated DeltaNet, Kimi Delta Attention, gated attention, and the hybrid models built from them (Qwen3-Next, Kimi Linear, Nemotron-H). Explained from zero with equations, paper screenshots and real runs, including three small models trained side by side.

In Part 2, each token stored less. In Part 3, each token looked at fewer earlier tokens. Both kept the same basic picture: a list of keys and values that grows by one entry for every token.

This final part changes that picture. The idea: give each layer a memory of fixed size, and update it once per token, the way a person keeps notes on one page instead of re-reading the whole book. A 1,000-token chat and a 1,000,000-token chat then need the same memory in those layers.

That idea is old (2020), but it only became good enough for frontier models in 2025, thanks to two additions: a smarter way to write (the delta rule) and a way to forget (gates). Today Qwen3-Next, Kimi Linear and Nemotron-H all use it for most of their layers, keeping only a few ordinary attention layers.

As in every part: each new word gets a yellow box, each equation gets a list explaining every symbol, every quote is checked against the paper, and every number comes from code I ran.

1. Gated attention: a small fix for softmax attention

Before replacing attention, here is the one improvement to ordinary attention that the newest hybrid models use.

Recall the attention sink from Part 1 and Part 3: softmax forces every token's weights to add up to 1, so a head that has nothing useful to say still has to put its weight somewhere. Trained models learn to dump it on the first token. In Part 3, cutting that first token off broke a model completely.

A 2025 paper from the Qwen team tested a simple fix: let each head turn its own output down when it has nothing useful.

The gate is applied right after attention, before the output matrix:

Y′=Y⊙σ(XWθ)Y' = Y \odot \sigma(X W_\theta)

where:

  • YY is the normal attention output (what Part 1 called "the weighted mix of values"), one vector per token and head;
  • XX is the layer's input (the token vectors);
  • WθW_\theta is a new learned matrix, so the gate is computed from the token itself;
  • σ\sigma is the sigmoid, so every gate value is between 0 and 1;
  • ⊙\odot means "multiply number by number" (each output number has its own gate);
  • Y′Y' is the gated output, which then goes through the usual output matrix WoW_o.

Here is that equation as it appears in the paper:

Equation 5 from the Gated Attention paper: Y prime equals g of Y, X, W theta, sigma, which equals Y elementwise times sigma of X W theta
Equation (5) of Qiu et al., 2025, "Gated Attention for Large Language Models".
xattention(softmax)Y×σ(x Wθ)× WooutputGated attention: each output number is multiplied by a gate between 0 and 1, computed from the token x.A gate near 0 lets a head say "nothing useful here" without dumping attention on the first token.
Gated attention. The attention output Y is multiplied, number by number, by a gate between 0 and 1 computed from the input x. Only then does it go through the output matrix.
arXiv page of Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free, with authors and abstract
The paper on arXiv (2505.06708). They tested 30 variants on models up to 15B parameters trained on 3.5 trillion tokens.

Why does it help with sinks? With a gate, a head can say "nothing here" by setting its gate near 0. It no longer needs the trick of parking its attention on the first token.

It is in a real model. Qwen3-Next's attention layers do exactly this. From its code in the Hugging Face transformers library, the query projection produces both the query and the gate, and the gate is applied right after attention:

python
query_states, gate = torch.chunk(
    self.q_proj(hidden_states).view(*input_shape, -1, self.head_dim * 2), 2, dim=-1
)
...
attn_output = attn_output.reshape(*input_shape, -1).contiguous()
attn_output = attn_output * torch.sigmoid(gate)

attn_output = self.o_proj(attn_output)

We will measure the sink effect ourselves in section 10.

2. Linear attention: remove softmax, get a fixed-size memory

Now the big idea. Look again at the attention formula without the mask:

output=softmax⁡ ⁣(QK⊤)V\text{output} = \operatorname{softmax}\!\left(QK^\top\right) V

Softmax sits between QK⊤QK^\top and VV, so we must first build the full T×TT \times T grid of scores. What if there were no softmax?

Without softmax, (QK⊤)V=Q(K⊤V)(QK^\top)V = Q(K^\top V). And K⊤VK^\top V is only d×dd \times d, whatever the length TT:

softmax: (Q Kᵀ) firstlinear: (Kᵀ V) firstQT × d×Kᵀd × T=T × Tgrows with T²Kᵀd × T×VT × d=d × dfixed sizeSame numbers, different order. With softmax in the middle you must build the T × T grid;without softmax you can multiply Kᵀ V first and only ever keep a small d × d matrix.
Same numbers, different order. Left: softmax attention must build the T × T grid, which grows with the square of the text length. Right: without softmax, Kᵀ V is computed first, and it is a small d × d matrix whose size never changes.

But we cannot just delete softmax: softmax made all weights positive, and that matters. Katharopoulos et al. (2020) replaced it with a feature map.

Linear attention for token tt is then:

ot=∑i≤t(ϕ(qt)⋅ϕ(ki)) vi∑i≤tϕ(qt)⋅ϕ(ki)o_t = \frac{\sum_{i \le t} \big(\phi(q_t) \cdot \phi(k_i)\big)\, v_i}{\sum_{i \le t} \phi(q_t) \cdot \phi(k_i)}

where:

  • qtq_t is the current token's query, and ki,vik_i, v_i are the key and value of an earlier token ii;
  • ϕ(qt)⋅ϕ(ki)\phi(q_t) \cdot \phi(k_i) is the (always positive) score, replacing eqt⋅kie^{q_t \cdot k_i};
  • the bottom line divides by the total score, so the weights still add up to 1, just like softmax did.

Now move the brackets. Because ϕ(qt)\phi(q_t) is the same for every ii, it can come out of the sum:

ot=ϕ(qt)⊤Stϕ(qt)⊤zt,St=∑i≤tϕ(ki) vi⊤,zt=∑i≤tϕ(ki)o_t = \frac{\phi(q_t)^\top S_t}{\phi(q_t)^\top z_t}, \qquad S_t = \sum_{i \le t} \phi(k_i)\, v_i^\top, \qquad z_t = \sum_{i \le t} \phi(k_i)

where:

  • StS_t is a dk×dvd_k \times d_v matrix: the memory. It is the sum of every token's key-value pair;
  • ztz_t is a vector of length dkd_k (the running total used for dividing);
  • ϕ(ki) vi⊤\phi(k_i)\, v_i^\top is an outer product.

This is exactly what the linear transformer paper writes (its Equation 5 is the version without the mask; Equation 9 below adds it):

The paper also gives the cost in multiplications (Section 3.2.1). With NN tokens, keys and queries of length DD and values of length MM:

softmax attention: O(N2max⁡(D,M))linear attention: O(N D M)\text{softmax attention: } O\big(N^2 \max(D, M)\big) \qquad\qquad \text{linear attention: } O\big(N\,D\,M\big)

where N2N^2 comes from comparing every pair of tokens, and D MD\,M is the size of the memory matrix that each token updates. When the text is longer than the head size (N>DN > D), linear attention does less work, and the gap grows with NN.

For text written left to right, the sums only run over earlier tokens, which the paper writes like this:

And here is the punchline: StS_t can be updated one token at a time:

St=St−1+ϕ(kt) vt⊤,zt=zt−1+ϕ(kt)S_t = S_{t-1} + \phi(k_t)\, v_t^\top, \qquad z_t = z_{t-1} + \phi(k_t)

That is an RNN: a fixed-size state, updated once per token. Each step costs the same, no matter how long the text is, and there is no KV cache that grows.

Softmax attention: a list that keeps growingLinear attention / Gated DeltaNet: one fixed-size memoryk1 v1k2 v2k3 v3k4 v4k5 v5k6 v6k7 v7k8 v8k9 v9… + one more per tokenEach new token reads ALL stored keys and values. Memory and work grow with the text.Nothing is ever forgotten, so recall is exact, but at 1M tokens the list is huge.memory Sd × d numberswrite: add orcorrect one factmemory Ssame sizeEvery token does one smallupdate. The size never changes,so old facts must share space.
Two kinds of memory. Softmax attention keeps every key and value in a list that grows forever. Linear attention (and everything after it in this part) keeps one fixed-size matrix and updates it once per token.
arXiv page of Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
The 2020 paper that started linear attention for transformers (arXiv 2006.16236).

Proof: the two forms give the same answer

I wrote both forms in part4_linear.py: the parallel one (build the masked T×TT \times T grid, like attention) and the recurrent one (one token at a time, keeping only SS and zz):

python
def linear_attention_parallel(q, k, v):
    """Like causal attention, but exp(q.k) is replaced by phi(q).phi(k), and there is no softmax."""
    Q, K = phi(q), phi(k)
    scores = (Q @ K.T).tril()                                  # (T, T), future set to 0
    return (scores @ v) / scores.sum(-1, keepdim=True)


def linear_attention_recurrent(q, k, v):
    """The same thing as an RNN: a fixed-size state S (d_k x d_v) and a normaliser z (d_k)."""
    S = torch.zeros(k.shape[-1], v.shape[-1], dtype=q.dtype)
    z = torch.zeros(k.shape[-1], dtype=q.dtype)
    out = []
    for t in range(q.shape[0]):
        S = S + torch.outer(phi(k[t]), v[t])                   # write: add this token's key-value pair
        z = z + phi(k[t])
        out.append((phi(q[t]) @ S) / (phi(q[t]) @ z))          # read: one small matrix-vector product
    return torch.stack(out)
plain text
1. Linear attention, parallel form vs recurrent form (T = 200): max |diff| = 4.4e-16
   the recurrent form keeps only 272 numbers, however long the text

Identical (the difference is rounding noise). And the recurrent form kept just 16×16+16=27216 \times 16 + 16 = 272 numbers for 200 tokens. For 200,000 tokens it would still be 272.

3. The catch: a fixed-size memory gets confused

A memory that never grows sounds too good to be true. It is. Softmax attention keeps every key and value separately, so it can always find an old fact exactly. Linear attention adds everything into one matrix, and things start to blur.

To see why, write a few facts into SS (I drop ϕ\phi and the divisor here to keep it simple) and read one back with its own key kik_i:

S ki  =  ∑jvj (kj⋅ki)  =  vi (ki⋅ki)⏟the fact you wanted  +  ∑j≠ivj (kj⋅ki)⏟noise from every other factS\,k_i \;=\; \sum_j v_j\,(k_j \cdot k_i) \;=\; \underbrace{v_i\,(k_i \cdot k_i)}_{\text{the fact you wanted}} \;+\; \underbrace{\sum_{j \ne i} v_j\,(k_j \cdot k_i)}_{\text{noise from every other fact}}

where S=∑jvjkj⊤S = \sum_j v_j k_j^\top is the memory (in the "value × key" layout used by the Gated DeltaNet paper), and kj⋅kik_j \cdot k_i is how much key jj overlaps with key ii.

There is a second problem: linear attention can only add. It has no way to change a fact.

Test 1: overwriting. I wrote x = 1, then y = 7, then x = 5 (each fact stored as an outer product of a key and a value, with perpendicular keys for x and y), and read x and y back:

plain text
2. Write "x = 1", "y = 7", then "x = 5". Read x and y back:
   linear attention: x = 6.00, y = 7.00
   delta rule:       x = 5.00, y = 7.00
Write "x = 1", "y = 7", then "x = 5". What comes back?read x, linear attention: 6.00read x, linear attention6.00read x, delta rule: 5.00read x, delta rule5.00read y, linear attention: 7.00read y, linear attention7.00read y, delta rule: 7.00read y, delta rule7.00correct x = 5Linear attention can only ADD, so x becomes 1 + 5 = 6. The delta rule replaces the old value: x = 5.
Overwriting a fact. Linear attention can only add, so x comes back as 1 + 5 = 6. The delta rule (next section) replaces the old value and returns 5.

Linear attention says x = 6: it added the new value on top of the old one.

Test 2: capacity. I stored nn random facts (random unit-length keys, random values) in a 64×6464 \times 64 memory, then read all of them back and measured the error (0 = perfect, 1 = as wrong as the value itself):

plain text
3. Store n random facts in a 64 x 64 memory, then read all of them back (relative error, 0 = perfect):
   n =   4: linear attention 0.21   delta rule 0.10
   n =   8: linear attention 0.32   delta rule 0.20
   n =  16: linear attention 0.49   delta rule 0.33
   n =  32: linear attention 0.70   delta rule 0.50
   n =  48: linear attention 0.86   delta rule 0.63
   n =  64: linear attention 1.00   delta rule 0.73
   n =  96: linear attention 1.22   delta rule 0.89
   n = 128: linear attention 1.41   delta rule 0.99
   n = 192: linear attention 1.74   delta rule 1.12
   n = 256: linear attention 2.01   delta rule 1.20
0.000.501.001.502.0041664256linear attention, 4: 0.21linear attention, 8: 0.32linear attention, 16: 0.49linear attention, 32: 0.70linear attention, 48: 0.86linear attention, 64: 1.00linear attention, 96: 1.22linear attention, 128: 1.41linear attention, 192: 1.74linear attention, 256: 2.01delta rule, 4: 0.10delta rule, 8: 0.20delta rule, 16: 0.33delta rule, 32: 0.50delta rule, 48: 0.63delta rule, 64: 0.73delta rule, 96: 0.89delta rule, 128: 0.99delta rule, 192: 1.12delta rule, 256: 1.20linear attention: 2.01delta rule: 1.20facts stored in a 64 × 64 memory (log scale)read-back error (0 = perfect)
Reading back n random facts from a 64 × 64 memory. The error grows as the memory fills up. The delta rule is always better than plain adding, but neither is perfect: a fixed-size memory has a fixed capacity.

Two lessons. Plain adding gets worse with every fact (by 64 facts, the error is as big as the signal). And even the better rule cannot escape the limit: a fixed-size memory has a fixed capacity. Keep this in mind; it is why hybrid models exist (section 9).

A 2023 study found that this is exactly where cheaper models lose to attention on real text:

4. The delta rule: write by correcting

The fix for overwriting is old: the delta rule (Widrow and Hoff, 1960), brought to transformers as "DeltaNet" (Schlag et al., 2021; Yang et al., 2024). The idea in plain words:

  1. Before writing, look up what the memory currently says for this key: St−1ktS_{t-1} k_t.
  2. Compare it with the value you want to store, vtv_t. The difference is the error ("delta").
  3. Change the memory by just that error, scaled by a writing strength βt\beta_t.

As an equation (memory in "value × key" layout, keys of length 1):

St  =  St−1+βt(vt−St−1kt) kt⊤  =  St−1(I−βt ktkt⊤)+βt vtkt⊤S_t \;=\; S_{t-1} + \beta_t \big(v_t - S_{t-1} k_t\big)\, k_t^\top \;=\; S_{t-1}\big(I - \beta_t\, k_t k_t^\top\big) + \beta_t\, v_t k_t^\top

where:

  • St−1ktS_{t-1} k_t is the old value the memory holds for key ktk_t;
  • vt−St−1ktv_t - S_{t-1} k_t is the error: what we want minus what is there;
  • βt\beta_t, between 0 and 1, is the writing strength (1 = replace completely, 0.5 = move halfway);
  • II is the identity matrix (the matrix that changes nothing); I−βtktkt⊤I - \beta_t k_t k_t^\top removes part of whatever was stored along the direction ktk_t;
  • βt vtkt⊤\beta_t\, v_t k_t^\top then writes the new value there.

With βt=1\beta_t = 1 and a unit-length key, reading right after writing gives exactly Stkt=vtS_t k_t = v_t: the old value is gone. That is why the test above returned x = 5.

Here is the beautiful part: the delta rule is one step of gradient descent, taken at every token, on the error "how wrong is my memory for this key?". The Kimi Linear paper writes it this way (in their layout, keys and values swap sides):

Equation from the Kimi Linear paper: S t equals S t minus 1 minus beta t times the gradient of L t, which equals I minus beta t k t k t transposed, times S t minus 1, plus beta t k t v t transposed
The delta rule as one gradient-descent step, from the Kimi Linear paper (Section 2). The loss being reduced is ½‖Sᵀk − v‖².
Lt(S)=12 ∥S kt−vt∥2\mathcal{L}_t(S) = \tfrac{1}{2}\,\big\lVert S\,k_t - v_t \big\rVert^2

where Lt\mathcal{L}_t is the squared error between what the memory returns for key ktk_t and the value vtv_t it should return, and βt\beta_t plays the role of the step size. So the memory is literally learning while it reads, one token at a time. This is why this family is sometimes called "test-time training".

In code it is one line:

python
def write_delta(S, k, v, beta=1.0):
    return S - beta * torch.outer(S @ k - v, k)                # S (I - beta k k^T) + beta v k^T

5. Forget gates: clearing out old memories

The delta rule fixes one fact at a time. But sometimes the model needs to forget a lot at once: a new paragraph, a new topic, a new document. For that, models add a forget gate.

Mamba-2 (2024) can be written as linear attention with exactly this gate. As written in the Gated DeltaNet paper:

Equation from the Gated DeltaNet paper: S t equals alpha t times S t minus 1 plus v t k t transposed, and o t equals S t q t
Mamba-2 as a gated linear attention, from the Gated DeltaNet paper (Section 2).
St=αt St−1+vtkt⊤,ot=St qtS_t = \alpha_t\, S_{t-1} + v_t k_t^\top, \qquad o_t = S_t\, q_t

where αt\alpha_t is the forget gate (computed from token tt), vtkt⊤v_t k_t^\top writes the new fact, and ot=Stqto_t = S_t q_t reads the memory with the query.

How fast do facts fade? It depends strongly on α\alpha:

plain text
4. Strength of a fact after t more tokens, with forget gate alpha (alpha^t):
   alpha = 0.9  : t=0: 1.000  t=10: 0.349  t=50: 0.005  t=100: 0.000  t=500: 0.000
   alpha = 0.99 : t=0: 1.000  t=10: 0.904  t=50: 0.605  t=100: 0.366  t=500: 0.007
   alpha = 0.999: t=0: 1.000  t=10: 0.990  t=50: 0.951  t=100: 0.905  t=500: 0.606
0.000.250.500.751.0010100250500α = 0.9, 10: 0.35α = 0.9, 50: 0.01α = 0.9, 100: 0.00α = 0.9, 500: 0.00α = 0.99, 10: 0.90α = 0.99, 50: 0.61α = 0.99, 100: 0.37α = 0.99, 500: 0.01α = 0.999, 10: 0.99α = 0.999, 50: 0.95α = 0.999, 100: 0.90α = 0.999, 500: 0.61α = 0.999: 0.61α = 0.99: 0.01α = 0.9: 0.00tokens since the fact was writtenstrength left (α^t)
How much of a stored fact is left after t more tokens. With α = 0.9 a fact is almost gone after 50 tokens; with α = 0.999 more than half survives 500 tokens. Because α is computed from each token, the model can choose: remember (α near 1) or wipe (α near 0).

The important word is data-dependent: αt\alpha_t is computed from token tt itself. So the model can keep α\alpha near 1 inside a paragraph and drop it towards 0 at a topic change.

6. Gated DeltaNet: correct AND forget

Gating is good at forgetting everything a bit. The delta rule is good at changing one fact exactly. Gated DeltaNet (Yang, Kautz and Hatamizadeh, ICLR 2025) uses both:

arXiv page of Gated Delta Networks: Improving Mamba2 with Delta Rule, by Songlin Yang, Jan Kautz and Ali Hatamizadeh, with the abstract
The Gated DeltaNet paper on arXiv (2412.06464).

The gated delta rule, as printed in the paper:

St=St−1(αt(I−βt ktkt⊤))+βt vtkt⊤,ot=St qtS_t = S_{t-1}\Big(\alpha_t\big(I - \beta_t\, k_t k_t^\top\big)\Big) + \beta_t\, v_t k_t^\top, \qquad o_t = S_t\, q_t

where:

  • StS_t is the memory, a dv×dkd_v \times d_k matrix (one per head);
  • αt∈(0,1)\alpha_t \in (0, 1) is the forget gate: shrink the whole memory a little (or a lot);
  • βt∈(0,1)\beta_t \in (0, 1) is the writing strength of the delta rule;
  • I−βtktkt⊤I - \beta_t k_t k_t^\top removes (part of) the old value stored at key ktk_t;
  • βt vtkt⊤\beta_t\, v_t k_t^\top writes the new value at key ktk_t;
  • qtq_t is the query; ot=Stqto_t = S_t q_t reads the memory;
  • if αt=1\alpha_t = 1 it becomes the plain delta rule; if βt=0\beta_t = 0 it only forgets.
One Gated DeltaNet step for token t1. forgetS ← αₜ · Sshrink every old fact2. look upold = S kₜwhat is stored at kₜ now?3. correctS ← S + βₜ (vₜ − old) kₜᵀmove it toward vₜ4. readoₜ = S qₜanswer the queryαₜ (between 0 and 1): how much of the old memory to keep. Near 1 = remember, near 0 = wipe.βₜ (between 0 and 1): how strongly to write. 1 = fully replace what was stored at kₜ.Both are computed from the token itself, so the model decides, token by token, what to forget and what to write.
One Gated DeltaNet step, in four moves: forget a little, look up what is stored at the key, correct it towards the new value, and read the answer with the query.

Every rule as "learning while reading"

There is a neat way to see all these memories as one family. At each token, the new memory StS_t is the answer to a tiny optimisation problem: stay close to the old memory, but fit the new fact.

For plain linear attention:

St=arg⁡min⁡S  ∥S−St−1∥F2⏟stay close  −  2 ⟨S kt,  vt⟩⏟point toward vt⟹St=St−1+vtkt⊤S_t = \arg\min_S \;\underbrace{\lVert S - S_{t-1} \rVert_F^2}_{\text{stay close}} \;-\; 2\,\underbrace{\langle S\,k_t,\; v_t \rangle}_{\text{point toward } v_t} \quad\Longrightarrow\quad S_t = S_{t-1} + v_t k_t^\top

For the delta rule, the second term instead points toward the error vt−St−1ktv_t - S_{t-1}k_t, scaled by βt\beta_t:

St=arg⁡min⁡S  ∥S−St−1∥F2  −  2 ⟨S kt,  βt(vt−St−1kt)⟩⟹St=St−1(I−βtktkt⊤)+βtvtkt⊤S_t = \arg\min_S \;\lVert S - S_{t-1} \rVert_F^2 \;-\; 2\,\big\langle S\,k_t,\; \beta_t (v_t - S_{t-1} k_t) \big\rangle \quad\Longrightarrow\quad S_t = S_{t-1}(I - \beta_t k_t k_t^\top) + \beta_t v_t k_t^\top

where:

  • arg⁡min⁡S\arg\min_S means "the SS that makes this smallest";
  • ∥A∥F2\lVert A \rVert_F^2 (the squared Frobenius norm) is the sum of the squares of all entries of a matrix: here, how much the memory changed;
  • ⟨a,b⟩\langle a, b \rangle is the dot product of two vectors;
  • setting the slope (gradient) to zero gives the update on the right. For the first one: 2(S−St−1)−2 vtkt⊤=02(S - S_{t-1}) - 2\,v_t k_t^\top = 0, so S=St−1+vtkt⊤S = S_{t-1} + v_t k_t^\top.

The Gated DeltaNet paper lists the whole family in one table:

Where α and β come from

Both are computed from the token by small learned layers, so the model decides token by token. In Qwen3-Next's code:

python
beta = b.sigmoid()
# If the model is loaded in fp16, without the .float() here, A might be -inf
g = -self.A_log.float().exp() * F.softplus(a.float() + self.dt_bias)

where:

  • b and a are produced from the token by a learned matrix;
  • beta = sigmoid(b) keeps βt\beta_t between 0 and 1;
  • g is log⁡αt\log \alpha_t, the forget gate stored as a logarithm, and αt=eg\alpha_t = e^{g};
  • softplus(x) =log⁡(1+ex)= \log(1 + e^x) is always positive and A_log.exp() is positive, so g is always negative and αt=eg\alpha_t = e^g is always between 0 and 1.

A real Gated DeltaNet layer adds a few more parts around this core: queries and keys are normalised to length 1 (so ktkt⊤k_t k_t^\top behaves well), a short convolution mixes each token with its 3 neighbours before the memory step, and the output passes through a norm and an output gate. The Gated DeltaNet paper describes the query and key path as "linear proj., shortconv., SiLU and L2 norm".

Proof: my version matches Qwen3-Next's own code

I wrote the gated delta rule from scratch, line by line from the equation:

python
def gated_delta_rule(q, k, v, alpha, beta):
    """S_t = alpha_t * S_{t-1} (I - beta_t k_t k_t^T) + beta_t v_t k_t^T ;  o_t = S_t q_t.
    q, k: (H, T, d_k) L2-normalised; v: (H, T, d_v); alpha, beta: (H, T) in (0, 1)."""
    H, T, dk = k.shape
    S = torch.zeros(H, v.shape[-1], dk, dtype=q.dtype)
    out = []
    for t in range(T):
        kt, vt = k[:, t], v[:, t]
        S = alpha[:, t, None, None] * S                                         # 1. forget a little (gate)
        S = S - beta[:, t, None, None] * torch.einsum('hv,hk->hvk', torch.einsum('hvk,hk->hv', S, kt) - vt, kt)   # 2. correct (delta)
        out.append(torch.einsum('hvk,hk->hv', S, q[:, t]))                       # 3. read
    return torch.stack(out, 1)

The transformers library ships Qwen3-Next's reference implementation in two versions: a step-by-step one (used when generating) and a chunked one (used for long inputs: it processes 64 tokens at a time with clever algebra, giving the same result). I compared mine with both, on 4 heads and 256 tokens of random inputs:

plain text
5. Our gated delta rule vs the Qwen3-Next reference code in transformers (float32 inside):
   vs step-by-step version: max |diff| = 3.0e-08
   vs chunked version:      max |diff| = 2.3e-07   (outputs up to 0.17)

Differences around 10−710^{-7} are the rounding noise of 32-bit numbers, which their code uses internally. So the equation above really is what runs inside Qwen3-Next.

7. Kimi Delta Attention: a forget rate for every channel

Kimi Linear (Moonshot AI, 2025) makes one more change. In Gated DeltaNet, αt\alpha_t is one number per head: the whole memory of a head fades at the same speed. Kimi Delta Attention (KDA) gives every channel its own forget rate.

Equation 1 from the Kimi Linear paper: S t equals I minus beta t k t k t transposed, times Diag of alpha t, times S t minus 1, plus beta t k t v t transposed; o t equals S t transposed q t
Equation (1) of the Kimi Linear paper: Kimi Delta Attention. (Their memory is stored as d_k × d_v, the transpose of the Gated DeltaNet layout, so keys and values swap sides.)
St=(I−βt ktkt⊤) Diag⁡(αt) St−1+βt ktvt⊤,ot=St⊤qtS_t = \big(I - \beta_t\, k_t k_t^\top\big)\, \operatorname{Diag}(\alpha_t)\, S_{t-1} + \beta_t\, k_t v_t^\top, \qquad o_t = S_t^\top q_t

where everything is as in Gated DeltaNet, except that αt\alpha_t is now a vector (one forget rate per key channel) instead of a single number.

Why would that help? Some channels can hold long-lived facts (α near 1) while others act as scratch space that is quickly cleared (α small), all inside one head.

Written out entry by entry, the forget step of KDA is

(Diag⁡(αt) St−1)ij=αt,i⋅(St−1)ij\big(\operatorname{Diag}(\alpha_t)\, S_{t-1}\big)_{ij} = \alpha_{t,i} \cdot (S_{t-1})_{ij}

where ii indexes the key channels (rows of KDA's dk×dvd_k \times d_v memory) and jj the value channels. Every entry in row ii fades at its own rate αt,i\alpha_{t,i}. In Gated DeltaNet the same step is αt⋅(St−1)ij\alpha_t \cdot (S_{t-1})_{ij}: one rate for the whole head.

arXiv page of Kimi Linear: An Expressive, Efficient Attention Architecture, with the abstract
The Kimi Linear paper on arXiv (2510.26692).

8. Speed and memory: the payoff

Each new token costs a linear layer one memory update, the same size every time. Softmax attention instead reads the whole cache. I timed one decode step of one layer, using Qwen3-Next's real layer shapes (gated attention: 16 query heads sharing 2 key/value heads of size 256; Gated DeltaNet: 32 heads with a 128×128128 \times 128 memory each), on an Apple M5 Pro GPU:

plain text
7. One decode step on mps (Qwen3-Next layer shapes):
     4096 tokens so far: gated attention 0.097 ms   Gated DeltaNet 0.180 ms
    16384 tokens so far: gated attention 0.284 ms   Gated DeltaNet 0.181 ms
    65536 tokens so far: gated attention 1.238 ms   Gated DeltaNet 0.180 ms
   262144 tokens so far: gated attention 5.831 ms   Gated DeltaNet 0.180 ms
0.002.004.006.004,09616,38465,536262,144gated attention, 4,096: 0.10gated attention, 16,384: 0.28gated attention, 65,536: 1.24gated attention, 262,144: 5.83Gated DeltaNet, 4,096: 0.18Gated DeltaNet, 16,384: 0.18Gated DeltaNet, 65,536: 0.18Gated DeltaNet, 262,144: 0.18gated attention: 5.83Gated DeltaNet: 0.18tokens so far (log scale)milliseconds per step
Time for one decode step of one layer. Attention grows with the conversation; the Gated DeltaNet step stays flat at 0.18 ms. Below about 10,000 tokens attention is actually faster in this test; beyond that, the fixed-size memory wins by more and more.

Note the honest detail: at 4,096 tokens, attention was faster (0.097 vs 0.180 ms). My Gated DeltaNet step is several small unfused operations, and a 32×128×12832 \times 128 \times 128 memory is not tiny. The advantage appears only once the conversation is long, and then it keeps growing: 32× faster at 262K tokens.

9. Hybrid models: a few attention layers keep recall sharp

Section 3 showed the weakness: a fixed-size memory has fixed capacity, so it cannot recall every detail of a long text exactly. Softmax attention can. The solution all three of today's big designs use: mostly linear layers, plus a few full attention layers.

Qwen3-Next-80B (48 layers)layer 1: linear (Gated DeltaNet / KDA)layer 2: linear (Gated DeltaNet / KDA)layer 3: linear (Gated DeltaNet / KDA)layer 4: full attentionlayer 5: linear (Gated DeltaNet / KDA)layer 6: linear (Gated DeltaNet / KDA)layer 7: linear (Gated DeltaNet / KDA)layer 8: full attentionlayer 9: linear (Gated DeltaNet / KDA)layer 10: linear (Gated DeltaNet / KDA)layer 11: linear (Gated DeltaNet / KDA)layer 12: full attentionlayer 13: linear (Gated DeltaNet / KDA)layer 14: linear (Gated DeltaNet / KDA)layer 15: linear (Gated DeltaNet / KDA)layer 16: full attentionlayer 17: linear (Gated DeltaNet / KDA)layer 18: linear (Gated DeltaNet / KDA)layer 19: linear (Gated DeltaNet / KDA)layer 20: full attentionlayer 21: linear (Gated DeltaNet / KDA)layer 22: linear (Gated DeltaNet / KDA)layer 23: linear (Gated DeltaNet / KDA)layer 24: full attentionlayer 25: linear (Gated DeltaNet / KDA)layer 26: linear (Gated DeltaNet / KDA)layer 27: linear (Gated DeltaNet / KDA)layer 28: full attentionlayer 29: linear (Gated DeltaNet / KDA)layer 30: linear (Gated DeltaNet / KDA)layer 31: linear (Gated DeltaNet / KDA)layer 32: full attentionlayer 33: linear (Gated DeltaNet / KDA)layer 34: linear (Gated DeltaNet / KDA)layer 35: linear (Gated DeltaNet / KDA)layer 36: full attentionlayer 37: linear (Gated DeltaNet / KDA)layer 38: linear (Gated DeltaNet / KDA)layer 39: linear (Gated DeltaNet / KDA)layer 40: full attentionlayer 41: linear (Gated DeltaNet / KDA)layer 42: linear (Gated DeltaNet / KDA)layer 43: linear (Gated DeltaNet / KDA)layer 44: full attentionlayer 45: linear (Gated DeltaNet / KDA)layer 46: linear (Gated DeltaNet / KDA)layer 47: linear (Gated DeltaNet / KDA)layer 48: full attentionKimi Linear 48B (27 layers)layer 1: linear (Gated DeltaNet / KDA)layer 2: linear (Gated DeltaNet / KDA)layer 3: linear (Gated DeltaNet / KDA)layer 4: full attentionlayer 5: linear (Gated DeltaNet / KDA)layer 6: linear (Gated DeltaNet / KDA)layer 7: linear (Gated DeltaNet / KDA)layer 8: full attentionlayer 9: linear (Gated DeltaNet / KDA)layer 10: linear (Gated DeltaNet / KDA)layer 11: linear (Gated DeltaNet / KDA)layer 12: full attentionlayer 13: linear (Gated DeltaNet / KDA)layer 14: linear (Gated DeltaNet / KDA)layer 15: linear (Gated DeltaNet / KDA)layer 16: full attentionlayer 17: linear (Gated DeltaNet / KDA)layer 18: linear (Gated DeltaNet / KDA)layer 19: linear (Gated DeltaNet / KDA)layer 20: full attentionlayer 21: linear (Gated DeltaNet / KDA)layer 22: linear (Gated DeltaNet / KDA)layer 23: linear (Gated DeltaNet / KDA)layer 24: full attentionlayer 25: linear (Gated DeltaNet / KDA)layer 26: linear (Gated DeltaNet / KDA)layer 27: full attentionNemotron-H-8B (52 layers)layer 1: Mamba-2layer 2: MLP only (no token mixing)layer 3: Mamba-2layer 4: MLP only (no token mixing)layer 5: Mamba-2layer 6: MLP only (no token mixing)layer 7: Mamba-2layer 8: full attentionlayer 9: MLP only (no token mixing)layer 10: Mamba-2layer 11: MLP only (no token mixing)layer 12: Mamba-2layer 13: MLP only (no token mixing)layer 14: Mamba-2layer 15: MLP only (no token mixing)layer 16: Mamba-2layer 17: MLP only (no token mixing)layer 18: Mamba-2layer 19: full attentionlayer 20: MLP only (no token mixing)layer 21: Mamba-2layer 22: MLP only (no token mixing)layer 23: Mamba-2layer 24: MLP only (no token mixing)layer 25: Mamba-2layer 26: MLP only (no token mixing)layer 27: Mamba-2layer 28: MLP only (no token mixing)layer 29: Mamba-2layer 30: full attentionlayer 31: MLP only (no token mixing)layer 32: Mamba-2layer 33: MLP only (no token mixing)layer 34: Mamba-2layer 35: MLP only (no token mixing)layer 36: Mamba-2layer 37: MLP only (no token mixing)layer 38: Mamba-2layer 39: MLP only (no token mixing)layer 40: Mamba-2layer 41: full attentionlayer 42: MLP only (no token mixing)layer 43: Mamba-2layer 44: MLP only (no token mixing)layer 45: Mamba-2layer 46: MLP only (no token mixing)layer 47: Mamba-2layer 48: MLP only (no token mixing)layer 49: Mamba-2layer 50: MLP only (no token mixing)layer 51: Mamba-2layer 52: MLP only (no token mixing)linear: Gated DeltaNet or KDAMamba-2full attentionMLP only
Layer patterns from the published configurations. Qwen3-Next: 3 Gated DeltaNet layers, then 1 gated attention layer, repeated (36 + 12). Kimi Linear: 3 KDA layers, then 1 MLA layer (20 + 7). Nemotron-H-8B: 24 Mamba-2 layers, 24 MLP-only layers, and just 4 attention layers. Hover over a layer to see its type.

Qwen3-Next-80B (Qwen, 2025). From its configuration on Hugging Face:

json
{
  "num_hidden_layers": 48,
  "full_attention_interval": 4,
  "num_attention_heads": 16,
  "num_key_value_heads": 2,
  "head_dim": 256,
  "linear_num_key_heads": 16,
  "linear_num_value_heads": 32,
  "linear_key_head_dim": 128,
  "linear_value_head_dim": 128,
  "linear_conv_kernel_dim": 4,
  "max_position_embeddings": 262144
}

full_attention_interval: 4 means every 4th layer is (gated) full attention: 12 of 48. The other 36 are Gated DeltaNet, each with 32 heads holding a 128×128128 \times 128 memory, and a short convolution of 4 tokens.

Kimi Linear 48B (Moonshot AI, 2025): 27 layers; layers 4, 8, 12, 16, 20, 24 and 27 are MLA (the compressed attention from Part 2), the other 20 are KDA.

Nemotron-H (NVIDIA, 2025) uses Mamba-2 (the gated linear attention from section 5) instead of a delta rule:

How much memory does the hybrid save?

For attention layers the memory grows with the text; for linear layers it is fixed. With 2 bytes per number:

memory=Lattn×2×Hkv×dh×n×2 B⏟attention layers: grows with n  +  Llin×H×dk×dv×2 B⏟linear layers: fixed\text{memory} = \underbrace{L_{\text{attn}} \times 2 \times H_{kv} \times d_h \times n \times 2\,\text{B}}_{\text{attention layers: grows with } n} \;+\; \underbrace{L_{\text{lin}} \times H \times d_k \times d_v \times 2\,\text{B}}_{\text{linear layers: fixed}}

where nn is the number of tokens, LattnL_{\text{attn}} and LlinL_{\text{lin}} count the two kinds of layers, and the other symbols are the head counts and sizes from the configuration.

plain text
6. Memory, 2 bytes per number
   Qwen3-Next-80B at 262144 tokens: 12 attention layers 6.00 GiB (all 48 as attention: 24.00 GiB), plus a fixed 36 MiB of Gated DeltaNet state
   Kimi Linear 48B: 8064 B per token with 7 MLA layers vs 31104 B if all 27 were MLA (74% less), plus a fixed 20 MiB of KDA state
0.006.0012.0018.0024.004,09616,38465,536262,144all 48 layers attention, 4,096: 0.38all 48 layers attention, 16,384: 1.50all 48 layers attention, 65,536: 6.00all 48 layers attention, 262,144: 24.00real hybrid (12 attention), 4,096: 0.13real hybrid (12 attention), 16,384: 0.41real hybrid (12 attention), 65,536: 1.54real hybrid (12 attention), 262,144: 6.04all 48 layers attention: 24.00real hybrid (12 attention): 6.04tokens in the conversation (log scale)memory, GiB
Memory for one conversation in Qwen3-Next-80B: if all 48 layers were attention (orange) versus the real hybrid (blue). At 262K tokens: 24 GiB versus about 6 GiB.
  • Qwen3-Next at 262K tokens: 24 GiB if every layer were attention, 6 GiB as built, plus a fixed 36 MiB for all the Gated DeltaNet memories together.
  • Kimi Linear: 7 MLA layers out of 27 means 74% less cache per token, matching the paper's "up to 75%".

10. Experiment: three small models, one difference

Equations and timings are one thing. Do these layers actually learn as well as attention? I trained three small models that are identical except for their token-mixing layers:

ModelIts 4 layers
softmax attention4 × softmax attention (with RoPE)
gated attention4 × softmax attention + output gate (section 1)
linear attention4 × linear attention, ϕ=elu⁡+1\phi = \operatorname{elu} + 1 (section 2)

Everything else is the same: 4 layers, width 256, 4 heads, the same feed-forward blocks, the same optimizer, the same random seed, the same data, the same number of steps. Each model has about 3.3 to 3.5 million parameters.

Task 1: predicting text. Each model reads public-domain books one byte at a time and learns to predict the next byte (1,500 steps of 32 × 256 bytes, from six books). It is then tested on a book it never saw, The Adventures of Tom Sawyer. The score is bits per byte: lower is better.

Task 2: recall. The model sees nn facts (a key and its value at each position), then nn questions (just a key, in a new order), and must answer each with the right value. Keys come from 1,024 possible tokens and values from 256. Each model trains for 3,000 steps on random sets of 32 to 256 facts, then is tested on fresh ones. This is a simplified, one-step version of the "multi-query associative recall" test from the Zoology paper (Arora et al., 2023), which found that recall explains most of the gap between attention and its cheaper rivals.

Result 1: predicting text

plain text
softmax  params 3.28M | books: 1.994 bits/byte (558 s)
gated    params 3.54M | books: 1.989 bits/byte (489 s)
linear   params 3.28M | books: 2.267 bits/byte (482 s)
softmax attention: 1.994 bits per bytesoftmax attention1.994gated attention: 1.989 bits per bytegated attention1.989linear attention: 2.267 bits per bytelinear attention2.267bits per byte on a held-out book (lower is better); bars start at 1.6
Bits per byte on a book none of the models saw during training (lower is better). Softmax and gated attention are almost tied; linear attention is clearly behind.
ModelBits per byte (lower is better)
softmax attention1.994
gated attention1.989
linear attention2.267
  • Gated attention is slightly better than plain softmax (1.989 vs 1.994). That is the direction the paper reports, but the gap here is tiny, and with one run per model I would not call it a real win at this scale.
  • Linear attention is clearly worse (2.267, about 14% more bits). Its fixed memory cannot keep exact track of the recent letters and words the way softmax attention can. This is the "fixed memory gets confused" problem from section 3, showing up in real training. It is exactly what the delta rule and gates were invented to fix.

Result 2: recall

plain text
softmax  recall accuracy  32 facts 1.000  64 facts 1.000  128 facts 1.000  256 facts 1.000
gated    recall accuracy  32 facts 0.159  64 facts 0.147  128 facts 0.122  256 facts 0.079
linear   recall accuracy  32 facts 0.998  64 facts 0.993  128 facts 0.956  256 facts 0.781
0.000.250.500.751.003264128256softmax attention, 32: 1.00softmax attention, 64: 1.00softmax attention, 128: 1.00softmax attention, 256: 1.00gated attention, 32: 0.16gated attention, 64: 0.15gated attention, 128: 0.12gated attention, 256: 0.08linear attention, 32: 1.00linear attention, 64: 0.99linear attention, 128: 0.96linear attention, 256: 0.78softmax attention: 1.00linear attention: 0.78gated attention: 0.08facts to remember (log scale)share of lookups answered correctly
Share of lookups answered correctly as the number of facts grows. Softmax attention stays at 100%. Linear attention starts near 100% but falls to 78% at 256 facts: its memory is full. Gated attention (this run) failed to learn the task in 3,000 steps; see the note below.
Facts to remember3264128256
softmax attention100%100%100%100%
linear attention99.8%99.3%95.6%78.1%
gated attention (seed 0)15.9%14.7%12.2%7.9%

Linear attention shows its capacity limit. With few facts it is nearly perfect, but as the facts pile up, its fixed memory (4 heads of 64×6464 \times 64 per layer) starts mixing them, exactly like the capacity test in section 3. Softmax attention keeps every fact separately and stays at 100%.

Result 3: attention sinks

The gated attention paper says the gate removes attention sinks. I measured how much attention the first token receives in the trained softmax and gated models (queries from position 64 on, averaged per layer):

plain text
softmax  first-token attention per softmax layer: 0.003 0.001 0.000 0.000
gated    first-token attention per softmax layer: 0.003 0.000 0.000 0.000

Both are near zero: these tiny models never formed an attention sink at all, so there was nothing for the gate to remove. Sinks are known to appear in larger models trained for much longer, and my byte-level texts start mid-sentence with no special first token. So this experiment cannot confirm or refute the paper's 46.7% → 4.8% result; it only shows that sinks are not automatic.

11. Proof of the runs

Both scripts, with their real outputs:

Terminal output of part4_linear.py: linear attention equivalence, overwrite test, capacity test, forget gate fading, comparison with Qwen3-Next reference code, memory and decode timing
Output of part4_linear.py on an Apple M5 Pro.
Terminal output of part4_train.py for the three small models: parameters, bits per byte, recall accuracy and first-token attention
Output of part4_train.py: the three models, each trained from scratch, plus the two extra recall checks.

The impact, and where you meet it

In 2024, linear attention was mostly a research topic. By late 2025, three major model families shipped it in most of their layers. The results the companies reported are what made the difference:

Use cases in one line each:

  • Very long documents and codebases (hundreds of thousands to millions of tokens): hybrids keep memory and per-token cost almost flat.
  • Reasoning models and agents that write very long outputs: each new token stays cheap, as in Kimi Linear's 6× faster decoding at 1M tokens.
  • High-volume serving: more requests per GPU (Nemotron-H's 2.4× throughput) means lower cost per answer.
  • Devices with little memory: a fixed-size state means a long conversation does not need a growing cache.
  • Where full attention still wins: exact lookups of details far back in the text. That is why every model above keeps some attention layers.

The whole series in one table

IdeaPartWhat it cutsUsed by
Multi-head attention1nothing (the baseline)every transformer
MQA / GQA2KV cache size: fewer key/value headsLlama 3, Qwen2.5, Gemma 3, Mistral
MLA2KV cache size: compressed latentDeepSeek-V2/V3, Kimi K2, Kimi Linear
Sliding window3tokens looked at, and cacheMistral 7B, Gemma 2/3, gpt-oss
DeepSeek Sparse Attention3tokens looked at (top-k)DeepSeek-V3.2
Gated attention4attention sinks (quality fix)Qwen3-Next
Linear attention, Mamba-24the cache itself: fixed-size memoryNemotron-H (Mamba-2)
Gated DeltaNet, KDA4the cache itself, with better recallQwen3-Next, Kimi Linear
Hybrids4most of the cache, keeping exact recallQwen3-Next, Kimi Linear, Nemotron-H

Summary

  • Gated attention multiplies each attention output by a sigmoid gate, Y′=Y⊙σ(XWθ)Y' = Y \odot \sigma(XW_\theta). The paper cut first-token attention from 46.7% to 4.8%; Qwen3-Next uses it.
  • Linear attention replaces softmax with ϕ(q)⋅ϕ(k)\phi(q) \cdot \phi(k). Moving the brackets turns attention into an RNN with a fixed memory St=St−1+ϕ(kt)vt⊤S_t = S_{t-1} + \phi(k_t) v_t^\top. The parallel and recurrent forms matched to 4.4×10−164.4 \times 10^{-16}.
  • A fixed memory gets confused: plain adding cannot overwrite (x came back as 6 instead of 5), and errors grow as facts pile up.
  • The delta rule writes by correcting: St=St−1+βt(vt−St−1kt)kt⊤S_t = S_{t-1} + \beta_t (v_t - S_{t-1}k_t) k_t^\top, one step of gradient descent per token. It overwrote correctly and stored random facts with less error.
  • Forget gates αt\alpha_t fade old memories (αt\alpha^t), chosen token by token.
  • Gated DeltaNet combines both: St=St−1 αt(I−βtktkt⊤)+βtvtkt⊤S_t = S_{t-1}\,\alpha_t(I - \beta_t k_t k_t^\top) + \beta_t v_t k_t^\top. My version matched Qwen3-Next's reference code to 3×10−83 \times 10^{-8}. KDA gives every channel its own forget rate.
  • One Gated DeltaNet decode step stayed at 0.18 ms at any length; attention grew to 5.8 ms at 262K tokens.
  • Hybrids keep a few full attention layers for exact recall: Qwen3-Next (36 + 12), Kimi Linear (20 + 7, 74% less cache), Nemotron-H (4 attention layers of 52).
  • In my three small models, linear attention was clearly worse at predicting text (2.267 vs 1.994 bits per byte) and its recall fell to 78% at 256 facts, while softmax attention stayed at 100%. Gated and plain softmax were too close to rank from single runs.

That is the end of the series. Attention started in 2014 as "let the decoder look back at every word"; in 2025 the best models look back with only a handful of layers, and remember everything else in memories that never grow.

Run it yourself
  • code/attention/part4_linear.py: linear attention forms, overwrite and capacity tests, forget-gate fading, the comparison with Qwen3-Next's reference code, memory and decode timing. Runs in about a minute.
  • code/attention/part4_train.py: the small models, trained on books and on the recall task. Pass model names to choose: python part4_train.py softmax gated linear (what this article reports). It can also train gdn and hybrid, but those are very slow without the special GPU kernels. SEED=1 changes the random start and RECALL_ONLY=1 skips the book task, as in the extra checks.
bash
pip install torch transformers
python part4_linear.py     # writes results/part4.json
python part4_train.py softmax gated linear   # writes results/part4_train.json

References

  1. A. Katharopoulos, A. Vyas, N. Pappas, F. Fleuret. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention. ICML 2020.
  2. I. Schlag, K. Irie, J. Schmidhuber. Linear Transformers Are Secretly Fast Weight Programmers. ICML 2021.
  3. S. Yang, B. Wang, Y. Zhang, Y. Shen, Y. Kim. Parallelizing Linear Transformers with the Delta Rule over Sequence Length. NeurIPS 2024.
  4. T. Dao, A. Gu. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality (Mamba-2). ICML 2024.
  5. S. Yang, J. Kautz, A. Hatamizadeh. Gated Delta Networks: Improving Mamba2 with Delta Rule. ICLR 2025.
  6. Z. Qiu et al. Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free. 2025.
  7. Kimi Team. Kimi Linear: An Expressive, Efficient Attention Architecture. 2025.
  8. NVIDIA. Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models. 2025.
  9. S. Arora et al. Zoology: Measuring and Improving Recall in Efficient Language Models. 2023.
  10. Model configurations: Qwen3-Next-80B-A3B-Instruct, Kimi-Linear-48B-A3B-Instruct, Nemotron-H-8B-Base-8K. Qwen3-Next code: Hugging Face transformers, models/qwen3_next/modeling_qwen3_next.py.
  11. Texts: Project Gutenberg books 1342, 1661, 11, 84, 2701, 98 for training and 74 for testing.