Emerging Frontiers · 7 min
Beyond Attention: State-Space Models
Attention's cost grows with the square of sequence length. State-space models like Mamba swap the quadratic attention matrix for linear-time recurrence — winning on long-sequence throughput but losing on exact recall, which is why the frontier runs hybrids.
Feed a Transformer a 100,000-token document and something ugly happens under the hood: it builds an attention matrix that compares every token to every other token. That's 100,000 × 100,000 entries, ten billion numbers, for a single layer at a single step. Double the input and the cost quadruples. This is the quadratic wall, and for a decade the field mostly paid the toll because attention worked so well. State-space models are a serious attempt to stop paying it.
The quadratic wall, concretely
Attention's power is also its bill. Every token attends to every prior token, so compute and memory during training scale as O(n²) in sequence length n. At inference, autoregressive decoding keeps a KV cache (the keys and values for every token seen so far) that grows linearly with the sequence and gets re-read on every new token. Generate token 40,000 and you're streaming 40,000 cached entries through memory just to produce one output. That cache is not free floor space: for a mid-size model it runs on the order of hundreds of kilobytes per token, so a long context can burn tens of gigabytes of GPU memory before the model emits a single word of reply. Long contexts don't just cost more. They cost disproportionately more, and they eat the exact resource you have least of.
A recurrent model doesn't have this problem. An old-school RNN carries a fixed-size hidden state h and updates it one token at a time:
h_t = f(h_{t-1}, x_t) # state has fixed size, independent of t
y_t = g(h_t)Cost per token is constant. Memory is constant. Process a million tokens and step one million looks exactly like step ten. The catch, historically, was twofold. RNNs were slow to train because that recurrence is inherently sequential (you can't compute step t until step t-1 finishes), and they forgot things: gradients vanished, long-range dependencies washed out. Attention won precisely because it sidestepped both problems. State-space models are the attempt to keep the RNN's linear cost while fixing what made RNNs bad.
What a state-space model actually is
A state-space model (SSM) comes from control theory. It maps an input signal to an output through a continuous linear system defined by matrices A, B, C:
h'(t) = A h(t) + B x(t) # how the hidden state evolves y(t) = C h(t) # how you read the state out
Discretize that for sequences and it becomes a linear recurrence: the next state is the current state pushed through A plus the new input through B. The elegant part is that a linear recurrence can also be unrolled and computed as a convolution. That means you can train it in parallel across the whole sequence (fast, like a Transformer) and then run it as a recurrence at inference (cheap, like an RNN). The S4 model (2021) showed this could actually model long-range dependencies well, by choosing A carefully using a structure called HiPPO that mathematically encodes "remember the input history." S4 beat Transformers on the Long Range Arena benchmark, a suite built specifically to stress sequences thousands of steps long. But on language, SSMs lagged. Something was missing.
Mamba: making the state selective
The missing piece was that classic SSMs are linear time-invariant: A, B, and C are the same at every step. The model processes "the" and "mitochondria" with identical dynamics. It can't choose to attend to one token and ignore another, because its behavior doesn't depend on the content flowing through it. Attention does exactly that. It's content-dependent by construction.
Mamba (Gu and Dao, 2023) fixes this with selectivity. It makes B, C, and the discretization step size functions of the input:
# Classic SSM: parameters fixed B, C, delta = constants # Mamba: parameters depend on the current token B_t = Linear_B(x_t) C_t = Linear_C(x_t) delta_t = softplus(Linear_delta(x_t)) # per-token "how much to update the state"
Now the model decides, per token, how much new information enters the state and how much old state persists. A filler word barely moves the state. A key entity strongly writes to it. That's a content-based gate, the thing attention had and prior SSMs didn't. (Note that A itself stays fixed; only B, C, and the step size become input-dependent.)
There's a cost. Once the parameters vary with the input, the "unroll it as a convolution" trick no longer applies. So Mamba's authors wrote a hardware-aware parallel scan: a GPU kernel that computes the recurrence in parallel across the sequence while keeping the big state tensors in fast SRAM instead of spilling to slow HBM. The result scales near-linearly in sequence length, keeps no KV cache, and matched or beat Transformers of similar size on language modeling in the original paper, while reporting up to 5× higher inference throughput on long sequences at constant memory per step.
The tradeoff nobody should skip
Here is the framing that separates people who understand this from people repeating hype. Mamba does not replace Transformers. It trades one weakness for another.
Attention keeps every past token available exactly, addressable by content. That makes it superb at what researchers call associative recall: "earlier the passphrase was 4471; what was the passphrase?" A Transformer looks it up. An SSM has to have written that fact into its fixed-size state and not overwritten it since. The state is a bottleneck. It's a lossy summary, not a transcript. This is the same skill a needle-in-a-haystack test probes: bury one fact in a huge context and ask for it back. Follow-up analyses (notably the work behind the Based architecture) showed pure SSMs measurably underperform attention on these recall-heavy, in-context-lookup tasks, precisely the skills that make in-context learning and long-document QA work. Compress history into a fixed vector and some retrieval will fail. That isn't a bug you can tune away. It's the nature of a bounded state.
So the honest scorecard: SSMs win on throughput and memory for long sequences; attention wins on exact recall and in-context lookup. Which matters more depends entirely on your workload. Summarizing a long log or transcribing audio leans on the SSM's strength. Answering "what account number did the customer give three thousand lines ago" leans on attention's. Most real applications need both at once, which is exactly why the field stopped treating this as a contest.
Hybrids, and Mamba's second act
If attention and SSMs have complementary strengths, the obvious move is to use both. That's what the strongest production designs do. Jamba (AI21, 2024) interleaves Mamba layers with a smaller number of attention layers and mixes in mixture-of-experts. Its block uses a 1:7 ratio, one attention layer for every seven Mamba layers, which lands a 52B-parameter model (12B active) that handles contexts up to 256K tokens with a KV cache a fraction the size of a comparable Transformer, because only the rare attention layers keep a cache. Later hybrids from other labs follow the same recipe: let cheap SSM layers carry the bulk of the sequence processing, and sprinkle in enough attention layers to preserve the exact-recall ability. A few attention layers go a surprisingly long way toward restoring what pure SSMs lose.
Mamba-3 (2026) is the inference-oriented successor to this line. Where the original Mamba was pitched as a general architecture, Mamba-3's design leans into how these models are actually deployed: decoding one token at a time, under memory and latency constraints, at long context. The through-line from 2023 to 2026 is consistent. Not "kill attention," but push the linear-time backbone far enough that hybrids can lean on it for most of the work. (Public detail on Mamba-3 is still thin; treat specifics as provisional until the paper and code are widely reproduced.)
What should you take away? When someone claims a new architecture "beats Transformers," ask at what. On long-context throughput and memory, SSMs genuinely change the economics. On associative recall, attention still holds. The reason hybrids dominate the frontier is that the field already internalized this tradeoff, and the smart bet isn't picking a winner. It's knowing which tool each layer of your model should be.
Sources
- Gu, Albert and Dao, Tri. "Mamba: Linear-Time Sequence Modeling with Selective State Spaces." 2023. https://arxiv.org/abs/2312.00752
- Gu, Albert; Goel, Karan; and Ré, Christopher. "Efficiently Modeling Long Sequences with Structured State Spaces" (S4). 2021. https://arxiv.org/abs/2111.00396
- Lieber, Opher et al. "Jamba: A Hybrid Transformer-Mamba Language Model." AI21 Labs, 2024. https://arxiv.org/abs/2403.19887
- Arora, Simran et al. "Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff" (Based). 2024. https://arxiv.org/abs/2402.18668
- Dao, Tri and Gu, Albert. "Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality" (Mamba-2). 2024. https://arxiv.org/abs/2405.21060
- Gu, Albert et al. "HiPPO: Recurrent Memory with Optimal Polynomial Projections." 2020. https://arxiv.org/abs/2008.07669