Get the app

The 3:1 Convergence: Why Qwen and GLM Just Killed the Pure Transformer

Released on the same day, Qwen3.8-Flash-Next and GLM-5.3-Flash prove the frontier of million-token inference isn't quadratic softmax—it's linear recurrence with sparse retrieval.

The standard, quadratic-attention Transformer architecture is officially a legacy artifact for production inference. Within hours of each other, two leading open-weights AI labs—Alibaba’s Qwen team and Zhipu AI (Z.ai)—released next-generation frontier models that independently arrived at the exact same architectural breakthrough: a strict 3:1 ratio of linear-attention layers to sparse-retrieval attention layers inside a massive Mixture-of-Experts (MoE) trunk.

On August 26, Zhipu AI open-sourced GLM-5.3-Flash (under a permissive MIT license, unmasking the viral "Ox Alpha" stealth model from OpenRouter), while Alibaba rolled out Qwen3.8-Flash-Next, an architectural preview of its upcoming Qwen4 series. While both models aim for extreme serving cost-efficiency at 1-million-token context windows, their near-identical architectural blueprints signal a permanent inflection point for machine learning systems: full quadratic self-attention across all layers is dead.

Here is a technical dissection of how these two architectures solve the crushing KV cache memory bottleneck, how they differ, and what this convergence means for the future of inference economics.


The KV Cache Wall and the Death of Monolithic Softmax

For the past four years, scaling context windows to 1M+ tokens in pure Transformer models hit an inescapable physical limit: the KV cache memory wall.

In standard Multi-Head Attention (MHA) or Grouped-Query Attention (GQA), storing the key-value states for every token across 40 to 80 layers consumes tens of gigabytes of high-bandwidth memory (HBM) per user session. Even with Multi-Head Latent Attention (MLA), long-context concurrency rapidly exhausts GPU VRAM, forcing batch sizes down and driving per-token serving costs through the roof.

Both the Qwen and GLM teams tackled this crisis by adopting a hybrid recurrent-retrieval paradigm:

  • Linear attention layers compress the unbounded history of a sequence into a compact, fixed-size recurrent hidden state in $O(N)$ time. Memory footprint stays flat regardless of whether the prompt is 1,000 tokens or 1,000,000 tokens.
  • Sparse retrieval layers sit interspersed at deliberate intervals, retaining a traditional KV cache but only activating over top-ranked token chunks retrieved by dedicated lightweight indexers.

Remarkably, both research teams converged on the exact same layer pacing: three cheap linear recurrent layers for every one expensive sparse retrieval layer.

Architectural Metric Z.ai GLM-5.3-Flash Alibaba Qwen3.8-Flash-Next
Total Parameters 320B 180B (125B Base + 51B N-gram + 4B MTP)
Active Parameters / Token 18B 6B
Total Layer Count 45 layers 48 layers
Layer Layout 34 KDA Linear : 11 Sparse MLA 36 GDN Linear : 12 QSA Sparse
Linear : Sparse Ratio ~3 : 1 3 : 1
Linear Recurrence Mechanism Kimi Delta Attention (KDA) Gated DeltaNet (GDN)
Sparse Retrieval Budget Top-2048 tokens via Lightning Indexer 512 micro-blocks (2048 tokens) via QSA
Residual Connection Manifold-Constrained Hyper-Connections (mHC) 4-branch Gated Residuals
Position Encoding NoPE (No RoPE in sparse layers) RoPE maintained in attention
Base Context Window 1,048,576 tokens (native) 262,144 tokens (extensible to 1M)
Open License MIT License qwen-community-1.0

Convergence Point 1: 3:1 Linear-to-Sparse Interleaving

In GLM-5.3-Flash, Zhipu stacks 45 layers arranged as repeating blocks of three linear attention layers followed by one sparse MLA layer and an MoE feed-forward network, culminating in a 34:11 split. The linear layers implement Kimi Delta Attention (KDA), an evolution of linear attention that incorporates fine-grained per-channel decay gates.

In Qwen3.8-Flash-Next, Alibaba organizes its 48-layer trunk into identical 4-layer modules: three Gated DeltaNet (GDN) layers followed by one Qwen Sparse Attention (QSA) layer. Gated DeltaNet replaces unbounded KV accumulation with a state-space-like recurrent matrix update that uses per-head gating signals as a dynamic "eraser" to purge irrelevant context while preserving salient state transitions.

Because 75% of the network operates as a linear recurrence with zero token-wise KV cache allocation, KV cache memory overhead is reduced by 70% to 77% compared to standard dense Transformers. Context processing latency drops by over 3x at 256k+ token lengths.


Convergence Point 2: Top-2048 Micro-Block Indexing

When information must be retrieved across long sequences, linear recurrent states can suffer from retrieval degradation over complex multi-hop dependencies. The remaining 25% of layers—the sparse retrieval layers—solve this by executing full attention, but only on the most critical tokens.

Both teams implemented near-identical approximate retrieval indexing algorithms:

  • GLM-5.3-Flash deploys a 32-head Lightning Indexer coupled with IndexPool (which pools 4 keys into 1). It computes an ultra-fast preliminary similarity pass over the sequence and routes only the top-2048 candidate tokens into the full Multi-Head Latent Attention module.
  • Qwen3.8-Flash-Next uses QSA (Qwen Sparse Attention), which partitions the context into 4-token micro-blocks. A compressed indexer scores these micro-blocks, selecting exactly 512 blocks (2,048 tokens total) for the exact attention calculation.

By restricting the quadratic computation $O(K^2)$ strictly to $K=2048$ tokens regardless of whether the prompt is 50,000 or 1,000,000 tokens long, long-context attention computation shifts from quadratic disaster to an effective flat line.


Divergent Bets: Qwen's 51B "Engram" vs. GLM's Native Multimodal Trunk

While the attention-layer math converged, the two labs made radically different bets regarding parameter allocation and multimodality:

1. Qwen’s 51B N-Gram Table (The "Engram")

Qwen3.8-Flash-Next only activates 6B parameters per token out of its 125B neural weights. To compensate for reduced active parameter capacity without inflating floating-point operations, Alibaba introduced an unprecedented 51-billion-parameter static N-gram lookup table (over 20 million N-gram entries).

Because N-gram table lookups are $O(1)$ memory fetches requiring zero tensor math, they provide vast factual recall and linguistic surface regularization with almost no computational penalty. Crucially, community frameworks like Unsloth and vLLM have already demonstrated that this 51B engram table can be memory-mapped (mmap) onto standard PCIe SSDs or host CPU RAM, enabling high-speed local inference on unified-memory and consumer-grade workstations.

2. GLM’s Unified Native Multimodality

Zhipu AI focused its compute on eliminating the adapter tax for multimodal data. Rather than prepending an external vision encoder via an MLP cross-attention projector, GLM-5.3-Flash embeds a 24-layer Vision Transformer (ViT) directly into the model trunk.

Image and video patches (processed via a 448px input resolution, patch size 14, and 2x2 spatial merging) are projected straight into the core 4096-dimensional embedding space. Text, code, static visuals, and video frames share the exact same KDA/MLA attention blocks and routing gates from layer 0 to 45.


The Commercial Reality: Benchmark Parity at Commodity Cost

Early evaluations demonstrate that these hybrid architectures give up virtually nothing in capability compared to traditional brute-force models:

  • On code generation and repo-level tasks, Qwen3.8-Flash-Next achieves 48.1% on NL2Repo-Bench and 73.9% on CoWorkBench, rivaling Claude Opus 4.6 and DeepSeek-V4.
  • GLM-5.3-Flash achieves an Elo of 1773 on Artificial Analysis GDPval-AA v2, outperforming GLM-5.2 while cutting token serving costs by almost 90%.
  • On QwenCloud, hosted Qwen3.8-Flash is priced at $0.15 per million input tokens and $0.47 per million output tokens for a default 1M context window—undercutting closed frontier APIs by up to 95%.

What This Means for the Future of LLM Architectures

The simultaneous emergence of GLM-5.3-Flash, Qwen3.8-Flash-Next, and DeepSeek’s hybrid lineage proves that the AI research ecosystem has broken out of the vanilla Transformer monoculture.

By blending linear recurrence ($O(N)$ efficiency) with sparse indexed attention ($O(1)$ KV footprint) and hyper-connected MoE routing, open-weight developers can now deploy trillion-token-capable models on commodity clusters and consumer hardware without sacrificing deep reasoning. For enterprise architectures and local inference builders alike, the 3:1 hybrid MoE is the new undisputed standard.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play