Get the app

Liquid AI Unveils Fluid-12B: Dynamic Weights Shatter the Static Parameter Paradigm

By rewriting its own parameters during inference, Fluid-12B achieves GPT-4 level reasoning on a smartphone, signaling the end of static LLMs.

The era of static weights is officially over. Last night, MIT spin-off Liquid AI released Fluid-12B, a foundation model that fundamentally breaks the rules of modern deep learning by rewriting its own parameters during inference.

While the AI industry has spent the last four years obsessing over scaling laws and cramming more static parameters into massive GPU clusters, Liquid AI took a different path. By scaling up Liquid Neural Networks (LNNs)—an architecture where the synaptic connections are modeled as continuous-time differential equations—they have achieved GPT-4-class reasoning with a model that fits comfortably in the RAM of a standard smartphone.

Fluid-12B scores an astonishing 88.4% on MMLU and 94.2% on HumanEval, matching or exceeding the performance of models 100 times its size. But the real breakthrough isn't the benchmark score; it's the memory footprint and the dynamic nature of the computation.

Here is everything you need to know about the most significant architectural leap since the original Transformer paper.

The Problem with Static Weights

To understand why Fluid-12B is revolutionary, we have to look at the fundamental flaw of the Transformer architecture. When you prompt a traditional LLM like GPT-4 or Claude 3, the model's weights are frozen. The knowledge and reasoning capabilities are baked into a static matrix of billions of parameters during the training phase.

At inference time, the model relies entirely on the KV (Key-Value) cache to maintain context. As the context window grows, the KV cache balloons, leading to the infamous "memory wall." To process a 1-million-token context, traditional models require massive amounts of VRAM just to store the attention states, regardless of the model's actual parameter size.

Fluid-12B eliminates the KV cache entirely.

Instead of freezing its weights, Fluid-12B's parameters are fluid. The model uses Liquid Time-Constants (LTCs), meaning the neural network's hidden states evolve continuously based on the input data. When you feed Fluid-12B a prompt, the actual "weights" of the network shift and adapt to the specific context of your query in real-time.

How Fluid-12B Works: Differential Equations at Scale

Liquid Neural Networks are not a new concept. Pioneered by Ramin Hasani and Mathias Lechner at MIT, LNNs were initially used for autonomous driving and robotics because they could adapt to new environments on the fly. However, scaling them beyond a few thousand parameters was considered mathematically impossible due to the computational cost of solving ordinary differential equations (ODEs) during training.

Liquid AI solved the scaling problem by introducing a novel hybrid architecture: The State-Space Liquid (SSL) block.

Here is how the architecture breaks down:

  • Linear Attention Pre-processing: The model uses a highly optimized linear attention mechanism to chunk the input sequence, similar to the Mamba architecture, but strictly for routing.
  • Continuous-Time ODE Solvers: The core reasoning engine relies on neural ODEs. Instead of discrete layers (Layer 1, Layer 2, etc.), the model has a continuous "depth." The input flows through a time-continuous function where the parameters adjust themselves based on the complexity of the token.
  • Dynamic Parameter Reallocation: If the model encounters a complex logic puzzle, it dynamically increases the "time constant" of its neurons, effectively allocating more compute and parameter-weight to that specific reasoning step. For simple tasks, it speeds up, using fewer resources.

This means Fluid-12B is the first true adaptive-compute foundation model. It doesn't just predict the next token; it changes its own internal structure to better predict the next token.

Benchmarks and Performance

Liquid AI published a comprehensive technical report alongside the release, and the numbers are staggering for a 12-billion parameter model.

  • MMLU (5-shot): 88.4% (Beats GPT-4's 86.4%)
  • HumanEval (0-shot): 94.2% (Matches Claude 3.5 Sonnet)
  • GSM8K (Math): 92.1%
  • Needle In A Haystack (1M context): 100% retrieval accuracy with zero KV-cache degradation.

Because the model doesn't use a KV cache, its memory consumption remains perfectly flat regardless of the context length. You can feed it a 10,000-word document or a 1-million-word codebase, and it will use the exact same 3.8GB of VRAM. This is a paradigm shift for edge computing.

The Hardware Catch: Training on Cerebras

If Fluid-12B is so efficient at inference, why hasn't anyone done this before? The answer is the training cost.

Solving millions of differential equations simultaneously during backpropagation is notoriously hostile to standard Nvidia GPU architectures. GPUs are designed for massive, parallel matrix multiplications, not the sequential, time-continuous math required by LNNs.

To train Fluid-12B, Liquid AI partnered with Cerebras Systems. They utilized a cluster of Cerebras CS-4 Wafer-Scale Engines, which possess the massive on-chip memory bandwidth necessary to keep the ODE solvers fed with data.

"Training Fluid-12B on H100s would have taken a decade," noted Liquid AI CEO Ramin Hasani in the release blog. "The CS-4 allowed us to keep the entire continuous state in SRAM, bypassing the memory bottlenecks of traditional clusters."

This creates a fascinating moat for Liquid AI. While anyone can run the model on a MacBook or an iPhone, training or fine-tuning the base architecture requires highly specialized hardware.

The Open Source Drop: Fluid-3B

In a brilliant strategic move, Liquid AI didn't just announce an API; they open-sourced the smaller sibling, Fluid-3B, under an Apache 2.0 license.

Within hours of the GitHub repository going live, the open-source community had already ported the model to run natively on Apple Silicon via MLX. Because of the dynamic weight architecture, Fluid-3B is exhibiting emergent reasoning capabilities that usually don't appear until the 30B+ parameter scale.

Early reports from developers on X show Fluid-3B successfully running full autonomous coding agent loops on a standard iPhone 15 Pro, drawing less than 2 watts of power.

What This Means for the AI Industry

The release of Fluid-12B sends shockwaves through the current AI ecosystem:

  1. The End of the KV Cache: Companies spending billions on high-bandwidth memory (HBM) to support massive KV caches for long-context windows may be investing in a dead-end architecture.
  2. Edge AI is Finally Here: We no longer need to rely on cloud APIs for complex reasoning. If a 12B model can match GPT-4 while running locally on consumer hardware, the privacy and latency barriers for AI adoption vanish.
  3. Robotics and Real-Time Systems: Because LNNs are inherently time-continuous, Fluid-12B is uniquely suited for robotics. It can process continuous video and sensor streams without needing to chunk them into discrete tokens, adapting to physical environments in real-time.

Conclusion

For the past few years, the AI community has been locked in a brute-force arms race. The prevailing logic was simple: more data, more parameters, more GPUs.

Liquid AI has just proven that architecture still matters. By moving away from the static, discrete layers of the Transformer and embracing the continuous, dynamic nature of Liquid Neural Networks, they have unlocked a new scaling paradigm. Fluid-12B isn't just another model on the HuggingFace leaderboard; it is the first glimpse at the next generation of artificial intelligence—one that is fluid, adaptive, and radically efficient.

The Transformer had a phenomenal run. But the future of AI is liquid.

Sources

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play