Thinking Without Speaking: BDH-CQ Shatters ARC-AGI Cost Frontiers
Pathway's 150M-parameter model fuses recurrent latent reasoning with in-context learning, achieving 29.5% on ARC-AGI-1 for less than a tenth of a cent per task.
The prevailing doctrine of frontier AI reasoning relies on a brute-force tax: serial narration. To solve a difficult logic puzzle or navigate abstract visual transformations, models like OpenAI's reasoning series, DeepSeek-R1, and Claude 3.7 Sonnet emit thousands of visible chain-of-thought (CoT) tokens. Every intermediate hypothesis must be projected through a discrete vocabulary, written to an ever-ballooning key-value cache, and parsed autoregressively before the model can take its next cognitive step.
BDH-CQ, a new reasoning architecture developed by researchers at Pathway, Bielik AI, and New York University (arXiv:2608.09888), fundamentally rejects this paradigm. By decoupling internal computation from natural-language emission and marrying recurrent latent reasoning with in-context learning, a lightweight 150M-parameter model scores 29.5% pass@2 on ARC-AGI-1 at a computed inference cost of just $0.00070 per task—less than one-tenth of a cent.
In doing so, BDH-CQ punches clean through the established ARC-AGI cost-accuracy Pareto frontier, proving that deep algorithmic reasoning does not require hundreds of billions of parameters or verbose token streams.
+--------------------------------------------------------------------------------+
| BDH-CQ DUAL-PHASE PIPELINE |
+--------------------------------------------------------------------------------+
[Demonstrations] -> D_1 -> D_2 -> ... -> D_K
│
▼
[Phase 1: Memory Update] ──> S_t = U_θ(S_{t-1}, D_t)
│
▼
[Compressed State: S_K]
│
[Query Input: x*] ──────────────────────────────────────┤
▼
[Latent Workspace: H_0]
│
┌─────────────┴─────────────┐
▼ │
[Phase 2: Recurrent Latent Steps] │
H_{r+1} = F_θ(H_r, S_K) │
(r = 0 ... R-1) │
│ │
└─────────────┬─────────────┘
▼
[Refined Latent: H_R]
│
▼
[One-Shot Output Decoder: ŷ]
The Flaw in Tokenized Chain-of-Thought
Over the past two years, test-time compute scaling has delivered massive performance jumps on verifiable benchmarks. However, forcing reasoning through natural language creates three severe bottlenecks:
- Computational Inefficiency: Natural language is low-bandwidth and high-overhead. Serializing intermediate algebraic states or coordinate matrices into ASCII characters burns GPU memory bandwidth on token projection rather than mathematical computation.
- Hypothesis Collapse: Autoregressive language generation forces the network to commit to a single discrete token at every step. Exploring superimposed or branching candidate rules in parallel requires explicit tree-search decoding or speculative branching, causing inference compute to explode.
- Memory Accumulation: As reasoning traces expand into tens of thousands of tokens, the quadratic memory consumption of Transformer attention caches degrades hardware utilization and inflates latency.
Cognitive neuroscience has long demonstrated that human conceptual reasoning and linguistic narration rely on distinct neurological pathways (Fedorenko & Varley, 2016). Thinking does not require speaking. BDH-CQ translates this neurological principle into machine learning architecture.
Under the Hood: The Baby Dragon Hatchling (BDH) Architecture
BDH-CQ builds upon the Dragon Hatchling (BDH) foundation, a post-transformer architecture inspired by scale-free biological neural networks. Unlike standard Transformers that rely on global softmax attention and dense matrix multiplications across a dynamic sequence, BDH structures computation through high-dimensional positive activations and low-rank synaptic communication.
Instead of managing an unbounded KV-cache, the model maintains a fixed-size associative memory state. When presented with a task, BDH-CQ executes a clean, two-phase reasoning cycle:
1. Contextual Memory Acquisition
Given a set of few-shot demonstrations $\mathcal{D} = {(x_t, y_t)}{t=1}^K$, the model ingests each input-output demonstration sequentially. A recurrent memory state $S_t$ updates continuously via a learned update operator $U\theta$:
$$S_t = U_\theta(S_{t-1}, D_t)$$
By the time all $K$ demonstration pairs are processed, the accumulated latent state $S_K$ contains a compressed representation of the transformation rule. Critically, this occurs entirely in-context—no weights are fine-tuned, and no test-time gradient descent is executed.
2. Iterative Latent Reasoning
To evaluate the test query $x_\star$, BDH-CQ encodes the query alongside the compressed rule $S_K$ to initialize a high-dimensional continuous workspace $H_0$:
$$H_0 = E_\theta(x_\star, S_K)$$
The network then executes $R$ recurrence iterations through an internal latent operator $F_\theta$, refining its continuous hypothesis without generating a single language token:
$$H_{r+1} = F_\theta(H_r, S_K) \quad \text{for } r = 0, \dots, R-1$$
Finally, the model decodes the refined latent representation $H_R$ into the final grid prediction $\hat{y} = G_\theta(H_R)$ in a single forward pass.
| Architecture Feature | Standard Frontier LLMs (o1, R1, Sonnet) | BDH-CQ (Pathway / NYU) |
|---|---|---|
| Reasoning Substrate | Discrete natural language tokens | Continuous high-dimensional latent space |
| Inference Compute Scaler | Number of generated CoT tokens | Recurrent latent update iterations ($R$) |
| Memory Footprint | Dynamic, expanding KV-cache | Fixed-size associative recurrent state |
| Few-Shot Adaptation | Token concatenation in context window | Continuous state updates ($S_t$) |
| ARC-AGI-1 Inference Cost | $0.50 – $10.00+ per task | $0.00070 per task |
Cracking ARC-AGI Without the Compute Tax
François Chollet's Abstraction and Reasoning Corpus (ARC) was intentionally designed to measure general fluid intelligence: the capacity to acquire novel abstract rules from minimal demonstrations without relying on memorized pretraining corpora.
Frontier LLMs have historically struggled with ARC because transforming spatial grids through token streams is inherently clumsy. Approaches that scored over 30% typically relied on generating thousands of Python programs via test-time tree search or fine-tuning task-specific adapters on the fly, pushing per-task inference costs into multi-dollar territories.
BDH-CQ's performance on the public ARC-AGI-1 evaluation set establishes a radical shift:
- Accuracy: Reaches 29.5% pass@2 using an ultra-compact 150M parameter model.
- Cost Efficiency: Operates at $0.00070 per task, representing a reduction of several orders of magnitude compared to frontier frontier-scale LLM search ensembles.
- Generalization: Neither task identifiers nor evaluation demonstrations were exposed during pretraining. The model synthesizes the underlying symmetry, color mapping, or topological translation strictly at test time.
Why Latent Reasoning Is the Next Frontier
BDH-CQ highlights a crucial turning point in AI systems design. While autoregressive chain-of-thought models remain essential for dialogue, document generation, and human-facing synthesis, they are fundamentally inefficient engines for raw symbolic and spatial planning.
As the industry seeks to deploy autonomous agents on local edge hardware, inside robotics controllers, and across massive high-throughput simulation environments, the economics of token generation become untenable. By demonstrating that recurrent latent computation can preserve the adaptive flexibilities of in-context learning while slashing inference compute by orders of magnitude, BDH-CQ provides a compelling blueprint for the post-Transformer, post-token reasoning era.
Sources
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.