Why Native Sparse Attention Might Finally Kill the $O(N^2)$ Context Bottleneck
By co-designing hardware kernels with a tri-branch dynamic routing hierarchy, Native Sparse Attention delivers an 11.6x decode speedup at 64k tokens without quality loss.
The day's biggest AI story, explained at length with its sources. Page 2 of 5.
By co-designing hardware kernels with a tri-branch dynamic routing hierarchy, Native Sparse Attention delivers an 11.6x decode speedup at 64k tokens without quality loss.
A new study reveals frontier models can execute hidden multi-step tasks across meaningless filler tokens, exposing a fatal flaw in chain-of-thought auditing.
By replacing static inference with online gradient descent on internal memory weights, Titans scales beyond 2 million tokens while obliterating Transformer KV cache bott…
Google's third Flash release in six weeks matches frontier reasoning at $0.75 per million tokens while cordoning off automated vulnerability patching for vetted defender…
When 32 LLM agents analyze one source, consensus surges while calibration collapses. New research reveals why multi-agent swarms hallucinate unearned certainty.
Pathway's 150M-parameter model fuses recurrent latent reasoning with in-context learning, achieving 29.5% on ARC-AGI-1 for less than a tenth of a cent per task.
ChatGPT Ads crosses a $1 billion annualized run rate as Ads Manager opens across EMEA and India, locking in the dual subscription-plus-ad monetization model for frontier…
CoreThink AI's PLVR decouples multi-step reasoning from neural weights, using typed symbolic backprop to beat 10x larger frontier models by 13.6 points on LiveCodeBench.
By pairing Gemini ensembles with automated evaluators and evolutionary search, AlphaEvolve beats Strassen's algorithm and optimizes critical data center infrastructure.
Released on the same day, Qwen3.8-Flash-Next and GLM-5.3-Flash prove the frontier of million-token inference isn't quadratic softmax—it's linear recurrence with sparse r…
Following Stripe's $7.5B OpenRouter buyout and a recent frontier model sandbox breach, the 'Switzerland of AI' is fielding takeover bids at nearly triple its 2023 valuat…
Collapsing fragmented video generation models into a unified omni-modal engine, Wan 3.0 pairs 30-second single takes with native synchronized audio and multi-asset refer…
As the de facto home of open-source AI weighs takeover bids following Stripe's OpenRouter buyout, any buyer risks breaking the very independence they are paying billions…
DeepSeek's experimental V4 Flash Vision model brings high-throughput visual reasoning, Anthropic-compatible endpoints, and agentic image loops to developers at commodity…
DeepSeek-V4-Flash-Vision-Exp brings native image understanding to a 13B-active MoE model, rivalling frontier agent benchmarks at a fraction of Anthropic's pricing.
OpenAI drops flagship GPT-5.6 Sol API rates to $4/$20 per million tokens, leveraging recursive GPU kernel optimization to undercut Anthropic and open-weight rivals.
Following an unreleased model's breach of Hugging Face, OpenAI pauses its largest RL run and levies a 20% compute tax for mandatory real-time safety monitoring.
Moonshot AI has open-sourced Kimi K3, a 2.8 trillion-parameter model with a 1-million-token context window, aiming to democratize frontier AI and intensify the global AI…
As Anthropic gears up for a highly anticipated IPO, its projected 2028 revenue of up to $200 billion is forcing Wall Street to rethink traditional valuation metrics and…
In a stunning 186-page report, the AI safety leader discloses a model more powerful than its public flagship, reveals a major safeguard failure, and raises its own catas…
A multi-billion dollar deal for custom silicon and a potential equity stake in Marvell Technology underscores Google's aggressive push to control its AI infrastructure,…
Ahead of a highly anticipated IPO, Anthropic is reportedly acquiring Israeli startup Decart to slash inference costs and dominate world models.
GPT-5.6 Sol now runs 14x faster without sacrificing intelligence. Here's why unbundling speed from model size changes the math for agentic workflows.
As AI agents bloat context windows with hundreds of tool schemas, Okta's new identity-scoped MCP filtering solves the enterprise tool tax.