Get the app
LLMs

Local SOTA Is Here: Run Top-Tier AI on Your Own GPU

New open-weight models now hit state-of-the-art benchmarks on RTX 3090/4090/5090, MacBook Pro, and DGX Spark — no cloud needed.

Local SOTA Is Here: Run Top-Tier AI on Your Own GPU

The gap between local and cloud AI just collapsed again. Qwen3.5 — Alibaba's latest open-weight family — is now fully optimized for consumer NVIDIA GPUs via Ollama and llama.cpp, hitting up to 256 tokens/sec on RTX 5090 and running comfortably on MacBook Pro configs from 24 to 96 GB unified memory.

The lineup covers every tier: Qwen3.5-9B fits in 10–16 GB VRAM, making it viable on an RTX 3090. The 35B-A3B MoE variant punches like a 14B dense model at 196 tokens/sec on a 4090. For DGX Spark owners, Qwen3.5-122B runs at ~40 tokens/sec with 128 GB unified memory. The 397B flagship is in range too — on multi-GPU or DGX Spark setups.

Benchmarks are serious: Qwen3.5-9B beats GPT-OSS-120B on GPQA Diamond (81.7 vs 71.5) and HMMT Feb 2025. It supports 256K context, vision, and 201 languages natively.

Why it matters: local SOTA is no longer theoretical — if you have the hardware, the model quality is there today.

Sources

Primary: the company, paper or repository

Independent coverage

Written by an AI pipeline from the sources above. Methodology · Report an error

The daily AI brief, on your phone.

Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.

Get it on Google Play