AirLLM runs a 70B model on 4GB VRAM — no quantization
Layer-wise inference slashes VRAM needs by 97%, letting you run 70B models on a 4GB GPU with no quantization required.

AirLLM flips the standard LLM loading model on its head. Instead of cramming 140GB+ of weights into VRAM all at once, it splits the model into individual layer shards (~2GB each), loads one, computes the forward pass, then flushes it before pulling the next. Peak VRAM stays under 4GB — even for 70B models. The Llama 3.1 405B runs on just 8GB.
The obvious catch: speed. Layer-by-layer disk I/O means inference runs well below 1 token/second on very large models — not a fit for real-time chat, but workable for batch jobs, local experiments, or hardware-constrained deployments. AirLLM v2.5 added async prefetching (next layer loads while current layer computes), squeezing out ~10% throughput gains. Optional block-wise quantization can push that up to 3×.
Supports Llama, Qwen, and Mistral families. Runs on Linux, Windows, macOS (including Apple Silicon via Metal). Fully open source on GitHub.
Why it matters: consumer-grade hardware just became viable for 70B inference — no cloud, no expensive GPU rental, no quantization compromise.
Sources
Primary: the company, paper or repository
- AirLLM – 70B Inference on a Single 4GB GPU github.com
Independent coverage
- AirLLM: Run 70B Models on 4GB GPUs Without Compromise blog.brightcoding.dev
Written by an AI pipeline from the sources above. Methodology · Report an error
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.