When Agents Read Before They Code: SkyPilot's Auto-Researcher Speeds Up Llama.cpp by 15%
Code-only AI agents fail when optimizations require domain knowledge. By adding a literature review phase, SkyPilot's agent found C++ operator fusions humans missed.
At some point, staring at a codebase isn't enough to make it faster. You have to read the literature, study the hardware constraints, and see what competing projects are doing. It turns out, AI coding agents are no different.
In the past 24 hours, SkyPilot published a fascinating breakthrough with their pi-autoresearch framework that fundamentally changes how autonomous coding agents operate. By forcing the agent to conduct a "literature review"—reading arXiv papers, studying competing forks, and analyzing different hardware backends—before writing a single line of code, the agent autonomously discovered five major optimizations in llama.cpp's CPU inference path.
The result? A 15% speedup on x86 and a 5% speedup on ARM for text generation (tested on TinyLlama 1.1B). It achieved this in just three hours, running on four cloud VMs, for a total cost of $29.
This marks a critical shift in agentic workflows: moving from localized code-editing loops to research-driven software engineering.
The Ceiling of Code-Only Context
To understand why this is a breakthrough, we have to look at how auto-coding agents have operated until now.
The foundation was laid by Andrej Karpathy's autoresearch, which demonstrated that an LLM could autonomously improve a neural network training script by iteratively proposing changes, running the training loop, and checking the validation loss. SkyPilot later scaled this to 16 GPUs, driving val_bpb down by 2.87% over 910 experiments.
Recently, Shopify CEO Tobi Lütke applied this exact code-only loop to Liquid, Shopify's Ruby template engine. The agent ran ~120 experiments and produced 93 commits that cut parse and render time by 53% with zero regressions across 974 unit tests.
But here is the catch: code-only context only works when the optimization surface is visible within the source code itself.
In the Liquid example, the agent could read the tokenizer, identify that StringScanner was the bottleneck, and brainstorm alternatives entirely from the codebase. However, when SkyPilot pointed their agent at the CPU inference path of llama.cpp, the code-only approach hit a brick wall.
Wave 1: The Micro-Optimization Trap
The optimization search space for a C++ inference engine isn't as simple as swapping out a standard library function. The answers live outside the source code—in hardware manuals, academic papers, and the domain knowledge of senior systems engineers.
When the agent initially attacked llama.cpp using only the source code as context, it fell into the classic junior engineer trap: it went straight for SIMD micro-optimizations in the quantized dot products. It tried:
- AVX2 prefetching in the
Q4_0dot product inner loop (+0.8% speedup). - 2x loop unrolling with dual accumulators (+0.9% speedup).
- Hoisting block boundary calculations (+0.6% speedup).
- Eliminating a temporary buffer in
mul_mat(which actually caused a 2.8% regression).
All of these results were within the margin of noise.
Fascinatingly, the agent's own automated postmortem diagnosed the failure perfectly: "Wave 1 results show that micro-optimizations in the compute path give negligible returns because text generation is memory-bandwidth bound, not compute bound."
A 606 MiB model generating ~49 tokens per second consumes roughly 30 GB/s of memory bandwidth. On the c6i.2xlarge instances used for testing, this is brushing right up against the physical DRAM limit. No amount of clever AVX2 loop unrolling will make a difference when the CPU is stalled, waiting for model weights to arrive from memory. But the source code alone doesn't tell you that. You need to understand the roofline model.
The Pivot: Adding a Literature Research Phase
Realizing that hypothesis quality was the bottleneck, the SkyPilot team altered the agent's loop. Before running any experiments, the agent was instructed to perform a literature search.
Instead of the standard edit code -> run experiment -> check metric loop, the new architecture looks like this:
- Research: Read relevant arXiv papers, study competing forks (like
ik_llama.cpp), and analyze other backends within the same project (CUDA, Metal). - Hypothesize: Formulate optimization strategies based on external domain knowledge.
- Implement: Edit the C++ code.
- Verify: Compile, run the benchmark, and check against the test suite.
The research phase completely changed the agent's trajectory. By studying the CUDA and Metal backends of llama.cpp, the agent realized that the GPU implementations were utilizing operator fusion—combining multiple operations into a single pass to avoid writing intermediate results back to memory. It noticed these fusions were completely absent from the CPU backend.
The Winning Optimizations: Fusing the Memory Wall
Armed with this new domain knowledge, the agent ran a new batch of 30+ experiments. This time, instead of tweaking SIMD instructions, it focused on reducing memory roundtrips. Five major optimizations successfully landed:
- Softmax Fusion: Combining the exponentiation and normalization steps of softmax to reduce memory passes.
- RMS Norm Fusion: Fusing the root mean square normalization directly into the subsequent operations.
- Adaptive
from_floatParallelization: Dynamically adjusting thread counts based on tensor sizes during type conversion. - Graph-level RMS_NORM + MUL Fusion: A higher-level graph optimization that caught inefficiencies the compiler missed.
- Flash Attention KQ Fusion: This was the biggest win of the run. The agent successfully fused three separate passes over the Flash Attention QK tile into a single AVX2 Fused Multiply-Add (FMA) loop.
By fusing these operations, the CPU didn't have to constantly fetch and flush intermediate tensors to DRAM. The agent directly attacked the memory bandwidth bottleneck, resulting in a +15% performance bump on x86 and +5% on ARM.
The Economics of Autonomous Systems Engineering
Perhaps the most staggering part of this experiment is the cost.
To achieve a 15% speedup in a highly optimized, heavily scrutinized open-source C++ project like llama.cpp, you would typically need to hire a senior systems engineer. They would spend days reading the codebase, profiling the execution with perf, and testing hypotheses.
The SkyPilot agent did this in ~3 hours. It utilized four cloud VMs in parallel. The total cost? $29 ($20 in CPU VM costs, and $9 in LLM API calls).
What This Means for Coding Agents
This experiment proves that we are moving past the era of "glorified autocomplete."
When an agent can read a paper on Flash Attention, study a CUDA implementation, and successfully translate those concepts into an AVX2-optimized C++ CPU backend, it is performing actual software engineering.
Here are the key takeaways for the AI ecosystem:
- Context is More Than Code: RAG (Retrieval-Augmented Generation) for coding agents shouldn't just index the repository. It needs to index the project's GitHub issues, competing repositories, and relevant academic literature.
- Cross-Pollination as a Feature: The agent succeeded because it looked at the CUDA backend and asked, "Why isn't the CPU doing this?" AI agents are uniquely positioned to spot these cross-platform discrepancies instantly.
- The End of the "Low-Hanging Fruit" Era: As these research-driven agents are deployed across open-source repositories, the baseline performance of all software is going to jump. The $30 automated optimization run will become a standard CI/CD step.
If your agent is just reading your code, it's flying blind. It's time to let them read the literature.
Sources
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.