AI news digest — August 28, 2026
6 items, each with its source.
Researchers train 2B LLM on RTX 5090 GPUs under five thousand dollars
Researchers released Puro-2B, an open-source pretraining recipe capable of training a 2-billion parameter model on up to 1.4 trillion tokens using consumer-grade RTX 5090 GPUs and FP8 precision. The best model configuration cost less than $6,900 in compute and approached Qwen2.5-1.5B performance under benchmark evaluations. The team also formulated a cost-scaling law demonstrating that $4,400 of consumer compute suffices to match earlier baseline architectures.
Why it matters. Full pretraining pipelines and curriculum research become accessible to academic labs without access to enterprise GPU clusters.
arxiv.orgTTPO enables label-free test-time policy optimization for mathematical reasoning models
Researchers introduced Test-Time Policy Optimization (TTPO), an asymmetric optimization framework that executes reinforcement learning and self-distillation during inference without ground-truth labels. The method distills rollouts matching pseudo-labels via on-policy self-distillation while penalizing disagreeing rollouts with grouped reinforcement learning. In benchmark evaluations, TTPO improved Qwen3-1.7B test-time reasoning accuracy from 38.0% to 45.2% on competition math sets.
Why it matters. Reasoning models can self-correct during inference on unseen math problems without degrading into majority-vote hallucinations.
arxiv.orgSingle prompt enables frontier LLMs to design near-optimal operations research algorithms
A new study demonstrated that frontier LLMs can automatically design algorithmic solvers for operations research problems such as inventory control, assortment optimization, and queueing network management. Given a single untuned problem specification and Python execution sandbox access, the model produced algorithms that matched or exceeded specialized hand-crafted baselines. The performance held even when algorithms were generated before evaluation instances were revealed.
Why it matters. Domain engineers can automate the design of customized heuristic dispatchers and inventory controllers without manual mathematical derivations.
arxiv.orgWikiSkill framework separates agent experience into persistent knowledge for skill evolution
Researchers introduced WikiSkill, a framework that consolidates raw agent execution traces into a persistent wiki to support reusable skill evolution across tasks. Across multiple agent benchmarks, the persistent accumulation of experience enabled smaller models equipped with evolved skills to outperform substantially larger unassisted models. The framework also enabled cross-model transfer of discovered skills across disparate model families.
Why it matters. Agentic systems can retain and reuse operational workflows across runs instead of repeating exploratory trial-and-error.
arxiv.orgPrefixing weak model trajectories prevents entropy collapse during verifiable reward training
A research study proposed using partial reasoning trajectories from smaller language models as prefix prompts during Reinforcement Learning with Verifiable Rewards (RLVR). The external prefixes disrupt overconfidence in the target policy and compel the model to explore distinct reasoning trajectories. Experiments across mathematical reasoning benchmarks showed consistent pass@k improvements without requiring auxiliary supervised fine-tuning or specialized reward modeling.
Why it matters. Reasoning models maintain diverse solution paths at high sample budgets instead of collapsing to brittle repetitive answers.
arxiv.orgDiscrete flow matching over electron rearrangements predicts complex chemical reaction mechanisms
Researchers developed MAELLE, a continuous-time Markov chain framework that models chemical reaction prediction through discrete flow matching over graph-structured electron occupation vectors. The architecture formulates reactant-to-product mapping via optimal transport without requiring step-by-step intermediate annotations. Benchmarks on USPTO-480K demonstrated high accuracy alongside robust generalization across complex out-of-distribution reactions.
Why it matters. Computational chemists can directly inspect step-by-step electron transfer trajectories to troubleshoot synthesis pathways and predict unwanted side products.
arxiv.orgFeed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.