AI news digest — October 2, 2026
6 items, each with its source.
Perplexity releases open-weight pplx-decider-v1-27b and launches low-cost Decisions API
Perplexity released Apache-2.0 weights for pplx-decider-v1-27b, a 27B-parameter multimodal decision model fine-tuned from Qwen3.8-27B. Alongside the weights, the company launched a hosted Decisions API priced at $0.04 per million input tokens with free output generation. The model produces calibrated probability distributions and discrete classification outputs in a single forward pass rather than generating autoregressive text.
Why it matters. Developers can replace slow, token-heavy LLM routing and classification prompts with calibrated categorical probabilities at a fraction of generative inference costs.
huggingface.coMIT researchers introduce VISTA visual harness achieving perfect score on ARC-AGI-3
Researchers from MIT introduced VISTA, a visual harness framework that equips multimodal language models with long-horizon visual perception and lossless memory. VISTA enables agents to retrieve past visual observations and reorganize their visual input dynamically during multi-step reasoning. Tested with Claude Opus 5.0 on the ARC-AGI-3 benchmark, the framework achieved a 100.00 score while requiring 57.4% fewer actions than first-time human participants.
Why it matters. Visual agent performance on complex interactive puzzles can be solved through structured observation memory harnesses without fine-tuning foundation model parameters.
arxiv.orgUC Berkeley researchers unveil RPG framework for weight-free robot skill self-improvement
UC Berkeley researchers introduced Reconstruct, Practice, Go Real (RPG), an autonomous skill acquisition framework that improves physical robot controllers without updating neural network weights. RPG extracts manipulation primitives from offline datasets, builds practice tasks in simulation, and generates reusable symbolic skills and system prompts guided by failure diagnosis. In hardware evaluation, the frozen system attained a 100% success rate across 30 physical manipulation trials.
Why it matters. Embodied AI systems can autonomously acquire and verify robust physical skills in simulation without requiring expensive real-world teleoperation data or model parameter updates.
arxiv.orgTACO optimizer slashes LLM fine-tuning memory by 174x via operator-norm steepest descent
Researchers introduced TACO, an optimization algorithm for full-parameter LLM fine-tuning that computes steepest descent under a dimension-normalized operator norm. By updating only the sign of the largest magnitude entry per weight matrix column, TACO reduces persistent optimizer memory by 174x compared to 8-bit AdamW on OPT-13B. The approach enables full-parameter fine-tuning of 30B-to-32B parameter models on a single 80GB H100 GPU without loss of training accuracy.
Why it matters. Full-parameter fine-tuning of 30B-scale models becomes accessible on single-GPU hardware without resorting to parameter-efficient adapters or multi-node clusters.
arxiv.orgAutoCompact trains coding agents to autonomously compress context during long-horizon tasks
Researchers developed AutoCompact, a reinforcement learning method that trains coding agents to autonomously manage and compress their context windows during long software engineering workflows. The model learns when to summarize intermediate progress and prune obsolete exploration trajectories as part of its internal action policy. On SWE-bench Verified and SWE-PolyBench Verified, AutoCompact boosted task resolution rates by 9.2% and 5.0% over base agent policies.
Why it matters. Long-horizon agent reliability no longer degrades as contexts lengthen because compaction becomes a learned, policy-driven action instead of a crude rule-based buffer clear.
arxiv.orgStanford and CMU researchers release ScholarCatalyst benchmark for scientific discovery agents
A multi-institution team released ScholarCatalyst, an evaluation benchmark measuring how effectively AI agents retrieve foundational literature that inspires novel scientific ideas. Built from annotations by 184 lead authors across 207 computer science papers, the benchmark tests agents on identifying key prior work available prior to project inception. State-of-the-art retrieval and reasoning agents scored at most 0.51 Recall@20, revealing significant gaps in exploratory scientific search.
Why it matters. Frontier models struggle to replicate human scientific intuition in cross-domain literature retrieval, exposing a major bottleneck for fully autonomous research systems.
arxiv.orgFeed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.