arXiv:2608.29460v1 Announce Type: new Abstract: When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or…
Category: cs.AI updates on arXiv.org
Toward Latent Language Model Skills Steering and Optimization: An Empirical Study
arXiv:2608.29459v1 Announce Type: new Abstract: Skills, as a useful abstraction for the procedural capabilities of large language models (LLMs), capture…
FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents
arXiv:2608.29372v1 Announce Type: new Abstract: Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out historical…
Evaluating Tiny Recursive Models Across Training for Code Generation
arXiv:2608.29376v1 Announce Type: new Abstract: Code generation increasingly relies on large transformer models, whose capability advances with scale. Yet…
Plant-Inspired AI: Plants as Inspiration for Novel Problem Formulations, and Two Case Studies
arXiv:2608.29356v1 Announce Type: new Abstract: Artificial Intelligence (AI) has long been inspired by studies of biological intelligence. Reinforcement…
Reviving our data foundations is the most disruptive step to data maturity
arXiv:2608.29368v1 Announce Type: new Abstract: The most disruptive step that enterprises of small-medium size and maturity can take to make the most of…
LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO
arXiv:2608.29357v1 Announce Type: new Abstract: Multimodal search agents answer visual questions by interleaving image understanding, web retrieval, tool…
APPSolver: Adaptive Patch Partitioning for Point-Wise Ship Flow Prediction on Unstructured Meshes
arXiv:2608.29355v1 Announce Type: new Abstract: Large non-uniform point sets make direct attention-based surrogate modeling costly for ship hydrodynamics.…
TRACER: Per-Tool Context Retention for LLM Agents via Consequence-Attributed Reinforcement Learning
arXiv:2608.29363v1 Announce Type: new Abstract: Enterprise data agents answer business queries by chaining many tool calls over multiple reasoning steps,…
BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations
arXiv:2608.29345v1 Announce Type: new Abstract: While recent Large Language Model (LLM)-based text-to-SQL systems achieve impressive performance on…
