arXiv:2608.23070v1 Announce Type: new Abstract: With the rapid progress of diffusion models and large-scale video generation, generative world models are…
Category: cs.AI updates on arXiv.org
LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications
arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with…
AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models
arXiv:2608.23078v1 Announce Type: new Abstract: Large language models increasingly operate over large collections of tools, functions, APIs, and…
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
arXiv:2608.23041v1 Announce Type: new Abstract: LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended…
Artificial Empathy: Towards a Framework for Unsupervised Agency Detection and Policy Reconstruction
arXiv:2608.23030v1 Announce Type: new Abstract: We study how an AI system can identify and model other agents in its environment from observation alone,…
PatchWrite: One Line, Not One Section — Compile-Gated, Validity-Preserving Editing for AI-Drafted Manuscripts
arXiv:2608.23001v1 Announce Type: new Abstract: Automated manuscript pipelines often regenerate an entire section to repair a local defect, allowing…
PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies
arXiv:2608.23028v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and…
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
arXiv:2608.23035v1 Announce Type: new Abstract: As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key…
Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B
arXiv:2608.22975v1 Announce Type: new Abstract: Embodied multimodal agents must answer from growing observation streams under a fixed per-decision token…
Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents
arXiv:2608.22963v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit…
