arXiv:2402.01767v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) significantly improves document-based question answering by…
Category: cs.AI updates on arXiv.org
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
arXiv:2608.23283v2 Announce Type: replace Abstract: General-purpose language models can reason and synthesize knowledge, but complex work also requires…
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks
arXiv:2608.23035v2 Announce Type: replace Abstract: As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key…
CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents
arXiv:2608.22577v2 Announce Type: replace Abstract: Long-horizon GUI agents can retain complete action histories as compact text, but only a few…
GenCoord: Skill-Path Commitments under Private Information
arXiv:2608.22055v2 Announce Type: replace Abstract: Suppose one embodied agent knows what must be built, while its teammate alone knows which…
Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models
arXiv:2608.22232v2 Announce Type: replace Abstract: Real-world situation appearances can deviate from their underlying physical states, challenging the…
SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing
arXiv:2608.22018v2 Announce Type: replace Abstract: Hate speech research has moved from coarse-grained classification towards structured parsing, where…
SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems
arXiv:2608.22979v2 Announce Type: replace Abstract: Deploying high-dimensional multimodal features in industrial recommender systems incurs substantial…
Auditing an AI-Generated Mathematical Proof: Human Assessment of OpenAI’s Quantum Parallel-Repetition Argument
arXiv:2608.14673v2 Announce Type: replace Abstract: We present an independent human assessment of the proof developed in Chapter 6 of OpenAI’s Ten…
Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies
arXiv:2608.16645v3 Announce Type: replace Abstract: Can a language model recover the true research idea of a published paper when given only that paper’s…
