arXiv:2608.22577v1 Announce Type: new Abstract: Long-horizon GUI agents can retain a complete interaction trace cheaply as textual action records, but…
Tag: AI
CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories
arXiv:2608.22533v1 Announce Type: new Abstract: Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated…
STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control
arXiv:2608.22538v1 Announce Type: new Abstract: Policy-governed agents must interpret case evidence while following an authorized procedure. We present…
ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation
arXiv:2608.22559v1 Announce Type: new Abstract: Rubrics aim to make language-model evaluation transparent by decomposing response quality into…
Google Takes Gemini Enterprise Into Big Law With Legal-Specific Agents
Google Cloud launched Gemini Enterprise for Legal on August 25, 2026, a purpose-built version of its Gemini Enterprise platform configured for law firms…
Scaling Curriculum Learning For Autonomous Driving
arXiv:2608.22549v1 Announce Type: new Abstract: Batched simulators for autonomous driving have recently enabled training reinforcement learning (RL)…
Small Reasoning Models are Instruction Followers in Function Calling
arXiv:2608.22472v1 Announce Type: new Abstract: Function calling represents the core capability of agentic large language models (LLMs). Existing research…
HANSARD: A Reference Architecture for Forensic Readiness, Runtime Witnessing, and Graded Attribution in Autonomous Multi-Agent AI Systems
arXiv:2608.22512v1 Announce Type: new Abstract: Autonomous multi-agent systems nowadays act in finance, software supply chains, and security operations.…
New Pharma Regs Have Created a Perfect Storm for the Industry – AI Is the Fix
Following a recent slew of regulations, pharma companies are increasingly struggling to roll out their products, falling at the commercialisation hurdle.…
When Does AI for PDEs Yield Scientific Evidence?
arXiv:2608.22504v1 Announce Type: new Abstract: Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy.…
