arXiv:2608.13625v1 Announce Type: new Abstract: Signal temporal logic (STL) provides a formal language for specifying real-time properties of real-valued…
Category: cs.AI updates on arXiv.org
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
arXiv:2608.13622v1 Announce Type: new Abstract: Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for…
A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
arXiv:2608.13626v1 Announce Type: new Abstract: A hidden state signal can be decodable or causally usable without supporting a reusable action map. We…
Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
arXiv:2608.13608v1 Announce Type: new Abstract: Agentic “Continual Learning Harnesses”, systems that pair an LLM with retrieval or memory to improve from…
SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data
arXiv:2608.13612v1 Announce Type: new Abstract: Natural-language interfaces to enterprise data must translate underspecified requests into governed,…
MobileMem: Learning from a Year of Mobile Experiences
arXiv:2608.13606v1 Announce Type: new Abstract: The next generation of AI agents is increasingly moving beyond systems that answer isolated questions…
No Universal Signal Predicts Sample-Level LLM Regression under Version Updates
arXiv:2608.13607v1 Announce Type: new Abstract: Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. But…
How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generating Clinical Compliance Insights
arXiv:2608.13617v1 Announce Type: new Abstract: Verifying whether clinical care follows evidence-based protocols is a natural neuro-symbolic problem, yet…
Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI Agents
arXiv:2608.13604v1 Announce Type: new Abstract: Detection of misunderstanding is an urgent problem to solve because communication has moved away from…
AI Evaluation Should Work With Humans
arXiv:2608.13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman…
