arXiv:2608.13719v1 Announce Type: new Abstract: Autonomous systems can fail in rare and heterogeneous ways, making real-world failure discovery difficult…
Category: AI
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
arXiv:2608.13667v1 Announce Type: new Abstract: LLM agents in the ReAct paradigm alternate between reasoning, acting, and observing, but deliberate…
Learning to Assemble Novel Structures with Unfamiliar Parts under Semantic Constraints
arXiv:2608.13684v1 Announce Type: new Abstract: This paper describes a neurosymbolic architecture for learning to assemble novel structures using evidence…
Algorithm Design and Physician Liability
arXiv:2608.13618v1 Announce Type: new Abstract: A single clinical algorithm can deliver unequal accuracy across patient groups, and concern about such…
Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning
arXiv:2608.13621v1 Announce Type: new Abstract: A hidden Markov model (HMM) combines three roles: inference of a hidden-state belief from observations,…
Reward Machines for Signal Temporal Logic
arXiv:2608.13625v1 Announce Type: new Abstract: Signal temporal logic (STL) provides a formal language for specifying real-time properties of real-valued…
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
arXiv:2608.13622v1 Announce Type: new Abstract: Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for…
A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
arXiv:2608.13626v1 Announce Type: new Abstract: A hidden state signal can be decodable or causally usable without supporting a reusable action map. We…
Evaluating Agentic Learning Harness Capabilities Without Labels via the Scaling Hypothesis
arXiv:2608.13608v1 Announce Type: new Abstract: Agentic “Continual Learning Harnesses”, systems that pair an LLM with retrieval or memory to improve from…
SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise Data
arXiv:2608.13612v1 Announce Type: new Abstract: Natural-language interfaces to enterprise data must translate underspecified requests into governed,…
