arXiv:2609.28587v1 Announce Type: cross Abstract: Large language models can interpret natural lan- guage, yet robust decisions remain challenging.…
Category: cs.AI updates on arXiv.org
The Fellowship of the Query: Learning Retrieval Actions
arXiv:2609.28653v1 Announce Type: cross Abstract: Retrieval-augmented question answering requires control decisions about when to decompose a question,…
Learning to Discover Interesting Mathematics
arXiv:2609.28603v1 Announce Type: cross Abstract: Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical…
UO-FIE: Combining Exact-Label Supervision with Graded Utility for Factivity Inference
arXiv:2609.28605v1 Announce Type: cross Abstract: The Factivity Inference Evaluation 2026 (FIE2026) classifies Chinese context-hypothesis pairs into nine…
Decision Hijacking: Prompt Injection Attacks on Jev’s Typed Probabilistic Decisions
arXiv:2609.28613v1 Announce Type: cross Abstract: Most studies of prompt injection focus on generative agents, leaving their effects on models with…
When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages
arXiv:2609.28565v1 Announce Type: cross Abstract: Post hoc explanation methods such as SHAP and LIME are widely used to interpret text classifiers, but…
Auditability Is Not One Property: Rule Overlap, Behavioural Agreement, and Composition in Reinforcement Learning
arXiv:2609.28581v1 Announce Type: cross Abstract: Reinforcement-learning (RL) policies are often distributed as opaque neural checkpoints, while training…
SGA: Uncertainty Quantification for Multi-Step Forecasting in Time Series Foundation Models
arXiv:2609.28582v1 Announce Type: cross Abstract: The recent emergence of Time Series Foundation Models (TSFMs) has significantly advanced multi-step…
Where Cyber Agents Struggle: Bottleneck Analysis of Multi-Stage LLM Agents
arXiv:2609.28572v1 Announce Type: cross Abstract: Multi-stage LLM-based cyber agents may complete attack workflows while remaining brittle, costly, or…
Persistent Billable State: Denial-of-Wallet Attacks and Defenses in Tool-Calling LLM Agents
arXiv:2609.28585v1 Announce Type: cross Abstract: Multi-step tool-calling LLM agents rely on host runtimes to preserve state across turns. When a runtime…
