arXiv:2609.18442v1 Announce Type: new Abstract: Safe motion planning in automated driving requires anticipating evolving traffic risks and deciding when…
Category: AI
What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models
arXiv:2609.18286v1 Announce Type: new Abstract: Chess has long served as a model domain for studying search, expertise, decision-making, and artificial…
Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents
arXiv:2609.18357v1 Announce Type: new Abstract: Large language model (LLM) pricing agents may respond to how market data is presented, even when its…
Visual Compliance via Executable Safety Rule Entailment
arXiv:2609.18328v1 Announce Type: new Abstract: Recent advances in LLMs and VLMs have enabled safety systems to reason beyond simple risk patterns toward…
Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
arXiv:2609.18366v1 Announce Type: new Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a…
Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition
arXiv:2609.18346v1 Announce Type: new Abstract: Large language models (LLM) deployed as autonomous pricing agents may sustain supracompetitive prices…
REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement
arXiv:2609.18262v1 Announce Type: new Abstract: Precise retrieval of scientific information is fundamentally constrained by long-tailed concepts and high…
Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI
arXiv:2609.18272v1 Announce Type: new Abstract: Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of…
Building Trust in Artificial Intelligence: A Necessity for Railway Applications
arXiv:2609.18278v1 Announce Type: new Abstract: Artificial Intelligence (AI) is currently only applied to non-safety critical applications due to the…
BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs
arXiv:2609.18270v1 Announce Type: new Abstract: Payment operations are a critical financial infrastructure, but the value of large language models in this…
