arXiv:2609.30050v1 Announce Type: new Abstract: We present NNV3, the latest version of the Neural Network Verification (NNV) tool, a MATLAB framework for…
Category: cs.AI updates on arXiv.org
Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
arXiv:2609.30027v1 Announce Type: new Abstract: Frontier language models are rarely used in clinical workflows because the realistic, longitudinal…
Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models
arXiv:2609.30048v1 Announce Type: new Abstract: If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one…
SciWalker: Synthesizing Scientific Coding Problems with Operator Graphs and Execution Feedback
arXiv:2609.30054v1 Announce Type: new Abstract: Improving the scientific coding capabilities of large language models (LLMs) requires high-quality…
How does Adversarial Influence Scale in Multi-Agent Systems?
arXiv:2609.30028v1 Announce Type: new Abstract: Multi-agent deliberation can improve performance, but what happens when some agents do not act in good…
Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes
arXiv:2609.29952v1 Announce Type: new Abstract: Before a product or policy change ships, the question that matters is how people will react to it. Augur…
ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation
arXiv:2609.29948v1 Announce Type: new Abstract: Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack…
Who Holds the Pen? Let Specifications, Not Agents, Sign Off
arXiv:2609.29921v1 Announce Type: new Abstract: Large language model agents increasingly combine generation, decision-making, execution, and…
Advancing Model Research in AgentX: Long-Horizon Autonomy for Industrial Recommender Systems
arXiv:2609.30001v1 Announce Type: new Abstract: Sustaining industrial recommendation research requires using the results of one experiment to decide what…
Neuro-symbolic AI for Industrial Configuration
arXiv:2609.29947v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown impressive performance on a wide range of generative tasks. Yet…
