arXiv:2609.24101v1 Announce Type: new Abstract: Automated biomedical evidence synthesis depends on retrieving published studies, but the biomedical…
Tag: cs.AI updates on arXiv.org
Self-Healing Harness for Runtime Oversight of Agent Self-Modification
arXiv:2609.24130v1 Announce Type: new Abstract: LLM agents can change their own future behavior, raising a basic control question of which self-generated…
Incremental Consistency Execution for Autonomous Intelligent Systems
arXiv:2609.24090v1 Announce Type: new Abstract: Long-horizon autonomous intelligent systems rely on heterogeneous components such as large language…
EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation
arXiv:2609.24115v1 Announce Type: new Abstract: Tool-calling LLM agents are increasingly deployed in enterprise applications. However, effective…
DocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents
arXiv:2609.24092v1 Announce Type: new Abstract: Real-world document processing systems rely on rigid, predefined schemas, yet critical target fields often…
Structured Decomposition for Reliable LLM-Generated Access Control Policies
arXiv:2609.24036v1 Announce Type: new Abstract: This paper presents an LLM-based system that translates natural-language access control policies (NLACPs)…
Representation-guided in-context learning for medical image interpretation with multimodal large language models
arXiv:2609.24057v1 Announce Type: new Abstract: Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal…
Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search
arXiv:2609.24025v1 Announce Type: new Abstract: We present a method for synthesizing reactive character behaviors for continuous games as compact,…
Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure
arXiv:2609.24012v2 Announce Type: new Abstract: Validation of generative social simulators often stops at face validity: emergent network structure is…
Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech
arXiv:2609.24016v1 Announce Type: new Abstract: Commercial large language models are increasingly deployed across African fintech infrastructure for fraud…
