arXiv:2609.05090v1 Announce Type: new Abstract: Evaluation of medical artificial intelligence agents remains predominantly answer-centric, assessing only…
Category: cs.AI updates on arXiv.org
LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28
arXiv:2609.05093v1 Announce Type: new Abstract: We present Discovery Loop, a lightweight system that uses a large language model (LLM) to iteratively…
Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?
arXiv:2609.05088v1 Announce Type: new Abstract: AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is…
MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning
arXiv:2609.05075v1 Announce Type: new Abstract: General continual learning (GCL) aims to learn from evolving data streams without task identities,…
Language models judge war differently when tested for alignment
arXiv:2609.05009v1 Announce Type: new Abstract: Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to…
Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications
arXiv:2609.05040v1 Announce Type: new Abstract: As evolutionary transfer optimization (ETO) scales to larger collections of related tasks, problem…
Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment
arXiv:2609.05036v1 Announce Type: new Abstract: AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism…
A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering
arXiv:2609.04981v1 Announce Type: new Abstract: Recent structured RAG methods leverage tree- or graph-based reasoning structures to improve multi-hop QA.…
TROVE: Adaptive Agent Skill Orchestration via Trace-Grounded Route Validation and Editing
arXiv:2609.05019v1 Announce Type: new Abstract: Agents tend to optimize, select, or constrain execution structures before decisive runtime outcomes are…
Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing
arXiv:2609.04915v1 Announce Type: new Abstract: Many long-horizon LLM deployments face tight prompt budgets: latency, cost, and context limits make…
