arXiv:2609.02620v1 Announce Type: new Abstract: Generative AI is changing how cultural artifacts are created and circulated, and with it our understanding…
Category: cs.AI updates on arXiv.org
Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting
arXiv:2609.02649v1 Announce Type: new Abstract: Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge…
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
arXiv:2609.02459v1 Announce Type: new Abstract: We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon,…
Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks
arXiv:2609.02399v1 Announce Type: new Abstract: Argumentation frameworks are useful tools for representing and reasoning with information in a variety of…
SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology
arXiv:2609.02292v1 Announce Type: new Abstract: The rapid proliferation of large language models (LLMs) and the growing diversity of their applications…
CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging
arXiv:2609.02273v1 Announce Type: new Abstract: Model merging provides an efficient paradigm for constructing multi-task large language models (LLMs)…
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
arXiv:2609.02302v1 Announce Type: new Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are…
SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning
arXiv:2609.02336v1 Announce Type: new Abstract: Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations.…
Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions
arXiv:2609.02371v1 Announce Type: new Abstract: With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is…
Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training
arXiv:2609.02244v1 Announce Type: new Abstract: Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete.…
