14 posts published in the last hour 08:32UTP-Bench: Uncertainty-aware Travel Planning Benchmark 08:32Collective creativity in hybrid societies 08:32Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting 08:32CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI 08:32Anthropic ramps up…
Author: script
UTP-Bench: Uncertainty-aware Travel Planning Benchmark
arXiv:2609.02421v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary…
Collective creativity in hybrid societies
arXiv:2609.02620v1 Announce Type: new Abstract: Generative AI is changing how cultural artifacts are created and circulated, and with it our understanding…
Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting
arXiv:2609.02649v1 Announce Type: new Abstract: Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge…
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
arXiv:2609.02459v1 Announce Type: new Abstract: We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon,…
Anthropic ramps up Claude infrastructure with $35 billion Lambda deal
Anthropic has signed a $35 billion cloud computing deal with Lambda, an Nvidia-backed cloud provider. The article Anthropic ramps up Claude infrastructure…
Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks
arXiv:2609.02399v1 Announce Type: new Abstract: Argumentation frameworks are useful tools for representing and reasoning with information in a variety of…
SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology
arXiv:2609.02292v1 Announce Type: new Abstract: The rapid proliferation of large language models (LLMs) and the growing diversity of their applications…
CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging
arXiv:2609.02273v1 Announce Type: new Abstract: Model merging provides an efficient paradigm for constructing multi-task large language models (LLMs)…
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
arXiv:2609.02302v1 Announce Type: new Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are…
