arXiv:2609.02786v1 Announce Type: new Abstract: The performance of LLM-based agents is jointly shaped by the base model and the harness used when…
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
arXiv:2609.02749v1 Announce Type: new Abstract: Autonomous agents are beginning to carry out machine-learning (ML) research end to end. These agents…
Door-in-the-Face Requests and Refusal Behaviour in Large Language Models
arXiv:2609.02707v1 Announce Type: new Abstract: Does the door-in-the-face technique work on language models? In humans, a large request that is refused…
AI News Brief Hourly Summary 2026-09-03 11h : 14 posts
14 posts published in the last hour 08:32UTP-Bench: Uncertainty-aware Travel Planning Benchmark 08:32Collective creativity in hybrid societies 08:32Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting 08:32CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI 08:32Anthropic ramps up…
UTP-Bench: Uncertainty-aware Travel Planning Benchmark
arXiv:2609.02421v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary…
Collective creativity in hybrid societies
arXiv:2609.02620v1 Announce Type: new Abstract: Generative AI is changing how cultural artifacts are created and circulated, and with it our understanding…
Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting
arXiv:2609.02649v1 Announce Type: new Abstract: Aggregating noisy, conflicting textual hypotheses into a reliable consensus is a fundamental challenge…
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
arXiv:2609.02459v1 Announce Type: new Abstract: We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon,…
Anthropic ramps up Claude infrastructure with $35 billion Lambda deal
Anthropic has signed a $35 billion cloud computing deal with Lambda, an Nvidia-backed cloud provider. The article Anthropic ramps up Claude infrastructure…
Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks
arXiv:2609.02399v1 Announce Type: new Abstract: Argumentation frameworks are useful tools for representing and reasoning with information in a variety of…
SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology
arXiv:2609.02292v1 Announce Type: new Abstract: The rapid proliferation of large language models (LLMs) and the growing diversity of their applications…
CoMerge: Conflict-Driven Preference Optimization for Multi-Task Model Merging
arXiv:2609.02273v1 Announce Type: new Abstract: Model merging provides an efficient paradigm for constructing multi-task large language models (LLMs)…
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
arXiv:2609.02302v1 Announce Type: new Abstract: A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are…
Training a coding model to paint watercolours with TRL and OpenEnv
This post has no text preview — click the link below to read the original article. This article has been indexed from Hugging Face – Blog Read the original article: Training a coding model to paint watercolours with TRL and…
SALA: Semantic-Aware Logical Alignment for Complex Reasoning in In-Context Learning
arXiv:2609.02336v1 Announce Type: new Abstract: Effective in-context learning (ICL) for complex reasoning relies on selecting the right demonstrations.…
AfterQuery reportedly becomes Y Combinator’s fastest-ever unicorn, now valued at $3.2B
AI model-training startup AfterQuery has reportedly raised a round that valued it at $3.2 billion, just five months after announcing its $30 million…
Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions
arXiv:2609.02371v1 Announce Type: new Abstract: With the proliferation of LLM agents, the ability to understand and diagnose failures in agents is…
AI News Brief Hourly Summary 2026-09-03 10h : 12 posts
12 posts published in the last hour 07:32Task-Level Natural Language Priors as Learning Signals for Low-Resource LLM Training 07:32Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality 07:32LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails…
