arXiv:2608.06949v1 Announce Type: new Abstract: Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death…
Category: cs.AI updates on arXiv.org
Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
arXiv:2608.06931v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear…
Fast LapSum: Exact Differentiable Top-k at Million Scale
arXiv:2608.06912v1 Announce Type: new Abstract: The top-$k$ operation is a fundamental building block of modern sparse computation, enabling token…
Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiation
arXiv:2608.06922v1 Announce Type: new Abstract: Negotiation is a demanding social task for LLM agents, requiring strategic reasoning, persuasion, and…
Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
arXiv:2608.06909v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user…
TRIBE: Predicting Team Performance via Communication Behavior Ensembles
arXiv:2608.06926v1 Announce Type: new Abstract: Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics,…
ReGraph: Learning to Generate Recipe Graphs from Food Images
arXiv:2608.06917v1 Announce Type: new Abstract: Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food…
CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems
arXiv:2608.06871v1 Announce Type: new Abstract: Complex systems, core objects of study in artificial life, model diverse phenomena through nonlinear,…
From Points to Edges: Edge-Conditioned Spectral Operators for Physics-Sensitive PDE Learning
arXiv:2608.06894v1 Announce Type: new Abstract: Neural operators have become a central tool for solving partial differential equations (PDEs), with…
SkillEval: Decomposing Agent Skill Quality into Interpretable Signals
arXiv:2608.06891v1 Announce Type: new Abstract: Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use…
