arXiv:2609.11028v1 Announce Type: cross Abstract: LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe…
Category: AI
What a Random Draw from the MCP Registry Contains, and What Tool-Use Benchmarks Contain Instead
arXiv:2609.10962v1 Announce Type: cross Abstract: Studies of the Model Context Protocol (MCP) server ecosystem draw their samples in ways that quietly…
Robust Multimodal Sentiment Analysis with Incomplete Modalities via Semantic-aware Completeness based Reconstruction
arXiv:2609.10950v1 Announce Type: cross Abstract: Recent multimodal sentiment analysis studies increasingly adopt text-centric fusion approaches to…
EGGROLL, Unrolled: Understanding and Improving Low-Rank Evolution Strategies at Scale
arXiv:2609.10980v1 Announce Type: cross Abstract: EGGROLL makes evolution strategies (ES) practical for LLMs by replacing dense Gaussian weight…
A Mathematical Theory of Pragmatic Information
arXiv:2609.10986v1 Announce Type: cross Abstract: We propose a pragmatic information theory unifying communication, control, and decision-making. Its core…
AI models’ written reasoning steps correspond to distinct internal patterns, a new study finds
Reasoning steps like calculation, formula retrieval, and deduction are clearly separable in a model’s internal states, especially in the middle layers.…
Importance Weighting for Unlabeled-unlabeled Learning under Distribution Shift
arXiv:2609.10994v1 Announce Type: cross Abstract: Unlabeled-unlabeled (UU) learning allows us to learn a binary classifier from two sets of unlabeled data…
Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training
arXiv:2609.10939v1 Announce Type: cross Abstract: Clinical education must prepare medical students to conduct safe and coherent patient interviews under…
ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
arXiv:2609.10895v1 Announce Type: cross Abstract: Reacting to sudden physical hazards (catching a slipping plate, dodging a falling knife) is both a…
DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents
arXiv:2609.10892v1 Announce Type: cross Abstract: When an indirect prompt injection succeeds against an LLM agent, the compromise is visible in the…
