arXiv:2608.18744v1 Announce Type: new Abstract: Agents improve quickly against a reliable automatic metric and stall without one, and the applications…
Category: cs.AI updates on arXiv.org
Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization
arXiv:2608.18719v1 Announce Type: new Abstract: Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document,…
RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training
arXiv:2608.18682v1 Announce Type: new Abstract: Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to…
Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction
arXiv:2608.18677v1 Announce Type: new Abstract: Amid concerns that generative AI may standardize art interpretation, this paper examines whether LLM-based…
Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference
arXiv:2608.18591v1 Announce Type: new Abstract: Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking…
Candidate-Fate Accounting for Transparent Sensor Diagnostic Pipeline Search
arXiv:2608.18665v1 Announce Type: new Abstract: Industrial sensor diagnostics relies on preprocessing, representation, and classification pipelines,…
CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence
arXiv:2608.18613v1 Announce Type: new Abstract: Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that…
Preference Reasoning under Indeterminacy in Large Language Models
arXiv:2608.18631v1 Announce Type: new Abstract: As large language models evolve into decision-making agents, the ability to reason over preferences…
Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson
arXiv:2608.18531v1 Announce Type: new Abstract: Industrial explainable-recommendation systems built on LLMs incur a substantial serving cost: each request…
Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
arXiv:2608.18521v1 Announce Type: new Abstract: Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered…
