arXiv:2608.23205v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have revolutionized reasoning in LLMs, and the increasing public…
Tag: cs.AI updates on arXiv.org
Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement
arXiv:2608.23098v1 Announce Type: new Abstract: Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words, imagery,…
POOL: Propagated Uncertainty Over Lookalikes
arXiv:2608.23086v1 Announce Type: new Abstract: Black-box large language models need confidence scores that can separate likely-correct from…
Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies
arXiv:2608.23061v1 Announce Type: new Abstract: Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains…
From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation
arXiv:2608.23045v1 Announce Type: new Abstract: Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks…
From Generation to Simulation: How Far Are World Models from Being True Simulators?
arXiv:2608.23070v1 Announce Type: new Abstract: With the rapid progress of diffusion models and large-scale video generation, generative world models are…
LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications
arXiv:2608.23058v1 Announce Type: new Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with…
AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models
arXiv:2608.23078v1 Announce Type: new Abstract: Large language models increasingly operate over large collections of tools, functions, APIs, and…
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
arXiv:2608.23041v1 Announce Type: new Abstract: LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended…
Artificial Empathy: Towards a Framework for Unsupervised Agency Detection and Policy Reconstruction
arXiv:2608.23030v1 Announce Type: new Abstract: We study how an AI system can identify and model other agents in its environment from observation alone,…
