arXiv:2608.20631v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring…
Category: cs.AI updates on arXiv.org
Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions
arXiv:2608.20649v1 Announce Type: new Abstract: Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces,…
Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance
arXiv:2608.20661v1 Announce Type: new Abstract: Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in…
Difficulty-Aware Semantic-ID Optimization for Generative Recommendation
arXiv:2608.20611v1 Announce Type: new Abstract: Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over…
Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
arXiv:2608.20569v1 Announce Type: new Abstract: Are frontier models able to introspect about their internal states? Recent work suggests that under…
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
arXiv:2608.20614v1 Announce Type: new Abstract: Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and…
FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
arXiv:2608.20574v1 Announce Type: new Abstract: Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or…
Dual-Cache Latent Space Communication between Heterogeneous Language Models
arXiv:2608.20617v1 Announce Type: new Abstract: Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in…
A Temporal Planning Approach for Intelligent Flood Response
arXiv:2608.20510v1 Announce Type: new Abstract: Effective response to multiple, simultaneously flooded areas requires coordinating appropriate actions in…
Consilience: Conformally Calibrated Communication Control for Hidden-Profile Multi-Agent Reasoning
arXiv:2608.20564v1 Announce Type: new Abstract: Multi-agent LLM systems can improve reasoning by pooling diverse perspectives, but their effectiveness…
