arXiv:2412.07255v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) demonstrate remarkable capabilities in generative tasks but pose…
Tag: cs.AI updates on arXiv.org
Unleash LLMs Potential for Sequential Recommendation by Coordinating Dual Dynamic Index Mechanism
arXiv:2409.09253v2 Announce Type: replace-cross Abstract: Owing to the unprecedented capability in semantic understanding and logical reasoning, large…
Abstention vs. Hallucination: Benchmarking LLM Source Attribution for Scientific Citations
arXiv:2405.02228v5 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly generate citation-backed responses, yet citation…
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
arXiv:2502.14254v3 Announce Type: replace-cross Abstract: Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made…
little m: An AI Agent for Industrial Process Optimization
arXiv:2609.16680v2 Announce Type: replace Abstract: Manufacturing consumes one third of global energy and still has significant room for improvement in…
AI Persuasion as a Threat to Human Control
arXiv:2609.14796v2 Announce Type: replace Abstract: The threat that AI persuasion poses to human control has been acknowledged in the literature, but not…
Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
arXiv:2609.15293v2 Announce Type: replace Abstract: When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results…
Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?
arXiv:2609.16814v2 Announce Type: replace Abstract: While Large Language Model (LLM)-based Natural Language Inference (NLI) systems achieve high accuracy,…
Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
arXiv:2609.14708v2 Announce Type: replace Abstract: A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly…
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
arXiv:2609.08149v2 Announce Type: replace Abstract: SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on…
