arXiv:2608.26346v1 Announce Type: cross Abstract: We investigate whether AI-generated impulsive sounds can be distinguished from real ones through group…
Category: AI
Co-Evolving Structured Knowledge and Reasoning in Language Models
arXiv:2608.26386v1 Announce Type: cross Abstract: Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge,…
CG4AI: A Column Generation Framework for Training AI Models Under Constraints
arXiv:2608.26375v1 Announce Type: cross Abstract: Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that…
Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries
arXiv:2608.26385v1 Announce Type: cross Abstract: Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that…
Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
arXiv:2608.26372v1 Announce Type: cross Abstract: Large language models are increasingly deployed as autonomous agents serving users on behalf of…
MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
arXiv:2608.26295v1 Announce Type: cross Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their…
How Unlikely Is “Unlikely”? Assessing Verbal Probability Perception Across Large Language Models
arXiv:2608.26327v1 Announce Type: cross Abstract: Large language models increasingly produce and interpret verbal probability expressions, yet whether…
On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study
arXiv:2608.26292v1 Announce Type: cross Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query,…
How Decathlon runs demand forecasting at scale with Chronos-2
Decathlon, one of the world’s largest sporting goods retailers, forecasts weekly demand for tens of thousands of products across multiple continents.…
Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
arXiv:2608.26317v1 Announce Type: cross Abstract: Frontier language models are increasingly marketed as omni systems that can perceive and respond across…
