arXiv:2608.26375v1 Announce Type: cross Abstract: Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that…
Category: cs.AI updates on arXiv.org
Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries
arXiv:2608.26385v1 Announce Type: cross Abstract: Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that…
Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
arXiv:2608.26372v1 Announce Type: cross Abstract: Large language models are increasingly deployed as autonomous agents serving users on behalf of…
MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
arXiv:2608.26295v1 Announce Type: cross Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their…
How Unlikely Is “Unlikely”? Assessing Verbal Probability Perception Across Large Language Models
arXiv:2608.26327v1 Announce Type: cross Abstract: Large language models increasingly produce and interpret verbal probability expressions, yet whether…
On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study
arXiv:2608.26292v1 Announce Type: cross Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query,…
Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
arXiv:2608.26317v1 Announce Type: cross Abstract: Frontier language models are increasingly marketed as omni systems that can perceive and respond across…
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation
arXiv:2608.26237v1 Announce Type: cross Abstract: Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of…
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
arXiv:2608.26222v1 Announce Type: cross Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust…
Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness Bugs
arXiv:2608.26209v1 Announce Type: cross Abstract: Data-driven software systems are increasingly deployed in high-stakes socio-economic domains, from…
