arXiv:2608.26372v1 Announce Type: cross Abstract: Large language models are increasingly deployed as autonomous agents serving users on behalf of…
Author: script
AI News Brief Hourly Summary 2026-08-28 19h : 14 posts
14 posts published in the last hour 16:33MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models 16:33How Unlikely Is “Unlikely”? Assessing Verbal Probability Perception Across Large Language Models 16:33On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY…
MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
arXiv:2608.26295v1 Announce Type: cross Abstract: Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their…
How Unlikely Is “Unlikely”? Assessing Verbal Probability Perception Across Large Language Models
arXiv:2608.26327v1 Announce Type: cross Abstract: Large language models increasingly produce and interpret verbal probability expressions, yet whether…
On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study
arXiv:2608.26292v1 Announce Type: cross Abstract: Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query,…
How Decathlon runs demand forecasting at scale with Chronos-2
Decathlon, one of the world’s largest sporting goods retailers, forecasts weekly demand for tens of thousands of products across multiple continents.…
Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
arXiv:2608.26317v1 Announce Type: cross Abstract: Frontier language models are increasingly marketed as omni systems that can perceive and respond across…
Spreading the load: How Salesforce met Multi-AZ HA with SageMaker Inference Components
Learn how Salesforce used Amazon SageMaker AI Inference Component placement (the SchedulingConfig parameter) to distribute model copies across multiple…
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation
arXiv:2608.26237v1 Announce Type: cross Abstract: Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of…
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
arXiv:2608.26222v1 Announce Type: cross Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust…
