arXiv:2609.02116v1 Announce Type: new Abstract: Reverse-logistics operators often decide how to inspect and route returned assets before their condition…
Sivers Commits $30M to Glasgow InP Laser Expansion for AI Datacenter Ramp
Sivers Semiconductors AB said on September 3, 2026, that it will invest USD 30 million to expand its Indium Phosphide (InP) manufacturing facility in…
EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision
arXiv:2609.02133v1 Announce Type: new Abstract: Empathetic response generation requires models to decide not only what to say, but also how to respond to…
MASkills: Continual Skills Optimization for Multi-Agent LLM Systems
arXiv:2609.02094v1 Announce Type: new Abstract: LLM-based multi-agent systems have shown strong performance on complex tasks, yet continual improvement…
CHIME: Credit-Aware Hierarchical Memory Evolution for Long-Horizon Agentic Planning
arXiv:2609.02074v1 Announce Type: new Abstract: Planning is a central capability that enables agents to decompose complex long-horizon tasks into…
ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
arXiv:2609.02067v1 Announce Type: new Abstract: Scientific benchmarks are commonly built by domain experts who write tasks and cross-check one another’s…
Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
Anthropic has released Claude Fable 5.1 and Claude Mythos 5.1, the same model behind two different safeguard layers. Fable 5.1 is generally available on…
READY or Not: Reliable Enterprise Agent Deployment
arXiv:2609.02095v1 Announce Type: new Abstract: An AI agent can perform well on benchmarks and still be unsuitable for deployment. Existing AI-agent…
Path to Astra: critical capabilities and frontier safeguards
Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework, with stronger safeguards for…
Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems
arXiv:2609.02092v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet…
AI News Brief Hourly Summary 2026-09-03 08h : 12 posts
12 posts published in the last hour 05:32DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents 05:32HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models 05:32MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity 05:32Anthropic’s Claude Fable 5.1 promises better…
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual…
HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
arXiv:2609.02029v1 Announce Type: new Abstract: Long-context inference retains a growing key–value (KV) cache during decoding, which consumes substantial…
MineTRACE: An Evidence-Grounded Interactive Reasoning System for Mineral Prospectivity
arXiv:2609.02060v1 Announce Type: new Abstract: Mineral exploration requires integrating heterogeneous geochemical, geophysical, and geological evidence,…
Anthropic’s Claude Fable 5.1 promises better coding and research at up to 45 percent less
Anthropic launches Claude Fable 5.1 and Mythos 5.1, its most capable AI models yet. Fable 5.1 doubles its predecessor’s score on Terminal-Bench-Science…
Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision
arXiv:2609.02057v1 Announce Type: new Abstract: Reliable web-agent monitoring is difficult when model-internal uncertainty signals such as token logits…
Anthropic’s new Fable release is cheaper, less restrictive
Fable 5.1 includes changes meant to reduce token cost and false-positive restrictions from the model’s safeguards.
ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
arXiv:2609.01992v1 Announce Type: new Abstract: Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from…
