arXiv:2609.03450v1 Announce Type: cross Abstract: An agent that inherits six one-line memories may pull at most one archived source record before acting;…
Author: script
OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits
According to an analysis by collusion.wiki, autonomous AI agents that identified themselves as OpenAI systems left roughly 18,000 posts in a 25-year-old…
When Users Don’t Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents
arXiv:2609.03467v1 Announce Type: cross Abstract: Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents,…
TabScope: Question-Adaptive Scope Selection for Table Question Answering
arXiv:2609.03395v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown strong performance on table question answering, yet their…
The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems
arXiv:2609.03425v1 Announce Type: cross Abstract: Humans are the transport layer between AI systems, losing context at every hop. We present the…
Spectral Convergence of Random Feature Method in Multiple Dimensions
arXiv:2609.03401v1 Announce Type: cross Abstract: We first prove spectral convergence of the random feature method (RFM) for multidimensional targets in…
Privacy, Robustness, and Fairness Trade-offs in Federated Intrusion Detection: Geometric Indistinguishability at the Aggregation Interface
arXiv:2609.03420v1 Announce Type: cross Abstract: Federated learning enables privacy-conscious collaboration for network intrusion detection without…
StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios
arXiv:2609.03414v1 Announce Type: cross Abstract: Audio enhancement in real-world scenarios involves complex distortion couplings and requires…
AI News Brief Hourly Summary 2026-09-04 15h : 13 posts
13 posts published in the last hour 12:33SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking 12:33ObserverBench: Testing Mechanistic Estimates for Intervention and Control 12:33FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience 12:33Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation…
SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking
arXiv:2609.03047v1 Announce Type: cross Abstract: Libraries and archives manage large collections with limited staff and computing budgets, yet common…
