arXiv:2609.00455v1 Announce Type: new Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in…
Category: AI
Anthropic Introduces Enterprise Frontier Safeguards (EFS): Zero-Data-Retention Privacy Plus Cross-Session Misuse Detection
Anthropic announced Enterprise Frontier Safeguards on September 1, 2026, an architecture that stores monitoring data in the customer’s own cloud account…
Validity-Aware Jailbreak Evaluation for Large Language Models
arXiv:2609.00498v1 Announce Type: new Abstract: Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing…
SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation
arXiv:2609.00427v1 Announce Type: new Abstract: The exponential growth of wireless devices is driving unprecedented spectrum demand, pushing spectrum…
Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations
arXiv:2609.00441v1 Announce Type: new Abstract: Effective manager-employee communication is critical for retaining high performers and developing…
SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents
arXiv:2609.00434v1 Announce Type: new Abstract: Evaluating task-oriented dialogue agents requires judging not merely whether a reply reads well but…
mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers
arXiv:2609.00453v1 Announce Type: new Abstract: Giving an agent a file about a named expert can supply hard-to-find material, produce a recognizable…
Dependency-Aware Chain-of-Thought Compression for Financial Reasoning
arXiv:2609.00413v1 Announce Type: new Abstract: Chain of thought prompting improves complex reasoning, but its long intermediate traces create substantial…
RestoreBench: Can AI Agents Restore Power Flow Convergence?
arXiv:2609.00384v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly automate multi-step engineering workflows through tool use,…
SlideBank: A Persistent Hierarchical Evidence Bank for Consistent Whole-Slide Reasoning
arXiv:2609.00342v1 Announce Type: new Abstract: Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant…
