arXiv:2609.10928v1 Announce Type: cross Abstract: Maximizing the area under the receiver operating characteristic curve (AUC) is a standard approach to…
Category: AI
GPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends
Overly long skill descriptions, blanket reading requirements, and rigid approval rules can get in GPT-6 Astra’s way, warns OpenAI’s Eric Provencher. More…
Does Linguistic Structure Enrichment Enhance Coherence Assessment? Not With Current Architectures
arXiv:2609.10893v1 Announce Type: cross Abstract: Recent advances in large language models have transformed human-computer interaction. Despite their…
No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vulnerabilities in MCP Servers
arXiv:2609.10854v1 Announce Type: cross Abstract: Conventional vulnerability analysis relies on either system access or dynamic interaction, all of which…
Counterfactual Marginalisation: Framework for Evaluating Robustness to Nuisance Variables
arXiv:2609.10778v1 Announce Type: cross Abstract: Machine learning models can achieve strong test performance while relying on demographic or…
Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble
arXiv:2609.10883v1 Announce Type: cross Abstract: Language models are trained to implement a helpful AI Assistant character (e.g., Claude). We explore how…
Are We Really Doing Few-Shot Learning? A Critical Examination of Pre-Training Assumptions
arXiv:2609.10851v1 Announce Type: cross Abstract: Few-shot learning is commonly evaluated under protocols that pre-train a model on a large auxiliary set…
Tapes Together Strong: The Co-evolution of Computation and Cooperation
arXiv:2609.10817v1 Announce Type: cross Abstract: How does cooperation evolve in complex agentic systems? Prior work in evolutionary game theory studies…
Temporal and Multimodal Deep Learning for Cyberattack Detection in LEO Satellite Systems
arXiv:2609.10746v1 Announce Type: cross Abstract: The growing reliance on Low-Earth Orbit (LEO) satellite communication systems has increased the need for…
When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents
arXiv:2609.10750v1 Announce Type: cross Abstract: LLM agents increasingly rely on external skills retrieved at runtime, making skill selection from large…
