OpenAI can disclose misalignment before fixes exist. Its 6 initial reports include fabricated data and leaked API keys.
Category: AI
Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum
arXiv:2609.18283v1 Announce Type: new Abstract: As telecommunication networks evolve toward autonomous 5G-Advanced and 6G operations, agentic artificial…
WFM: Wiki Foundation Model for Complex Agentic Reasoning
arXiv:2609.18182v1 Announce Type: new Abstract: Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e.,…
Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost
arXiv:2609.18126v1 Announce Type: new Abstract: Agentic AI systems often approach the same task through multiple workflows that differ in reasoning…
Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment
arXiv:2609.18249v1 Announce Type: new Abstract: Real-world recommendation scenarios are commonly grounded in shared physical environments during…
Time-Aligned Evolving Concept Graphs for Scientific Relation Forecasting
arXiv:2609.18163v1 Announce Type: new Abstract: Forecasting scientific relations can guide discovery by identifying promising connections before they…
Symbolic Temporal Supervision of LLM Agents Using Contracts
arXiv:2609.18128v1 Announce Type: new Abstract: Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web…
When Is Graph Structure Worth Its Cost? The Case for Structure Pricing in Retrieval-Augmented Generation
arXiv:2609.18099v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (RAG) can help answer questions that require information from…
AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
arXiv:2609.18123v1 Announce Type: new Abstract: Large language model agents tune GPU kernels and serving engines through a closed loop of propose,…
The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
arXiv:2609.18063v1 Announce Type: new Abstract: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is…
