arXiv:2512.13979v2 Announce Type: replace Abstract: Large reasoning models achieve strong performance on diverse tasks by producing extended chains of…
Tag: AI
OpenAI’s first custom chip “Jalapeño” reportedly beats Nvidia’s Blackwell and Rubin in inference benchmarks
OpenAI showed off “Jalapeño,” its first in-house inference chip, with benchmarks at the Hot Chips conference. According to SemiAnalysis tests, the chip…
Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models
arXiv:2602.02304v3 Announce Type: replace Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as…
Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light
arXiv:2507.11482v5 Announce Type: replace Abstract: Artificial learning systems are graduating from passive learners to increasingly autonomous agents,…
Efficient LLM Collaboration via Planning
arXiv:2506.11578v5 Announce Type: replace Abstract: Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to…
LEMMA-RCA: A Large Multi-modal Multi-domain Dataset for Root Cause Analysis
arXiv:2406.05375v4 Announce Type: replace Abstract: Root cause analysis (RCA) is crucial for enhancing the reliability and performance of complex systems.…
Introducing the Admin plugin for ChatGPT Work and Codex
Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.
Olapa-MCoT: Enhancing the Chinese Mathematical Reasoning Capability of LLMs
arXiv:2312.17535v2 Announce Type: replace Abstract: In the past two years, the outstanding performance of ChatGPT in multilingual and multitasking has led…
Claude Cowork finally remembers what you told the app in chat
Anthropic is giving Claude a shared memory across chat and Cowork, so users no longer have to repeatedly brief the AI on projects, preferences, and other…
Adaptive GR(1) Specification Repair for Liveness-Preserving Shielding in Reinforcement Learning
arXiv:2511.02605v3 Announce Type: replace Abstract: Shielding is widely used to enforce safety in reinforcement learning (RL), ensuring that an agent’s…
