arXiv:2609.18366v2 Announce Type: replace Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a…
Advisory Group on Mathematics and Artificial Intelligence
OpenAI is working with an independent Advisory Group on Mathematics and Artificial Intelligence to guide the review and communication of emerging AI…
Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning
arXiv:2609.18461v2 Announce Type: replace Abstract: Personalized agents are required to reason over long-term history interactions to infer both explicit…
Prices go up in 7 days — get your Disrupt ticket now
Current ticket pricing ends September 25 at 11:59 p.m. PT. Join 10,000+ founders, investors, and tech leaders at Disrupt and save up to $200 on your…
A Unified Evaluation Framework for Trustworthy Large Language Models, Agentic AI, and Multimodal Systems
arXiv:2609.19524v2 Announce Type: replace Abstract: Benchmark scores alone provide an incomplete basis for assessing the trustworthiness of modern…
AI News Brief Hourly Summary 2026-09-21 20h : 19 posts
19 posts published in the last hour 17:33Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings 17:33A visual large language foundational model for medical image recognition using clinician-contributed online resources 17:33Fraglingo: Molecular Design via Attachment-Aware…
Planetary Prediction Engine: Autonomous Geospatial Prediction via Intelligent Data Selection and Foundation Model Embeddings
arXiv:2608.26088v2 Announce Type: replace Abstract: Addressing critical global challenges, from food security and disaster risk to disease outbreaks and…
A visual large language foundational model for medical image recognition using clinician-contributed online resources
arXiv:2609.06914v3 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing…
Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment Generation
arXiv:2609.13519v2 Announce Type: replace Abstract: We introduce Fraglingo, an autoregressive molecular generator that constructs molecules step by step…
Building standards for the next phase of AI
OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety.
Balance of Benchmarks: Semantic Density Reweighting for Task-Conditioned Model Comparison
arXiv:2608.30044v3 Announce Type: replace Abstract: Model comparison increasingly relies on large collections of publicly reported benchmark scores, yet…
Expanding OpenAI Academy with new learning paths
Explore new OpenAI Academy learning paths for employees, developers, leaders, educators, and students to build and demonstrate practical AI skills.
MOSCOPT: Mixture-of-Skills Collective Optimization for LLM Agents
arXiv:2609.14399v2 Announce Type: replace Abstract: Natural language prompts and skills serve as the strategic backbone of LLM-based agents. Recent…
Intent-Governed Tool Authorization for AI Agents
arXiv:2606.22916v4 Announce Type: replace Abstract: Tool-using AI agents commonly operate under integration credentials whose static permissions exceed a…
xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6
xAI has released Grok 4.7, its most capable model yet. But on the Artificial Analysis Intelligence Index, it scores just 46 points, landing mid-pack and…
Run Positron on Amazon SageMaker AI for data science workflows
Positron, Posit’s IDE for data science, now runs on Amazon SageMaker AI. This post shows how a data scientist explores an Amazon Athena table, validates…
Proposed EU Scheme Would Auto-Generate Sustainability Labels From 2027
The European Commission on September 21, 2026 proposed a common Union rating scheme that will assign data centers with a capacity above 500 kW…
BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
arXiv:2605.09134v4 Announce Type: replace Abstract: Reinforcement learning for program repair is hindered by sparse execution feedback and coarse…
