Paper2Agent, published in Nature, converts papers into validated MCP tools, scoring 91.2% on 300 questions across 74 papers.
Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead
arXiv:2609.17394v1 Announce Type: cross Abstract: Small differences on coding-agent leaderboards are often read as an ordering of systems. We audit…
AI News Brief Hourly Summary 2026-09-17 00h : 16 posts
16 posts published in the last hour 21:58AI News Brief Roundup: 2026-09-16 21:57AI News Brief Daily Summary 2026-09-16 21:32Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record…
AI News Brief Daily Summary 2026-09-16
200 posts published today 21:32Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record 21:32Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs 21:32Knowledgator…
Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record
arXiv:2609.17226v1 Announce Type: cross Abstract: An agent that learns from rewards has to trust whatever reports those rewards. When the reports suddenly…
Vroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
arXiv:2609.17327v1 Announce Type: cross Abstract: This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying…
Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens
GLiFormer Large scores 91.10 F1 on nested JSON, near GPT-5.6-luna’s 91.96, while grounding every value in source spans.
Mo’ Models, Mo’ Problems: How to best select model pools when designing Multi-Agent Systems
arXiv:2609.17306v1 Announce Type: cross Abstract: Multi-agent Systems (MAS) combine multiple model outputs to solve complex reasoning tasks. However,…
Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecedented access, but warn meaningful…
After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind
arXiv:2609.17274v1 Announce Type: cross Abstract: AI agents increasingly act through agent skills, i.e., natural-language instructions, that direct a host…
5 Free Microsoft GitHub Courses to Learn Data Science and Artificial Intelligence
Explore five free Microsoft GitHub courses covering data science, machine learning, artificial intelligence, generative AI, LLMs, RAG, fine-tuning, and AI…
FROD: Feature Matching Residual Denoising Oracle Bone Decipher
arXiv:2609.17227v1 Announce Type: cross Abstract: Oracle bone script (OBS), one of the earliest Chinese writing systems, plays an important role in the…
MUMINS: Metadata-conditioned Uncertainty-aware Medical Image Next-state Synthesis
arXiv:2609.17169v1 Announce Type: cross Abstract: Forecasting anatomical changes such as tumor growth and neurodegeneration is a challenging generative…
A unified framework for global and local interpretability using adaptive derivative-ordered random explanation
arXiv:2609.17171v1 Announce Type: cross Abstract: The interpretability of complex machine learning models is of paramount importance, especially in…
FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence
arXiv:2609.17210v1 Announce Type: cross Abstract: Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning…
Grounding SWE-Agent Decisions in Architecture-0 Design: Navigating Unknown Unknowns through Physical Mapping
arXiv:2609.17221v1 Announce Type: cross Abstract: Autonomous Software Engineering Agents (SWE-Agents) excel in deterministic coding tasks but struggle…
Multimodal Cultural Heritage Architectural Style Classification for Residential Buildings in the UAE Based on CLIP Embeddings and SVM
arXiv:2609.17181v1 Announce Type: cross Abstract: The analysis and classification of cultural heritage architectural styles remain challenging due to the…
AI News Brief Hourly Summary 2026-09-16 23h : 14 posts
14 posts published in the last hour 20:32Kernel-Based Metrics Learning for Uncertain Opponent Vehicle Trajectory Prediction in Autonomous Racing 20:32ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers 20:32Continual Learning for Traversability Prediction with Uncertainty-Aware Adaptation 20:32After…
