arXiv:2609.17499v1 Announce Type: cross Abstract: Uncertainty estimation for Vision-Language-Navigation (VLN) models is a critical task since it can help…
When Should LLMs Abstain? Chain-of-Self-Questioning for Selective Risk Control
arXiv:2609.17516v1 Announce Type: cross Abstract: Large language models can produce fluent answers when their factual support is weak. This paper…
OpenAI Launches Misalignment Reporting Framework With Six Incident Reports
OpenAI published a framework for tracking, investigating, and disclosing instances of model misalignment on September 16, 2026, alongside six reports on…
PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
arXiv:2609.17521v1 Announce Type: cross Abstract: Interactive control for video generation is moving from coarse prompts toward fine-grained, physically…
AI News Brief Hourly Summary 2026-09-17 01h : 13 posts
13 posts published in the last hour 22:32Decomposition Buys Integrity, Not Yield 22:32Det-LIME: Detector-Aware, Multi-Instance Local Interpretable Model-Agnostic Explanations for Automated Marine Mammal Detection 22:32Evaluating Verified Autonomy in Quantum Engineering 22:32CareMirror: Bringing Caregiver Wellbeing into the Dementia Care Ecosystem 22:32Our…
Decomposition Buys Integrity, Not Yield
arXiv:2609.17464v1 Announce Type: cross Abstract: Multi-agent systems split a task across a tree of agents and justify the split with folklore: smaller…
Det-LIME: Detector-Aware, Multi-Instance Local Interpretable Model-Agnostic Explanations for Automated Marine Mammal Detection
arXiv:2609.17479v1 Announce Type: cross Abstract: Despite the rapid uptake of black-box object detectors in marine mammal research and monitoring,…
Evaluating Verified Autonomy in Quantum Engineering
arXiv:2609.17439v1 Announce Type: cross Abstract: Reliable quantum engineering is essential for turning quantum phenomena into practical technologies. As…
CareMirror: Bringing Caregiver Wellbeing into the Dementia Care Ecosystem
arXiv:2609.17434v1 Announce Type: cross Abstract: Family caregivers of people living with dementia shoulder emotional and practical responsibilities, yet…
Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback
arXiv:2609.17474v1 Announce Type: cross Abstract: Large language model (LLM) distillation aims to transfer the capabilities of a powerful teacher to a…
Where Should a Document Live: Context, Representations, or Parameters?
arXiv:2609.17346v1 Announce Type: cross Abstract: To answer questions outside of their pre-training data, large language models (LLMs) need access to new…
Learning-Guided Planning in Large Dynamic Action Spaces: Budgeted Tree Search for One-to-Many Mobile Charging
arXiv:2609.17429v1 Announce Type: cross Abstract: Many learned sequential decision systems map the current state directly to an action. That shortcut…
Tracking the Unseen: An Occlusion-Robust Framework for Target Tracking Under Full and Long-Term Occlusion
arXiv:2609.17427v1 Announce Type: cross Abstract: Real-time multi-object tracking systems remain highly vulnerable to full and long-term occlusion, where…
CTAN: Cycle-Temporal Attention Network for Embodied Audio-Visual Navigation
arXiv:2609.17420v1 Announce Type: cross Abstract: Audio-visual embodied navigation equips robots with the capability to infer the locations of sound…
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data
Paper2Agent, published in Nature, converts papers into validated MCP tools, scoring 91.2% on 300 questions across 74 papers.
Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead
arXiv:2609.17394v1 Announce Type: cross Abstract: Small differences on coding-agent leaderboards are often read as an ordering of systems. We audit…
AI News Brief Hourly Summary 2026-09-17 00h : 16 posts
16 posts published in the last hour 21:58AI News Brief Roundup: 2026-09-16 21:57AI News Brief Daily Summary 2026-09-16 21:32Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record…
