arXiv:2609.19848v1 Announce Type: new Abstract: Point-feature label placement on interactive maps must reconcile geometric validity, display yield, local…
Reproducibility is not construct validity: LLM measurement of institutionally situated communication
arXiv:2609.19866v1 Announce Type: new Abstract: High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the…
Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization
arXiv:2609.19830v1 Announce Type: new Abstract: Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment…
Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents
arXiv:2609.19843v1 Announce Type: new Abstract: LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with…
Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy
arXiv:2609.19820v1 Announce Type: new Abstract: Regularized self-play — the family behind DeepNash’s Stratego play — drives a two-player zero-sum policy…
Rethinking Multi-Agent Collaboration: When More Is Less
arXiv:2609.19759v1 Announce Type: new Abstract: The rapid advancement of large language models and single-agent harnesses has reshaped the landscape of…
TorchCraft: Unified binder design by inverting an all-atom structure predictor
arXiv:2609.19770v1 Announce Type: new Abstract: All-atom structure predictors model diverse molecular interactions, but using their learned structural…
Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems
arXiv:2609.19789v1 Announce Type: new Abstract: Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative…
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data
Paper2Agent, published in Nature, converts papers into validated MCP tools, scoring 91.2% on 300 questions across 74 papers.
Integrating knowledge from case reports: a medical ontology based multimodal information system with structured summary
arXiv:2609.19775v1 Announce Type: new Abstract: Published medical case reports serve as a crucial medical information carrier, documenting discoveries in…
AI News Brief Hourly Summary 2026-09-18 09h : 12 posts
12 posts published in the last hour 06:33Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions 06:33AutoData: Agentic Search for Pre-training Data Selection 06:33When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning…
Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions
arXiv:2609.19654v1 Announce Type: new Abstract: Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight…
AutoData: Agentic Search for Pre-training Data Selection
arXiv:2609.19754v1 Announce Type: new Abstract: LLM agents have recently shown promise in automating machine learning engineering by editing model and…
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
arXiv:2609.19671v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks but exhibit systematic…
LearnActCoder: Role-Aware Error Memory for Adaptive Clinical Coding Agents
arXiv:2609.19721v1 Announce Type: new Abstract: Clinical coding agents repeatedly encounter the same failure modes, including unsupported codes, missed…
Knowledgator Releases GLiFormer: A 575M-Parameter Encoder That Hits 91.10 F1 on Nested JSON Extraction Without Generating Tokens
GLiFormer Large scores 91.10 F1 on nested JSON, near GPT-5.6-luna’s 91.96, while grounding every value in source spans.
FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA
arXiv:2609.19680v1 Announce Type: new Abstract: Financial QA systems are typically improved before deployment through better retrieval, prompting, or…
