Cognition announced on September 10, 2026, that Jonathan Kelley and the entire Dioxus team are joining the company to work on Devin, its autonomous…
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
arXiv:2609.08149v1 Announce Type: new Abstract: SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on…
OpenAI’s GPT-Live-1 Arrives in the API at $0.05 Per Minute
OpenAI launched GPT-Live-1 in the API on September 10, 2026, making its full-duplex voice model available to developers at $0.05 per minute for the…
Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering
arXiv:2609.08173v1 Announce Type: new Abstract: Sparse autoencoder (SAE)-based steering has been widely used to address knowledge conflicts by guiding…
AI News Brief Hourly Summary 2026-09-10 20h : 18 posts
18 posts published in the last hour 17:33Build more natural voice experiences with GPT‑Live‑1 in the API 17:33RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts 17:33A Candid Abacus AI Review:…
Build more natural voice experiences with GPT‑Live‑1 in the API
GPT‑Live‑1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.
RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts
arXiv:2609.08090v2 Announce Type: new Abstract: Assistive devices for people with mobility impairments, such as powered exoskeletons, rely on accurate…
A Candid Abacus AI Review: The All-in-One AI Platform for Professionals & Enterprises
If you’re paying for ChatGPT, Claude, and another AI tool simultaneously, this review is for you. It covers what an AI platform like Abacus AI actually…
CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information
arXiv:2609.08094v1 Announce Type: new Abstract: Large Language Models are increasingly deployed in public-sector settings, where incorrect guidance can…
Pony.ai Starts Fully Driverless Robotaxi Passenger Tests in Zagreb
Pony AI Inc. and Verne announced on September 10, 2026, the start of fully driverless robotaxi test rides carrying invited passengers on public roads in…
Automated Design of Inventory Policy with Large Language Models: An Exploratory Study
arXiv:2609.08071v1 Announce Type: new Abstract: Firms making inventory decisions have access to operational data, optimization tools, and large language…
T. Rowe Price Expands Claude Across Investment Teams and Developers
T. Rowe Price and Anthropic announced on September 10, 2026, the expansion of Claude across the asset manager’s investment organization. Portfolio…
Inference-Time Nash Alignment
arXiv:2609.08082v1 Announce Type: new Abstract: Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large preference…
Val Bercovici, Chief AI Officer at WEKA – Interview Series
Val Bercovici, Chief AI Officer at WEKA, is an AI and data infrastructure executive focused on advancing the technologies that underpin next-generation…
Artificial Intelligence-Assisted Digital Inventory of Cultural Heritage & Traditional Knowledge: Case for Indonesian Open Digital Library of Culture
arXiv:2609.08105v1 Announce Type: new Abstract: The Indonesian Digital Library of Culture (Perpustakaan Digital Budaya Indonesia, PDBI;…
A Layered Analysis of Disagreement And Answer Quality in Multi-Agent LLM Debate
arXiv:2609.08016v1 Announce Type: new Abstract: Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to…
From Version Conflicts to Decision Conflicts: Selective Revalidation for Long-Running AI Agents
arXiv:2609.08015v1 Announce Type: new Abstract: Long-running AI agents may read state, reason, wait for tools or human approval, and perform an external…
Eliciting Self-Verification in Multimodal Reasoning Agents with Reinforcement Learning
arXiv:2609.08025v1 Announce Type: new Abstract: Reasoning agents increasingly rely on external tools such as web search to answer complex queries.…
