6 posts published in the last hour 19:33Do Dynamic Routers Need Memory? HeRo: History-Aware Routing for Efficient LLM Inference 19:33A Better Spur Should Start From Each Objective 19:33TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs 19:33Qiushi Engine…
Do Dynamic Routers Need Memory? HeRo: History-Aware Routing for Efficient LLM Inference
arXiv:2609.08189v1 Announce Type: new Abstract: Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip…
A Better Spur Should Start From Each Objective
arXiv:2609.08211v1 Announce Type: new Abstract: Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward…
TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs
arXiv:2609.08226v1 Announce Type: new Abstract: Temporal graph learning models the evolution of dynamic systems, where both structural interactions and…
Qiushi Engine on AstaBench E2E-Bench-Hard
arXiv:2609.08196v1 Announce Type: new Abstract: This report analyzes Qiushi Engine v0.8 across all 40 test tasks in AstaBench E2E-Bench-Hard, a benchmark…
Vision: Data-Centric Anchoring for Robust and Interpretable Agentic AI
arXiv:2609.08216v1 Announce Type: new Abstract: Agentic AI systems built on large language models fail in two persistent ways that scaling does not fix:…
AI News Brief Hourly Summary 2026-09-10 21h : 21 posts
21 posts published in the last hour 18:34Safe Harness Self-Evolution: A Theoretical Analysis of Feasibility and Limits 18:34Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment 18:34OpenAI Launches ChatGPT for Financial Services With Built-In Data 18:34Does Deeper Reasoning Compromise…
Safe Harness Self-Evolution: A Theoretical Analysis of Feasibility and Limits
arXiv:2609.08175v1 Announce Type: new Abstract: Harness self-evolution is the process by which an agent modifies its prompts, tools, code, or…
Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment
arXiv:2609.08188v1 Announce Type: new Abstract: Vision-language models (VLMs) augmented with retrieval-augmented generation (RAG) benefit from access to…
OpenAI Launches ChatGPT for Financial Services With Built-In Data
OpenAI on September 10, 2026 launched ChatGPT for Financial Services, a tailored ChatGPT Work experience that pairs built-in financial data with the…
Does Deeper Reasoning Compromise Alignment? Revealing and Mitigating of Alignment Collapse in Large Reasoning Models
arXiv:2609.08186v1 Announce Type: new Abstract: The emergence of Chain-of-Thought (CoT) has established a robust foundation for Large Reasoning Models…
Introducing ChatGPT for Financial Services
Introducing ChatGPT for Financial Services, combining built-in financial data and GPT-6 Astra for research, modeling, and client-ready materials.
Less Is Personal: Learning Minimal Sufficient User Profiles for Personalized Language Models
arXiv:2609.08180v1 Announce Type: new Abstract: Retrieval-augmented personalization enables large language models to produce more accurate and…
Amazon Quick is now generally available on desktop
Your teams get an AI assistant that handles real work while your data stays in your environment and your conversations stay private Today, the Amazon…
OntologyBench: Can Dense Retrieval Satisfy Structured Biomedical Constraints?
arXiv:2609.08174v1 Announce Type: new Abstract: We introduce OntologyBench, a tiered biomedical retrieval benchmark comprising 471,854 training and…
WorldAgen: Unified State-Action Prediction with Test-Time World Model Training
arXiv:2609.08162v1 Announce Type: new Abstract: How can vision-language-action (VLA) models adapt to new environments where world dynamics shift? While…
Anthropic reveals rogue AI agents hate CAPTCHAs, just like you
Come inside the mind of a bot trying to convince the internet it’s human.
OpenAI’s GPT-Live-1 API lets developers build apps that talk and listen at the same time
OpenAI releases GPT-Live-1 as a developer API. The full-duplex speech model scores 80.1 percent in interactivity tests, up from 45.4 percent for its…
