arXiv:2609.09702v1 Announce Type: new Abstract: Candidate decision correctness and rationale grounding are different objectives. We examine…
Author: script
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though…
Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
arXiv:2609.09678v1 Announce Type: new Abstract: Clinical diagnosis agents must decide not only what test to request next, but also when to diagnose or…
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
arXiv:2609.09754v1 Announce Type: new Abstract: As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic…
Google Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation
Google Research has released ToolGrad, an ACL 2026 Findings framework that inverts tool-use dataset generation: it builds a verified API chain first, then…
Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety
arXiv:2609.09735v1 Announce Type: new Abstract: Healthcare systems, mental health, and public well-being are increasingly affected by cyberbullying and…
AI News Brief Hourly Summary 2026-09-11 08h : 10 posts
10 posts published in the last hour 05:32RobustSGPO: Search-Space Control for Agent Harness Evolution 05:32Seven Sources of Physical AI Capability Formation 05:32PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations 05:32RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems 05:32Black-Box…
RobustSGPO: Search-Space Control for Agent Harness Evolution
arXiv:2609.09646v1 Announce Type: new Abstract: Semantic-gradient-based prompt optimization (SGPO) improves agent harnesses using execution feedback, but…
Seven Sources of Physical AI Capability Formation
arXiv:2609.09627v1 Announce Type: new Abstract: Capabilities relevant to Physical AI can arise from materially different formation histories, yet existing…
PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations
arXiv:2609.09664v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as personalized assistants that interact with users…
