arXiv:2608.24076v1 Announce Type: new Abstract: Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users,…
Author: script
Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagnosis, and Cost-Aware Routing
arXiv:2608.24112v1 Announce Type: new Abstract: Modern text-to-image (T2I) models often have similar total scores but different strengths, making…
ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation
arXiv:2608.24103v1 Announce Type: new Abstract: Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two…
Advancing price-performance for developers with GPT‑5.6 in Kiro
GPT‑5.6 is now available in Kiro, helping developers plan, build, review, and test software with better price-performance.
Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
arXiv:2608.24099v1 Announce Type: new Abstract: GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to…
AI News Brief Hourly Summary 2026-08-26 11h : 16 posts
16 posts published in the last hour 08:33Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment 08:33Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems 08:32Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval…
Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment
arXiv:2608.24046v1 Announce Type: new Abstract: When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of…
Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems
arXiv:2608.24069v1 Announce Type: new Abstract: LLM-based multi-agent trading systems, in which specialized agents collaborate through structured…
Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding
arXiv:2608.24024v1 Announce Type: new Abstract: Confidence-based voting aggregates parallel LLM rollouts by weighting each with internal signals such as…
Introducing new Ray capabilities on SageMaker HyperPod
Amazon SageMaker HyperPod now offers managed Ray support on Amazon EKS. Create and monitor Ray clusters, connect JupyterLab and Code Editor notebooks to…
