arXiv:2609.20779v1 Announce Type: cross Abstract: Safety evaluations for large language models rely on surface-form classifiers that report declining harm…
Category: AI
Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights
arXiv:2609.20768v1 Announce Type: cross Abstract: Generative agents are increasingly used to select and narrate video highlights, but they typically…
A shared agentic platform for Wood Mackenzie, on Amazon Bedrock AgentCore
Wood Mackenzie built APEX, a shared agentic AI platform on Amazon Bedrock AgentCore so every team can ship production agents without rebuilding runtime,…
Quantifying Overclaiming Propensity in Frontier LLM Agents
arXiv:2609.20812v1 Announce Type: cross Abstract: Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent’s…
How MRH Trowe enabled secure self-service AI agents in financial services
Learn how MRH Trowe, one of Germany’s leading commercial and industrial insurance brokers, gave about 400 employees secure, self-service access to AI…
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
arXiv:2609.20784v1 Announce Type: cross Abstract: Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per…
Selecting a vector store for Amazon Bedrock Knowledge Bases
Choosing the right vector store for your Amazon Bedrock Knowledge Bases RAG application affects performance and cost. This post compares Amazon OpenSearch…
GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies
arXiv:2609.20776v1 Announce Type: cross Abstract: Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA)…
Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation
arXiv:2609.20758v1 Announce Type: cross Abstract: Evaluating an AI system requires disaggregated assessment, as performance varies across domains such as…
Chronicle: Cut-Point Replay for Regression Testing of LLM Agents
arXiv:2609.20625v1 Announce Type: cross Abstract: Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a…
