Letting an AI do your shopping might not get you the best deal. Researchers at the Wharton School show how erratic AI shopping agents really are: a single…
Category: AI
Task-Adaptive Rubrics for GUI Reward Modeling
arXiv:2608.24174v1 Announce Type: new Abstract: Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome…
Gatik raises $200M to scale AI-powered autonomous freight
Autonomous trucking company Gatik has raised $200 million in Series D funding to expand its driverless freight operations across North America. The round…
Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping
arXiv:2608.24135v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has emerged as a pivotal technique for enhancing the…
Raised on AI
When my oldest child was born, I immediately set up Gmail and Twitter accounts in her name. I broadly announced her birth online and proceeded to plaster…
AHEAD: Adaptive Hindsight with Environment-Augmented Distillation for Agentic RL
arXiv:2608.24114v1 Announce Type: new Abstract: Training multi-turn LLM agents with reinforcement learning typically relies on trajectory-level rewards,…
Previewing the Model Hardware Standard
We’re opening a research preview of the Model Hardware Standard (MHS), a shared specification for AI agents to safely operate physical devices, to a first…
OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
arXiv:2608.24160v1 Announce Type: new Abstract: Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and…
OpenAI rallies 100+ companies to sign open letter warning AI-powered cyberattacks on critical infrastructure are imminent
OpenAI, together with more than 100 companies including Microsoft, Google, Anthropic, Deutsche Telekom, and SAP, has published an open letter on…
Paritok-4B: Intent-Conditioned Context Compression for Coding Agents
arXiv:2608.24188v1 Announce Type: new Abstract: Coding agents re-send large file reads and tool outputs to a frontier LLM every turn, and this context…
