arXiv:2608.23643v1 Announce Type: new Abstract: Artificial intelligence is increasingly being introduced into healthcare workflows, yet most evaluations…
Category: AI
AI Agents Push Humans Out of the Loop
arXiv:2608.23642v1 Announce Type: new Abstract: AI agents pose significant risks as they are granted increasing autonomy. A commonly proposed solution is…
OpenAI is building AI agents for everything. Will everyone use them?
Inside the frontier lab’s push to bring AI agents from software engineers to the masses.
How much of a measured AI preference is the model, and how much is the instrument?
arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit…
SpaceXAI Puts NVIDIA’s Vera CPU at the Center of Gigawatt-Scale Buildout
SpaceXAI will deploy NVIDIA’s Vera CPUs to run the CPU-intensive work behind its next generation of agentic AI applications, expanding an AI…
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
arXiv:2608.23640v1 Announce Type: new Abstract: When a large language model (LLM) is asked to write a person’s life, how much of what it writes actually…
Valor, Point72 back General Intuition at $6B valuation as AI startup pushes into robotics
General Intuition, the startup building a foundation model that trains generalized AI agents how to move through space and time, is in talks to raise at a…
Function-Level Execution Feedback for Code Preference Optimization
arXiv:2608.23632v1 Announce Type: new Abstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed…
ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence
arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on…
LLM Agents Perform Controlled Experiments Using Simulation Models
arXiv:2608.23622v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many…
