arXiv:2609.07803v1 Announce Type: new Abstract: Model pruning is widely used to compress deep neural networks, reducing memory and computational…
Claude Fable 5.1’s language is less “load-bearing” than its predecessor’s
Arena.ai analyzed how Claude’s writing changed from Fable 5 to Fable 5.1 across tens of thousands of benchmark responses. Fable 5.1 writes more…
Do Large Language Models Know What They Don’t Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty
arXiv:2609.07879v1 Announce Type: new Abstract: Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question…
Now everyone can put data to work
Meet the Data agent in ChatGPT Work. Connect company data, uncover insights, and build interactive dashboards with AI using natural language.
When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems
arXiv:2609.07741v1 Announce Type: new Abstract: Persistent AI assistants are intended to extend human attention, memory, and coordination across changing…
Expanding AI access and cyber defense for federal, state, local, and tribal governments
OpenAI and GSA will offer eligible federal, state, local, and tribal governments $0 license fees, 50% off usage, and expanded cyber defense support.
What Does an LLM-Agent Leaderboard Rank Actually Compare?
arXiv:2609.07785v1 Announce Type: new Abstract: An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent.…
Aegix Pulse: A Traceable Three-Stage Architecture for Personalized Content Generation and Context-Preserving Revision
arXiv:2609.07672v1 Announce Type: new Abstract: Production content-generation systems must integrate a user’s immediate task, long-term brand identity,…
APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI Agents
arXiv:2609.07712v1 Announce Type: new Abstract: Mobile GUI agents can execute tasks from natural-language instructions, but their evaluation remains…
The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
arXiv:2609.07713v1 Announce Type: new Abstract: Generative and agentic AI are reshaping both the production and evaluation of scientific research. These…
The Profit Alignment Problem: How Profit Mandates Induce Alignment Failures in LLMs
arXiv:2609.07731v1 Announce Type: new Abstract: We show that ordinary business language — “maximize profitability” — induces profit-oriented ambiguity…
AI agents are flooding public services with new requests
“The vast majority of cases we find are people who are entitled to claim for something, claiming for that thing,” the researcher told TechCrunch.
A radiographic world model for clinical reasoning and evidence generation
arXiv:2609.07719v1 Announce Type: new Abstract: Medical imaging artificial intelligence (AI) is commonly developed as separate mappings from radiographs…
AI News Brief Hourly Summary 2026-09-10 17h : 13 posts
13 posts published in the last hour 14:34AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era 14:34Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best 14:34From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction 14:34FinCUABuild:…
AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era
arXiv:2609.07611v1 Announce Type: new Abstract: Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence,…
Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
arXiv:2609.07627v1 Announce Type: new Abstract: AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue…
From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction
arXiv:2609.07573v1 Announce Type: new Abstract: Multi-agent LLM deliberation has been explored as a scalable way to simulate public deliberation. For such…
FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?
arXiv:2609.07603v1 Announce Type: new Abstract: Financial scenarios are diverse and complex, spanning varying data conditions, tool configurations, and…
