arXiv:2505.15276v2 Announce Type: replace Abstract: Large reasoning models (LRMs) have achieved remarkable success on complex tasks, yet their tendency to…
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy
arXiv:2505.21907v3 Announce Type: replace Abstract: AI copilots represent a new generation of AI-powered systems designed to assist users, particularly…
Accel reportedly in talks to lead $1B round for Thinking Machines at $40B valuation
The high-profile startup’s annual revenue run rate stands at at over $100 million.
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
arXiv:2609.02849v1 Announce Type: cross Abstract: Competitive programming has become a key test of large language model reasoning, with international…
Anthropic Released Claude Commerce Agents: An Apache-2.0 Blueprint for Shopping and Merchant Agents Across Retail, Travel, Telecom and Entertainment
Most teams building a shopping assistant or agent rebuild the same scaffolding: an agent loop, a tool layer over the catalog, an approval gate, and an…
AI Mathematician: Towards Fully Automated Frontier Mathematical Research
arXiv:2505.22451v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have made significant progress in mathematical capabilities in recent…
AI News Brief Hourly Summary 2026-09-03 22h : 16 posts
16 posts published in the last hour 19:32From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution 19:32HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design 19:32Untangling the Mechanisms of Misleading Context…
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
arXiv:2609.02771v1 Announce Type: cross Abstract: Training data attribution (TDA) aims to identify training examples that shape model behavior, but its…
HiPoly: a hierarchical polymer-native AI framework for property prediction and generative design
arXiv:2609.02746v1 Announce Type: cross Abstract: Polymeric materials are central to modern technologies, with applications ranging from energy to health…
Untangling the Mechanisms of Misleading Context in Medical Question Answering
arXiv:2609.02754v1 Announce Type: cross Abstract: Large language models now answer medical questions with expert-level performance. However, the context…
GPT-6 Astra is the first model making OpenAI willing to declare the “AGI era”
OpenAI has released GPT-6 Astra, its most capable model yet. President Greg Brockman says it marks the start of the “AGI era.” Astra tops benchmarks in…
frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study
arXiv:2609.02804v1 Announce Type: cross Abstract: For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its…
AI Efficiency Could Cost Us the Next Generation of Experts
A little over a decade ago, I led the controls design for a first-of-its-kind full digital-control system for a U.S. nuclear plant. It was, on paper, a…
Dutch Books for Language Models
arXiv:2609.02797v1 Announce Type: cross Abstract: People increasingly use language models to support life decisions. Many such decisions involve a…
Language Models Can Control Their Own Attention
arXiv:2609.02737v1 Announce Type: cross Abstract: Language models spend most of their attention on a small fraction of context, yet they read the entire…
RVSD: Retrieval Vision Sparse Decoding for Mitigating Visual Hallucinations in Large Vision-Language Models
arXiv:2609.02731v1 Announce Type: cross Abstract: Large vision-language models have achieved remarkable success in vision-language tasks. However, they…
Meta AI Released Muse Spark 1.3: An Agentic Coding Model That Uses ~20% Fewer Tool Calls and ~25% Fewer Tokens Than Muse Spark 1.2
Perplexity has shipped hybrid compute for its Mac app, splitting a single Perplexity Computer task between frontier models in the cloud and a compact…
TaRA: Training-Aware Low-Rank Adaptation Initialization
arXiv:2609.02639v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become a de facto standard for parameter-efficient fine-tuning (PEFT),…
