arXiv:2608.29460v2 Announce Type: replace Abstract: When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or…
Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling
arXiv:2608.29291v3 Announce Type: replace Abstract: Unified multimodal models jointly support understanding and generation, but incur substantial…
India’s richest man now wants to turn aging computers into AI-ready PCs
Jio is betting it can turn an aging computer into an AI-ready PC for as little as about $11 for two months.
Automated Researchers Can Mitigate Well-characterized Alignment Failures
arXiv:2608.28945v3 Announce Type: replace Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard…
Quantifying User Behavior Patterns to Build Better Predictive Features
Simply knowing that a 35-year-old male in Seattle clicked 12 times last month tells you almost nothing about his intent.
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
arXiv:2608.27463v2 Announce Type: replace Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models’…
AI News Brief Hourly Summary 2026-09-04 02h : 11 posts
11 posts published in the last hour 23:32FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training 23:32FemWear: A Parameter-Efficient Wearable Foundation Model for Women’s Health 23:32SKILL.state: Scalable Long-Horizon Agent Skills 23:32Nova: An End-to-End MLIR Compiler for Deep Learning…
FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training
arXiv:2608.20574v3 Announce Type: replace Abstract: We introduce FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned…
FemWear: A Parameter-Efficient Wearable Foundation Model for Women’s Health
arXiv:2608.08244v2 Announce Type: replace Abstract: General-purpose wearable foundation models are pretrained on broad sensor streams and populations, but…
SKILL.state: Scalable Long-Horizon Agent Skills
arXiv:2608.26263v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running…
Nova: An End-to-End MLIR Compiler for Deep Learning
arXiv:2608.00029v3 Announce Type: replace Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level…
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
arXiv:2608.08389v2 Announce Type: replace Abstract: Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and…
Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
arXiv:2607.24814v2 Announce Type: replace Abstract: Access to specialist clinical expertise remains severely limited across sub-Saharan Africa, where…
LivingArena: Do LLMs Know What Other LLMs Don’t? Peer-Probing as Scalable Evaluation
arXiv:2607.24780v2 Announce Type: replace Abstract: Fixed benchmarks are costly to renew and cannot adapt their questions to model-specific failures. We…
Adaptive Graph-of-Islands Evolution for Automatic Feature Engineering with LLMs
arXiv:2607.23286v2 Announce Type: replace Abstract: Automatic feature engineering (AutoFE) for tabular data requires discovering informative…
Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance
arXiv:2607.22667v2 Announce Type: replace Abstract: This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for…
PIE-APT: Abductive Planning over Temporal Dynamic Knowledge Graphs via Incremental Reasoning
arXiv:2607.27287v2 Announce Type: replace Abstract: Planning over Temporal Dynamic Knowledge Graphs (TDKGs) presents theoretical challenges in open-world…
AI News Brief Hourly Summary 2026-09-04 01h : 16 posts
16 posts published in the last hour 22:32CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning 22:32Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models 22:32Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation 22:32Facilitating AI integration…
