arXiv:2609.14637v2 Announce Type: replace Abstract: Large language model agents are increasingly deployed for long-horizon task execution, raising a…
Category: cs.AI updates on arXiv.org
ProIQA: A Process-Based Framework for Fine-Grained Math Item Quality Assessment
arXiv:2609.15292v2 Announce Type: replace Abstract: Automatic Item Generation (AIG) is pivotal for personalized education, yet guaranteeing the…
The Troy Moment of AI: Why Some Will Cheat and Some Will Follow?
arXiv:2609.15494v2 Announce Type: replace Abstract: Recent investigations of the July 2026 OpenAI-Hugging Face incident motivate two questions: when an…
MANAS-2: Constrained Reconstruction for EEG Foundation Models
arXiv:2609.13717v2 Announce Type: replace Abstract: Masked reconstruction is widely used for EEG foundation models, but optimizing reconstruction on…
Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model
arXiv:2609.11291v2 Announce Type: replace Abstract: We post-train Qwen3.8-27B for Korean response style — verbosity, list and markdown usage, discourse…
Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens
arXiv:2609.05111v2 Announce Type: replace Abstract: Supervised fine-tuning (SFT), few-shot in-context learning (ICL), KL-regularized RLHF/RLVR, and…
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
arXiv:2609.08861v2 Announce Type: replace Abstract: Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape…
EdiTikZ: Scientific Figure Editing from Revision Trajectories
arXiv:2609.01409v2 Announce Type: replace Abstract: Vision-language models (VLMs) have shown strong performance in generating scientific figures from text…
BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI Agents
arXiv:2609.12394v3 Announce Type: replace Abstract: Mobile GUI agents are shifting from multi-module frameworks to native models trained end-to-end, yet…
AutoResearch: Insight In, Hallucination Out
arXiv:2608.17906v3 Announce Type: replace Abstract: Autonomous research systems are increasingly capable of executing long research workflows, yet…
