arXiv:2609.11801v1 Announce Type: cross Abstract: Humans and machines often solve harder problems by spending more time on computation. In deep learning,…
Category: cs.AI updates on arXiv.org
Logit Refiner: Improving Visual Autoregressive Models via Intra-Scale Dependency Modeling
arXiv:2609.11804v1 Announce Type: cross Abstract: Visual Autoregressive Models (VAR) generate images through next-scale prediction, producing all tokens…
RetroThinker: Enabling Retrospective Thinking in Speech LLMs
arXiv:2609.11864v1 Announce Type: cross Abstract: Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that…
Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport
arXiv:2609.11842v1 Announce Type: cross Abstract: Diffusion and flow-matching schedules control the signal and noise coefficients that mix data and noise…
ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI
arXiv:2609.11737v1 Announce Type: cross Abstract: Collective intelligence depends not only on the capabilities of individual members, but also on how…
LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation
arXiv:2609.11739v1 Announce Type: cross Abstract: Large language model serving costs scale directly with output sequence length, yet standard preference…
Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech
arXiv:2609.11786v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error…
Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing
arXiv:2609.11769v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to analyze and rewrite news, yet current framing…
Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations
arXiv:2609.11725v1 Announce Type: cross Abstract: Text-to-speech (TTS) models commonly address text–speech alignment by expanding phone-level encoder…
ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies
arXiv:2609.11697v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in…
