arXiv:2608.28640v1 Announce Type: cross Abstract: In this paper, we present PromptKWS, a novel Prompt-guided keyword spotting (KWS) framework to improve…
Category: cs.AI updates on arXiv.org
Do MLLMs Really Understand Low-Resource Khmer Documents? A Pilot Study on Khmer Document VQA
arXiv:2608.28635v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have advanced document understanding, visual question…
PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation
arXiv:2608.28633v1 Announce Type: cross Abstract: Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often…
Redesigning and Auditing Deep Research Writing for Faithful Reports
arXiv:2608.28643v1 Announce Type: cross Abstract: Rubric-based evaluations of deep-research (DR) systems often obscure fine-grained factual failures in…
Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework
arXiv:2608.28630v1 Announce Type: cross Abstract: Compared with half-duplex dialogue systems where the system waits for user turn completion before it…
MA-RAG: Multi-Agent Retrieval-Augmented Generation for Query-Driven Summarization of Longitudinal Parkinson’s Disease Assessments
arXiv:2608.28624v1 Announce Type: cross Abstract: Accurate interpretation of single-visit and longitudinal clinical assessments for Parkinson’s disease is…
Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects
arXiv:2608.28626v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of…
Intelligent Identification and Repair of Design Defects in BIM via Domain-Specific Large Language Models
arXiv:2608.28629v1 Announce Type: cross Abstract: Existing methods lack a generalized approach to efficiently identify and resolve the diversity of design…
Asymmetric Within-Document Predictive Learning for Scientific Document Representation
arXiv:2608.28625v1 Announce Type: cross Abstract: We study predictive pretraining for scientific document representation using the discourse structure of…
Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure
arXiv:2608.28623v1 Announce Type: cross Abstract: Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating…
