arXiv:2608.14047v1 Announce Type: cross Abstract: This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose…
Category: cs.AI updates on arXiv.org
HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation
arXiv:2608.14032v1 Announce Type: cross Abstract: Existing multimodal RAG methods often flatten structured documents into isolated text and image units,…
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
arXiv:2608.14022v1 Announce Type: cross Abstract: Action-conditioned video world models require low-latency causal generation and reliable responses to…
Musical Mirrors: The LLM as Sounding Board in Songwriting
arXiv:2608.13944v1 Announce Type: cross Abstract: This paper examines a use of AI in creative practice as an interpretive sounding board for…
CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification
arXiv:2608.13939v1 Announce Type: cross Abstract: Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework…
EchoRec: Multi-Item Prediction-Empowered Generative Recommendation via Cycle-Consistent Preference Alignment
arXiv:2608.14011v1 Announce Type: cross Abstract: Generative recommendation autoregressively generates the semantic IDs of the target item, unifying…
MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning
arXiv:2608.14015v1 Announce Type: cross Abstract: Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what…
Content Based Video Narration of Gameplay with Vision Language Models
arXiv:2608.14016v1 Announce Type: cross Abstract: Live game commentary is scarce: it exists for professional esports broadcasts and almost nowhere else.…
Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning
arXiv:2608.13914v1 Announce Type: cross Abstract: Electrocardiogram (ECG) recordings are sensitive biomedical data, limiting the ability of hospitals and…
CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing
arXiv:2608.13925v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a…
