arXiv:2505.11326v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image…
Category: cs.AI updates on arXiv.org
HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement
arXiv:2505.15740v2 Announce Type: replace-cross Abstract: Formal methods play a crucial role in ensuring the reliability of critical systems through…
Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation
arXiv:2506.21599v5 Announce Type: replace-cross Abstract: Advancing large language models (LLMs) for the next point-of-interest (POI) recommendation task…
CollaFuse: Collaborative Diffusion Models
arXiv:2406.14429v5 Announce Type: replace-cross Abstract: In the landscape of generative artificial intelligence, diffusion-based models have emerged as a…
The BS-meter: Detecting Politics and Labour through ChatGPT’s Language
arXiv:2411.15129v3 Announce Type: replace-cross Abstract: What can we learn about language from studying how it is used by ChatGPT and other large…
Recurrent Reinforcement Learning with Memoroids
arXiv:2402.09900v4 Announce Type: replace-cross Abstract: Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially…
Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
arXiv:2504.05216v5 Announce Type: replace-cross Abstract: Dense retrieval is a crucial task in Information Retrieval (IR), serving as the basis for…
Communication styles and reader preferences of LLM- and human-authored COVID-19 information explanations: a case study
arXiv:2505.08143v2 Announce Type: replace-cross Abstract: With the wide adoption of large language models (LLMs) in information assistance, it is…
Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents
arXiv:2608.22963v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where…
ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation
arXiv:2608.22559v2 Announce Type: replace Abstract: Rubrics aim to make language-model evaluation transparent by decomposing response quality into…
