arXiv:2604.27269v3 Announce Type: replace Abstract: Biomedical knowledge graphs (KGs) are widely used in the life sciences, yet many are derived from…
Category: cs.AI updates on arXiv.org
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
arXiv:2605.24661v4 Announce Type: replace Abstract: Despite remarkable progress on reasoning benchmarks, current LLM evaluation practice remains anchored…
From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds
arXiv:2606.03557v2 Announce Type: replace Abstract: As generative AI capabilities expand, AI-driven virtual worlds face a growing architectural challenge.…
Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules
arXiv:2606.16337v4 Announce Type: replace Abstract: Predictive modeling for clinical decision support requires both strong predictive performance and…
FormalEvolve: Neuro-Symbolic Evolutionary Search for Diverse Autoformalization
arXiv:2603.19828v4 Announce Type: replace Abstract: Autoformalization aims to produce formal statements that compile and faithfully preserve the intended…
UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents
arXiv:2604.11557v3 Announce Type: replace Abstract: Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external…
From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems
arXiv:2604.02198v2 Announce Type: replace Abstract: While Artificial Intelligence (AI) offers transformative potential for operational performance, its…
TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning
arXiv:2603.03072v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to assist scientists across diverse workflows. A…
BUZZY: Contrastive Scoring to Mitigate Text-Induced Bias in Multimodal Multiple-Choice QA
arXiv:2603.28026v3 Announce Type: replace Abstract: Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively…
Edit Knowledge, Not Just Facts via Multi-Step Reasoning over Background Stories
arXiv:2602.02028v3 Announce Type: replace Abstract: Enabling artificial intelligence systems, particularly large language models, to update knowledge and…
