arXiv:2609.24620v1 Announce Type: new Abstract: Answering epidemiological questions from real-world clinical data requires medical coding, schema-aware…
Category: cs.AI updates on arXiv.org
The Endless Exam: Mathematical Constructions from Today’s Models toward Superintelligence
arXiv:2609.24555v1 Announce Type: new Abstract: We introduce the Endless Exam, a benchmark for measuring mathematical progress from today’s models toward…
Not All Task Vectors Need Equal Rank: Energy-Proportional Allocation for Model Merging
arXiv:2609.24517v1 Announce Type: new Abstract: Model merging aims to combine multiple fine-tuned models derived from a common pretrained model into a…
DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
arXiv:2609.24662v1 Announce Type: new Abstract: LLM-based agents increasingly operate in environments where they interact with users, tools, and external…
Custom Named Entity Recognition and Topic Classification for Global Health Publications
arXiv:2609.24625v1 Announce Type: new Abstract: How should natural language processing models be selected and adapted for global health literature in…
Predicting Postprandial Glycemic Response from Meal Images, Clinical Variables, and Gut Microbiome Information
arXiv:2609.24453v1 Announce Type: new Abstract: Predicting postprandial glycemic response (PPGR) is fundamental to personalized nutrition and type 2…
LADDER: Graph-Guided Diffusion Language Models for Efficient Multi-Hop Reasoning
arXiv:2609.24346v1 Announce Type: new Abstract: Graph Retrieval-Augmented Generation (GraphRAG) has remarkably enhanced large language models on complex…
Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards
arXiv:2609.24480v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) in healthcare requires robust performance across two complementary…
Few-Shot Demonstrations Elicit the Use of In-Context World Representations in LLMs
arXiv:2609.24352v1 Announce Type: new Abstract: Large language models (LLMs), when acting as agents, are expected to take observed data in context, infer…
VLM-in-Sandbox: Visual Workspaces for Agentic Visual Reasoning
arXiv:2609.24362v1 Announce Type: new Abstract: Sandboxed computer environments support multi-step reasoning with tools, executable programs, and…
