arXiv:2608.23809v1 Announce Type: cross Abstract: Multilingual language models can solve the same mathematical problem in different languages, but it…
Category: cs.AI updates on arXiv.org
Restoring Without Forgetting: Continual Learning Across Image Degradations
arXiv:2608.23799v1 Announce Type: cross Abstract: Recent progress in image restoration has converged on all-in-one architectures that jointly handle…
Predicting Radiologist Expertise from 3D Gaze Patterns During CT Interpretation
arXiv:2608.23836v1 Announce Type: cross Abstract: Accurate interpretation of volumetric CT requires efficient navigation of 3D image volumes and attention…
EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$
arXiv:2608.23758v1 Announce Type: cross Abstract: Recent large audio language models (LALMs) have achieved impressive progress in audio understanding.…
What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development
arXiv:2608.23766v1 Announce Type: cross Abstract: Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are…
Too much of a good thing — when knowledge distillation promotes overfitting, and how to avoid it
arXiv:2608.23752v1 Announce Type: cross Abstract: The growing size of Convolutional Neural Networks has led to increasingly large and costly models.…
Disentangled Skill Representations for Predictive Human Modeling
arXiv:2608.23776v1 Announce Type: cross Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist people.…
TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers
arXiv:2608.23763v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model…
Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling
arXiv:2608.23653v1 Announce Type: cross Abstract: AI agents are increasingly used for simulation-driven engineering. Physical system modeling presents…
The Limits of Automatic Evaluation of Creativity in Large Language Models
arXiv:2608.23705v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly capable of generating text that challenges human…
