arXiv:2608.23766v1 Announce Type: cross Abstract: Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are…
Category: cs.AI updates on arXiv.org
Restoring Without Forgetting: Continual Learning Across Image Degradations
arXiv:2608.23799v1 Announce Type: cross Abstract: Recent progress in image restoration has converged on all-in-one architectures that jointly handle…
Disentangled Skill Representations for Predictive Human Modeling
arXiv:2608.23776v1 Announce Type: cross Abstract: Understanding human skill is important for AI systems that collaborate with, coach, or assist people.…
EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$
arXiv:2608.23758v1 Announce Type: cross Abstract: Recent large audio language models (LALMs) have achieved impressive progress in audio understanding.…
Too much of a good thing — when knowledge distillation promotes overfitting, and how to avoid it
arXiv:2608.23752v1 Announce Type: cross Abstract: The growing size of Convolutional Neural Networks has led to increasingly large and costly models.…
The Limits of Automatic Evaluation of Creativity in Large Language Models
arXiv:2608.23705v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly capable of generating text that challenges human…
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
arXiv:2608.23663v1 Announce Type: cross Abstract: Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device…
TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers
arXiv:2608.23763v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model…
Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling
arXiv:2608.23653v1 Announce Type: cross Abstract: AI agents are increasingly used for simulation-driven engineering. Physical system modeling presents…
Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail
arXiv:2608.23651v1 Announce Type: cross Abstract: Agent harnesses record a failed tool call and its error message in the transcript and ask the model to…
