arXiv:2608.20975v1 Announce Type: new Abstract: Effective social interaction requires agents to translate mental state inferences into coordinated…
Category: cs.AI updates on arXiv.org
Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance
arXiv:2608.21036v1 Announce Type: new Abstract: The transport of dangerous goods by sea is a high-consequence activity governed by the International…
The Cost of a Physics Prior Is Bounded by the Ablation Gap
arXiv:2608.21059v1 Announce Type: new Abstract: Shape-constrained and physics-informed learning reports an accuracy cost of enforcing a prior and treats…
Generalizing Soft Tissue Deformation and Force Prediction Across Material Stiffness and Geometry
arXiv:2608.20967v1 Announce Type: new Abstract: Accurate soft tissue simulation is essential for surgical training, pre-operative planning, and haptic…
TreeWY: Speculative Verification for Gated DeltaNet Hybrids
arXiv:2608.20961v1 Announce Type: new Abstract: Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a…
Deep Learning Models Also Recall Features
arXiv:2608.20970v1 Announce Type: new Abstract: Recent work in mechanistic interpretability has studied how large language models recall facts stored in…
Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning
arXiv:2608.20960v1 Announce Type: new Abstract: Language models (LMs) are trained on static scientific corpora, whereas scientific knowledge continuously…
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
arXiv:2608.20958v1 Announce Type: new Abstract: E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where…
No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators
arXiv:2608.20938v1 Announce Type: new Abstract: Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems…
UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists
arXiv:2608.20918v1 Announce Type: new Abstract: Organizations maintain task-specific adapters for open-weight language models, and each new base-model…
