arXiv:2608.09307v1 Announce Type: new Abstract: We present a linearized form of 2-simplicial attention by rewriting the trilinear score as an inner…
Category: cs.AI updates on arXiv.org
CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation
arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements,…
P$^{3}$: Joint Program-and-Proof Planning for Verified Code Generation
arXiv:2608.09277v1 Announce Type: new Abstract: Verified code generation asks a large language model (LLM) to generate both an executable program and a…
Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline
arXiv:2608.09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different:…
Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
arXiv:2608.09263v1 Announce Type: new Abstract: Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens.…
Entropy-based Code Adversarial Translation for Real-world Repository Migration
arXiv:2608.09273v1 Announce Type: new Abstract: LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating…
MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence
arXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks…
Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution
arXiv:2608.09248v1 Announce Type: new Abstract: Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet…
An Explainable GNN Framework for Component-Level Anomaly Diagnosis
arXiv:2608.09246v1 Announce Type: new Abstract: Industrial processes are complex systems composed of multiple interacting sensors that generate…
SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge
arXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective…
