arXiv:2608.09277v1 Announce Type: new Abstract: Verified code generation asks a large language model (LLM) to generate both an executable program and a…
Category: cs.AI updates on arXiv.org
Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline
arXiv:2608.09254v1 Announce Type: new Abstract: LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different:…
Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
arXiv:2608.09263v1 Announce Type: new Abstract: Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens.…
Entropy-based Code Adversarial Translation for Real-world Repository Migration
arXiv:2608.09273v1 Announce Type: new Abstract: LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating…
MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence
arXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks…
Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution
arXiv:2608.09248v1 Announce Type: new Abstract: Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet…
An Explainable GNN Framework for Component-Level Anomaly Diagnosis
arXiv:2608.09246v1 Announce Type: new Abstract: Industrial processes are complex systems composed of multiple interacting sensors that generate…
SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge
arXiv:2608.09230v1 Announce Type: new Abstract: Industrial-safety understanding requires more than detecting workers, equipment, and personal protective…
SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance
arXiv:2608.09253v1 Announce Type: new Abstract: LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and…
Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models
arXiv:2608.09227v1 Announce Type: new Abstract: Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint…