arXiv:2607.05915v3 Announce Type: replace Abstract: PCB routing is the task of connecting the nets of a board with copper traces under strict design…
Tag: cs.AI updates on arXiv.org
Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
arXiv:2606.31002v2 Announce Type: replace Abstract: Lean verifies that a generated declaration is well typed, but not that it expresses the statement a…
GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models
arXiv:2606.12821v2 Announce Type: replace Abstract: Environmental scientists spend disproportionate effort on data wrangling rather than analysis. New AI…
Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
arXiv:2606.31413v3 Announce Type: replace Abstract: Composing independently trained LoRA adapters into a single large language model is useful for…
StatefulDiscovery: Evidence-Calibrated Claim Formation in Open-Ended Scientific Discovery
arXiv:2606.11851v2 Announce Type: replace Abstract: Open-ended scientific discovery asks agents to move beyond executing analyses for predefined…
AIP: A Graph Representation for Learning and Governing Agent Skills
arXiv:2606.04781v2 Announce Type: replace Abstract: Agent Skills today consist largely of free-form prose requiring the agent to read, interpret, and…
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
arXiv:2605.28553v2 Announce Type: replace Abstract: In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate…
CoMAP: Co-Evolving World Models and Agent Policies for LLM Agents
arXiv:2606.02372v2 Announce Type: replace Abstract: Equipping language agents with world models enables them to anticipate environment dynamics and…
Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models
arXiv:2606.02914v3 Announce Type: replace Abstract: Background: Oral diseases affect nearly 3.5 billion people worldwide, yet the comparative clinical…
MIRA: A Bilingual Benchmark for Medical Information Response Audit
arXiv:2605.28025v2 Announce Type: replace Abstract: Existing safety evaluations for large language models overlook whether responses preserve comparable…
