arXiv:2608.13456v2 Announce Type: replace Abstract: World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan,…
Tag: AI
K-Bench: measuring model performance on real scientific agent requests
arXiv:2608.21601v2 Announce Type: replace Abstract: Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice…
Nvidia confirms it will buy Hugging Face for $12.9 billion
Nvidia said Hugging Face hosts over 3 million models and is used by over 18 million developers.
Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
arXiv:2607.20058v2 Announce Type: replace Abstract: Large language models can answer scientific questions, yet a correct output does not reveal whether…
SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data
arXiv:2606.16276v3 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no…
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
arXiv:2607.05915v3 Announce Type: replace Abstract: PCB routing is the task of connecting the nets of a board with copper traces under strict design…
Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
arXiv:2606.31002v2 Announce Type: replace Abstract: Lean verifies that a generated declaration is well typed, but not that it expresses the statement a…
GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models
arXiv:2606.12821v2 Announce Type: replace Abstract: Environmental scientists spend disproportionate effort on data wrangling rather than analysis. New AI…
Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
arXiv:2606.31413v3 Announce Type: replace Abstract: Composing independently trained LoRA adapters into a single large language model is useful for…
StatefulDiscovery: Evidence-Calibrated Claim Formation in Open-Ended Scientific Discovery
arXiv:2606.11851v2 Announce Type: replace Abstract: Open-ended scientific discovery asks agents to move beyond executing analyses for predefined…
