arXiv:2606.16276v3 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no…
Category: AI
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
arXiv:2607.05915v3 Announce Type: replace Abstract: PCB routing is the task of connecting the nets of a board with copper traces under strict design…
Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
arXiv:2606.31002v2 Announce Type: replace Abstract: Lean verifies that a generated declaration is well typed, but not that it expresses the statement a…
GeoNatureAgent Benchmark: Benchmarking LLM Agents for Environmental Geospatial Analysis Across Frontier and Open-Weight Foundation Models
arXiv:2606.12821v2 Announce Type: replace Abstract: Environmental scientists spend disproportionate effort on data wrangling rather than analysis. New AI…
Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
arXiv:2606.31413v3 Announce Type: replace Abstract: Composing independently trained LoRA adapters into a single large language model is useful for…
StatefulDiscovery: Evidence-Calibrated Claim Formation in Open-Ended Scientific Discovery
arXiv:2606.11851v2 Announce Type: replace Abstract: Open-ended scientific discovery asks agents to move beyond executing analyses for predefined…
AIP: A Graph Representation for Learning and Governing Agent Skills
arXiv:2606.04781v2 Announce Type: replace Abstract: Agent Skills today consist largely of free-form prose requiring the agent to read, interpret, and…
Deploy a multimodal WhatsApp ordering assistant with Amazon Bedrock AgentCore
Learn how to deploy a multimodal WhatsApp ordering assistant that takes customer orders through text, voice notes, and real-time voice calls on a single…
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Activations
arXiv:2605.28553v2 Announce Type: replace Abstract: In this paper, we investigate whether refusal behavior can be predicted from LLM intermediate…
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
This post has no text preview — click the link below to read the original article. This article has been indexed from Hugging Face – Blog Read the original article: Fine-tuning a 350M Model for Better Structured Outputs in 100…
