arXiv:2609.19244v1 Announce Type: new Abstract: Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search…
Tag: cs.AI updates on arXiv.org
MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
arXiv:2609.19391v1 Announce Type: new Abstract: LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly…
The syntax and semantics of goals
arXiv:2609.19448v1 Announce Type: new Abstract: In both cognitive science and computer science, goals are conceptualized as cognitive states that flexibly…
Do AI Agents Understand Computer Architecture?
arXiv:2609.19387v1 Announce Type: new Abstract: Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports…
Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer
arXiv:2609.19203v1 Announce Type: new Abstract: AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems.…
BioPhys-Bridge: A Benchmark for Interdisciplinary Scientific Reasoning in Physics-Grounded Biological Research
arXiv:2609.19180v1 Announce Type: new Abstract: Language models face unique challenges in analyzing interdisciplinary scientific research literature. In…
Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes
arXiv:2609.19170v1 Announce Type: new Abstract: Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its…
What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis
arXiv:2609.19212v1 Announce Type: new Abstract: Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is…
What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks
arXiv:2609.19182v1 Announce Type: new Abstract: Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet…
Extracting Probabilistic Knowledge from Large Language Models for Bayesian Network Parameterization
arXiv:2505.15918v3 Announce Type: replace-cross Abstract: In this work, we evaluate the potential of Large Language Models (LLMs) in building Bayesian…
