arXiv:2609.23201v1 Announce Type: new Abstract: Benchmark scores increasingly influence the development, marketing, and selection of large language models…
Tag: cs.AI updates on arXiv.org
Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events
arXiv:2609.23074v1 Announce Type: new Abstract: Forecasters often know an event is imminent but not the shape, size, or timing of its effect. We introduce…
CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine
arXiv:2609.23142v1 Announce Type: new Abstract: Building gameplay features in a game engine requires more than code, as code that compiles and runs does…
From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving
arXiv:2609.23130v1 Announce Type: new Abstract: Large language model (LLM) inference is evolving from an engine-local optimization problem into a…
Expansion Counts under Standard A* Tie-Breaking Strategies on the Final Plateau
arXiv:2609.23293v1 Announce Type: new Abstract: In the A* search algorithm, the tie-breaking strategies for nodes with the same $f$-value determines which…
Tutoring Large Language Models to be Domain-adaptive, Precise and Safe
arXiv:2609.23071v1 Announce Type: new Abstract: This thesis proposes a framework for “responsible intelligence” to address AI’s critical challenges in…
Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
arXiv:2609.23038v1 Announce Type: new Abstract: Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical…
FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics
arXiv:2609.23064v1 Announce Type: new Abstract: Understanding the physical world requires more than object recognition, scene description, and short-term…
LazyAgent: Demand-Driven Materialization and Physical Optimization of Agentic Programs
arXiv:2609.23058v1 Announce Type: new Abstract: Current agent runtimes that plan before acting generally execute a step once it becomes ready. We present…
Enforcing Narrative Reliability and Epistemic Pacing in LLM-Driven Detective Games via Structured Knowledge Trees
arXiv:2609.23043v1 Announce Type: new Abstract: Large Language Models (LLMs) enable open-ended dialogue in interactive games, but their non-deterministic…
