arXiv:2609.23363v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have led to the emergence of coding agents capable of…
Tag: AI
Leaky-integrator reconstruction: taming error accumulation in recursive differenced time-series forecasting
arXiv:2609.23378v1 Announce Type: new Abstract: We introduce leaky-integrator reconstruction, a training-free method that cures the error accumulation of…
Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment
arXiv:2609.23640v1 Announce Type: new Abstract: Human-feedback alignment has made language models useful assistants and is commonly described as aligning…
PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI
arXiv:2609.23695v1 Announce Type: new Abstract: Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as…
AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction
arXiv:2609.23512v1 Announce Type: new Abstract: Large language model agents are typically deployed with predefined configurations, although the required…
Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation
arXiv:2609.23201v1 Announce Type: new Abstract: Benchmark scores increasingly influence the development, marketing, and selection of large language models…
Event Signature Transfer: Model-Agnostic Forecast Scenario Construction from Historical Events
arXiv:2609.23074v1 Announce Type: new Abstract: Forecasters often know an event is imminent but not the shape, size, or timing of its effect. We introduce…
CraftBench-UE: Deterministic Evaluation for Coding Agents in Unreal Engine
arXiv:2609.23142v1 Announce Type: new Abstract: Building gameplay features in a game engine requires more than code, as code that compiles and runs does…
From Inference Engine to Inference Control Plane: Connecting vLLM, llm-d, and the Evolution of Efficient Distributed LLM Serving
arXiv:2609.23130v1 Announce Type: new Abstract: Large language model (LLM) inference is evolving from an engine-local optimization problem into a…
AI Agents Are Becoming a New Malware Distribution Channel
By Farukh Rakhimov, Head of Compliance, Data Protection and Information Security at AdTech Holding Roughly 7,600 fake GitHub repositories, 6,600…
