arXiv:2609.19906v1 Announce Type: cross Abstract: Closed-loop robot policies require observation processing, state management, and situation-dependent…
Category: cs.AI updates on arXiv.org
KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms
arXiv:2609.19916v1 Announce Type: cross Abstract: Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language…
PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces
arXiv:2609.19883v1 Announce Type: cross Abstract: Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific…
ClashBench: Conflicts Leading Agents to Seize and Harm
arXiv:2609.19892v1 Announce Type: cross Abstract: As agent systems become more widely used, multiple agent sessions increasingly run alongside…
Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision
arXiv:2609.19846v1 Announce Type: cross Abstract: As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain…
PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance
arXiv:2609.19853v1 Announce Type: cross Abstract: Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and…
Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies
arXiv:2609.19844v1 Announce Type: cross Abstract: AI-generated RTL verification plans can satisfy a provider schema yet fail at the boundary to trusted…
Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles
arXiv:2609.19831v1 Announce Type: cross Abstract: In this reproducibility study, we investigate the transparency and scrutability of recommender systems…
A Functional Pilot for Certified Freshness-Aware Semantic–Spatial Range Retrieval
arXiv:2609.19855v1 Announce Type: cross Abstract: Geographic applications need every object inside a radius that satisfies a semantic threshold, yet…
SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes
arXiv:2609.19705v1 Announce Type: cross Abstract: Autonomous large language model (LLM) agents are moving rapidly into high-stakes domains, yet existing…
