arXiv:2609.21229v1 Announce Type: cross Abstract: Learning robot manipulation policies typically requires substantial demonstration data, which are costly…
Tag: cs.AI updates on arXiv.org
Hallucination-R1: Robustness-Oriented Paraphrase Generation for Factual Consistency
arXiv:2609.21227v1 Announce Type: cross Abstract: Factual hallucination is commonly defined by incorrect factual outputs. We study a paraphrase-induced…
Verify, Don’t Trust: Agentic Model Development for Video Discovery Retrieval at Scale
arXiv:2609.21257v1 Announce Type: cross Abstract: Large language model (LLM) agents can propose, implement, and evaluate model changes. Autoresearch loops…
FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models
arXiv:2609.21228v1 Announce Type: cross Abstract: Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong…
Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies
arXiv:2609.21216v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated…
Visual Navigation Transformer with Pose Attention
arXiv:2609.21212v1 Announce Type: cross Abstract: Learned navigation policies typically consume observations as a temporally ordered history, with…
EnSol: an environment-aware graph neural network for molecular solubility prediction
arXiv:2609.21151v1 Announce Type: cross Abstract: Molecular solubility directly affects key aspects of molecular development such as reaction feasibility,…
SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
arXiv:2609.21190v1 Announce Type: cross Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering.…
The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators
arXiv:2609.21133v1 Announce Type: cross Abstract: SQL has been augmented with AI operators, enabling modern data analytics platforms to derive insights…
How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?
arXiv:2609.21058v1 Announce Type: cross Abstract: Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations…
