arXiv:2609.04094v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most…
Tag: cs.AI updates on arXiv.org
Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable
arXiv:2609.04127v1 Announce Type: new Abstract: Large language models are increasingly used to support organizational decisions, yet users often lack a…
Spurious Advantage Hidden in GRPO
arXiv:2609.04063v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable…
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
arXiv:2609.04098v1 Announce Type: new Abstract: Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose…
LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening
arXiv:2609.04013v1 Announce Type: new Abstract: Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine…
Instruction Duplication as an Inference-Time Control Primitive
arXiv:2609.04024v1 Announce Type: new Abstract: Procedural instruction following is a basic requirement for controllable language-model systems,…
InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models
arXiv:2609.04014v1 Announce Type: new Abstract: For trained operators, gauge reading requires little specialized knowledge, low cognitive effort, and high…
The Dually Flat Geometry of Planning as Inference
arXiv:2609.04005v1 Announce Type: new Abstract: We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by…
FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models
arXiv:2609.04021v1 Announce Type: new Abstract: Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more…
More Criticism Does Not Make a Better Review: EquiReview-R
arXiv:2609.03943v1 Announce Type: new Abstract: AI reviewers can now produce many specific criticisms, but more criticism is not necessarily a better…
