arXiv:2609.00065v1 Announce Type: cross Abstract: A language-model agent asked to analyse an experiment will usually return working code. Whether the…
Category: cs.AI updates on arXiv.org
Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
arXiv:2609.00067v1 Announce Type: cross Abstract: External text can override conflicting image evidence in multimodal large language models, a failure we…
Medical Causal Hypothesis Verification with Large Language Models
arXiv:2609.00063v1 Announce Type: cross Abstract: The growing use of large language models (LLMs) for search and information retrieval underscores the…
OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization
arXiv:2609.00066v1 Announce Type: cross Abstract: NVFP4 is an efficient microscaling format for low-bit inference, but activation outliers can still…
ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration
arXiv:2609.00061v1 Announce Type: cross Abstract: Reward post-training of diffusion generators inevitably concentrates probability mass on a few…
A Formal Analysis of Agent Payment Protocols
arXiv:2609.00060v1 Announce Type: cross Abstract: Agent payment protocols are emerging as a key transaction layer for autonomous commerce, enabling AI…
CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
arXiv:2609.00058v1 Announce Type: cross Abstract: Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation,…
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
arXiv:2609.00062v1 Announce Type: cross Abstract: Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical…
DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction
arXiv:2609.00059v1 Announce Type: cross Abstract: Materials property prediction remains difficult in low-data settings, where many target properties are…
REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent
arXiv:2609.00049v1 Announce Type: cross Abstract: Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict…
