arXiv:2608.30050v1 Announce Type: new Abstract: Building a blockchain digital twin largely requires translating domain knowledge and specific system…
Category: cs.AI updates on arXiv.org
Mitigating Over-Optimization in PRM-Guided Search in Mathematical Reasoning by Optimizing the Guide
arXiv:2608.30051v1 Announce Type: new Abstract: Process reward models (PRMs) provide dense step-level guidance for search-based reasoning, enabling…
Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks
arXiv:2608.30047v1 Announce Type: new Abstract: Recent AI systems promise autonomous scientific discovery, claiming to discover algorithms and produce…
Balance of Benchmarks: Semantic Density Reweighting for Benchmark Multiplicity and Task-Conditioned Evaluation
arXiv:2608.30044v1 Announce Type: new Abstract: Language models are commonly compared by averaging scores across a benchmark list with equal weight. Such…
Beyond Uncertainty: Multi-Solver Disagreement Rewards for Self-Evolving Reasoning Curricula
arXiv:2608.30035v1 Announce Type: new Abstract: Self-evolving reasoning frameworks train a Challenger to generate questions exposing a Solver’s…
AutoCRAT: Within-trajectory Joint Control of Stochasticity and Compute for LLM Reasoning
arXiv:2608.29988v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong reasoning performance, which depends critically on…
Interpreting and Steering for Safe and Correct Code Generation
arXiv:2608.30025v1 Announce Type: new Abstract: Large language models (LLMs) frequently generate source code containing vulnerabilities, yet little work…
Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models
arXiv:2608.30022v1 Announce Type: new Abstract: Introduction: NICE guidelines provide evidence-based recommendations for clinical care but remain largely…
An Open-Source, Event-Driven Pipeline for Cryptocurrency Market Data: Ingestion, Forecasting, and On-Chain Fraud Detection
arXiv:2608.29973v1 Announce Type: new Abstract: Cryptocurrency markets generate high-frequency, multi-source data that is expensive to work with unless a…
AcrossWAM1.0:A Modular Latent World-Action Stack for Compact Robot Policies
arXiv:2608.29937v1 Announce Type: new Abstract: Latent world-action models avoid rendering future pixels by predicting an action-relevant visual subgoal…
