arXiv:2608.09325v1 Announce Type: new Abstract: Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines…
Category: cs.AI updates on arXiv.org
OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks
arXiv:2608.09380v1 Announce Type: new Abstract: Long-horizon complex tasks require agents to repeatedly observe states, formulate plans, invoke tools,…
LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling
arXiv:2608.09343v1 Announce Type: new Abstract: Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most…
CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits
arXiv:2608.09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image. A solver must ground…
Control-Oriented Scenario Tree Construction through Reinforcement Learning
arXiv:2608.09335v1 Announce Type: new Abstract: Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario…
ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management
arXiv:2608.09315v1 Announce Type: new Abstract: While mathematical models act as vital decision support systems for operational Air Traffic Flow and…
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
arXiv:2608.09324v1 Announce Type: new Abstract: On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive…
ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons
arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a…
Linearized 2-Simplicial Attention
arXiv:2608.09307v1 Announce Type: new Abstract: We present a linearized form of 2-simplicial attention by rewriting the trilinear score as an inner…
CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation
arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements,…