arXiv:2608.15138v1 Announce Type: new Abstract: Designing an ABR algorithm for one network scenario takes an engineer months, and large language models…
Category: cs.AI updates on arXiv.org
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads
arXiv:2608.15117v1 Announce Type: new Abstract: Analytical models of peak VRAM consumption for LLM inference decompose memory into weight-storage,…
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models
arXiv:2608.15145v1 Announce Type: new Abstract: Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain…
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
arXiv:2608.15089v1 Announce Type: new Abstract: Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may…
Validation-Frontier Representation Selection under Constrained Observation
arXiv:2608.15095v1 Announce Type: new Abstract: AI systems deployed outside clean benchmark settings often rely on observations that are incomplete,…
Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents
arXiv:2608.15109v1 Announce Type: new Abstract: Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical…
Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Systems
arXiv:2608.15082v1 Announce Type: new Abstract: Cold chain logistics has advanced technologically, yet most deployed systems remain reactive monitors, not…
Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation
arXiv:2608.15101v1 Announce Type: new Abstract: Policy evaluation often estimates direct benefits and costs while treating the institutional environment…
Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
arXiv:2608.15071v1 Announce Type: new Abstract: Learning from experience is critical for developing capable, self-improving large language model (LLM)…
GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG
arXiv:2608.15056v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (RAG) systems often rely on long unstructured contexts or…
