arXiv:2608.12788v1 Announce Type: new Abstract: The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure…
Category: cs.AI updates on arXiv.org
PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs
arXiv:2608.12762v1 Announce Type: new Abstract: Schedulability analysis is essential for certifying real-time systems, but existing tests are often…
Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
arXiv:2608.12674v1 Announce Type: new Abstract: Maintaining price consistency and executing an Every Day Low Price strategy is critical for global…
On the Expressive Power of Transformers
arXiv:2608.12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in…
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
arXiv:2608.12675v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely used to improve the performance of Large Language Models…
The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis
arXiv:2608.12677v1 Announce Type: new Abstract: Detecting infection-related behavioral changes in mosquitoes from video data is challenging because…
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The…
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or…
@skills: Attention is all you have
arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery…
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.…
