arXiv:2608.12677v1 Announce Type: new Abstract: Detecting infection-related behavioral changes in mosquitoes from video data is challenging because…
Tag: cs.AI updates on arXiv.org
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
arXiv:2608.12679v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The…
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
arXiv:2608.12654v1 Announce Type: new Abstract: Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or…
@skills: Attention is all you have
arXiv:2608.12610v1 Announce Type: new Abstract: There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery…
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
arXiv:2608.12645v1 Announce Type: new Abstract: LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling.…
General Probabilities of Causation with Causal Knowledge
arXiv:2608.12657v1 Announce Type: new Abstract: Probabilities of causation (PoCs) characterize individual causal responses that cannot be directly…
Designing AI Pipelines for Decision-Ready ITSM Intelligence
arXiv:2608.12670v1 Announce Type: new Abstract: IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are…
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
arXiv:2608.12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective…
DiG-bench: Discovery in Games
arXiv:2608.12593v1 Announce Type: new Abstract: Discovery—formulating novel generalizations—is a central part of the scientific process. Despite its…
Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues
arXiv:2608.12599v1 Announce Type: new Abstract: Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not…
