arXiv:2609.22592v1 Announce Type: new Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment…
Tag: cs.AI updates on arXiv.org
Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation
arXiv:2609.22620v1 Announce Type: new Abstract: Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split…
GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment
arXiv:2609.22619v1 Announce Type: new Abstract: Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation…
EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability
arXiv:2609.22537v1 Announce Type: new Abstract: Enterprise AI assistants must produce responses that are verifiable and traceable to source evidence.…
MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent…
IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law
arXiv:2609.22529v1 Announce Type: new Abstract: International law provides the normative framework through which states coordinate action, regulate armed…
Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
arXiv:2609.22478v1 Announce Type: new Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement…
Goal-driven Variant Categorization
arXiv:2609.22475v1 Announce Type: new Abstract: Process discovery rarely yields a single coherent process structure. For analysis, a common step is to…
Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that…
The Wisdom of Artificial Deliberative Crowds
arXiv:2609.22497v1 Announce Type: new Abstract: The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as…
