arXiv:2608.18409v1 Announce Type: new Abstract: Combinatorial scheduling poses a significant challenge for language models, requiring them to identify…
Category: AI
When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification
arXiv:2608.18397v1 Announce Type: new Abstract: Wearable stress classifiers can achieve strong average performance while failing completely for a…
The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
arXiv:2608.18300v1 Announce Type: new Abstract: LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another…
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam’s 2025 Convex Marking Scheme
arXiv:2608.18336v1 Announce Type: new Abstract: When evaluating language models on human exams, benchmarks typically score each response as right or wrong…
Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair
arXiv:2608.18324v1 Announce Type: new Abstract: Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier…
SESSE: Sketch, Expand, Sort, Summarize, Evaluate — LLM-as-Judge Evaluation via Structured Decomposition
arXiv:2608.18303v1 Announce Type: new Abstract: LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice,…
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
arXiv:2608.18307v1 Announce Type: new Abstract: Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic…
Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
arXiv:2608.18289v1 Announce Type: new Abstract: The extraction of structured information from unstructured documents represents a critical component of…
Redakto – The Incognito Tab for LLMs
arXiv:2608.18260v1 Announce Type: new Abstract: Large Language Models (LLMs) are being increasingly used in everyday applications. A major challenge in…
Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study
arXiv:2608.18261v1 Announce Type: new Abstract: Serving a 235B-parameter Mixture-of-Experts (MoE) model on a single 8 GB GPU is bottlenecked not by…
