arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent…
Author: script
AI News Brief Hourly Summary 2026-09-23 07h : 10 posts
10 posts published in the last hour 04:32IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law 04:32Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation 04:32Goal-driven Variant Categorization 04:32Agreement Overstates Evidence: Error Dependence in LLM…
IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law
arXiv:2609.22529v1 Announce Type: new Abstract: International law provides the normative framework through which states coordinate action, regulate armed…
Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
arXiv:2609.22478v1 Announce Type: new Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement…
Goal-driven Variant Categorization
arXiv:2609.22475v1 Announce Type: new Abstract: Process discovery rarely yields a single coherent process structure. For analysis, a common step is to…
Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that…
The Wisdom of Artificial Deliberative Crowds
arXiv:2609.22497v1 Announce Type: new Abstract: The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as…
An Affordable AI-Integrated Smart Cane for Multimodal Mobility Assistance of Visually Impaired Users
arXiv:2609.22277v1 Announce Type: new Abstract: Visual impairment affects over 2.2 billion people worldwide, yet conventional white canes cannot detect…
Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models
arXiv:2609.22161v1 Announce Type: new Abstract: Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and…
PAANI : On Device Visual Evidence Fusion and Explainable Guidance for River Robot Simulation
arXiv:2609.22353v1 Announce Type: new Abstract: Mobile river monitoring robots must interpret obstacles and water boundaries that geographic waypoints…
