arXiv:2609.22619v1 Announce Type: new Abstract: Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation…
Author: script
SpeakON Ships a MagSafe AI Voice Button With Its Own Microphone
Voice input on phones has been solved for years. What has not been solved is the output. Speak into most dictation tools and you get back exactly what you…
EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability
arXiv:2609.22537v1 Announce Type: new Abstract: Enterprise AI assistants must produce responses that are verifiable and traceable to source evidence.…
“We’re already fighting yesterday’s battle”: Greece’s prime minister gets candid about AI
Most leaders on a trade mission stick to the pitch, but when I interviewed Greek Prime Minister Kyriakos Mitsotakis this week, he also admitted that no…
MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent…
AI News Brief Hourly Summary 2026-09-23 07h : 10 posts
10 posts published in the last hour 04:32IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law 04:32Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation 04:32Goal-driven Variant Categorization 04:32Agreement Overstates Evidence: Error Dependence in LLM…
IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law
arXiv:2609.22529v1 Announce Type: new Abstract: International law provides the normative framework through which states coordinate action, regulate armed…
Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
arXiv:2609.22478v1 Announce Type: new Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement…
Goal-driven Variant Categorization
arXiv:2609.22475v1 Announce Type: new Abstract: Process discovery rarely yields a single coherent process structure. For analysis, a common step is to…
Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that…
