arXiv:2609.22695v1 Announce Type: new Abstract: The term “linear representation hypothesis” (LRH) has appeared across diverse subfields of artificial…
Tag: AI
AutoGym: Blueprint-First Generation of Verifiable Agent Gyms
arXiv:2609.22592v1 Announce Type: new Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment…
Splitting Documents at Lower Cost: Multi-Split Boundary Decisions for LLM-Based Page Stream Segmentation
arXiv:2609.22620v1 Announce Type: new Abstract: Scanned mail, uploaded PDFs, and consolidated attachments often arrive as page streams that must be split…
GaitVista: Reliability-Aware AI Measurement toward Accessible Longitudinal Gait Assessment
arXiv:2609.22619v1 Announce Type: new Abstract: Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation…
SpeakON Ships a MagSafe AI Voice Button With Its Own Microphone
Voice input on phones has been solved for years. What has not been solved is the output. Speak into most dictation tools and you get back exactly what you…
EvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability
arXiv:2609.22537v1 Announce Type: new Abstract: Enterprise AI assistants must produce responses that are verifiable and traceable to source evidence.…
“We’re already fighting yesterday’s battle”: Greece’s prime minister gets candid about AI
Most leaders on a trade mission stick to the pitch, but when I interviewed Greek Prime Minister Kyriakos Mitsotakis this week, he also admitted that no…
MAWILE: Multi-Axis Workbench for Inspecting LLM Evaluators
arXiv:2609.22599v1 Announce Type: new Abstract: Large language model (LLM) judges provide a flexible and scalable method for evaluating model and agent…
IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law
arXiv:2609.22529v1 Announce Type: new Abstract: International law provides the normative framework through which states coordinate action, regulate armed…
Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation
arXiv:2609.22478v1 Announce Type: new Abstract: Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement…
