arXiv:2609.22611v1 Announce Type: cross Abstract: Humanoid robots can acquire complex skills by imitating kinematic humanoid motion references, yet…
Author: script
LLaDA-PRM: A Bidirectional Step-Level Reasoning Evaluator
arXiv:2609.22700v1 Announce Type: cross Abstract: Step-level reasoning evaluators are commonly based on autoregressive language models, whose causal…
From Capability to Assurance in Autonomous Penetration-Testing Harnesses: A Framework and Reference Implementation
arXiv:2609.22664v1 Announce Type: cross Abstract: Research on large language model agents for penetration testing is evaluated almost entirely by…
Anthropic says its biology lab has already found something big
But maybe the biggest reveal is that Anthropic has not let Claude run lose in its biology lab. Humans are still, so far, in the loop.
Math2Visual-X: A Modular Framework for Pedagogically Aligned Lower-Primary Math Visuals Generation
arXiv:2609.22647v1 Announce Type: cross Abstract: Visual representations can help lower-primary learners understand Math Word Problems, but generating…
Fairly Compensated Distributed Information Retrieval and Augmentation for AI Agents
arXiv:2609.22601v1 Announce Type: cross Abstract: The increasing reliance of autonomous AI agents on external and distributed knowledge sources introduces…
From Documented Strengths to Force Limits: Material-Informed Robotic Insertion for Construction Assembly
arXiv:2609.22609v1 Announce Type: cross Abstract: Insertion is a fundamental operation in robotic construction assembly, where variations in material…
Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation
arXiv:2609.22603v1 Announce Type: cross Abstract: Summarization ships in countless production systems, making model selection a routine decision that…
Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act
arXiv:2609.22588v1 Announce Type: cross Abstract: Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make…
SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning
arXiv:2609.22586v1 Announce Type: cross Abstract: Modern audio-language models are no longer judged only on what words they can transcribe, but on whether…
