arXiv:2609.22664v1 Announce Type: cross Abstract: Research on large language model agents for penetration testing is evaluated almost entirely by…
Tag: AI
Anthropic says its biology lab has already found something big
But maybe the biggest reveal is that Anthropic has not let Claude run lose in its biology lab. Humans are still, so far, in the loop.
Math2Visual-X: A Modular Framework for Pedagogically Aligned Lower-Primary Math Visuals Generation
arXiv:2609.22647v1 Announce Type: cross Abstract: Visual representations can help lower-primary learners understand Math Word Problems, but generating…
Fairly Compensated Distributed Information Retrieval and Augmentation for AI Agents
arXiv:2609.22601v1 Announce Type: cross Abstract: The increasing reliance of autonomous AI agents on external and distributed knowledge sources introduces…
From Documented Strengths to Force Limits: Material-Informed Robotic Insertion for Construction Assembly
arXiv:2609.22609v1 Announce Type: cross Abstract: Insertion is a fundamental operation in robotic construction assembly, where variations in material…
Preserving What Matters: Semantic Scaffolds Beyond Saturation in Summarization Evaluation
arXiv:2609.22603v1 Announce Type: cross Abstract: Summarization ships in countless production systems, making model selection a routine decision that…
Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act
arXiv:2609.22588v1 Announce Type: cross Abstract: Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make…
SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning
arXiv:2609.22586v1 Announce Type: cross Abstract: Modern audio-language models are no longer judged only on what words they can transcribe, but on whether…
Do Student LLMs Inherit OOD Robustness? Invariance-Weighted Distillation for Reliable Knowledge Transfer
arXiv:2609.22566v1 Announce Type: cross Abstract: Knowledge distillation (KD) aims to compress high-performance teacher LLMs into lightweight students.…
FRAMES: Failure Recovery And Monitoring of Embodied Skills for Humanoid Loco-Manipulation
arXiv:2609.22538v1 Announce Type: cross Abstract: Large language model (LLM) planners can decompose natural-language instructions and select reusable…
