arXiv:2609.08228v1 Announce Type: new Abstract: Modern LLM agents increasingly rely on reusable skills, yet as skill libraries scale to thousands of…
Tag: AI
Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts
arXiv:2609.08236v1 Announce Type: new Abstract: Automatic safety judges — systems such as Llama Guard or a GPT-4o grading prompt that decide whether a…
Meta’s AI agent Muse is now the No. 2 app in the US
Meta’s newest app Muse is off to a slower start than the company’s other apps, like Meta AI or Threads.
CircuTutor: Transforming Static Circuit Problems into Intelligent and Dynamic Tutoring
arXiv:2609.08254v1 Announce Type: new Abstract: Learning direct current circuit concepts requires learners to connect invisible physical quantities, such…
Do Dynamic Routers Need Memory? HeRo: History-Aware Routing for Efficient LLM Inference
arXiv:2609.08189v1 Announce Type: new Abstract: Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip…
A Better Spur Should Start From Each Objective
arXiv:2609.08211v1 Announce Type: new Abstract: Real-world Multi-Objective Reinforcement Learning (MORL) often suffers from sparse rewards, reward…
TTGBench: Benchmarking Topological Evolution and Semantic Drift in Text-attributed Temporal Graphs
arXiv:2609.08226v1 Announce Type: new Abstract: Temporal graph learning models the evolution of dynamic systems, where both structural interactions and…
Qiushi Engine on AstaBench E2E-Bench-Hard
arXiv:2609.08196v1 Announce Type: new Abstract: This report analyzes Qiushi Engine v0.8 across all 40 test tasks in AstaBench E2E-Bench-Hard, a benchmark…
Vision: Data-Centric Anchoring for Robust and Interpretable Agentic AI
arXiv:2609.08216v1 Announce Type: new Abstract: Agentic AI systems built on large language models fail in two persistent ways that scaling does not fix:…
Safe Harness Self-Evolution: A Theoretical Analysis of Feasibility and Limits
arXiv:2609.08175v1 Announce Type: new Abstract: Harness self-evolution is the process by which an agent modifies its prompts, tools, code, or…
