arXiv:2608.08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models…
Tag: cs.AI updates on arXiv.org
What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files
arXiv:2608.08453v1 Announce Type: new Abstract: Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting…
Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation
arXiv:2608.08469v1 Announce Type: new Abstract: Existing streaming multimodal models process observations incrementally but still follow a turn-based…
Yesterday’s Shield, Today’s Spear: A Self-Evolving Safety Guardrail in Production
arXiv:2608.08471v1 Announce Type: new Abstract: Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new…
Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
arXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the…
TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Personalized Generation
arXiv:2608.08446v1 Announce Type: new Abstract: Personalized generation systems retrieve user history by request–memory relevance and inject it into the…
Estimating Uncertainty in Galaxy Morphology Classification
arXiv:2608.08398v1 Announce Type: new Abstract: Astronomers classify galaxy morphology to investigate cosmic evolution. While deep foundation models are…
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
arXiv:2608.08389v1 Announce Type: new Abstract: Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and…
Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective
arXiv:2608.08445v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of…
CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through…