arXiv:2608.29753v1 Announce Type: new Abstract: Multi-hop question answering in retrieval-augmented gener?ation (RAG) often benefits from retrieving…
Category: cs.AI updates on arXiv.org
FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production
arXiv:2608.29814v1 Announce Type: new Abstract: Modern video generators excel at synthesizing individual clips, but complete video production requires…
Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment
arXiv:2608.29696v1 Announce Type: new Abstract: Evaluating research ideas generated by LLMs is difficult because their scientific value cannot be fully…
Perceive to Hypothesize, Verify to Ground: An Agentic Reasoning Framework for Open-World Geo-Localization
arXiv:2608.29880v1 Announce Type: new Abstract: Open-world geo-localization requires models to reason over ambiguous visual cues through multi-step…
Towards a Systems Foundation for Agentic Skills: Architecture, Lifecycle, and Security
arXiv:2608.29596v1 Announce Type: new Abstract: Autonomous large language model (LLM) agents increasingly face reliability, context consumption, and…
Detect Before You Attribute: Cascade Failure Attribution for Multi-Agent Systems
arXiv:2608.29646v1 Announce Type: new Abstract: Large language model (LLM)-based agents have shown strong potential in solving complex tasks through…
Not Safe for All: Auditing the Dialect Penalty in Text-to-Image Safety Pipelines
arXiv:2608.29589v1 Announce Type: new Abstract: Text-to-image (T2I) safety guardrails fail to generalize equitably to non-standard dialects. Evaluating…
LLMs Interpret, Embeddings Organize, Graphs Emerge: Agent-Driven Compilation of Scientific Knowledge
arXiv:2608.29612v1 Announce Type: new Abstract: Sustained scientific work requires a knowledge substrate that carries interpretation across tasks and…
Call Neighbours Yourself: Graph Walks with Destination-Conditioned On-Policy Self-Distillation
arXiv:2608.29588v1 Announce Type: new Abstract: Reasoning over text-attributed graphs (TAGs) requires large language models (LLMs) to combine a node’s…
EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
arXiv:2608.29387v1 Announce Type: new Abstract: Large language models can generate interactive web interfaces, but reliable generative UI requires…
