arXiv:2609.02242v1 Announce Type: new Abstract: AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before…
Category: AI
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
arXiv:2609.02246v1 Announce Type: new Abstract: Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score…
Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
arXiv:2609.02264v1 Announce Type: new Abstract: Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and…
BenchMIRT: What are LLM benchmarks actually measuring?
This post has no text preview — click the link below to read the original article. This article has been indexed from Hugging Face – Blog Read the original article: BenchMIRT: What are LLM benchmarks actually measuring?
APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering
arXiv:2609.02253v1 Announce Type: new Abstract: Deep research agents augment large language models with external tools to answer complex, long-horizon…
PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion
arXiv:2609.02216v1 Announce Type: new Abstract: Inductive knowledge graph completion (IKGC) aims to predict missing links involving entities unseen during…
ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models
arXiv:2609.02215v1 Announce Type: new Abstract: Safety alignment trains large language models to refuse harmful requests stated plainly, but that training…
PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment
arXiv:2609.02231v1 Announce Type: new Abstract: Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging…
PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks
arXiv:2609.02236v1 Announce Type: new Abstract: Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in…
Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one…
