arXiv:2609.02244v1 Announce Type: new Abstract: Large language models (LLMs) often struggle when low-resource training data are ambiguous or incomplete.…
Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality
arXiv:2609.02242v1 Announce Type: new Abstract: AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before…
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
arXiv:2609.02246v1 Announce Type: new Abstract: Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score…
Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems
arXiv:2609.02264v1 Announce Type: new Abstract: Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and…
BenchMIRT: What are LLM benchmarks actually measuring?
This post has no text preview — click the link below to read the original article. This article has been indexed from Hugging Face – Blog Read the original article: BenchMIRT: What are LLM benchmarks actually measuring?
APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering
arXiv:2609.02253v1 Announce Type: new Abstract: Deep research agents augment large language models with external tools to answer complex, long-horizon…
PEARL: Path-Entity Aligned Relational Learning with Contextual Subgraphs for Inductive Knowledge Graph Completion
arXiv:2609.02216v1 Announce Type: new Abstract: Inductive knowledge graph completion (IKGC) aims to predict missing links involving entities unseen during…
ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models
arXiv:2609.02215v1 Announce Type: new Abstract: Safety alignment trains large language models to refuse harmful requests stated plainly, but that training…
PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment
arXiv:2609.02231v1 Announce Type: new Abstract: Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging…
PGPO: Potential-Guided Policy Optimization for Multi-Turn Agentic Tasks
arXiv:2609.02236v1 Announce Type: new Abstract: Group-based reinforcement learning (RL) has become an effective paradigm for LLM post-training, but in…
Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon
Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one…
SkillGLoW: Procedural-Family Skill Consolidation for Self-Improving Agents on Long-Horizon Task Streams
arXiv:2609.02217v1 Announce Type: new Abstract: LLM agents increasingly self-improve by writing and reusing textual skills, kept either as one global…
AI News Brief Hourly Summary 2026-09-03 09h : 17 posts
17 posts published in the last hour 06:32Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents 06:32Anthropic opens Claude AI text detection to regulators, media, fact-checkers, and others 06:32FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs 06:32Google’s Android update…
Beyond Context Windows: Persistent Discovery Context for Data-Centric Agents
arXiv:2609.02129v1 Announce Type: new Abstract: Data-centric agents repeatedly perform a discovery step before planning or execution: identifying the data…
Anthropic opens Claude AI text detection to regulators, media, fact-checkers, and others
Anthropic is launching an API that lets regulators, media outlets, and researchers check whether text carries Claude’s digital watermark. The EU AI Act…
FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs
arXiv:2609.02168v1 Announce Type: new Abstract: Fragmented safety evaluation undermines the governance of dangerous AI capabilities. We present a modular…
Google’s Android update tackles motion sickness, accessibility, and more
While some of the features see Google playing catch-up to Apple, which already offers similar features for iPhone users, others specifically leverage…
Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning
arXiv:2609.02191v1 Announce Type: new Abstract: Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We…
