arXiv:2608.30362v1 Announce Type: new Abstract: As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a…
Tag: cs.AI updates on arXiv.org
Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents
arXiv:2608.30322v1 Announce Type: new Abstract: Professional agent tasks often depend on conventions that are absent from public corpora, yet benchmarks…
Augmenting Human Performance with an XR Agent Learning from Online Behavior and BCI Evidence
arXiv:2608.30369v1 Announce Type: new Abstract: We present OLIVE, a framework for adapting a foundation model to provide real-time assistance in…
Answer Probing-Guided Search for Diverse Solution Exploration of LLMs
arXiv:2608.30345v1 Announce Type: new Abstract: Generating multiple diverse and high-quality solutions is valuable for many applications, such as…
LLM-Based Knowledge Graph Completion Combining Discrete Structural Coding with Similar Entity Information
arXiv:2608.30235v1 Announce Type: new Abstract: Knowledge graph completion requires models to use both textual descriptions and relational structure.…
CoLa-ICD: A Knowledge-Enhanced Framework for Long-Tail Automated Medical Coding
arXiv:2608.30234v1 Announce Type: new Abstract: Automatic medical coding assigns ICD codes to clinical notes, but it remains challenging due to long…
Generating Workflow DAGs from Natural Language with Non-Reasoning LLMs
arXiv:2608.30250v1 Announce Type: new Abstract: This paper addresses the problem of translating natural-language routing rules written by business…
Rethinking the Test-Time Prompt Tuning Objective from the Perspective of Calibration
arXiv:2608.30230v1 Announce Type: new Abstract: Test-time prompt tuning (TPT) has emerged as a powerful paradigm, refining prompts for each test sample…
SimCRAFT: Distilling Remote Sensing Agents via Synthetic Trajectories and Contextual Retrieval-Augmented Fine-Tuning
arXiv:2608.30277v1 Announce Type: new Abstract: The unprecedented surge in Earth observation data volume and diversity has exposed a critical bottleneck…
A.X K2 Technical Report
arXiv:2608.30181v1 Announce Type: new Abstract: We introduce A.X K2, a 688B-parameter Mixture-of-Experts (MoE) language model trained from scratch as a…
