arXiv:2608.26194v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to…
Category: cs.AI updates on arXiv.org
ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
arXiv:2608.26204v1 Announce Type: cross Abstract: Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on…
Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model
arXiv:2608.26221v1 Announce Type: cross Abstract: As generative AI gains traction, researchers are investigating its potential to serve as proxies for…
Investigating the Influence of Prompt and Response Languages on LLM Content Generation
arXiv:2608.26186v1 Announce Type: cross Abstract: This study examines how prompt and response language influence the behavior of large language models.…
PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants
arXiv:2608.26180v1 Announce Type: cross Abstract: Shopping assistants are shifting from ranked product lists toward structured decision support, where…
Comparing Chunking and Embedding Strategies for Turkish RAG Systems
arXiv:2608.26192v1 Announce Type: cross Abstract: How documents are segmented into retrievable chunks and how those chunks are embedded strongly affect…
A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs
arXiv:2608.26177v1 Announce Type: cross Abstract: Long-form generation exposes fundamental limitations of large language models. Even 70B-parameter models…
When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models
arXiv:2608.26187v1 Announce Type: cross Abstract: Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of…
Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention
arXiv:2608.26168v1 Announce Type: cross Abstract: The lifecycle of hallucination in LLMs is a concept that enables building solid frameworks on the…
Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning
arXiv:2608.26166v1 Announce Type: cross Abstract: Advancing reasoning capabilities allow large language models (LLMs) to tackle increasingly complex…
