arXiv:2608.20434v1 Announce Type: cross Abstract: The coordination of multi-scale tasks is an effective strategy for computational materials discovery,…
Category: cs.AI updates on arXiv.org
Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources
arXiv:2608.20438v1 Announce Type: cross Abstract: Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent…
ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib
arXiv:2608.20432v1 Announce Type: cross Abstract: Formal proofs in Lean 4 that pass the kernel’s type checker can nonetheless vary widely in quality. We…
LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine
arXiv:2608.20402v1 Announce Type: cross Abstract: Biomedical knowledge graphs (KGs) are pivotal for knowledge organization, yet traditional binary…
Six misconceptions about large language models: A minimal model and diagnostic taxonomy
arXiv:2608.20421v1 Announce Type: cross Abstract: Large language models (LLMs) are now embedded in scientific, educational, and governance workflows, with…
From Thermal Preference Prediction to Adaptive Thermal Intervention: A Reinforcement Learning Approach Using Physiological and Environmental Sensing
arXiv:2608.20423v1 Announce Type: cross Abstract: Personalised thermal comfort is essential for occupant wellbeing and for the development of more…
Rigorous Evaluation of Large Language Models for Malaria Drug Discovery: Trade-offs in Performance, Scale, and Resource Utility
arXiv:2608.20418v1 Announce Type: cross Abstract: We introduce Malaria-Instruct, a curated instruction-following dataset derived from the ChEMBL Legacy…
Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI
arXiv:2608.20393v1 Announce Type: cross Abstract: Agentic large language models (LLMs) deployed in fact-sensitive applications such as customer support…
Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants
arXiv:2608.20392v1 Announce Type: cross Abstract: LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding…
EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators
arXiv:2608.20381v1 Announce Type: cross Abstract: Automating slide editing requires simultaneously satisfying modification accuracy, preservation…
