arXiv:2608.23706v1 Announce Type: new Abstract: A large language model (LLM) trained on synthetic limit order book (LOB) data achieves near perfect scores…
Tag: cs.AI updates on arXiv.org
A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification
arXiv:2608.23817v1 Announce Type: new Abstract: SHAP and LIME are now standard tools for interpreting black-box predictions, yet their outputs can vary…
Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search
arXiv:2608.23811v1 Announce Type: new Abstract: Automated fact-checking is essential for ensuring the reliability of public health information, yet the…
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
arXiv:2608.23691v1 Announce Type: new Abstract: We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which…
Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering
arXiv:2608.23666v1 Announce Type: new Abstract: Sycophancy and hallucination are persistent failure modes of Large Language Models (LLMs) across domains.…
Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value
arXiv:2608.23644v1 Announce Type: new Abstract: Large language models (LLMs) are becoming routine instruments of scientific research, assisting with…
Automata from Agent Traces: Failure and Next-Step Prediction
arXiv:2608.23670v1 Announce Type: new Abstract: LLM-based agents execute multi-step tasks, but their behavioral structure remains opaque: long…
MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models
arXiv:2608.23646v1 Announce Type: new Abstract: Molecular embedding models can serve as foundational infrastructure for computational chemistry and drug…
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
arXiv:2608.23640v1 Announce Type: new Abstract: When a large language model (LLM) is asked to write a person’s life, how much of what it writes actually…
How much of a measured AI preference is the model, and how much is the instrument?
arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit…
