arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit…
Tag: cs.AI updates on arXiv.org
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
arXiv:2608.23640v1 Announce Type: new Abstract: When a large language model (LLM) is asked to write a person’s life, how much of what it writes actually…
Function-Level Execution Feedback for Code Preference Optimization
arXiv:2608.23632v1 Announce Type: new Abstract: Process supervision has improved mathematical reasoning, where intermediate steps are naturally expressed…
ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence
arXiv:2608.23569v1 Announce Type: new Abstract: State-of-the-art Natural Language to SQL (NL2SQL) models report execution accuracy exceeding 89 percent on…
LLM Agents Perform Controlled Experiments Using Simulation Models
arXiv:2608.23622v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many…
A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts
arXiv:2608.23626v1 Announce Type: new Abstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived…
TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery
arXiv:2608.23631v1 Announce Type: new Abstract: Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can…
RENDER: Controlling Reader-Facing Evidence in LLM Memory Evaluation
arXiv:2608.23568v1 Announce Type: new Abstract: Memory and RAG evaluations often treat the answering model’s input as an implementation detail, even…
Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction
arXiv:2608.22071v1 Announce Type: cross Abstract: Turn-taking is a basic organizational feature of human conversation and remains difficult to model in…
ADMIL: Attention-Distilled Multiple Instance Learning for Selective Foundation Model Inference in Pathology
arXiv:2608.22066v1 Announce Type: cross Abstract: Attention-based multiple instance learning (ABMIL) using pathology foundation model embeddings is…
