arXiv:2608.29198v1 Announce Type: new Abstract: As Large Language Models (LLMs) increasingly encourage users to disclose personal profiles for tailored…
Category: AI
An Explainable Coherence Score for Detecting Temporal Inconsistencies in Political News
arXiv:2608.29175v1 Announce Type: new Abstract: Temporal inconsistencies, such as mandates attributed outside their real interval, events presented as…
Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation
arXiv:2608.29210v1 Announce Type: new Abstract: Text-to-Image (T2I) models have recently achieved impressive visual fidelity, yet their evaluation remains…
Benevolent Bias in Multi-Turn Human-Agent Dialogue
arXiv:2608.29206v1 Announce Type: new Abstract: Bias in human-agent interaction can manifest not only through hostile language but also as benevolent…
Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling
arXiv:2608.29207v1 Announce Type: new Abstract: Protein structure modeling rests on a single computational primitive: the interaction between what a…
Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges
arXiv:2608.29127v1 Announce Type: new Abstract: We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality…
Emergent Misalignment Is Not Magical
arXiv:2608.29118v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a…
APIFlow-Bench: Measuring Whether Agents Survive Long, Dependent API Workflows
arXiv:2608.29128v1 Announce Type: new Abstract: Tool-using agents are commonly evaluated by a single bit: whether an end-to-end workflow completed. This…
JudgePanel: A Compact Judge with Panel Deliberation via Adaptive Multi-Reward Reinforcement Learning
arXiv:2608.29168v1 Announce Type: new Abstract: The LLM-as-a-Judge paradigm has emerged as a scalable alternative to human evaluation. However,…
Renesas Joins Autoware Foundation to Bring Open-Source AI to ADAS Platforms
The Autoware Foundation said on September 1, 2026, that Renesas Electronics has joined the organization as a member of its highest membership tier, in a…
