arXiv:2608.24237v1 Announce Type: new Abstract: Longitudinal radiology report generation (LRRG) requires identifying both current findings and their…
Category: cs.AI updates on arXiv.org
Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks
arXiv:2608.24258v1 Announce Type: new Abstract: AI systems are increasingly evaluated for legally accountable settings, where correct outputs must also be…
Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding
arXiv:2608.24273v1 Announce Type: new Abstract: Continual knowledge graph embedding updates entity and relation representations as a graph grows. Existing…
Evaluating Multiple LLM Generations with Validated Task Coverage
arXiv:2608.24228v1 Announce Type: new Abstract: Many LLM applications are most useful when they provide several candidate outputs for comparison,…
MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG
arXiv:2608.24214v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue searching…
Constraint-Guided Enterprise Data Mapping with Large Language Models
arXiv:2608.24218v1 Announce Type: new Abstract: Enterprise entity alignment must handle semi-structured records, implicit attributes, and unit or…
Preference Data Selection for Mitigating the Alignment Tax in Large Language Models
arXiv:2608.24192v1 Announce Type: new Abstract: Aligning large language models to human preferences is crucial for real-world deployment but frequently…
TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models
arXiv:2608.24232v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even…
Task-Adaptive Rubrics for GUI Reward Modeling
arXiv:2608.24174v1 Announce Type: new Abstract: Recent studies on GUI agents have increasingly focused on outcome reward modeling, which assigns outcome…
OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
arXiv:2608.24160v1 Announce Type: new Abstract: Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and…
