arXiv:2608.08032v2 Announce Type: replace Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in…
Category: cs.AI updates on arXiv.org
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
arXiv:2608.11210v2 Announce Type: replace Abstract: Bayesian calibration of process-based models requires a prior distribution for each model parameter.…
LLM-Guided Graph Generation for Structure-Based Local Improvement Methods
arXiv:2608.13333v3 Announce Type: replace Abstract: Large neighborhood search normally selects a random subset of decision variables for iterative…
VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge
arXiv:2608.07994v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA),…
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
arXiv:2608.03744v2 Announce Type: replace Abstract: Clinical decision support is moving toward committees of language-model agents deliberating on a…
JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
arXiv:2607.16207v2 Announce Type: replace Abstract: Public open-weight language models are often fine-tuned on private or domain-specific data before…
G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution
arXiv:2608.01324v2 Announce Type: replace Abstract: Deep search has become a fundamental capability of large language models (LLMs) for solving…
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
arXiv:2607.27155v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly expected to assist users in completing tasks.…
FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation
arXiv:2607.12982v3 Announce Type: replace Abstract: Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large…
The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents
arXiv:2607.07436v2 Announce Type: replace Abstract: A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge…
