arXiv:2609.03026v1 Announce Type: cross Abstract: Mechanistic interpretability is increasingly used to guide interventions such as activation steering,…
Category: cs.AI updates on arXiv.org
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
arXiv:2609.03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile:…
Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression
arXiv:2609.03189v1 Announce Type: cross Abstract: This article presents a structured framework of behavioral indicators that may signal progression toward…
Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data
arXiv:2609.03391v1 Announce Type: cross Abstract: Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language…
Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning
arXiv:2609.02967v1 Announce Type: cross Abstract: Topology-guided safeguards for LLM-based multi-agent systems (MAS) train a GNN over the inter-agent…
Evaluating Graph Neural Networks for Change-Criticality Classification in Maritime Navigation Charts
arXiv:2609.02996v1 Announce Type: cross Abstract: Graph neural networks (GNNs) are a class of neural networks suitable for learning on graph-structured…
When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization
arXiv:2609.02964v1 Announce Type: cross Abstract: This paper focuses on defending generative search engines against malicious Generative Engine…
Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy
arXiv:2609.02990v1 Announce Type: cross Abstract: To scale up collective decision-making, participatory democracy platforms such as Polis and Remesh…
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation
arXiv:2609.02998v1 Announce Type: cross Abstract: On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a…
Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation
arXiv:2609.02942v1 Announce Type: cross Abstract: LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption…
