12 posts published in the last hour 19:33Non-Stationarity Breaks Permutation Surrogates in Multi-Agent Reinforcement Learning: Diagnosis and Remedies 19:32CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification 19:32Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations 19:32Zero-shot…
Author: script
Non-Stationarity Breaks Permutation Surrogates in Multi-Agent Reinforcement Learning: Diagnosis and Remedies
arXiv:2604.23716v4 Announce Type: replace Abstract: Reporting guidance for information-theoretic measures is rarely tested against ground truth. We test…
CoGReV: A Confidence-Gated Post-Hoc Non-Monotonic Belief Revision Framework for Phishing Website Classification
arXiv:2604.25512v3 Announce Type: replace Abstract: In phishing detection, machine learning classifiers act as a first line of defense, but the false…
Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations
arXiv:2605.14175v2 Announce Type: replace Abstract: In a long conversation, an LLM can produce a plausible continuation that rests on premises the…
Zero-shot World Models Are Developmentally Efficient Learners
arXiv:2604.10333v2 Announce Type: replace Abstract: Young children demonstrate early abilities to understand their physical world, estimating depth,…
Cultural Binding Heads in Language Models
arXiv:2605.28543v3 Announce Type: replace Abstract: LLMs often default to equal treatment across cultural groups, even though context warrants…
Reinforcement learning for Quantum Tiq-Taq-Toe
arXiv:2411.06429v2 Announce Type: replace Abstract: Quantum Tiq-Taq-Toe is a well-known benchmark and playground for both quantum computing and machine…
ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork
arXiv:2505.23686v3 Announce Type: replace Abstract: Learning to collaborate with previously unseen partners is a fundamental generalization challenge,…
MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
arXiv:2604.10169v5 Announce Type: replace Abstract: Trajectory prediction is a key component of autonomous driving systems because future motions directly…
Reinforcement Learning with Temporal-Logic-Based Causal Diagrams
arXiv:2306.13732v2 Announce Type: replace Abstract: We study a class of reinforcement learning (RL) tasks where the objective of the agent is to…
