arXiv:2609.03781v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in multilingual settings, yet their safety is still…
Tag: AI
Cross-Dataset Transfer and Reliability of Explainable Artificial Intelligence for RhythmFormer Remote Photoplethysmography
arXiv:2609.03663v1 Announce Type: cross Abstract: Background. Remote photoplethysmography estimates the cardiovascular pulse from facial video, and its…
Enhancing Financial Question Answering: A Novel Benchmark Dataset of Banks’ financial statements
arXiv:2609.03654v1 Announce Type: cross Abstract: The comparative analysis of banks’ financial statements poses significant challenges for automated…
Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning
arXiv:2609.03667v1 Announce Type: cross Abstract: Generalising to unseen tasks remains a fundamental challenge in offline multi-agent reinforcement…
Local Updates, Global Learning (LUGL): Playing Games with non-incremental Learners
arXiv:2609.03660v1 Announce Type: cross Abstract: The dominance of Neural Networks (NNs) in RL is partially due to their incremental learning capability,…
Doesn’t Stop Reasoning: Analysis of Spurious CoT Termination
arXiv:2609.03633v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large reasoning models (LRMs) on complex tasks but often…
Test-time adaptation for speech enhancement with an autoregressive speech prior
arXiv:2609.03622v1 Announce Type: cross Abstract: Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under…
EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders
arXiv:2609.03629v1 Announce Type: cross Abstract: Recent advances in text-to-video (T2V) diffusion models have demonstrated remarkable generative…
Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation
arXiv:2609.03619v1 Announce Type: cross Abstract: Multi-agent debate (MAD) improves the reasoning capabilities of large language models by having multiple…
FailBench: How Reliable are VLMs at Judging Robot Task Success?
arXiv:2609.03611v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but…
