arXiv:2609.05146v1 Announce Type: new Abstract: This study introduces an intelligent framework that integrates machine learning and deep neural network…
Category: AI
ProCA: Progressive Contrastive Alignment for Robust EEG Visual Decoding
arXiv:2609.05094v1 Announce Type: new Abstract: Electroencephalogram (EEG) visual decoding aims to recover visual semantics from non-invasive neural…
Compact Bellman-Grounded Cognitive Maps for Cost-Aware Navigation
arXiv:2609.05104v1 Announce Type: new Abstract: Biological agents navigate familiar environments not by re-solving routes for each new goal, but by…
Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens
arXiv:2609.05111v1 Announce Type: new Abstract: Large language models are now trained and evaluated under a diverse set of paradigms: supervised…
TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
arXiv:2609.05079v1 Announce Type: new Abstract: Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write…
Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent
arXiv:2609.05090v1 Announce Type: new Abstract: Evaluation of medical artificial intelligence agents remains predominantly answer-centric, assessing only…
LLM-Guided Program Evolution for Circle Packing: Breaking 10 Packomania Records for $28
arXiv:2609.05093v1 Announce Type: new Abstract: We present Discovery Loop, a lightweight system that uses a large language model (LLM) to iteratively…
Measuring AI Accountability Through Argumentation Analysis: Can Model Reasoning Withstand Scrutiny?
arXiv:2609.05088v1 Announce Type: new Abstract: AI oversight methods rely on ground truth for validation, but what constitutes appropriate AI behavior is…
MePo++: Unifying Representation Refinement and Reconciliation for General Continual Learning
arXiv:2609.05075v1 Announce Type: new Abstract: General continual learning (GCL) aims to learn from evolving data streams without task identities,…
Language models judge war differently when tested for alignment
arXiv:2609.05009v1 Announce Type: new Abstract: Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to…
