200 posts published today 21:33MIRA: A Bilingual Benchmark for Medical Information Response Audit 21:33Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs 21:33CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery 21:33A Comparative Study in Surgical AI: Potential and…
Author: script
MIRA: A Bilingual Benchmark for Medical Information Response Audit
arXiv:2605.28025v2 Announce Type: replace Abstract: Existing safety evaluations for large language models overlook whether responses preserve comparable…
Towards Affordable Energy: A Gymnasium Environment for Electric Utility Demand-Response Programs
arXiv:2605.12462v2 Announce Type: replace Abstract: Extreme weather and volatile wholesale electricity markets expose residential consumers to…
CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery
arXiv:2604.01658v3 Announce Type: replace Abstract: Large language model (LLM)-based evolution is a promising approach for open-ended discovery, where…
A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling
arXiv:2603.27341v5 Announce Type: replace Abstract: Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several…
AI compute provider Nscale is looking for $3.5B in pre-IPO financing
Nscale, which recently struck a $45 billion deal with Anthropic, is in talks to raise additional funds in anticipation of an upcoming IPO.
Causal Probing for Internal Visual Representations in Multimodal Large Language Models
arXiv:2605.05593v3 Announce Type: replace Abstract: Despite the remarkable success of Multimodal Large Language Models (MLLMs) across diverse tasks, the…
Auditing Multi-Agent LLM Reasoning Trees Outperforms Majority Vote and LLM-as-Judge
arXiv:2602.09341v2 Announce Type: replace Abstract: Multi-agent systems (MAS) can substantially extend the reasoning capacity of large language models…
Not All Preferences Deserve Gradients: Understanding Gradient Utility in Offline Reasoning Alignment
arXiv:2602.01207v2 Announce Type: replace Abstract: Offline preference optimization aligns reasoning models from fixed chosen–rejected pairs, yet…
NeuroWeaver: An Autonomous Evolutionary Agent for Exploring the Programmatic Space of EEG Analysis Pipelines
arXiv:2602.13473v3 Announce Type: replace Abstract: Although foundation models have achieved remarkable success in general domains, applying them to…
