arXiv:2608.14426v2 Announce Type: replace Abstract: AI is increasingly being used to help with AI R&D. Under certain conditions this feedback loop might…
Tag: cs.AI updates on arXiv.org
What You Can’t See Is What You Learn: Slot-Selective Evidence Masking Favors Compositional Generalization in Shared-Genome Language-Model Societies
arXiv:2608.20054v3 Announce Type: replace Abstract: Multi-module neural systems often expose every module to the full input. We test whether a…
ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory Forecasting
arXiv:2608.20009v2 Announce Type: replace Abstract: Understanding object dynamics requires not only predicting future trajectories but also examining…
SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents
arXiv:2608.08253v2 Announce Type: replace Abstract: We present SuperLocalMemory 4.0, a governed, local-first memory operating system for AI agents,…
Atomic Units of X: The Compression Layer of Intelligence
arXiv:2607.12634v2 Announce Type: replace Abstract: This paper proposes a theoretical and empirical framework for understanding intelligence as a process…
Self-Evolving Scientific Agent Designs Physically-Reasoned Whitebox Fluid Control
arXiv:2606.08405v3 Announce Type: replace Abstract: While data-intensive deep reinforcement learning can optimize complex control policies, scientific…
Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
arXiv:2608.09696v4 Announce Type: replace Abstract: A primary goal of science is to learn mechanistic or causal world models from data. These models can…
EngiAI: Capability-Based Evaluation of Tool-Connected LLM Agents for Engineering Design
arXiv:2605.19743v3 Announce Type: replace Abstract: Engineering-agent systems are proliferating, but differences in tasks, tools, and success criteria…
Strategic Exploitation in LLM Agent Markets: A Simulation Framework for E-Commerce Trust
arXiv:2605.10059v3 Announce Type: replace Abstract: Agent-based modeling (ABM) has long been used in economics to study human behavior, and large language…
ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams
arXiv:2604.15994v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) excel at recognizing individual visual elements and reasoning…
