arXiv:2609.27756v1 Announce Type: new Abstract: Large language models are increasingly asked to analyze data and report what the results mean, a task…
Category: AI
Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions
arXiv:2609.27749v1 Announce Type: new Abstract: The surge in AI-assisted generation of educational materials has outpaced our capacity to validate their…
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
arXiv:2609.27490v1 Announce Type: new Abstract: AI research agents need reliable knowledge of how their experiments change outcomes. We introduce…
SHRAV: State-Hypothesis-Reason-Action-Verify Framework for Physical Modeling and Inverse Design
arXiv:2609.27621v1 Announce Type: new Abstract: Physical modeling and inverse design require computation that can continue from reusable state. We…
State-Grounded Conditioning: Wrapping User-Facing LLM Agents Where Direction Depends on Live State
arXiv:2609.27606v1 Announce Type: new Abstract: We introduce State-Grounded Conditioning (SGC), a design principle for user-facing LLM agents that must…
Not What You Meant: Can LLMs Follow a Specified Negation Semantics?
arXiv:2609.27517v1 Announce Type: new Abstract: Negation does not carry a uniform interpretation across domains. In legal, regulatory, and medical…
BiCFlow-MER: Orchestrating Discriminative and Generative Multimodal Emotion Recognition via Conditional Transport
arXiv:2609.27615v1 Announce Type: new Abstract: In multimodal emotion recognition (MER), human affective states are inferred by integrating complementary…
Emergi-PersonaOS: A Persona Agent Operating System for Situational Adaptation and Controllable Evolution
arXiv:2609.27417v1 Announce Type: new Abstract: Symbiosis between humans and digital beings offers a vision for the future of human–machine interaction.…
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
arXiv:2609.27334v1 Announce Type: new Abstract: Agentic memory systems reuse past experience to improve future performance, yet most existing designs…
CART: Closed-Loop Adaptive Red Teaming for Large Language Models
arXiv:2609.27336v1 Announce Type: new Abstract: Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn…
