arXiv:2609.09882v1 Announce Type: new Abstract: Language models fine-tuned on customer behavior can predict outcomes and generate explanations, but these…
Tag: AI
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model
arXiv:2609.09815v1 Announce Type: new Abstract: Compound LLM systems often solve a coordination problem by adding a higher-level LLM. The resulting…
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward
arXiv:2609.09776v1 Announce Type: new Abstract: Frontier gains in language-model reasoning come from reinforcement learning on reasoning traces and are…
Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks
arXiv:2609.09774v1 Announce Type: new Abstract: Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine…
The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents
arXiv:2609.09853v1 Announce Type: new Abstract: LLM agents for enterprise systems of record cannot be evaluated on customer production data, and no…
Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields
arXiv:2609.09864v1 Announce Type: new Abstract: Affective computing has largely followed an individual-state paradigm, extracting discrete emotion labels…
Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation
arXiv:2609.09702v1 Announce Type: new Abstract: Candidate decision correctness and rationale grounding are different objectives. We examine…
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though…
Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
arXiv:2609.09678v1 Announce Type: new Abstract: Clinical diagnosis agents must decide not only what test to request next, but also when to diagnose or…
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
arXiv:2609.09754v1 Announce Type: new Abstract: As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic…
