arXiv:2609.00267v1 Announce Type: cross Abstract: Autonomous LLM agents increasingly act on a user’s behalf: they hold credentials, call tools and…
Category: cs.AI updates on arXiv.org
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
arXiv:2609.00250v1 Announce Type: cross Abstract: Many people now see AI systems as not just productivity tools but as social companions. Researchers are…
CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
arXiv:2609.00242v1 Announce Type: cross Abstract: Long-tail autonomous driving failures are often framed as rare-object recognition errors. We argue that…
Don’t Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes
arXiv:2609.00227v1 Announce Type: cross Abstract: LLM agents increasingly diagnose incidents and propose remediations. In a GitOps workflow, applying a…
WHALE: A Simple Recipe for Joint Harness-Weight Optimization
arXiv:2609.00196v1 Announce Type: cross Abstract: Agent performance depends jointly on the model parameters and the executable harness code that manages…
QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization
arXiv:2609.00224v1 Announce Type: cross Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large…
Rock, Paper, Scissors, … Dynamite – A Model of Disruption from New Technologies
arXiv:2609.00207v1 Announce Type: cross Abstract: We seek to understand the effect of adding disruptive highly-capable new technologies to competitions by…
Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost
arXiv:2609.00193v1 Announce Type: cross Abstract: We study federated online reinforcement learning with linear function approximation. While recent…
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
arXiv:2609.00206v1 Announce Type: cross Abstract: Despite their growing use in video moderation, multimodal large language models (MLLMs) exhibit a…
Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
arXiv:2609.00191v1 Announce Type: cross Abstract: Crisis helplines assess suicide risk through structured interviews, a process that is slow and dependent…
