14 posts published in the last hour 20:32Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer’s Disease Detection 20:32Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems 20:32CompanionSim: Synthetic Data for Evaluating…
Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer’s Disease Detection
arXiv:2609.00276v1 Announce Type: cross Abstract: Speech-based Alzheimer’s disease (AD) detection increasingly relies on speech-enhanced and curated…
Delegation Without Trust: An Empirical Gap Analysis of Identity, Authorization, and Runtime Governance in Multi-Agent LLM Systems
arXiv:2609.00267v1 Announce Type: cross Abstract: Autonomous LLM agents increasingly act on a user’s behalf: they hold credentials, call tools and…
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
arXiv:2609.00250v1 Announce Type: cross Abstract: Many people now see AI systems as not just productivity tools but as social companions. Researchers are…
OpenAI’s new reasoning technique alarms AI safety experts
OpenAI’s new Astra model will use “recurrent depth,” a technique that allows the model to operate outside of the sequential thinking that characterizes…
CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
arXiv:2609.00242v1 Announce Type: cross Abstract: Long-tail autonomous driving failures are often framed as rare-object recognition errors. We argue that…
OneStream Adds SensibleAI Tools to FedRAMP High Authorization
OneStream said on September 2, 2026, that its SensibleAI Forecast and SensibleAI Studio capabilities have been added to the company’s FedRAMP High…
Don’t Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes
arXiv:2609.00227v1 Announce Type: cross Abstract: LLM agents increasingly diagnose incidents and propose remediations. In a GitOps workflow, applying a…
WHALE: A Simple Recipe for Joint Harness-Weight Optimization
arXiv:2609.00196v1 Announce Type: cross Abstract: Agent performance depends jointly on the model parameters and the executable harness code that manages…
QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization
arXiv:2609.00224v1 Announce Type: cross Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large…
Rock, Paper, Scissors, … Dynamite – A Model of Disruption from New Technologies
arXiv:2609.00207v1 Announce Type: cross Abstract: We seek to understand the effect of adding disruptive highly-capable new technologies to competitions by…
Provably Efficient Federated Reinforcement Learning with Linear Function Approximation and Logarithmic Communication Cost
arXiv:2609.00193v1 Announce Type: cross Abstract: We study federated online reinforcement learning with linear function approximation. While recent…
Meta Launches Muse Spark 1.3, Citing Gains in Coding and Agentic Tasks
Meta on September 2, 2026 released Muse Spark 1.3, the latest model from Meta Superintelligence Labs, making it available the same day in Muse Code and in…
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
arXiv:2609.00206v1 Announce Type: cross Abstract: Despite their growing use in video moderation, multimodal large language models (MLLMs) exhibit a…
AI News Brief Hourly Summary 2026-09-02 22h : 12 posts
12 posts published in the last hour 19:32Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models 19:32Intelligent Edge Computing 19:32Do General NLP Embeddings Capture Ontological Reasoning? 19:32Lingua Franca or Probing Artifact? Rethinking…
Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
arXiv:2609.00191v1 Announce Type: cross Abstract: Crisis helplines assess suicide risk through structured interviews, a process that is slow and dependent…
Intelligent Edge Computing
arXiv:2609.00181v1 Announce Type: cross Abstract: The number of edge devices in large-scale edge systems is rapidly increasing. Edge devices have limited…
Do General NLP Embeddings Capture Ontological Reasoning?
arXiv:2609.00177v1 Announce Type: cross Abstract: General-purpose NLP embedding models perform well on linguistic tasks, but their ability to capture…
