Learn how Salesforce used Amazon SageMaker AI Inference Component placement (the SchedulingConfig parameter) to distribute model copies across multiple…
Category: AI
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation
arXiv:2608.26237v1 Announce Type: cross Abstract: Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of…
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
arXiv:2608.26222v1 Announce Type: cross Abstract: Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust…
Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness Bugs
arXiv:2608.26209v1 Announce Type: cross Abstract: Data-driven software systems are increasingly deployed in high-stakes socio-economic domains, from…
A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers
arXiv:2608.26194v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to…
ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
arXiv:2608.26204v1 Announce Type: cross Abstract: Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on…
Anthropic gets its first court win over the Pentagon’s supply-chain risk label
A federal judge ruled the Trump administration illegally labeled Anthropic a supply-chain risk, handing the AI company a victory as its second Pentagon…
Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model
arXiv:2608.26221v1 Announce Type: cross Abstract: As generative AI gains traction, researchers are investigating its potential to serve as proxies for…
Investigating the Influence of Prompt and Response Languages on LLM Content Generation
arXiv:2608.26186v1 Announce Type: cross Abstract: This study examines how prompt and response language influence the behavior of large language models.…
PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants
arXiv:2608.26180v1 Announce Type: cross Abstract: Shopping assistants are shifting from ranked product lists toward structured decision support, where…
