A recap of all the incidents involving LLMs made by Anthropic, Meta, and OpenAI, which went rogue and attacked real companies and individuals on the…
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization
arXiv:2608.20768v3 Announce Type: replace Abstract: Specialist language models are usually understood through endpoint gains: the generalist scores lower,…
Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations
arXiv:2608.03611v2 Announce Type: replace Abstract: Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet…
Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models
arXiv:2607.04562v2 Announce Type: replace Abstract: Large language models (LLMs) generate fluent outputs that can be wrong. Unlike humans, who often…
What We Can Learn From Google Engineers’ Indispensible Prompts
Hey, Google Engineers: What prompt do you personally refuse to work without, and why?
Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
arXiv:2607.25529v2 Announce Type: replace Abstract: As neural network models for image classification advance, neurons play critical roles in pruning,…
Ransomware Operator Ran Cursor Agent Inside Ten Victim Networks
Gambit Security’s threat intelligence team has published a detailed account of the Aurora ransomware operation, including six weeks of session logs…
Rethinking the Evaluation of Harness Evolution for Agents
arXiv:2607.12227v2 Announce Type: replace Abstract: We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution…
When Consumers Ask AI: Rethinking Brand Visibility in the Age of AI Recommendations
The Search Result Is Becoming a Recommendation For years, digital marketing was built around a simple transaction: a consumer searched, a search engine…
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
arXiv:2607.15439v2 Announce Type: replace Abstract: Our previous ARC-AGI-3 agent bundled executable world modeling, prompted simplification, and exact…
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
arXiv:2606.18037v3 Announce Type: replace Abstract: Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous…
Learning the ARTS of Search for Automated Discovery
arXiv:2606.21891v2 Announce Type: replace Abstract: Scientific discovery can be formulated as an iterative search process over the space of hypotheses and…
From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems
arXiv:2605.23955v4 Announce Type: replace Abstract: Deploying machine learning in regulated financial environments — credit risk, fraud detection, and…
Nomad: Autonomous Exploration and Discovery
arXiv:2603.29353v3 Announce Type: replace Abstract: We introduce Nomad, a system for autonomous data exploration and insight discovery. Given a corpus of…
When AI Is Everywhere, What Becomes the Competitive Advantage?
For the past few years, access to AI has been an advantage in itself. The companies that moved early could automate faster, build new capabilities, and…
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
arXiv:2606.11637v4 Announce Type: replace Abstract: Touch is a key modality for embodied agents to understand the physical world. Although recent work has…
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
arXiv:2510.10002v4 Announce Type: replace Abstract: As agentic AI systems are deployed in advisory and evaluative roles, understanding how multi-agent…
Learning to Predict, Discover, and Reason in High-Dimensional Event Sequences
arXiv:2603.16313v3 Announce Type: replace Abstract: Electronic control units (ECUs) embedded within modern vehicles generate a large number of…
