Multi-agent systems fail in ways traditional monitoring misses. This post presents a dual-layer approach to monitoring production agents: Amazon Bedrock…
Tag: AI
Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs
arXiv:2609.10439v1 Announce Type: cross Abstract: Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable…
Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload
Comparing models on dollars per million tokens misses what production workloads actually pay for: outcomes. This post shares an open-source benchmarking…
Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization
arXiv:2609.10464v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the…
Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs
arXiv:2609.10346v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image,…
PACE: Perceived-Latency-Aware Cascading Service Routing and Filler Control for QoE-Efficient Retrieval-Augmented Dialogue Serving
arXiv:2609.10372v2 Announce Type: cross Abstract: We present the PACE, a framework for retrieval-augmented dialogue serving that formalizes Perceived…
OmniMed-FL: A Robust Multimodal Federated Learning Framework for Clinical Diagnosis
arXiv:2609.10364v1 Announce Type: cross Abstract: Simultaneous assessment of medical imaging and patient records is often required in clinical diagnosis.…
MOONWALK: Mediating Operations with Intent-Evidence-Action Alignment Across Junior-Supervisor Review Workflows in Animation/VFX Pre-Production
arXiv:2609.10385v1 Announce Type: cross Abstract: Animation and VFX pre-production review requires teams to translate loosely specified creative…
Ex-Deepmind VP Vinyals says AI self-improvement is coming but won’t trigger an intelligence explosion
Oriol Vinyals, until recently head of research at Google DeepMind, thinks a sudden AI intelligence explosion through recursive self-improvement is…
Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization
arXiv:2609.10410v1 Announce Type: cross Abstract: The growing complexity of content moderation policies presents a critical challenge for their consistent…
