Amazon SageMaker Feature Store now supports two new APIs: BatchWriteRecord writes up to 25 records across multiple feature groups in a single call, and…
PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?
arXiv:2608.26882v1 Announce Type: cross Abstract: Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked…
Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic
arXiv:2608.26860v1 Announce Type: cross Abstract: Connected and automated vehicle (CAV) platooning offers a promising approach to improving road safety…
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
arXiv:2608.26856v1 Announce Type: cross Abstract: Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical…
Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay
arXiv:2608.26846v1 Announce Type: cross Abstract: Interactive language-model agents use confidence signals to decide whether to answer immediately,…
GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture
Z.ai and Qwen independently shipped near-identical architectures: 3:1 linear hybrids, compressed indexers, gated residuals, and Muon training.
Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory
arXiv:2608.26807v1 Announce Type: cross Abstract: Travel planning agents assist users in generating personalized travel plans by modeling their individual…
An Anthropic researcher just gave us a peek at self-improving AI
Given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading…
MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
arXiv:2608.26848v1 Announce Type: cross Abstract: Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet…
LiveVVT: High-Fidelity Video Virtual Try-On in Real Time
arXiv:2608.26714v1 Announce Type: cross Abstract: Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional…
Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning
arXiv:2608.26732v1 Announce Type: cross Abstract: Graph neural networks (GNNs) are typically conceptualized as message-passing neural networks, yet it…
Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research
arXiv:2608.26753v1 Announce Type: cross Abstract: LLM agents used for scientific experimentation must do more than generate executable code: they must…
Google Deepmind’s AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers
Google Deepmind has expanded Co-Scientist from a hypothesis generator into a research system that’s integrated into the lab. Across three disciplines,…
FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs
arXiv:2608.26746v1 Announce Type: cross Abstract: Generated operational programs are often validated with either a few hand-written examples or exhaustive…
Blue Owl Funds Lead $2.4B AI Factory Equipment Financing for IREN
Blue Owl Capital announced on August 28, 2026 that funds it manages are leading a $2.4 billion compute equipment financing for IREN Limited, the…
Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
arXiv:2608.26733v1 Announce Type: cross Abstract: Agent skills bundle instructions, reference data, and executable helpers that let a general agent…
AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability
arXiv:2608.26713v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have extended Image Aesthetic Assessment…
FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models
arXiv:2608.26676v1 Announce Type: cross Abstract: Pruning is a practical approach to compress large language models (LLMs), but it can amplify text…
