arXiv:2608.24275v1 Announce Type: new Abstract: Safeguarding language model agents requires assessing complete execution trajectories under…
Author: script
Barret Zoph, the Thinking Machines co-founder who defected to OpenAI, is now at Google
Zoph, who co-founded Thinking Machines Lab alongside Mira Murati and also served as the startup’s CTO, led a brief stint at OpenAI and is now at Google.
VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models
arXiv:2608.24302v1 Announce Type: new Abstract: Long-video understanding depends critically on how a limited model context is constructed from a much…
AI News Brief Hourly Summary 2026-08-27 22h : 14 posts
14 posts published in the last hour 19:32Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing 19:32STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation 19:32SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction 19:32Beyond Accuracy: A Dual-Judge…
Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing
arXiv:2608.24263v1 Announce Type: new Abstract: Change data synthesis provides a cost-effective solution for expanding training data and improving the…
STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation
arXiv:2608.24237v1 Announce Type: new Abstract: Longitudinal radiology report generation (LRRG) requires identifying both current findings and their…
SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction
arXiv:2608.24252v1 Announce Type: new Abstract: LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful…
Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks
arXiv:2608.24258v1 Announce Type: new Abstract: AI systems are increasingly evaluated for legally accountable settings, where correct outputs must also be…
Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding
arXiv:2608.24273v2 Announce Type: new Abstract: Continual knowledge graph embedding updates entity and relation representations as a graph grows. Existing…
MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG
arXiv:2608.24214v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue searching…
