Zoph, who co-founded Thinking Machines Lab alongside Mira Murati and also served as the startup’s CTO, led a brief stint at OpenAI and is now at Google.
VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models
arXiv:2608.24302v1 Announce Type: new Abstract: Long-video understanding depends critically on how a limited model context is constructed from a much…
AI News Brief Hourly Summary 2026-08-27 22h : 14 posts
14 posts published in the last hour 19:32Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing 19:32STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation 19:32SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction 19:32Beyond Accuracy: A Dual-Judge…
Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing
arXiv:2608.24263v1 Announce Type: new Abstract: Change data synthesis provides a cost-effective solution for expanding training data and improving the…
STRIVE: Multi-Agent Structured Temporal Reasoning with Integrated Verification for Longitudinal Radiology Report Generation
arXiv:2608.24237v1 Announce Type: new Abstract: Longitudinal radiology report generation (LRRG) requires identifying both current findings and their…
SA-Bench: Evaluating Semantic Alignment in LLM-Based Paper Reproduction
arXiv:2608.24252v1 Announce Type: new Abstract: LLM agents can generate paper reproduction code, yet often produce scientifically unfaithful…
Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks
arXiv:2608.24258v1 Announce Type: new Abstract: AI systems are increasingly evaluated for legally accountable settings, where correct outputs must also be…
Matched Excess-Outranker Regularization for Candidate-Set Interference in Continual Knowledge Graph Embedding
arXiv:2608.24273v2 Announce Type: new Abstract: Continual knowledge graph embedding updates entity and relation representations as a graph grows. Existing…
MetaRAG: Belief-Action Aligned Policy Optimization for Agentic RAG
arXiv:2608.24214v1 Announce Type: new Abstract: Agentic retrieval-augmented generation (RAG) requires language models to decide when to continue searching…
Constraint-Guided Enterprise Data Mapping with Large Language Models
arXiv:2608.24218v1 Announce Type: new Abstract: Enterprise entity alignment must handle semi-structured records, implicit attributes, and unit or…
Introducing India cross-Region inference for OpenAI GPT-5.6 models on Amazon Bedrock
Amazon Bedrock now supports the OpenAI GPT-5.6 models, Terra and Luna, in India with India geographic cross-Region inference. If you have local data…
Preference Data Selection for Mitigating the Alignment Tax in Large Language Models
arXiv:2608.24192v1 Announce Type: new Abstract: Aligning large language models to human preferences is crucial for real-world deployment but frequently…
Canada Is Luring AI and Science Talent as Trump Upends U.S. Research
For decades, the United States benefited from one of the most powerful competitive advantages in science: many of the world’s best researchers wanted to…
Evaluating Multiple LLM Generations with Validated Task Coverage
arXiv:2608.24228v1 Announce Type: new Abstract: Many LLM applications are most useful when they provide several candidate outputs for comparison,…
AI models flub these intelligence tests. Can you fare any better?
Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic…
TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models
arXiv:2608.24232v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even…
AI News Brief Hourly Summary 2026-08-27 21h : 19 posts
19 posts published in the last hour 18:32How loveholidays is making everyone a builder with Codex 18:32AI shopping agents aren’t ready to buy on your behalf, study finds 18:32Task-Adaptive Rubrics for GUI Reward Modeling 18:32Gatik raises $200M to scale AI-powered…
How loveholidays is making everyone a builder with Codex
Discover how loveholidays uses OpenAI Codex to make software development accessible across the business, helping teams turn ideas into products faster.
