arXiv:2608.17319v1 Announce Type: new Abstract: Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally…
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
arXiv:2608.17336v1 Announce Type: new Abstract: Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic…
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models
arXiv:2608.17299v1 Announce Type: new Abstract: Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for…
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
arXiv:2608.17330v1 Announce Type: new Abstract: Large language models for medical consultation are often evaluated after a clinical problem has already…
ChatGPT Ads expands across Europe
ChatGPT Ads is expanding to 31 European markets. Learn how advertisers can reach people as they explore, compare options, and make decisions.
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning
arXiv:2608.17301v1 Announce Type: new Abstract: Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable…
AI News Brief Hourly Summary 2026-08-19 08h : 14 posts
14 posts published in the last hour 05:32ASI-Bench: At the Dawn of Artificial Superintelligence 05:32PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs 05:32Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for…
ASI-Bench: At the Dawn of Artificial Superintelligence
arXiv:2608.17271v1 Announce Type: new Abstract: Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward…
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs
arXiv:2608.17289v1 Announce Type: new Abstract: Group-relative policy optimization has emerged as a key paradigm for training agentic large language…
Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking
arXiv:2608.17270v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for scientific hypothesis generation. However,…
Build OpenClaw agents that transact with Amazon Bedrock AgentCore payments
Give an autonomous agent a wallet and spending guardrails so it can pay for paywalled APIs, MCP servers, and web content. This post connects OpenClaw to…
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation
arXiv:2608.17282v1 Announce Type: new Abstract: Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing…
Groq raises $350M to fuel its pivot from AI chips to neocloud
Groq raised $350 million at a $3.5 billion valuation as the former AI chipmaker pivots to a neocloud business and expands its Nvidia-powered data center…
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification
arXiv:2608.17247v1 Announce Type: new Abstract: Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried…
Fool’s Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
arXiv:2608.17202v1 Announce Type: new Abstract: Safety alignment in open-weight language models is trivially removable: abliteration projects a…
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models
arXiv:2608.17183v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive…
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn
arXiv:2608.17150v1 Announce Type: new Abstract: To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must…
A Constitution for the New Enterprise: 15 Rules for AI Governance
These essays were supposed to be about architecture. Read them back and notice what they actually did: they kept issuing rules. Approvals are not data;…
