arXiv:2608.23373v1 Announce Type: new Abstract: Forecasting long-range influenza-like illness (ILI) matters for public health readiness. Publicly…
Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning
arXiv:2608.23318v1 Announce Type: new Abstract: Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a…
Russia used ChatGPT to run a covert influence campaign pushing pro-Kremlin narratives across the West
OpenAI has disrupted a covert Russian influence campaign that used ChatGPT to generate social media posts by banning a cluster of accounts. The operators…
MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction
arXiv:2608.23397v1 Announce Type: new Abstract: Interactive clinical agents must gather decisive evidence and convert it into grounded actions under…
Agentic observability with Amazon OpenSearch Service MCP Apps
Amazon OpenSearch Service now supports MCP Apps, which return interactive visualizations alongside your AI agent’s text responses. Learn how a single,…
SkillAlchemy: Open-World Agent Skill Creation
arXiv:2608.23417v1 Announce Type: new Abstract: Agent skills are reusable procedural artifacts that extend language agents with specialized workflows,…
AI News Brief Hourly Summary 2026-08-25 21h : 15 posts
15 posts published in the last hour 18:32EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models 18:32Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance 18:32Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data 18:32Automated…
EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models
arXiv:2608.23313v1 Announce Type: new Abstract: Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses,…
Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
arXiv:2608.23264v1 Announce Type: new Abstract: Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these…
Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data
arXiv:2608.23256v1 Announce Type: new Abstract: Recent work proposes next-chunk reasoning RL for leveraging no-CoT data—corpora such as worked solutions…
Automated Construction of FAIR Digital Object Knowledge Graphs from Flat Cultural Heritage Records
arXiv:2608.23263v1 Announce Type: new Abstract: The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent…
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
arXiv:2608.23283v1 Announce Type: new Abstract: General-purpose language models can reason and synthesize knowledge, but complex work also requires…
What is mathematics now, and what should it be?
arXiv:2608.23218v1 Announce Type: new Abstract: Advances in neural theorem provers have been impressive, but the successes obscure a broader vision of…
Mikhail Yatsuha, CEO and Co-Founder of CaseCraft.AI
Mikhail Yatsuha, CEO and Co-Founder of CaseCraft.AI, is a UK solicitor and legal technology entrepreneur with more than a decade of experience in legal…
AI emotional support is better only when chosen, but shifts preferences even when it is not
arXiv:2608.23196v1 Announce Type: new Abstract: People increasingly face a novel decision when seeking emotional support: human or AI. In existing…
OpenAI’s first custom chip “Jalapeño” reportedly beats Nvidia’s Blackwell and Rubin in inference benchmarks
OpenAI showed off “Jalapeño,” its first in-house inference chip, with benchmarks at the Hot Chips conference. According to SemiAnalysis tests, the chip…
Cognitive Profiling of LRMs’ Reasoning Traces Using Bloom’s Taxonomy
arXiv:2608.23205v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have revolutionized reasoning in LLMs, and the increasing public…
Introducing the Admin plugin for ChatGPT Work and Codex
Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.
