arXiv:2608.26235v1 Announce Type: new Abstract: Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment…
LLM Agents for Time-Series: A Survey
arXiv:2608.26226v1 Announce Type: new Abstract: LLM-based agents are increasingly being developed for time-series problems, but their design choices vary…
Same Model, Different Harness: Different Coding-Agent Results
arXiv:2608.26218v1 Announce Type: new Abstract: A coding agent combines a model with a harness, which decides what the model sees, which tools it can use,…
AI News Brief Hourly Summary 2026-08-28 09h : 13 posts
13 posts published in the last hour 06:32Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs 06:32Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling 06:32Predicting Consequences and Reinforcing Navigation Policies with Latent World Models 06:32Agentic…
Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs
arXiv:2608.26191v1 Announce Type: new Abstract: Incident risk prediction from longitudinal electronic health records (EHRs) is challenging because…
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
arXiv:2608.26199v1 Announce Type: new Abstract: We ask whether AI agents powered by locally deployed large language models can reliably automate…
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
arXiv:2608.26190v1 Announce Type: new Abstract: World models enable agents to reason about future outcomes and learn policies from their knowledge of…
Agentic AI for operating scientific instruments for nanoscale characterization
arXiv:2608.26198v1 Announce Type: new Abstract: Operating a scientific instrument such as an atomic force microscope (AFM) requires continuous expert…
NVIDIA Posts $96.2B Quarter as Data Center Revenue Hits $89B
NVIDIA reported revenue of $96.2 billion for the second quarter of fiscal 2027, ended July 26, 2026, up 18% from the previous quarter and up 106% from a…
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
arXiv:2608.26193v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet…
Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion
arXiv:2608.26185v1 Announce Type: new Abstract: Equal participation in co-located discussion is important for effective collaboration, yet people often…
TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education
arXiv:2608.26184v1 Announce Type: new Abstract: AI programming tutors provide scalable support, yet lack the behavioral context human tutors rely on to…
Is Your Neighborhood Safe? Place-based Stigma in Large Language Models’ Urban Safety Judgments
arXiv:2608.26188v1 Announce Type: new Abstract: Large language models are increasingly used to inform safety decisions in cities, such as where it is safe…
Invocation-Level Reliability of Tool-Using Agents
arXiv:2608.26189v1 Announce Type: new Abstract: Tool-using agents fail two ways: choosing the wrong tool, or forming wrong arguments, and an early failure…
Deep Cogito Raises $43M Series A to Build the Post-Training Engine for Self-Improving AI
Deep Cogito has raised a $43 million Series A as the San Francisco AI lab looks to scale an increasingly important part of the artificial intelligence…
Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI
arXiv:2608.26182v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for verbal interaction in social robots, yet prompt…
AI News Brief Hourly Summary 2026-08-28 08h : 18 posts
18 posts published in the last hour 05:33A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving 05:33Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain…
A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving
arXiv:2608.26164v1 Announce Type: new Abstract: Large language models can interpret natural-language chemistry questions, but their internal reasoning is…
