arXiv:2609.29508v1 Announce Type: new Abstract: Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over…
Tag: AI
BiGraph-Diffuse: A Bidirectional Diffusion Language Model with Graph-Structured Retrieval For Mental Health Counseling
arXiv:2609.29519v1 Announce Type: new Abstract: Mental health disorders affect hundreds of millions of people around the world, yet access to professional…
Ruby on Rails creator DHH says he’s done writing code by hand
David Heinemeier Hansson, creator of Ruby on Rails, has quit writing code by hand after 25 years. He says he hasn’t typed a single line since March 2026.…
Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets
arXiv:2609.29509v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as multi-turn agents that must sustain goals, use…
Airbnb widens access to GPT-6 Astra and OpenAI frontier models
Learn how Airbnb is expanding access to GPT-6 Astra and OpenAI frontier models to help engineering teams solve bugs, design systems, and ship faster.
Clinical Knowledge Graphs for Chest X-Ray Device Reasoning
arXiv:2609.29536v1 Announce Type: new Abstract: Chest radiographs are routinely used to verify the position of catheters, tubes, and other support…
RD-JEPA: Predictive latent pretraining for few-trajectory transfer across reaction–diffusion equations
arXiv:2609.29403v1 Announce Type: new Abstract: Learning surrogates for time-dependent partial differential equations often requires a new simulation…
An auditable conditional-strategy framework for open-ended decision-making in complex lung cancer
arXiv:2609.29381v1 Announce Type: new Abstract: Complex lung cancer decisions can involve several defensible pathways whose eligibility, sequencing and…
The Pentagon wants $30 million to build an AI-powered lie detector
The US government wants to spend $30.3 million over the next five years on an improved form of lie detector, according to a Department of Defense budget…
SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories
arXiv:2609.29465v1 Announce Type: new Abstract: Large language model based coding agents have made substantial progress on repository-level software…
