arXiv:2605.20555v2 Announce Type: replace-cross Abstract: We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT)…
Author: script
Cognition helps Devin test its own work with GPT‑6 Astra
GPT‑6 Astra improves Devin’s ability to test software and show that it works, with the goal of helping engineers review less code and ship more.
“What Are You Really Trying to Do?”: Co-Creating Life Goals from Everyday Computer Use
arXiv:2605.00497v2 Announce Type: replace-cross Abstract: Recent advances in user modeling make it feasible to conduct open-ended inference over a…
DGCPath: Distribution-Aware Generative Contrastive Framework for Self-supervised Path Representation Learning — Extended Version
arXiv:2609.07316v2 Announce Type: replace Abstract: Due to the proliferation of vehicle trajectory data enabled by advanced sensing technologies, path…
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
arXiv:2604.10701v2 Announce Type: replace-cross Abstract: Credit assignment is a central challenge in reinforcement learning (RL). Classical actor-critic…
Where is the Mind? Persona Vectors and LLM Individuation
arXiv:2604.17031v3 Announce Type: replace-cross Abstract: The individuation problem for large language models asks which entities associated with them, if…
Spec-Harness: Measuring and Improving Behavioral Adequacy of LLM-Synthesized Formal Specifications
arXiv:2604.00280v2 Announce Type: replace-cross Abstract: Formal specifications play a central role in ensuring software reliability, yet automatically…
The Biggest Risk of Embodied AI is Governance Lag
arXiv:2604.21938v2 Announce Type: replace-cross Abstract: Embodied AI is widely discussed as a job-displacement problem. The deeper risk, however, is…
AI News Brief Hourly Summary 2026-09-12 02h : 14 posts
14 posts published in the last hour 23:32MOSAIC: A Universal Agent-Level Interface for Cross-Paradigm Agent Mixing and Human-AI Collaboration 23:32False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UK 23:32City Editing: Hierarchical Agentic Execution for…
MOSAIC: A Universal Agent-Level Interface for Cross-Paradigm Agent Mixing and Human-AI Collaboration
arXiv:2603.01260v3 Announce Type: replace-cross Abstract: Existing infrastructure cannot deploy agents from different decision-making paradigms within the…
