15 posts published in the last hour 23:32ContextSniper: AntTrail’s Token-Efficient Code Memory for Repository-Level Program Repair 23:32ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System 23:32A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice 23:32OpenAI…
Author: script
ContextSniper: AntTrail’s Token-Efficient Code Memory for Repository-Level Program Repair
arXiv:2607.01916v5 Announce Type: replace Abstract: Large language model agents can repair real repository issues, but they often spend large context…
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
arXiv:2607.14178v3 Announce Type: replace Abstract: Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex…
A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice
arXiv:2606.06081v2 Announce Type: replace Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration.…
OpenAI Closes on Anthropic in Ramp’s Business Spending Data
OpenAI is growing faster than Anthropic among U.S. businesses so far this quarter, according to new Ramp spending data posted August 20, 2026 by the…
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
arXiv:2606.21654v2 Announce Type: replace Abstract: Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop…
Replit expands access to software creation with GPT-5.6 Luna
Replit introduces Free Mode, powered by GPT-5.6 Luna, so anyone can turn ideas into working software without worrying about token costs.
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
arXiv:2606.16149v4 Announce Type: replace Abstract: Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer;…
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
arXiv:2605.02782v2 Announce Type: replace Abstract: Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech.…
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
arXiv:2605.22664v5 Announce Type: replace Abstract: LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts…
