StepFun has released Step 5 Preview, a sparse Mixture-of-Experts model with 600B total parameters and 27B active per token. It supports a 1M-token context…
Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education
arXiv:2609.21600v1 Announce Type: new Abstract: Access to academic support is a key determinant of student success, yet students experience it unequally:…
AI News Brief Hourly Summary 2026-09-21 09h : 10 posts
10 posts published in the last hour 06:32PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design 06:32Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks 06:32The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models 06:32Dual-Interest Sequential Product…
PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
arXiv:2609.21493v1 Announce Type: new Abstract: Multimodal large language models, or MLLMs, perform well at visual understanding and structured…
Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks
arXiv:2609.21519v1 Announce Type: new Abstract: Artificial Intelligence (AI) is becoming a fundamental design principle of future AI-native communication…
The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
arXiv:2609.21509v1 Announce Type: new Abstract: When language models reason in chain-of-thought or exchange free-text intermediates, they serialize…
Dual-Interest Sequential Product Recommendation With Multi-Granular SSM
arXiv:2609.21548v1 Announce Type: new Abstract: Sequential recommendation aims to predict the next item a user will interact with based on their…
LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers
arXiv:2609.21492v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models…
Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving
arXiv:2609.21470v1 Announce Type: new Abstract: Sparse representation formulates the environment perception for the end-to-end driving system as a set of…
Offline Multimodal Large Language Models for Decision Support in Air Operations
arXiv:2609.21390v1 Announce Type: new Abstract: Air operations rely on complex rules, established procedures, and time-critical analysis under limited…
GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation
arXiv:2609.21432v1 Announce Type: new Abstract: Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of…
Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous Driving
arXiv:2609.21486v1 Announce Type: new Abstract: Multimodal trajectory prediction improves behavioral coverage in end-to-end autonomous driving, but…
DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
arXiv:2609.21423v1 Announce Type: new Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert…
AI News Brief Hourly Summary 2026-09-21 08h : 10 posts
10 posts published in the last hour 05:33CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition 05:33LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces 05:33Efficient Benchmarking in Production: A Study of an Evolving LLM Agent 05:32GameASG-Bench: Benchmarking Autonomous Software…
CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
arXiv:2609.21259v1 Announce Type: new Abstract: Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI)…
LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces
arXiv:2609.21325v1 Announce Type: new Abstract: Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete…
Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
arXiv:2609.21267v1 Announce Type: new Abstract: Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to…
GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development
arXiv:2609.21293v1 Announce Type: new Abstract: Autonomous software generation (ASG) aims to turn human requirements into executable applications, but…
