StepFun has released Step 5 Preview, a sparse Mixture-of-Experts model with 600B total parameters and 27B active per token. It supports a 1M-token context…
Tag: AI
Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education
arXiv:2609.21600v1 Announce Type: new Abstract: Access to academic support is a key determinant of student success, yet students experience it unequally:…
PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design
arXiv:2609.21493v1 Announce Type: new Abstract: Multimodal large language models, or MLLMs, perform well at visual understanding and structured…
Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks
arXiv:2609.21519v1 Announce Type: new Abstract: Artificial Intelligence (AI) is becoming a fundamental design principle of future AI-native communication…
The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models
arXiv:2609.21509v1 Announce Type: new Abstract: When language models reason in chain-of-thought or exchange free-text intermediates, they serialize…
Dual-Interest Sequential Product Recommendation With Multi-Granular SSM
arXiv:2609.21548v1 Announce Type: new Abstract: Sequential recommendation aims to predict the next item a user will interact with based on their…
LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers
arXiv:2609.21492v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning has been shown to improve the performance of large language models…
Risk-Aware Occupancy for Safety-Oriented End-to-End Autonomous Driving
arXiv:2609.21470v1 Announce Type: new Abstract: Sparse representation formulates the environment perception for the end-to-end driving system as a set of…
Offline Multimodal Large Language Models for Decision Support in Air Operations
arXiv:2609.21390v1 Announce Type: new Abstract: Air operations rely on complex rules, established procedures, and time-critical analysis under limited…
GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation
arXiv:2609.21432v1 Announce Type: new Abstract: Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of…
