Nvidia is moving its Groq 3 LPX inference chip into full production and reports 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras.…
Category: AI
Where World Models Break: Natural-Input Failure Discovery
arXiv:2608.22421v1 Announce Type: new Abstract: World models predict action-conditioned futures and serve as critical internal simulators for downstream…
HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory
arXiv:2608.22310v1 Announce Type: new Abstract: Long-term memory is crucial for personalized responses and long-horizon agent interactions. Existing…
Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture
arXiv:2608.22347v1 Announce Type: new Abstract: A cognitive architecture is more than the module that reasons: it must also decide how long to think and…
Read Less, Solve More: Token-Efficient Sparse Reading for AI Agents
arXiv:2608.22237v1 Announce Type: new Abstract: Long-horizon agents increasingly rely on repeated access to external artifacts, yet current reading…
Clarify User Expertise: Towards Proactive Conversational Agents Tailoring Responses to User Proficiency
arXiv:2608.22266v1 Announce Type: new Abstract: In the context of information seeking, conversational agents are undergoing an evolution from reactive…
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
This post has no text preview — click the link below to read the original article. This article has been indexed from Hugging Face – Blog Read the original article: Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision…
Addressing the Selection Problem in Explainable AI
arXiv:2608.22356v1 Announce Type: new Abstract: Explainable AI (XAI) research has produced a plethora of explanation techniques, yet user studies…
Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents
arXiv:2608.22191v1 Announce Type: new Abstract: Software-engineering agents solve repository-level tasks through long, stochastic tool-use trajectories,…
MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning
arXiv:2608.22167v1 Announce Type: new Abstract: Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language…
