arXiv:2602.03702v2 Announce Type: replace-cross Abstract: Large language models are increasingly trained in continual or open-ended settings, where the…
Author: script
Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together
Model cards report quality under server-class, full-precision conditions. Those numbers rarely predict how the same model behaves on a phone. This week,…
You Can Learn Tokenization End-to-End with Reinforcement Learning
arXiv:2602.13940v3 Announce Type: replace-cross Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Large…
AI News Brief Hourly Summary 2026-08-27 11h : 11 posts
11 posts published in the last hour 08:32Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes 08:32Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model 08:32Ad Insertion in LLM-Generated Responses 08:32TangramPuzzle: Evaluating Multimodal Large Language…
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
arXiv:2601.07737v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in mainstream…
Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model
arXiv:2601.10034v3 Announce Type: replace-cross Abstract: Decision making often exhibits context dependence that is difficult to accommodate within a…
Ad Insertion in LLM-Generated Responses
arXiv:2601.19435v2 Announce Type: replace-cross Abstract: Sustainable monetization of large language models (LLMs) remains a critical open challenge.…
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
arXiv:2601.16520v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in visual recognition…
Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility
arXiv:2601.17027v2 Announce Type: replace-cross Abstract: While synthetic data has proven effective for improving scientific reasoning in the text domain,…
CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
arXiv:2511.01870v3 Announce Type: replace-cross Abstract: Studying the cellular architecture of the human cerebral cortex is essential for understanding…
