16 posts published in the last hour 15:33xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems 15:33Cohere Debuts Open-Weight 218B Mixture-of-Experts Machine Translation Model 15:33Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging 15:33Claude…
Author: script
xDailyBench: Benchmarking LLMs on Professional Consultation for Real-Life Problems
arXiv:2609.07784v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only…
Cohere Debuts Open-Weight 218B Mixture-of-Experts Machine Translation Model
Cohere released North Small Translate on September 10, 2026, an open-weight mixture-of-experts machine translation model with 218 billion total parameters…
Understanding the Impact of Model Pruning on Long-Tail Forgetting and Explanation Reliability in Medical Imaging
arXiv:2609.07803v1 Announce Type: new Abstract: Model pruning is widely used to compress deep neural networks, reducing memory and computational…
Claude Fable 5.1’s language is less “load-bearing” than its predecessor’s
Arena.ai analyzed how Claude’s writing changed from Fable 5 to Fable 5.1 across tens of thousands of benchmark responses. Fable 5.1 writes more…
Do Large Language Models Know What They Don’t Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty
arXiv:2609.07879v1 Announce Type: new Abstract: Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question…
Now everyone can put data to work
Meet the Data agent in ChatGPT Work. Connect company data, uncover insights, and build interactive dashboards with AI using natural language.
When Intelligence Becomes Agency: A Theory of Governed, Proactive Agency for Symbiotic AI Systems
arXiv:2609.07741v1 Announce Type: new Abstract: Persistent AI assistants are intended to extend human attention, memory, and coordination across changing…
Expanding AI access and cyber defense for federal, state, local, and tribal governments
OpenAI and GSA will offer eligible federal, state, local, and tribal governments $0 license fees, 50% off usage, and expanded cyber defense support.
What Does an LLM-Agent Leaderboard Rank Actually Compare?
arXiv:2609.07785v1 Announce Type: new Abstract: An LLM-agent leaderboard invites a familiar inference: an agent ranked above another is the better agent.…
