arXiv:2608.23941v1 Announce Type: new Abstract: Pre-execution oversight is core to trusted monitoring in AI control: a fallible LLM monitor vets planned…
Tag: AI
How to encourage smarter AI use in the classroom
This article is from Making AI Work, MIT Technology Review’s limited-run newsletter examining how to apply LLMs across industries. To receive it in your…
More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving
arXiv:2608.23962v1 Announce Type: new Abstract: When an LLM serving deployment runs out of KVcache room, there are two well-established ways out. Tensor…
Evolutionary Recurrent Decision Model in Developing Adaptive and Maladaptive Behaviors
arXiv:2608.23932v1 Announce Type: new Abstract: This study introduces the evolutionarily recurrent decision model (ERDM), a computational reinforcement…
Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining
arXiv:2608.23922v1 Announce Type: new Abstract: Data mixing is a central design problem in large language model pretraining: given a fixed token budget,…
PROOF-Gen: From Optimized Data to Better Distillation
arXiv:2608.23911v1 Announce Type: new Abstract: Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling…
Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2×2 factorial experiment
arXiv:2608.23908v1 Announce Type: new Abstract: Tax-loss harvesting demonstrates consistent benefits to long-term portfolio growth; yet implementing it…
OpenAI loses a top data center exec as stream of high-profile departures continues
In a statement to TechCrunch about Malone’s departure, OpenAI said it had “recently reorganized” its “infrastructure organization to support the scale and…
MARS: Multi-Specialist LLM Relay System for Competitive Programming
arXiv:2608.23918v1 Announce Type: new Abstract: Large Language Models excel at code generation, yet competitive programming exposes a persistent failure…
BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification
arXiv:2608.23898v1 Announce Type: new Abstract: We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks…
