arXiv:2608.13786v1 Announce Type: cross Abstract: Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to…
Category: AI
Capacity-Dependent Effects of Data Selection for Reasoning
arXiv:2608.13721v1 Announce Type: cross Abstract: In reasoning supervised fine-tuning, candidate responses for the same instruction can differ…
Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring Development Cost
arXiv:2608.13730v1 Announce Type: cross Abstract: Empirical reports on the true cost of AI-intensive software development remain scarce, and the few that…
AirTag reveals how Amazon destroys rare books for AI training
Amazon buys large quantities of printed books, scans them as AI training data, and destroys them in the process. The article AirTag reveals how Amazon…
Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation
arXiv:2608.13712v1 Announce Type: cross Abstract: Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations…
The Hidden Cost of AI: How “Black Box” Models Are Eroding Trust, Budgets, and the Environment
The AI industry has a model design problem. Enterprises are starting to feel it in cost, performance, and reliability. AI can talk – but can it listen?…
TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials
arXiv:2608.13708v1 Announce Type: cross Abstract: Automatically generating textbook-grounded assessment items can reduce science teachers’ workload, but…
Axiom Math’s AI Verifies the 246 Prime-Gaps Theorem in Lean
Axiom Math says its AxiomProver system has produced a machine-checked Lean 4 proof of the strongest known result on gaps between prime numbers: the…
Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline
arXiv:2608.13742v1 Announce Type: cross Abstract: In LLM-based code generation, Non-Functional Requirements (NFRs) are often specified as terse one-line…
Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT
arXiv:2608.13681v1 Announce Type: cross Abstract: Translating C code into safe, idiomatic Rust is a longstanding software-engineering goal because it can…
