arXiv:2608.13760v1 Announce Type: cross Abstract: Which reasoning behaviors are associated with correct answers in reasoning models, and does…
Category: cs.AI updates on arXiv.org
CutClean: Neural Network Pruning for Privacy-Preserving Inference
arXiv:2608.13773v1 Announce Type: cross Abstract: Neural networks are increasingly deployed in high-stakes applications with growing privacy leakage…
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
arXiv:2608.13786v1 Announce Type: cross Abstract: Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to…
Capacity-Dependent Effects of Data Selection for Reasoning
arXiv:2608.13721v1 Announce Type: cross Abstract: In reasoning supervised fine-tuning, candidate responses for the same instruction can differ…
Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring Development Cost
arXiv:2608.13730v1 Announce Type: cross Abstract: Empirical reports on the true cost of AI-intensive software development remain scarce, and the few that…
Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation
arXiv:2608.13712v1 Announce Type: cross Abstract: Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations…
TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials
arXiv:2608.13708v1 Announce Type: cross Abstract: Automatically generating textbook-grounded assessment items can reduce science teachers’ workload, but…
Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and Structured Interventions against a Natural-Language Baseline
arXiv:2608.13742v1 Announce Type: cross Abstract: In LLM-based code generation, Non-Functional Requirements (NFRs) are often specified as terse one-line…
Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT
arXiv:2608.13681v1 Announce Type: cross Abstract: Translating C code into safe, idiomatic Rust is a longstanding software-engineering goal because it can…
From BERT to Frontier Agents: Eight Years of Language-Model Progress, the Collapse of the Capability-Cost Curve, and the Rise of Task-Targeted Models
arXiv:2608.13675v1 Announce Type: cross Abstract: Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents…
