arXiv:2609.22243v1 Announce Type: cross Abstract: Input-conditioned neural interventions raise a runtime question: what persists when one behavioral…
Category: cs.AI updates on arXiv.org
Do Chess Explanations Reflect Model Decisions? Behavioral and Token-Level Tests of LLM Reasoning Faithfulness
arXiv:2609.22245v1 Announce Type: cross Abstract: Large language models can produce fluent explanations for chess moves, but plausible language does not…
The Corroboration Illusion: When More News Makes LLM Forecasts Less True
arXiv:2609.22246v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to forecast real-world events by retrieving and…
CHART: A Harness-Rotation Curriculum for Harness-Robust Search Agents
arXiv:2609.22247v1 Announce Type: cross Abstract: Search agents are usually trained under a single harness. But once an agent is deployed in a real…
Universal Observatory Graphs for Distributed Sky Coverage and Artificial Intelligence Based Interplanetary Routing
arXiv:2609.22244v1 Announce Type: cross Abstract: This research proposes the Universal Observatory Graph (UOG), an AI-driven framework for distributed…
Team DArgk at the 2026 ELOQUENT lab for evaluating generative language model quality: Residuals of Humanity: AI Detection Evasion via GRPO Fine-Tuning
arXiv:2609.22221v1 Announce Type: cross Abstract: Large language models (LLMs) can generate fluent and coherent text that is increasingly difficult to…
Knowledge Graph-Augmented Ambient AI for Clinical Note Generation
arXiv:2609.22239v1 Announce Type: cross Abstract: Ambient AI is increasingly adopted in healthcare to automatically generate clinical notes from…
Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks
arXiv:2609.22237v1 Announce Type: cross Abstract: Merging low-rank adapters (LoRAs) promises to eliminate the overhead of swapping task-specific weights…
H2LooP Telecom Model v1: From Telecom Comprehension to Autonomous Issue and PR Resolution
arXiv:2609.22241v1 Announce Type: cross Abstract: We present H2LooP Telecom Model v1, a domain-specialized large language models fine-tuned for the…
Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark
arXiv:2609.22222v1 Announce Type: cross Abstract: Large language models can generate executable data-analysis code, but successful execution is not…
