arXiv:2607.09770v2 Announce Type: replace Abstract: Industrial agentic AI systems increasingly exhibit a gap between prototype capability and production…
Tag: cs.AI updates on arXiv.org
SGA: Plug&Play Geometric Verification for Educational Video Synthesis
arXiv:2607.18116v2 Announce Type: replace Abstract: Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical…
Relevance Is Not Permission: Localizing and Controlling Metric-Facing Attention Contributions
arXiv:2606.30139v3 Announce Type: replace Abstract: Attention identifies items relevant to a current query, but does not separately determine whether…
Cognitive Digital Twins: Ethical Risks and Governance for AI Systems That Model the Mind
arXiv:2606.23094v2 Announce Type: replace Abstract: As AI systems become increasingly persistent and personalized, they make possible a class of…
Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP
arXiv:2606.13720v2 Announce Type: replace Abstract: Arditi et al. (2024) has shown that refusal in safety fine-tuned chat models is mediated by a single…
Some hypotheses on how chatbots work in problem-solution-driven conversations: Large Language Models as confirmation of the Innovation Illusion
arXiv:2606.07722v5 Announce Type: replace Abstract: We discuss the nature of chatbots as conversation partners in problem-solving conversations. What can…
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale
arXiv:2604.17406v5 Announce Type: replace Abstract: The convergence of large language models and agents is catalyzing a new era of scientific discovery:…
Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
arXiv:2604.15210v2 Announce Type: replace Abstract: Humor is one of the few cognitive tasks where getting the reasoning right matters as much as getting…
A Lightweight Multi-Agent Framework for Automated Concrete Barrier Design
arXiv:2606.12040v3 Announce Type: replace Abstract: The design of reinforced concrete (RC) highway barriers is a safety-critical engineering task that…
Rescaling Confidence: What Scale Design Reveals About LLM Metacognition
arXiv:2603.09309v3 Announce Type: replace Abstract: Verbalized confidence, in which LLMs report a numerical certainty score, is widely used to estimate…
