arXiv:2609.21325v1 Announce Type: new Abstract: Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete…
Tag: cs.AI updates on arXiv.org
Efficient Benchmarking in Production: A Study of an Evolving LLM Agent
arXiv:2609.21267v1 Announce Type: new Abstract: Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to…
GameASG-Bench: Benchmarking Autonomous Software Generation for Game Development
arXiv:2609.21293v1 Announce Type: new Abstract: Autonomous software generation (ASG) aims to turn human requirements into executable applications, but…
PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking
arXiv:2609.21263v1 Announce Type: new Abstract: Automated macro placement remains a fundamental challenge in VLSI physical design. Despite decades of…
Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation
arXiv:2609.21208v1 Announce Type: new Abstract: Self-play methods that co-train a single language model as both coder and test author promise to move…
Ability-Residual Decoupled Modeling for Affective Cognitive Diagnosis
arXiv:2609.21214v1 Announce Type: new Abstract: Cognitive diagnosis infers students’ concept mastery from response logs. However, students’ responses are…
A Fully Differentiable Neuro-Soft-Symbolic Framework for Perceptual Task Planning
arXiv:2609.21221v1 Announce Type: new Abstract: Perceptual planning tasks require two key capabilities: accurately perceiving uncertain scenes and…
AI-GRACE: A Use-Case Operationalization Framework for Agentic AI: From Organizational Objectives and Obligations to Deployment Capabilities and Architecture
arXiv:2609.21192v1 Announce Type: new Abstract: Organizations deploying agentic artificial intelligence must determine more than whether a model is…
Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks
arXiv:2609.21181v1 Announce Type: new Abstract: The Abstraction and Reasoning Corpus and related benchmarks evaluate whether AI models can solve novel…
TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers
arXiv:2609.21139v1 Announce Type: new Abstract: Replacing attention in a pretrained language model is a compatibility problem: a plausible substitute may…
