arXiv:2609.08944v1 Announce Type: new Abstract: Agent skills provide a lightweight way to equip frozen language-model agents with domain knowledge and…
Tag: cs.AI updates on arXiv.org
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
arXiv:2609.08861v1 Announce Type: new Abstract: Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape public…
Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
arXiv:2609.08966v1 Announce Type: new Abstract: Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that…
Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course
arXiv:2609.08832v1 Announce Type: new Abstract: Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a…
GoAnt: Quality-Diversity Multi-Agent Search for Alpha Factor Discovery in Market Microstructure Data
arXiv:2609.08719v1 Announce Type: new Abstract: Automated alpha factor discovery searches symbolic trading signals from price-volume panels and order-book…
CLAMP: Constrained Decoding for Vision-Language Embodied Planning
arXiv:2609.08602v1 Announce Type: new Abstract: Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and…
It’s All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction
arXiv:2609.08772v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly being investigated for physiological time-series prediction,…
When Can One Obtain Certificates of Optimality Using Positivstellensaetze?
arXiv:2609.08736v1 Announce Type: new Abstract: We study certificates of positivity and optimality for learning problems whose objectives and constraints…
Application of curiosity driven exploration methods for hardware interference identification
arXiv:2609.08729v1 Announce Type: new Abstract: The transition from single-core to multi-core architectures in safety-critical embedded systems introduces…
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
arXiv:2609.08572v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing…
