arXiv:2608.20686v1 Announce Type: new Abstract: Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain…
Category: AI
DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning
arXiv:2608.20717v1 Announce Type: new Abstract: Reliable confidence estimation is essential for using large language models in mathematical reasoning, but…
SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL
arXiv:2608.20630v1 Announce Type: new Abstract: SQL systems increasingly expose AI functions for tasks such as classification, extraction, filtering,…
Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work
arXiv:2608.20622v1 Announce Type: new Abstract: Frontier models have collapsed the cost of writing custom code: a niche problem a specialist sees in their…
Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents
arXiv:2608.20631v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring…
Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions
arXiv:2608.20649v1 Announce Type: new Abstract: Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces,…
Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance
arXiv:2608.20661v1 Announce Type: new Abstract: Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in…
Difficulty-Aware Semantic-ID Optimization for Generative Recommendation
arXiv:2608.20611v1 Announce Type: new Abstract: Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over…
Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
arXiv:2608.20569v1 Announce Type: new Abstract: Are frontier models able to introspect about their internal states? Recent work suggests that under…
Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
arXiv:2608.20614v1 Announce Type: new Abstract: Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and…
