arXiv:2608.18534v1 Announce Type: new Abstract: Large language models are increasingly used to support financial operations, but their apparent reasoning…
Category: cs.AI updates on arXiv.org
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
arXiv:2608.18580v1 Announce Type: new Abstract: Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal…
Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement
arXiv:2608.18543v1 Announce Type: new Abstract: Modern e-commerce platforms often operate search, recommendation, personalization, and CRM systems…
UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval
arXiv:2608.18504v1 Announce Type: new Abstract: Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both…
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations
arXiv:2608.18389v1 Announce Type: new Abstract: AI code agents are increasingly deployed to resolve real software issues, yet their reliability under…
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
arXiv:2608.18423v1 Announce Type: new Abstract: Language model agents now execute bounded tasks reliably. Whether they can sustain effective…
Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions
arXiv:2608.18409v1 Announce Type: new Abstract: Combinatorial scheduling poses a significant challenge for language models, requiring them to identify…
When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification
arXiv:2608.18397v1 Announce Type: new Abstract: Wearable stress classifiers can achieve strong average performance while failing completely for a…
The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
arXiv:2608.18300v1 Announce Type: new Abstract: LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another…
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam’s 2025 Convex Marking Scheme
arXiv:2608.18336v1 Announce Type: new Abstract: When evaluating language models on human exams, benchmarks typically score each response as right or wrong…
