arXiv:2608.28691v1 Announce Type: cross Abstract: Wearable VLM pipelines promise continuous multimodal assistance from egocentric visual capture: a user…
Category: cs.AI updates on arXiv.org
The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents
arXiv:2608.28795v1 Announce Type: cross Abstract: Modern artificial-intelligence coding agents can be equipped with tools for checking their own work e.g.…
ASTRA – Agentic System for Ticket Resolution and Analysis
arXiv:2608.28790v1 Announce Type: cross Abstract: Technical operations teams resolve large volumes of incidents by synthesizing fragmented evidence from…
RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction
arXiv:2608.28718v1 Announce Type: cross Abstract: Video world models increasingly serve as data engines, action planners, and simulators for embodied AI,…
Cross Lingual Transfer in Tulu Legal Comprehension: Script-Dependent Improvement and RAG-Induced Knowledge Conflict
arXiv:2608.28645v1 Announce Type: cross Abstract: Low-resource languages without an adequate training corpus often use a related, higher-resource language…
Can Large Language Models Identify Meaningful Touchpoints in Conversion Attribution?
arXiv:2608.28649v1 Announce Type: cross Abstract: Touchpoint selection in conversion attribution, namely identifying meaningful touchpoints contributing…
GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon
arXiv:2608.28667v1 Announce Type: cross Abstract: The rapid proliferation of Large Language Models (LLMs) has raised concerns about their environmental…
Measuring Similarity between Artistic and AI Generated Images using Siamese Neural Networks
arXiv:2608.28671v1 Announce Type: cross Abstract: AI-generated art has sparked debates around potential plagiarism, as these images may closely resemble…
Test-Time Scaling for Scientific Equation Discovery
arXiv:2608.28660v1 Announce Type: cross Abstract: Test-time scaling (TTS) improves language model reasoning by allocating additional test-time compute,…
Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
arXiv:2608.28641v1 Announce Type: cross Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect…
