arXiv:2608.13563v1 Announce Type: cross Abstract: Early-stage teams often lack users, time, and budget to run repeated UX studies, yet still need…
Category: AI
PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments
arXiv:2608.14441v1 Announce Type: new Abstract: Self-evolving agents improve future behavior from interaction experience, yet existing evaluations…
SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning
arXiv:2608.14452v1 Announce Type: new Abstract: Spreadsheets are widely used to organize, analyze, and manipulate semi-structured data, yet automated…
Shift Aware Transfer Learning with Adaptive Dual-Encoder Fusion for PM Forecasting in Data-Limited Environments
arXiv:2608.14456v1 Announce Type: new Abstract: Short-horizon forecasting of fine particulate matter (PM2.5) remains difficult when observations from the…
The Dynamics of Intelligence Explosions
arXiv:2608.14426v1 Announce Type: new Abstract: AI is increasingly being used to help with AI R&D. Under certain conditions this feedback loop might be…
Anthropic watermarks Claude’s output, but critics question the tradeoffs
Anthropic’s text watermarking for Claude is supposed to make AI-generated content detectable. But critics doubt that word choice stays unaffected, and…
Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports
arXiv:2608.14446v1 Announce Type: new Abstract: In the current artificial intelligence-driven innovation era, the pace of knowledge growth is…
Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations
arXiv:2608.14425v1 Announce Type: new Abstract: LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after…
The Past and Future of AI Scientists
arXiv:2608.14407v1 Announce Type: new Abstract: We present a survey of the past and future of AI Scientists: machines capable of automating science. AI…
Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons
arXiv:2608.14392v1 Announce Type: new Abstract: Neuron- and path-level interventions offer the finest-grained route to defending large language models…
