arXiv:2608.10101v1 Announce Type: cross Abstract: Code review is credited with substantially changing a patch’s code between its first submission and the…
Category: AI
Pay with confidence: How Solv Labs built verifiable, auditable agent payments on Amazon Bedrock AgentCore payments
Solv Labs built a governed agent-payments workflow on Amazon Bedrock AgentCore payments, where every transaction is authorized, attested in an AWS Nitro…
MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation
arXiv:2608.10166v1 Announce Type: cross Abstract: Digital watermarking has emerged as a critical technique for provenance and copyright attribution in…
Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons
arXiv:2608.10045v1 Announce Type: cross Abstract: The problem of learning from pairwise comparisons has been widely studied across many domains such as…
UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs
arXiv:2608.10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users’ behalf, but existing benchmarks usually focus on…
Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds
arXiv:2608.10056v1 Announce Type: cross Abstract: Following a target human in crowded environments involves an inherent conflict between staying close to…
Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review
arXiv:2608.10047v1 Announce Type: cross Abstract: In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and…
Status Association Does Not Reliably Predict Decision Leakage
arXiv:2608.10089v1 Announce Type: cross Abstract: Bias evaluations often move too quickly from evidence that a model encodes a social association to…
DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents
arXiv:2608.10037v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks,…
Nvidia’s Nemotron 4 aims for one trillion parameters, a scale Chinese labs already surpassed
Nvidia is working on Nemotron 4, a new open-weight model designed to rival the world’s best freely available models. The article Nvidia’s Nemotron 4 aims…
