arXiv:2608.13228v1 Announce Type: new Abstract: Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful…
Tag: AI
TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems
arXiv:2608.13221v1 Announce Type: new Abstract: The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet…
SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents
arXiv:2608.13173v1 Announce Type: new Abstract: Agent skills are crucial external instructions that enable language agents to execute long procedural…
Google AI health coach to use Abbott glucose data
Abbott and Google are linking continuous glucose monitoring data with Google’s AI-powered health coaching tools, giving the Gemini-powered service access…
Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents
arXiv:2608.13179v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for…
Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes
Is Anthropic’s new watermarking system a travesty? Some have taken to social media to complain that it is.
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing
arXiv:2608.13156v1 Announce Type: new Abstract: Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint…
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
arXiv:2608.13120v1 Announce Type: new Abstract: Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently…
Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement
arXiv:2608.13129v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain…
SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference
arXiv:2608.13076v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable success in natural language understanding and…
