arXiv:2608.23635v1 Announce Type: cross Abstract: Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to…
Category: AI
Preparing data for supervised fine-tuning Part 2: Advanced data strategies
The advanced side of supervised fine-tuning data prep. This second post in a two-part series covers evaluating data readiness with learning curves,…
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
arXiv:2608.23663v2 Announce Type: cross Abstract: Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device…
When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs
arXiv:2608.23623v1 Announce Type: cross Abstract: Tool-using agents must decide when to stop. Existing systems already gate terminal success, certify…
Preparing data for supervised fine-tuning Part 1: Formatting and quality
Data preparation determines the ceiling of any supervised fine-tuning project. This first post in a two-part series covers the foundations of SFT data…
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
arXiv:2608.23660v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural…
AI Isn’t Ready for the Real Work: Why Models Flunk Complex Tasks
Investors in the AI bubble beg white collar professionals to hand over their hardest problems and promise workers that they’ll get their afternoons back…
Elastic KV Cache for LLM Serving:A Working Reclamation Mechanism, and Why Chunked Prefill Already Closes the Gap
arXiv:2608.23658v1 Announce Type: cross Abstract: An LLM serving engine sizes its key-value (KV) cache once, at startup, permanently setting aside a…
Connect Amazon Bedrock AgentCore to cross-account knowledge bases
Learn how Amazon Bedrock AgentCore agents in one account can generate answers from an Amazon Bedrock knowledge base backed by Amazon Redshift Serverless…
Macro-Operator Generation and Predicate Selection for TAMP Operator Learning
arXiv:2608.23629v1 Announce Type: cross Abstract: Creating symbolic operators by hand is one of the main bottlenecks in deploying Task and Motion Planning…
