In this post, you will learn how GoDaddy migrated from their legacy business intelligence (BI) tool to Amazon Quick. This was a two-year transformation…
Tag: AI
Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail
arXiv:2608.23651v1 Announce Type: cross Abstract: Agent harnesses record a failed tool call and its error message in the transcript and ask the model to…
Natera’s intelligent appointment scheduling with Amazon Bedrock AgentCore
Learn how Natera built an automated voice agent on Amazon Bedrock AgentCore that lets patients book mobile phlebotomy appointments through natural…
ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents
arXiv:2608.23635v1 Announce Type: cross Abstract: Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to…
Preparing data for supervised fine-tuning Part 2: Advanced data strategies
The advanced side of supervised fine-tuning data prep. This second post in a two-part series covers evaluating data readiness with learning curves,…
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
arXiv:2608.23663v2 Announce Type: cross Abstract: Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device…
When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs
arXiv:2608.23623v1 Announce Type: cross Abstract: Tool-using agents must decide when to stop. Existing systems already gate terminal success, certify…
Preparing data for supervised fine-tuning Part 1: Formatting and quality
Data preparation determines the ceiling of any supervised fine-tuning project. This first post in a two-part series covers the foundations of SFT data…
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
arXiv:2608.23660v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural…
AI Isn’t Ready for the Real Work: Why Models Flunk Complex Tasks
Investors in the AI bubble beg white collar professionals to hand over their hardest problems and promise workers that they’ll get their afternoons back…
