Now you can get more intelligent speech-to-text transcription with Gemini 3.5 Transcribe.
Tag: AI
The Limits of Automatic Evaluation of Creativity in Large Language Models
arXiv:2608.23705v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly capable of generating text that challenges human…
How GoDaddy transformed its analytics with Amazon Quick
In this post, you will learn how GoDaddy migrated from their legacy business intelligence (BI) tool to Amazon Quick. This was a two-year transformation…
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model
arXiv:2608.23663v1 Announce Type: cross Abstract: Aligning deployed language models requires knowing when their outputs can be trusted, yet on-device…
Natera’s intelligent appointment scheduling with Amazon Bedrock AgentCore
Learn how Natera built an automated voice agent on Amazon Bedrock AgentCore that lets patients book mobile phlebotomy appointments through natural…
TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers
arXiv:2608.23763v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has emerged as the standard layer connecting Large Language Model…
Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling
arXiv:2608.23653v1 Announce Type: cross Abstract: AI agents are increasingly used for simulation-driven engineering. Physical system modeling presents…
Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail
arXiv:2608.23651v1 Announce Type: cross Abstract: Agent harnesses record a failed tool call and its error message in the transcript and ask the model to…
Preparing data for supervised fine-tuning Part 2: Advanced data strategies
The advanced side of supervised fine-tuning data prep. This second post in a two-part series covers evaluating data readiness with learning curves,…
ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling Agents
arXiv:2608.23635v1 Announce Type: cross Abstract: Large language models (LLMs) rely on tool calling as a fundamental agent capability, enabling them to…
