Come Partner with Us

Agent Evaluation Metric for multi-turn conversations

Favorite Multi-turn agents fail in ways that single-turn evaluation misses: one early mistake quietly corrupts every later turn. This post introduces the Agent Evaluation Metric (AEM), a decomposable, turn-level way to measure agent quality. We apply it to its first dimension, correctness. We show how AEM pinpoints the one turn

Read More
Shared by AWS Machine Learning September 11, 2026

Model-agnostic PII detection with LLMs

Favorite A configurable, instruction-driven detector that runs on any large language model (LLM) managed on Amazon Bedrock, evaluated on five public PII corpora across nine LLM-based detectors, including the OpenAI PrivacyFilter. Fine-tuning a model on real-world text creates a personally identifiable information (PII) detection problem. Training corpora are full of

Read More
Shared by AWS Machine Learning September 11, 2026

Amazon Quick is now generally available on desktop

Favorite Your teams get an AI assistant that handles real work while your data stays in your environment and your conversations stay private Today, the Amazon Quick desktop application is generally available on macOS and Windows. We’re also adding a new activity feed to the mobile experience on iOS and

Read More
Shared by AWS Machine Learning September 11, 2026

Automate user-level custom permissions for Amazon Quick

Favorite As Amazon Quick environments scale and new AI-powered capabilities expand what users can do, automating user-level custom permissions becomes critical to maintaining the principle of least privilege. To address this, with custom permissions in Quick, you can enforce fine-grained access control by toggling specific features on or off for

Read More
Shared by AWS Machine Learning September 10, 2026