Favorite Multi-tenant agentic chat assistants have become a frequent request for large-scale customers, and document chat sits at the top of the list. A user uploads a contract, a report, or a product manual, and then researches or asks questions about it immediately or in the future. The conversational interface
Read More
Shared by AWS Machine Learning September 1, 2026
Favorite Teams that add Retrieval Augmented Generation (RAG) to a foundation model usually start with a single retrieval step against a single knowledge base. That works until the questions get harder, when the answer spans several sources, or the system has to decide which source to consult before it can
Read More
Shared by AWS Machine Learning September 1, 2026
Favorite Most organizations scaling their use of agents and tools hit the same challenges. Teams build in isolation, with no shared record of what exists, who owns it, or whether it’s been reviewed. The problem has moved from building agents and tools to discovering and governing them. AWS Agent Registry
Read More
Shared by AWS Machine Learning September 1, 2026
Favorite We’re excited to share that AWS has been recognized as a Leader in The Forrester Wave: AI Infrastructure Solutions, Q4 2025. In this evaluation of 13 providers, AWS received the highest score in the Strategy category. We believe this recognition reflects our commitment to delivering flexible, cost-efficient AI infrastructure
Read More
Shared by AWS Machine Learning September 1, 2026
Favorite Model Context Protocol (MCP) servers allow foundation models to access external data and tools, supporting standardized, secure access to files, databases, and APIs. They give AI agents the ability to interact with real-world applications, reduce hallucinations with accurate context, and offer stateful, multi-turn capabilities. Industry-standard architectures quickly evolved and
Read More
Shared by AWS Machine Learning September 1, 2026
Favorite When Salesforce set out to make Agentforce (Salesforce’s AI foundation for agents) highly available (HA) across multiple Availability Zones (AZs), the team faced a gap. Amazon SageMaker AI Inference Components (ICs) could cut GPU costs, but their default placement didn’t guarantee the Multi-AZ resilience Salesforce’s compliance bar required. For
Read More
Shared by AWS Machine Learning August 29, 2026
Favorite This post is co-written with Vianney Bruned, Filippo Giruzzi, Belkiss Saidi, and Carlos Ramirez from Decathlon. Decathlon is one of the world’s largest sporting goods retailers, with more than 100,000 teammates and 400 million users worldwide. The company relies on accurate demand forecasting at scale to support the availability
Read More
Shared by AWS Machine Learning August 29, 2026
Favorite Amazon SageMaker Feature Store is a fully managed, purpose-built repository to store, share, and manage features for machine learning (ML) models. It provides low-latency online serving for real-time inference, an offline store for historical retention and training feature data, and supports both streaming and batch ingestion patterns. As ML
Read More
Shared by AWS Machine Learning August 29, 2026
Favorite This post is a collaboration between AWS, NVIDIA and Heidi. Reducing automatic speech recognition (ASR) inference costs on Amazon Elastic Compute Cloud (Amazon EC2) becomes critical when GPU utilization per request is low but latency requirements are strict. A single ASR inference request typically uses only 15–20 percent of a
Read More
Shared by AWS Machine Learning August 28, 2026
Favorite Self-hosted speech AI has historically carried an observability trade-off. The service can tell you an endpoint is up and how many requests it served. The questions that actually drive capacity planning and cost management stay locked inside the vendor’s container: what you are billed for, which features your traffic
Read More
Shared by AWS Machine Learning August 28, 2026