The Hallway Track
Product Launches

Amazon SageMaker AI Async Inference now supports inline request payloads

AWS Machine Learning Blog · Jun 17, 2026 · Product Launches

SageMaker AI Async Inference now accepts inline request payloads up to 128,000 bytes, removing the mandatory S3 upload step.

“For payloads up to 128,000 bytes, this removes an entire network round-trip, simplifies client-side code, and reduces the operational surface area of asynchronous inference workloads.”

AWS added a Body parameter to SageMaker Async Inference's InvokeEndpointAsync API, letting customers send small payloads inline instead of uploading to S3 first. It's a developer-convenience improvement for asynchronous ML workloads with small inputs, but it is an incremental infrastructure update rather than a major AI industry signal.

aws sagemaker inference mlops

Watch / read the original source →