Amazon SageMaker AI Async Inference now supports inline request payloads
SageMaker AI Async Inference now accepts inline request payloads up to 128,000 bytes, removing the mandatory S3 upload step.
“For payloads up to 128,000 bytes, this removes an entire network round-trip, simplifies client-side code, and reduces the operational surface area of asynchronous inference workloads.”
AWS added a Body parameter to SageMaker Async Inference's InvokeEndpointAsync API, letting customers send small payloads inline instead of uploading to S3 first. It's a developer-convenience improvement for asynchronous ML workloads with small inputs, but it is an incremental infrastructure update rather than a major AI industry signal.