Getting access to the right GPUs when you need them is one of the biggest challenges in training or customizing AI models at scale. During peak demand periods, your preferred GPU may not be immediately available – and when your job is tied to one specific GPU configuration, the only option is to wait or manually try alternatives. This slows down experimentation and pulls engineering focus away from model development. What if you could submit a single job with a list of suitable GPU options and have Amazon SageMaker AI automatically find available capacity from your list – reducing wait times and getting your teams back to building? Today, we’re excited to announce Instance preference lists for Amazon SageMaker AI Training Jobs and Amazon SageMaker Processing Jobs , helping you secure on-demand capacity faster by automatically checking across your preferred instance types. With this feature, you can specify an ordered list of up to five acceptable instance types when creating a training or processing job. Amazon SageMaker AI automatically evaluates your list in priority order and launches on the first type with available capacity – making it faster to secure GPU resources and start training. This removes the manual retry loops, complex monitoring scripts, and additional time teams sometimes invest in building systems to manage job submission for jobs that can run across multiple instance types. The result is faster job starts, higher capacity utilization, and more time spent building models rather than managing constraints. Customer challenges During peak demand periods, sec
Source: AWS Artificial Intelligence
