Hello Model

AWS for machine learning

The service to use for each part of an ML system on AWS, and why it helps your model.

ComponentAWS serviceWhy it helps your model
Data storageAmazon S3
  • Practically unlimited, highly durable storage for datasets and model files
  • SageMaker, Athena and Glue read straight from S3, so you train without copying data
  • Versioning keeps old dataset versions; lifecycle rules move cold data to cheaper tiers
NotebooksSageMaker Studio notebooks
  • Managed JupyterLab with nothing to install
  • Switch the instance from CPU to GPU when you need more power
  • Built-in access to S3 data, experiments and the model registry
GPU trainingSageMaker Training on ml.g5.xlarge (A10G) / ml.p4d for large jobs
  • Billed only while the training job runs; machines shut down automatically
  • Managed Spot Training can cut training costs dramatically
  • Scales from one GPU to distributed multi-GPU training
ML platform & registryAmazon SageMaker (experiments, model registry)
  • One place for experiments, the model registry, approvals and deployment
  • Lineage shows which data and code produced each model
  • Access controlled with IAM roles
Serverless servingAWS Lambda (container image) or SageMaker Serverless Inference
  • Scales to zero, so you pay nothing when nobody is using the model
  • Handles spiky traffic automatically
  • No servers to patch or manage
GPU servingSageMaker real-time endpoint on ml.g5 / Inferentia2
  • Low-latency real-time predictions with autoscaling
  • Inferentia2 chips can lower the cost per prediction
  • Production variants let you A/B test two models
Batch predictionsSageMaker Batch Transform, scheduled by EventBridge
  • Scores millions of rows without keeping a server running
  • Reads input from S3 and writes results back to S3
  • You pay only for the job's duration
PipelinesSageMaker Pipelines or Step Functions
  • Repeatable train → evaluate → register workflows
  • Conditional steps, e.g. only deploy if accuracy beats a threshold
  • Retraining on a schedule or when new data arrives
Vector databaseAmazon OpenSearch Serverless or RDS PostgreSQL + pgvector
  • Managed similarity search for RAG chatbots
  • Hybrid keyword + vector search in OpenSearch
  • pgvector keeps embeddings next to your relational data
LLM accessAmazon Bedrock (Claude and others)
  • Claude and other models through one API, with no GPUs to manage
  • Prompts and data stay in your AWS account and region
  • Built-in guardrails, knowledge bases and agents
MonitoringCloudWatch + SageMaker Model Monitor
  • Alerts on latency, errors and traffic
  • Model Monitor detects data drift against a training baseline
  • Central logs make debugging predictions easier
Privacy controlsVPC endpoints, KMS encryption, HIPAA-eligible services, keep data in one region
  • VPC endpoints keep traffic off the public internet
  • Encrypt data and models with your own KMS keys
  • Many HIPAA-eligible services and compliance certifications