Hello Model

Cloud comparison

Which service to use for each part of an ML system, on each cloud, and why it helps your model. Pick a cloud to see its advantages side by side, or open “Advantages” under any service.

ComponentAWSGoogle CloudMicrosoft AzureOpen-source / self-hosted
Data storageAmazon S3
Advantages
  • Practically unlimited, highly durable storage for datasets and model files
  • SageMaker, Athena and Glue read straight from S3, so you train without copying data
  • Versioning keeps old dataset versions; lifecycle rules move cold data to cheaper tiers
Cloud Storage (GCS)
Advantages
  • One global namespace with strong consistency
  • Vertex AI and BigQuery read directly from Cloud Storage
  • Autoclass moves rarely used data to cheaper storage automatically
Azure Blob Storage / Data Lake Gen2
Advantages
  • Data Lake Gen2 adds folders and fast analytics on big datasets
  • Connects to Azure ML datastores, Synapse and Databricks
  • Hot, cool and archive tiers to balance speed and cost
Local disk / MinIO (S3-compatible) / PostgreSQL
Advantages
  • No cloud bill, and data never leaves your hardware
  • MinIO speaks the S3 API, so code moves to the cloud unchanged later
  • PostgreSQL is a solid home for structured data
NotebooksSageMaker Studio notebooks
Advantages
  • Managed JupyterLab with nothing to install
  • Switch the instance from CPU to GPU when you need more power
  • Built-in access to S3 data, experiments and the model registry
Vertex AI Workbench / Colab Enterprise
Advantages
  • Familiar Colab / Jupyter experience on managed machines
  • Query BigQuery with SQL and analyse in Python in one place
  • Idle shutdown keeps costs under control
Azure ML compute instance notebooks
Advantages
  • Managed notebooks that also open in VS Code
  • Attach datastores and compute in a few clicks
  • Auto-shutdown schedules stop forgotten machines
JupyterLab or VS Code
Advantages
  • Free, and runs on any laptop or server
  • Full control over packages and extensions
  • Works offline
GPU trainingSageMaker Training on ml.g5.xlarge (A10G) / ml.p4d for large jobs
Advantages
  • Billed only while the training job runs; machines shut down automatically
  • Managed Spot Training can cut training costs dramatically
  • Scales from one GPU to distributed multi-GPU training
Vertex AI custom training with L4 (g2) / A100 (a2) GPUs
Advantages
  • Wide choice of accelerators: L4, A100, H100 GPUs and TPUs
  • Spot VMs make long training runs much cheaper
  • Billed only while the job runs
Azure ML compute clusters on NC-series (T4 / A100) GPUs
Advantages
  • NC / ND GPUs from T4 up to A100 and H100
  • Low-priority VMs reduce training costs
  • Compute clusters scale to zero when idle
A local NVIDIA GPU, or rented GPUs (RunPod, Lambda, Vast.ai)
Advantages
  • An owned GPU has no hourly cost after purchase
  • GPU rental marketplaces are often cheaper than the big clouds
  • Choose the exact GPU model you need
ML platform & registryAmazon SageMaker (experiments, model registry)
Advantages
  • One place for experiments, the model registry, approvals and deployment
  • Lineage shows which data and code produced each model
  • Access controlled with IAM roles
Vertex AI (experiments, model registry)
Advantages
  • Datasets, training, registry and endpoints in one product
  • Experiments and TensorBoard built in
  • AutoML for teams without ML specialists
Azure Machine Learning (MLflow-based registry)
Advantages
  • MLflow-native tracking and registry, portable to other platforms
  • Responsible AI dashboard for fairness and explanations
  • Designer and AutoML for low-code model building
MLflow (tracking + model registry)
Advantages
  • Open-source, vendor-neutral experiment tracking and registry
  • The same MLflow API works locally and on Databricks or Azure ML
  • Easy to self-host with Docker
Serverless servingAWS Lambda (container image) or SageMaker Serverless Inference
Advantages
  • Scales to zero, so you pay nothing when nobody is using the model
  • Handles spiky traffic automatically
  • No servers to patch or manage
Cloud Run (scales to zero)
Advantages
  • Deploy any container and scale to zero
  • A generous free tier for small projects
  • Can attach an L4 GPU when you need one
Azure Functions (container) or Azure Container Apps (scale to zero)
Advantages
  • Container Apps scale to zero between requests
  • Pay per request or per second of use
  • Trigger predictions from queues or new files with Functions
Docker container on a small VM (or Fly.io / Render)
Advantages
  • Simple and cheap for low or steady traffic
  • A container runs the same anywhere
  • Predictable flat monthly price
GPU servingSageMaker real-time endpoint on ml.g5 / Inferentia2
Advantages
  • Low-latency real-time predictions with autoscaling
  • Inferentia2 chips can lower the cost per prediction
  • Production variants let you A/B test two models
Vertex AI Endpoint with GPU, or Cloud Run with L4 GPU
Advantages
  • Autoscaling endpoints with traffic splitting for safe rollouts
  • Cloud Run GPUs scale to zero, so idle GPUs don't cost money
  • Use prebuilt containers or your own
Azure ML managed online endpoint on GPU SKUs
Advantages
  • Managed online endpoints with blue/green deployments
  • Autoscaling and built-in monitoring
  • Secure with keys or Microsoft Entra ID
Triton Inference Server / vLLM / BentoML on a GPU box
Advantages
  • Request batching squeezes the most throughput out of a GPU
  • vLLM is a leading engine for serving LLMs fast
  • No per-request fees
Batch predictionsSageMaker Batch Transform, scheduled by EventBridge
Advantages
  • Scores millions of rows without keeping a server running
  • Reads input from S3 and writes results back to S3
  • You pay only for the job's duration
Vertex AI Batch Prediction, scheduled by Cloud Scheduler
Advantages
  • Predict over files in Cloud Storage or whole BigQuery tables
  • No always-on endpoint to pay for
  • Results land in BigQuery for analysis and dashboards
Azure ML batch endpoints, scheduled by Azure ML schedules
Advantages
  • Process large datasets in parallel on a cluster
  • Clusters shut down when the job finishes
  • Results written to Blob Storage
Cron or Prefect/Airflow job running a Python script
Advantages
  • Just a scheduled Python script, easy to understand
  • No vendor lock-in
  • Runs on servers you already have
PipelinesSageMaker Pipelines or Step Functions
Advantages
  • Repeatable train → evaluate → register workflows
  • Conditional steps, e.g. only deploy if accuracy beats a threshold
  • Retraining on a schedule or when new data arrives
Vertex AI Pipelines (Kubeflow)
Advantages
  • Managed Kubeflow / TFX pipelines
  • Tracks the lineage of every dataset, model and metric
  • Schedule retraining with Cloud Scheduler
Azure ML pipelines
Advantages
  • Reusable components shared across projects
  • Schedules and triggers for automatic retraining
  • Lineage between data, runs and registered models
Prefect, Dagster or Airflow
Advantages
  • Open source with large communities
  • Workflows are plain Python
  • Run locally or move to any cloud later
Vector databaseAmazon OpenSearch Serverless or RDS PostgreSQL + pgvector
Advantages
  • Managed similarity search for RAG chatbots
  • Hybrid keyword + vector search in OpenSearch
  • pgvector keeps embeddings next to your relational data
AlloyDB / Cloud SQL + pgvector or Vertex AI Vector Search
Advantages
  • Vertex AI Vector Search handles billions of vectors at low latency
  • AlloyDB / Cloud SQL with pgvector keeps vectors beside app data
  • Managed backups and high availability
Azure AI Search or Azure Database for PostgreSQL + pgvector
Advantages
  • Azure AI Search combines keyword, vector and semantic ranking
  • Plugs straight into Azure AI Foundry for RAG
  • PostgreSQL + pgvector for relational and vector data together
Qdrant, Chroma or PostgreSQL + pgvector
Advantages
  • Free and open source
  • Chroma is great for prototypes; Qdrant scales to production
  • pgvector reuses an existing PostgreSQL database
LLM accessAmazon Bedrock (Claude and others)
Advantages
  • Claude and other models through one API, with no GPUs to manage
  • Prompts and data stay in your AWS account and region
  • Built-in guardrails, knowledge bases and agents
Vertex AI Model Garden (Claude, Gemini and others)
Advantages
  • Claude, Gemini and open models in one catalogue
  • Enterprise data governance: your prompts aren't used to train the models
  • Tuning and evaluation tools built in
Azure AI Foundry (Claude, OpenAI and others)
Advantages
  • Claude, OpenAI and open models under Azure governance
  • Content safety filters built in
  • Private networking and regional deployments
Ollama or vLLM serving an open-weights model (Llama, Qwen, Mistral)
Advantages
  • Prompts and documents never leave your network
  • No per-token charges, only hardware costs
  • Pick, and even fine-tune, any open-weights model
MonitoringCloudWatch + SageMaker Model Monitor
Advantages
  • Alerts on latency, errors and traffic
  • Model Monitor detects data drift against a training baseline
  • Central logs make debugging predictions easier
Cloud Monitoring + Vertex AI Model Monitoring
Advantages
  • Model Monitoring alerts on feature skew and drift
  • Dashboards for latency and errors in Cloud Monitoring
  • Export logs to BigQuery to analyse predictions
Azure Monitor / Application Insights + Azure ML model monitoring
Advantages
  • Application Insights traces each request end to end
  • Azure ML watches for data drift and prediction quality
  • Alerts can go to email or Teams
Prometheus + Grafana + Evidently AI
Advantages
  • Industry-standard open-source monitoring
  • Evidently produces ready-made drift and quality reports
  • Build a dashboard for any metric
Privacy controlsVPC endpoints, KMS encryption, HIPAA-eligible services, keep data in one region
Advantages
  • VPC endpoints keep traffic off the public internet
  • Encrypt data and models with your own KMS keys
  • Many HIPAA-eligible services and compliance certifications
VPC Service Controls, CMEK encryption, data residency regions
Advantages
  • VPC Service Controls build a perimeter against data leaks
  • Customer-managed encryption keys (CMEK)
  • Choose the region where data is stored
Private endpoints, customer-managed keys, regional data residency
Advantages
  • Private endpoints keep traffic on Microsoft's network
  • Customer-managed encryption keys
  • Fine-grained access with Microsoft Entra ID
Everything stays on your hardware; encrypt disks and restrict network access
Advantages
  • Full control over where data lives
  • Works in air-gapped environments
  • You set the encryption and access rules (and own the security work)