| Data storage | Amazon S3
Advantages- Practically unlimited, highly durable storage for datasets and model files
- SageMaker, Athena and Glue read straight from S3, so you train without copying data
- Versioning keeps old dataset versions; lifecycle rules move cold data to cheaper tiers
| Cloud Storage (GCS)
Advantages- One global namespace with strong consistency
- Vertex AI and BigQuery read directly from Cloud Storage
- Autoclass moves rarely used data to cheaper storage automatically
| Azure Blob Storage / Data Lake Gen2
Advantages- Data Lake Gen2 adds folders and fast analytics on big datasets
- Connects to Azure ML datastores, Synapse and Databricks
- Hot, cool and archive tiers to balance speed and cost
| Local disk / MinIO (S3-compatible) / PostgreSQL
Advantages- No cloud bill, and data never leaves your hardware
- MinIO speaks the S3 API, so code moves to the cloud unchanged later
- PostgreSQL is a solid home for structured data
|
| Notebooks | SageMaker Studio notebooks
Advantages- Managed JupyterLab with nothing to install
- Switch the instance from CPU to GPU when you need more power
- Built-in access to S3 data, experiments and the model registry
| Vertex AI Workbench / Colab Enterprise
Advantages- Familiar Colab / Jupyter experience on managed machines
- Query BigQuery with SQL and analyse in Python in one place
- Idle shutdown keeps costs under control
| Azure ML compute instance notebooks
Advantages- Managed notebooks that also open in VS Code
- Attach datastores and compute in a few clicks
- Auto-shutdown schedules stop forgotten machines
| JupyterLab or VS Code
Advantages- Free, and runs on any laptop or server
- Full control over packages and extensions
- Works offline
|
| GPU training | SageMaker Training on ml.g5.xlarge (A10G) / ml.p4d for large jobs
Advantages- Billed only while the training job runs; machines shut down automatically
- Managed Spot Training can cut training costs dramatically
- Scales from one GPU to distributed multi-GPU training
| Vertex AI custom training with L4 (g2) / A100 (a2) GPUs
Advantages- Wide choice of accelerators: L4, A100, H100 GPUs and TPUs
- Spot VMs make long training runs much cheaper
- Billed only while the job runs
| Azure ML compute clusters on NC-series (T4 / A100) GPUs
Advantages- NC / ND GPUs from T4 up to A100 and H100
- Low-priority VMs reduce training costs
- Compute clusters scale to zero when idle
| A local NVIDIA GPU, or rented GPUs (RunPod, Lambda, Vast.ai)
Advantages- An owned GPU has no hourly cost after purchase
- GPU rental marketplaces are often cheaper than the big clouds
- Choose the exact GPU model you need
|
| ML platform & registry | Amazon SageMaker (experiments, model registry)
Advantages- One place for experiments, the model registry, approvals and deployment
- Lineage shows which data and code produced each model
- Access controlled with IAM roles
| Vertex AI (experiments, model registry)
Advantages- Datasets, training, registry and endpoints in one product
- Experiments and TensorBoard built in
- AutoML for teams without ML specialists
| Azure Machine Learning (MLflow-based registry)
Advantages- MLflow-native tracking and registry, portable to other platforms
- Responsible AI dashboard for fairness and explanations
- Designer and AutoML for low-code model building
| MLflow (tracking + model registry)
Advantages- Open-source, vendor-neutral experiment tracking and registry
- The same MLflow API works locally and on Databricks or Azure ML
- Easy to self-host with Docker
|
| Serverless serving | AWS Lambda (container image) or SageMaker Serverless Inference
Advantages- Scales to zero, so you pay nothing when nobody is using the model
- Handles spiky traffic automatically
- No servers to patch or manage
| Cloud Run (scales to zero)
Advantages- Deploy any container and scale to zero
- A generous free tier for small projects
- Can attach an L4 GPU when you need one
| Azure Functions (container) or Azure Container Apps (scale to zero)
Advantages- Container Apps scale to zero between requests
- Pay per request or per second of use
- Trigger predictions from queues or new files with Functions
| Docker container on a small VM (or Fly.io / Render)
Advantages- Simple and cheap for low or steady traffic
- A container runs the same anywhere
- Predictable flat monthly price
|
| GPU serving | SageMaker real-time endpoint on ml.g5 / Inferentia2
Advantages- Low-latency real-time predictions with autoscaling
- Inferentia2 chips can lower the cost per prediction
- Production variants let you A/B test two models
| Vertex AI Endpoint with GPU, or Cloud Run with L4 GPU
Advantages- Autoscaling endpoints with traffic splitting for safe rollouts
- Cloud Run GPUs scale to zero, so idle GPUs don't cost money
- Use prebuilt containers or your own
| Azure ML managed online endpoint on GPU SKUs
Advantages- Managed online endpoints with blue/green deployments
- Autoscaling and built-in monitoring
- Secure with keys or Microsoft Entra ID
| Triton Inference Server / vLLM / BentoML on a GPU box
Advantages- Request batching squeezes the most throughput out of a GPU
- vLLM is a leading engine for serving LLMs fast
- No per-request fees
|
| Batch predictions | SageMaker Batch Transform, scheduled by EventBridge
Advantages- Scores millions of rows without keeping a server running
- Reads input from S3 and writes results back to S3
- You pay only for the job's duration
| Vertex AI Batch Prediction, scheduled by Cloud Scheduler
Advantages- Predict over files in Cloud Storage or whole BigQuery tables
- No always-on endpoint to pay for
- Results land in BigQuery for analysis and dashboards
| Azure ML batch endpoints, scheduled by Azure ML schedules
Advantages- Process large datasets in parallel on a cluster
- Clusters shut down when the job finishes
- Results written to Blob Storage
| Cron or Prefect/Airflow job running a Python script
Advantages- Just a scheduled Python script, easy to understand
- No vendor lock-in
- Runs on servers you already have
|
| Pipelines | SageMaker Pipelines or Step Functions
Advantages- Repeatable train → evaluate → register workflows
- Conditional steps, e.g. only deploy if accuracy beats a threshold
- Retraining on a schedule or when new data arrives
| Vertex AI Pipelines (Kubeflow)
Advantages- Managed Kubeflow / TFX pipelines
- Tracks the lineage of every dataset, model and metric
- Schedule retraining with Cloud Scheduler
| Azure ML pipelines
Advantages- Reusable components shared across projects
- Schedules and triggers for automatic retraining
- Lineage between data, runs and registered models
| Prefect, Dagster or Airflow
Advantages- Open source with large communities
- Workflows are plain Python
- Run locally or move to any cloud later
|
| Vector database | Amazon OpenSearch Serverless or RDS PostgreSQL + pgvector
Advantages- Managed similarity search for RAG chatbots
- Hybrid keyword + vector search in OpenSearch
- pgvector keeps embeddings next to your relational data
| AlloyDB / Cloud SQL + pgvector or Vertex AI Vector Search
Advantages- Vertex AI Vector Search handles billions of vectors at low latency
- AlloyDB / Cloud SQL with pgvector keeps vectors beside app data
- Managed backups and high availability
| Azure AI Search or Azure Database for PostgreSQL + pgvector
Advantages- Azure AI Search combines keyword, vector and semantic ranking
- Plugs straight into Azure AI Foundry for RAG
- PostgreSQL + pgvector for relational and vector data together
| Qdrant, Chroma or PostgreSQL + pgvector
Advantages- Free and open source
- Chroma is great for prototypes; Qdrant scales to production
- pgvector reuses an existing PostgreSQL database
|
| LLM access | Amazon Bedrock (Claude and others)
Advantages- Claude and other models through one API, with no GPUs to manage
- Prompts and data stay in your AWS account and region
- Built-in guardrails, knowledge bases and agents
| Vertex AI Model Garden (Claude, Gemini and others)
Advantages- Claude, Gemini and open models in one catalogue
- Enterprise data governance: your prompts aren't used to train the models
- Tuning and evaluation tools built in
| Azure AI Foundry (Claude, OpenAI and others)
Advantages- Claude, OpenAI and open models under Azure governance
- Content safety filters built in
- Private networking and regional deployments
| Ollama or vLLM serving an open-weights model (Llama, Qwen, Mistral)
Advantages- Prompts and documents never leave your network
- No per-token charges, only hardware costs
- Pick, and even fine-tune, any open-weights model
|
| Monitoring | CloudWatch + SageMaker Model Monitor
Advantages- Alerts on latency, errors and traffic
- Model Monitor detects data drift against a training baseline
- Central logs make debugging predictions easier
| Cloud Monitoring + Vertex AI Model Monitoring
Advantages- Model Monitoring alerts on feature skew and drift
- Dashboards for latency and errors in Cloud Monitoring
- Export logs to BigQuery to analyse predictions
| Azure Monitor / Application Insights + Azure ML model monitoring
Advantages- Application Insights traces each request end to end
- Azure ML watches for data drift and prediction quality
- Alerts can go to email or Teams
| Prometheus + Grafana + Evidently AI
Advantages- Industry-standard open-source monitoring
- Evidently produces ready-made drift and quality reports
- Build a dashboard for any metric
|
| Privacy controls | VPC endpoints, KMS encryption, HIPAA-eligible services, keep data in one region
Advantages- VPC endpoints keep traffic off the public internet
- Encrypt data and models with your own KMS keys
- Many HIPAA-eligible services and compliance certifications
| VPC Service Controls, CMEK encryption, data residency regions
Advantages- VPC Service Controls build a perimeter against data leaks
- Customer-managed encryption keys (CMEK)
- Choose the region where data is stored
| Private endpoints, customer-managed keys, regional data residency
Advantages- Private endpoints keep traffic on Microsoft's network
- Customer-managed encryption keys
- Fine-grained access with Microsoft Entra ID
| Everything stays on your hardware; encrypt disks and restrict network access
Advantages- Full control over where data lives
- Works in air-gapped environments
- You set the encryption and access rules (and own the security work)
|