Hello Model

Google Cloud for machine learning

The service to use for each part of an ML system on Google Cloud, and why it helps your model.

ComponentGoogle Cloud serviceWhy it helps your model
Data storageCloud Storage (GCS)
  • One global namespace with strong consistency
  • Vertex AI and BigQuery read directly from Cloud Storage
  • Autoclass moves rarely used data to cheaper storage automatically
NotebooksVertex AI Workbench / Colab Enterprise
  • Familiar Colab / Jupyter experience on managed machines
  • Query BigQuery with SQL and analyse in Python in one place
  • Idle shutdown keeps costs under control
GPU trainingVertex AI custom training with L4 (g2) / A100 (a2) GPUs
  • Wide choice of accelerators: L4, A100, H100 GPUs and TPUs
  • Spot VMs make long training runs much cheaper
  • Billed only while the job runs
ML platform & registryVertex AI (experiments, model registry)
  • Datasets, training, registry and endpoints in one product
  • Experiments and TensorBoard built in
  • AutoML for teams without ML specialists
Serverless servingCloud Run (scales to zero)
  • Deploy any container and scale to zero
  • A generous free tier for small projects
  • Can attach an L4 GPU when you need one
GPU servingVertex AI Endpoint with GPU, or Cloud Run with L4 GPU
  • Autoscaling endpoints with traffic splitting for safe rollouts
  • Cloud Run GPUs scale to zero, so idle GPUs don't cost money
  • Use prebuilt containers or your own
Batch predictionsVertex AI Batch Prediction, scheduled by Cloud Scheduler
  • Predict over files in Cloud Storage or whole BigQuery tables
  • No always-on endpoint to pay for
  • Results land in BigQuery for analysis and dashboards
PipelinesVertex AI Pipelines (Kubeflow)
  • Managed Kubeflow / TFX pipelines
  • Tracks the lineage of every dataset, model and metric
  • Schedule retraining with Cloud Scheduler
Vector databaseAlloyDB / Cloud SQL + pgvector or Vertex AI Vector Search
  • Vertex AI Vector Search handles billions of vectors at low latency
  • AlloyDB / Cloud SQL with pgvector keeps vectors beside app data
  • Managed backups and high availability
LLM accessVertex AI Model Garden (Claude, Gemini and others)
  • Claude, Gemini and open models in one catalogue
  • Enterprise data governance: your prompts aren't used to train the models
  • Tuning and evaluation tools built in
MonitoringCloud Monitoring + Vertex AI Model Monitoring
  • Model Monitoring alerts on feature skew and drift
  • Dashboards for latency and errors in Cloud Monitoring
  • Export logs to BigQuery to analyse predictions
Privacy controlsVPC Service Controls, CMEK encryption, data residency regions
  • VPC Service Controls build a perimeter against data leaks
  • Customer-managed encryption keys (CMEK)
  • Choose the region where data is stored