← CasesML platform

cloud provider2022-2025

An on-demand ML platform on ClearML

The provider sells an ML platform as a service: every client gets a private installation in any region, with a free two-week trial. We moved platform delivery from ClickOps to IaC and automated tests.

Numbers

8 h → 24 minto deploy the platform for a client
~19Helm charts installed by Terraform in the right order
6 minof automated tests before handing it to the client

Before

  • ClickOps: Managed Kubernetes, file storage, S3 and the container registry were clicked together by hand in the console.
  • About 19 Helm charts were installed in a strict order, with credentials and addresses copied into them by hand.
  • Building the infrastructure by hand took about two days, the whole platform deployment 8 hours. Snowflake servers and knowledge in one person's head.

After

  • The platform is deployed for a client in 24 minutes, 20 times faster: 30 seconds for code checks, 17 minutes to deploy, 6 minutes of tests, 1 minute to tear down.
  • The client configuration lives in one file, and every branch gets its own isolated platform.
  • The same pipeline later delivered the inference platform to clients, about 20 minutes per deployment.

What we did

  1. An open source platform on Managed Kubernetes: ClearML for experiments and queues, JupyterHub, Gitea, Keycloak, Istio, prometheus-stack, the GPU operator and KServe for serving.

  2. Three iterations of IaC. GitLab downstream pipelines, then a monolithic Terraform job, then separate jobs with their own state per component and terraform_remote_state. Modules are versioned in the GitLab Registry.

  3. Terraform installs the Helm charts. The order comes from depends_on, and variables from S3, DNS and the rest of the infrastructure go straight into values.

  4. Automated tests before handing it to the client. pytest checks training on GPU and CPU, artifact uploads to S3 and image pushes to the registry, Selenium checks the web UIs. No green tests, no merge.

  5. A branch is its own platform. GitLab environments spin up an isolated installation for every branch and remove it after the merge. Phoenix servers instead of snowflake servers.

  6. GPU sharing inside the platform. Automatic MIG and TimeSlicing partitioning: several Jupyter instances on one card and ClearML experiment queues.

Stack

ClearMLJupyterHubKServeGiteaKeycloakIstioTraefikprometheus-stackNVIDIA GPU OperatorMIGTimeSlicingManaged KubernetesS3TerraformGitLab CIpytestSelenium

My role

DevOps/MLOps engineer in the Data/ML products department. I automated platform delivery (IaC, CI/CD, automated tests) and owned GPUs: the GPU operator and sharing.

Discuss a similar taskTelegrama consultation via getMentor or Telegram

↑↓ select · Enter open · Esc close