JobRaahGet matched free

Jobs

ML Ops Engineer

Anaplan · London, United Kingdom · On-site

Posted Aug 13, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

At Anaplan, we are a team of innovators focused on optimizing business decision-making through our leading AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market. What unites Anaplanners across teams and geographies is our collective commitment to our customers’ success and to our Winning Culture. Our customers rank among the who’s who in the Fortune 50. Coca-Cola, LinkedIn, Adobe, LVMH and Bayer are just a few of the 2,400+ global companies who rely on our best-in-class platform. Our Winning Culture is the engine that drives our teams of innovators. We champion diversity of thought and ideas, we behave like leaders regardless of title, we are committed to achieving ambitious goals, and we love celebrating our wins – big and small. Supported by operating principles of being strategy-led, values -based and disciplined in execution, you’ll be inspired, connected, developed and rewarded here. Everything that makes you unique is welcome; join us and let’s build what’s next - together! Role Overview We are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting-edge AI-infused scenario planning platform. You will work closely with Data Scientists, ML Engineers, and Cloud Infrastructure teams to streamline model training, deployment, and inference while ensuring optimal GPU utilisation, reliability, and cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed networking, and storage for efficient model training and low-latency inference. Build and maintain robust CI/CD and MLOps pipelines for continuous model training, evaluation, packaging, and production deployment. Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM). Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks. Monitor and optimise cloud spend across high-cost GPU/CPU clusters across AWS, GCP, or Azure. Implement auto-scaling strategies, spot instance policies, and dynamic resource allocation to eliminate infrastructure waste. Establish benchmarking and telemetry to track unit economics and throughput for training and serving AI models. Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your Skills Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure. …