JobRaahGet matched free

Jobs

ML Ops Engineer

cmcmarkets · London · United Kingdom · On-site

Posted Sep 8, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

ML Ops Engineer London We’re hiring an ML Ops Engineer to build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services. You’ll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams. This is not a research role. It is a hands-on engineering role focused on making ML systems reproducible, scalable, secure and dependable, from model packaging and release through to serving, monitoring, retraining and incident response. What you’ll work on ML lifecycle and platform engineering Build repeatable workflows for model training, validation, promotion, deployment and retraining. Productionise models through packaging, versioning, model registry integration, deployment automation and safe rollback. Design CI/CD pipelines for ML systems, including automated testing, validation, release controls and environment promotion. Manage experiment tracking, model metadata and reproducibility across research and production. Build reusable tooling and platform capabilities that support multiple models and engineering teams. Model serving and observability Deploy and operate batch and online inference services in containerised cloud environments. Define and meet availability, latency, throughput and recovery objectives for ML services. Monitor service health, infrastructure, data-quality signals, data drift, prediction drift and model performance decay. Establish dashboards, alerting and operational runbooks so failures are detected and resolved quickly. Support automated or controlled retraining, model promotion, rollback and model retirement. Debug production issues across model, application, infrastructure and critical data-dependency layers. Reliability, security and engineering quality Improve system robustness, scalability and cost efficiency through automation, observability and infrastructure as code. Write production-grade Python for long-running services, deployment tooling and ML workflows. Establish testing, validation, release and incident-management practices for ML systems. Collaborate with platform, security and data engineering teams on reliable model inputs, access controls, secrets, resilience and compliance. Make explicit trade-offs between research flexibility, delivery speed, operational risk and production stability. Additional responsibilities Maintain personal/professional development to meet the changing demands of the role, including all relevant regulatory and legislative training When dealing with all customers, clients or colleagues ensure that we provide a clear, fair and consistent high quality service that presents a professional and positive image of CMC Markets Take all reasonable steps to ensure appropriate confidentiality Undertake such other duties, training and/or hours of work as may …