Senior Azure ML Infrastructure Engineer
stblaw · New York, NY · United States · On-site
Pay: USD 160,000 – 180,000 a year
Posted Oct 5, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Simpson Thacher & Bartlett LLP is looking for a Senior Azure ML Infrastructure Engineer to lead the design, development, and optimization of scalable ML infrastructure on Microsoft Azure. In this role, you will be the technical lead for deploying and maintaining robust (Machine Learning) ML Ops frameworks, ensuring efficient collaboration between data science, engineering, and DevOps teams. You’ll be instrumental in scaling our machine learning capabilities from experimentation to production across multiple use cases.
ESSENTIAL JOB DUTIES & RESPONSIBILITIES
Infrastructure Architecture & Engineering
Lead the architecture and implementation of production-grade ML infrastructure using Azure Machine Learning, AKS, Azure Data Lake, Azure Databricks, and related services.
Design scalable training and inference environments for deep learning and traditional ML workloads, optimizing performance and cost.
MLOps Strategy & Execution
Define and implement MLOps best practices: versioning, CI/CD for ML pipelines, monitoring, and model governance.
Automate end-to-end ML workflows using tools such as MLFlow, Azure ML Pipelines, or Kubeflow.
Build reusable templates and frameworks to standardize ML deployment across teams.
Cross-Functional Leadership
Collaborate with data scientists to productionize models, offering guidance on infrastructure, deployment strategies, and performance optimization.
Partner with DevOps and platform engineering teams to align infrastructure with broader cloud strategies and compliance standards.
Mentor junior ML and platform engineers, sharing best practices and driving engineering excellence.
Security, Reliability, and Observability
Implement enterprise-grade security and compliance controls using Azure Active Directory, RBAC, and data encryption strategies.
Integrate observability tooling (e.g., Azure Monitor, Prometheus, Grafana) for end-to-end monitoring of ML systems.
Ensure systems are highly available, reliable, and scalable to meet the demands of production ML workloads.
EDUCATION
Bachelor’s degree in Computer Science, Information Systems, or a related field, or equivalent practical experience in lieu of formal education.
Legal IT experience a plus but not required
SKILLS AND EXPERIENCE
Required
5+ years of experience in ML infrastructure, cloud engineering, or MLOps
2+ years of experience working in Azure environments.
Deep hands-on experience with Azure cloud services relevant to ML, including Azure Machine Learning, AKS, Blob Storage, Databricks, Azure Data Factory, and Synapse.
Strong expertise in containerization (Docker) and orchestration (Kubernetes, preferably AKS).
Proficient in Python and scripting languages (e.g., Bash, PowerShell).
Advanced knowledge of CI/CD tools such as Azure DevOps, GitHub Actions, or Jenkins for ML workloads.
Solid understanding of IaC tools: Terraform, Bicep, or ARM templates.
Preferred
Microsoft Azure certifications (e.g., Azure AI…