MLOps Support Engineer
CloudFactory · Remote · Medellín, Medellin, Colombia · Remote
Posted Sep 23, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale.
More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.
Our Culture
At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:
Mission-Driven: We focus on creating economic and social impact.
People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
Innovative: We embrace change and find better ways to do things together.
Globally Connected: We foster collaboration between diverse cultures and perspectives.
If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!
About the role:
The MLOps Support Engineer is an operations-first role, focused on ensuring AI/ML systems remain stable, observable, and supportable in production environments. This is not a data science or feature development role.
The primary objective is to maintain continuous performance of ML models and associated pipelines with minimal disruption to both internal and client-facing services. You will provide Tier 1 and Tier 2 support, escalating to Tier 3 Engineering as needed.
What you’ll do:
Provide Tier 1 / Tier 2 operational support for AI/ML solutions.
Identify failed jobs, degraded pipelines, or performance anomalies.
Triage incidents, investigate issues, and coordinate escalation to Tier 3 Engineering.
Participate in on-call rotas once established.
Validate that pipelines and jobs complete successfully.
Monitor data pipeline health, model execution, and basic performance metrics.
Identify operational issues before they impact customers
Respond or alert customers when there has been an outage or issue with one of their models.
Support incident management, rollback, and recovery activities.
Use and maintain runbooks and operational documentation.
Work with Engineering to improve supportability and observability.
Contribute to knowledge sharing to reduce single points of failure.
Work within defined SLAs and support processes as the service matures
Build quarterly business reviews to provide updates on the health of the ML Models.
Evaluate champion/challenger models to see if a new model should be promoted.
Monitor for model drift and performance degradation, while validating that updates (new champion models or added data) do not introduce bias.
Requirements
Essential
Experience in operations, DevOps, SRE, or platform support roles.
Strong troubleshooting…