Site Reliability Engineer
Intermedia · Portugal, Portugal · Remote
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
*ALL CANDIDATES MUST BE LOCATED IN PORTUGAL* About Intermedia Are you looking for a company where YOUR VOICE is heard? Where can you MAKE A DIFFERENCE? Do you THRIVE in a FAST-PACED work environment? Do you wake every morning EXCITED to work with GREAT PEOPLE and create SUCCESS TOGETHER? Then Intermedia is the place for you. Intermedia has established itself as a leading provider of cloud communications and collaboration tech that allows companies to connect better. We have a strong track record of growth, profitability, and creating an environment where everyone matters. Everyone. While we are fast-paced and admittedly a bit intense, we promise that you won’t be bored. You will find Intermedia is a place where you can indulge your passion for creating and supporting great cloud technology. What’s more, we always look to promote from within and have many employees who have been with us 10, 15, and 20+ years! Culture at Intermedia is built on teamwork and transparency. We hold each other accountable and always have each other’s back! While primarily remote, this role requires occasional visits to the office in Coimbra or in Aveiro. We plan to open an office in Porto in the future. This approach gives team members the flexibility to work remotely while also coming together in the office for collaboration and teamwork. Are you ready to make your mark?
About the Role
We are looking for a Site Reliability Engineer (SRE) to improve the reliability of our AI and analytics platforms and the services that depend on them. As Intermedia expands its global product deployments, this role will strengthen SRE practices across production infrastructure, data pipelines, machine-learning services, customer-facing analytics, and AI-powered Voice and Unified Communications capabilities. You will partner with AI, data, product, and platform teams to build scalable, observable systems that deliver dependable insights and resilient customer experiences.
Run and improve production environments that support AI and AI workloads, data pipelines, analytics applications, and customer-facing services.
Build software and automation to manage cloud infrastructure, data platforms, model-serving infrastructure, and application services.
Define and measure service level indicators, service level objectives, and error budgets for AI and analytics services, including availability, data freshness, pipeline completion, and inference latency.
Build end-to-end observability that correlates metrics, logs, and traces with data-quality signals, AI-service performance, and customer impact.
Monitor and optimize the reliability, performance, capacity, and cost of batch and streaming workloads, analytics queries, and inference services.
Partner with data engineering and machine-learning teams to make ingestion, transformation, feature, training, deployment, and reporting workflows production-ready.
Automate CI/CD and production-readiness checks for data pipelines, model and prompt…