Site Reliability Engineer
Melza
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Site Reliability Engineer - Job Description
A tech company with a business-driven DNA, founded in Barcelona, with the vision of becoming a global leader in the creation of Digital Hubs for major corporations worldwide.
We dont just build digital products we help shape the digital business strategy behind them. Our teams operate at the intersection of technology and business, acting as high-level strategic partners for the companies we work with. Were trusted not only to deliver, but to lead driving innovation, challenging assumptions, and defining the future of digital for global organizations.
We work with agile methodologies, embracing continuous iteration as our mantra. Our multidisciplinary teams manage the entire end-to-end lifecycle of our digital products, and we are dynamic and business oriented.
Job Summary
We are looking for a Site Reliability Engineer (SRE) to join our Platform & Foundations teams and help us build and maintain the infrastructure that supports our digital products.
You will work closely with cross-functional squads, ensuring the scalability, reliability, and performance of our cloud-native systems. Your role will be crucial in optimizing developer experience, automating infrastructure operations, and evolving our Lakehouse and Kubernetes- based environments.
Responsibilities
Collaborate with engineering teams throughout the software development lifecycle to deploy, maintain, and monitor cloud infrastructure.
Design, improve, and manage our Kubernetes platform and related observability tools.
Support the migration of workloads to Kubernetes and implement scalable infrastructure solutions.
Improve Infrastructure as Code practices and CI/CD workflows using tools like Terraform and
Azure DevOps.
Research, test, and introduce new technologies to improve system performance and team efficiency.
Contribute to a developer-friendly environment by mentoring and enabling squads to manage and operate their own services.
Ensure operational excellence through automation, monitoring, and proactive incident management.
Requirements
5+ years of hands-on experience as an SRE or DevOps Engineer.
Expertise with PaaS services in at least one major cloud provider (Azure preferred).
Experience working with Terraform and Infrastructure as Code practices.
Strong knowledge of Kubernetes, Helm, and containerized workloads.
Solid understanding of observability stacks (Grafana, Prometheus, Loki, Opsgenie...).
Familiarity with message brokers (RabbitMQ or similar).
Scripting skills in Bash, Python, PowerShell, or similar languages.
Comfort with Linux systems and networking fundamentals.
Experience operating large-scale, complex cloud environments.
Nice to Have
Experience with .NET and Azure DevOps ecosystems.
Knowledge of mobile release lifecycles.
Familiarity with OpenTelemetry and distributed tracing.
Understanding of data technologies like Cosmos DB, MongoDB, or Redis.
Experience…