JobRaahGet matched free

Jobs

Site Reliability Engineering (SRE)

adidas · Bogota, Distrito Capital de Bogota, CO · Colombia · Hybrid

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Site Reliability Engineer Purpose of the Role As a Site Reliability Engineer, you will help improve the stability, reliability, performance, and operational readiness of business-critical digital solutions. This is not a traditional application support or service management position. You will combine software development, cloud infrastructure, DevOps, integration, and observability practices to identify recurring issues, improve existing solutions, automate operational processes, and prevent incidents before they affect users. You will work with complex, high-load, and highly integrated front-end and back-end systems in a global technology environment. Key Responsibilities Position located in Bogota, Hybrid Role (3 days onsite).  Reliability and Software Engineering Develop and improve software solutions that increase system stability, reliability, and performance. Troubleshoot complex technical issues across applications, integrations, infrastructure, and cloud environments. Identify recurring operational problems and implement sustainable technical solutions rather than temporary fixes. Refactor and enhance existing code, scripts, services, and automation. Contribute to the technical design and continuous improvement of highly integrated digital solutions. Cloud and DevOps Support and improve cloud-based solutions running in AWS environments. Work with CI/CD pipelines to automate software build, testing, release, and deployment activities. Support containerized applications and orchestration environments using Kubernetes. Contribute to infrastructure configuration, secrets management, deployment processes, and operational readiness. Work with tools such as Jenkins or comparable CI/CD technologies. Observability and Monitoring Design and implement monitoring, alerting, logging, and observability solutions. Proactively identify reliability risks, performance issues, and abnormal system behavior. Create meaningful dashboards and alerts that enable teams to detect and resolve issues effectively. Work with tools such as Grafana, Elastic Stack, Kibana, Logstash, Dynatrace, Datadog, or comparable observability and application performance management technologies. Continuously improve alert quality and reduce unnecessary operational noise. Integration and API Reliability Investigate and resolve issues across integration layers, APIs, gateways, and distributed systems. Support the reliability of solutions with multiple internal and external integrations. Collaborate with application, platform, infrastructure, and integration teams to diagnose end-to-end system issues. Contribute to improvements in API performance, availability, monitoring, and error handling. Experience with integration technologies such as TIBCO, Kong, or similar platforms would be beneficial. Incident Management and Operational Support Take ownership of complex incidents and drive them toward…