Site Reliability Engineer (SRE)
Social Discovery Group · Worldwide · Remote
Posted Oct 6, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Social Discovery Group (SDG) is a group of social discovery companies. SDG solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal. SDG’s products redefine the way people interact and connect with one another.
Our portfolio includes social entertainment platforms designed to connect people online across different cultures and regions of the world.
We bring together a team of like-minded people and IT professionals who specialize in creating and developing globally impactful social discovery products. Our international team of digital nomads works remotely from all over the world.
We’re proud to be a two-time “Great Place to Work” winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).
We are looking for a Site Reliability Engineer (SRE) passionate about infrastructure reliability, automation, and the development of scalable production systems.
Your main tasks will be:
Own and improve production infrastructure reliability and stability
Prepare, execute, and support deployments and infrastructure changes
Build and maintain Infrastructure-as-Code solutions using Ansible and Terraform
Support and optimize Kubernetes-based and containerized environments
Develop automation scripts and internal operational tooling
Monitor system health, investigate incidents, and proactively improve observability
Participate in CI/CD improvements together with Development, QA, DevOps, and SRE teams
Work with monitoring and alerting systems to reduce downtime and improve system performance
Maintain technical documentation, runbooks, and operational procedures
Support DNS, WAF, CDN, and caching infrastructure where required
We expect from you:
3+ years of experience in SRE, DevOps, System Administration, or Build/Release Engineering
Strong Linux administration and troubleshooting skills
Hands-on experience with Kubernetes and containerization technologies (Docker/Podman)
Experience with CI/CD pipelines, preferably GitLab CI
Practical experience with Infrastructure-as-Code and configuration management tools (Ansible and/or Terraform)
Experience with observability and monitoring tools such as Prometheus, Grafana, Zabbix, or VictoriaMetrics
Good understanding of networking fundamentals, DNS, HTTP/HTTPS, load balancing, and troubleshooting
Experience with Git and modern software delivery workflows
Ability to work independently, take ownership, and proactively improve infrastructure
Fluent Russian level for technical documentation and team communication
Nice to have:
AWS or GCP experience
RabbitMQ / AMQP experience
Cloudflare, Akamai, WAF, CDN experience
Experience with tracing and advanced observability tooling
What do we offer:
REMOTE OPPORTUNITY to work full-time;
Vacation 28 calendar days per year;
7 wellness days per year (time off) that can be used to deal with household issues, to lie down and recover without taking…