Senior Site Reliability Engineer
Hive · Berlin · Germany · On-site
Posted Oct 1, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Total compensation: 90.000€ - 120000€ (depending on experience)
The position
If you are excited about growing with Hive and love building reliable, scalable systems, this might be the right position for you!
You'll join as a Senior Site Reliability Engineer , helping power our infrastructure, ensure platform reliability, and drive operational excellence across our systems.
Hive is building the leading operations platform for independent commerce. It's time for us to take the next step and deeply invest in our infrastructure and reliability foundations as a key enabler to enhance the value of our multi-product offering. Better commerce operations for merchants, consumers, and our fulfillment network.
What you'll do
Some exemplary topics of what we expect a Senior Site Reliability Engineer to cover are below — an important note is that we're not looking for someone who purely optimizes the technical aspects of what they're doing, but that they're doing it with their users (engineers, product teams, analysts, operations, …) and our customers in mind.
Keep production reliable: own SLOs, alerting and observability across metrics, logs and traces, take part in our paid on-call rotation, and turn incidents into lasting fixes through post-incident reviews.
Evolve our infrastructure: design, build and optimize our Kubernetes and public cloud platform, managed as code and delivered through GitOps, with performance, cost efficiency and self-service for engineering teams in mind.
Keep our data layer fast and healthy: own the performance, capacity, replication and upgrades of our PostgreSQL fleet, and the pipelines that feed our analytical backends.
Build AI as a platform capability: build AI templates and, with the product engineers, the coding standards and guardrails for AI agents. Design the data access model our apps build on, so teams can go from prototype to production safely. This is hands-on building, not advisory work.
Connect platform and product : you're the person who knows how platform and product fit together. You embed reliability into how teams ship, and make targeted changes in our Rails/React codebases when that's the fastest path.
Harden security: least-privilege access across cloud and Kubernetes, production access controls, secrets management, and timely remediation of infrastructure and container vulnerabilities.
Your profile
We know – sometimes, you can't tick every box. We would still love to hear from you if you think you're a good fit!
You run production on cloud and Kubernetes: hands-on production experience with AWS, Kubernetes (ideally Amazon EKS) and infrastructure as code (Terraform).
You've operated PostgreSQL in production: query and index tuning, replication, upgrades, and backup and recovery.
You own reliability end to end: observability (Prometheus, Grafana, Loki or similar), SLOs, incident response, and on-call for systems that matter.
You write real code: Python, Ruby, TypeScript…