Reliability & Observability Engineer (m/f/d)
Certivity GmbH · München - HQ · On-site
Posted Oct 7, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About the Job
Our crawlers collect regulatory updates from 80+ sources and push them through an automated pipeline: extraction, embedding, search, classification, consolidation and translation. As we add regions, keeping it healthy has become a job of its own. We want an engineer who makes production tell us what's wrong before customers do, fixes what can be fixed automatically, and builds AI agents that diagnose the rest. When an issue does reach a developer, it should arrive with a root cause and a suggested fix. Not a ticket-driven ops role: you'll write production Python from week one and own platform reliability alongside a core-team developer.
What you will do
Cut the noise. Separate transient failures (network blips, timeouts, errors that vanish on rerun) from real ones, and build alerting the team trusts.
Build agentic incident response: agents that gather Sentry issues, logs, metrics and recent deploys, classify the failure, apply known fixes or open draft PRs, and brief the right developer.
Monitor the data, not just the infrastructure. A job that succeeds but extracts nothing is still a failure. Track freshness and completeness, e.g. "every source checked on time" and "every document produced text and embeddings".
Make the pipeline self-healing: retries with backoff, idempotent and resumable jobs, dead-letter handling and clear escalation when automation gives up.
Keep agents safe: scoped permissions, audit trails, human approval for risky actions and tracking of how often they're right.
Continuously audit our infrastructure to find ways to make it more efficient and reduce costs.
Catch memory, timeout and cost problems across Cloud Run before they become silent OOM kills or surprise bills.
Add structured logging, metrics, tracing and sensible Sentry grouping, with infrastructure as code and CI/CD.
Our Tech Stack
Python, Docker, GCP (Cloud Run, Cloud Logging, Cloud Monitoring), Sentry, MongoDB, Azure Blob Storage, LLM APIs. Some Go and React .
Your profile
5+ years in software engineering, SRE or platform roles, with real production Python experience.
A track record of turning an ignored alert channel into one people act on.
Hands-on experience building with LLMs or agents in production, and the judgment to know when a plain if-statement is better.
Solid experience with GCP (or similar), containers and CI/CD.
A strong grasp of distributed-system failure modes: retries, idempotency, partial failure and backpressure.
Pragmatism and clear communication in a small team.
Why us?
Flexible working hours.
26 + 4 vacation days per year (4 fixed "company rest days" over Christmas).
30 days of "workation" per year, within the EU and selected countries.
High autonomy and flat hierarchies.
EGYM Wellpass for unlimited access to fitness courses and gyms.
Udemy access for online courses
Closing
We welcome candidates from all backgrounds and value diversity in our team. We especially encourage women and non-binary engineers to…