JobRaahGet matched free

Jobs

Distributed Systems Engineer III

Mozn · Remote · Egypt · Remote

Posted Sep 20, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About Mozn MOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence. We’re a diverse, collaborative team of innovators united by a shared purpose: to build AI that delivers tangible business value, builds trust, and empowers people and organizations with augmented intelligence. Our culture is built on the relentless pursuit of excellence and meaningful impact. If you’re passionate about working alongside exceptional talent on world-class AI, and you want the autonomy and runway to do the best work of your career, join us in shaping the future of intelligent enterprises. About the role We are looking for a highly motivated Distributed Systems Engineer III to join our Cloud Platform Engineering team. This role focuses on building and operating reliable, scalable, and resilient cloud platforms for distributed and data-intensive workloads. You will work across Kubernetes, cloud infrastructure, messaging systems, databases, automation, and platform reliability. The role requires strong hands-on experience with Kafka, Kubernetes, and at least one relational database such as MySQL or PostgreSQL, along with a solid understanding of distributed-systems fundamentals. As our platform evolves, you will also contribute to AI and data infrastructure, helping build the underlying platform capabilities required to run data-intensive and AI What you'll do Cloud Platform & Distributed Systems Build, operate, and continuously improve production cloud-native platforms running distributed workloads. Work hands-on with Kubernetes, including upgrades, node pools, workload lifecycle, troubleshooting, and platform operations. Deploy and manage workloads using ArgoCD, Helm, GitOps, Terraform, and automation. Design and operate systems with a focus on scalability, availability, resilience, performance, and operational simplicity. Troubleshoot complex issues across Kubernetes, cloud infrastructure, networking, storage, applications, and distributed services. Kafka & Data Infrastructure Operate and troubleshoot Apache Kafka in production across high-throughput and distributed workloads. Work with topics, partitions, replication, consumer groups, retention, throughput, latency, and failure recovery. Integrate Kafka with databases and applications using technologies such as Kafka Connect, Debezium, or similar CDC/event-streaming platforms. Operate and troubleshoot MySQL and/or PostgreSQL, including replication, high availability, backup, recovery, performance, and migrations. Support data-intensive workloads and analytical platforms such as StarRocks, ClickHouse, Apache Doris, or similar technologies. Reliability, DR & Multi-Tenant Platforms Design and operate platforms that remain resilient across node, service, zone, and infrastructure failures. Implement and validate backup, recovery, disaster recovery,…