Software Engineer, Distributed Systems
ebay · San Jose · On-site
Pay: USD 172,000 – 229,600 a year
Posted Sep 28, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
At eBay, we're more than a global ecommerce leader — we’re changing the way the world shops and sells. Our platform empowers millions of buyers and sellers in more than 190 markets around the world. We’re committed to pushing boundaries and leaving our mark as we reinvent the future of ecommerce for enthusiasts.
Our customers are our compass, authenticity thrives, bold ideas are welcome, and everyone can bring their unique selves to work — every day. We're in this together, sustaining the future of our customers, our company, and our planet.
Join a team of passionate thinkers, innovators, and dreamers — and help us connect people and build communities to create economic opportunity for all.
About the team and the role:
The Observability Platform team builds and operates the infrastructure that helps eBay teams monitor, fix, and improve the reliability of large-scale distributed systems. This platform supports the telemetry and reliability needs of thousands of microservices across eBay and operates at hyperscale, processing billions of time series and petabytes of log data using modern open-source technologies including Prometheus, ClickHouse, OpenTelemetry, and related tools.
As a Software Engineer on this team, you will design and build scalable distributed systems that power metrics, logs, traces, and related observability workflows across the stack—from ingestion and storage to query and visualization. You will partner closely with SREs, platform engineers, and service owners to solve complex reliability challenges, improve operational excellence, and help shape the future of observability at eBay.
This role offers the opportunity to work on critical systems, contribute to open-source technologies, and grow through direct exposure to some of eBay’s most complex infrastructure challenges. This role also includes participation in the team’s on-call rotation in support of production reliability.
What you will accomplish:
Design and deliver scalable, fault-tolerant observability infrastructure that improves reliability while reducing operational overhead for platform and engineering teams
Build and optimize high-throughput services for ingesting, transforming, storing, and querying telemetry data across logs, metrics, and traces
Strengthen the resilience of Kubernetes-based production systems through self-healing, autoscaling, and robust operational design
Partner with SREs, platform teams, and service owners to translate observability needs into tools and platform capabilities that improve incident response and operational excellence
Contribute to architecture reviews, production readiness discussions, and post-incident findings to drive continuous improvement across the platform
Expand your technical breadth by working across distributed systems, cloud-native infrastructure, and optionally user-facing observability experiences.
What you will bring:
7+ years of experience in software engineering, infrastructure…