Senior Observability Platform Engineer
Nscale · US · United States · On-site
Pay: USD 160,000 – 230,000 a year
Posted Aug 21, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
Senior Observability Platform Engineer – Nscale
About Nscale
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency while contributing to the technology that powers the future.
About the Role
As a Senior Observability Platform Engineer, you’ll play a key role in designing, building, and scaling Nscale’s observability platform. You’ll focus on delivering reliable, high-quality visibility into GPU clusters, AI workloads, and the infrastructure that powers them.
You approach observability as a product—balancing usability, scalability, and operational efficiency. You build systems that reduce cognitive load for engineers, surface meaningful signals, and enable fast, confident debugging when things go wrong.
You’ll contribute to platform direction, implement critical systems, and collaborate closely with SRE, infrastructure, and AI/ML teams to ensure observability is embedded into everything we run.
This is a hands-on engineering role with meaningful influence over platform design and evolution.
What You’ll Do
Design, build, and operate scalable observability systems across metrics, logs, traces, and alerting
Contribute to architectural decisions around tooling, data pipelines, storage, and retention strategies
Improve signal quality by reducing noise, managing cardinality, and refining alerting practices
Help identify and address observability gaps before they impact reliability
Partner with SRE, infrastructure, and AI/ML teams to integrate observability into services and platforms
Develop reusable patterns, libraries, and best practices that improve consistency across teams
Participate in incident response and postmortems, driving actionable improvements
Evaluate and adopt tools that improve developer experience, scalability, and operational efficiency
Support and mentor engineers within the team through code reviews and knowledge sharing
About You
5+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles
Experience operating and scaling observability systems in production environments
Strong understanding of monitoring concepts: metrics, logs, traces, alerting, and SLOs
Hands-on experience with several of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic
Solid programming skills (Python, Go, or similar) with the ability to build and maintain production systems
Experience working with Kubernetes-based…