JobRaahGet matched free

Jobs

Staff DevOps Engineer

Nexxa · SF Bay area · United States · Remote

Posted Aug 27, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

Nexxa http://Nexxa.ai is building the best AI systems for heavy industries — enabling machines, systems, and operations to think, decide, and act autonomously across manufacturing, large-scale infrastructure, logistics, and legacy environments. Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry. ABOUT THE ROLE We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient — from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments. This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations — not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale. WHAT YOU'LL DO - Own and evolve Nexxa's core infrastructure — compute, networking, storage, and deployment systems — end-to-end - Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams - Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments - Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling - Partner with data and AI teams to support the infrastructure behind: - Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks) - Feature stores, embedding indices, and retrieval pipelines - Model training, evaluation, and serving infrastructure - Define and drive observability practices — metrics, logging, tracing, and alerting — across distributed systems - Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations - Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations - Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity - Collaborate with engineering leadership to define infrastructure roadmap and platform strategy - Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org REQUIRED QUALIFICATIONS - 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles - Deep hands-on experience with: - Cloud platforms (AWS, GCP, or Azure) at…