Staff Software Engineer - Fleet Management
Nscale · US · United States · On-site
Pay: USD 220,000 – 320,000 a year
Posted Aug 7, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
.
About Nscale
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
About the role
Nscale is hiring a Staff Software Engineer to build Fleet Manager — the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale.
This role sits at the intersection of distributed systems, infrastructure automation, and physical hardware. You'll own domain-level architecture within Fleet Manager: Python-based systems that manage the entire operational lifecycle of our compute infrastructure, from initial device enrollment through multi-day burn-in testing to ongoing health monitoring and automated remediation. The problems are challenging and the stakes are high — the software you design and build determines how quickly and how reliably Nscale scales its GPU fleet to meet demand.
This is an opportunity to shape a foundational platform early, setting the patterns and standards that engineers across Fleet Manager build on.
What you'll work on
Device provisioning and enrollment: automation that takes bare-metal GPU nodes and network switches from first power-on to production-ready — BMC configuration, DHCP reservations, and provisioning state machines.
Burn-in and validation: multi-day testing workflows that qualify hardware before it enters, and re-enters, the fleet.
Workflow orchestration: durable, event-driven state machines that span multiple days, survive crashes, resume from checkpoints, support human-in-the-loop approval gates, and let thousands of concurrent idempotent workflows run without stepping on each other.
GPU health monitoring and self-healing: detection, diagnosis, and automated remediation workflows that keep nodes healthy in production.
Network configuration: switch lifecycle automation and network state management across the fleet.
Integrations: keeping Fleet Manager consistent with datacenter inventory tooling (DCIM, NetBox), bare-metal provisioning systems (MAAS, Ironic, IPMI), credential stores, and monitoring infrastructure.
Observability: structured logging, metrics, distributed tracing, and tooling that lets operators troubleshoot effectively.
Responsibilities
Domain-level technical direction.…