JobRaahGet matched free

Jobs

Systems Operations Support Engineer — Linux

Vastai · Los Angeles · United States · On-site

Pay: USD 90,000 – 150,000 a year

Posted Jul 30, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About Us Vast.ai http://Vast.ai's cloud powers AI projects and businesses all over the world. We are democratizing and decentralizing AI computing — reshaping our future for the benefit of humanity. Our mission is to organize, optimize, and orient the world's computation. We value elegance, ownership, integrity, and continuous learning. You'll have the opportunity to dive into state-of-the-art AI systems while collaborating with a globally distributed team. About the Role This is a systems operations support role focused on deep-diving into escalated infrastructure issues that go beyond frontline triage. You’ll be the engineering resource our L1 support team relies on when tickets become complex, investigating and resolving issues across the full infrastructure stack—including hardware, BIOS and firmware, networking, Ubuntu, Docker, NVIDIA CUDA and GPUs, and KVM virtual machines. You’ll own complex escalations end-to-end: gathering evidence, reproducing issues, identifying the root cause, proposing solutions, and working with the appropriate teams to bring each issue to resolution. The best engineers in this role don’t just resolve individual tickets—they identify recurring patterns, improve operational tooling, and build runbooks that prevent future incidents. You’ll collaborate directly with the engineering and host support teams on systemic infrastructure issues. Strong Linux systems knowledge, technical depth, and support experience are the primary requirements. You should be comfortable working autonomously in Ubuntu environments, troubleshooting hardware, networking, containers, virtual machines, and GPU workloads, and clearly communicating your findings and proposed solutions to both technical and non-technical audiences. Vast.ai http://Vast.ai users or hosts strongly preferred. LOCATION AND SCHEDULE This is a full-time position based in our Westwood, Los Angeles office. Available schedules: - Monday–Friday: Fully on-site - Sunday–Thursday: Four days on-site and one day working from home Key Responsibilities - Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration - Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting - Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads - Provide coverage for L1 support overflow during peak periods or incidents - Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and KVM virtualization environments - Troubleshoot network-layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines - Investigate performance issues involving GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O…