JobRaahGet matched free

Jobs

Cloud Systems Engineer

Alarm.com · Tysons, Virginia · United States · On-site

Pay: USD 100,000 – 135,000 a year

Posted Sep 25, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

We are seeking a Cloud Systems Engineer to support and operate large-scale AI and high-performance computing (HPC) environments. This role will be responsible for the deployment, maintenance, performance, and lifecycle management of GPU-accelerated compute infrastructure that powers critical AI, machine learning, and data-intensive workloads. The ideal candidate is a hands-on infrastructure professional with strong Linux administration skills, deep hardware troubleshooting experience, and expertise supporting enterprise-class compute platforms. This individual will work closely with infrastructure, networking, storage, and AI engineering teams to ensure the reliability, scalability, and operational excellence of our AI infrastructure. Responsibilities: AI Infrastructure Operations Deploy, configure, and maintain GPU-accelerated compute infrastructure. Manage operating system, firmware, BIOS, BMC, driver, and software lifecycle updates. Monitor system health, performance, utilization, and capacity across AI infrastructure environments. Support infrastructure utilized for AI model training, inference, and data processing workloads. Develop and maintain operational standards, runbooks, and maintenance procedures. Participate in on-call support and incident response activities. Linux Systems Administration Administer enterprise Linux environments, including Ubuntu and Red Hat-based distributions. Perform system patching, hardening, and operating system lifecycle management. Troubleshoot operating system, kernel, storage, networking, and application-level issues. Develop automation to streamline deployment, monitoring, and operational processes. Support security and compliance initiatives across AI infrastructure platforms. Hardware and Datacenter Operations Install, configure, maintain, and troubleshoot enterprise compute hardware. Diagnose and resolve issues involving GPUs, CPUs, memory, storage, power, and networking components. Perform firmware upgrades and hardware lifecycle management activities. Coordinate hardware replacements, vendor support engagements, and warranty services. Participate in rack-and-stack deployments, datacenter expansions, and technology refresh projects. Maintain accurate asset inventories and operational documentation. Support high-performance networking technologies, including Ethernet and InfiniBand environments. Collaborate with networking, storage, cloud, and AI engineering teams on infrastructure design and operations. Assist with scalability, resiliency, and performance optimization initiatives. Perform root-cause analysis of infrastructure failures and develop preventative measures. Other duties as assigned. Required Qualifications Experience Bachelor’s degree required 3-5 years of Linux systems administration experience in production environments. 3-5 years of experience supporting enterprise server infrastructure. Experience supporting…