JobRaahGet matched free

Jobs

Kernel Engineer (Custom Silicon), Hardware

River AI Inc. · Palo Alto, CA; Austin, TX · United States · On-site

Pay: USD 200,000 – 420,000 a year

Posted Jun 29, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

At River, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructure, next-generation UIs, and frontier deep learning research. Who we are We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models. About the Role We are looking for exceptional performance and kernel generation engineers to build the foundational compute engine for our high-performance custom silicon. In this role, you will design and implement robust kernel generators that programmatically emit optimized low-level assembly code for our greenfield hardware architecture. You will bridge the gap between high-level compilation and raw hardware capability, pushing our custom architecture to its absolute theoretical limits for critical deep learning operations (including GEMMs, FlashAttention, and custom activations). You will collaborate closely up and down the stack with compiler engineers, silicon architects, and deep learning researchers to unlock maximum compute efficiency. What You’ll Do Kernel Generator Development : Design and build C++ code-generation frameworks and meta-programming toolchains that automatically emit optimized custom ISA assembly code. Low-Level Compute Optimization : Author and optimize core deep learning primitives (GEMM/MatMul, Attention mechanisms, Convolutions, and element-wise layers) directly targeted at our custom hardware. Microarchitectural Tuning : Hand-craft and automate instruction scheduling, register allocation, and software pipelining to maximize ALU utilization and hide execution latency on our silicon. Memory Hierarchy Management : Design sophisticated tiling, double-buffering, and data-movement strategies to optimize on-chip SRAM utilization and minimize memory bandwidth bottlenecks. HW/SW Co-Design : Partner with the RTL and architecture teams to evaluate hardware simulations, provide feedback on the ISA, and influence the design of future compute units based on kernel execution profiles. Performance Profiling & Validation : Benchmark generated assembly against hardware simulators and silicon, utilizing hardware performance counters to eliminate performance gaps and ensure mathematical correctness. Minimum Qualifications: Bachelor’s degree in Computer Engineering, Computer Science, Electrical Engineering, or a related field, and 5+ years of practical industry experience in low-level performance programming. Deep understanding of hardware programming models (e.g., CUDA, Triton, CUTLASS, or custom accelerator assembly) and a proven track record of shipping highly optimized kernels. Advanced knowledge of Computer Architecture, including vector units,…