JobRaahGet matched free

Jobs

Performance Engineer - Inference

Zenteiqai · Bengaluru Head Office · India · On-site

Posted Sep 25, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About the Role We are looking for a Performance Engineer, Inference to understand and improve the systems that serve our foundation models. Inference is a tightly coupled system spanning model execution, serving runtimes, distributed systems, accelerators, scheduling, memory, and reliability. You will measure the system end to end, identify the highest-leverage performance gaps, and work across teams to close them while preserving correctness. Responsibilities • Run cross-layer performance investigations across throughput, latency, memory efficiency, reliability, and cost. • Build profiling, benchmarking, and observability tools that make inference performance measurable and explainable. • Identify bottlenecks across model servers, batching and scheduling, distributed execution, memory systems, and accelerators. • Partner with model, platform, and infrastructure teams to prioritize and land high-impact optimizations. • Validate that performance improvements preserve model quality and numerical correctness. Minimum Qualifications • Hands-on experience profiling and optimizing ML systems or other performance-critical production systems. • Production experience with at least one modern inference stack such as vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, or an equivalent serving runtime. • Experience serving or operating large models across accelerator-backed infrastructure, including multi-GPU or multi-accelerator systems. • Strong Python skills and the ability to read, instrument, and modify large production codebases. • Solid understanding of transformer inference, distributed systems, latency/throughput trade-offs, and accelerator performance fundamentals. Preferred Qualifications • Experience with large-scale or multi-node inference, including tensor, pipeline, data, or expert parallelism. • Experience with GPUs, TPUs, NPUs, or other ML accelerators and associated profiling tools. • Experience with quantization, low-precision inference, KV-cache optimization, speculative decoding, or long-context serving. • Experience contributing to or modifying inference runtimes, kernels, compilers, or distributed serving components. • Experience optimizing inference for constrained or on-device environments.