Performance Engineer - Inference
Zenteiqai · Bengaluru Head Office · India · On-site
Posted Sep 25, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About the Role
We are looking for a Performance Engineer, Inference to understand and improve the systems that
serve our foundation models. Inference is a tightly coupled system spanning model execution, serving
runtimes, distributed systems, accelerators, scheduling, memory, and reliability. You will measure the
system end to end, identify the highest-leverage performance gaps, and work across teams to close
them while preserving correctness.
Responsibilities
• Run cross-layer performance investigations across throughput, latency, memory efficiency,
reliability, and cost.
• Build profiling, benchmarking, and observability tools that make inference performance
measurable and explainable.
• Identify bottlenecks across model servers, batching and scheduling, distributed execution, memory
systems, and accelerators.
• Partner with model, platform, and infrastructure teams to prioritize and land high-impact
optimizations.
• Validate that performance improvements preserve model quality and numerical correctness.
Minimum Qualifications
• Hands-on experience profiling and optimizing ML systems or other performance-critical production
systems.
• Production experience with at least one modern inference stack such as vLLM, SGLang,
TensorRT-LLM, NVIDIA Dynamo, or an equivalent serving runtime.
• Experience serving or operating large models across accelerator-backed infrastructure, including
multi-GPU or multi-accelerator systems.
• Strong Python skills and the ability to read, instrument, and modify large production codebases.
• Solid understanding of transformer inference, distributed systems, latency/throughput trade-offs,
and accelerator performance fundamentals.
Preferred Qualifications
• Experience with large-scale or multi-node inference, including tensor, pipeline, data, or expert
parallelism.
• Experience with GPUs, TPUs, NPUs, or other ML accelerators and associated profiling tools.
• Experience with quantization, low-precision inference, KV-cache optimization, speculative
decoding, or long-context serving.
• Experience contributing to or modifying inference runtimes, kernels, compilers, or distributed
serving components.
• Experience optimizing inference for constrained or on-device environments.