JobRaahGet matched free

Jobs

ML Systems Engineer – Training & Inference

Zenteiqai · Head Office - Bengaluru · India · On-site

Posted Sep 4, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About ZenteiQ ZenteiQ is a deep-tech company born out of IISc Bangalore, building Scientific Intelligence Infrastructure: physics-native AI for engineering, manufacturing, energy, mobility and national systems. Rather than wrapping general-purpose language models, we train foundation models (BrahmAI) from scratch to reason over thermal, electromagnetic, structural and materials domains, and put them to work through industrial platforms (KogneX) and a talent OS for engineers and researchers (AhamX). We're backed by the IndiaAI Mission (MeitY) and work closely with IISc, ARTPARK and a national AI Hub Network — building the sovereign AI infrastructure that India's engineering and industrial systems will run on. About the Role You will build the systems that let our researchers train, evaluate, and serve large foundation models reliably at scale. This role sits at the intersection of model research and infrastructure, with a focus on accelerator-based (TPU) training and inference, performance, reproducibility, and researcher velocity. You will own the paved path that turns expensive, long-running model runs into a repeatable, observable, and cost-efficient process. What You'll Do ● Build and improve the paved path for distributed training, evaluation, experiment tracking, checkpoint management, model release, and inference on Cloud TPUs. ● Operate long-running ML workloads with strong observability, failure detection, automated recovery, and practical operational tooling. ● Profile and remove bottlenecks across compute, HBM and host memory, data input, networking and collectives, XLA compilation, checkpointing, and serving. ● Build validation, representative-scale testing, CI/CD, reproducibility, and lineage systems that catch problems before expensive model runs or production releases. ● Create reusable APIs, abstractions, and self-service tooling that help researchers move quickly while preserving useful low-level controls. ● Own accelerator capacity workflows, including TPU provisioning, quotas, reservations, priorities, topology-aware placement, utilization, and cost efficiency. ● Partner with research teams to debug model and systems failures, translate recurring pain points into durable platform improvements and define operational standards. What We're Looking For ● Strong software engineering skills and experience owning systems used by researchers or engineers in production or research-critical environments. ● Hands-on experience with ML training or inference infrastructure at meaningful scale, including large accelerator or distributed compute workloads. ● Strong distributed-systems fundamentals and experience with Kubernetes, GKE, or comparable orchestration for long-running compute jobs. ● Ability to debug across model code, data pipelines, runtimes, cluster services, storage, networking, and accelerator behaviour. ● Strong observability, reliability, and performance-engineering…