ML Systems Engineer – Training & Inference
Zenteiqai · Head Office - Bengaluru · India · On-site
Posted Sep 4, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About ZenteiQ
ZenteiQ is a deep-tech company born out of IISc Bangalore, building Scientific Intelligence
Infrastructure: physics-native AI for engineering, manufacturing, energy, mobility and national
systems. Rather than wrapping general-purpose language models, we train foundation models
(BrahmAI) from scratch to reason over thermal, electromagnetic, structural and materials
domains, and put them to work through industrial platforms (KogneX) and a talent OS for
engineers and researchers (AhamX). We're backed by the IndiaAI Mission (MeitY) and work
closely with IISc, ARTPARK and a national AI Hub Network — building the sovereign AI
infrastructure that India's engineering and industrial systems will run on.
About the Role
You will build the systems that let our researchers train, evaluate, and serve large foundation
models reliably at scale. This role sits at the intersection of model research and infrastructure,
with a focus on accelerator-based (TPU) training and inference, performance, reproducibility,
and researcher velocity. You will own the paved path that turns expensive, long-running model
runs into a repeatable, observable, and cost-efficient process.
What You'll Do
● Build and improve the paved path for distributed training, evaluation, experiment tracking,
checkpoint management, model release, and inference on Cloud TPUs.
● Operate long-running ML workloads with strong observability, failure detection, automated
recovery, and practical operational tooling.
● Profile and remove bottlenecks across compute, HBM and host memory, data input,
networking and collectives, XLA compilation, checkpointing, and serving.
● Build validation, representative-scale testing, CI/CD, reproducibility, and lineage systems
that catch problems before expensive model runs or production releases.
● Create reusable APIs, abstractions, and self-service tooling that help researchers move
quickly while preserving useful low-level controls.
● Own accelerator capacity workflows, including TPU provisioning, quotas, reservations,
priorities, topology-aware placement, utilization, and cost efficiency.
● Partner with research teams to debug model and systems failures, translate recurring
pain points into durable platform improvements and define operational standards.
What We're Looking For
● Strong software engineering skills and experience owning systems used by researchers
or engineers in production or research-critical environments.
● Hands-on experience with ML training or inference infrastructure at meaningful scale,
including large accelerator or distributed compute workloads.
● Strong distributed-systems fundamentals and experience with Kubernetes, GKE, or
comparable orchestration for long-running compute jobs.
● Ability to debug across model code, data pipelines, runtimes, cluster services, storage,
networking, and accelerator behaviour.
● Strong observability, reliability, and performance-engineering…