JobRaahGet matched free

Jobs

Developer Advocate, MAX Inference & Serving

Modular · United States - Remote · Remote

Pay: USD 155,400 – 233,000 a year

Posted Aug 7, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About Modular At Modular, a Qualcomm company , we’re on a mission to revolutionize AI infrastructure by systematically rebuilding the AI software stack from the ground up. Our team, made up of industry leaders and experts, is building cutting-edge, modular infrastructure that simplifies AI development and deployment. By rethinking the complexities of AI systems, we’re empowering everyone to unlock AI’s full potential and tackle some of the world’s most pressing challenges. If you’re passionate about shaping the future of AI and creating tools that make a real difference in people’s lives, we want you on our team. You can read about our culture and careers to understand how we work and what we value. About the role: We are looking for a Developer Advocate to evangelize the MAX Platform's inference and serving capabilities with our user base and developer community. This involves creating technical content such as user guides and blog posts as well as giving talks at conferences, leading workshops, all with the goal of enabling our community of builders deploying models in production. Join our world-leading product team and be part of redefining how AI infrastructure is built and deployed. LOCATION: Candidates based in the US or Canada are welcome to apply. To support growth and collaboration, those in earlier career stages work in a hybrid capacity at our Los Altos, CA. More senior staff can work out of our office in Los Altos, CA or remotely from home. Onboarding for new hires is conducted in-person in our Los Altos, CA office. Additionally, this role requires travel to conferences and developer events, which may be as often as once per month, as well as travel for team and company events (typically 2-4 times per year). What you will do: Build and publish reproducible benchmarks comparing MAX against vLLM, SGLang, TensorRT-LLM, and Triton Inference Server, including methodology, harness code, and hardware configurations so others can verify the numbers. Deploy and profile real inference workloads on MAX across CPU and GPU targets, and investigate performance gaps in latency, throughput, and cost per token. Write and maintain technical content grounded in that work: performance deep dives, serving architecture explainers, and posts that show how MAX handles batching, KV cache management, quantization, and multi-GPU serving. Author runnable tutorials, examples, and video walkthroughs covering model deployment on MAX, from pip install modular to a served endpoint under load. Provide technical support to engineers evaluating MAX for inference and serving, including teams at some of the world's largest companies, and debug their deployment and performance issues directly. Answer technical questions across GitHub, Discord, X, and LinkedIn, reproducing reported issues and filing them with enough detail for engineering to act on. Translate community and customer findings into prioritized inference and serving feedback…