Senior Software Engineer, Machine Learning Platform
Chime Financial, Inc · San Francisco, CA, USA · United States · On-site
Pay: USD 187,000 – 259,000 a year
Posted Sep 11, 2026
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
About the role
Chime’s Machine Learning Platform (MLP) team builds and operates the infrastructure, tooling, and developer experience that powers machine learning across the company. We enable data scientists and ML engineers to develop, train, deploy, and monitor models reliably and efficiently.
As a Senior Software Engineer on the Machine Learning Platform team, you will design and build scalable systems spanning traditional machine learning and emerging AI workloads, including model training, feature computation, real-time inference, foundation-model access, evaluation, and agentic orchestration. You’ll work at the intersection of distributed systems, cloud infrastructure, applied machine learning, and AI product engineering.
This role focuses on creating secure, reliable, and reusable platform capabilities that help teams choose the right approach, from conventional predictive models to LLM-powered and multi-step agentic systems, while maintaining strong standards for evaluation, observability, governance, privacy, and cost efficiency.
In this role, you can expect to
Design, build, and operate scalable ML and AI infrastructure on AWS.
Design and operate shared platform capabilities for LLM and agentic workloads, including model access, prompt and configuration lifecycle, retrieval, tool integration, state management, and workflow orchestration.
Build evaluation frameworks for non-deterministic AI systems, including offline benchmarks, regression testing, online quality signals, human feedback, and failure analysis.
Establish observability, reliability, and governance for models and agents, covering traces, model and prompt versions, tool calls, latency, token usage, quality, safety, privacy, and cost.
Help teams make principled architecture decisions across traditional ML, LLM-powered applications, and agentic workflows, and contribute to the platform’s technical roadmap.
Build distributed training, batch inference, and large-scale processing systems using frameworks such as Ray or Spark.
Build and maintain infrastructure as code using Terraform.
Support and evolve the feature store and feature pipelines.
Develop data ingestion and streaming systems using technologies such as Kinesis, Kafka, Flink, or Spark.
Improve CI/CD workflows for ML models, AI applications, and platform components.
Partner closely with Data Science and ML Engineering teams to improve developer experience.
Participate in on-call rotations to support production systems.
To thrive in this role, you have
Knowledge of the machine learning development lifecycle, including data preprocessing, model training, evaluation, deployment, and monitoring.
Experience designing distributed systems and large-scale data or compute platforms on AWS using frameworks such as Spark or Ray.
5+ years of experience in ML or AI infrastructure, platform engineering, distributed systems, or production ML systems.
Working knowledge of LLM application…