JobRaahGet matched free

Jobs

Software Engineer, Full-Stack

Ema · San Francisco Bay Area · United States · On-site

Posted Oct 1, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

ABOUT EMA Ema builds AI Employees for HR, IT and Finance. Our AI Employees take on the busy work across the employee experience, from recruiting, onboarding and benefits to IT support, invoice processing and payroll, so people can spend their time on work that needs them. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs. We are backed by industry leading investors including Accel, Naspers/Prosus, Creaegis, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production. We ship real systems that run real business processes at scale. BUILD AGENTS THAT LEARN FROM THE WORK THEY DO. Ema builds AI employees that carry out complex workflows across enterprise applications. Our ML team works on the loop that makes them better: production traces become data, data becomes training and evaluation, and better agents produce better traces. The hard part is deciding which intervention will improve behavior in the next real workflow. THE PROBLEM SPACE - Harnesses and inference-time compute. Design context, tools, skills and orchestration for multi-step agents, including work across documents, slides, images, audio and video. Test where extra reasoning, search or verification earns its latency and cost. Build self-improvement loops with explicit permissions, evaluation gates and rollback. - Agent post-training. Curate trajectories for SFT, optimize preferences, or run RL on real agent tasks. Investigate methods such as DPO, GRPO or DAPO where they fit; compare process and outcome supervision, shape rewards, and distill useful frontier behavior into smaller models. Measure whether gains transfer beyond the training environment. - Environments and rewards. Turn enterprise workflows into reproducible training and evaluation environments: fixture tenants, simulated users who may get impatient and leave, and rewards grounded in verifiable outcomes. Find the shortcuts an agent can exploit before a training run optimizes for them. - Data engines and evaluation. Mine production agent-steps for failures; build curated corpora and useful synthetic augmentation. Calibrate judges against human labels, construct behavior-level benchmarks from real workflows, and quantify data quality, performance uplift and reliability across stochastic runs. - Retrieval, memory and context graphs. Connect enterprise information with user- and tenant-level learnings. Separate failures of retrieval from failures to use retrieved context; test what to retain, update and retrieve so that past experience improves the next decision. - Quality per dollar. Build and evaluate routing, ensembles, caching and small-model specialization. Measure downstream task success alongside latency and cost; a…