JobRaahGet matched free

Jobs

Senior Site Reliability Engineer — Voice AI Platform

Skit · Bangalore · India · On-site

Posted Jul 5, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About the Role Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/ Job Title: Senior Site Reliability Engineer — Voice AI Platform Type: Full-time Location: Bangalore Why this role exists: We run a voice AI platform that places and answers up to ~1 million calls per hour for regulated enterprises in banking, telecom, and collections. Unlike most SaaS, our workload is real-time and conversational: every call is a live media session where an extra few hundred milliseconds anywhere in the ASR → LLM → TTS loop is the difference between a natural exchange and a caller hanging up. Traffic is also bursty — outbound campaigns spin up huge concurrency inside narrow calling windows — and it runs across multiple clouds for resilience and data residency. We are hiring a Senior SRE to own the reliability and performance of that system: the SLOs, the observability that makes problems visible, the capacity that absorbs campaign spikes, and the incident response that keeps regulated clients online. This is a systems-reliability role — latency, uptime, saturation, and the health of the telephony and serving path. (Model quality and evaluation live with a separate AI Observability role; you'll partner with them, not own their signals.) If you want reliability problems that are genuinely hard — real-time media, sub-second budgets, six-figure concurrency, multi-cloud failover — this is that. What you'll own: SLOs and error budgets. Define and defend service-level objectives for availability and latency across the call path, and use error budgets to steer the balance between shipping and stability. Observability. Own the metrics, tracing, and logging stack so failures surface fast and root cause is minutes not hours — distributed traces across signaling, ASR, LLM, TTS, and infra, with dashboards and alerting that page on real problems and stay quiet otherwise. The real-time media path. Keep SIP signaling and RTP media healthy at scale — concurrency, jitter, packet loss, session setup — and the reliability of the components that carry them. Capacity and autoscaling. Plan for peak (campaign windows that push toward the platform's concurrency ceiling), pre-warm capacity ahead of demand, and tune autoscaling so we neither drop calls nor burn money…