JobRaahGet matched free

Jobs

Site Reliability Engineer — Multi-Cloud Infrastructure

Skit · Bangalore · India · On-site

Posted Jul 5, 2026

Apply with JobRaah

Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.

About the Role Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/ Job Title: Site Reliability Engineer — Multi-Cloud Infrastructure Type: Full-time Location: Bangalore Why this role exists: We run a voice AI platform for regulated enterprises in banking, telecom, and collections, spread across AWS, GCP, and Azure — for resilience, for cost, and because client data-residency rules leave us no choice. That's a lot of surface area: compute, networking, storage, identity, clusters, pipelines, and supporting services, all needing to stay healthy across three providers. This role owns the day-to-day reliability and operations of that estate. It's the generalist counterpart to our real-time-platform SRE: where they go deep on the latency-critical call path, you go broad — keeping the whole infrastructure dependable, well-automated, and cost-sane, and sharing the on-call load. If you like knowing how everything fits together and making the boring parts reliable and self-serve, this is a good seat. What you'll own: Multi-cloud operations. Provision, operate, and keep healthy compute, networking, storage, and identity across AWS, GCP, and Azure — with sensible consistency instead of three snowflakes. Infrastructure as code. Manage the estate through Terraform (or equivalent) and version control — reproducible environments, reviewed changes, no undocumented hand-tweaks. CI/CD and delivery. Keep build and deploy pipelines fast and reliable so engineers ship safely and often. Clusters and workloads. Run Kubernetes/container platforms and the supporting services (databases, queues, caches, internal tooling) that everything depends on. Monitoring and on-call. Maintain monitoring and alerting for infrastructure health, take a turn in the rotation, and respond to and mitigate incidents with clear communication and blameless follow-up. Cost and hygiene. Keep an eye on cloud spend, rightsizing, and waste; own the unglamorous but essential hygiene — patching, backups, secrets, and access. Automation and toil reduction. Replace manual, repetitive operations with automation and self-service so the team scales without headcount scaling with it. What the first year looks like: First 90 days. Learn the estate across all three clouds. Take a…