Workplace
hybrid
Employment
Internship
Published
Aug 11, 2026
Closes
No date supplied
The role
We are seeking a Senior Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB.
As a Senior Site Reliability Engineer on this new team, you will be responsible for enabling deployment at scale of AI applications and improving the performance, scalability, and reliability of the distributed systems infrastructure for this new product. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. MongoDB engineering teams pride themselves on building high-quality software and living MongoDB cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition.
We are looking to speak to candidates who are based in Gurugram for our hybrid working model.
Position Expectations
• Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads
• Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing
• Identify and configure key metrics to detect incidents and quantify service health, availability, and performance
• Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure
• Mentor early-career SREs and contribute to the team’s operational practices as it grows
Qualifications
• Strong background in software development and operating distributed systems
• 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language
• Experience operating Kubernetes in production and debugging below the abstract
Requirements
Department: Magenta