Senior Site Reliability Engineer
Descripción del puesto
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in India.
This is a senior engineering role focused on building and operating highly available, resilient distributed systems at global scale.
You will architect reliability solutions across multi-region microservices, APIs, and authentication infrastructure running on AWS and GCP.
The role combines hands-on engineering with technical leadership across observability, disaster recovery, Kubernetes, and infrastructure automation.
You will help define reliability standards, including SLIs, SLOs, error budgets, and high-availability objectives.
You will also lead major incident response and turn production learnings into lasting systemic improvements.
Beyond reliability, you will drive cloud cost optimization, automation, and AI-assisted engineering practices.
This is an opportunity to mentor engineers, shape platform architecture, and raise the technical bar in a fast-moving, remote-first environment.
Accountabilities
- Architect, scale, and continuously improve the reliability, availability, and performance of multi-region microservices, APIs, and authentication infrastructure across AWS and GCP.
- Design and maintain disaster recovery and business continuity strategies, including multi-region failover automation, recovery dashboards, and validation processes aligned with defined RTO and RPO targets.
- Establish and enforce SLI, SLO, and error-budget frameworks across engineering teams to improve reliability and accountability.
- Lead the observability strategy using platforms such as Datadog, with actionable Golden Signals monitoring designed to reduce MTTD, MTTR, and alert fatigue.
- Lead on-call escalation and major incident management, ensuring rapid response to production issues and strong adherence to availability objectives.
- Facilitate blameless post-incident reviews and drive root-cause remediation to prevent recurring failures.
- Architect and operate production Kubernetes environments, including EKS/GKE, networking, RBAC, ingress/egress, and GitOps workflows using tools such as Argo CD and Kargo.
- Design and maintain modular, enterprise-grade Terraform infrastructure across complex multi-account and multi-region cloud environments.
- Build FinOps dashboards and cost-optimization initiatives covering resource utilization, right-sizing, cost allocation, and multi-cloud spend visibility.
- Develop production-grade Python or Go tooling, automation, integrations, and platform capabilities to eliminate operational toil.
- Champion AI-assisted engineering tools to accelerate automation, runbook creation, incident triage, and engineering productivity.
- Create operational runbooks and architecture documentation while mentoring junior and mid-level engineers and contributing to technical standards.
- 8+ years of professional experience in Site Reliability Engineering, DevOps, Platform Engineering, or software engineering supporting 24/7 mission-critical distributed systems.
- Bachelor’s degree in Computer Science, Software Engineering, or a comparable technical discipline, or equivalent professional experience.
- Advanced Python or Go development skills, with experience building internal SRE platforms, automation tools, and API integrations.
- Deep hands-on Kubernetes expertise, including production EKS/GKE environments, cluster lifecycle management, networking, RBAC, ingress/egress, and GitOps tooling such as Argo CD.
- Strong Terraform expertise, including module architecture, state management, refactoring, and infrastructure deployment across multi-account AWS environments.
- Strong AWS and/or GCP expertise, including services and concepts such as IAM, VPC, Transit Gateway, ALB/NLB, Route53, and cloud networking.
- Proven experience designing and testing multi-region disaster recovery architectures, automating failover, and monitoring recovery health.
- Strong knowledge of SLI/SLO frameworks, production observability, PagerDuty, monitoring strategy, and reliability engineering practices.
- Experience designing and operating enterprise service meshes such as Istio or Linkerd and production ingress/proxy technologies such as HAProxy or NGINX.
- Demonstrated FinOps and cloud cost-optimization experience, including right-sizing, cost allocation, workload optimization, and financial visibility.
- Experience with technical leadership, architectural discussions, RFCs/design documentation, and mentoring engineering peers.
- Strong troubleshooting, communication, collaboration, and problem-solving skills, with the ability to work effectively on complex distributed systems.
- Fluency in written and spoken English and the ability to collaborate with global teams.
- Willingness and ability to participate in an on-call rotation and respond during assigned shifts.
- Preferred experience includes chaos engineering, secrets management tools such as Vault or AWS Secrets Manager, DevSecOps practices, automated infrastructure vulnerability remediation, and identity or IAM-focused platforms.
- Fully remote work from India.
- Remote-first working environment with collaboration across global teams.
- Opportunity to work on highly available, mission-critical distributed systems at significant scale.
- Exposure to modern cloud, Kubernetes, GitOps, Infrastructure-as-Code, observability, FinOps, and AI-assisted engineering technologies.
- High level of technical ownership and autonomy in a fast-moving SaaS environment.
- Opportunities to influence architecture, engineering practices, and reliability standards.
- Technical mentorship and career development opportunities.
- Collaborative culture that values connection, innovation, continuous improvement, and ambitious problem-solving.
- Equal-opportunity workplace committed to an inclusive and diverse workforce.
- Full-time employment with participation in an on-call rotation as part of the engineering role.
Requirements
Benefits
Habilidades mencionadas
Postúlate a este puesto
Use the application link supplied with this listing to apply to jobgether. Check the destination before entering personal information.
