Senior DevOps Engineer – B2B SaaS
Job Overview
We are looking for a hands-on Senior DevOps Engineer with strong experience in Kubernetes, cloud infrastructure, automation, and production troubleshooting.The role involves managing end-to-end customer deployments across cloud, BYOC, air-gapped, and data-center environments. You will work closely with Engineering, Product, and Customer Success teams to troubleshoot production issues, improve reliability, and ensure successful customer deployments.
This is a customer-facing role requiring strong ownership and a hands-on production troubleshooting mindset.
Key Responsibilities
- Own end-to-end customer deployments across cloud, BYOC, air-gapped, and data-center environments.
- Deploy, operate, and scale production environments using Kubernetes, Helm, and Terraform.
- Troubleshoot Kubernetes, networking, infrastructure, application, and deployment issues.
- Manage production incidents and customer escalations through to resolution.
- Build and maintain monitoring, alerting, logging, metrics, and observability systems.
- Automate deployment and operational processes using Python or Golang.
- Work with Engineering, Product, and Customer Success teams on production issues.
- Support high-availability and reliable production environments.
- Participate in on-call and shift rotations for critical customer escalations.
- 6+ years of hands-on DevOps experience.
- Experience working with enterprise B2B SaaS or production environments.
- Strong hands-on experience with Kubernetes.
- Good knowledge of Kubernetes workloads, networking, storage, RBAC, and security.
- Practical experience with Helm charts, templating, deployment, and troubleshooting.
- Hands-on Terraform experience for infrastructure provisioning and automation.
- Strong Linux administration and troubleshooting skills.
- Experience with at least one major cloud platform:
- AWS
- GCP
- Azure
- OCI
- Experience with monitoring, logging, metrics, alerting, and observability.
- Strong production incident-management and troubleshooting skills.
- Ability to independently handle customer escalations and production issues.
- Experience with BYOC, air-gapped, restricted-network, or enterprise data-center deployments.
- Prometheus and Grafana.
- OpenTelemetry.
- ELK / OpenSearch.
- Datadog.
- GitOps and CI/CD.
- Python or Golang scripting.
- SRE practices, including SLIs, SLOs, DR, HA, and reliability engineering.
- Multi-tenant SaaS environments.
- AI/ML infrastructure or iPaaS platforms.
The candidate should be comfortable handling urgent production issues, customer escalations, deployments, and on-call responsibilities.
Interview Process
- Technical Round 1
- Technical Round 2
- Final Cultural Fit Round
Application Question(s)
- How many years of total DevOps experience do you have?
- How many years of hands-on Kubernetes experience do you have?
- Which cloud platforms have you worked with hands-on?
Azure
GCP
OCI
Multiple
- Do you have hands-on experience with Helm?
- Do you have hands-on experience with Terraform?
- Have you independently handled production troubleshooting involving Kubernetes, networking, infrastructure, and application issues?
- Have you worked on customer deployments in any of the following environments?
- Which observability tools have you worked with?
- Do you have experience with Python or Golang for DevOps automation?
- Are you comfortable participating in on-call / shift rotations and handling critical production escalations outside regular working hours?
- What is your current CTC?
Variable CTC: ₹___ LPA
- What is your expected CTC?
- If your expected hike is above 40%, please explain the reason.
- Where are you currently based?
- What is your official notice period?