AI DevOps Engineer — Mid/Senior Level
<p><strong>Job Title: </strong>AI DevOps Engineer — Mid/Senior Level</p><p><strong>Experience:</strong> 4–7 Years</p><p><strong>Location:</strong> Raidurg Main Road, Hyderabad.</p><p><strong>Work Mode:</strong> On-site</p><p><strong>Work Hours:</strong> 2-11 PM</p><p><strong>Notice Period:</strong> Immediate Joiner (15-30 days)</p><p><br></p><p><strong>About us: </strong>At Stackular, we are more than just a team – we are a product development community driven by a shared vision. Our values shape who we are, what we do, and how we interact with our peers and our customers. We're not just seeking any regular engineer; we want individuals who identify with our core values and are passionate about software development.</p><p><br></p><p><strong>About the Role</strong></p><p>We are looking for a <strong>Mid-Level AI DevOps Engineer</strong> with <strong>4-7 years of experience</strong> in DevOps, cloud infrastructure, automation, and production deployment environments. This role will focus on building, maintaining, and improving scalable infrastructure and deployment pipelines for AI and machine learning applications.</p><p><br></p><p>The ideal candidate should have strong hands-on experience with <strong>cloud platforms, CI/CD, Docker, Kubernetes, infrastructure as code, monitoring, and automation</strong>, along with a working understanding of AI/ML deployment workflows.</p><p><br></p><p><strong>Key Responsibilities</strong></p><p><strong>Cloud Infrastructure & DevOps</strong></p><ul><li>Design, deploy, and manage cloud-based infrastructure for AI and software applications.</li><li>Work with cloud platforms such as <strong>AWS, Azure, or GCP</strong>.</li><li>Build and maintain infrastructure using tools such as <strong>Terraform, CloudFormation, Ansible</strong>.</li><li>Support scalable, secure, and reliable environments for production workloads.</li><li>Optimize infrastructure for performance, cost, availability, and operational efficiency.</li></ul><p><br></p><p><strong>CI/CD & Automation</strong></p><ul><li>Build and maintain CI/CD pipelines for application and AI service deployments.</li><li>Automate build, testing, deployment, and rollback processes.</li><li>Improve deployment reliability and reduce manual operational tasks.</li><li>Work with tools such as <strong>Azure DevOps, GitHub Actions, Jenkins</strong>.</li><li>Create reusable scripts, templates, and automation workflows for engineering teams.</li></ul><p><br></p><p><strong>Containerization & Orchestration</strong></p><ul><li>Deploy and manage containerized applications using <strong>Docker</strong>.</li><li>Work with <strong>Kubernetes</strong> for application deployment, scaling, networking, and troubleshooting.</li><li>Manage Helm charts and Kubernetes manifests.</li><li>Support production deployments and ensure application availability.</li><li>Troubleshoot container, cluster, and infrastructure-related issues.</li></ul><p><br></p><p><strong>AI / MLOps Support</strong></p><ul><li>Support deployment and monitoring of AI/ML models in production environments.</li><li>Collaborate with data scientists, ML engineers, and backend engineers to streamline model deployment workflows.</li><li>Assist with model versioning, model serving, and release automation.</li><li>Work with MLOps tools such as <strong>MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, Airflow, or similar platforms</strong>.</li><li>Support infrastructure for AI services, APIs, and model inference workloads.</li></ul><p><br></p><p><strong>Monitoring, Logging & Reliability</strong></p><ul><li>Implement and maintain monitoring, logging, tracing, and alerting systems.</li><li>Use tools such as <strong>Prometheus, Grafana, ELK Stack, Datadog, New Relic, CloudWatch, or Azure Monitor</strong>.</li><li>Monitor application and infrastructure performance.</li><li>Participate in incident response, root cause analysis, and production support.</li><li>Help improve system reliability, uptime, and operational visibility.</li></ul><p><br></p><p><strong>Security & Compliance</strong></p><ul><li>Apply DevSecOps practices across infrastructure and deployment pipelines.</li><li>Manage access controls, IAM roles, secrets, and secure configuration.</li><li>Support vulnerability scanning, patching, and security hardening.</li><li>Ensure cloud and deployment environments follow security best practices.</li><li>Work with tools such as <strong>HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, or GCP Secret Manager</strong>.</li></ul><p><br></p><p><strong>Required Qualifications</strong></p><ul><li><strong>4-7 years of experience</strong> in DevOps, Cloud Engineering, Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering.</li><li>Strong hands-on experience with at least one cloud platform: <strong>AWS, Azure, or GCP</strong>.</li><li>Experience building and managing CI/CD pipelines.</li><li>Strong proficiency on skills like <strong>Go, Python, Java, Ansible, Terraform, Pulumi, Shell Scripting, Bash, or PowerShell.</strong></li><li>Strong experience with <strong>Docker</strong> and containerized deployments.</li><li>Working experience with <strong>Kubernetes</strong> in production or near-production environments.</li><li>Experience with infrastructure-as-code tools such as <strong>Terraform, Ansible, CloudFormation</strong>.</li><li>Experience with monitoring and logging tools such as <strong>Prometheus, Grafana, ELK, Datadog, New Relic, or CloudWatch</strong>.</li><li>Good understanding of networking, Linux systems, security, and cloud architecture.</li><li>Familiarity with AI/ML workflows, model deployment, or MLOps concepts.</li><li>Experience supporting production applications and troubleshooting infrastructure issues.</li></ul><p><br></p><p><strong>Preferred Qualifications</strong></p><ul><li>Experience supporting AI/ML applications or model deployment pipelines.</li><li>Exposure to <strong>LLM applications, vector databases, RAG pipelines, or generative AI infrastructure</strong>.</li><li>Experience with GPU-based workloads or AI inference infrastructure.</li><li>Familiarity with tools such as <strong>MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, Airflow, or Argo Workflows</strong>.</li><li>Experience with Helm, service mesh, or Kubernetes operators.</li><li>Knowledge of DevSecOps practices and cloud security controls.</li><li>Cloud, Kubernetes, or DevOps certifications are a plus.</li></ul><p><br></p><p><strong>Required Technical Skills</strong></p><p><strong>Cloud Platforms:</strong> AWS, Azure, GCP</p><p><strong>Containers & Orchestration:</strong> Docker, Kubernetes, Helm</p><p><strong>Infrastructure as Code:</strong> Terraform, Ansible, CloudFormation</p><p><strong>CI/CD:</strong> GitHub Actions, Jenkins, Azure DevOps</p><p><strong>Scripting:</strong> Python, Bash, PowerShell</p><p><strong>Monitoring & Logging:</strong> Prometheus, Grafana, ELK Stack, Datadog, New Relic, CloudWatch</p><p><strong>MLOps / AI Tools:</strong> MLflow, Kubeflow, SageMaker, Vertex AI, Azure ML, Airflow</p><p><strong>Security:</strong> IAM, secrets management, vulnerability scanning, DevSecOps</p><p><strong>Operating Systems:</strong> Linux, Unix-based systems</p>