Like Minded People
Work Together

Back to Career

Lead Platform Engineer/Lead DevOps Engineer

Location: India, Hyderabad (Hybrid)

Work Experience: 13–15 Years (including 4+Years in a lead role)

Requirements:

Mandatory Skills:

  • Strong hands-on experience in Kubernetes cluster management, administration, scaling, and troubleshooting.  
  • AWS Bedrock
  • Strong experience in managing AWS EKS and Azure AKS production clusters.  
  • Strong experience in Kubernetes upgrades, patching, node management, autoscaling, ingress, storage, networking, RBAC, and secrets management.  
  • Working experience with Azure or private cloud environments.  
  • Strong hands-on experience with Terraform for infrastructure provisioning and automation.  
  • Hands-on experience in AWS Bedrock or exposure to AI/GenAI platform services is preferred.
  • Experience supporting AI/ML or GenAI workloads on cloud platforms is an added advantage.
  • Experience in Terraform modules, remote backend, state management, workspaces, variables, outputs, and environment-based deployments.  
  • Hands-on experience with CI/CD tools such as Jenkins, GitHub Actions, or Azure DevOps.  
  • Experience in designing and maintaining CI/CD pipelines for application and infrastructure deployments.  
  • Strong automation and scripting skills using Bash, Shell, and Python.
  • Strong understanding of cloud security, IAM, RBAC, compliance, governance, encryption, secrets management, and network security.  

Desired Skills:

  • CKA certification, AWS Solutions Architect or AWS DevOps Engineer certification.
  • Azure Administrator or Azure Solutions Architect certification.  
  • Experience with multi-cloud environments including AWS, Azure, and private cloud.  
  • Hands-on experience with Helm charts, Kustomize, and Kubernetes operators.  
  • Experience with Kubernetes add-ons such as Karpenter, Cluster Autoscaler, External DNS, Cert Manager, CSI drivers, Metrics Server, and Ingress Controllers.  
  • Experience in cloud cost optimization, tagging governance, budget alerts, and FinOps reporting.  
  • Exposure to AWS Bedrock or AI/ML platform services is an added advantage.  
  • Exposure to MLOps tools such as MLflow, model registry, feature store, or AI workload operations is good to have.  
  • Experience with configuration management tools such as Ansible, Chef, Puppet, or similar platforms.  

Qualifications: Bachelor's or Master's degree in Computer Science, Information Technology, or a related field.

Job Description:

  • Design, build, and manage scalable platform engineering solutions across AWS, Azure, and Kubernetes environments.
  • Lead implementation and operations of Kubernetes platforms such as EKS, AKS, and other container orchestration platforms.
  • Define platform standards, reusable templates, golden paths, and self-service capabilities for application teams.
  • Build and maintain CI/CD and GitOps-driven deployment models using Jenkins, GitHub Actions, GitLab CI, Azure DevOps, Argo CD, or Flux.
  • Develop and manage Infrastructure as Code using Terraform modules, remote state, reusable patterns, and environment-based deployments.
  • Drive Kubernetes best practices for RBAC, namespaces, secrets, ingress, autoscaling, storage, networking, security, and upgrades.
  • Improve developer productivity by creating standardized deployment workflows and automation frameworks.
  • Implement observability, logging, monitoring, and alerting for cloud-native platforms.
  • Ensure high availability, scalability, reliability, security, and cost optimization of platform services.
  • Troubleshoot complex Kubernetes, cloud, CI/CD, networking, and infrastructure issues.
  • Define DevSecOps practices including image scanning, code scanning, vulnerability management, secrets management, and policy enforcement.
  • Collaborate with development, security, infrastructure, architecture, and operations teams.
  • Mentor junior and mid-level engineers on Kubernetes, cloud, Terraform, CI/CD, and platform engineering best practices.
  • Support AI/ML and GenAI platform requirements, including AWS Bedrock exposure where applicable.
  • Drive continuous improvement in automation, reliability, operational maturity, and platform adoption.