Design, implement, and manage cloud infrastructure on Amazon Web Services (AWS) and Google Cloud Platform (GCP) to support business and operational requirements
Build, manage, and optimize Kubernetes platforms for production environments, ensuring scalability, high availability, and reliability
Develop, maintain, and optimize CI/CD pipelines to enable fast, secure, and reliable application deployments
Implement Infrastructure as Code (IaC) using Ansible to automate infrastructure provisioning, configuration, and management
Perform monitoring, troubleshooting, capacity planning, and performance tuning for cloud infrastructure, applications, and platforms
Manage cloud security, including Identity and Access Management (IAM), secrets management, and the implementation of DevSecOps best practices
Collaborate closely with Development, QA, Security, and Infrastructure teams to support the software development and delivery lifecycle
Design and ensure the implementation of High Availability (HA), Disaster Recovery (DR), backup strategies, and business continuity solutions in accordance with organizational standards
Optimize cloud resource utilization and implement cost optimization strategies to improve operational efficiency
Lead incident response, problem management, and conduct Root Cause Analysis (RCA) for critical incidents to drive continuous improvement
Create and maintain technical documentation, platform standards, and DevOps best practices to ensure consistency and knowledge sharing
Provide technical guidance, mentoring, and code reviews to team members, fostering engineering excellence and continuous skill development
Requirements
Person We Are Looking For:
Bachelor's degree in Computer Science, Information Systems, Information Technology, Software Engineering, or a related field
Minimum 7 years of experience as a DevOps Engineer, Site Reliability Engineer (SRE), or Cloud Engineer
Minimum 2–3 years of experience in a Senior or Lead role
Proven experience managing production environments and high-availability infrastructures
Mandatory Skills Cloud Platforms: Amazon Web Services (AWS) and Google Cloud Platform (GCP). Kubernetes (EKS, GKE, and self-managed Kubernetes). Docker and container platforms. Infrastructure as Code (IaC) using Terraform and Ansible.CI/CD tools such as GitLab CI, GitHub Actions, and Jenkins. Linux administration (Ubuntu, RHEL, Rocky Linux). Networking fundamentals, including TCP/IP, DNS, VPN, Load Balancers, Reverse Proxies, and TLS/SSL.
Monitoring and observability tools such as Prometheus, Grafana, ELK Stack, and OpenSearch.
Scripting using Bash and Python.Git and GitOps practices, including Argo CD.3. Technical Competencies
Design and manage scalable, secure, and highly available cloud infrastructure.
Build, operate, and optimize production-grade Kubernetes environments.
Develop, maintain, and optimize CI/CD pipelines and implement and manage Infrastructure as Code (IaC).Troubleshoot issues across cloud platforms, Kubernetes, Linux, networking, and application deployments
Implement DevSecOps practices, including Identity and Access Management (IAM), Secrets Management, and container security
Lead the implementation and continuous improvement of DevOps practices and cloud platforms
Conduct code reviews and architecture reviews to ensure engineering quality and best practices
Mentor and provide technical guidance to engineering teams and lead incident response, problem management, and Root Cause Analysis (RCA) for critical incidents
Collaborate effectively with Development, QA, Security, and Infrastructure teams to support the software delivery lifecycle
Knowledge of Service Mesh technologies such as Istio and Cilium
Experience with Chaos Engineering practices and Crossplane
Familiarity with AI-powered developer tools such as Claude Code, GitHub Copilot, Codex CLI, and Gemini CLI.6.
Cloud certifications are considered a strong advantage (AWS Certified Solutions, Professional AWS Certified DevOps Engineer – Professional Google Professional Cloud Architect Google Professional Cloud DevOps Engineer)