Architect - Cloud Automation
We are seeking an experienced Cloud Automation Engineer to design, develop, and implement automated solutions for managing and scaling our infrastructure. As a Cloud Engineer, you will leverage your software engineering skills to build tools, frameworks, and pipelines that streamline infrastructure provisioning, configuration, deployment, and monitoring. You'll partner closely with SREs, DBREs, CloudOps, and application teams to ensure Reliability, Scalability, Resilience, Security, Cost optimisation and Operational Excellence.
Responsibilities:
- Design, implement, and maintain Infrastructure as Code (IaC) solutions using tools such as Terraform, CRD, and Pulomi to automate Infra provisioning and management.
- Build and manage CI/CD framework, platforms and amp; pipelines (e. g., Github Actions, JenkinsX, GitLab, ArgoCD) for infrastructure and applications deployments, ensuring seamless and safe delivery processes across environments.
- Automate across hybrid cloud (e. g., AWS, Azure, GCP), ensuring consistency, scalability, and security.
- Develop automation for monitoring setup, alerting, and self-healing/fault response, integrating with tools like Prometheus, Grafana, Observe or NewRelic.
- Improve infrastructure reliability by automating routine maintenance, patching, and backup processes.
- Work with application, platform, and SRE teams to define requirements, develop solutions, and hand off automation to operations
- Create and maintain comprehensive documentation for automation workflows, infrastructure patterns, and runbooks; contribute to developing Cloud engineering best practices.
Requirements:
- Bachelor's degree in computer science, Engineering, or related field (or equivalent practical experience).
- 10+ years of experience in infrastructure software engineering, systems engineering, or Reliability Engineering roles.
- Infrastructure as Code: Strong hands-on experience in Terraform, Pulumi, CRD, etc.
- Proficiency in Coding languages (Go or Python).
- Experience with container management and orchestration (Docker, Kubernetes, Helm).
- Working knowledge of CI/CD pipelines, version control (Git), and related tooling.
- Observability as Code: Good knowledge of monitoring, logging, and alerting stacks.
- Solid understanding of networking, security, and troubleshooting.
- Strong troubleshooting skills and a passion for automation and process improvement.
- Excellent communication skills and ability to work in a collaborative team environment.
- Exposure to hybrid/multi-cloud architectures.
- Background in performance tuning, cost optimisation, or incident response automation.