Job description
As an Infrastructure Engineer at Zero One Group, you’ll build and maintain reliable, scalable, and secure infrastructure that supports our growing products and engineering teams. You’ll work closely with engineers to improve deployment, observability, performance, and infrastructure efficiency.
Responsibilities
- Assist with cloud infrastructure provisioning via Terraform on AWS; support GitLab CI/CD pipeline maintenance; monitor server logs and health metrics
- Understands Linux command-line tools, networking basics, Docker containerization, and AWS core concepts.
- Write modular Infrastructure as Code using Terraform / OpenTofu / Terragrunt; build GitLab CI/CD deployment pipelines; manage core AWS services (VPC, EC2, ECS, RDS, CloudFront).
- Demonstrates clear judgment on cost vs. availability trade-offs and multi-environment state management.
- Architect high-availability, fault-tolerant AWS infrastructure; standardize Terragrunt modules across teams; lead incident response for cloud outages; drive FinOps cost optimizations.
- Experienced in cloud scale (auto-scaling, multi-account structures, disaster recovery). Leads incident communications calmly.
Requirements
- Strong proficiency in Linux, AWS, Docker, Git, and Bash scripting.
- Experience with AWS services such as VPC, EC2, ECS, and RDS.
- Experience with Terraform/OpenTofu and Terragrunt for Infrastructure as Code.
- Experience building and maintaining CI/CD pipelines using GitLab CI/CD.
- Familiarity with Prometheus, Grafana, and CloudWatch for monitoring and observability.
- Understanding of AWS Multi-Account Architecture, security, and infrastructure scalability.
- Knowledge of FinOps and cloud cost optimization.
- Understanding of Disaster Recovery and high-availability strategies.
- Strong incident management and troubleshooting skills, with the ability to lead critical incidents.
- Understanding of infrastructure best practices and a continuous improvement mindset.
Nice to have
- Familiarity with Python, Ansible, and CloudWatch dashboards.
- Experience with Kubernetes (EKS), Helm, and HashiCorp Vault.
- Familiarity with AWS Lambda/Serverless.
- Understanding of Zero Trust, Chaos Engineering, and Service Mesh (Istio).
- Experience with SOC 2 compliance implementation.