This is a remote position.
Position Requirements:
Must Have:
- Experience in designing, building and maintaining highly scalable, distributed, cloud-based CI/CD pipelines and automation solutions
- Strong expertise in the end-to-end implementation of DevOps solutions including build automation, release management and deployment pipelines in AWS /Azure
- Hands-on experience with Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, Ansible and configuration management using
Python / Java Basics
- Experience in translating complex functional and technical requirements into detailed CI/CD pipeline and infrastructure design
- Should be aware of implementing deployment strategies such as blue-green, canary and rolling deployments
- Should have experience with cloud security, IAM roles, secrets management and access controls design
- Deep understanding of containerization and orchestration using Docker and Kubernetes
- Proficient in monitoring and logging tools such as Prometheus, Grafana, ELK Stack or Splunk
- Strong AWS hands-on expertise with a scripting background preferably Python/Bash/Shell
- Good knowledge of build automation tools - Experience in Maven, Gradle or npm is mandatory
- Strong experience in AWS/Azure compute services like EC2, ECS and EKS and storage services like S3, EBS & EFS
- Good experience with any one of the artifact/package repositories like Nexus, Artifactory or Amazon ECR
- Troubleshooting and Performance tuning experience across build, deployment and production environments
- Experience in one of the workflow/orchestration tools like Airflow or similar
- Good knowledge of Application DevOps tools (Git, CI/CD Frameworks) - Experience in Jenkins or GitLab with rich experience in source code management like AWS CodePipeline, CodeBuild and CodeCommit
- Experience with AWS CloudWatch, AWS Cloud Trail, AWS Account Config, AWS Config Rules
- Strong understanding of cloud infrastructure provisioning, migration processes and project lifecycle
- Good analytical & problem-solving skills
- Good communication and presentation skills
Desired Knowledge / Skills:
- Experience in implementing Site Reliability Engineering (SRE) practices and incident management
- Experience with service mesh and observability tools, such as Istio, Datadog, or New Relic
- Worked in Offshore / Onsite Engagements
- Experience with GitOps tools such as ArgoCD or FluxCD
- Experience in AWS services like Step Functions & Lambda for automation