DP - JD - Cloud Data Engineer
Job Description
We are on the hunt for a dynamic and proficient Cloud Data Engineer to join our DataPlatform within the Data Team.
Duties and Responsibilities
Quickly learn and adapt to new technologies as the Data Team's technology stack evolves,demonstrating the ability to tackle new challenges.
Consider all aspects of usability, scalability, deployment, integration, maintenance, andautomation when integrating new technology stacks.
Demonstrate strong programming skills in at least one language (Python, Scala, Java) and theability to learn additional languages as needed.
Build and maintain ETL pipelines and data-driven systems utilizing technologies such asApache Spark, AWS Glue, Athena, Redshift, and AWS Batch.
Expertise in writing complex SQL queries is essential.
Manage code on GitHub, with a comprehensive understanding of advanced git operations,including git-flow, rebasing, and squashing.
Implement infrastructure as code using Terraform and utilize AWS Analytics and DataServices like Glue, S3, Lambda, AWS Batch, Athena, Redshift, DynamoDB, CloudWatch,Kinesis, SQS, SNS, and DMS.
Use Jenkins to implement deployment pipelines and engage in requirements gathering toestimate efforts for integrating new technology stacks.
Design and architect solutions for ML, Data Governance, Deployment/IntegrationAutomations, and Data Analytics.
Explore and learn additional AWS services such as ECS, ECR, and EC2, along with DataModeling.
Requirements
Past Job Experience
A minimum of 8 years of experience in software development, with at least 3-5 years focusedon building scalable and stable data pipelines using the AWS tech stack.
Proven experience in constructing Data Pipelines in the AWS Cloud, gained through jobexperience or personal projects.
Data and AWS Tools
Strong programming skills and proficiency in SQL.
Familiarity with a range of AWS Analytics Ecosystem components, including but not limited toApache Airflow, Apache Spark, S3, Glue, Kafka, AWS Athena, Lambda, Redshift, LakeFormation, AWS Batch, ECS - Fargate, Kinesis, Flink, DynamoDB, and SageMaker.
Should have experience in using IaC while deploying the Data Pipelines in AWS. Example -terraform, Cloudformation, etc.
Experience with Docker, Kubernetes, ECR, EC2, VPC, SNS, SQS, CloudWatch is highly valued.
Experience in building Jenkins deployment pipelines is highly valued.
Proficiency in using collaboration tools like JIRA, Confluence, GitHub is beneficial.
Exposure to NoSQL databases is an advantage.