We are seeking a highly motivated and talented Associate Data Engineer to join our growing data team. This entry-level role is ideal for a recent graduate or early-career professional with a strong foundation in data principles and a passion for building robust data pipelines. The successful candidate will be expected to quickly ramp up and contribute at a level comparable to someone with a year of professional experience, focusing on developing and maintaining data infrastructure, handling diverse data types including unstructured data, and ensuring data quality and accessibility.
Roles & Responsibilities
- Design, develop, and maintain scalable and efficient data pipelines using various ETL/ELT tools and techniques.
- Ingest, transform, and load diverse datasets, including structured, semi-structured, and unstructured data, from various sources.
- Implement robust data quality checks and monitoring mechanisms to ensure data integrity and reliability.
- Collaborate with Data Scientists, Analysts, and other engineering teams to understand data requirements and deliver solutions.
- Optimize data storage and processing for performance and cost-efficiency.
- Troubleshoot and resolve issues in data pipelines and data systems.
- Stay updated with emerging data engineering technologies and best practices, contributing to continuous improvement.
Required Skills
- Strong proficiency in at least one programming language commonly used in data engineering (e.g., Python, Java, Scala).
- Solid understanding of SQL and relational databases (e.g., PostgreSQL, MySQL, SQL Server).
- Familiarity with data warehousing concepts and ETL/ELT processes.
- Basic understanding of handling unstructured data (e.g., using NoSQL databases like MongoDB, Cassandra, or object storage like AWS S3, Azure Data Lake Storage, Google Cloud Storage).
- Knowledge of data structures, algorithms, and software development best practices.
- Experience with version control systems (e.g., Git).
- Excellent problem-solving skills and analytical thinking.
- Strong communication and collaboration abilities.
- Eagerness to learn new technologies and adapt to a fast-paced environment.
Preferred Skills
- Exposure to cloud platforms (AWS, Azure, GCP) and their data services (e.g., S3, Redshift, Snowflake, BigQuery, Databricks).
- Familiarity with big data technologies (e.g., Apache Spark, Hadoop).
- Experience with workflow orchestration tools (e.g., Apache Airflow, Luigi).
- Understanding of data streaming concepts (e.g., Kafka, Kinesis).
- Basic understanding of containerization (e.g., Docker).
Benefits
- Comprehensive health Insurance
- Professional development opportunities and tuition reimbursement
- Flexible work arrangements
- Collaborative and inclusive company culture
Work Location: In person