Role description
We are seeking an experienced Data Engineer to design, develop, and maintain scalable data pipelines and data processing solutions in a cloud-based environment. The ideal candidate will have strong hands-on experience with Google Cloud Platform (GCP), Apache Airflow DAGs, Java, Apache Spark, and SQL. The Data Engineer will work closely with data architects, software engineers, analysts, and business stakeholders to build reliable and high-performance data solutions. This role requires strong programming, data engineering, troubleshooting, and problem-solving skills, with a focus on developing production-grade batch and data processing pipelines. Key Responsibilities Design, develop, test, and maintain scalable data pipelines and ETL/ELT workflows. Develop and manage Apache Airflow DAGs for workflow orchestration, scheduling, monitoring, and dependency management. Build and optimize Apache Spark jobs using Java for large-scale data processing and transformation. Develop data processing solutions using GCP services and cloud-native technologies. Write complex and optimized SQL queries for data extraction, transformation, validation, and analysis. Work with large and complex datasets across structured and semi-structured data sources. Implement data ingestion, transformation, cleansing, and integration pipelines. Monitor production pipelines and troubleshoot failures, performance issues, and data-quality problems. Optimize Spark jobs, SQL queries, and data pipelines for performance, scalability, and cost efficiency. Implement appropriate error handling, logging, ing, retry mechanisms, and operational monitoring. Develop reusable and maintainable data engineering components following coding and design best practices. Collaborate with application developers, data analysts, data scientists, and other engineering teams. Participate in code reviews and contribute to development standards and best practices. Ensure data pipelines meet requirements for reliability, availability, scalability, and data quality. Support deployment and release activities across development, testing, and production environments. Document data pipelines, workflows, technical designs, and operational procedures. Required Skills and Experience Google Cloud Platform (GCP) Strong hands-on experience developing and deploying data engineering solutions on GCP. Experience working with GCP data and compute services such as: Google Cloud Storage (GCS) BigQuery Dataproc Cloud Composer Pub/Sub Cloud Logging and Monitoring Understanding of cloud-based data architecture and distributed data processing. Experience designing scalable, secure, and reliable data pipelines in GCP. Ability to troubleshoot and optimize cloud-based data workloads. Apache Airflow / DAG Development Strong practical experience with Apache Airflow. Hands-on experience designing, developing, scheduling, and maintaining Airflow DAGs. Experience creating task dependencies, scheduling workflows, retries, sensors, operators, and failure-handling mechanisms. Experience integrating Airflow with cloud services and data processing frameworks. Ability to troubleshoot failed DAGs and identify root causes of pipeline issues. Understanding of Airflow best practices for maintainability, scalability, monitoring, and operational support. Experience with Cloud Composer or another managed Airflow environment is highly desirable. Java and Apache Spark Strong programming experience in Java. Hands-on experience developing Apache Spark jobs using Java. Strong understanding of Spark architecture and distributed data processing. Experience working with Spark transformations, actions, DataFrames/Datasets, partitioning, joins, and aggregations. Ability to optimize Spark applications for performance and efficient resource utilization. Experience processing large volumes of data using Spark. Strong understanding of distributed computing concepts and common Spark performance considerations. SQL Strong hands-on experience with SQL. Ability to write complex SQL queries involving: Joins Subqueries CTEs Window functions Aggregations Data transformations Experience working with large datasets and optimizing SQL queries for performance. Strong understanding of relational database concepts, data modeling, and query optimization. Experience with BigQuery SQL is highly desirable. Ability to perform data validation, reconciliation, and troubleshooting using SQL.
Skills
Apache Airflow, Google Cloud Platform, Java, SQL
About UST
UST is a global digital transformation solutions provider. For more than 20 years, UST has worked side by side with the world’s best companies to make a real impact through transformation. Powered by technology, inspired by people and led by purpose, UST partners with their clients from design to operation. With deep domain expertise and a future-proof philosophy, UST embeds innovation and agility into their clients’ organizations. With over 30,000 employees in 30 countries, UST builds for boundless impact—touching billions of lives in the process.