Senior Data QA Engineer – Big Data & Azure Databricks
Experience: 5–10 Years
Location: Remote
CTC: Up to ₹22 LPA
Position: Senior Engineer
Environment: 60% On-Premise Big Data | 40% Azure Cloud
Role Overview
We are looking for an experienced Senior Data QA Engineer / Big Data Test Engineer with strong hands-on expertise in Hadoop, Hive, PySpark, Advanced SQL, Python, and Azure Databricks.
The role involves validating large-scale data pipelines, performing source-to-target reconciliation, testing ETL/data transformations, and automating data validation across on-premise Hadoop ecosystems and Azure cloud environments.
Key Responsibilities
- Perform end-to-end ETL/Data Warehouse and Big Data testing across source, transformation, and target layers.
- Validate large datasets within Hadoop/HDFS, including CSV and Parquet files.
- Write complex HiveQL and Advanced SQL queries for data validation, reconciliation, and transformation testing.
- Use PySpark/Spark for large-scale data processing and validation.
- Perform data testing and validation within Azure Databricks.
- Navigate Databricks workspace, manage files using commands such as %fs, and work with Unity Catalog.
- Validate data movement and consistency between on-prem Hadoop/Hive and Azure/cloud environments.
- Develop Python-based automation scripts for data validation and reconciliation.
- Perform source-to-target, record-level, aggregate-level, and apple-to-apple data comparison.
- Work with Snowflake or similar cloud data warehouse platforms.
- Use UNIX/Linux shell scripting for file validation, log analysis, and ETL testing activities.
- Monitor and validate scheduled ETL/data pipeline jobs using Control-M.
- Work with database environments including SQL Server/SSMS.
- Track requirements, user stories, defects, and test activities using Jira and Confluence.
Must-Have Skills
- Hadoop / HDFS – Strong hands-on experience
- Apache Hive / HiveQL – Strong hands-on experience
- PySpark / Apache Spark – Strong hands-on experience
- Advanced SQL – Complex queries and data validation
- Python – Strong scripting/automation experience; critical skill
- ETL / Data Warehouse Testing – Strong practical experience
- Azure Databricks – Mandatory
- Snowflake – Important
- Strong source-to-target data validation and reconciliation experience
Good-to-Have Skills
- Unity Catalog
- Databricks %fs and file management
- Python data-comparison/reconciliation libraries such as DataComPy
- Control-M
- UNIX/Linux Shell Scripting
- SQL Server / SSMS
- PyCharm
- Jira & Confluence
- Cross-platform Hadoop–Databricks connectivity
- Hive Metastore integration
Pay: ₹1,000,000.00 - ₹2,200,000.00 per year
Benefits:
- Flexible schedule
- Health insurance
- Life insurance
- Paid sick time
- Paid time off
- Provident Fund
- Work from home
Application Question(s):
- Are you based in Noida?
- How many years of experience do you have in Hadoop?
- How many years of experience do you have in Hive?
- How many years of experience do you have in Pysprak?
- How many years of experience do you have in Advanced SQL?
- How many years of experience do you have in Python?
- How many years of experience do you have in Data warehouse - Snowflake?
- What's your current CTC?
Education:
Experience:
- Data QA Engineer: 5 years (Required)
Work Location: Remote