About the Role
You will own automated testing for our GCP data pipelines. This is not manual QA and it is not UI testing. You will write code that validates data pipelines — schema, row counts, transformation logic, reconciliation between source and target — and run those tests continuously across development, testing, and staging environments before anything reaches production.
You will work inside our staging environment, building and maintaining the test suites that gate every release. When a pipeline breaks in staging, your framework should catch it before a client ever sees it.
Key Responsibilities
- Design, build, and maintain automated test scripts for data pipelines on GCP, covering ingestion, transformation, and load stages.
- Execute and manage test runs across development, testing, and staging environments; maintain environment parity and test data readiness.
- Build data validation checks: row count reconciliation, schema drift detection, null and duplicate checks, referential integrity, business rule validation, and source-to-target mapping verification.
- Write regression, smoke, and integration test suites that run on every code change and on scheduled nightly cycles.
- Integrate test execution into CI/CD pipelines so that failing tests block promotion from staging to production.
- Create and manage test data sets, including synthetic and masked data where production data cannot be used.
- Investigate failures, isolate root cause between data issues, pipeline logic, and infrastructure, and log defects with reproducible detail.
- Build test reporting and dashboards so engineering leads and clients can see coverage, pass rates, and failure trends at a glance.
- Work directly with data engineers to define acceptance criteria before development begins, not after.
- Contribute to and improve the reusable QA automation framework used across client projects.
Required Skills
- Strong SQL — you can write complex analytical queries and reason about query plans, not just SELECT statements.
- Strong Python for test automation, including pytest or an equivalent framework.
- Hands-on GCP data stack experience: BigQuery, Dataform
- Practical experience testing ETL or ELT pipelines, including source-to-target reconciliation and transformation logic validation.
- Data quality tooling: Great Expectations, Soda, dbt tests, or a comparable framework.
- Version control with Git, and working knowledge of CI/CD tooling such as Cloud Build, GitHub Actions, GitLab CI, or Azure DevOps.
- Experience working in multi-environment setups (dev, test, staging, prod) and understanding of promotion and release gating.
- Defect tracking and test management discipline using Jira or similar.
Good to Have
- Experience with GCP.
- Infrastructure as code exposure — Terraform.
- Familiarity with data observability and monitoring using Cloud Logging, Cloud Monitoring, or third-party tools.
- Experience building a QA automation framework from scratch rather than only using an existing one.
- Performance and load testing of data pipelines.
Benefits:
- Flexible schedule
- Paid sick time
- Provident Fund
Work Location: In person