Role Summary
We are looking for a hands-on Data Engineer with strong PySpark and Python expertise to build and maintain data marts and ETL pipelines within a banking environment. The ideal candidate owns the full SDLC, from build through UAT, bug fixing, production deployment, and postproduction support.
Key Responsibilities
-
Design, build, and maintain ETL pipelines and data marts using PySpark and Python
-
Write clean, maintainable, and robust production-grade code
-
Own end to end SDLC activities: build, UAT, UAT bug fixes, production deployment, and postproduction support
-
Perform Oracle query analysis and PySpark code debugging
-
Work across structured, semi structured, and unstructured data sources
-
Apply software engineering best practices to production pipelines
-
Support CI/CD processes and data testing/validation activities
Required Skills & Experience
-
5+ years commercial experience in a data-driven role
-
Hands-on experience building data marts and ETL pipelines
-
Expert level Python for ETL scripting
-
Strong PySpark experience
-
Analytical expertise in Oracle SQL and data analysis
-
Understanding of software engineering concepts and best practices for production pipelines
-
Banking client or banking domain knowledge
-
Strong Data Warehousing fundamentals
-
Familiarity with query languages and both SQL and NoSQL database technologies
-
CI/CD exposure, including testing and validation of data pipelines
Daily Tech Stack
-
Python
-
Spark / PySpark
-
Jupyter
-
SQL and NoSQL DBMS
-
Hadoop / MapReduce / Hive
-
Pandas