Job Description-
- Perform data analysis to understand source systems and business requirements
- Translate requirements into scalable data models and pipeline designs
- Build and maintain data ingestion pipelines from multiple sources (batch and incremental/delta loads) into Microsoft Fabric
- Experience in writing scripts from multiple sources including database tables, flat files and APIs
- Ingest and land raw data into bronze layer, ensuring reliability, scalability, and traceability
- Transform raw data into clean, standardised silver datasets (deduplication, conformance, validation)
- Design and develop gold layer models (fact/dimension tables, aggregates) for analytics and reporting
- Design end-to-end ETL/ELT pipelines across bronze, silver, and gold layers
- Support creation of semantic models and reporting-ready datasets
- Implement data quality rules and validation checks (completeness, accuracy, freshness, reconciliation)
- Implement robust error handling mechanisms including retry logic, logging, exception handling, and alerting
- Design self-healing pipelines (auto-retry, checkpointing, idempotent processing, and failure recovery)
- Optimise pipeline performance and ensure reliable, production-grade data processing
- Monitor pipelines and proactively identify and resolve failures or performance bottlenecks
- Work with CI/CD and DevOps practices for deployment across environments (Dev/Test/Prod)
- Able to write, amend and debug scripts using both PySpark and SQL
- Have a good working knowledge of software design standards
- Collaborate with business, BI, and architecture teams to deliver end-to-end data solutions.
Pay: ₹130,000.00 - ₹160,000.00 per month
Work Location: Remote