Senior Data Engineer
Experience: 6–10+ Years
Engagement: Freelance / Project-Based
Location: Remote
Role Overview
We are looking for an experienced Senior Data Engineer to design and develop highly scalable, fault-tolerant data pipelines for a large-scale logistics data platform. The role will focus on real-time ingestion, distributed data processing, cloud data platforms, legacy-system integration and building reliable data infrastructure capable of operating at massive scale.
Key Responsibilities
- Design and develop high-throughput batch and real-time data pipelines for large-scale data volumes.
- Build streaming pipelines for IoT telemetry, satellite feeds and enterprise data sources.
- Develop distributed data processing solutions using technologies such as Spark, Kafka, Flink or equivalent platforms.
- Design and implement data ingestion frameworks capable of handling heterogeneous and continuously changing data sources.
- Develop Change Data Capture (CDC) pipelines for synchronizing legacy systems with modern data platforms.
- Design data lake/lakehouse architectures using open table formats and support ACID transactions at scale.
- Implement mechanisms for automated schema evolution and schema-drift management without pipeline downtime.
- Build resilient orchestration workflows with failure recovery and automated rerouting.
- Integrate legacy databases/mainframes with modern cloud data platforms.
- Implement data quality, governance, lineage and observability across the data lifecycle.
- Collaborate with architects, DevOps/SecOps teams and analytics teams to deliver production-ready data infrastructure.
- Optimize pipeline performance, reliability and infrastructure utilization.
Required Skills
- Strong hands-on experience in Data Engineering and distributed data processing.
- Strong Python and SQL expertise.
- Experience with Apache Spark/PySpark and real-time streaming technologies such as Kafka, Flink or equivalent.
- Strong experience with cloud data platforms across AWS, Azure or GCP.
- Experience with Data Lakes/Lakehouse architectures.
- Strong understanding of ETL/ELT, CDC, data ingestion and data integration.
- Experience with schema evolution/schema management.
- Experience with Airflow or equivalent workflow orchestration platforms.
- Experience with open table formats such as Delta Lake, Apache Iceberg or Apache Hudi.
- Strong understanding of data quality, governance, lineage and observability.
- Experience designing fault-tolerant and highly available data pipelines.
Preferred
Experience with mainframe/legacy-system integration, IoT data, logistics/supply-chain platforms, satellite data or other high-volume real-time data environments.
Pay: ₹1,000,000.00 - ₹1,500,000.00 per year
Work Location: Remote