Project Role : Data Engineer
Project Role Description : Design, develop and maintain data solutions for data generation, collection, and processing. Create data pipelines, ensure data quality, and implement ETL (extract, transform and load) processes to migrate and deploy data across systems.
Must have skills : Data Engineering
Good to have skills : NA
Minimum 7.5 year(s) of experience is required
Educational Qualification : 15 years full time education
Data Engineering Lead – Azure Data Platform, Lakehouse
Experience: 12+ years overall, including 5+ years leading data engineering teams or platform initiatives
Location: India
Role Summary
A hands-on technical lead who still builds: personally designing and writing core ETL/ELT pipelines and graph data models — both property-graph and RDF/semantic-graph — for the highest-priority integrations, while setting the lakehouse, medallion, data mesh, and catalog standards the rest of the team builds against. Owns the enterprise data platform architecture for an organization integrating data from many source systems, and leads by pairing and debugging alongside the team, not only through design sign-off.
Key Responsibilities
Personally design, build, and maintain ETL/ELT pipelines and transformation logic for the highest-priority and most complex data integrations — extract, transform (PySpark/Spark SQL), and load into bronze/silver/gold lakehouse tables
Write and review transformation logic for incremental loads, change data capture (CDC), deduplication, slowly changing dimensions (SCD), and schema evolution across the team's pipelines
Write and review Gremlin graph models and traversal queries in Azure Cosmos DB for Apache Gremlin for knowledge-graph and relationship-heavy use cases
Write and review SPARQL queries and Turtle (.ttl) ontology/data files for RDF-based knowledge graph and semantic layer work, using a triplestore such as Graphwise GraphDB
Set enterprise data architecture standards — lakehouse design, medallion layering conventions, incremental-load/CDC patterns, and data-store selection (relational, NoSQL, graph) — by building the reference pipelines the team follows
Own platform strategy and hands-on build-out for a modern Azure data platform (e.g., Microsoft Fabric) — unified lake storage, Data Factory/Dataflow Gen2 pipelines, Spark-based data engineering, and real-time data flows
Own the enterprise data catalog and governance strategy (e.g., Microsoft Purview) — lineage, classification, glossary, and stewardship model
Lead, mentor, and grow a team of data engineers through pairing, code review, and hands-on pipeline debugging — not only design sign-off
Partner with enterprise/solution architects, AI engineering leads, and business stakeholders to translate data strategy into a delivery roadmap
Own data security, access control, and compliance posture — RBAC, classification/sensitivity labeling, and audit — across the data platform
Required Skills & Experience
12+ years in data engineering/data platform roles, including 5+ years leading data engineering teams — while remaining hands-on and writing production code
Deep, current hands-on experience building ETL/ELT pipelines — incremental loads, CDC, deduplication, SCD, and schema evolution — on a modern Azure data platform (e.g., Microsoft Fabric, or Azure Synapse/Databricks-equivalent)
Strong Python and Spark (PySpark) for transformation development strong SQL for analytical and transactional workloads
Proven track record building and operating medallion (bronze/silver/gold) lakehouse pipelines at enterprise scale, not only designing them on paper
Hands-on experience writing graph data models and Gremlin queries in Azure Cosmos DB for Apache Gremlin
Hands-on experience with RDF/semantic graph technologies — writing SPARQL queries and authoring Turtle (.ttl) files, using a triplestore such as Graphwise GraphDB (or equivalent, e.g., Amazon Neptune, Stardog)
Experience designing and rolling out data mesh operating models — domain ownership, data-as-a-product, self-serve platforms, and federated governance
Experience owning enterprise data catalog / governance tooling (e.g., Microsoft Purview) — data map, unified catalog, lineage, stewardship
Track record integrating and rationalizing data from many heterogeneous enterprise source systems (ERP, CRM, SaaS, APIs, legacy databases)
Experience mentoring engineers through hands-on pairing and code review, and presenting data architecture/strategy to senior stakeholders
Preferred
Exposure to the broader semantic web stack — RDF/RDFS, OWL, SHACL, SKOS — and ontology-driven knowledge graph / semantic layer architecture
Experience with pipeline orchestration tools (e.g., Apache Airflow, or Fabric's Airflow-based orchestration)
Experience across multiple major cloud data platforms (AWS, Azure, and GCP) — data lake, warehouse, and serverless/ETL services
Familiarity with vector stores and data preparation patterns for RAG/agentic AI consumption
Experience with real-time/streaming architectures (e.g., Eventstream, KQL, or equivalent)
Exposure to regulated-industry data governance (e.g., life sciences, healthcare, financial services)
15 years full time education