SENIOR DATA ENGINEER
Lakehouse engineering, dimensional modelling and data quality — the senior technical voice on the engagement, guiding junior engineers and working directly with clients.
Experience: 3–4 years, hands-on data engineering
Location: Remote (India) — HQ: Pune
Reports into: Engagement / Delivery Lead
Progression: Technical leadership scope grows with demonstrated competency
THE ROLE
This is the senior technical seat on a lakehouse build, and it stays hands-on. Routine construction — standard ingestion, repeatable loads, straightforward transformations — is carried by junior engineers working to your design and your review. Your value is in the work they cannot yet do: choosing the model, finding the defect nobody has reported, and raising the problem while it is still cheap to fix. We are looking for someone who reads a new dataset and can tell, early, where it will break the pipeline six weeks from now — and who says so, with evidence, before it becomes a client escalation.
WHAT YOU OWN
- Getting ahead of problems. Profile new sources before build starts; identify grain, key, hierarchy and format defects up front; quantify the impact and put a decision in front of the client or vendor early rather than absorbing it silently downstream.
- Modelling decisions. Declare and enforce fact grain. Composite natural keys, surrogate keys, SCD1/SCD2, conformed and degenerate dimensions, hierarchy validation, source-to-target mappings that are defensible.
- The hard transformations. Incremental and idempotent Silver/Gold logic — MERGE upserts, CDC and watermarking, deterministic deduplication, explicit date parsing, DECIMAL for monetary values, multi-currency with rate and rate-date retained.
- Quality architecture. Design the gates: primary-key uniqueness, referential integrity, containment and ratio bounds, outlier detection, volume alerts, raw-versus-curated reconciliation and quarantine handling.
- Performance and cost. Partitioning, file sizing, skew and join strategy, cluster and job tuning.
- Client conversations. Run technical working sessions, present options and findings, raise and track open items with the client and their vendors.
WHAT YOU DESIGN, DIRECT AND REVIEW
- Ingestion build. Junior engineers implement landing-to-Bronze against the schema contract, typing rules, quarantine behaviour and lineage standard you define.
- Repeatable pipeline work. Standard extracts, config-driven loads, reference-data refreshes and routine orchestration in ADF and Databricks Workflows.
- Test and validation coverage. Extending the quality framework you designed to new tables and sources.
- Their development. Scope work so they can both deliver and learn; pair on the parts that stretch them; review every pull request with reasoning rather than verdicts.
- Standards and reuse. Set coding, naming and testing standards and hold them in review; turn recurring fixes into reusable components so the same problem is not solved twice.
- Knowledge retention. Keep design notes, runbooks and onboarding material current; run short internal sessions on Spark behaviour, modelling decisions and debugging technique.
PLATFORM OUTLOOK
We hire for core data engineering ability rather than for a single vendor. Databricks is mandatory today because it is what our current delivery runs on, but the landscape moves — Snowflake or another engine may be the answer on the next engagement. We invest in training and give people time to learn; what we ask in return is genuine willingness to pick up a new tool and reach working proficiency quickly. Depth in one product without the underlying principles will not suit this role.
ESSENTIAL SKILLS AND EXPERIENCE
Education
B.E. / B.Tech (or M.Tech) in Computer Science, Information Technology or a related engineering discipline, from a premium engineering institution — IIT, NIT, BITS, IIIT or an institute of comparable standing. This is a firm requirement for this role.
Databricks
Live production build experience. Delta Lake — transaction log, time travel, schema evolution and enforcement, OPTIMIZE / Z-ORDER or liquid clustering, VACUUM. Medallion architecture, Workflows, job versus all-purpose compute, Auto Loader, Unity Catalog, Delta Live Tables and expectations.
Data engineering foundations
Advanced SQL and strong Python beyond notebooks — modules, packaging, tests. Batch pipeline design: idempotency, incremental versus full loads, watermarking, replay and backfill, dependency and failure handling. Warehousing fundamentals that transfer across platforms.
Spark engineering
Advanced PySpark and Spark SQL; window functions; join strategy and broadcast control; skew and small-file handling; partitioning and pruning; adaptive query execution; use of the Spark UI and query plans to diagnose performance.
Data modelling
Dimensional modelling in production: star schema, grain declaration, fact types, SCD1/SCD2, natural versus surrogate keys, conformed dimensions, late-arriving dimensions. Able to prove a candidate key holds and defend a model decision with evidence.
Diagnostic instinct
Profiling and root-cause analysis on imperfect source data — fill rates, cardinality, duplicates, hierarchy breaches, format inconsistencies. Distinguishes a source defect from a pipeline defect, and converts either into a control or a documented request to the provider.
Azure and orchestration
Azure Data Factory — parameterised pipelines, Copy activity, ForEach and lookup patterns, triggers, integration runtimes; ADLS Gen2 zone and partition layout; Key Vault; file-based and API-based ingestion.
DevOps
Azure DevOps or GitHub — branching strategy, pull-request gates, YAML CI/CD, dev/test/prod promotion; unit and data testing (pytest, chispa, Great Expectations or equivalent); logging and observability; Databricks Asset Bundles or Terraform.
Client-facing ability
Wants client exposure, not shielding from it. Runs a working session, presents a recommendation to a technical or business audience, and holds a position with evidence under challenge. Concise written communication — findings, options and open items a client can act on.
Developing others
Record of bringing junior engineers up to standard through onboarding, pairing and structured code review. Explains reasoning rather than issuing instructions, and treats the team's output as their own responsibility.
Adaptability
Evidence of learning a new platform, language or toolset to working depth and applying it on live delivery within a supported ramp-up.
Pay: ₹1,200,000.00 - ₹1,800,000.00 per year
Work Location: Hybrid remote in Pune, Maharashtra (Pune, Pune District)