Designation – Data Engineer
Location: Bengaluru
About the role: We're hiring a Data Engineer in Bangalore to build the platform-level data infrastructure that powers FirstHive's Customer Data Platform across every client integration — connector frameworks, CDC ingestion pipelines, transformation services, and data quality tooling. You'll work in Java and Spring Boot on Kafka, Kafka Connect, MongoDB, and our analytical warehouse stack (StarRocks, Snowflake, BigQuery), deployed on multi-cloud Kubernetes.
The role is framework-building, not client work — you design once for the next ten clients. FirstHive ingests customer data from ERPs, CRMs, mobile apps, PoS systems, social, voice, and customer care across hundreds of enterprise deployments. The bottleneck is not data volume — it is the variety of source systems and the cost of onboarding each new client. We need someone who builds the systems that turn client onboarding into configuration, not engineering.
What you'll build
Organized by where it sits in the pipeline. Within each group, the most architectural work comes first.
Ingestion frameworks — the first thing every new client hits
- Pluggable connector framework over Kafka and Kafka Connect for databases, APIs, file feeds, and event streams — custom SMTs, DLQ patterns, reusable connector configurations.
- CDC ingestion pipelines from MongoDB and relational sources via Debezium and Kafka Connect — multi database routing, schema change handling, ordering guarantees.
- Automated schema mapping, detection, and inference tooling — so onboarding a new client is configuration, not engineering.
Transformation and quality — what makes the data trustworthy
- Transformation layer as composable Spring Boot modules: cleaning, deduplication, normalization, identity resolution, enrichment. Not glue scripts.
- Data quality framework — profiling, validation gates, anomaly detection, lineage tracking — wired into the pipeline so bad data is caught at ingestion, not at the dashboard.
- Schema evolution handling — backward compatibility across Kafka topics, transformations, and warehouse tables when client source schemas change.
Warehouse layer — where it lands and gets queried
- Data models for StarRocks, Snowflake, and BigQuery — partitioning, clustering / bucketing, materialization strategy, primary-key vs. duplicate vs. aggregate table design.
- Optimized SQL and stored procedures for mixed workloads: point lookups, high-concurrency customer profile dashboards, and large batch ETL.
- Metadata layer driving per-client schema definitions, mapping rules, and transformation logic — controlled by configuration, not code changes.
What we need
Grouped by where it matters. The first bullet of each group is the non-negotiable.
4+ years building data systems — not running them
You've designed and shipped framework-level data systems in production. You can point to ones still running.
Production-grade Java and Spring Boot
- Real microservices: error handling, observability, testing, lifecycle management. Not scripts.
- Framework-builder instinct — reusable tooling for the next ten clients, not the next ticket.
SQL fluency and data modeling depth
- Complex joins, window functions, CTEs (including recursive), and a real instinct for performance and cost.
- Star schema, SCD types, event sourcing, EAV patterns — and judgment on when each is the right answer.
Real depth on the streaming and warehouse stack
- Kafka and Kafka Connect at depth: connector configuration, custom transforms and converters, consumer group design, DLQ patterns, exactly-once vs.
- At-least-once tradeoffs. At least one analytical warehouse at architecture level — StarRocks, Snowflake, or BigQuery — covering data modeling, performance tuning, partitioning / clustering, and cost optimization.
- MongoDB or similar document store — schema design, compound indexing, change streams, CDC tradeoffs.
- Workflow orchestration in production — Airflow, Argo Workflows, dbt, or similar.
Bonus, not gating
These don't decide the hire, but they shape the shortlist:
- Debezium at production scale — buffer / lock tuning, multi-database capture, snapshot strategies.
- StarRocks, ClickHouse, Druid, or similar MPP / OLAP engines.
- Open table formats — Apache Iceberg, Hudi — and lakehouse architectures.
- Multi-cloud Kubernetes (GKE, EKS) and object storage (GCS, S3).
- Identity resolution — deterministic / probabilistic matching, graph-based stitching.
- CDP, MarTech, or AdTech domain exposure.
Pay: ₹1,800,000.00 - ₹3,500,000.00 per year
Work Location: In person