Job Title: Data Modeler – Databricks Delta Lake / Lakebase & Oracle Cloud Migration
Experience: 8–12 Years
Employment Type: Full-Time
Location: Hybrid
Department: Data Engineering & Cloud Modernization
Role Summary
We are seeking an experienced Data Modeler with strong expertise in Databricks Delta Lake, Databricks Lakebase, Oracle database migration, and cloud-native data platforms. The ideal candidate will be responsible for designing scalable enterprise data models, leading Oracle-to-cloud migration initiatives, and building modern Lakehouse architectures that support analytics, business intelligence, and AI/ML workloads.
The successful candidate should possess hands-on experience in Databricks Unity Catalog, Delta Sharing, cloud migration, metadata governance, ETL/ELT pipelines, and enterprise data modeling, along with a deep understanding of cloud platforms such as Azure, AWS, or GCP.
Key ResponsibilitiesData Modeling & Architecture
- Design and develop enterprise logical, conceptual, and physical data models for modern Lakehouse architectures.
- Build scalable data models supporting analytical, transactional, and AI/ML workloads.
- Design dimensional, normalized, and hybrid data models for enterprise applications.
- Develop data models optimized for Delta Lake, Lakebase, and cloud-native databases.
- Define enterprise data standards, naming conventions, and modeling best practices.
Oracle Database Migration
- Lead migration of enterprise workloads from Oracle to modern cloud platforms including:
- Databricks Delta Lake
- Databricks Lakebase
- Snowflake
- Azure SQL Database
- Azure Synapse Analytics
- PostgreSQL
- MongoDB and other modern databases
- Perform schema conversion, dependency analysis, and data validation.
- Develop migration strategies with minimal downtime.
- Optimize migrated databases for performance, scalability, and cost efficiency.
- Support data reconciliation and migration testing.
Databricks Lakehouse Development
- Design enterprise Lakehouse architectures using Databricks Delta Lake.
- Implement Delta Lake best practices for ACID transactions, versioning, and optimization.
- Configure and manage Databricks Unity Catalog.
- Implement Delta Sharing for secure enterprise data sharing.
- Develop secure and governed data access models.
- Support Databricks Lakebase implementations for transactional workloads.
Cloud Data Engineering
- Design cloud-native data platforms on Azure, AWS, or GCP.
- Develop scalable data ingestion and transformation pipelines.
- Collaborate with cloud engineering teams to modernize enterprise data platforms.
- Optimize cloud storage, compute resources, and data processing performance.
ETL / ELT Development
- Build and optimize ETL/ELT pipelines using:
- Azure Data Factory (ADF)
- dbt
- Informatica
- Fivetran
- Other enterprise integration tools
- Design data transformation workflows.
- Automate data ingestion and orchestration.
- Ensure data quality throughout ingestion and transformation processes.
Data Governance & Metadata Management
- Implement enterprise data governance frameworks.
- Manage metadata using:
- Microsoft Purview
- Collibra
- Alation
- Informatica Enterprise Data Catalog (EDC)
- Define business glossaries and metadata standards.
- Support data lineage, cataloging, and compliance initiatives.
- Ensure regulatory compliance and secure data access.
Data Quality & Master Data Management
- Design and implement data quality validation processes.
- Develop data profiling and cleansing strategies.
- Support Master Data Management (MDM) initiatives.
- Implement automated data quality monitoring and reporting.
- Ensure consistency and accuracy across enterprise data assets.
AI/ML Data Enablement
- Design data models optimized for AI and Machine Learning workloads.
- Build feature-ready datasets for predictive analytics.
- Support feature engineering pipelines.
- Optimize data structures for GenAI and advanced analytics.
- Collaborate with Data Scientists and ML Engineers to deliver AI-ready data products.
Performance Optimization
- Optimize Delta Lake performance through partitioning, indexing, and file optimization.
- Improve SQL query performance and storage efficiency.
- Tune cloud data pipelines for scalability and reliability.
- Monitor and resolve performance bottlenecks.
Collaboration
- Work closely with:
- Enterprise Architects
- Data Engineers
- Data Scientists
- Cloud Engineers
- Business Analysts
- DevOps Teams
- Application Development Teams
- Participate in Agile ceremonies including Sprint Planning, Daily Stand-ups, Sprint Reviews, and Retrospectives.
- Produce technical documentation for data models, migration strategies, and governance standards.
Required Technical SkillsData Modeling
- Conceptual Data Modeling
- Logical Data Modeling
- Physical Data Modeling
- Dimensional Modeling
- Star Schema
- Snowflake Schema
- Data Architecture
- Data Normalization
Databricks
- Databricks Delta Lake
- Databricks Lakebase
- Unity Catalog
- Delta Sharing
- Databricks SQL
- Lakehouse Architecture
Database Technologies
- Oracle Database
- PostgreSQL
- Azure SQL Database
- Snowflake
- Azure Synapse Analytics
- MongoDB
Cloud Platforms
- Microsoft Azure
- Amazon Web Services (AWS)
- Google Cloud Platform (GCP)
ETL / ELT Tools
- Azure Data Factory (ADF)
- dbt
- Informatica
- Fivetran
- ETL/ELT Pipeline Development
Data Governance
- Microsoft Purview
- Collibra
- Alation
- Informatica Enterprise Data Catalog (EDC)
- Metadata Management
- Data Lineage
- Master Data Management (MDM)
- Data Quality
Programming & Query Languages
- SQL
- Python (Preferred)
- PySpark
- Spark SQL
DevOps & Version Control
- Git
- Azure DevOps
- CI/CD Pipelines
- Infrastructure as Code (Preferred)
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, Data Science, or a related field.
- 8–12 years of experience in data modeling, database architecture, or cloud data engineering.
- Hands-on experience with Databricks Delta Lake and enterprise Lakehouse implementations.
- Strong experience migrating Oracle databases to cloud-native platforms.
- Expertise in SQL and enterprise data modeling techniques.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Experience implementing enterprise metadata management and governance frameworks.
- Strong understanding of ETL/ELT processes and cloud data integration.
- Experience working in Agile software development environments.
Preferred Qualifications
- Databricks Certified Data Engineer Associate or Professional.
- Microsoft Certified: Azure Data Engineer Associate (DP-203).
- Snowflake SnowPro Certification.
- Experience with Oracle-to-Databricks modernization programs.
- Experience in healthcare, financial services, insurance, or other highly regulated industries.
- Knowledge of AI/ML feature engineering and modern data products.
- Experience with enterprise Data Mesh or Data Fabric architectures.
- Familiarity with Infrastructure as Code (Terraform or Bicep).
Soft Skills
- Strong analytical and problem-solving abilities.
- Excellent communication and stakeholder management skills.
- Ability to work independently and collaboratively in cross-functional teams.
- Strong attention to detail and data quality.
- Ownership mindset with a focus on scalable, secure, and high-quality data solutions.
- Ability to manage multiple priorities in a fast-paced Agile environment.
- Passion for cloud modernization, data governance, and emerging AI technologies.
- Continuous learning mindset with the ability to adapt to evolving cloud and data engineering practices.
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)