Job Summary
We are looking for a skilled Data Engineer with 3–5 years of experience in building scalable data pipelines and integration solutions. The ideal candidate should have strong hands-on expertise in Databricks (must-have), Python, and API-based data integration.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Databricks and Python.
- Build and optimize data workflows using Apache Spark within the Databricks environment.
- Integrate data from internal systems and third-party platforms through REST APIs and other integration mechanisms.
- Develop reusable data ingestion frameworks and automation scripts.
- Perform data transformation, cleansing, validation, and enrichment for analytics and reporting.
- Work with structured and semi-structured data sources such as JSON, CSV, APIs, databases, and cloud storage.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Collaborate with Data Analysts, Data Scientists, and application teams to deliver reliable datasets.
- Implement monitoring, logging, and error-handling mechanisms for production pipelines.
- Participate in code reviews, testing, and deployment activities following engineering best practices.
Required Skills & Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
- 3–5 years of experience in Data Engineering or Big Data development.
- Strong hands-on experience with Databricks (mandatory).
- Proficiency in Python for data processing and automation.
- Experience with REST APIs, API authentication, and data integration.
- Strong understanding of Apache Spark and distributed data processing.
- Experience writing complex SQL queries and working with relational databases.
- Knowledge of data lake/lakehouse concepts and Delta Lake.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Familiarity with Git, CI/CD pipelines, and Agile development practices.
Preferred Skills
- Experience with Azure Databricks and Azure Data Factory.
- Knowledge of orchestration tools such as Airflow or Databricks Workflows.
- Exposure to streaming technologies such as Kafka.
- Understanding of data governance, security, and access control.
Key Technical Stack
- Databricks (Must Have)
- Python
- Apache Spark / PySpark
- REST APIs
- SQL
- Delta Lake
- Azure / AWS / GCP
- Git
- Airflow / Databricks Workflows
Pay: ₹1,000,000.00 - ₹1,400,000.00 per year
Benefits:
Application Question(s):
- What is your Current CTC?
- What is your Expected CTC?
- How soon can you join if selected?
Work Location: Remote