Experience: 5+ Years
Location: Remote
Employment Type: Full-Time
Job Summary
We are looking for an experienced Python Data Engineer with strong expertise in PySpark, Apache Airflow, and Python-based API frameworks (FastAPI/Django/Flask). The ideal candidate will be responsible for designing, developing, and maintaining scalable data pipelines, distributed data processing solutions, workflow orchestration, and backend services to support business-critical applications and analytics platforms.
Key Responsibilities
- Design, develop, and optimize large-scale data pipelines using PySpark.
- Build and manage ETL/ELT workflows for processing structured and unstructured data.
- Develop RESTful APIs and microservices using FastAPI, Django, or Flask.
- Create, schedule, monitor, and troubleshoot data workflows using Apache Airflow.
- Work with large datasets in distributed computing environments.
- Integrate data from multiple sources, including databases, APIs, and cloud services.
- Implement data quality, validation, and monitoring frameworks.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Collaborate with Data Scientists, Analysts, and Product teams to deliver data solutions.
- Develop CI/CD pipelines and automate deployment processes.
- Ensure security, reliability, and maintainability of data platforms.
Required SkillsCore Technologies
- Strong proficiency in Python (5+ years).
- Hands-on experience with PySpark and Apache Spark.
- Experience with Apache Airflow for workflow orchestration.
- Strong knowledge of at least one Python web framework:
- FastAPI (Preferred)
- Django
- Flask
Databases
- SQL: PostgreSQL, MySQL, SQL Server.
- NoSQL: MongoDB, Cassandra, DynamoDB (good to have).
Cloud & Big Data
- Experience with AWS, Azure, or GCP.
- Knowledge of distributed systems and big data technologies.
- Experience with Databricks, EMR, or Synapse is a plus.
DevOps & Tools
- Git, GitHub/GitLab.
- Docker and Kubernetes.
- CI/CD tools such as Jenkins, Azure DevOps, or GitHub Actions.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering, or related field.
- 5+ years of experience in Python development.
- 3+ years of hands-on experience with PySpark and large-scale data processing.
- Strong understanding of REST API development.
- Experience working in Agile/Scrum environments.
- Excellent problem-solving and communication skills.
Preferred Skills
- Experience with Kafka or other streaming technologies.
- Knowledge of Delta Lake, Iceberg, or Lakehouse architecture.
- Experience with data warehousing concepts and tools.
- Familiarity with ML pipelines and MLOps practices.
- Cloud certifications (AWS/Azure/GCP).
Pay: ₹1,000,000.00 - ₹2,200,000.00 per year
Benefits:
Application Question(s):
- What is your ctc and ectc?
- What is your notice period?
- What is your current location?
Experience:
- total: 4 years (Preferred)
- Python: 4 years (Preferred)
- Django: 4 years (Preferred)
- Pyspark: 4 years (Preferred)
- Airflow: 4 years (Preferred)
Work Location: Remote