We are looking for a highly skilled Data Engineer with 4–6 years of experience in designing and developing large-scale data processing systems, real-time streaming solutions, and distributed data platforms. The ideal candidate should possess strong expertise in Python, PySpark, Apache Kafka, Redis, REST APIs, and Distributed File Systems. The role involves building scalable, high-performance data pipelines that support both batch and real-time analytics workloads.
- Design, develop, and maintain robust data pipelines using Python and PySpark.
- Build and manage real-time data streaming applications using Apache Kafka.
- Develop and integrate RESTful APIs for data ingestion, processing, and retrieval.
- Implement caching and high-speed data access solutions using Redis.
- Work with distributed file systems and large-scale storage frameworks to manage structured and unstructured data.
- Design ETL/ELT processes for ingesting, transforming, and loading data from multiple sources.
- Optimize data processing jobs for performance, scalability, and reliability.
- Monitor production data pipelines and resolve performance bottlenecks.
- Collaborate with Data Architects, Data Scientists, Analysts, and Business teams to deliver data-driven solutions.
- Ensure data quality, security, governance, and compliance standards are met.
Programming & Data Processing
- Strong hands-on experience in Python.
- Advanced knowledge of PySpark.
- Experience in distributed data processing and optimization.
- Strong understanding of data structures and algorithms.
- Strong experience in designing and consuming REST APIs.
- API authentication (OAuth, JWT, token-based authentication).
- Data exchange using JSON/XML.
- API performance monitoring and troubleshooting.
Databases
- Strong SQL skills.
- Query optimization.
- Relational databases (PostgreSQL, MySQL, SQL Server).
- Exposure to NoSQL databases is an advantage.