Job Title: AI Data Engineer
Experience: 5–8 Years
Employment Type: Full-Time
Position Summary
We are seeking an experienced AI Data Engineer to build and maintain scalable data platforms that power AI, Machine Learning, and Generative AI solutions. The ideal candidate will have strong expertise in Apache Spark (Python/PySpark), SQL, Azure Cloud, Airflow, Kafka, and processing unstructured data such as JSON, logs, documents, and PDFs. The role involves designing modern data architectures, building high-performance ETL/ELT pipelines, and delivering reliable data solutions for AI applications.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Apache Spark (PySpark) and SQL.
- Build batch and real-time data processing solutions on Azure Cloud.
- Develop ETL/ELT workflows for structured, semi-structured, and unstructured data.
- Process and transform JSON, application logs, PDFs, documents, and other unstructured data sources for AI/ML applications.
- Develop and manage workflow orchestration using Apache Airflow.
- Build streaming data pipelines using Apache Kafka.
- Optimize data ingestion, transformation, and storage for performance and scalability.
- Design data models and data lakes to support AI and analytics workloads.
- Integrate data from multiple enterprise systems, APIs, cloud storage, and databases.
- Ensure data quality, governance, security, and compliance across data platforms.
- Collaborate with AI/ML Engineers and Data Scientists to prepare feature-ready datasets.
- Monitor, troubleshoot, and optimize data pipelines in production environments.
- Automate data workflows using CI/CD and Infrastructure as Code (IaC) best practices.
- Document data architecture, pipeline designs, and operational procedures.
Required Skills
Technical Skills (Must Have)
- Apache Spark
- Python / PySpark
- SQL
- Azure Cloud
- Azure Data Lake Storage (ADLS)
- Azure Data Factory (ADF)
- Azure Databricks
- Apache Airflow
- Apache Kafka
- ETL/ELT Development
- Data Pipeline Development
- Data Warehousing
- Git
Data Engineering Expertise
- Batch and Streaming Data Processing
- Data Modeling
- Data Lake Architecture
- Data Transformation
- Performance Optimization
- Data Validation
- Data Quality Management
Unstructured Data Processing
- JSON
- Application Logs
- Documents (Word, Excel, Text)
- PDFs
- XML
- CSV
- API Data Integration
Good to Have
- Azure Synapse Analytics
- Delta Lake
- MLflow
- Azure Machine Learning
- Docker
- Kubernetes
- Terraform
- REST APIs
- Generative AI and LLM data pipelines
- Vector Databases (Azure AI Search, Pinecone, FAISS, ChromaDB)
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Data Engineering, or a related field.
- 5–8 years of experience in Data Engineering.
- Strong hands-on experience with Apache Spark (PySpark), SQL, and Azure Cloud services.
- Experience building batch and streaming data pipelines using Airflow and Kafka.
- Expertise in processing structured, semi-structured, and unstructured data.
- Strong understanding of cloud data architecture, distributed computing, and performance optimization.
- Excellent analytical, problem-solving, and communication skills.
Preferred Qualifications
- Experience supporting AI/ML and Generative AI workloads.
- Knowledge of Retrieval-Augmented Generation (RAG) data pipelines and document ingestion frameworks.
- Experience with Azure AI services, Azure Machine Learning, or Azure AI Search.
- Familiarity with DevOps, CI/CD pipelines, and Infrastructure as Code (Terraform/Bicep).
- Microsoft Azure Data Engineer Associate (DP-203) or equivalent cloud certification is preferred.
What We Offer
- Opportunity to build enterprise-scale AI and data engineering platforms.
- Exposure to modern Azure, big data, and Generative AI technologies.
- Collaborative and innovation-driven work environment.
- Career growth through challenging cloud and AI transformation projects.
- Competitive compensation and comprehensive employee benefits.
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)