Job Title: AI Data Engineer
Experience: 5–8 Years
Employment Type: Full-Time
Position Summary
We are seeking an experienced AI Data Engineer to build and maintain scalable data pipelines that power AI/ML and Generative AI solutions. The ideal candidate will have strong expertise in Apache Spark (Python/PySpark), SQL, Azure Cloud, Airflow, Kafka, and processing unstructured data. This role involves designing modern data architectures, building ETL/ELT pipelines, managing large-scale structured and unstructured datasets, and enabling AI/ML teams with high-quality, reliable data.
Key Responsibilities
- Design, develop, and optimize scalable data pipelines for AI/ML and Generative AI applications.
- Build batch and real-time data processing solutions using Apache Spark (PySpark) and Kafka.
- Develop ETL/ELT workflows to ingest, transform, and process data from multiple sources.
- Process and manage structured, semi-structured, and unstructured data, including JSON, logs, documents, PDFs, and text data.
- Develop workflows using Apache Airflow for scheduling, orchestration, and monitoring.
- Design and optimize data models for analytics and AI workloads.
- Work with Azure cloud services to build secure, scalable, and high-performance data platforms.
- Collaborate with Data Scientists, ML Engineers, and AI teams to prepare feature-ready datasets.
- Optimize Spark jobs for performance, scalability, and cost efficiency.
- Implement data quality checks, validation, and monitoring processes.
- Ensure data governance, security, and compliance across data platforms.
- Troubleshoot production issues and optimize existing data pipelines.
- Maintain technical documentation, data lineage, and architecture diagrams.
- Stay current with emerging technologies in AI data engineering and big data ecosystems.
Required Technical Skills
- Strong expertise in Python and PySpark.
- Advanced SQL development and query optimization.
- Apache Spark for distributed data processing.
- Azure Cloud services (Azure Data Lake Storage, Azure Databricks, Azure Data Factory, Azure Synapse Analytics, Azure Blob Storage).
- Apache Airflow for workflow orchestration.
- Apache Kafka for real-time data streaming.
- Experience handling unstructured and semi-structured data (JSON, XML, logs, PDFs, text, documents).
- Data ingestion, transformation, and pipeline development.
- Experience with Delta Lake and Parquet file formats.
- Git and version control.
- CI/CD pipelines for data engineering workflows.
- Linux/Unix command-line experience.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or a related field.
- 5–8 years of experience in Data Engineering or Big Data development.
- Hands-on experience building enterprise-scale data pipelines.
- Experience working with cloud-native data platforms, preferably Microsoft Azure.
- Strong understanding of distributed computing and big data architecture.
- Experience working in Agile/Scrum environments.
Preferred Skills
- Experience with Azure Databricks.
- Knowledge of Generative AI, LLMs, and Retrieval-Augmented Generation (RAG) data pipelines.
- Experience with Vector Databases such as Pinecone, Milvus, Azure AI Search, or FAISS.
- Familiarity with Azure Machine Learning and MLflow.
- Experience with data lakehouse architecture.
- Knowledge of data cataloging, metadata management, and data governance.
- Experience with Docker and Kubernetes.
- Exposure to Infrastructure as Code (Terraform or ARM/Bicep).
- Experience with monitoring tools such as Azure Monitor, Prometheus, or Grafana.
Preferred Certifications
- Microsoft Certified: Azure Data Engineer Associate (DP-203)
- Microsoft Azure Fundamentals (AZ-900)
- Databricks Certified Data Engineer Associate or Professional
- Apache Spark Certification (Preferred)
Key Competencies
- AI Data Engineering
- Apache Spark (PySpark)
- SQL Optimization
- Azure Cloud Data Services
- Workflow Orchestration (Airflow)
- Real-Time Data Streaming (Kafka)
- Unstructured Data Processing
- ETL/ELT Pipeline Development
- Data Modeling
- Performance Optimization
- Problem Solving
- Collaboration and Communication
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)