Job Title: AI Data Engineer
Experience:5–8 Years
Employment Type: Full-Time
Position Summary
We are seeking an experienced **AI Data Engineer** to build and maintain scalable data pipelines that power AI/ML and Generative AI solutions. The ideal candidate will have strong expertise in **Apache Spark (Python/PySpark), SQL, Azure Cloud, Airflow, Kafka, and processing unstructured data**. This role involves designing modern data architectures, building ETL/ELT pipelines, managing large-scale structured and unstructured datasets, and enabling AI/ML teams with high-quality, reliable data.
Key Responsibilities
* Design, develop, and optimize scalable data pipelines for AI/ML and Generative AI applications.
* Build batch and real-time data processing solutions using Apache Spark (PySpark) and Kafka.
* Develop ETL/ELT workflows to ingest, transform, and process data from multiple sources.
* Process and manage structured, semi-structured, and unstructured data, including JSON, logs, documents, PDFs, and text data.
* Develop workflows using Apache Airflow for scheduling, orchestration, and monitoring.
* Design and optimize data models for analytics and AI workloads.
* Work with Azure cloud services to build secure, scalable, and high-performance data platforms.
* Collaborate with Data Scientists, ML Engineers, and AI teams to prepare feature-ready datasets.
* Optimize Spark jobs for performance, scalability, and cost efficiency.
* Implement data quality checks, validation, and monitoring processes.
* Ensure data governance, security, and compliance across data platforms.
* Troubleshoot production issues and optimize existing data pipelines.
* Maintain technical documentation, data lineage, and architecture diagrams.
* Stay current with emerging technologies in AI data engineering and big data ecosystems.
Required Technical Skills
* Strong expertise in Python and PySpark.
* Advanced SQL development and query optimization.
* Apache Spark for distributed data processing.
* Azure Cloud services (Azure Data Lake Storage, Azure Databricks, Azure Data Factory, Azure Synapse Analytics, Azure Blob Storage).
* Apache Airflow for workflow orchestration.
* Apache Kafka for real-time data streaming.
* Experience handling unstructured and semi-structured data (JSON, XML, logs, PDFs, text, documents).
* Data ingestion, transformation, and pipeline development.
* Experience with Delta Lake and Parquet file formats.
* Git and version control.
* CI/CD pipelines for data engineering workflows.
* Linux/Unix command-line experience.
### Required Qualifications
* Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or a related field.
* 5–8 years of experience in Data Engineering or Big Data development.
* Hands-on experience building enterprise-scale data pipelines.
* Experience working with cloud-native data platforms, preferably Microsoft Azure.
* Strong understanding of distributed computing and big data architecture.
* Experience working in Agile/Scrum environments.
Preferred Skills
* Experience with Azure Databricks.
* Knowledge of Generative AI, LLMs, and Retrieval-Augmented Generation (RAG) data pipelines.
* Experience with Vector Databases such as Pinecone, Milvus, Azure AI Search, or FAISS.
* Familiarity with Azure Machine Learning and MLflow.
* Experience with data lakehouse architecture.
* Knowledge of data cataloging, metadata management, and data governance.
* Experience with Docker and Kubernetes.
* Exposure to Infrastructure as Code (Terraform or ARM/Bicep).
* Experience with monitoring tools such as Azure Monitor, Prometheus, or Grafana.
Preferred Certifications
* Microsoft Certified: Azure Data Engineer Associate (DP-203)
* Microsoft Azure Fundamentals (AZ-900)
* Databricks Certified Data Engineer Associate or Professional
* Apache Spark Certification (Preferred)
Key Competencies
* AI Data Engineering
* Apache Spark (PySpark)
* SQL Optimization
* Azure Cloud Data Services
* Workflow Orchestration (Airflow)
* Real-Time Data Streaming (Kafka)
* Unstructured Data Processing
* ETL/ELT Pipeline Development
* Data Modeling
* Performance Optimization
* Problem Solving
* Collaboration and Communication
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)