Senior Data Engineer – GenAI & Unstructured Data Pipelines
Job Type: Full-time
Experience: 6–8 Years
Location: Offshore (Remote)
Job Summary
We are seeking an experienced Senior Data Engineer – GenAI & Unstructured Data Pipelines to build next-generation AI data platforms that power Large Language Model (LLM) applications. The ideal candidate will have strong expertise in designing scalable data pipelines for unstructured and multi-modal data while enabling Retrieval-Augmented Generation (RAG), embeddings, vector search, and AI copilots.
This role is ideal for professionals passionate about modern data engineering, cloud-native architectures, and Generative AI technologies.
Key Responsibilities
- Design, build, and maintain scalable data pipelines for structured, unstructured, and multi-modal data.
- Own and support the end-to-end machine learning lifecycle, including data ingestion, feature engineering, model training, evaluation, deployment, monitoring, retraining, and rollback.
- Develop production-grade ML pipelines using Azure native services with CI/CD automation and best practices.
- Utilize Azure Machine Learning and MLflow for experiment tracking, model registry, and governed deployments across environments.
- Design and implement Generative AI solutions using Azure OpenAI, embeddings, vector search, and Retrieval-Augmented Generation (RAG).
- Build Agentic AI workflows with multi-step reasoning, tool integration, observability, guardrails, reliability, and cost optimization.
- Develop scalable batch and streaming data pipelines using Azure Databricks, Apache Spark, and Kafka.
- Create efficient RAG pipelines, including document chunking, embedding generation, indexing, and retrieval workflows.
- Design and manage vector databases and search solutions such as Azure AI Search and Pinecone.
- Build high-performance batch and real-time data ingestion pipelines using Spark and Kafka.
Required Skills & Qualifications
- 6–8 years of experience in Data Engineering.
- Strong hands-on expertise with:
- Python
- PySpark
- SQL
- Apache Spark
- Apache Airflow
- Apache Kafka
- Experience processing and managing unstructured data, including JSON, logs, documents, and PDFs.
- Strong experience with cloud platforms, preferably Microsoft Azure.
- Hands-on exposure to Generative AI technologies, including:
- Large Language Models (LLMs)
- Retrieval-Augmented Generation (RAG)
- Embeddings
- Vector Databases
- Azure OpenAI
- Azure AI Search
- Pinecone
- Experience working with Azure Databricks and modern data engineering frameworks.
- Knowledge of CI/CD pipelines, MLflow, and Azure Machine Learning is highly preferred.
Preferred Skills
- Experience building AI copilots and Agentic AI applications.
- Understanding of data governance, monitoring, and production-grade ML operations.
- Excellent analytical, problem-solving, and communication skills.
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)