Job Title: Senior Data Engineer – GenAI & Unstructured Data Pipelines
Experience: 6–8 Years
Employment Type: Full-Time
Role Summary
We are seeking an experienced Senior Data Engineer to design and build next-generation GenAI-powered data platforms that enable Large Language Model (LLM) applications. The ideal candidate will have extensive expertise in developing scalable data pipelines for unstructured and multi-modal data, implementing Retrieval-Augmented Generation (RAG) architectures, embeddings, vector search, and AI copilot solutions. The role requires strong experience in Azure cloud technologies, modern data engineering frameworks, and production-grade machine learning pipelines.
Key Responsibilities
- Own the complete machine learning lifecycle including:
- Data ingestion
- Feature engineering
- Model training
- Evaluation
- Deployment
- Monitoring
- Retraining
- Rollback strategies
- Design, develop, and maintain production-grade ML pipelines using Azure-native services with robust CI/CD automation.
- Utilize Azure Machine Learning and MLflow for experiment tracking, model registry, and governed model promotion across Dev, Test, and Production environments.
- Design and implement Generative AI solutions leveraging:
- Azure OpenAI
- Embeddings
- Vector Search
- Retrieval-Augmented Generation (RAG)
- Build Agentic AI workflows with:
- Multi-step reasoning
- Tool integrations
- Guardrails
- Observability
- Reliability
- Cost optimization
- Develop scalable batch and streaming data pipelines using Azure Databricks.
- Build scalable pipelines for processing:
- Text
- Documents
- Logs
- PDFs
- Multi-modal datasets
- Design and implement RAG pipelines including:
- Data chunking
- Embedding generation
- Retrieval workflows
- Design, optimize, and manage vector search platforms such as Azure AI Search and Pinecone.
- Develop both batch and real-time data ingestion pipelines using Apache Spark and Kafka.
Required Technical SkillsProgramming
Big Data Technologies
- Apache Spark
- Apache Airflow
- Apache Kafka
Cloud Platforms
- Microsoft Azure (Preferred)
Azure Services
- Azure Databricks
- Azure Machine Learning
- Azure OpenAI
- Azure AI Search
Machine Learning & MLOps
- MLflow
- CI/CD for ML pipelines
- Model Registry
- Model Monitoring
- Experiment Tracking
Generative AI
- Large Language Models (LLMs)
- Retrieval-Augmented Generation (RAG)
- Embeddings
- Vector Databases
- AI Copilots
- Agentic AI Workflows
Data Engineering
- Batch Processing
- Streaming Pipelines
- Feature Engineering
- Unstructured Data Processing
Data Types
- JSON
- Logs
- Documents
- PDFs
- Multi-modal Data
Required Experience
- 6–8 years of hands-on experience in Data Engineering.
- Strong expertise in Python, PySpark, and SQL.
- Experience with Apache Spark, Airflow, and Kafka.
- Hands-on experience processing unstructured data such as JSON, logs, documents, and PDFs.
- Experience working on Microsoft Azure cloud platforms.
- Exposure to GenAI technologies including RAG pipelines, embeddings, and vector databases.
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or a related field.
- Azure certifications in Data Engineering or AI are an advantage.
- Experience designing enterprise-scale AI and analytics platforms.
- Familiarity with MLOps best practices and production AI deployments.
Soft Skills
- Excellent analytical and problem-solving abilities.
- Strong communication and stakeholder management skills.
- Ability to work independently and within cross-functional teams.
- Strong ownership mindset with attention to quality and scalability.
- Passion for emerging AI and cloud technologies.
Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)