Job Title: Senior Data Engineer – GenAI & Unstructured Data Pipelines
Experience: 6–8 Years
Open Positions: 2
Location: Offshore
Employment Type: Full-Time
Job Summary
We are looking for an experienced Senior Data Engineer to build next-generation GenAI data platforms that power Large Language Model (LLM)-based applications. The ideal candidate will have strong expertise in data engineering, scalable data pipelines, Azure cloud services, and Generative AI technologies, including RAG, embeddings, vector databases, and AI copilots.
Key Responsibilities
- Own the end-to-end ML lifecycle, including data ingestion, feature engineering, model training, evaluation, deployment, monitoring, retraining, and rollback.
- Design, build, and maintain production-grade ML pipelines using Azure native services with CI/CD automation.
- Utilize Azure Machine Learning and MLflow for experiment tracking, model registry, and environment promotion (Dev/Test/Prod).
- Design and implement Generative AI solutions using Azure OpenAI, embeddings, vector search, and Retrieval-Augmented Generation (RAG) pipelines.
- Develop Agentic AI workflows with multi-step reasoning, tool integration, guardrails, observability, reliability, and cost optimization.
- Build scalable batch and streaming data pipelines using Azure Databricks.
- Design and optimize pipelines for processing unstructured and multi-modal data, including text, documents, logs, images, and PDFs.
- Develop robust RAG pipelines involving document chunking, embedding generation, vector indexing, and semantic retrieval.
- Design and manage vector databases and search solutions such as Azure AI Search, Pinecone, or similar platforms.
- Build batch and real-time data ingestion pipelines using Apache Spark and Apache Kafka.
- Collaborate with data scientists, ML engineers, and software development teams to deliver AI-powered solutions.
- Ensure scalability, performance, security, and reliability of enterprise data platforms.
Required Skills
- 6–8 years of experience in Data Engineering.
- Strong programming skills in Python, PySpark, and SQL.
- Hands-on experience with:
- Apache Spark
- Apache Airflow
- Apache Kafka
- Experience processing unstructured data, including JSON, logs, documents, and PDFs.
- Strong knowledge of Azure cloud services (preferred).
- Experience building GenAI/LLM pipelines, including:
- Retrieval-Augmented Generation (RAG)
- Embeddings
- Vector Databases
- Vector Search
- Experience with Azure Machine Learning and MLflow.
- Knowledge of CI/CD pipelines and DevOps best practices.
- Strong problem-solving and analytical skills.
Preferred Qualifications
- Experience with Azure OpenAI Service.
- Knowledge of Azure Databricks, Azure AI Search, or Pinecone.
- Experience working with multi-modal AI applications.
- Familiarity with MLOps, model monitoring, and production deployment.
- Excellent communication and collaboration skills.
Why Join Us?
- Work on cutting-edge Generative AI and LLM-based enterprise solutions.
- Opportunity to build scalable AI-powered data platforms.
- Collaborative and innovative work environment.
- Career growth with exposure to modern AI and cloud technologies.
- Competitive compensation and learning opportunities.
Employment Type: Full-Time
Location: Offshore
Experience Required: 6–8 Years
Open Positions: 2
Job Types: Full-time, Fresher, Internship
Work Location: Hybrid remote in Noida, Uttar Pradesh