Data Engineer
Experience: 6+ Years
Work Hours: 1:00 to 10:00 pm IST
Duration: 6 months (Extension possible)
Overall, they need to know:
Amazon Neptune
Vector DB like OpenSearch and RAG patterns,
Knowledge of Strands Agents, Ontology design and knowledge representation
Glue Catalog, S3 Tables, Apache Iceberg format
TTL files and Databricks familiarity is good
Python and data engineering skills
LLM application development (Amazon Bedrock, prompt engineering, RAG patterns),
Knowledge graph integration — graph-augmented generation, subgraph retrieval for context injection, Python/TypeScript development for AI pipelines (LangChain, LlamaIndex, or equivalent),
Text2SQL experience
Skills: expertise in ontology design and knowledge representation (RDF, OWL, SKOS, taxonomy modeling) along with strong graph database skills (Neptune or Neo4j) for modeling rule entities, dependencies, and lineage. Python and data engineering skills
Optional Data Engineer skills
Ingestion from enterprise and vendor sources — APIs, flat files, SFTP drops, CDC, batch and incremental loads
Pipeline orchestration — Airflow, Glue, Step Functions; dependency management, retries, backfills, idempotent reruns
Transformation at engineering standard — dbt or equivalent, modular SQL, tested and version-controlled
Strong SQL and Python; performance tuning on large tables
AWS data stack — S3, Glue Catalog, Athena, Redshift, Iceberg/Parquet, partitioning and file-layout decisions
Data contracts and schema drift detection on third-party feeds — automated checks, alerting, quarantine of bad loads
Data quality engineering — Great Expectations or similar, completeness/freshness/conformance checks as pipeline gates
Prepares data for agent consumption — chunking and embedding pipelines, vector store loading, metadata and catalog design so retrieval has something to point at
Pay: ₹90,000.00 - ₹110,000.00 per month
Work Location: Remote