Bei Roche kannst du ganz du selbst sein und wirst für deine einzigartigen Qualitäten geschätzt. Unsere Kultur fördert persönlichen Ausdruck, offenen Dialog und echte Verbindungen. Hier wirst du für das, was du bist, wertgeschätzt, akzeptiert und respektiert. Dies schafft ein Umfeld, in dem du sowohl persönlich als auch beruflich wachsen kannst. Gemeinsam wollen wir Krankheiten vorbeugen, stoppen und heilen und sicherstellen, dass jeder Zugang zur Gesundheitsversorgung hat – heute und in Zukunft. Werde Teil von Roche, wo jede Stimme zählt.
Die Position
Job Description:
The Data Engineer – Clinical Study Design sits at the intersection of data architecture, clinical science, and technology delivery, helping build the data foundation behind Study Designer, a digital product transforming how studies are designed. This role combines strong data engineering skills, an understanding of clinical study data and workflows, and technical curiosity to translate complex clinical data structures into reliable, scalable pipelines and models that power the platform's insights. The ideal candidate is curious, collaborative, AI-minded, and passionate about using data to enable smarter, faster, and more effective clinical study design.
Key Responsibilities
- Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles (PubMed, CTIS, ClinicalTrials.gov) - handling PDF parsing, text extraction, and structured data normalization
- Design and implement data models inAmazon Aurora (relational) andGraphDB (knowledge graph) to represent trial design entities: endpoints, eligibility criteria, study arms, interventions, therapeutic areas, and their relationships
- Develop embedding and vectorization pipelines to prepare extracted clinical text for RAG-based retrieval in LangGraph agentic workflows - chunking strategies, metadata enrichment, and vector store population
- Build and maintain ETL/ELT workflows that transform unstructured clinical content into queryable, linked data across both relational and graph stores
- Implement data quality validation specific to clinical data - protocol section classification accuracy, entity extraction completeness, cross-reference integrity (NCT IDs, EudraCT numbers, MeSH terms)
- Build data serving APIs (Python/FastAPI) that expose curated datasets to the Angular frontend and LangGraph agent layer
- Set up data lineage tracking and audit trails to support regulatory traceability of AI-generated trial design recommendations
Preferred Qualifications:
- Education: Bachelor's degree in Computer Science, Data Engineering, or a related discipline.
- 5-8 years of experience building production grade data platforms and pipelines.
- Experience with biomedical knowledge graphs (e.g., linking drugs - > targets - > diseases - > trials)
- Prior work with PubMed/MEDLINE data, ClinicalTrials.gov API, or EMA/CTIS data
- Apache Spark or Databricks for batch processing of large document
Required Skills
- Python — Primary language; experience with PDF/document parsing libraries (PyMuPDF, pdfplumber, unstructured.io, or similar)
- SQL — Advanced PostgreSQL-compatible SQL (Aurora); schema design, migrations, query optimization, indexing strategies for clinical data volumes
- Graph Databases — Hands-on with Neptune, Neo4j, or similar; SPARQL or Cypher query language; ontology/knowledge graph modeling for biomedical entities
- AWS — Aurora (PostgreSQL), S3, Lambda, Step Functions, SQS/SNS, IAM; infrastructure for data pipeline orchestration
- NLP / Document Processing — Text extraction from PDFs, section classification, named entity recognition for clinical/biomedical text; familiarity with embedding models and vector stores (OpenSearch, pgvector, or Pinecone)
- FastAPI — Building data serving endpoints; async patterns; integration with the application backend
- AI/ML Data Infrastructure — Preparing data for LangChain/LangGraph consumption; RAG pipeline design (chunking, retrieval, reranking); prompt-data integration patterns
- Pipeline Orchestration — Experience with workflow orchestration tools (Airflow, Prefect, Step Functions, or Temporal); designing DAGs for multi-stage data pipelines with dependency management, retry logic, and monitoring
- CI/CD& IaC — Terraform or CDK, Docker, Git; automated pipeline testing and deployment on AWS
#Hyderabad2026
Wer wir sind
Eine gesündere Zukunft treibt uns zur Innovation an. Mehr als 100.000 Mitarbeiter weltweit arbeiten gemeinsam daran, wissenschaftliche Fortschritte zu erzielen und sicherzustellen, dass jeder Zugang zur Gesundheitsversorgung hat – heute und für zukünftige Generationen. Durch unser Engagement werden über 26 Millionen Menschen mit unseren Medikamenten behandelt und mehr als 30 Milliarden Tests mit unseren Diagnostik-Produkten durchgeführt. Wir ermutigen uns gegenseitig, neue Möglichkeiten zu erkunden, Kreativität zu fördern und hohe Ziele zu setzen, um lebensverändernde Gesundheitslösungen zu liefern.
Gemeinsam können wir eine gesündere Zukunft gestalten.
Roche ist ein Arbeitgeber, der die Chancengleichheit fördert.