Key Skill Data Science Mandatory Skills GenAI / LLM (NLP, RAG, Prompt Eng.) Python (ML / Data Science) Machine Learning — Supervised & Unsupervised Healthcare Data — HL7 FHIR / eCQM / HEDIS SQL & Data Wrangling Model Evaluation & MLOps basics Data Visualization Desired Skills Databricks (Delta Lake, ML Runtime) Spark / PySpark Azure ML / AWS SageMaker Clinical NLP (de-identification, ICD/CPT coding) Vector DBs (Pinecone / Weaviate / pgvector) JIRA / Agile delivery Job Description "Required Design, develop, and deploy machine learning and GenAI models for clinical and operational healthcare use cases Build and maintain RAG pipelines, LLM prompt chains, and fine-tuned models for healthcare NLP abstraction Apply data science methods to eCQM, HEDIS, and HCC risk adjustment datasets using FHIR-structured data Develop end-to-end ML pipelines from feature engineering through model deployment and monitoring Collaborate with product and engineering teams to translate analytical requirements into scalable solutions Create data visualizations and dashboards to communicate model insights to clinical and business stakeholders Perform exploratory data analysis on large clinical datasets — EHR, claims, lab, pharmacy Ensure model fairness, explainability, and compliance with HIPAA data standards Maintain clear documentation of models, experiments, and data lineage Stay current with advances in GenAI, clinical NLP, and healthcare data standards Desired Experience with Databricks Lakehouse platform — Delta Lake, Workflows, ML Runtime Familiarity with large-scale data processing using Apache Spark / PySpark Prior work on clinical NLP tasks: named entity recognition, de-identification, ICD/CPT auto-coding Exposure to cloud ML platforms (Azure ML, AWS SageMaker, GCP Vertex AI) Knowledge of vector databases and semantic search for healthcare document retrieval Strong problem-solving and analytical mindset with a track record of production model delivery"