Project Role : AI Infrastructure Architect
Project Role Description : Architect and build custom Artificial Intelligence (AI) infrastructure/hardware solutions. Optimize AI infrastructure/hardware performance, power consumption, cost and scalability of computational stack. Advise on AI infrastructure technology and vendor evaluation, selection and full stack integration.
Must have skills : Databricks Unified Data Analytics Platform
Good to have skills : Machine Learning Operations
Minimum 5 year(s) of experience is required
Educational Qualification : 15 years full time education
Role Summary / Description
AI Powered Tech Talent
As a hands-on Engineer in AI Infrastructure Architecture, you will design, build, automate, monitor and optimize Databricks-based AI/ML infrastructure for scalable data, feature engineering, model training, experimentation and production deployment workloads. you will work on moderately complex platform components under guidance from senior architects and engineers, contributing to cluster configuration, ML pipelines, model serving, CI/CD, observability, governance and operational reliability for AI-driven business solutions.
Key Responsibilities
Write, review and debug code, notebooks, scripts and infrastructure-as-code for Databricks AI/ML infrastructure, automation, monitoring and deployment tooling.
Configure and manage Databricks workspaces, clusters, jobs, model registry, MLflow tracking, feature engineering workflows and integrations with cloud storage and compute services.
Support deployment automation and CI/CD pipelines for Databricks workloads using tools such as Git, Databricks Asset Bundles, Terraform, Docker, Kubernetes and workflow orchestration tooling where applicable.
Deploy and operate ML pipelines, model-serving components and data/feature pipelines while applying reliability, security, cost-efficiency and scalability practices.
Monitor platform, cluster and job health troubleshoot issues across notebooks, jobs, compute clusters, libraries, storage access, networking and model-serving layers.
Collaborate with data scientists, ML engineers, platform engineers and architects to integrate Databricks AI models and pipelines into enterprise systems while meeting compliance and operational requirements.
Document reusable patterns, configuration standards and runbooks for Databricks-based AI infrastructure.
Required Qualifications
Bachelor's degree in Computer Science, Computer Engineering, Information Technology or a related engineering field.
Minimum 2 years of experience coding, building, monitoring or troubleshooting AI/ML infrastructure, data platforms, model deployment pipelines or cloud/platform engineering solutions.
Strong understanding of AI/ML concepts and the compute, storage, networking, security and deployment foundations required to run AI workloads.
Minimum 2 years of proficiency in programming or scripting languages such as Python, Java, C++, Bash or PowerShell.
Experience with CI/CD, infrastructure-as-code, containers, Kubernetes, workflow orchestration and operational monitoring tools.
Strong problem-solving ability, communication skills and collaboration mindset in a fast-paced engineering environment.
Required Skills/ Experience
Hands-on experience with Databricks workspaces, clusters, jobs, notebooks, MLflow, model registry, feature engineering workflows and cloud storage integrations.
Experience designing or operating scalable data and ML pipelines, distributed processing workloads, model-serving patterns and production ML operations.
Working knowledge of Python, SQL, Spark, Terraform/Databricks Asset Bundles, Git-based CI/CD and observability practices.
Ability to optimize clusters, jobs and pipelines for performance, reliability, scalability, cost and security.
Understanding of MLOps patterns including experiment tracking, model registry, model deployment, monitoring and rollback approaches.
Good to Have Skills
Databricks certification such as Databricks Data Engineer, Machine Learning Associate/Professional or related lakehouse credentials.
Exposure to industry use cases in BFSI, healthcare, retail/e-commerce, telecom, manufacturing or public sector where data/AI platforms must meet compliance, reliability and data-governance expectations.
Familiarity with large language model workflows, vector search, retrieval pipelines, feature stores, model optimization or GPU-backed training patterns.
Knowledge of Unity Catalog, data governance, FinOps practices, incident management and production support processes for enterprise AI platforms.
15 years full time education