[About Yotta](https://yotta.com/)
Yotta Data Services is India’s leading sovereign AI infrastructure, cloud platform and data centre services company, enabling enterprises, governments, startups, and digital platforms to build, deploy, and scale next-generation AI and digital workloads securely within India.
With hyperscale data centre campuses in Navi Mumbai and Greater Noida (Delhi NCR), advanced GPU-powered AI infrastructure, and a comprehensive ecosystem of cloud, AI, hosting, cybersecurity, and managed platform services, Yotta delivers high-performance, scalable, and compliant digital infrastructure built for the AI era.
Yotta is at the forefront of powering India’s sovereign AI and digital transformation journey through world-class infrastructure, deep technology partnerships, and fully India-hosted enterprise-grade platforms.
Job Scope
We are looking for an enthusiastic Junior Data Platform Engineer to support and manage our Apache Spark, Apache Airflow, and JupyterHub environments. This role is ideal for someone with a strong foundation in Python and Linux, who is eager to build a career in big data engineering and data platform administration.
You will work closely with senior engineers to ensure smooth operation, deployment, and optimization of our data processing ecosystem.
Total /Relevant Experience
years experience
in the setup, monitoring, and maintenance of Apache Spark clusters, Apache Hive, Hadoop and Airflow environments.
the development and scheduling of data pipelines using Airflow DAGs and Python scripts.
manage and configure JupyterHub for multi-user access and integration with Spark.
cluster health and performance under guidance and assist in troubleshooting Spark job failures.
and maintain Python automation scripts for data workflows, ETL, and process automation.
in code reviews, documentation, and deployment activities.
and follow best practices for distributed data processing, CI/CD, and DevOps workflows.
with senior engineers and data scientists to implement improvements and new features.
Must-have skill
understanding of Apache Spark, Apache Hive and Hadoop File System.
with Apache Airflow (understanding of DAGs, scheduling, and task dependencies).
experience with Python scripting (data processing, automation, or API interaction).
working in Linux and container environments (command line, system logs, process management).
understanding of data processing concepts, including ETL and distributed computing.
knowledge of Git and version control.
Good-to-Have Skills
to Jupyter / JupyterHub for collaborative notebook environments.
of Docker or Kubernetes.
of Hadoop and Apache Spark Cluster.
with SQL and working with structured/unstructured data.
with cloud platforms (AWS, GCP, or Azure) is a plus.
in big data (Hadoop), DevOps, and data pipeline automation.
Qualifications Criteria
or any relevant Degree.
Certification Criteria
Shifts Timing, (If rotational, please specify)
Shift
Number of Interview Rounds
Name of interviewer
rounds
Behavioral Attributes:
of skilful conversation
& Problem Solving
on fly
Acumen
Trust
Focus
Horsepower (Functional Skills)
Orientation & Accountability
Excellence
Planning & organizing
Sensing, Observing
Direct Reports
Collaborative Relationships