Location: Bangalore, India
nference is an AI-first healthcare technology company dedicated to accelerating biomedical discovery and transforming healthcare through cutting-edge artificial intelligence and large-scale computing. Our platforms help pharmaceutical companies, healthcare providers, and research organizations unlock insights from complex clinical, molecular, and imaging data to drive better decisions and improve patient outcomes.
Founded with the vision of solving some of healthcare's most challenging problems, nference combines deep expertise in software engineering, machine learning, data science, and life sciences to build products that make a meaningful impact on global healthcare. Recognized by The Washington Post as the "Google of Biomedicine," we continue to push the boundaries of innovation by partnering with leading healthcare institutions and industry pioneers.
Our people are our greatest strength. At nference, you'll work alongside exceptional software engineers, AI researchers, physicians, scientists, and product leaders who are passionate about solving meaningful, real-world problems.
We foster a culture built on ownership, curiosity, collaboration, and continuous learning. Engineers are encouraged to challenge assumptions, contribute ideas, influence technical direction, and take ownership of their work from concept to production. Whether you're improving platform reliability, automating cloud infrastructure, or enhancing developer productivity, you'll collaborate with some of the brightest minds in technology and healthcare.
We're looking for a Site Reliability Engineer (SRE) who is passionate about building and maintaining reliable, scalable, and highly available cloud infrastructure. In this role, you'll help ensure the stability, performance, and operational excellence of the platforms that power our AI-driven healthcare products.
You'll work closely with Software Engineering, Platform Engineering, AI, Data Science, and Product teams to improve system reliability through automation, observability, and modern infrastructure practices. This role is ideal for someone who enjoys solving operational challenges, automating repetitive tasks, and continuously improving platform resilience and developer experience.
-
Monitor the health, availability, and performance of cloud infrastructure, Kubernetes clusters, CI/CD systems, and platform services.
-
Assist in production incident response by collecting logs, metrics, and diagnostic information to support rapid troubleshooting and recovery.
-
Configure, maintain, and improve monitoring dashboards, alerting systems, and observability platforms.
-
Develop automation scripts and operational tooling to eliminate manual processes and improve infrastructure efficiency.
-
Contribute to Infrastructure-as-Code (IaC) modules for provisioning and managing cloud resources.
-
Support the reliability, maintenance, and continuous improvement of CI/CD pipelines and deployment workflows.
-
Participate in production operations while learning and applying Site Reliability Engineering principles, including SLIs, SLOs, error budgets, and incident management.
-
Collaborate with software engineers to improve system reliability, scalability, and deployment processes.
-
Troubleshoot infrastructure, networking, and platform-related issues across cloud environments.
-
Maintain operational documentation, runbooks, and knowledge repositories to improve incident response and operational consistency.
-
Continuously identify opportunities to improve platform reliability, automation, and operational excellence.
-
Bachelor's or Master's degree in Computer Science, Information Technology, Software Engineering, or a related technical discipline.
-
1–3 years of professional experience in Site Reliability Engineering, DevOps, Cloud Infrastructure Engineering, or a related systems engineering role.
-
Strong understanding of Linux systems administration, command-line tools, process management, system diagnostics, and troubleshooting.
-
Proficiency in at least one scripting or programming language such as Python, Bash, or Shell for automation and tooling.
-
Hands-on experience with at least one major cloud platform (AWS, GCP, or Azure) and familiarity with core infrastructure services including compute, networking, and storage.
-
Working knowledge of containerization and orchestration technologies such as Docker and Kubernetes.
-
Familiarity with Infrastructure-as-Code (IaC) concepts and tools such as Terraform or similar automation frameworks.
-
Experience with monitoring, logging, and observability platforms such as Prometheus, Grafana, ELK Stack, Datadog, or similar tools.
-
Good understanding of networking fundamentals including DNS, TCP/IP, load balancing, and distributed systems concepts.
-
Familiarity with CI/CD pipelines and modern software delivery practices.
-
Experience using version control systems such as Git.
-
Strong analytical, troubleshooting, and problem-solving skills.
-
Excellent verbal and written communication skills.
-
Exposure to incident management and production support in cloud-native environments.
-
Familiarity with reliability engineering concepts including SLIs, SLOs, and error budgets.
-
Experience with configuration management or automation tools such as Ansible.
-
Knowledge of messaging technologies such as Kafka or RabbitMQ.
-
Exposure to distributed systems and microservices architectures.
-
Experience supporting highly available production environments.
-
Familiarity with security best practices for cloud infrastructure.
-
Contributions to open-source projects are a plus.
At nference, you'll work on technology that has the potential to transform healthcare and improve lives across the world. Every engineering challenge you solve contributes to advancing biomedical research, accelerating drug discovery, and enabling better clinical decisions.
You'll be part of a team that values innovation, ownership, and collaboration, where you'll have the opportunity to work with cutting-edge cloud technologies, modern infrastructure platforms, and automation tools while growing alongside some of the brightest minds in software engineering, artificial intelligence, and life sciences.
If you're excited about building reliable platforms that enable engineering teams to innovate faster while delivering meaningful real-world impact, nference is the place for you.
-
Industry Prestige: Build your career at the "Google of Biomedicine" (as recognized by The Washington Post), working alongside exceptional software engineers, physicians, scientists, and researchers.
-
Cutting-Edge Innovation: Solve complex healthcare challenges using advanced AI, machine learning, and large-scale clinical, molecular, and imaging datasets.
-
Meaningful Impact: Contribute to technologies that accelerate drug discovery and biomedical research, with opportunities to be recognized as a contributing author on high-impact scientific publications where applicable.
-
Growth & Flexibility: Thrive in a collaborative, innovation-driven culture with continuous learning opportunities and a hybrid work model for eligible employees after successful completion of the three-month probation period.
-
Wellness & Perks: Enjoy reimbursements for gym memberships, technology gadgets, high-speed internet, professional development, comprehensive health insurance, and complimentary breakfast, lunch, and snacks at our Bangalore office.
nference is committed to building a diverse, equitable, and inclusive workplace where everyone has the opportunity to thrive. We celebrate different perspectives, backgrounds, and experiences because they strengthen our teams and drive innovation. We are proud to be an equal opportunity employer and welcome applicants from all backgrounds to join us in building technology that transforms healthcare.