EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
We are seeking an experienced Lead/Senior SRE Engineer with strong Dynatrace expertise to join our team. In this role, you will drive the implementation of DevOps and SRE practices, shape the technology roadmap, and ensure the reliability, performance, and observability of our production systems. You will collaborate closely with application teams and product owners to foster a culture of operational excellence and continuous improvement.
Responsibilities
-
Help to implement DevOps & SRE practices
-
Drive discussions for the technology roadmap for the SRE team
-
Identify, craft, and maintain SLIs and SLOs for teams, as well as metrics such as MTTR, Lead time for change, Deployment Frequency, and Change Failure Rate
-
Design, develop, and manage monitoring, alerting, operability, and observability for applications using Dynatrace, Splunk, and Grafana
-
Perform performance assessments and monitoring, and recommend performance enhancements
-
Enforce application teams to meet performance and availability SLAs
-
Partner with product owners to manage error budget, prioritize toil backlog, and validate against team, application, and incident metrics
-
Participate in an on-call rotation for production events or outages
-
Strive for continuous improvement for continuous integration & continuous deployment (CI/CD Pipeline)
-
Apply troubleshooting techniques, incident management, and root cause analysis
-
Encourage and build automated processes wherever possible
-
Implement cybersecurity measures by continuously performing vulnerability assessment and risk management
-
Manage periodic reporting on the progress to management and the customer
-
Collaborate with application teams to ease their adoption of the platform, and coordinate communication within the team and with customers
-
Analyze the current system and develop plans for enhancements and improvements
Requirements
-
Bachelor's degree in Computer Science or related fields and/or equivalent work experience
-
5-12 years of general IT experience, including 5+ years working within DevOps or SRE teams
-
Background in supporting production infrastructure
-
Strong knowledge of CI/CD
-
Understanding of observability (monitoring, logging, and tracing)
-
Expertise in Dynatrace and Splunk
-
Familiarity with a leading cloud provider (AWS, Azure, or GCP)
-
Proficiency in operating high-availability, fault-tolerant, scalable, distributed software in production
-
Capability to work independently and as part of a team, with strong organizational and interpersonal skills and experience developing a culture of operational maturity
-
Strong analytical and problem-solving mindset, with strategic thinking and experience troubleshooting under pressure
-
Flexibility to adjust quickly to new technologies
-
English B2 or higher
We offer
-
Opportunity to work on technical challenges that may impact across geographies
-
Vast opportunities for self-development: online university, knowledge sharing opportunities globally, learning opportunities through external certifications
-
Opportunity to share your ideas on international platforms
-
Sponsored Tech Talks & Hackathons
-
Unlimited access to LinkedIn learning solutions
-
Possibility to relocate to any EPAM office for short and long-term projects
-
Focused individual development
-
Benefit package:
-
Health benefits
-
Retirement benefits
-
Paid time off
-
Flexible benefits
-
Forums to explore beyond work passion (CSR, photography, painting, sports, etc.)