Responsible for establishing monitoring systems, managing incident responses, automating infrastructure tasks, analyzing performance metrics, enabling continuous deployment, conducting reliability tests, ensuring security compliance, optimizing system performance, leading regional projects, and promoting knowledge sharing across teams to maintain robust and efficient systems.
1. Plan, manage, and oversee all aspects of a Production Environment for Enterprise Data Accessibility.
2. Engage with product support team to identify and provide self-service tooling and automation that drive efficiencies, scale, and overall customer experience.
3. Define strategies for Application Performance Monitoring, Unit Cost and Chaos Engineering aspects.
4. Ensure that architecture and deployment models are sufficient to support SLA commitments.
5. Leverage cloud technology and platform capabilities to provide operationally sustainable solutions that are robust and cost effective.
6. Provide regular updates and recommendations to senior leadership on health and performance of production systems
7. Ability to understand MTTR, SLO, SLI definitions and apply them to services.
8. Respond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.
9. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data, services and infrastructures.
10. Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
11. Ability to isolate problems between hardware and software and work with appropriate team(s) and vendors until a resolution has been reached.
12. Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.
13. Analyze ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.
14. Support services before they go live through activities such as system design consulting, capacity planning and launch reviews and maintain services once they are live by measuring and monitoring availability, latency and overall system health.
15. Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
16. Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead team in DevOps automation and best practices.
17. Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize mean time to recover.
Education: Bachelor's degree or equivalent in Computer Science, Electrical / Electronics Engineering, MIS or related discipline
Experience: Five (5) years of work experience in managing reliability and performance of infrastructure, applications and systems at scale.
Knowledge, Skills and Abilities
- Fluency in English
- Problem Solving Skills
- Communication
- Collaboration
- Adaptability
- Ability to operate in a 24x7 environment encompassing global timezones
Preferred Qualifications:
Pay Transparency:
Pay:
Additional Details:
FedEx was built on a philosophy that puts people first, one we take seriously. We are an equal opportunity/affirmative action employer and we are committed to a diverse, equitable, and inclusive workforce in which we enforce fair treatment, and provide growth opportunities for everyone.
All qualified applicants will receive consideration for employment regardless of age, race, color, national origin, genetics, religion, gender, marital status, pregnancy (including childbirth or a related medical condition), physical or mental disability, or any other characteristic protected by applicable laws, regulations, and ordinances.
FedEx is one of the world's largest express transportation companies and has consistently been selected as one of the top 10 World’s Most Admired Companies by "Fortune" magazine. Every day FedEx delivers for its customers with transportation and business solutions, serving more than 220 countries and territories around the globe. We can serve this global network due to our outstanding team of FedEx team members, who are tasked with making every FedEx experience outstanding.
The People-Service-Profit philosophy (P-S-P) describes the principles that govern every FedEx decision, policy, or activity. FedEx takes care of our people; they, in turn, deliver the impeccable service demanded by our customers, who reward us with the profitability necessary to secure our future. The essential element in making the People-Service-Profit philosophy such a positive force for the company is where we close the circle, and return these profits back into the business, and invest back in our people. Our success in the industry is attributed to our people. Through our P-S-P philosophy, we have a work environment that encourages team members to be innovative in delivering the highest possible quality of service to our customers. We care for their well-being, and value their contributions to the company.
Our culture is important for many reasons, and we intentionally bring it to life through our behaviors, actions, and activities in every part of the world. The FedEx culture and values have been a cornerstone of our success and growth since we began in the early 1970’s. While other companies can copy our systems, infrastructure, and processes, our culture makes us unique and is often a differentiating factor as we compete and grow in today’s global marketplace.