Noida, Uttar Pradesh
Job Summary
HCLTech is seeking a seasoned Network Reliability Engineer (NRE) with deep expertise in designing, operating, and optimizing hybrid & cloud network environments. You will focus on ensuring high availability, scalability, performance, and resilience of enterprise networks spanning on-premises data centers, SD-WAN, and major cloud platforms (AWS, Azure).
This role combines network engineering with Reliability Engineering principles—emphasizing automation, observability, proactive incident prevention, and rapid recovery to meet stringent SLAs for mission-critical applications.
Key Responsibilities
Reliability & Availability : Design and implement highly reliable network architectures. Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
Incident Management : Lead major incident response, root cause analysis (RCA), and post-mortem reviews. Implement blameless post-mortems and drive corrective actions to prevent recurrence.
Automation & Self-Healing : Build automation scripts, tools, and self-healing mechanisms for network provisioning, configuration management, monitoring, and failover.
Monitoring & Observability : Develop and maintain comprehensive dashboards, alerts, and logging using tools like Prometheus, Django, Grafana, Datadog, Splunk etc.
Capacity Planning & Performance : Conduct network capacity planning, performance tuning, and chaos engineering to validate resilience.
Security & Compliance : Collaborate on network security posture, zero-trust models, firewall policies, and compliance requirements.
Cross-functional Collaboration : Work with DevOps, SRE, Cloud, Security, and Application teams to embed reliability into the development lifecycle.
Mentorship : Guide junior engineers and contribute to knowledge sharing within the global team.
Skill Requirements
Required Skills & Experience (15 Years Profile)
Core Networking (10+ years hands-on):
Deep expertise in Routing (BGP, OSPF, EIGRP), Switching (STP, VLANs, VXLAN), Firewalls, Load Balancers, VPNs, and SD-WAN.
Strong troubleshooting of complex Layer 2/3/4 issues.
SRE/NRE Practices :
Experience applying SRE principles (error budgets, toil reduction, automation).
Proficiency in scripting (Python) and Infrastructure as Code (Terraform, Ansible, etc.).
Tools & Technologies :
Monitoring: Prometheus, Grafana, Datadog, etc.
CI/CD & Automation: Jenkins, GitOps, Ansible.
Packet analysis: Wireshark, tcpdump.
Soft Skills : Excellent problem-solving, communication, and stakeholder management. Ability to work in a global, 24x7 on-call rotation.
Education : Bachelor’s or Master’s in Computer Science, Engineering, or related field (or equivalent experience).
Preferred Certifications :
CCIE (Enterprise, Data Center, or Security)/CCNP (R&S)
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-