Description:
Position Summary
The NOC Engineer is part of a newly established Network Operations Center that moves IT operations from a reactive support model toward proactive, centralized monitoring across our US and India environment. The role spans the full operational picture — platform and service health, business applications, data ingestion and integration jobs, and security alerting — not infrastructure alone.
This is not a help desk role that waits for tickets. The NOC Engineer detects issues through monitoring and alerts, assesses business impact, performs Tier 1 triage and resolution, escalates the rest to the appropriate engineering, application, data, or security teams, and retains ownership until service is restored. Desktop and end-user support is carried in parallel with these duties on the same shift.
The NOC is the operational front door. It does not replace specialized engineering teams: infrastructure, network, cloud, security, application, and data teams retain ownership of advanced engineering work and permanent corrective actions.
Key Responsibilities
Monitoring and Alert Response
- Monitor platform, service-health, and application dashboards throughout the shift; validate every alert and classify it by severity and business impact before acting.
- Respond to alerts covering availability, performance, CPU, memory, disk, connectivity, service health, and application errors.
- Write and modify log queries (KQL) to investigate incidents and surface operational trends; tune alert thresholds and routing rules to reduce noise and false positives.
- Monitor connectivity and latency between sites, VPN paths, and cloud services; monitor virtual desktop session health and user experience.
- Monitor backups, endpoint health, patching, device compliance, and scheduled operational jobs; perform documented daily health checks for anything not yet covered by automated monitoring.
- Create and maintain the service-health dashboards and workbooks the team works from, and keep them current as the environment changes.
Map what is monitored today, identify critical systems and services not yet covered, and onboard them into monitoring so that manual checks are progressively replaced by automated alerting
Requirements:
Tier 1 Application and Data Ingestion Support
- Monitor business-critical applications for availability, response time, error rates, and failed transactions; perform Tier 1 triage and resolution within documented procedures.
- Monitor data ingestion, integration, and interface jobs — scheduled transfers, feeds, and pipelines — for failures, delays, backlogs, and record-count or reconciliation mismatches.
- Restart failed jobs, re-trigger transfers, and clear known error conditions where a runbook exists; escalate anything outside the runbook to the application or data teams with full diagnostic detail.
- Track ingestion issues through to confirmed data delivery, not just to job restart.
Security Alert Monitoring (Tier 1)
- Monitor and triage security alerts — suspicious sign-ins, account lockouts, MFA anomalies, endpoint protection detections, and reported phishing — as first line of review.
- Validate alerts against documented criteria, capture supporting evidence, and escalate to the Security team promptly with a clear summary; deeper investigation and remediation remain with Security.
- Follow defined containment steps (such as user or device isolation) only where explicitly authorised by runbook.
Incident Response and Ownership
- Perform first-level triage, determine business impact, resolve within established procedures, and escalate complex incidents to the Network, Server, Cloud, Application, Data, or Security teams as appropriate.
- Create and update incidents in Freshservice with accurate categorisation, timelines, and diagnostic detail.
- Own escalated incidents through to service restoration; support incident coordination, communication, and stakeholder updates during major incidents.
- Maintain a consistent shift-handover log and end-of-shift summaries in Teams and Freshservice.
- Identify recurring incidents, contribute to root-cause analysis and post-incident reviews rather than repeatedly resolving the same symptoms.
- Maintain targeted watchlists for known problem areas and review operational trends to spot issues before they generate alerts.
Desktop and End-User Support (carried in parallel)
- Provide day-to-day desktop and end-user support — desktops, laptops, peripherals, printers, conference rooms, and business applications — alongside monitoring duties on the same shift.
- Manage the Freshservice queue to agreed response and resolution targets, prioritising between desktop requests and NOC alerts by business impact.
- Support endpoint lifecycle activities: imaging, deployment, onboarding and offboarding, compliance, patching, and asset accuracy.
- Provide the Help Desk with clear pass/fail checks, documented escalation steps, and defined NOC tasks they can support during each shift.
- Train and cross-train Help Desk staff to acknowledge, validate, and escalate alerts correctly, and act as a technical point of reference for them during the shift.
Documentation, Automation, and Reporting
- Develop and maintain runbooks, SOPs, troubleshooting guides, and escalation paths — every alert should map to a documented action.
- Build automation with PowerShell and Azure Automation to reduce repetitive monitoring and remediation work.
- Contribute to operational reporting on service availability, alert trends, recurring issues, response times, and incident resolution, giving leadership clear visibility into the health of the technology environment.
Required Qualifications
- Bachelor’s degree in a technical discipline, or equivalent practical experience.
- 2–4 years in IT support or NOC monitoring, with exposure beyond desktop-only troubleshooting.
- Sound networking fundamentals — TCP/IP, subnetting, DNS, DHCP, VPN, and VLAN concepts — with the ability to isolate where in the path a fault sits using ping, traceroute, nslookup, and port checks. Device configuration depth is not required.
- Working knowledge of cloud monitoring on Microsoft Azure — the portal, Azure Monitor, Log Analytics, Resource Health, and Service Health: reading metrics and logs, understanding alert rules and action groups, and basic KQL. Demonstrable KQL familiarity will be prioritised.
- Ability to read application and job logs to identify failures, and to distinguish an application fault from a platform fault before escalating.
- Awareness of common security alert types — suspicious sign-in, account lockout, malware detection, phishing — and the discipline to escalate rather than investigate beyond defined scope.
- Strong Windows endpoint and server administration, including Active Directory, Entra ID, and Microsoft 365, plus proven hands-on desktop support experience.
- Experience with an ITSM ticketing platform such as Freshservice or ServiceNow, and disciplined documentation of every action taken.
- A methodical approach to troubleshooting, and good spoken and written English — able to explain a technical problem clearly and simply, both on calls and in written updates.
- Availability for a 24×7 rotational shift, including nights, weekends, and holidays on rotation.
Preferred Qualifications
- Experience building alert rules, workbooks, or dashboards from scratch; intermediate KQL (joins, aggregations, time-series operators).
- PowerShell scripting and Azure Automation.
- Exposure to Microsoft Sentinel, Microsoft Defender for Cloud, Microsoft Intune, or Defender for Endpoint.
- Exposure to Power BI or Microsoft Fabric for operational reporting.
- Exposure to application or integration monitoring, scheduled job orchestration, SQL basics, or API and file-transfer troubleshooting.
- Prior experience supporting a healthcare or other regulated environment.
Certifications such as AZ-900, AZ-104, SC-900, MS-900, CCNA, or ITIL Foundation will be an added advantage