We are looking for an experienced systems operations engineer to own the day-to-day health of a mixed on-premises and multi-cloud estate — and to steadily make it better. This is a hands-on role carrying lead-level responsibility. You will keep servers, identity, mail and endpoints running reliably, run the monthly maintenance cycle, support end users when they need you, and bring forward the automation, tooling and security improvements that stop the same problems recurring.
-
Administer and maintain Windows and Linux servers across on-premises Proxmox and cloud (AWS, Azure)
-
Own identity and collaboration: Active Directory, Microsoft 365, and Google Workspace / Gmail
-
Monitor availability, capacity, backups and alerts — and act before users notice
-
Plan and execute the monthly maintenance cycle: patching, updates, backup verification, reporting
-
Manage the endpoint protection estate and enforce security baselines
-
Provide escalated end-user support and mentor junior engineers
-
Identify, propose and implement automation and infrastructure improvements
Cloud
AWS, Microsoft Azure
Virtualisation
Proxmox VE, Containers
Servers
Windows Server, Linux (RHEL / Ubuntu)
Identity
Active Directory
Collaboration
Google Workspace (Gmail)
Endpoint & security
Antivirus / EDR, Patch management
Automation
PowerShell, Bash, Python; IaC tooling
-
One to two years in systems or infrastructure operations, spanning both cloud and on-premises environments
-
Demonstrated depth in at least two of the platforms listed above
-
Proven scripting and automation ability in PowerShell, Bash or Python
-
A track record of proposing and delivering an improvement — not only keeping the lights on
-
Clear written English and the discipline to document what you build
-
Build, configure, patch and decommission Windows and Linux servers across Proxmox, AWS and Azure
-
Manage compute, storage, networking and spend in AWS and Azure; identify and remove waste
-
Maintain the Proxmox cluster: host health, VM lifecycle, snapshots, storage and high-availability behaviour
-
Own backup and recovery — schedules, retention, and restore tests that actually get performed
-
Administer Active Directory: users, groups, GPO, DNS/DHCP, and the joiner-mover-leaver process
-
Administer Google Workspace: mailboxes, licensing, sharing, retention, mail routing, SPF/DKIM/DMARC
-
Run periodic access reviews and clear down stale accounts and excess permissions
-
Keep monitoring and alerting meaningful — tune out noise, add coverage where we are blind
-
Track capacity and performance trends and raise needs before they become urgent
-
Run the monthly maintenance window end to end: patch planning, change notes, execution, verification, written report
-
Lead incident response for infrastructure issues and write short, honest post-incident notes with follow-up actions
-
Maintain endpoint protection coverage; investigate detections and respond
-
Enforce hardening baselines, MFA, least privilege, and vulnerability remediation timeframes
-
Track emerging threats and current standards such as CIS Benchmarks and ISO 27001-style controls, and tell us where we fall short
-
Script away recurring manual work: onboarding, reporting, patching, routine health checks
-
Evaluate tools and practices from outside our sector and present short recommendations with a clear before-and-after
-
Move suitable workloads toward infrastructure-as-code and repeatable builds
-
Keep runbooks and documentation current so the team is never dependent on one person
-
Handle escalated tickets across desktop, mail, access and connectivity
-
Mentor junior engineers, review their work, and raise the overall capability of the team
-
Build out self-service knowledge-base content to reduce repeat tickets
-
First 3 months. Estate mapped and documented, monitoring gaps identified, and the monthly maintenance cycle running as a repeatable, reported process.
-
By 6 months. At least two recurring manual tasks automated; patching and backup verification consistently on schedule; endpoint protection coverage at or close to 100%.
-
By 12 months. A measurable fall in repeat incidents, a security baseline defined and largely enforced, and at least one externally-inspired improvement delivered with its impact demonstrated.