Overview:
Senior Cloud Infrastructure Engineer
We're looking for a Senior Cloud Infrastructure Engineer to own the design, resilience, and day-to-day health of HealthEdge's infrastructure across AWS and our hybrid on-prem/Azure/GCP estate. This role sits at the intersection of cloud engineering and infrastructure engineering spanning AWS migration execution, disaster recovery, patching and platform currency, and the security/compliance controls that keep our security obligations intact. It's a hands-on senior IC role for someone who wants deep ownership of infrastructure resilience across a large, multi-account, multi-platform environment that's actively migrating off legacy on-prem infrastructure.
Areas of Responsibility:
Infrastructure and Hybrid Cloud Architecture
-
Design, build, and maintain infrastructure across AWS (primary), with supporting work in Azure and GCP, plus the on-prem estate still in active retirement
-
Own infrastructure across the environments including secure configuration baselines and patch management for OS images and on-prem hardware
-
Build and maintain reusable, auditable Infrastructure as Code for cloud deployments, and for remaining on-prem server deployment.
-
Support cloud networking execution, VPC provisioning, security group standards, and related connectivity work.
-
Manage storage across cloud and legacy on-prem storage as workloads migrate; contribute to on-prem retirement and datacenter decommissioning efforts.
-
Support AWS migration execution for in-flight waves, including server deployments and resource change requests via IaC.
Disaster Recovery & Resilience
-
Own disaster recovery architecture and execution across cloud environments.
-
Maintain DR solutions and backup strategy and run DR drills on a regular cadence; document gaps and drive remediation.
-
Design for resilience from the start and treat recoverability as a first-class requirement, not an afterthought.
Security, Compliance & Vulnerability Management
-
Contribute to vulnerability management triage across infrastructure teams, threat detection, and infra security findings review.
-
Support PHI/PII classification scanning, penetration test coordination, and security exception approvals.
-
Maintain compliance controls; support HIPAA and SOC 2 audit readiness, access review and recertification, evidence collection, and change freeze coordination.
-
Maintain EKS container runtime security sensor coverage as part of ongoing platform hardening.
Cloud Infrastructure Operations
-
Design and manage roles and cross-account access controls following least-privilege principles across multi-account, multi-cloud environments.
-
Own cloud execution: load balancers, DNS (), VPC provisioning, and security group standards.
-
Manage compute resources at scale with an eye toward right-sizing and long-term maintainability.
-
Administer cloud storage and database services with attention to cost, performance, and resilience.
-
Own infrastructure health, cost, and performance monitoring using native and third-party tooling, building the observability that lets issues surface before they become incidents.
-
Administer and harden Linux and Windows Server environments across cloud and on-prem, including patching, performance tuning, troubleshooting, Active Directory integration, Group Policy, DNS, and certificate services.
-
Manage hybrid identity and authentication across on-prem and cloud workloads, and maintain OS-level security baselines and hardening standards across the estate.
CI/CD, Automation & Delivery
-
Build and evolve CI/CD pipelines for secure, repeatable infrastructure deployments.
-
Write automation to reduce manual toil and enforce operational consistency across cloud and on-prem environments.
-
Take solutions from proof-of-concept to production with an eye toward long-term maintainability, not just getting it working once.
Reliability, Monitoring & Incident Response
-
Monitor, scale, and maintain production infrastructure with availability, performance, and security as top priorities.
-
Participate in on-call rotation; serve as L2 escalation point for cross-team infrastructure support; lead root cause analysis and drive incident retrospectives to closure.
-
Author and maintain runbooks that hold up under pressure, not just at handoff.
Cost & Tagging Governance
-
Contribute to FinOps efforts: identify and remediate cost anomalies, own tagging remediation against enterprise tagging standards, and make pragmatic cost/performance/resilience tradeoffs.
AI-Enabled Engineering
-
Use AI coding assistants to accelerate IaC development, scripting, and troubleshooting.
-
Use AI tooling to draft first-pass runbooks, DR documentation, and incident retrospectives to validate and refine before publishing.
Collaboration & Documentation
-
Document architecture, DR runbooks, and standard operating procedures others can actually follow under pressure.
-
Provide technical guidance to product teams on infrastructure resilience, migration sequencing, and recovery design.
Required Qualifications
-
5+ years of hands-on cloud infrastructure engineering experience, with deep expertise in AWS and in other cloud environments.
-
Direct experience with disaster recovery design and execution.
-
Strong Infrastructure as Code experience (CDK, Terraform, or CloudFormation).
-
Experience with containerized environments and Kubernetes/EKS, including version upgrade and lifecycle management.
-
Linux and Windows Server administration experience, including patching and OS lifecycle management at scale.
-
Solid IAM design experience, including cross-account access and least-privilege enforcement.
-
Strong scripting ability (Python, Bash, or PowerShell).
-
Experience building and maintaining CI/CD pipelines.
-
Comfortable being the primary on-call and L2 escalation point for infrastructure incidents.
Preferred Qualifications
-
Experience operating in regulated environments (FedRAMP, HIPAA, SOC 2) and understanding of what that means for infrastructure and DR design specifically.
-
AWS certification (Solutions Architect or SysOps, Associate or Professional).
-
Experience with hybrid infrastructure, bridging on-prem virtualization with cloud-native services during active migration.
-
Familiarity with DISA STIG or CIS benchmark hardening, and vulnerability management/triage workflows.
-
FinOps or cost governance experience, including tagging standards enforcement.
-
Healthcare technology or digital health platform background.
-
Experience with AI-assisted engineering workflows as part of daily practice.
Behaviors & Traits
-
Raises risk early rather than waiting for it to become an incident.
-
Comfortable with ambiguity in a large, multi-account, evolving cloud environment.
-
Strong sense of ownership; closes gaps rather than escalating and waiting.
-
Communicates technical tradeoffs clearly to both engineers and non-technical stakeholders.