Business Unit:
Data Strategy Group
Industry:
No Industry
Overview
Houlihan Lokey, Inc. (NYSE:HLI) is a leading global investment bank recognized for delivering independent strategic and financial advice to corporations, financial sponsors, and governments. With uniquely deep industry expertise, broad international reach, and a partnership approach rooted in trust, the firm provides innovative, integrated solutions across mergers and acquisitions, capital solutions, financial restructuring, and financial and valuation advisory. Our unmatched transaction volumes provide differentiated, data-driven perspectives that help our clients achieve their most critical goals. To learn more about Houlihan Lokey, please visit HL.com.
We are seeking a highly accomplished Principal Platform Engineer to serve as the technical authority and architectural leader for our cloud-native platform and infrastructure within the Financial Platform organization. In this role, you will define the platform architecture roadmap, design mission-critical infrastructure patterns, and lead the technical execution of platforms that power large-scale, distributed financial applications.
Platform Architecture & Strategic Direction: Define and drive the target-state architecture for our next-generation, cloud-native platform (Kubernetes, multi-region cloud topologies, and AI/workload enablement). Translate overarching business goals into scalable, resilient, and decoupled platform capabilities.
Hands-On Core Engineering: Lead by example by writing production-grade code, building foundational custom tooling/operators (Golang, Python), authoring reusable Terraform modules, and designing reference implementations for engineering teams.
Technical Leadership & Mentorship: Serve as a mentor and technical coach to junior and mid-level engineers. Facilitate design reviews, establish architecture RFC processes, conduct deep code/infrastructure reviews, and foster a strong culture of engineering excellence.
Container Platforms & Orchestration: Architect resilient, multi-cluster Kubernetes topologies at scale. Define best practices for workload scheduling, networking (service mesh), custom controllers/operators, GitOps workflows (ArgoCD/Flux), and multi-tenant isolation.
Enterprise CI/CD & Developer Experience (DevEx): Architect automated, self-service golden paths and progressive delivery pipelines (canary, blue-green) that minimize cognitive load for application developers while ensuring strict deployment guardrails.
Zero-Trust Security & Policy-as-Code: Embed DevSecOps at the architectural level—designing policy-as-code frameworks (OPA/Kyverno), centralized secrets management (HashiCorp Vault), secure supply chain standards (SBOM, image signing, Trivy), and compliance baselines (SOC 2, ISO 27001).
Observability, SRE & Chaos Engineering: Design holistic telemetry strategies utilizing OpenTelemetry. Define enterprise SLI/SLO standards, error budget policies, resilience testing/chaos engineering, and lead complex root-cause analyses (RCA) for critical production incidents.
FinOps & Performance Optimization: Architect cost-efficient cloud topologies. Establish capacity planning models, multi-region traffic routing, right-sizing strategies, and automated cost-governance guardrails.
Experience: 7+ years of hands-on experience in Platform Engineering, SRE, Distributed Systems, or Cloud Architecture, with a proven track record of designing and operating large-scale, production-grade distributed systems.
Architecture & System Design: Demonstrated expertise in architecting complex, highly available, fault-tolerant, and secure cloud topologies (AWS, Azure, or GCP).
Deep Hands-on Coding & Automation: Proficiency in at least one modern systems programming language (e.g., Golang, Python) along with advanced scripting (Bash). Strong foundation in data structures, algorithms, and systems software design.
Infrastructure-as-Code Mastery: Expert-level proficiency with Terraform(custom provider/module development, state locking, drift detection) and Policy-as-Code engines (e.g., OPA, Rego, Kyverno).
Kubernetes Internals: Comprehensive understanding of Kubernetes internals (control plane components, CNI/CSI, custom controllers/operators, RBAC, API extensions, and scheduling primitives).
Observability & Reliability: Hands-on experience with modern observability platforms, distributed tracing with OpenTelemetry, Prometheus/Grafana ecosystems, and enterprise SRE practices.
Experience architecting platforms within highly regulated domains (e.g., Financial Services, FinTech, Banking).
We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability, gender identity, sexual orientation, protected veteran status, or any other characteristic protected by law.
#LI-111419