As an Engineering Manager - Control Plane, you will lead the design and development of Manageability solutions for the DDN Infinia AI Data Platform. This role is responsible for building foundational capabilities that enable centralized control, automated operations, and intelligent support across large-scale hybrid (OnPrem + cloud) environments. You will lead a team delivering cloud-native, API-first, and AI/ML-powered systems that ensure operational excellence, proactive incident management, and seamless user experiences at petabyte scale. This is a ground-up platform leadership role focused on scalability, reliability, automation, and innovation.
Lead a high-performing engineering team across distributed systems, cloud infrastructure, and AI/ML.
Collaborate with cross-functional teams (product, engineering, SRE, security, and customer
success) to align platform capabilities with business and customer needs.
Establish engineering best practices, development standards, and operational excellence
frameworks.
Implement policy-driven infrastructure management and Infrastructure-as-Code (IaC)
frameworks.
Develop self-service tooling and role-based access control (RBAC) for enterprise customers.
Design API-first management interfaces for integration with external tools and automation
workflows.
Drive proactive capacity planning and performance optimization for large-scale deployments.
Build self-healing systems that reduce manual intervention and improve system resilience.
Develop predictive analytics capabilities for capacity planning, performance forecasting, and
failure prevention.
Integrate intelligent recommendations and prescriptive insights into operational workflows.
Define and enforce an API-first, cloud-native architecture across all components.
Ensure systems are highly scalable, resilient, secure, and capable of operating at petabyte
scale.
Promote automation-first principles across development, testing, deployment, and
operations.
Oversee the design of distributed systems with high availability and fault tolerance.
15+ years of experience in software engineering, distributed systems, or cloud platforms
5+ years in technical leadership or management roles
Proven experience building large-scale platform management or infrastructure systems
Strong background in distributed systems architecture and cloud-native technologies
Experience with APIs, microservices, and infrastructure-as-code (IaC)
Familiarity with AI/ML concepts applied to operational analytics or automation• Experience managing teams delivering production-grade, enterprise-scale systems
Experience in storage systems, data platforms, or high-performance computing
environments
Background in building AI-driven operations or AIOps platforms
Experience with hybrid cloud and OnPrem deployments
Knowledge of security, compliance, and enterprise governance requirements
Familiarity with DevOps, SRE practices, and CI/CD pipelines