Job Description: Software Engineer – Infrastructure Platform (Puppet & DDI)Role Summary
We are seeking an experienced Platform Engineer to lead the modernization of our core
infrastructure services. You will be responsible for redesigning our global Puppet
ecosystem—migrating from legacy Puppet 5 to Puppet 7+—while simultaneously architecting our
next-generation IPAM/DDI solution. You will work across hardware, networking, and platform teams
to build scalable, automated, and reliable infrastructure services. This role requires strong Python
development skills to build robust, API-driven tools that orchestrate our infrastructure workflows.Key
ResponsibilitiesPuppet Ecosystem & Configuration Management
? Redesign & Consolidate: Lead the architectural redesign of the Puppet ecosystem, migrating
legacy Puppet 5 systems to Puppet 7+.
? Technology Consolidation: Consolidate multiple configuration management technologies
(Ansible, Puppet) into a unified, enterprise-standard, and at-scale solution.
? IaC Implementation: Maintain reusable playbooks/modules and templates to enforce
configuration standards, security baselines, and repeatable deployments across
environments.
DDI & Network Services
? IPAM/DDI Architecture: Design and implement the corporate IP Address Management
(IPAM) and DDI (DNS, DHCP, IPAM) platforms.
? Lifecycle Management: Design scalable strategies for IP allocation, DNS, and DHCP across
multiple datacenters and environments, including full lifecycle management (assignment,
tracking, reclamation, and quarantine).
? Integration: Integrate IPAM/DDI systems with provisioning, bootstrapping, and CMDB
workflows, ensuring data accuracy and automation-driven updates.
Python Automation & Platform Integration
? API-First Development: Develop enterprise-grade Python tools and APIs to orchestrate
infrastructure workflows (compute, storage, network, cloud).
? Self-Service Workflows: Build automation to support self-service workflows, reducing
manual operational tasks and improving deployment script efficiency.
? Orchestration: Integrate internal/external APIs to facilitate automated service provisioning
and dynamic IP/resource management.
Observability & Reliability
? Monitoring: Implement and improve monitoring, alerting, and dashboards for infrastructure
health (e.g., Prometheus, Grafana, ELK/Nagios).
? Incident Response: Participate in incident response, perform root cause analysis, and
implement long-term fixes/runbooks. Define and track key metrics (availability, latency,
capacity) to drive data-based improvements.
Required Skills & ExperienceCore Technical Skills
? Python Programming: 5+ years of enterprise software development experience, focusing on
automation scripts, internal tooling, and REST API integrations.
? Puppet Expertise: Deep, hands-on knowledge of Puppet 5/7 ecosystems, including both user
and development experience.
? Network (DDI): Strong enterprise-scale knowledge of IPAM, DNS, and DHCP. Experience in
designing and maintaining DDI solutions as code.
? Configuration Management: Strong experience with IaC tools such as Ansible, Terraform, or
Puppet.
? Infrastructure: Proficiency with Linux systems administration, bare-metal provisioning, and
hypervisor deployment. Familiarity with CMDB systems as a source of truth.
Collaboration & Mindset
? Strong operational mindset with a focus on reliability, correctness, and supportability.
? Experience working cross-functionally with hardware, software, storage, and SRE teams.
? Ability to drive ambiguous technical problems to clear, automated, and scalable solutions.
Plus Points
? CI/CD Pipeline Experience: Exposure to tooling such as Jenkins, GitHub Actions, ArgoCD, or
GitLab CI.
? SRE/DevOps: Background in SRE methodologies, observability, and incident management
(e.g., Incident Commander systems).
? Cloud & Kubernetes: Experience with large-scale internal platforms, Kubernetes deployment,
and hybrid/multi-datacenter environments.
? Infrastructure as Product: Experience operating self-service platforms and developer
enablement.