We are looking for a highly skilled SRE Engineer to own and operate our company-wide
JFrog Artifactory ecosystem. As the primary administrator for our artifact
management platform, you will ensure high availability, security, and performance for
critical CI/CD pipelines. This role is ideal for an expert who deeply understands the
JFrog ecosystem—specifically Artifactory, Xray, and Curation—and who enjoys
optimizing complex infrastructure, automating operational toil, and leading incident
response for mission-critical core services.
You will need:
? Deep expertise in managing enterprise-grade JFrog Artifactory, Xray, and Curation
clusters.
? Strong hands-on experience with artifact scanning security policies and
infrastructure reliability.
? Proficient in Python and shell scripting, with a focus on automation and API-driven
workflows.
? Ability to manage complex distributed systems, including troubleshooting message
queues (e.g., RabbitMQ) and infrastructure scaling.
? Familiarity with on-call support, incident response, and vendor management (JFrog
support).
What You Will Do
? Manage and Architect Artifactory Infrastructure
? Deploy, scale, and maintain large-scale, distributed Artifactory, Xray, and
Curation clusters across multiple datacenters.
? Design and enforce security policies (e.g., scan-before-download) while
maintaining high availability and minimizing performance impact.
? Plan and execute version upgrades, node expansions, and infrastructure
migration (on-prem to Kubernetes/EKS).
? Manage complex dependencies, including RabbitMQ/event-streaming
backends and database synchronization.
? Build Automation and Observability
? Develop Python-based automation tools to manage artifact lifecycles,
configuration drift, and self-service repository creation.
? Implement rigorous monitoring, alerting, and observability (e.g., Elastic,
Prometheus, Grafana) to proactively detect degradation in Xray/Artifactory
health.
? Develop runbooks for datacenter failover, cluster resets, and recovery
procedures.
? Configuration Management and IaC
? Manage infrastructure as code (e.g., Terraform, Ansible, Puppet) to ensure
environment consistency across development, staging, and production.
? Enforce security baselines and configuration standards for repository
management and user access.
? Reliability Engineering and Support
? Lead incident response for Artifactory/Xray-related outages, performing root
cause analysis (RCA) and implementing long-term preventative measures.
? Partner with CI/CD and engineering teams to optimize build/test performance
and resolve artifact retrieval issues.
? Participate in an on-call rotation, providing expert-level support for
high-priority (SEV) incidents.
Minimum Qualifications
? 5+ years of experience in SRE, DevOps, or Systems Engineering roles supporting
critical infrastructure.
? Proven hands-on expertise with JFrog Artifactory, Xray, and Curation in large-scale
production environments.
? 3+ years of Python development for automation scripts, REST API integrations, and
operational tooling.
? Experience with configuration management or IaC tools (e.g., Terraform, Ansible).
? Strong Linux systems administration knowledge and understanding of distributed
systems (clustering, high availability).
? Familiarity with message queue systems (e.g., RabbitMQ) and troubleshooting
complex backend service dependencies.
? Strong problem-solving skills, ownership mindset, and clear written and verbal
communication.
Preferred Qualifications
? Hands-on experience with Kubernetes/EKS and migrating on-prem services to
containerized environments.
? Knowledge of CI/CD pipeline orchestration (e.g., Jenkins, GitLab CI, ArgoCD).
? SRE/DevOps philosophy: experience with observability, SLO/SLI definition, and
postmortem/RCA processes.
? Experience with global traffic management or datacenter failover strategies for core
services.
? Experience managing Security/Compliance requirements within an artifact
management system (e.g., managing vulnerability scan thresholds).