We are seeking a highly motivated Software Engineer II to transform and modernize the hardware break/fix operations that support one of the world's largest cloud storage fleets.
This role is not a traditional operations position. The successful candidate will initially immerse themselves in the end-to-end hardware break/fix lifecycle, partnering with engineering, datacenter operations, vendors, and support organizations to understand current workflows, operational pain points, and reliability risks. They will then drive automation and AI-driven solutions to eliminate manual effort, accelerate hardware recovery, reduce operational risk, and improve service reliability.
This is a unique opportunity to work at the intersection of cloud infrastructure, software engineering, automation, AI, and operational excellence.
Own and Improve Hardware Break/Fix Workflows
- Develop deep expertise in the end-to-end hardware break/fix process for a large-scale cloud storage fleet.
- Coordinate and support hardware replacement activities involving storage controllers, motherboards, disks, networking components, and related infrastructure.
- Partner with engineering, datacenter operations, support teams, and hardware vendors to identify operational bottlenecks and opportunities for improvement.
- Analyze incidents and recurring failure patterns to improve repair workflows and fleet reliability.
Drive Automation and Modernization
- Design and build automation solutions that eliminate manual operational tasks.
- Develop tools and services that automate hardware diagnostics, repair orchestration, case management, approvals, escalations, and reporting.
- Create self-service platforms and workflow automation that improve operational efficiency and reduce mean time to repair (MTTR).
- Integrate data from multiple operational systems to provide a unified view of hardware health and repair status.
Leverage AI to Transform Operations
- AI and machine learning technologies to improve operational decision making.
- Build intelligent systems that assist with:
- Incident triage
- Failure prediction
- Repair prioritization
- Root cause analysis
- Knowledge management
- Automated workflow execution
- Explore and implement generative AI solutions to improve troubleshooting and operational productivity.
Engineering and Software Development
- Design, build, test, and maintain software services and automation frameworks.
- Develop scalable backend services, APIs, dashboards, and automation pipelines.
- modern software engineering practices including CI/CD, testing, observability, and reliability engineering principles.
- Write high-quality, maintainable, and production-ready code.
Reliability and Operational Excellence
- Identify systemic risks across the storage fleet and drive engineering solutions to address them.
- Develop metrics and dashboards to measure operational health, repair performance, automation effectiveness, and reliability outcomes.
- Partner with reliability engineering teams to drive continuous improvements in service availability and customer experience.
- Contribute to incident reviews and convert operational learnings into engineering investments.
Preferred Qualifications
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Experience with infrastructure automation tools and scripting.
- Familiarity with distributed systems, storage systems, networking, or datacenter infrastructure.
- Experience with AI, machine learning, LLMs, or intelligent automation platforms.
- Experience with reliability engineering, DevOps, SRE, or production engineering.
- Experience building dashboards, analytics solutions, or workflow orchestration systems.
- Knowledge of incident management and operational excellence practices.
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, or related technical field.
- 2+ years of software development experience.
- Strong programming skills in one or more languages such as Python, C#, Java, Go, or similar.
- Experience developing automation solutions, backend services, or operational tooling.
- Strong problem-solving and debugging skills.
- Ability to learn and understand complex operational systems and workflows.
- Excellent communication and collaboration skills.
At NetApp, we embrace a hybrid working environment designed to strengthen connection, collaboration, and culture for all employees. This means that most roles will have some level of in-office and/or in-person expectations, which will be shared during the recruitment process.
Equal Opportunity Employer:
NetApp is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination based on age, race, color, gender, sexual orientation, gender identity, national origin, religion, disability or genetic information, pregnancy, protected veteran status, and any other protected classification.
Why You'll Thrive at NetApp
At NetApp, you won't wait for the perfect moment—you'll make it. The early planning, the extra thought, the bold idea that turns good into great: That's how our people operate and how we continue to push the boundaries of data infrastructure.
NetApp is the trusted partner for organizations transforming data into opportunity. As the only enterprise-grade storage service natively embedded in Google Cloud, AWS, and Microsoft Azure, we empower customers to run everything from traditional workloads to enterprise AI with unmatched performance, resilience, and security.
Our culture
We celebrate mold breakers, bold thinkers, and problem solvers. We reward initiative, impact, and ownership. We provide flexibility so you can balance professional ambition with your personal life. Here, differences are not just welcomed—they drive everything we do.
If you're ready to innovate, rise to the challenge, and own every moment - make your next move your best one. Apply now.
Submitting an Application
To ensure a streamlined and fair hiring process for all candidates, our team only reviews applications submitted through our company website. This practice allows us to track, assess, and respond to applicants efficiently. Emailing our employees, recruiters, or Human Resources personnel directly will not influence your application.
AI Disclosure
For select roles, some stages of our hiring process may use artificial intelligence tools to help evaluate applications and candidate selection. These tools support—rather than replace—human decision-making.