Overview
Microsoft Silicon and Cloud Hardware Infrastructure Engineering (SCHIE) is the team behind Microsoft’s expanding cloud infrastructure, powering the company's Intelligent Cloud and AI platforms through industry-leading silicon, systems, and hardware innovation.
Are you passionate about building cutting-edge technologies in an environment that embraces a growth mindset? Do you want to help shape the future of AI infrastructure while contributing to Microsoft's mission to empower every person and every organization on the planet to achieve more?
The Firmware Center of Excellence (FW CoE) within SCHIE is responsible for delivering the hardware and firmware technologies that power Azure infrastructure. As part of this organization, we are seeking a Principal Firmware Engineer to lead the design and evolution of Microsoft's AI accelerator validation and performance characterization frameworks, supporting the MAIA roadmap and future generations of Azure AI silicon.
In this role, you will drive the architecture and development of scalable stress, validation and performance characterization solutions used across Microsoft's AI accelerator programs. You will work at the intersection of system architecture, firmware, software, performance engineering, AI workloads, and silicon validation to enable high-confidence silicon bring-up, characterization, and deployment at hyperscale. You will partner closely with teams spanning these disciplines to develop and operationalize validation frameworks and workload solutions that exercise compute, memory, networking, and platform infrastructure under both synthetic stress conditions and representative customer AI workloads. Your work will directly influence silicon readiness, platform quality, performance optimization, and production deployment of Microsoft's next-generation AI infrastructure.
We are looking for a technically strong leader who combines deep systems expertise with a passion for solving complex engineering problems, driving cross-organizational impact, and mentoring teams to deliver world-class solutions at cloud scale. You will be working on latest state-of-the art technologies, in a fun environment with a talented group of individuals with diverse backgrounds and skillsets and located in different geographic locations.
ResponsibilitiesDesign and development of highly-performant validation and stress workloads spanning GPU Compute engines, memory, networking, PCIe, and DMA subsystems.
Architect and develop end-to-end validation strategies from pre-silicon environments through post-silicon bring-up, characterization, and production deployment with a goal to identify hardware, firmware, and system-level reliability issues as early as possible
Develop performance analysis, telemetry, and observability solutions to measure compute utilization, memory bandwidth, network throughput, power, thermal behavior and end-to-end workload performance.
Drive adoption of profiling, performance monitoring, and other platform observability technologies to accelerate debugging, tuning, and characterization.
Partner closely with Architecture, AI software, Firmware, Silicon Validation, Manufacturing, Performance Engineering, and Cloud Infrastructure teams to influence platform requirements and readiness.
Mentor engineers across workload development, performance optimization, debugging, automation, and validation, while establishing best practices for software quality, CI/CD, telemetry, and large-scale system validation.
Qualifications
12+ years of experience developing complex software, firmware, system software, or validation frameworks.
Experience with one or more of these: Accelerators or GPUs, DMA Engines, PCIe, Memory (DDR, HBM), Networking
Proven experience in one or more areas: Silicon validation, Firmware development, Platform diagnostics, Performance engineering, Stress and reliability testing
Experience debugging issues across hardware, firmware, drivers, SDKs, and applications.
Knowledge of or experience with AI models/kernels such as GEMM, GPT, Gemma, MoE, Llama, or similar
Good understanding of AI accelerators, such as GPUs, NPUs, FPGAs, with an understanding of their architectures including Data pipelines, Data formats, Memory hierarchies including HBM/LPDDR, Tensor core architecture etc.
Ability to work closely with diverse customers and collaborators across varied disciplines (silicon architecture, FW, SW dev, validation engineers) to reconcile requirements from understanding their needs to resolving their problems.
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings: Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process.