MulticoreWare is a global software solutions & products company with its HQ in San Jose, CA, USA. With worldwide offices, it serves its clients and partners in North America, EMEA and APAC regions. Started by a group of researchers, MulticoreWare has grown to serve its clients and partners on HPC & Cloud computing, GPUs, Multicore & Multithread CPUS, DSPs, FPGAs and a variety of AI hardware accelerators.
MulticoreWare was founded by a team of researchers that wanted a better way to program for heterogeneous architectures. With the advent of GPUs and the increasing prevalence of multi-core, multi-architecture platforms, our clients were struggling with the difficulties of using these platforms efficiently.
We started as a boot-strapped services company and have since expanded our portfolio to span products and services related to compilers, machine learning, video codecs, image processing and augmented/virtual reality. Our hardware expertise has also expanded with our team; we now employ experts on HPC and Cloud Computing, GPUs, DSPs, FPGAs, and mobile and embedded platforms. We specialize in accelerating software and algorithms, so if your code targets a multi-core, heterogeneous platform, we can help.
We are seeking a talented engineer to implement and optimize machine learning, computer vision,
and numeric libraries for target hardware architecture, including CPUs, GPUs, DSPs, and other
accelerators. Your expertise will be instrumental in enabling efficient and high-performance
execution of algorithms on these hardware platforms.
Key Responsibilities:
Implement and optimize machine learning, computer vision, and numeric libraries for
target hardware architectures, including CPUs, GPUs, DSPs, and other accelerators.
- Work closely with software and hardware engineers to ensure optimal performance on
target platforms.
- Implement low-level optimizations, including algorithmic modifications, parallelization,
vectorization, and memory access optimizations, to fully leverage the capabilities of the
target hardware architectures.
- Work with customers to understand their requirements and implement libraries to meet
their needs.
- Develop performance benchmarks and conduct performance analysis to ensure the
optimized libraries meet the required performance targets.
- Stay current with the latest advancements in machine learning, computer vision, and
high-performance computing.
Qualifications:
BTech/BE/MTech/ME/MS/PhD degree in CSE/IT/ECE- > 4 years of experience working in Algorithm Development, Porting, Optimization &
Testing
- Proficient in programming languages such as C/C++, CUDA, OpenCL, or other relevant
languages for hardware optimization.
- Hands-on experience with hardware architectures, including CPUs, GPUs, DSPs, and
accelerators, and familiarity with their programming models and optimization techniques.
- Knowledge of parallel computing, SIMD instructions, memory hierarchies, and cache
optimization techniques.
- Experience with performance analysis tools and methodologies for profiling and
optimization.
- Knowledge of deep learning frameworks and techniques is good to have
- Strong problem-solving skills and ability to work independently or within a team.
Algorithm Engineer (CPU/GPU/DSP/FPGA)
Job Summary
We are looking for an Algorithm Engineer with experience in optimizing software for CPUs, GPUs, DSPs, and FPGAs. The role involves developing high-performance algorithms for machine learning, computer vision, and numerical applications.
Key Responsibilities
-
Develop and optimize algorithms for CPUs, GPUs, DSPs, and FPGAs.
-
Work on machine learning, computer vision, and numerical libraries.
-
Improve application performance through optimization techniques.
-
Work closely with software and hardware teams.
-
Analyze performance and fix bottlenecks.
-
Support customer requirements and deliver optimized solutions.
Required Skills
-
B.E./B.Tech./M.E./M.Tech./MS in Computer Science, Electronics, ECE, or related field.
-
4+ years of experience in algorithm development or software optimization.
-
Strong programming skills in C/C++.
-
Experience with CUDA, OpenCL, or similar parallel programming technologies.
-
Experience working with CPUs, GPUs, DSPs, or FPGAs.
-
Hands-on experience with FPGA development using VHDL, Verilog, HLS, Xilinx Vivado/Vitis, Intel Quartus, or similar tools.
-
Good understanding of performance optimization, parallel computing, and memory optimization.
-
Good problem-solving and debugging skills.
Good to Have
-
Experience with TensorFlow, PyTorch, or other AI/ML frameworks.
-
Knowledge of LLVM, MLIR, or TVM.
-
Experience in embedded systems or AI acceleration.
This version is much simpler, recruiter-friendly, and broad enough to attract candidates with either Algorithm Optimization or FPGA acceleration experience.