We are seeking a Senior AI Research Scientist with deep expertise in modern foundation models and advanced neural-network architectures. The ideal candidate understands large language models, diffusion and flow-based generative models, state-space models (SSMs), Transformer-SSM hybrids, mixture-of-experts systems, synthetic-data training, and distributed neural-network inference.
This is a hands-on research role for someone who can move comfortably between mathematical theory, rigorous experimentation, prototype implementation, and production collaboration. You will investigate new architectures and training methods while helping build systems that operate efficiently across GPUs, CPUs, NPUs, personal computers, edge devices, private servers, and distributed clusters.
Your work should lead to measurable improvements in model quality, reasoning, latency, throughput, memory use, energy efficiency, privacy, reliability, and total cost of operation. You will have significant influence over product's long-term technical direction and research roadmap.
Key Research Areas
The role will contribute across several of the following areas, with deep specialization expected in at least two:
- Large language models and multimodal foundation models
- Transformer alternatives and attention-efficient architectures
- State-space models, including selective SSMs and Mamba-style architectures
- Hybrid Transformer-SSM, recurrent, sparse, and modular model designs
- Diffusion, discrete diffusion, flow matching, and multimodal generative systems
- Mixture-of-experts models, expert routing, modular networks, and conditional computation
- Synthetic data, model-generated supervision, self-training, and knowledge distillation
- Distributed inference across heterogeneous and intermittently available devices
- Memory-efficient inference, KV-cache management, long-context execution, and model sharding
- Quantization, sparsity, pruning, low-rank adaptation, and dynamic adapter routing
- Agentic models, tool use, planning, reasoning, and multi-agent coordination
- Continual learning, personalization, privacy-preserving learning, and edge AI