A
Senior Research Engineer
AssemblyAI · Remote, Nigeria · Remote
Salary
$33750k – $38750k
Location
Remote
Posted
7 hours ago
Required skills
About the role
Senior Research Engineer
AssemblyAI is seeking a Senior Research Engineer to join our Research team, developing and improving the systems behind large-scale distributed training, data processing, and inference. This role is critical in raising the team's experimental velocity and ensuring that our models are accurate, efficient, and scalable.
Responsibilities
- Raise the team's experimental velocity — make it faster to launch an experiment job, get a number back you can trust, and know what to try next.
- Maintain and evolve our JAX training framework, keeping it scalable and efficient for large-scale distributed training runs on TPU.
- Improve the data our models learn from: investigating quality issues, building the tooling to surface them, and turning what you find into measurable accuracy gains.
- Analyze the accuracy of production models, build evaluation harnesses, and work out which improvements will matter most to customers.
- Translate research prototypes into production-ready systems, refactoring and modernizing model architectures and infrastructure along the way.
- Optimize production inference for speech language models, both from a serving architecture perspective and through advanced techniques such as quantization and speculative decoding.
- Investigate and resolve performance bottlenecks across the stack, from low-level kernels (XLA, Pallas) to high-level system design.
- Partner with researchers, infrastructure, and production engineering to trace problems to their real source and ship fixes that hold.
Requirements
- Expert-level proficiency with JAX and TPUs, including the surrounding ecosystem (Flax, Optax, the XLA compilation pipeline).
- Measurement discipline: You define what success looks like before you start, you stay skeptical of your own results until they hold up, and you treat an unexplained improvement as a problem rather than a win.
- Appetite for the whole pipeline: Your core strength might be JAX and TPU performance, but when a customer issue traces back to a data problem or an evaluation blind spot, you want to go find it yourself.
- Strong experience optimizing inference systems for production, ideally with LLMs or speech models.
- Deep understanding of distributed training at scale, modern deep learning systems, and ML infrastructure best practices.
- Familiarity with modern inference optimization techniques: continuous batching, KV-cache management, sharding strategies, quantization.
- Enthusiasm for refactoring and improving existing systems — you thrive in environments where you're constantly improving and refining.
How to Apply Apply directly via our careers page.
Apply NowThis role is hiring worldwide — you finish on We Work Remotely, and the application is saved to your Aremu account.