AI Jobs Map

NEXadept · Singapore, Singapore

LLM Inference Performance Engineer

seniorfull timePosted 3 days ago
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

llmcudac++python

We are looking for an LLM Inference Performance Engineer to optimize large-scale LLM inference performance across TPU/GPU and other AI accelerators.

You will work across LLM inference, kernels, compilers, and runtime systems, improving latency, throughput, scalability, and overall inference efficiency.

Responsibilities

- Optimize LLM inference performance on TPU/GPU and other AI accelerators

- Develop and optimize inference backends, kernels, and runtime components

- Optimize Attention, GEMM, KV Cache, Sampling, fused kernels and other performance-critical workloads

- Work with technologies such as JAX, XLA, Pallas, CUDA, Triton or related frameworks

- Build benchmarking and profiling tools to identify performance bottlenecks

- Collaborate with model, inference, compiler, and hardware teams to improve production performance

Requirements

- Bachelor’s degree or equivalent experience in CS, Engineering, ML, Systems, or related fields

- Experience in LLM inference, ML systems, hardware acceleration, or performance optimization

- Strong programming skills in C++ and/or Python

- Understanding of GPU/TPU architecture, memory behavior, kernels, or ML workload performance

- Experience with performance profiling and benchmarking

More jobs at NEXadept