Gen AI Inferencing Engineer — "Infra/MLOps-first engineer"
Looking for: an ML Platform/infra engineer, not primarily an app developer.
- Background running models in production at scale — deployment via vLLM or Triton Inference Server, containerized (Docker/K8s), with real throughput/latency tuning experience
- Strong MLOps chops: CI/CD for ML pipelines, fine-tuning workflows, inference framework internals
- Comfortable owning infrastructure other data science teams build on top of (shared tooling, not one-off notebooks)
- RAG knowledge is a plus but secondary — they should be stronger on "how do I serve this efficiently" than "how do I build the retrieval logic"
- Good fit: someone from an ML platform, SRE-for-ML, or MLOps background who's touched GenAI serving specifically (not just traditional ML model serving)