AI Jobs Map

Rebellions · Seongnam, Gyeonggi, South Korea

Infrastructure Support Lead

directorfull timePosted 6 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

pytorchvllmartificial-intelligenceembedded-systemslinuxkubernetespythonllm

ABOUT REBELLIONS

Rebellions builds high-performance AI systems with full-stack software natively supporting PyTorch and vLLM – familiar, open, and ready for production. Our products deliver the best performance per dollar per watt, purpose-built for inference from the ground up, not retrofitted from training. With proven commercial deployments already live across enterprises and governments, and a chiplet architecture built for the most demanding AI workloads, Rebellions is making good on a simple belief: every organization and nation should be able to own and control their AI, not just access it.

About the Job

Rebellions' systems are in production at customer sites, and the installed base is growing across multiple regions. We are hiring a support lead to own technical support for that customer installed base.

The role has two sides. You will be our support presence during Asia hours — the person a customer in the region reaches, and the one who drives an issue to resolution or routes it to the right engineer. Alongside that, you will help design and stand up the support structure underpinning Rebellions' global expansion: the operating model, escalation paths, service workflows, and the processes connecting customers, partners, and internal teams.

This is a highly cross-functional role. You will work across our global delivery and support partners, and with technical and business teams inside Rebellions, from escalation through to resolution.

Responsibilities and Opportunities

- Manage the support team and hold delivery to a consistent standard

- Serve as the senior technical point of contact for customer issues, driving them to resolution or routing them to the right engineer. Supporting a global installed base requires flexibility across time zones, including on-call and out-of-hours escalation coverage

- Build and operate an AI-forward support system: intake, triage, escalation paths, and the workflows connecting customers, partners, and technical and business teams

- Own the internal and external knowledge base — runbooks, troubleshooting guides, and customer-facing documentation — and drive deflection and first-contact resolution

- Implement and operate ticketing and field service management tooling

- Coordinate daily across our global support partners and with technical and business teams inside Rebellions

- Drive hardware service workflows including RMA, spares, and dispatch

Key Qualifications

- Over 8 years in technical support, field service, or service delivery for hardware deployed at customer sites

- Over 2 years working directly with AI/ML infrastructure or accelerated compute environments

- Over 2 years leading a support or field engineering team, including engineers you did not directly employ

- Able to lead and drive debugging on a technical bridge call: read logs, reason about the hardware/software boundary, and judge whether an issue is customer configuration, firmware, or silicon

- Hands-on depth in Linux, Kubernetes, and containerized inference serving, along with Python for automation and tooling

- Able to explain technical AI concepts clearly to non-technical stakeholders, both customers and internal teams

- Experience in an externally facing customer support or service function handling both hardware faults and software or performance issues, routing each correctly — in support of external customers, not internal IT users

- Has owned major incidents for enterprise or government customers

- Experience with hardware service operations — installed base and entitlement, warranty, RMA, spares, dispatch — either end-to-end, or substantial exposure to several parts with the appetite to take on the rest

- Experience building a support process or function, not only operating one someone else designed

- Experience delivering service through third parties: helping set the standard and holding partners to it

Ideal Qualifications

- AI/ML, HPC, or data center infrastructure background — accelerators or GPU systems, Linux, Kubernetes, inference serving

- Distributed LLM inference orchestration — vLLM disaggregated serving, NVIDIA Dynamo, llm-d, or equivalent KV-cache-aware systems

- Python and PyTorch; inference optimization and performance tuning

- Diagnosing multi-layered performance bottlenecks spanning software, middleware, networking, and infrastructure

- Data center integration experience — rack, power, cooling, and network requirements

- Has implemented ticketing and field service management systems, not just worked in them

- Experience with formal knowledge management practices such as KCS

- AI-assisted support operations — deflection, agent assist, automated triage

- Support for air-gapped or otherwise access-restricted sites

- Cross-border movement of controlled hardware for repair and replacement

- Service delivery in structurally different markets, including at least one emerging or access-restricted environment

- Experience shaping product supportability requirements before release

Location & Language Requirements

- Based in Korea; up to 25% international travel

- Professional Korean fluency required

- Professional English fluency required

Rebellions is committed to fostering a diverse and inclusive workplace. We are an equal opportunity employer and value diversity within our company. We do not discriminate based on personal identity. Applicants who would like to contact us regarding the accessibility of our website or who need special assistance or a reasonable accommodation for any part of the application or hiring process may contact us at: [email protected].

More jobs at Rebellions