Senior AI Engineer – On-Premise & Air-Gapped LLM Systems
Location: Fully remote – United States
Salary: $120,000 - $130,000 base salary plus bonus and benefits
Travel: Up to 15%
Employment: Permanent, full-time
The opportunity
We are working confidentially with an innovative US technology organisation that creates immersive, AI-powered training and simulation products for customers operating in secure and high-stakes environments.
They are looking for a hands-on Senior AI Engineer to take ownership of building and deploying generative AI systems on private, locally managed GPU infrastructure.
This is not a role focused solely on consuming third-party APIs or connecting applications to hosted models. You will be responsible for building AI solutions that can operate securely and independently within on-premise and air-gapped environments.
You will work closely with senior technical leadership, backend engineers, real-time development teams and product specialists to take AI solutions from early prototype through to production deployment.
What you’ll be doing
- Develop and maintain an on-premise LLM technology stack.
- Evaluate and select models based on performance, hardware and product requirements.
- Deploy, optimise and manage models across local GPU infrastructure.
- Apply quantisation and inference optimisation techniques.
- Build production RAG pipelines against specialist and proprietary data.
- Design retrieval, chunking and evaluation strategies that improve accuracy.
- Establish practical methods for measuring and reducing hallucinations.
- Build secure AI solutions capable of operating without cloud connectivity.
- Develop local speech pipelines covering automatic speech recognition and text-to-speech.
- Optimise AI systems for latency, natural interaction and concurrent users.
- Create integration layers between AI models and wider software products.
- Support real-time interactive, training and simulation experiences.
- Help define AI engineering standards, governance and responsible-use practices.
- Work directly with technical leadership to shape the organisation’s wider AI strategy.
What we’re looking for
- Approximately 3–5 + years of experience across machine learning, AI engineering, automation or technical scripting.
- At least 1–2 years of recent hands-on generative AI experience.
- Proven experience deploying and managing a production on-premise or air-gapped LLM system.
- Strong understanding of local model deployment, model selection, quantisation and inference optimisation.
- Experience with inference frameworks such as vLLM, llama.cpp, TGI or comparable technologies.
- Practical experience managing AI workloads on local GPU infrastructure.
- Production experience building RAG systems against custom data.
- Knowledge of retrieval evaluation, prompt design, chunking and hallucination measurement.
- Strong Python and software-engineering fundamentals.
- A builder’s mentality and the ability to prototype and solve complex technical problems personally.
Cloud-only AI experience will not be sufficient for this position.
Desirable experience
- Fine-tuning or training machine-learning and generative-AI models.
- Local ASR and TTS technologies, including platforms such as Whisper or comparable open-source tooling.
- Stable Diffusion or other generative media technologies.
- Multi-tenant LLM architectures.
- Real-time applications, simulation platforms or game-engine-adjacent products.
- Secure deployments within defence, education or other regulated environments.
- Experience supporting products serving multiple simultaneous AI interactions.
Why join?
- Take ownership of a technically ambitious AI platform.
- Work directly with experienced, hands-on technology leadership.
- Build genuine private AI infrastructure rather than API-only integrations.
- Deliver AI systems used in secure, real-world training and simulation environments.
- Influence technical architecture, standards and longer-term AI strategy.
- Fully remote working with occasional travel for project installations.
- Benefits include medical, dental and vision insurance, 401(k) and bonus eligibility.
If you have personally built and deployed production LLM solutions on private GPU infrastructure, I would be keen to hear about what you built, the models and inference stack you selected, and how you approached performance, security and accuracy.