AI Jobs Map

TALENT Software Services · Cincinnati, OH

LLM Engineer

Remoteseniorfull timePosted 15 days ago
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

llmagentic-airagobservabilitykubernetesdockerkubeflowmlflowtransformerspytorchvllmpythontypescriptjavascriptmicroservicesci/cdgithubazuredevopsetl

Job Details

- Job Title: LLM Engineer

- Location: Cincinnati, OH

- Work Location: Remote - USA

- Duration: 1 year

- Experience Required: 8+ years

- Role Category: AI and Automation

- Education: Any degree

- Start Date: 15-September-2026

Role Overview

- Design, build, optimize, deploy, and operate Large Language Model (LLM) and Small Language Model (SLM) capabilities.

- Build secure, reliable, reusable, and enterprise-ready AI capabilities.

- Support:

- Agentic AI workflows

- AI for SDLC

- Knowledge retrieval

- Model evaluation

- Private AI hosting

- AgentOps

- Work closely with:

- Principal AI Architect

- AI Engineering Lead

- Platform Engineers

- Security teams

- Enterprise Architecture

- Product Owners

- Domain teams

Must-Have Technical Skills

LLM & Generative AI

- Large Language Models (LLMs)

- Small Language Models (SLMs)

- Prompt Engineering

- Context Engineering

- Retrieval-Augmented Generation (RAG)

- Embeddings

- Semantic Search

- Agentic AI Patterns

- Multi-Agent Workflows

- Tool Calling

- Function Calling

- Model Evaluation

- LLM Observability

Model Engineering

- Fine-Tuning

- Supervised Fine-Tuning

- LoRA

- QLoRA

- Quantization

- Distillation

- Model Compression

- Synthetic Data Generation

- Model Benchmarking

- Model Selection

- Model Routing

Model Hosting & Serving

- Private LLM Hosting

- On-Prem Model Deployment

- GPU-Based Inference

- Model Serving APIs

- High-Availability Inference

- Autoscaling

- Load Balancing

- Caching

- Batch and Real-Time Inference

AI Infrastructure & Frameworks

- Kubernetes

- Docker

- Kubeflow

- KServe

- Ray Serve

- MLflow

- Hugging Face

- Transformers

- PyTorch

- PEFT

- DeepSpeed

- NVIDIA NIM

- Triton Inference Server

- TensorRT-LLM

- vLLM

- TGI

- SGLang

Programming & Engineering

- Python

- TypeScript or JavaScript

- REST APIs

- Microservices

- CI/CD

- GitHub or Azure DevOps

- API Design

- Distributed Systems

- Cloud-Native Engineering

- Test Automation

Data & Knowledge Systems

- Vector Databases

- Knowledge Graphs

- Document Processing

- Metadata Management

- Data Pipelines

- Object Storage

- Enterprise Search

- Structured and Unstructured Data Integration

Roles & Responsibilities

LLM Application Engineering

- Build enterprise-grade LLM-powered applications and intelligent agent capabilities.

- Design reusable LLM patterns, services, APIs, and accelerators.

- Develop model interaction patterns for:

- Reasoning

- Summarization

- Classification

- Extraction

- Planning

- Decision support

- Build reusable prompt, context, retrieval, memory, and evaluation components.

- Support AI-for-SDLC agents across:

- Requirements

- Design

- Coding

- Testing

- Security Review

- Deployment

- Operations

- Convert AI use cases into scalable production solutions.

Agent Factory Intelligence Layer

- Build core intelligence services for enterprise agents.

- Develop reusable capabilities for:

- Planning

- Task decomposition

- Reasoning

- Tool usage

- Agent collaboration

- Enable agent-to-agent interaction and multi-agent orchestration.

- Integrate LLMs with:

- Agent runtimes

- Tool registries

- Workflow engines

- MCP-based gateways

- Support human-in-the-loop, approval, escalation, and feedback workflows.

- Improve agent quality, accuracy, safety, and task completion.

Prompt & Context Engineering

- Design reusable prompt engineering standards, templates, and libraries.

- Create:

- System prompts

- Task prompts

- Role prompts

- Guardrail prompts

- Evaluation prompts

- Develop context engineering strategies for better grounding, relevance, and personalization.

- Optimize:

- Token usage

- Context windows

- Memory injection

- Retrieval inputs

- Establish prompt versioning, testing, and governance practices.

Retrieval-Augmented Generation (RAG)

- Design and implement enterprise RAG architectures.

- Build retrieval pipelines using:

- Enterprise documents

- Knowledge repositories

- Structured data

- Metadata

- Optimize:

- Chunking

- Embeddings

- Indexing

- Ranking

- Reranking

- Retrieval strategies

- Improve grounding, citation quality, precision, recall, and factual accuracy.

- Build reusable retrieval services for agents and business domains.

- Partner with data and knowledge management teams to onboard trusted data sources.

LLM / SLM Model Engineering

- Evaluate, build, fine-tune, deploy, and optimize LLMs and SLMs.

- Support domain-specific model development using approved datasets.

- Build supervised fine-tuning and model adaptation pipelines.

- Apply:

- LoRA

- QLoRA

- Distillation

- Quantization

- Model compression

- Evaluate commercial, open-source, and internally hosted models.

- Select models based on:

- Accuracy

- Latency

- Cost

- Data residency

- Security

- Operational requirements

Private AI & On-Prem Hosting

- Build and support private AI capabilities for LLM/SLM hosting.

- Deploy models across:

- On-premises

- Hybrid

- Private cloud environments

- Support GPU-enabled model hosting.

- Optimize latency, throughput, concurrency, resiliency, and GPU utilization.

- Build secure inference endpoints for internal applications and agents.

- Support air-gapped and restricted AI environments.

- Partner with infrastructure and platform teams on private AI hosting.

Model Serving & Inference Optimization

- Implement scalable model serving using modern inference frameworks.

- Build high-availability inference architectures.

- Optimize:

- Token throughput

- Response latency

- Cost efficiency

- Inference performance

- Implement:

- Model routing

- Load balancing

- Caching

- Fallback strategies

- Support batch and real-time inference.

- Develop reusable deployment templates for different model families.

LLMOps / ModelOps / AgentOps

- Build operational practices for managing models and agents throughout their lifecycle.

- Implement observability for:

- Prompts

- Retrieval

- Model responses

- Latency

- Cost

- Failures

- Develop evaluation pipelines for regression testing and continuous quality improvement.

- Monitor:

- Model drift

- Response quality

- Hallucination indicators

- Safety risks

- Support CI/CD and release management for:

- Prompts

- Models

- Agents

- Retrieval pipelines

- Build dashboards and metrics for AI quality, reliability, adoption, and operational readiness.

AI Evaluation & Benchmarking

- Define and implement LLM evaluation frameworks.

- Measure:

- Accuracy

- Groundedness

- Relevance

- Hallucination rate

- Toxicity risk

- Safety compliance

- Task completion

- User satisfaction

- Build automated test suites for prompts, agents, tools, and RAG pipelines.

- Benchmark models across enterprise use cases.

- Compare cloud, open-source, and on-prem models based on performance, cost, quality, and risk.

- Establish quality gates for production AI releases.

Responsible AI, Security & Governance

- Implement Responsible AI controls in LLM applications and agent workflows.

- Develop guardrails for:

- Safe output

- Tool usage

- Data access

- Enterprise policy compliance

- Support:

- Model risk management

- Auditability

- Transparency

- Traceability

- Ensure sensitive data is handled according to security and privacy requirements.

- Collaborate with Security, Enterprise Architecture, Risk, and Compliance teams.

- Support model and agent approval and production-readiness governance.

Preferred / Additional Experience

- Experience deploying open-source models such as:

- Llama

- Mistral

- Mixtral

- Phi

- Gemma

- Qwen

- DeepSeek

- Granite

- Falcon

- Domain-specific models

- Experience with GPU infrastructure such as:

- NVIDIA H100

- H200

- B200

- B300

- A100

- L40S

- GH200

- AMD MI300X

- Experience with:

- Private AI

- Hybrid AI

- Air-gapped AI environments

- Experience in regulated industries such as:

- Healthcare

- Financial Services

- Insurance

- Experience building:

- Enterprise copilots

- AI assistants

- Agent platforms

- Experience with:

- MCP

- Tool registries

- Agent runtimes

- Enterprise integration patterns

- Experience with Responsible AI, model governance, model risk management, and AI compliance.

- Experience optimizing AI workloads for:

- Cost

- Performance

- Latency

- Security

Role / Skills Summary

- Role: LLM Engineer

- Essential Skill: LLM

- Primary Skill: AI and Automation

- Experience: 8-10+ years

- Work Location: Remote USA

- Duration: 1 year

- Core Technologies: LLM, SLM, RAG, Generative AI, Agentic AI, Python, Kubernetes, MLflow, Hugging Face, PyTorch, vLLM, KServe, Docker, Vector Databases

More jobs at TALENT Software Services