Lead AI Engineer
Hybrid | Lewisville, TX
Permanent Role
The client is seeking a Lead AI Engineer to serve as the hands-on technical lead for AI-accelerated engineering within its Data, Analytics & AI organization.
This role will develop reusable AI skills, workflows, and standards that can be adopted across engineering teams.
The position is a senior individual contributor role with no direct reports, focused on building, operating, and scaling AI-assisted development practices.
This is not a data engineering or model-training role. The focus is on AI-enabled software development, agentic workflows, automation, governance, and operational excellence.
Responsibilities:
- Develop and maintain reusable AI skills, prompts, and workflows that align with the client's engineering standards.
- Design and operate AI-assisted and agentic workflows for developing, refactoring, testing, and supporting data and software products.
- Establish and manage the AI development lifecycle, including versioning, evaluation, deployment, monitoring, and continuous improvement.
- Build automated testing, evaluation, and quality gates into AI-assisted development workflows.
- Define appropriate levels of AI autonomy and implement human review for critical or irreversible actions.
- Lead adoption of AI engineering practices across development teams and mentor engineers on prompt engineering and effective use of AI.
- Maintain shared, version-controlled AI capabilities and prioritize requests from engineering teams.
- Develop AI-driven capabilities that turn platform data and telemetry into actionable insights and recommendations.
- Enable secure, natural-language access to governed data using semantic layers and trusted data sources.
- Monitor the reliability, performance, cost, and observability of LLM-driven workflows.
- Implement guardrails, prompt-injection defenses, data-leakage protections, access controls, and audit trails.
- Support ML model operations, including versioning, validation, serving, and drift monitoring.
- Optimize AI workloads through model routing, caching, right-sizing, and token/cost management.
- Track adoption and measurable improvements in delivery speed, quality, and cost.
Required Skills & Qualifications:
- Bachelor's degree in Computer Science, Data Science, Artificial Intelligence, Engineering, or a related field; master's preferred.
- 8+ years of experience in software, data, or ML engineering, including recent hands-on experience delivering AI or agentic systems in production.
- Strong Python and modern software engineering skills with experience building production-grade, tested, and version-controlled applications.
- Hands-on experience with AI-assisted development, agentic workflows, prompt engineering, and reusable AI skill development.
- Experience with LLM/agentic technologies such as LangChain, LangGraph, LlamaIndex, MCP, retrieval, and vector databases.
- Experience with model serving, orchestration, evaluation, and observability tools such as MLflow, vLLM, Langfuse, Phoenix, or OpenTelemetry.
- Experience with containerization and orchestration technologies such as Docker and Kubernetes, as well as API frameworks such as FastAPI.
- Knowledge of model, inference, prompt, and context optimization, including caching, batching, model routing, quantization, and latency/throughput optimization.
- Experience managing AI and data workload costs, including token budgeting, compute rightsizing, cost monitoring, and FinOps practices.
- Experience deploying and operating AI/LLM workflows end to end, including CI/CD, testing, evaluation, monitoring, and human review controls.
- Strong understanding of AI governance, security, data privacy, and production reliability.
- Demonstrated technical leadership and mentoring experience across multiple engineering teams.
- Strong communication, ownership, problem-solving, and ability to work effectively in an ambiguous environment.
Preferred Experience:
- Experience with Snowflake, Azure, Cortex, Snowpark ML, or Snowpark Container Services.
- Experience with dbt, Coalesce, Apache Airflow, Apache Spark, or Apache Iceberg.
- Experience operating open-weight or self-hosted models alongside managed AI services.
- Experience building reusable AI skill libraries and engineering standards in Git.
- Experience implementing AI governance and security controls in an enterprise or regulated environment.
- Familiarity with PyTorch, TensorFlow, scikit-learn, XGBoost, LightGBM, Ray, or Triton.
- Experience building internal AI platforms or enablement capabilities used across multiple engineering teams.