Who we are
Xenoss is an AI engineering and integration services company, helping medium to large enterprises run AI transformation end-to-end, from situation analysis and goals framing to data discovery and preparation, pipeline building, model development, retraining pipeline design, solution deployment, and support.
We build a broad spectrum of AI solutions such as user behaviour prediction, content generation, NLP, audience segmentation, pathfinding solutions, AI assistants, edge computer vision, fraud detection, and others.
We work with prominent companies such as Microsoft, Toshiba, AstraZeneca, Activision Blizzard, Verve Group, Voodoo Games, and Telefonica, among others.
We’re included in the top 100 software companies on the Inc. 5000 list.
What is the project
We’re hiring a Staff AI Engineer / AI Solution Architect to lead the AI architecture of a long-term In-Call Assistant initiative for a world-leading financial services company.
The project focuses on building a real-time conversational AI system that supports front-office employees during live customer conversations. The system identifies customer needs, objections, buying signals, and required process steps, and provides concise, context-aware recommendations.
The solution combines low-latency signal detection, context preparation, specialist recommendation generation, RAG over approved product and policy knowledge, confidence management, and compliance guardrails.
You will help define how the AI architecture, models, evaluation framework, and feedback loops are designed and evolved from the initial offline version to live production use
What will you do
You’ll lead the applied AI architecture across the In-Call Assistant lifecycle, from data and taxonomy design to model training, evaluation, and production readiness.
Core work includes:
- Designing the end-to-end AI architecture for the In-Call Assistant
- Defining signal and trigger taxonomies for live conversations
- Designing training strategies for signal detection and specialist recommendation models
- Shaping data preparation, annotation, and SME validation workflows
- Evaluating fine-tuning, post-training, RAG, and hybrid approaches
- Designing low-latency signal detection, routing, context preparation, and confidence management
- Designing evaluation frameworks, golden datasets, and model improvement cycles
- Defining grounding, guardrails, abstention, and policy-compliance behavior
- Making trade-offs between model quality, latency, cost, explainability, and governance
- Partnering with AI engineers, data engineering, MLOps, and client SMEs
Technology landscape
You’ll operate across the modern applied AI and ML ecosystem, including:
- LLM and smaller-model training for conversational AI
- SFT, DPO / preference optimization, LoRA / QLoRA, and PEFT
- PyTorch and Hugging Face ecosystem
- Signal extraction and multi-label classification
- RAG and knowledge-grounded recommendation generation
- Embeddings, retrieval, and context preparation
- Low-latency model serving and inference optimization
- Golden dataset creation, annotation, and SME validation
- Model evaluation, confidence calibration, and error analysis
- MLOps, monitoring, feedback loops, and model governance
Scope of ownership and delivery context
At Staff/Architect level, you’ll own the applied AI architecture and evaluation strategy for a complex enterprise AI program.
Core ownership
- Define the AI approach for the conversation intelligence PoC
- Establish the event / intent / insight taxonomy
- Define the golden dataset strategy and annotation workflow
- Establish evaluation frameworks and acceptance criteria
- Drive trade-offs between accuracy, explainability, latency, cost, and governance
- Decide which modeling approaches are appropriate for each use case
- Act as an escalation point for AI architecture, evaluation, and data strategy decisions
Team and delivery context
- Work within a cross-functional team spanning AI engineering, data engineering, MLOps, solution architecture, and client stakeholders
- Partner with domain SMEs on taxonomy, labeling, and validation
- Mentor engineers working on extraction, evaluation, and data pipelines
- Translate ambiguous business use cases into testable AI hypotheses and validation plans
What should you bring
Must have
- Strong hands-on experience with applied AI / ML systems in production-oriented environments
- Experience with NLP, conversational AI, or transcript-based intelligence systems
- Ability to design evaluation frameworks, not just run experiments
- Experience building or validating structured datasets from unstructured text
- Strong understanding of LLM-based extraction, classification, RAG, and fine-tuning trade-offs
- Practical knowledge of classical ML or predictive modeling
- Understanding of probability-based prediction, calibration, and outcome evaluation
- Comfort working with messy enterprise data and incomplete labels
- Ability to communicate with both technical teams and business stakeholders
- Strong ownership of ambiguity, scope control, and PoC validation strategy
Nice to have
- Financial services domain exposure
- Experience with sales, call center, or customer conversation analytics
- Speech / ASR pipeline familiarity
- Model governance and auditability experience
- Experience with real-time AI systems or low-latency inference
- Experience combining unstructured conversation signals with structured CRM, transaction, or customer profile data
- Experience designing golden datasets and SME review workflows
Operating model
- Engagement structure: FTE-equivalent via long-term B2B contract
- Work location: On-site or closely aligned with the client team in New York
- Infrastructure: Client environment only, no external training or data processing environments
- Data residency: All work executed within the client perimeter
- Delivery mode: PoC-first, with a path toward production-grade conversation intelligence and prediction systems