Data scientist / Principal AI Engineer
Location: Tel Aviv, Israel Team: BCG X Type: Full-time
About the role
BCG X is building an AI-native platform for the insurance industry — focused on the agentic transformation of underwriting and claims, in co-development with major global carriers. We are not running pilots. We are deploying production agents into live policy and claims systems at carrier scale.
The next phase is scale: moving from systems that win the deal to agent fleets running 24/7 across multiple carriers, geographies, and regulators. That scale problem is the reason this role exists.
The mission
- Architect a multi-tenant agentic platform that maximizes reuse across clients without becoming a rigid framework no squad wants to use.
- Raise the engineering bar across all product streams: shared evals, shared guardrails, shared observability, shared security baseline, shared deployment patterns.
- Be credible at the C-level with the world's largest carriers — and with their engineering teams in co-development.
- Own build-vs-buy-vs-partner decisions on the core stack, end-to-end.
What you'll do
- Own the reference architecture for the agentic backbone: cloud-native multi-tenant infrastructure, agent orchestration, integration into enterprise core systems, document understanding, and voice pipelines.
- Drive the engineering playbook: canonical patterns for RAG, multi-agent orchestration, tool use, evaluation, guardrails, prompt management, observability, and rollback — adopted across all squads as the default.
- Lead the hardest technical problems personally: production agent failure modes, latency and cost optimization at scale, hybrid LLM + classical ML for high-stakes decisions, multilingual voice quality, regulated-decision auditability.
- Co-design with senior client engineers and architects in joint build mode; translate technical trade-offs into business consequences the C-suite can act on.
- Shape the platform's stance on regulated AI use cases (EU AI Act, GDPR) — built in from day one, not retrofitted.
What you bring
Background — non-negotiable
You have shipped production AI or large-scale distributed systems at one of:
- Hyperscalers (preferred): AWS, Azure, GCP
- Frontier AI labs: Anthropic, OpenAI, DeepMind, Mistral, or equivalent
- AI-first scale-ups: Databricks, Scale, Cohere, Hugging Face, or equivalent
- Big Tech core engineering: Meta, Google, Amazon, Apple, or equivalent
The role requires reflexes that come from operating at this engineering bar, not from working adjacent to it.
Production AI track record
- 10+ years building software, 4+ years shipping LLM-based or agentic systems to production.
- Verifiable examples of agents or AI systems you have put live and kept live — not POCs, not internal demos. You can describe what broke, how you found it, and what you did about it.
- Deep familiarity with the gap between "demo at 95% accuracy" and "production at 70% with a long tail" — and the eval, monitoring, and rollback discipline that closes it.
Platform and infra
- Production experience on a major cloud at scale, with security, networking, IAM, and KMS as second nature.
- Kubernetes, Terraform or equivalent IaC, CI/CD for ML and agents.
- SRE and observability discipline applied to AI systems: distributed tracing, structured logging, latency budgets, on-call.
Agentic AI
- Strong Python. Hands-on with LangGraph or equivalent agent frameworks (we care about depth, not the brand).
- RAG, tool use, structured outputs, and multi-agent orchestration at production quality.
- Document AI: layout understanding, OCR, multimodal pipelines for complex enterprise documents.
- Voice: comfortable with the modern telephony / ASR / TTS stack — you understand the distance between a demo bot and a contact-center-grade voice agent.
Modelling, evaluation, and ops
- LLM evaluation and observability in anger (LangSmith, Langfuse, Braintrust, or custom harnesses you've built yourself). You have opinions about what "eval" actually means in production.
- Prompt engineering and structured-output design at a serious level — and the judgment to know when to fine-tune, distill, route to a smaller model, or fall back to classical ML.
- Guardrails, red-teaming, and safety patterns for high-stakes decisions.
- A/B testing and model performance monitoring at scale.
Posture
- Effective and credible in front of C-suite stakeholders at large enterprises, and in working sessions with their senior engineering teams.
- Tolerance for ambiguity, structured communication, and squad-based delivery.
Strong bonus
- Open source contributions to agent frameworks, LLM eval tooling, or LLM serving infrastructure.
- Conference talks or written work on production agentic AI.
- Working awareness of EU AI Act and GDPR as applied to automated decisioning — or the ability to ramp fast.