Role: Principal AI Architect — Multimodal Video Intelligence
Location: India Remote, with overlap with Singapore working hours
Employment Type: Full-time
Reporting to: Founder / CEO
Function: AI Architecture, Multimodal AI, Video Intelligence, Media Representation
About The Client
The client is building an AI-native media intelligence platform that transforms long-form video into structured, searchable, reusable and monetisable media intelligence.
The Platform Is Not Simply a Video-clipping Tool. We Are Developing a Persistent Intelligence Layer For Media, Where Video, Audio, Speech, Text, Objects, Scenes, Events, Entities, Emotions, Narrative Arcs And Commercial Signals Are Processed Into a Reusable Representation That Can Support Multiple Downstream Use Cases, Including
- short-form clip generation;
- semantic search;
- scene and narrative understanding;
- contextual advertising;
- shoppable video;
- creator and content analytics;
- automated editing workflows;
- future media-intelligence APIs.
We are looking for a Principal AI Architect who can define and guide the AI architecture behind this platform.
Role Summary
The Principal AI Architect — Multimodal Video Intelligence will own the technical architecture for AI systems, including multimodal video understanding, persistent media representation, model orchestration, evaluation frameworks, and production AI design.
This is a hands-on architecture role. The ideal candidate can move between research papers, model selection, system design, data schemas, prototype review, engineering trade-offs, and implementation guidance.
You will work closely with the Founder / CEO, senior AI engineers, computer vision engineers, backend engineers and external vendors to convert the product and IP vision into a robust technical system.
Key Responsibilities
- AI System Architecture
- Define the end-to-end AI architecture for long-form video understanding.
- Design the processing pipeline from video ingest to structured media intelligence.
- Define how vision, audio, speech, text, metadata and user signals should be fused.
- Design the architecture for reusable media intelligence rather than one-time clip generation.
- Ensure the system can support multiple downstream applications from the same processed media layer.
- Persistent Media Representation
- Design persistent media representation layer across multiple levels, including frame, object, shot, scene, segment, entity, event and full-video levels.
- Define what intelligence must be stored permanently versus computed on demand.
- Design schemas for temporal, spatial, semantic, narrative and commercial metadata.
- Define provenance, confidence, model versioning and evidence-tracking requirements.
- Ensure the representation remains usable even when underlying AI models are replaced or upgraded.
- Multimodal Model Strategy
- Select and evaluate appropriate models for video, image, audio, speech, OCR, entity extraction, scene understanding, action recognition, embeddings, reranking and LLM/VLM reasoning.
- Decide where to use open-source models, commercial APIs, fine-tuning or custom models.
- Define model interfaces so models can be swapped without breaking downstream systems.
- Guide model benchmarking for accuracy, latency, cost and scalability.
- Prevent over-dependence on any single model vendor or API.
- Temporal and Narrative Intelligence
- Design approaches for understanding long-form video structure, including scenes, events, story arcs, character/entity continuity and engagement peaks.
- Define methods to identify clip-worthy moments across different content types.
- Support narrative scoring, highlight ranking, scene segmentation and coherence validation.
- Ensure that clips are not only visually interesting but contextually and narratively coherent.
- Evaluation and Benchmarking
- Define objective evaluation frameworks for AI outputs.
- Build or guide creation of benchmark datasets and UAT criteria.
- Define metrics for clip quality, scene accuracy, entity continuity, timestamp alignment, hallucination control, ranking quality, retrieval precision and cost efficiency.
- Establish model and prompt evaluation processes.
- Create regression-testing methodology when models, prompts, schemas or scoring logic change.
- Search, Retrieval and Knowledge Layer
- Design hybrid search architecture across transcript, visual events, metadata, embeddings and structured knowledge.
- Define when to use relational storage, vector databases, graph databases and object storage.
- Design queryable media intelligence for downstream APIs and applications.
- Support knowledge-graph or ontology-based representation where useful.
- Ensure retrieved outputs are evidence-backed and timestamp-grounded.
- Production AI Architecture
- Work with AI engineers to convert architecture into deployable services.
- Guide decisions on batching, GPU inference, model serving, queues, retries, observability and cost controls.
- Review pipeline designs involving FFmpeg, GStreamer, DeepStream, TensorRT, Triton, ONNX, cloud services and model APIs.
- Define failure-handling, reprocessing, versioning and rollback mechanisms.
- Support scalable design without premature overengineering.
- IP and Technical Differentiation
- Help translate AI architecture into defensible technical differentiation.
- Support patent-related technical disclosures where required.
- Identify what is proprietary versus commodity.
- Avoid building a generic wrapper over existing models.
- Ensure the architecture reinforces the core thesis of persistent, reusable media intelligence.
- Team Guidance
- Provide technical direction to senior AI engineers and computer vision engineers.
- Review designs, experiments, evaluation results and architecture decisions.
- Mentor engineers without becoming a pure people manager.
- Help define technical milestones for the first 90, 180 and 365 days.
- Support hiring, technical interviews and vendor evaluation where needed.
Required Experience
The ideal candidate should have:
- 8+ years of AI/ML experience, with significant exposure to computer vision, video AI, multimodal AI, retrieval systems or production ML architecture.
- Strong experience designing AI systems, not only implementing isolated models.
- Hands-on experience with video understanding, temporal modelling, multimodal pipelines, VLMs, LLMs, embeddings, ranking or retrieval.
- Experience taking AI systems from prototype to production.
- Strong knowledge of Python and modern AI/ML frameworks such as PyTorch, TensorFlow, Hugging Face or equivalent.
- Experience with model evaluation, benchmarking, error analysis and dataset design.
- Understanding of production architecture: APIs, queues, databases, cloud, model serving, observability and deployment trade-offs.
- Ability to work with founders and engineers in a high-ambiguity startup environment.
Strongly Preferred Experience
- Video understanding, action recognition, scene segmentation, event detection or video retrieval.
- Multimodal AI involving video, audio, speech, text and metadata.
- LLM/VLM orchestration for structured outputs.
- Prompt/version management, schema validation and hallucination control.
- Embedding search, vector databases, reranking and retrieval evaluation.
- Knowledge graphs, ontologies, entity resolution or temporal knowledge representation.
- Model serving using TensorRT, Triton, ONNX, vLLM, DeepStream or similar.
- Experience with long-form video, OTT, sports media, entertainment, creator platforms, advertising technology or social commerce.
- Experience contributing to patents, technical disclosures or investor diligence.
Technical Areas
The candidate should be comfortable discussing and making architecture decisions across:
- Computer vision;
- video AI;
- multimodal fusion;
- speech-to-text;
- OCR;
- image/video embeddings;
- VLMs and LLMs;
- semantic search;
- vector databases;
- graph databases;
- temporal reasoning;
- ranking and scoring systems;
- prompt orchestration;
- model evaluation;
- model versioning;
- data lineage;
- GPU inference;
- cloud AI deployment.
What This Role Is Not
This is not a role for someone who has only built:
- chatbots;
- basic RAG demos;
- LangChain prototypes;
- prompt-engineering workflows;
- simple OpenAI/Gemini API wrappers;
- dashboards over model outputs;
- classical computer vision demos without production architecture;
- MLOps pipelines without AI system-design depth.
The role requires architectural depth in AI systems, not just familiarity with AI tools.
First 90-Day Expectations
First 30 Days
- Review product thesis, patent direction, prototype plans and existing technical assumptions.
- Assess current team capability and architecture gaps.
- Define the first version of AI architecture.
- Identify immediate technical risks and validation priorities.
First 60 Days
- Deliver a detailed architecture document covering media representation, model stack, pipeline design, storage strategy, evaluation framework and implementation roadmap.
- Define the canonical media-intelligence schema.
- Define model-selection and benchmarking criteria.
- Guide senior engineers on first implementation milestones.
First 90 Days
- Help the team implement and validate the first working version of the persistent media-intelligence layer.
- Establish evaluation datasets and UAT metrics.
- Review prototype outputs and improve architecture based on evidence.
- Produce a 6-month AI roadmap with technical risks, milestones and resourcing needs.
Success Metrics
The Principal AI Architect Will Be Successful If
- They have a clear AI architecture that the engineering team can execute.
- The platform does not collapse into a generic clip-generation pipeline.
- The media representation is reusable across multiple use cases.
- Models, prompts and schemas are versioned and testable.
- AI outputs are measurable through objective benchmarks.
- Snehashish, Abhishek and other engineers have clear technical direction.
- The architecture supports both product execution and investor/IP defensibility.
Candidate Personality Fit
The Right Candidate Should Be
- intellectually strong but practical;
- hands-on enough to review code and experiments;
- comfortable with ambiguity;
- willing to challenge assumptions with evidence;
- able to simplify complex AI architecture for engineers and investors;
- disciplined about evaluation, cost and production constraints;
- not attached to one model, tool or vendor;
- able to work in a founder-led early-stage startup.
Skills:- Multi-modal AI, Artificial Intelligence (AI), Computer Vision, Machine Learning (ML) and Python