AI Jobs Map

Cutshort · Keesara, Telangana, India

Principal AI Architect

Hybridexecutivefull timePosted yesterday
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

system-designcomputer-visionvector-databasesllmobservabilityonnxartificial-intelligencepythonpytorchtensorflowhugging-facevllmspeech-recognitionraglangchainopenaigeminimlopsmachine-learning

Role: Principal AI Architect — Multimodal Video Intelligence

Location: India Remote, with overlap with Singapore working hours

Employment Type: Full-time

Reporting to: Founder / CEO

Function: AI Architecture, Multimodal AI, Video Intelligence, Media Representation

About The Client

The client is building an AI-native media intelligence platform that transforms long-form video into structured, searchable, reusable and monetisable media intelligence.

The Platform Is Not Simply a Video-clipping Tool. We Are Developing a Persistent Intelligence Layer For Media, Where Video, Audio, Speech, Text, Objects, Scenes, Events, Entities, Emotions, Narrative Arcs And Commercial Signals Are Processed Into a Reusable Representation That Can Support Multiple Downstream Use Cases, Including

- short-form clip generation;

- semantic search;

- scene and narrative understanding;

- contextual advertising;

- shoppable video;

- creator and content analytics;

- automated editing workflows;

- future media-intelligence APIs.

We are looking for a Principal AI Architect who can define and guide the AI architecture behind this platform.

Role Summary

The Principal AI Architect — Multimodal Video Intelligence will own the technical architecture for AI systems, including multimodal video understanding, persistent media representation, model orchestration, evaluation frameworks, and production AI design.

This is a hands-on architecture role. The ideal candidate can move between research papers, model selection, system design, data schemas, prototype review, engineering trade-offs, and implementation guidance.

You will work closely with the Founder / CEO, senior AI engineers, computer vision engineers, backend engineers and external vendors to convert the product and IP vision into a robust technical system.

Key Responsibilities

- AI System Architecture

- Define the end-to-end AI architecture for long-form video understanding.

- Design the processing pipeline from video ingest to structured media intelligence.

- Define how vision, audio, speech, text, metadata and user signals should be fused.

- Design the architecture for reusable media intelligence rather than one-time clip generation.

- Ensure the system can support multiple downstream applications from the same processed media layer.

- Persistent Media Representation

- Design persistent media representation layer across multiple levels, including frame, object, shot, scene, segment, entity, event and full-video levels.

- Define what intelligence must be stored permanently versus computed on demand.

- Design schemas for temporal, spatial, semantic, narrative and commercial metadata.

- Define provenance, confidence, model versioning and evidence-tracking requirements.

- Ensure the representation remains usable even when underlying AI models are replaced or upgraded.

- Multimodal Model Strategy

- Select and evaluate appropriate models for video, image, audio, speech, OCR, entity extraction, scene understanding, action recognition, embeddings, reranking and LLM/VLM reasoning.

- Decide where to use open-source models, commercial APIs, fine-tuning or custom models.

- Define model interfaces so models can be swapped without breaking downstream systems.

- Guide model benchmarking for accuracy, latency, cost and scalability.

- Prevent over-dependence on any single model vendor or API.

- Temporal and Narrative Intelligence

- Design approaches for understanding long-form video structure, including scenes, events, story arcs, character/entity continuity and engagement peaks.

- Define methods to identify clip-worthy moments across different content types.

- Support narrative scoring, highlight ranking, scene segmentation and coherence validation.

- Ensure that clips are not only visually interesting but contextually and narratively coherent.

- Evaluation and Benchmarking

- Define objective evaluation frameworks for AI outputs.

- Build or guide creation of benchmark datasets and UAT criteria.

- Define metrics for clip quality, scene accuracy, entity continuity, timestamp alignment, hallucination control, ranking quality, retrieval precision and cost efficiency.

- Establish model and prompt evaluation processes.

- Create regression-testing methodology when models, prompts, schemas or scoring logic change.

- Search, Retrieval and Knowledge Layer

- Design hybrid search architecture across transcript, visual events, metadata, embeddings and structured knowledge.

- Define when to use relational storage, vector databases, graph databases and object storage.

- Design queryable media intelligence for downstream APIs and applications.

- Support knowledge-graph or ontology-based representation where useful.

- Ensure retrieved outputs are evidence-backed and timestamp-grounded.

- Production AI Architecture

- Work with AI engineers to convert architecture into deployable services.

- Guide decisions on batching, GPU inference, model serving, queues, retries, observability and cost controls.

- Review pipeline designs involving FFmpeg, GStreamer, DeepStream, TensorRT, Triton, ONNX, cloud services and model APIs.

- Define failure-handling, reprocessing, versioning and rollback mechanisms.

- Support scalable design without premature overengineering.

- IP and Technical Differentiation

- Help translate AI architecture into defensible technical differentiation.

- Support patent-related technical disclosures where required.

- Identify what is proprietary versus commodity.

- Avoid building a generic wrapper over existing models.

- Ensure the architecture reinforces the core thesis of persistent, reusable media intelligence.

- Team Guidance

- Provide technical direction to senior AI engineers and computer vision engineers.

- Review designs, experiments, evaluation results and architecture decisions.

- Mentor engineers without becoming a pure people manager.

- Help define technical milestones for the first 90, 180 and 365 days.

- Support hiring, technical interviews and vendor evaluation where needed.

Required Experience

The ideal candidate should have:

- 8+ years of AI/ML experience, with significant exposure to computer vision, video AI, multimodal AI, retrieval systems or production ML architecture.

- Strong experience designing AI systems, not only implementing isolated models.

- Hands-on experience with video understanding, temporal modelling, multimodal pipelines, VLMs, LLMs, embeddings, ranking or retrieval.

- Experience taking AI systems from prototype to production.

- Strong knowledge of Python and modern AI/ML frameworks such as PyTorch, TensorFlow, Hugging Face or equivalent.

- Experience with model evaluation, benchmarking, error analysis and dataset design.

- Understanding of production architecture: APIs, queues, databases, cloud, model serving, observability and deployment trade-offs.

- Ability to work with founders and engineers in a high-ambiguity startup environment.

Strongly Preferred Experience

- Video understanding, action recognition, scene segmentation, event detection or video retrieval.

- Multimodal AI involving video, audio, speech, text and metadata.

- LLM/VLM orchestration for structured outputs.

- Prompt/version management, schema validation and hallucination control.

- Embedding search, vector databases, reranking and retrieval evaluation.

- Knowledge graphs, ontologies, entity resolution or temporal knowledge representation.

- Model serving using TensorRT, Triton, ONNX, vLLM, DeepStream or similar.

- Experience with long-form video, OTT, sports media, entertainment, creator platforms, advertising technology or social commerce.

- Experience contributing to patents, technical disclosures or investor diligence.

Technical Areas

The candidate should be comfortable discussing and making architecture decisions across:

- Computer vision;

- video AI;

- multimodal fusion;

- speech-to-text;

- OCR;

- image/video embeddings;

- VLMs and LLMs;

- semantic search;

- vector databases;

- graph databases;

- temporal reasoning;

- ranking and scoring systems;

- prompt orchestration;

- model evaluation;

- model versioning;

- data lineage;

- GPU inference;

- cloud AI deployment.

What This Role Is Not

This is not a role for someone who has only built:

- chatbots;

- basic RAG demos;

- LangChain prototypes;

- prompt-engineering workflows;

- simple OpenAI/Gemini API wrappers;

- dashboards over model outputs;

- classical computer vision demos without production architecture;

- MLOps pipelines without AI system-design depth.

The role requires architectural depth in AI systems, not just familiarity with AI tools.

First 90-Day Expectations

First 30 Days

- Review product thesis, patent direction, prototype plans and existing technical assumptions.

- Assess current team capability and architecture gaps.

- Define the first version of AI architecture.

- Identify immediate technical risks and validation priorities.

First 60 Days

- Deliver a detailed architecture document covering media representation, model stack, pipeline design, storage strategy, evaluation framework and implementation roadmap.

- Define the canonical media-intelligence schema.

- Define model-selection and benchmarking criteria.

- Guide senior engineers on first implementation milestones.

First 90 Days

- Help the team implement and validate the first working version of the persistent media-intelligence layer.

- Establish evaluation datasets and UAT metrics.

- Review prototype outputs and improve architecture based on evidence.

- Produce a 6-month AI roadmap with technical risks, milestones and resourcing needs.

Success Metrics

The Principal AI Architect Will Be Successful If

- They have a clear AI architecture that the engineering team can execute.

- The platform does not collapse into a generic clip-generation pipeline.

- The media representation is reusable across multiple use cases.

- Models, prompts and schemas are versioned and testable.

- AI outputs are measurable through objective benchmarks.

- Snehashish, Abhishek and other engineers have clear technical direction.

- The architecture supports both product execution and investor/IP defensibility.

Candidate Personality Fit

The Right Candidate Should Be

- intellectually strong but practical;

- hands-on enough to review code and experiments;

- comfortable with ambiguity;

- willing to challenge assumptions with evidence;

- able to simplify complex AI architecture for engineers and investors;

- disciplined about evaluation, cost and production constraints;

- not attached to one model, tool or vendor;

- able to work in a founder-led early-stage startup.

Skills:- Multi-modal AI, Artificial Intelligence (AI), Computer Vision, Machine Learning (ML) and Python

More jobs at Cutshort

Principal AI Architect at Cutshort (Keesara, Telangana, India) | AI Jobs Map