AI Jobs Map

KAPDAA · Kingston upon Thames

Systems Architect Engineer

full timePosted 2 days ago
Apply on IndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

grpcpostgresqlapache-kafkaredisfastapiwebsocketsreactnext.jsdockeransiblegithub-actionsazureobservabilitysystem-designlinuxpythonllmanthropicci/cd

Job Overview

You will own the architecture of the entire AI4Fibres software platform from cameras and GPU inference on the line, through the on-site tier that keeps production running, to the cloud that carries analytics. You set the throughput and latency budgets, the data contracts, the storage and cost model, and the failure behaviour. The scale you are designing for: several parallel sorting lanes, each carrying multiple items per second, with every per-item decision time-bounded; a continuous, very high volume of image and hyperspectral data, where storage, transaction and bandwidth cost are first-class design constraints rather than an afterthought; and a site that must keep sorting when the WAN is down, and reconcile perfectly when it comes back.

Duties

- System architecture design the production-scale platform end to end: edge line controllers,a site tier that runs the system independently of the cloud, and the cloud tier above it. Decide what belongs in each tier and defend where the line falls.

- Capacity planning & latency budgeting size the system from first principles: per-item time budgets, camera and GPU load, network sizing, headroom and degraded modes.

- ​Data flow & contracts contract-first event schemas and schema registry, gRPC/Protocol Buffers, streaming and message queues, batching, backpressure and idempotency.

- Storage & cost engineering high-volume image and sensor storage: retention, sampling,tiering, packing, and the cost model that goes with them.

- Operational data PostgreSQL under sustained high ingest: schema design, partitioning, indexing, and separating operational from analytical workloads before they collide.

- Edge→cloud resilience store-and-forward, CDC (Kafka + Debezium), Redis, replay, and continuous reconciliation: everything detected is routed, rejected or accounted for, checked against physical mass balance.

- Services & UI layer own the design of FastAPI services (REST, WebSockets, validation, authentication) and React/Next.js operator dashboards with live data views, and build them where needed.

- Deployment & operations Docker, Ansible, GitHub Actions, Azure; versioned rollouts, monitoring, observability and remote diagnostics across an edge fleet; diagnose latency and throughput problems end to end.

- Technical leadership set engineering standards, review designs and code, and present and defend architectural decisions to management, the board and non-technical stakeholders.

Required Skills

Listed in order of expected depth expert command of the core, hands-on proficiency in the rest, and tooling you can pick up here.

CORE EXPERTISE:

●​ Distributed real-time system design, proven in production capacity planning, latency budgeting, backpressure, graceful degradation, failure-mode analysis. You have owned an architecture, not only contributed to one.

●​ Hard-deadline systems experience where a missed deadline could not be retried away, and the system had to degrade safely instead.

●​ Data-intensive architecture at scale, event streaming and CDC, high-ingest relational data, object storage, and explicit cost engineering across storage, transactions and bandwidth.

●​ Edge↔cloud systems an offline-capable control path, store-and-forward, reconciliation, and operating a fleet of remote Linux devices in production.

WORKING PROFICIENCY:

●​ PostgreSQL (schema design, indexing, partitioning, optimisation); object storage and lifecycle/tiering

●​ Kafka + Debezium (CDC), Redis; gRPC + Protocol Buffers, REST, WebSockets

●​ Python, FastAPI, Pydantic, authentication and security basics

●​ React, Next.js; live/streaming data UIs

●​ Docker, GitHub Actions, Ansible, Azure; Linux fundamentals and troubleshooting

●​ Monitoring and observability metrics, tracing, and load/latency testing on the hot path

●​ AI-assisted development: skilled with coding agents and LLM-based workflows (Claude Code, Codex) in everyday engineering without dropping the bar on architecture, tests, security or reliability

WHAT GOOD LOOKS LIKE

●​ Justify decisions in numbers throughput, latency, cost and failure modes, not preference or fashion.

●​ High engineering standards design and code review, secure coding, meaningful tests

including load and latency tests, reliable CI/CD.

●​ Communication & presentation explains an architecture clearly to engineers and to

management, and defends it under challenge.

●​ Start-up mindset, fast-paced environment, shifting requirements, broad ownership, and

willing to build as well as design.

Work Location: In person

More jobs at KAPDAA