About the Role
We are looking for a Lead Backend / Infrastructure Engineer to lead a team of engineers building core Neo Cloud infrastructure services across GPU virtualization, networking, storage, and firewall integrations.
You will own the architecture, execution, and technical direction of backend control-plane services and automation platforms that power large-scale GPU cloud infrastructure. This role requires strong backend engineering expertise, distributed systems knowledge, cloud infrastructure experience, and the ability to lead engineers working across deeply technical infrastructure domains.
You will work closely with virtualization, networking, storage, security, DevOps/SRE, product, and platform teams to deliver reliable, secure, high-performance cloud services for GPU-intensive workloads.
Key Responsibilities
Technical Leadership
-
Lead and mentor a team of backend and infrastructure engineers.
-
Provide technical direction across GPU virtualization, networking, storage, and firewall integration initiatives.
-
Drive architecture reviews, design reviews, code reviews, and engineering best practices.
-
Break down complex infrastructure problems into clear engineering plans.
-
Partner with product, infrastructure, security, and operations teams to define roadmap and execution priorities.
-
Own projects from requirements and architecture through implementation, rollout, operations, and long-term ownership.
Backend & Control-Plane Architecture
-
Design and build scalable backend services for Neo Cloud infrastructure management.
-
Define service architecture, APIs, data models, integration patterns, and platform strategy.
-
Build control-plane components for provisioning, orchestration, lifecycle management, telemetry, and automation.
-
Design REST/gRPC APIs and internal service interfaces across cloud infrastructure services.
-
Drive architectural decisions around scalability, availability, consistency, fault tolerance, and operational simplicity.
GPU Virtualization & Compute Infrastructure
-
Lead backend systems supporting GPU provisioning, allocation, scheduling, lifecycle management, and virtualization workflows.
-
Build integrations for GPU passthrough, vGPU, MIG, device plugins, drivers, and hardware telemetry where applicable.
-
Design systems that support multi-tenant GPU workloads with strong isolation, performance, and reliability.
-
Collaborate with platform and infrastructure teams to optimize GPU resource utilization and operational visibility.
Networking, Storage & Firewall Integrations
-
Lead engineering efforts across cloud networking, tenant networking, routing, load balancing, and network automation.
-
Build integrations with storage platforms including block, file, and object storage services.
-
Design APIs and orchestration workflows for volume provisioning, attachment, snapshots, replication, and performance-sensitive storage use cases.
-
Lead firewall and network security integrations including security groups, ACLs, network policies, distributed firewalls, and tenant isolation.
-
Collaborate on Kubernetes networking, CNI, service discovery, ingress/load balancing, VLAN/VXLAN, and network policy enforcement.
Distributed Systems & Data
-
Design event-driven and asynchronous architectures for processing infrastructure events and telemetry at scale.
-
Build high-throughput data pipelines for metrics, logs, topology, inventory, GPU telemetry, network state, storage state, and operational events.
-
Ensure services are observable, debuggable, resilient, and reliable in 24x7 production environments.
Cloud & Platform Engineering
-
Build and operate cloud-native backend services using Docker and Kubernetes.
-
Contribute to automation and Infrastructure-as-Code using tools such as Terraform and Ansible.
-
Drive CI/CD, automated testing, release engineering, and production-readiness practices.
-
Improve reliability, security, scalability, and operational excellence across Neo Cloud platform services.
Required Qualifications
-
8+ years of backend/software engineering experience, with significant experience in distributed systems, infrastructure platforms, or cloud services.
-
Experience leading technical initiatives and mentoring engineering teams.
-
Strong proficiency in at least one backend/system programming language such as Go, Java, Python, C++, or Rust.
-
Strong understanding of distributed systems, concurrency, scalability, fault tolerance, and high availability.
-
Proven experience designing and implementing RESTful and/or gRPC APIs and microservices.
-
Strong Linux fundamentals and production troubleshooting experience.
-
Hands-on experience with containers and Kubernetes.
-
Experience with cloud infrastructure, infrastructure automation, or control-plane systems.
-
Strong understanding of networking fundamentals including TCP/IP, L2/L3 networking, routing, switching, load balancing, IPv4/IPv6, VLAN, and VXLAN.
-
Experience with relational and/or NoSQL databases.
-
Strong analytical, debugging, and problem-solving skills.
-
Ability to lead complex projects across multiple technical domains.
Good to Have
-
Experience with GPU infrastructure, GPU virtualization, vGPU, MIG, device plugins, or accelerator scheduling.
-
Knowledge of Kubernetes networking and CNI technologies such as Cilium, Calico, or equivalent.
-
Experience with storage systems including Ceph, NVMe/TCP, iSCSI, NFS, object storage, or cloud block storage.
-
Familiarity with firewall platforms, security groups, ACLs, network policies, and distributed firewall architectures.
-
Experience building 24x7 production infrastructure, cloud platforms, or network management systems.
-
Familiarity with TLS, PKI, authentication/authorization, zero-trust architectures, and multi-tenant security models.