About the Role
Our client is a world-leading technology company operating at the intersection of high-performance computing, Bitcoin mining, and AI cloud services, with a globally distributed infrastructure footprint spanning multiple continents and a multi-gigawatt energy portfolio. Headquartered in Singapore, they are scaling rapidly and investing heavily in next-generation datacenter and cloud capabilities. This role sits at the deepest technical tier of the operations centre, owning the most complex customer escalations and platform incidents as the critical bridge between front-line support and Engineering/SRE teams. It is an exceptional opportunity for a seasoned infrastructure engineer to drive permanent, systemic improvements rather than just resolving individual tickets. You will directly shape the reliability and quality of a cloud platform serving high-demand AI and HPC workloads.
Key Responsibilities
Own complex, escalated customer issues and platform incidents end-to-end, driving each case through to full resolution
Conduct deep-dive troubleshooting across
GPU compute
, networking/SDN, storage, drivers, control plane, and billing systems
Serve as the primary liaison to SRE, Compute, and R&D teams — leading root-cause analysis and ensuring permanent fixes are implemented
Lead and support incident response and post-incident reviews, translating findings into improved runbooks and monitoring enhancements
Identify recurring escalation patterns and convert them into lasting platform or process improvements
Mentor L1 and L2 support engineers, raising escalation quality and expanding the team knowledge base
Participate in an on-call escalation rotation, providing incident leadership during critical platform events
Collaborate cross-functionally with engineering and product teams to feed operational learnings back into the platform roadmap
Requirements
Minimum 5 years of hands-on experience in
cloud infrastructure
technical support, escalation engineering, or SRE-adjacent roles
Strong practical proficiency in Linux, networking, and cloud infrastructure; experience with GPU/CUDA/HPC environments
(is a bonus)
Demonstrated track record of resolving complex production incidents and conducting thorough root-cause analysis
Excellent written English communication skills, with the ability to work effectively across technical and non-technical stakeholders
Comfortable with on-call responsibilities and experienced in taking ownership during high-pressure incident scenarios
Familiarity with SDN, distributed storage systems, or control plane architecture
(is a bonus)
If you are passionate about technology and meet the above requirements, please don't hesitate to apply. Please note that only shortlisted candidates will be contacted. Appreciate your understanding. Data provided is for recruitment purposes only.
Dada Consultants Pte Ltd
Website:
www.dadaconsultants.com
EA License No.: 18S9037
Business Registration Number: 201735941W
Do you still think the boutique agency cannot advanced your own tech recruitment career?
Are you considering to change your career path from big agency to enjoy the best commission structure and working with best in class tech clients in Singapore?
Join US!
We are the Singapore based technology recruitment agency and specialised in providing talent solution for internet and technology industry. Our client portfolio are including the global internet technology giants, technology MNC and Tech startup. Whatever technology clients you are think of? Yes we do have!
Due to confidence and client demand in Singapore, we're looking for an experienced tech recruitment Consultant to join our team to service our clients across Asia Pacific.
In this role you'll help our clients identify & hire talent people not only with the right skills and experience, but also the right motivational and cultural fit - people who'll stay and perform.