Jobs › US jobs › Senior Site Reliability Engineer, DGX Cloud
NVIDIA · US, CA, Santa Clara
NVIDIA DGX Cloud is developing and managing large-scale GPU infrastructure for AI research and production workloads. We are looking for Senior Reliability Engineers to help build the automation, tooling, and operational systems that make GPU clusters reliable, scalable, and safe to run. This role is part of a production engineering team passionate about Kubernetes-based infrastructure, GPU cluster operations, reliability, automation, GitOps, and Day 2 operability across DGX…
Read the full advert and apply on NVIDIA's site →
Collected from NVIDIA's own careers site (Workday). Posted 18 Sep 2026. TUNAI shows an excerpt and the facts it read from the advert; the employer's page has the full description and the application.
Does your CV match this job? Paste both and see exactly which of these skills it shows.
Check my CV against this job →