Own the physical reality of the platform in one of our seven regions. You bring up new GPU racks, validate InfiniBand fabric end-to-end, and keep the cluster running at the SLA. This is a hands-on role;
iFrame seeks a hands-on Cluster Site Reliability Engineer to own the physical reality of a GPU-centric platform across seven regions. This on-site role in Toronto involves bringing up new racks, validating InfiniBand fabric, and ensuring the SLA
Cluster Site Reliability Engineer Cluster & SRE Own the physical reality of the platform in one of our seven regions. You bring up new GPU racks, validate InfiniBand fabric end-to-end, and keep the cluster running at
Join a cutting-edge team as a Site Reliability Engineer, specializing in managing robust compute clusters and InfiniBand infrastructure across Cologix regions. This full-time, hands-on position is ideal for experienced engineers passionate about hardware and reliability. In