Expertise · AI infrastructure recruiting
GPU and HPC Infrastructure Recruiting
Critical Path Hiring recruits the engineers who make accelerated compute usable: GPU cluster engineering, HPC scheduling, hardware deployment and GPU fleet operations, fleet telemetry and observability, storage throughput and capacity and power-aware scheduling across Kubernetes and Slurm environments. These candidates are screened on the scale they have actually operated and the failure modes they have handled, not on the tooling listed on a resume.
Who this is for
GPU cloud and compute platforms, AI labs, HPC centers, enterprises standing up internal training clusters, and data center operators adding compute depth to an existing facilities organization.
The hiring problems it addresses
- Hardware has a delivery date and no one owns making the cluster usable when it lands.
- Job descriptions written for cloud infrastructure return candidates who have never run accelerated nodes at scale.
- Utilization is low and no one can say whether the cause is scheduling, storage, fabric or job design.
- Fleet operations work is absorbed by researchers, which slows the science and burns the team.
Role families
What we recruit, and how each pool is screened.
GPU cluster engineering
Cluster engineer, HPC systems engineer, accelerated compute engineer
Screened on node count and topology operated, driver and firmware lifecycle, and recovery of failed jobs on shared clusters.
Scheduling and capacity
Scheduler engineer, capacity engineer, Slurm or Kubernetes platform engineer
Assessed on queue design, fair-share policy, preemption, and power-aware and thermally constrained scheduling.
Hardware deployment and fleet operations
Hardware deployment engineer, fleet operations engineer, field systems engineer
Evaluated on rack and burn-in process, RMA cycles, spares strategy and the discipline of running a fleet in production.
Fleet telemetry and observability
Telemetry engineer, infrastructure observability engineer
Assessed on GPU-level metrics, per-job attribution, thermal and power telemetry, and turning signals into operational action.
Storage and data path
Storage infrastructure engineer, parallel filesystem engineer
Screened on throughput at training scale, checkpoint behavior and the interaction between storage and interconnect.
Environments
Technologies and environments.
- NVIDIA GPU fleets, DGX and OEM reference systems
- Slurm, Kubernetes, Kubeflow and job schedulers
- Parallel and object storage, high-throughput checkpointing
- Power and thermal constraints at rack and row level
- Firmware, driver and image lifecycle management
Calibration
Common role-calibration mistakes.
Each of these costs a hiring cycle, and each is fixable before outreach starts.
- Requiring a specific scheduler when the underlying skill is queue and capacity design.
- Screening for years of Kubernetes rather than for operating accelerated workloads under contention.
- Combining fleet operations, scheduling and platform engineering into one requisition.
- Setting a compensation band from general cloud engineering data rather than accelerated compute evidence.
How we work
Founder-led search, written research, client-owned pipeline.
Czarina Tabayoyong calibrates the scorecard with the people who will run the panel, maps the named market, runs direct outreach, reports in writing every week, and transfers the research and pipeline to your team. AI assists. People decide.
FAQ
Buyer questions.
- How do you tell a GPU cluster engineer from a cloud infrastructure engineer?
- By the evidence. A GPU cluster engineer can describe node counts and topology they have run, how they handled a failed collective, what firmware and driver drift did to a training run, and how the scheduler behaved under contention. Cloud infrastructure engineers often have none of that exposure even with strong Kubernetes depth.
- Do you recruit for Slurm environments as well as Kubernetes?
- Yes. Many training environments run both, and the transferable skill is queue policy, capacity allocation and failure recovery. We calibrate on that and treat the specific scheduler as a preference rather than a filter unless your environment genuinely requires it.
- What does GPU fleet operations cover?
- Physical and logical fleet health: rack and burn-in, firmware and driver lifecycle, spares and RMA cycles, node drain and return-to-service, and the telemetry that shows whether a node should be trusted with a long job.
- Can you help decide what to hire first?
- Yes. A typical order is one operations-capable engineer who can keep the cluster usable, then fabric and networking depth, then platform engineering to remove toil, then reliability ownership as availability becomes a commitment.
- How do you handle compensation in this market?
- We report the expectations heard in live conversations against your approved band, then work the levers you control: scope, hardware access, ownership, remote policy and equity structure.
Network and fabric recruiting
InfiniBand and high-performance Ethernet engineering for training clusters.
Platform, SRE and observability recruiting
The internal surface and the reliability ownership above the cluster.
AI infrastructure recruiting
The full practice across compute, platform, reliability and model delivery.
Have a role that keeps returning the wrong candidates?
Bring the scorecard and the deadline. You get a market read and a recommended engagement, even if the answer is to solve it internally.