Expertise · AI infrastructure recruiting
Platform, SRE and Observability Recruiting
Critical Path Hiring recruits platform engineers, site reliability engineers and infrastructure observability engineers for teams running AI and GPU compute in production. Platform engineering, reliability ownership and observability are three separate talent markets that share a vocabulary. We calibrate each on what the person owned, the scale they carried and the incidents they ran, then screen against that evidence.
Who this is for
GPU cloud and compute platforms, AI product companies, enterprises operating internal platforms, and operators whose availability targets have become customer commitments.
The hiring problems it addresses
- Researchers and product engineers are maintaining their own infrastructure surface.
- Availability moved from an internal aspiration to contract language with no reliability owner.
- Observability exists as dashboards but not as per-job or per-tenant attribution.
- One requisition is asking for platform, reliability and deployment engineering at once.
Role families
What we recruit, and how each pool is screened.
Platform engineering
Platform engineer, infrastructure engineer, developer platform lead
Assessed on the internal surface they built, adoption by other teams and the toil they removed, not on tool familiarity.
Site reliability
SRE, production engineer, reliability lead, incident commander
Screened on error budget practice, on-call design, incident command and postmortem discipline under real customer-facing targets.
Infrastructure observability
Observability engineer, telemetry platform engineer
Evaluated on cardinality and cost control, trace coverage, and attribution of resource use to jobs, tenants and teams.
Distributed systems
Distributed systems engineer, control plane engineer
Assessed on consistency, coordination and failure reasoning in systems they have actually operated.
Environments
Technologies and environments.
- Kubernetes control planes and multi-tenant clusters
- Terraform, service meshes and internal developer platforms
- Prometheus, OpenTelemetry and trace-based observability
- Error budgets, SLOs and on-call rotations
- Capacity, quota and cost attribution across shared compute
Calibration
Common role-calibration mistakes.
Each of these costs a hiring cycle, and each is fixable before outreach starts.
- Writing one job description covering platform, reliability and customer-facing engineering.
- Hiring reliability ownership before there is a service and a target to be reliable against.
- Screening SRE candidates on coding puzzles rather than incident ownership evidence.
- Expecting a platform hire to also own model training pipelines.
How we work
Founder-led search, written research, client-owned pipeline.
Czarina Tabayoyong calibrates the scorecard with the people who will run the panel, maps the named market, runs direct outreach, reports in writing every week, and transfers the research and pipeline to your team. AI assists. People decide.
FAQ
Buyer questions.
- What is the difference between platform engineering and SRE?
- Platform engineers build and own the internal surface other teams consume. SREs own availability outcomes: error budgets, on-call, incident response and the reliability work that follows a postmortem. Some organizations combine them, but the assessment evidence is different and should be written down before sourcing.
- Where does observability sit?
- Usually inside platform or reliability, but the hire is distinct. Observability engineers own signal quality, cost and attribution, which is what makes GPU utilization and tenant behavior legible.
- Do you recruit for on-call heavy roles?
- Yes, and we surface on-call reality in the first conversation. Hiding rotation load produces offers that fail late or attrition inside a year.
- Can you calibrate the scorecard with our engineers?
- Yes. Calibration with the engineers who will run the panel is part of every engagement, and the agreed scorecard is written down before outreach begins.
Have a role that keeps returning the wrong candidates?
Bring the scorecard and the deadline. You get a market read and a recommended engagement, even if the answer is to solve it internally.