Hiring strategy

Designing an AI infrastructure team: five functions, one sequence.

Network, platform, SRE, MLOps and capacity planning are distinct functions with distinct talent pools. Hiring them by title, all at once, is how orgs end up with three platform engineers and no one who owns the fabric.

Czarina Tabayoyong

Founder & Principal Recruiter · Published September 14, 2026

The short answer

An AI infrastructure org has five functions: fabric and networking, platform engineering, site reliability, MLOps, and capacity planning. What you hire first depends on your constraint: hardware landing means fleet and fabric first; researchers blocked on tooling means platform; customer SLAs mean SRE; training pipelines in production mean MLOps; power-limited growth means capacity planning. Sequence against the constraint, not the org chart.

The five functions and what each owns

Fabric and networking owns the interconnect: topology, congestion, collective-communication performance. Platform engineering owns the internal surface other teams consume: scheduling, tooling, self-service. SRE owns availability as a commitment: error budgets, incident response, observability. MLOps owns the model lifecycle: training pipelines, artifacts, serving. Capacity planning owns the future: power, supply and demand reconciled into a buildable plan.

The functions are not seniority levels and they are not interchangeable. Each draws from a different talent pool, hires on different evidence, and clears at a different comp band. Treating them as one requisition family is the single most expensive mistake in this org design, because it produces a team that all thinks in the same layer while the failure modes live in the layers nobody owns.

  • Fabric/networking: InfiniBand or RoCE topology, congestion, collectives
  • Platform: scheduling, developer surface, adoption of internal tooling
  • SRE: error budgets, incidents, observability under physical failure modes
  • MLOps: training pipelines, checkpoints, serving latency, GPU utilization
  • Capacity planning: power- and supply-constrained growth modeling

Where each pool actually comes from

Sourcing follows the pool, not the title. Fabric depth sits in HPC labs, federal research computing, hyperscale network engineering and the silicon ecosystem, a small, largely non-searching population that responds to specificity about the cluster, not to the word 'exciting'. Platform engineers come out of internal developer platform teams at scaled software companies and convert well, provided the interview tests adoption rather than tool familiarity.

SRE is the deepest pool of the five and the most likely to be miscast: an SRE from a pure-software environment has rarely debugged an incident whose root cause was thermal, electrical or optical. MLOps splits into a research-adjacent half and a serving-adjacent half, and the two are not substitutes. Capacity planning is the thinnest pool of all and often comes from data center development, utility interconnection or supply chain rather than from engineering at all.

  • Fabric: HPC and research computing, hyperscale networking, silicon ecosystem
  • Platform: internal developer platform teams at scaled software companies
  • SRE: broad pool, but screen explicitly for physical failure modes
  • MLOps: research-adjacent (training) and serving-adjacent (inference) halves
  • Capacity planning: development, interconnection and supply chain backgrounds

The comparison that keeps interviews honest

These candidates share vocabulary and tooling, so interviews must test ownership. Ask what they were paged for, what they built that others adopted, and what number moved because of their work. A platform engineer answers in adoption and toil removed; an SRE in incidents and budgets; an MLOps engineer in training throughput and serving latency.

The strongest single question in this whole org is: what did you personally do the last time the cluster was degraded and the cause was not in your layer? A candidate with real fleet exposure describes coordination with facilities, vendors and the network team, and can name the signal that told them where to look. A candidate whose experience is entirely abstracted describes escalating and waiting.

  • Fabric: which collectives regressed, and what in the topology explained it
  • Platform: what percentage of teams adopted the surface, and why the rest did not
  • SRE: the error budget policy they enforced and the release it stopped
  • MLOps: checkpoint recovery time and GPU utilization before and after
  • Capacity planning: the plan they signed that turned out wrong, and why

Sequencing by stage

First cluster landing: one operations-capable engineer and fabric depth. Researchers losing time to infrastructure: platform. Availability in contracts: SRE. Training and serving stabilizing: MLOps. Growth gated by power or supply: capacity planning. Hiring ahead of the constraint burns cash; hiring behind it burns the team's credibility with the people they support.

The sequence matters more than the headcount because each hire changes what the next one can be. A strong fabric engineer hired first makes the platform hire easier to scope and easier to sell, because the candidate can see what is already solid. A platform hire made first into an unstable fleet spends a year on firefighting they were not hired for, and leaves.

  • Stage 1, hardware landing: fleet and fabric
  • Stage 2, researchers blocked: platform engineering
  • Stage 3, external commitments: SRE and on-call structure
  • Stage 4, pipelines in production: MLOps
  • Stage 5, power- or supply-limited growth: capacity planning

Four ways this org design fails

Almost every troubled AI infrastructure org I look at has one of four shapes. Recognizing yours is faster than re-running the org chart exercise.

  • Three platform engineers and no fabric owner: everything works until the interconnect degrades, and then nobody can read it.
  • SRE hired as a title, not a mandate: on-call exists, error budgets do not, and reliability stays a matter of individual heroics.
  • MLOps absorbed into platform by accident: training pipelines get whatever attention is left after the developer surface, which is none.
  • Capacity planning owned by finance alone: the model is accurate about money and silent about power, so growth commitments outrun the buildable plan.

What to build internally and what to hire outside

Platform and MLOps promote well from strong internal software engineers who already know the workloads, provided someone senior owns the standard. Fabric and capacity planning rarely do: both depend on pattern recognition built over multiple builds and multiple failures, and neither can be learned on your schedule without an expensive teacher already in the building.

That asymmetry should drive the search plan. Spend external search budget on the functions you cannot grow, and spend management attention on the ones you can. Inverting that is how teams end up paying a premium for the easiest hire in the org.

How to run the first ninety days of hiring

Name the binding constraint in a number before opening any requisition, researcher queue hours, utilization, incident hours, or the gap between contracted and sellable capacity. Scope the first two seats against that number, and write the scorecard in evidence terms rather than tool lists. Then run both searches in parallel rather than in sequence, because the pools do not overlap and competing for them one at a time simply doubles the calendar.

Review the plan at day sixty against the same number you started with. If it has not moved, the problem is usually scope, not sourcing: the seat as written asks for a combination of evidence that fewer than a dozen people in the country hold.

By market

How this plays out where you are building.

Supply, compensation and sequencing are local. These reads cover the metros where this pattern shows up most.

Questions

What employers ask about this.

Can one senior hire cover two of these functions early on?
Yes, deliberately: platform and MLOps pair well early, as do fleet operations and fabric. What fails is accidental double-coverage, a hire made for one function who is silently expected to carry another they were never screened for.
Where does security fit in this design?
Usually alongside platform early, as a dedicated function once customer commitments include compliance language. It is a distinct pool again; do not fold it silently into SRE.
How do we know which constraint is actually binding?
Measure it: researcher queue time, utilization, incident hours, or the gap between contracted demand and sellable capacity. The constraint that shows up in a number is the one to hire against first.
What does the first infrastructure hire look like for a team with one cluster?
One operations-capable engineer with real fabric exposure, scoped broadly and hired senior. At one cluster you are buying judgment about what is actually wrong, not coverage of an org chart, and a narrow specialist in the wrong layer leaves you blind in the others.
What controls the pace of these searches?
Platform and SRE searches typically produce a signed offer inside six to ten weeks. Fabric and capacity planning run longer, often twelve weeks or more, because the pools are small, largely passive, and rarely competing on comp alone.
Should these roles be onsite, hybrid or remote?
Fleet and fabric work benefits materially from proximity to the hardware, especially during bring-up. Platform, MLOps and most SRE work is genuinely location-flexible, and insisting otherwise narrows an already thin market for no operational gain.

Want this applied to your own hiring plan?

A hiring diagnostic produces a written read on your roles, market and sequence, whether or not we run the search.