Menu

Use case / 04 · iFabric

Find the constraint. Plan capacity.

GPU demand, seen with the infrastructure around it.

All use cases

Configured evidence · Defined workloads & planning assumptions

GPU capacity and workload bottlenecksA workload slows or demand grows. Relate GPU, network, storage, power and cooling evidence to workload health, identify the limiting resource and review a capacity recommendation. Recommendations require review. Any capacity or placement change follows the configured approval process. Evidence sources converge on workload context; they are not sequential processing stages. Missing evidence leads to collection and operator review. The diagram ends at recommendation handoff for change approval, without provisioning capacity or moving workloads. Animation timing is illustrative.Configured infrastructure evidenceIllustrative sources · No live readingsiFabric / Workload + capacity contextevidence to contextevidence to contextevidence to contextevidence to contextevidence to contextcontext to analyzeanalyze to forecastforecast to recommendrecommend to reviewConstraint foundConstraint foundEvidence gapsEvidence gapsreview to handoffCollect available GPU, network, storage, power and cooling evidence from configured integrations. These are illustrative evidence sources, not live readings.GPU: Usage & healthGPUUsage & healthNetwork: Fabric congestionNetworkFabric congestionStorage: I/O behaviorStorageI/O behaviorPower: Power limitsPowerPower limitsCooling: Cooling capacityCoolingCooling capacityiFabric relates infrastructure signals and dependencies to the affected workload. Compare relevant workload demand and health over a representative period.RelateWorkload contextInvestigate the limiting resource across GPU, fabric, storage and facility constraints. A busy or underused GPU alone does not establish the cause. Evidence gaps go to collection and operator review.Find constraintEvidence across resourcesUse representative demand trends and stated assumptions to evaluate capacity needs against available GPU, network, storage, power and cooling headroom. The forecast is conditional on those inputs.ForecastDemand + headroomPrepare a recommendation grounded in the identified constraint. Growing GPU demand may call for a capacity plan; a storage constraint may call for storage investigation or a supported change instead of more GPUs.RecommendCapacity or constraintThe capacity owner evaluates the evidence, assumptions and proposed impact. An accepted recommendation can enter the configured change approval process; this workflow does not provision capacity or move workloads.ReviewCapacity ownerAn accepted recommendation enters the configured change approval process. Provisioning and workload movement are outside this workflow.Change approvalConfigured change systemCollect missing evidence and keep the capacity question with the operator for review.Collect evidenceOperator reviewRecommendation → configured approvalIllustrative sequence · RepeatsGPU capacity and workload bottlenecksA workload slows or demand grows. Relate GPU, network, storage, power and cooling evidence to workload health, identify the limiting resource and review a capacity recommendation. Recommendations require review. Any capacity or placement change follows the configured approval process. Evidence sources converge on workload context; they are not sequential processing stages. Missing evidence leads to collection and operator review. The diagram ends at recommendation handoff for change approval, without provisioning capacity or moving workloads. Animation timing is illustrative.Configured infrastructure evidenceIllustrative sources · No live readingsiFabric / Workload + capacity contextevidence to contextevidence to contextevidence to contextevidence to contextevidence to contextcontext to analyzeanalyze to forecastforecast to recommendrecommend to reviewanalyze to recommendGapsGapsreview to handoffCollect available GPU, network, storage, power and cooling evidence from configured integrations. These are illustrative evidence sources, not live readings.GPU: Usage & healthGPUUsage & healthNetwork: Fabric congestionNetworkFabric congestionStorage: I/O behaviorStorageI/O behaviorPower: Power limitsPowerPower limitsCooling: Cooling capacityCoolingCooling capacityiFabric relates infrastructure signals and dependencies to the affected workload. Compare relevant workload demand and health over a representative period.RelateWorkload contextInvestigate the limiting resource across GPU, fabric, storage and facility constraints. A busy or underused GPU alone does not establish the cause. Evidence gaps go to collection and operator review.Find constraintEvidence across resourcesUse representative demand trends and stated assumptions to evaluate capacity needs against available GPU, network, storage, power and cooling headroom. The forecast is conditional on those inputs.ForecastDemand + headroomPrepare a recommendation grounded in the identified constraint. Growing GPU demand may call for a capacity plan; a storage constraint may call for storage investigation or a supported change instead of more GPUs.RecommendCapacity or constraintThe capacity owner evaluates the evidence, assumptions and proposed impact. An accepted recommendation can enter the configured change approval process; this workflow does not provision capacity or move workloads.ReviewCapacity ownerAn accepted recommendation enters the configured change approval process. Provisioning and workload movement are outside this workflow.Change approvalConfigured change systemCollect missing evidence and keep the capacity question with the operator for review.CollectEvidence / reviewRecommendation → configured approvalIllustrative sequence · Repeats

Demand grows. Review a capacity plan.

EvidenceRecommendationReview

Recommendations require review. Any capacity or placement change follows the configured approval process.

Illustrative scenario: workload demand and available infrastructure evidence point to GPU capacity pressure. Review a forecast against network, storage, power and cooling constraints, then hand the recommendation to the configured change approval process.

iFabricRelate evidence & investigate
Capacity ownerReview assumptions & recommendation
Change authorityApprove any resulting change

Where it applies

AI Factory · AI Grid

Demand grows or a workload slows.

Measures to define

  • Usable GPU utilization
  • Workload throughput under defined conditions
  • Time to identify a bottleneck
  • Forecast error
Workflow & operating boundaries

A workload slows or demand grows. Relate GPU, network, storage, power and cooling evidence to workload health, identify the limiting resource and review a capacity recommendation.

  1. Gather. Collect available GPU, network, storage, power and cooling evidence from configured integrations. These are illustrative evidence sources, not live readings.
  2. Relate. iFabric relates infrastructure signals and dependencies to the affected workload. Compare relevant workload demand and health over a representative period.
  3. Find constraint. Investigate the limiting resource across GPU, fabric, storage and facility constraints. A busy or underused GPU alone does not establish the cause. Evidence gaps go to collection and operator review.
  4. Forecast. Use representative demand trends and stated assumptions to evaluate capacity needs against available GPU, network, storage, power and cooling headroom. The forecast is conditional on those inputs.
  5. Recommend. Prepare a recommendation grounded in the identified constraint. Growing GPU demand may call for a capacity plan; a storage constraint may call for storage investigation or a supported change instead of more GPUs.
  6. Review. The capacity owner evaluates the evidence, assumptions and proposed impact. An accepted recommendation can enter the configured change approval process; this workflow does not provision capacity or move workloads.

This is an illustrative operating workflow, not live telemetry or a measured customer result. Define the workload, observation period, evidence quality, capacity constraints and forecast assumptions for each environment. Source availability depends on the configured integrations. nLLM can support infrastructure investigation where deployed. A recommendation that changes placement also needs the configured placement policy and authorization; gFabric can supply that governance where deployed. Changes use a separate, approved customer or partner process. Animation timing is illustrative. Repetition restarts the explanation, not a provisioning or placement action.

Scenarios

Capacity pressure. Illustrative scenario: workload demand and available infrastructure evidence point to GPU capacity pressure. Review a forecast against network, storage, power and cooling constraints, then hand the recommendation to the configured change approval process.

Storage bottleneck. Illustrative scenario: workload and storage evidence indicate that data delivery is limiting progress. Prepare a recommendation to investigate or address the storage constraint. The capacity owner reviews it before any proposed change enters approval.

Insufficient evidence. Illustrative scenario: missing or inconsistent workload and infrastructure evidence prevents a supported conclusion. Record the gap, collect relevant evidence and return the question to the operator. No capacity recommendation is treated as approved.