Menu

Use case / 01 · iFabric

From incident to recovery.

A known service failure. A governed response.

All use cases

Known failures · Configured integrations

Known failure recovery workflowAn illustrative iFabric workflow for a VM or service that stops responding. Correlate operational evidence, propose a supported recovery, authorize execution and verify service health using fresh telemetry. Automatic action requires sufficient evidence and authorization. Insufficient evidence, denied authorization or failed verification lead to human review. nLLM and gFabric participate where deployed. Animation repeats the illustration, not the recovery action.observe to correlatecorrelate to reasonreason to authorizeauthorize to actact to verifyverify to healthyInsufficient evidenceInsufficient evidenceAuthorization deniedAuthorization deniedRecovery failedRecovery failedDetect an unresponsive VM or service from operational signals.01ObserveTelemetryRelate the signal to infrastructure topology, logs and affected services.02CorrelateTopology + logsUse evidence and approved runbooks to propose a recovery for a known failure. nLLM can assist where deployed; an uncertain diagnosis goes to human review.03ReasonPropose recoveryCheck the proposed action against configured identity and action policy. gFabric can provide this authorization where deployed. Denial stops automated execution.04AuthorizeConfigured policyInvoke only the authorized action through a supported operational integration.05ActConfigured executorCheck fresh service-health signals against the defined recovery criteria. An executor acknowledgement alone is not proof of recovery. Failed verification goes to human review.06VerifyFresh telemetryHuman review. An operator evaluates the evidence and decides the next step.Human reviewOperator decidesService health confirmed from fresh telemetryHealth confirmediFabric / Agentic operationsIllustrative sequence · RepeatsKnown failure recovery workflowAn illustrative iFabric workflow for a VM or service that stops responding. Correlate operational evidence, propose a supported recovery, authorize execution and verify service health using fresh telemetry. Automatic action requires sufficient evidence and authorization. Insufficient evidence, denied authorization or failed verification lead to human review. nLLM and gFabric participate where deployed. Animation repeats the illustration, not the recovery action.observe to correlatecorrelate to reasonreason to authorizeauthorize to actact to verifyverify to healthyUncertainUncertainDeniedDeniedFailedFailedDetect an unresponsive VM or service from operational signals.01ObserveTelemetryRelate the signal to infrastructure topology, logs and affected services.02CorrelateTopology + logsUse evidence and approved runbooks to propose a recovery for a known failure. nLLM can assist where deployed; an uncertain diagnosis goes to human review.03ReasonPropose recoveryCheck the proposed action against configured identity and action policy. gFabric can provide this authorization where deployed. Denial stops automated execution.04AuthorizeConfigured policyInvoke only the authorized action through a supported operational integration.05ActConfigured executorCheck fresh service-health signals against the defined recovery criteria. An executor acknowledgement alone is not proof of recovery. Failed verification goes to human review.06VerifyFresh telemetryHuman review. An operator evaluates the evidence and decides the next step.Human reviewOperator decidesService health confirmed from fresh telemetryHealth confirmediFabric / Agentic operationsIllustrative sequence · Repeats

Detect. Diagnose. Authorize. Recover. Verify.

EvidenceProposalActionReview

Automatic action requires sufficient evidence and authorization.

A supported failure pattern has sufficient evidence and an authorized recovery action. The configured executor acts; fresh telemetry confirms service health.

Where it applies

AI Factory · AI Grid · Edge

Infrastructure and service operations.

Measures to define

  • Time to diagnose
  • Time to restore
  • Failed recoveries
  • Escalation rate
Workflow & operating boundaries

An illustrative iFabric workflow for a VM or service that stops responding. Correlate operational evidence, propose a supported recovery, authorize execution and verify service health using fresh telemetry.

  1. Observe. Detect an unresponsive VM or service from operational signals.
  2. Correlate. Relate the signal to infrastructure topology, logs and affected services.
  3. Reason. Use evidence and approved runbooks to propose a recovery for a known failure. nLLM can assist where deployed; an uncertain diagnosis goes to human review.
  4. Authorize. Check the proposed action against configured identity and action policy. gFabric can provide this authorization where deployed. Denial stops automated execution.
  5. Act. Invoke only the authorized action through a supported operational integration.
  6. Verify. Check fresh service-health signals against the defined recovery criteria. An executor acknowledgement alone is not proof of recovery. Failed verification goes to human review.

This is a reference operating scenario, not a customer case study or a live incident. Supported telemetry sources, runbooks, authorization rules, executor integrations and service-health criteria are agreed for each deployment. Animation timing does not represent processing latency. Repetition restarts the illustration; it does not repeatedly restart a restored service.

Exception routes

Uncertain cause. Reasoning cannot establish a supported recovery with sufficient evidence. The proposal goes to human review. No automated action is taken.

Action denied. The proposed action fails the configured authorization check. The workflow goes to human review without invoking the executor.

Recovery failed. The action is authorized and executed, but fresh telemetry does not satisfy the recovery criteria. The operator receives the evidence for review; no automatic retry is implied.