AI operations

The AI Continuity Plan: Design the Workflow to Survive Its Model, Tool, and Vendor

How to preserve an essential business outcome, its evidence, and Human authority when an AI dependency changes, degrades, or disappears.

Current Operating Signal

What changed: current model retirements and a recent regional cloud incident make dependency change concrete.

Executive question: can the essential business outcome continue safely when the preferred AI path cannot?

An AI workflow is not production-ready until the business knows how to continue without its preferred AI path.

A model can retire. A region can fail. A tool can change its interface, a provider can alter access, or a rate limit can arrive during the hour when the work matters most. None of these events needs to destroy the business outcome. They become existential only when the organization has mistaken a dependency for the operating system.

That distinction is immediate, not theoretical. Amazon Bedrock currently lists several model end-of-life dates in September 2026, including Claude 3 Haiku on September 10, Nova Premier and Nova Sonic on September 14, and Nova Canvas and Nova Reel on September 30. AWS states that migration is not automatic and that requests to an end-of-life version will fail on or soon after the date.1 These dates apply to Bedrock and should be checked again before anyone acts on them. Their broader lesson is durable: availability has a lifecycle.

Infrastructure can also fail across several dependencies at once. Google reported that an August 20 incident in its us-west1 region degraded a wide range of services for two hours and twenty-two minutes. The reported cause began with reduced network capacity and included failed automated rerouting; downstream effects reached compute, storage, identity, databases, and other services.2 A list of alternate AI models would not, by itself, have solved an outage affecting the surrounding environment.

Model redundancy protects a call. Continuity protects the outcome the business is responsible for.

Define Continuity at the Business Boundary

Technical teams often begin continuity planning by inventorying components. That is necessary; it is not the first decision. The first decision is what the business must still accomplish when normal operation is no longer available.

An accounts-receivable workflow, for example, may normally collect records, classify exceptions, draft a customer message, and recommend the next action. If its model is unavailable, continuity does not necessarily mean reproducing every step through another model. It may mean preserving the queue, showing the evidence, identifying urgent cases, and returning the decision to an authorized person. The experience is slower; the essential promise remains intact.

AWS frames graceful degradation around the core business function and predictable, recoverable states. It also warns that failure paths need testing and should be significantly simpler than the primary path.3 Google similarly describes a degraded system as one that remains useful with reduced performance or accuracy while partial errors, overload, monitoring, and recovery are managed deliberately.4 The business question beneath both forms of guidance is plain: what can be reduced without violating the promise?

External model, cloud, identity, and tool dependencies connect to a business-owned evidence core that reroutes work after one dependency fails
Editorial Visualization · Dependency Is Not OwnershipKeep the Outcome Above the Dependency Graph.

A provider can supply capability. The business still owns the context, evidence, authority, degraded path, and decision to recover.

Write an AI Continuity Contract

I use the term AI Continuity Contract for a compact operating agreement attached to one consequential workflow. It is not a vendor contract or a promise of uninterrupted service. It is the business-owned record of what must continue, which compromises are safe, who can take control, what evidence must survive, and how normal operation earns its return.

01 · Outcome

Name the Essential Result

State the smallest business result that must remain available. Separate it from conveniences, preferred interfaces, and the model used today.

02 · Degraded Mode

Choose the Safe Reduction

Define what the workflow may omit, delay, simplify, or route differently without producing an unacceptable customer, financial, legal, or operational consequence.

03 · Authority

Name the Manual Owner

Identify the person or role that can pause automation, accept a manual result, communicate the limitation, and decide whether work should continue.

04 · Evidence

Preserve the Reconstruction Record

Keep the request, accepted sources, system state, completed steps, unresolved items, and Human decisions available even when a dependency is not.

05 · Recovery

Require Proof Before Return

Define the checks, sample work, reconciliation, and responsible acceptance needed before restored automation can regain its former authority.

This contract begins with business impact, not infrastructure preference. NIST contingency guidance similarly starts by identifying essential functions and recovery priorities, then develops strategies, testing, and plan maintenance around them.5 That publication addresses federal information systems. I am borrowing its disciplined sequence, not claiming that it prescribes this AI model for every private organization.

Operate Through Four Visible States

A continuity plan becomes useful when people can recognize which state the workflow occupies and what authority belongs in that state. Ambiguous failure is dangerous because the interface can continue looking complete while the evidence, tools, or checks underneath it are incomplete.

  1. Normal: the approved path is healthy; current dependencies, checks, and operating limits are available.
  2. Degraded: a dependency or quality condition has failed; a tested, simpler path provides a reduced result within explicit boundaries.
  3. Manual: automation cannot preserve the acceptable boundary; work pauses or transfers to the named Human owner with the evidence available.
  4. Recovery: the dependency appears healthy; reconciliation, verification, and limited test work occur before normal authority returns.

These states are not merely labels for an incident channel. They should change the interface, allowed actions, evidence requirements, customer communication, and acceptance path. A visible degraded state is more honest than a normal-looking system producing results from incomplete context.

Test the Decision, Not Only the Endpoint

Teams regularly test whether a backup endpoint responds. The harder question is whether the workflow still supports the decision the business needs to make. A second model can return fluent text while using different context limits, tool behavior, safety policies, regional dependencies, or economic assumptions. The endpoint is alive; the operating contract may still be broken.

NIST's initial TEVV-Athlon draft describes a customizable method for measuring AI systems against organizational objectives and real-world impacts. It explicitly includes large language models and agentic systems.6 The document remains an initial public draft. Its useful implication here is that continuity tests should be constructed around the organization's objective, not the provider's generic benchmark.

A practical continuity exercise should answer five questions with observable evidence:

  1. Detection: did the system recognize the dependency or quality failure before presenting an ordinary result?
  2. Transition: did it enter the correct degraded or manual state without expanding authority?
  3. Preservation: can a person reconstruct completed, incomplete, and uncertain work?
  4. Usefulness: did the reduced path preserve the essential business outcome within its accepted boundary?
  5. Recovery: did reconciliation and Human acceptance occur before normal operation resumed?
Operators compare preserved before-and-after evidence while a restored automated workflow passes through a verification gate beside a manual control station
Editorial Visualization · Recovery Is a DecisionRestored Service Must Earn Restored Authority.

A green provider status is an input. Recovery is complete only after interrupted work is reconciled, the path is tested, and an accountable person accepts the return.

Portability Is More Than a Second Provider

Provider choice matters. A well-designed system should be able to route suitable work to another compatible model or tool when the economics, availability, or requirements justify it. Yet portability fails when the organization cannot move the accepted context, permissions, evidence format, quality checks, Human review, and recovery history with the work.

This is where an assurance architecture sits above individual capabilities. In the public description of the Contextual Pipeline Framework, the important responsibilities are accepted context, bounded execution, evidence continuity, verification, correction, and Human acceptance. The article does not disclose private routing mechanics, schemas, scoring, or compression methods. The relevant principle is enough: dependencies may change while the business-owned contract remains coherent.

That principle also prevents false confidence in multi-provider designs. Two model vendors may depend on the same identity service, cloud region, data connector, or Human reviewer. A complete dependency map includes the surrounding systems and the organizational capacity required to use them.

Make the Manual Path a First-Class Product

Manual operation is often treated as an emergency improvisation. That makes the worst moment the first time people discover missing permissions, inaccessible records, undocumented exceptions, or a decision owner who is unavailable.

A first-class manual path has a visible entry condition, a bounded queue, readable evidence, qualified owners, and a way to record what happened. It should not imitate the speed of automation. Its value is that it remains understandable under pressure.

NIST's Generative AI Profile recommends incident response and recovery plans, communication of workarounds and alternate processes, and mechanisms to supersede, disengage, or deactivate AI systems.7 The profile is voluntary guidance. It reinforces a practical standard: the ability to stop or route around AI is part of operating AI responsibly.

What This Plan Cannot Promise

An AI Continuity Contract cannot guarantee zero downtime, eliminate correlated failures, preserve a service that has no safe reduced mode, or turn an untested manual process into reliable recovery. Some workflows should stop completely when a required source, verifier, specialist, or control is unavailable. Continuity sometimes means preserving the queue and communicating honestly, not completing the transaction.

Provider documentation is valuable for understanding stated lifecycle and architecture behavior; it is not independent proof of resilience. A recent incident illustrates one failure pattern; it does not predict the next one. NIST guidance supplies a disciplined vocabulary; it does not certify this operating model or any VerShep implementation.

Design for the Day the Preferred Path Is Gone

The mature AI question is not only which model performs best today. It is whether the organization can protect the result, the customer, and the responsible people when today's model is unavailable tomorrow.

Choose one consequential workflow. Write its essential outcome, safe degraded mode, manual authority, evidence-preservation rule, and recovery proof. Then interrupt a real dependency in a controlled exercise. The gaps you find are not reasons to abandon AI; they are the work required to own it.

Start With One Essential Outcome

Test the Workflow Before the Dependency Tests the Business.

I can help your team map one consequential workflow, define its continuity contract, and test whether the degraded, manual, and recovery paths preserve what the business is actually responsible for.

Research Record

References and Evidence

Sources were reviewed on September 7, 2026. Dates and service states are attributed to current provider documentation rechecked on publication day. The incident description is limited to Google's published account of one regional event. NIST materials are voluntary guidance or an initial public draft; they do not endorse the AI Continuity Contract, CPF, or VerShep. The contract, four operating states, worked example, and testing questions are my architectural interpretation, not measured customer results or a guarantee of availability.

  1. Model lifecycle: Amazon BedrockVendor Documentation · Amazon Web Services; reviewed September 7, 2026

    AWS documents Active, Legacy, and End-of-Life states. It states that migration is not automatic and requests to an EOL model will fail on or soon after the EOL date. The September dates cited in this essay are specific to Amazon Bedrock, were rechecked on publication day, and may change.

  2. Google Cloud us-west1 multi-service incidentVendor Incident Report · Google Cloud Service Health; incident August 20 and report August 27, 2026

    Google reports a two-hour-and-twenty-two-minute disruption across many services in us-west1, caused by reduced network capacity and failed automated rerouting. Services outside us-west1 were reported as unaffected. This is Google Cloud's account of one regional event, not a general failure-rate estimate.

  3. Implement graceful degradation to transform applicable hard dependencies into soft dependenciesVendor Architecture Guidance · AWS Well-Architected Reliability Pillar

    AWS recommends identifying core business functionality, designing predictable and recoverable degraded states, and testing simpler failure paths. The guidance also cautions that fallback strategies should generally be avoided unless designed carefully.

  4. Design for graceful degradationVendor Architecture Guidance · Google Cloud Architecture Framework

    Google describes graceful degradation as continuing to provide useful service with reduced performance or accuracy under stress, supported by overload handling, partial-error handling, monitoring, and testing. It is provider guidance, not independent evidence that a specific design is resilient.

  5. Contingency Planning Guide for Federal Information SystemsGovernment Contingency Guidance · NIST Special Publication 800-34 Revision 1

    NIST presents a contingency-planning process built around business impact analysis, recovery priorities, strategies, testing, and plan maintenance. It is federal information-system guidance; this essay adapts its continuity logic rather than claiming it directly governs private-business AI workflows.

  6. The TEVV-Athlon Framework for Evaluating AI SystemsInitial Government Standards Draft · NIST AI 200-2 Initial Public Draft; August 2026

    NIST describes a customizable four-stage method for measuring whether AI systems meet organizational goals while minimizing negative impacts. The draft includes LLMs and agentic systems and remains open for public comment through October 6, 2026; it is not a final standard.

  7. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileVoluntary Government Guidance · NIST AI 600-1; July 2024

    The profile recommends incident response and recovery planning, communication of workarounds and alternate processes, and mechanisms to supersede, disengage, or deactivate AI systems. It is voluntary guidance and does not validate the continuity contract proposed here.

Read nextBefore Your Business Makes an AI Claim, Build the Receipt

Continue the conversation

Good ideas improve under pressure.

If this model resembles something you are seeing in practice, or fails to account for it, I'd value the conversation.