An institution can retain control of its signing keys while losing the ability to complete a digital asset service. A transaction policy service may be unavailable, a provider may stop returning reliable status, or a funding route may miss its operating window. The continuity problem is the business service and its dependencies, not only access to cryptographic material.
A useful plan begins with the obligation the institution must meet and the time available to meet it. It then identifies what must remain working, what can be substituted and who can approve a change. This article proposes a vendor continuity process for payment, transaction and data services. It treats the examples as operating scenarios, not accounts of named vendor failures.
Define continuity at the business-service level
Name the service in terms an operating manager can assess. A payment service needs to accept an authorized instruction, execute it through an approved route and produce usable evidence. A reporting service needs to deliver an interpretable record by a deadline. A system can be reachable while failing those outcomes.
The Basel Committee's March 2021 operational resilience principles address banks' delivery of critical operations through disruption. They include dependency mapping, testing and third-party contingency arrangements. This is banking guidance, used here as a source for organizing the operating questions rather than asserting that every digital asset firm is subject to the same requirements.
Set the institution's own disruption tolerance for each service through its governance process. Include the deadline, acceptable degraded output and the point where a service must be suspended. An uptime percentage alone may not capture the impact of a short outage that occurs immediately before a settlement or payroll deadline.
Record the people and information needed to make the continuity decision. A fallback can be technically available yet unusable because the approver is absent or the operator cannot retrieve the approved instruction. The plan should include those dependencies alongside the technology.
Map the dependencies that deliver the outcome
Follow the workflow through its providers. A transaction might depend on identity access, destination verification, a policy engine, signing, network submission and status retrieval. Funding may depend on another account or conversion service. Map which provider performs each step and which internal team owns the relationship.
Identify shared dependencies between primary and alternative routes. Two providers can use the same cloud region, upstream data source or account relationship. Ask for evidence relevant to the claim that a fallback is independent. The institution needs to understand whether the alternative survives the particular scenario it is meant to address.
Our article on RPC infrastructure dependencies examines one layer of that map. A continuity plan should extend to the surrounding service, including authorization and reconciliation. Switching an endpoint can restore network access without resolving an unavailable approval or status service.
Keep the map maintained. Vendor architecture, account arrangements and subproviders can change during the contract. Use material-change reviews to ask whether the fallback assumptions remain true. A diagram created during procurement becomes a weaker basis for action if no one updates it after the operating route changes.
Distinguish temporary disruption from supplier exit
A short service interruption may call for queueing work or using a preapproved alternative. A longer interruption or the end of a supplier relationship may require migration. Define separate triggers and owners for those decisions. A plan that can bridge an hour does not automatically provide a workable exit from the service.
NIST's contingency planning guide, updated in November 2010, addresses priorities and relationships among information-system contingency plans. The source is dated and federal in scope. Its practical relevance here is the need to connect recovery choices to the operation's requirements, rather than applying one recovery target to every component.
For a potential exit, identify what the institution must retrieve: instructions, policy settings, historical evidence, account records and configuration definitions. Determine how long the alternative setup takes and which tasks still depend on supplier cooperation. Contractual access can support the plan, while the operating team must test whether the exported information is sufficient.
Prepare a fallback before the incident
A fallback route should have completed the institution's relevant approval, security and operational reviews. Staff need the receiving details, account access, supported assets and capacity information in advance. An emergency should activate the route, not begin the entire vendor onboarding process.
Define what can move to the fallback. It may support a narrower set of networks, lower volumes or a different approval method. Document those limits in the service plan. The institution can make a deliberate degraded-service decision, but the operator should not discover the limits after redirecting a critical instruction.
Verify the funds and data needed for the alternative. A backup payment provider may be available while its account remains unfunded. A backup data feed may cover live observations without historical identifiers needed for reconciliation. Test the output that the business actually needs before calling the route ready.
Where manual processing forms part of the plan, determine its sustainable capacity. Review staffing, supervision and the evidence produced. A manual procedure that handles one test instruction may not meet the deadline for an entire queue. Continuity planning should preserve control as well as throughput.
Resolve pending work before changing routes
The difficult boundary is an instruction whose execution outcome is uncertain. A vendor can stop responding after accepting it. Repeating the instruction through another route may create a duplicate if the original later completes. The plan should establish who investigates that state and what evidence is needed before resubmission.
Maintain an institution-controlled instruction ledger where appropriate to the service. It should connect the business intent with provider references and observed outcomes. During disruption, the operating team needs a way to distinguish new work from already submitted work. That distinction should survive the loss of the vendor's portal.
If the first route cannot be resolved within the business deadline, escalate the decision. The designated owner should understand the possible duplicate exposure, any cancellation capability and the process for reconciling a later result. Record the approved action and its assumptions. A hurried switch by the wallet operator should not silently decide the institution's commercial position.
Include the supplier in exercises
NIST's supply chain quick-start guide calls for relevant suppliers to participate in incident planning, response and recovery. For a continuity exercise, agree what the supplier will simulate, which evidence it will provide and how its escalation channel will operate. A test limited to the institution's staff cannot establish that the supplier's promised handoff works.
Choose scenarios that affect a real deadline. Test an unavailable authorization service, a stale status feed, an unresponsive support route and a supplier interruption that lasts longer than the plan's initial assumption. Include an unavailable internal decision-maker. Ask the team to complete the supported service outcome, with evidence, inside the intended window.
Separate observation from intervention during the exercise. If a vendor engineer quietly repairs the situation in a way the production process would not support, record that limitation. The result should describe what the operating team could do with its actual access and escalation rights.
Our coverage of digital asset recovery drills addresses recovery of the ability to move assets. Vendor continuity adds the service dependencies needed to decide, execute and explain that movement. Exercises should connect those capabilities while preserving their separate acceptance criteria.
Restore the primary route with a reconciliation gate
A supplier announcing restoration is an input to the institution's decision. Confirm the service's state through the relevant operating checks. Identify the backlog, changes made during recovery and any missing or revised records. A restored interface may still contain stale information that cannot support a safe resumption.
Reconcile work handled during the interruption before moving all traffic back. Track which instructions used the alternative and which remained on the original route. Check for late completions and duplicates. Set a clear boundary for new instructions so two operating teams do not resume different routes simultaneously.
Retain the incident decision log and compare the result with the plan's assumptions. Did the fallback have the required capacity? Did staff retrieve the instruction record? Was supplier support available through the agreed channel? Assign corrective actions to the owner of the relevant dependency, with evidence needed to close them.
The continuity plan becomes credible when the institution can still deliver its defined service through a plausible supplier disruption and explain every instruction handled along the way. Possession of keys remains important. The service also needs people, usable records, approved routes and a tested decision process that keeps them working together.