Skip to content

TechDirectArchive

Hands-on IT, Cloud, Security, Veeam & DevOps

  • Home
  • About
  • Advertise With US
  • Reviews
  • Contact
  • Toggle search form

Your DR Network Is Probably the Part You Haven’t Tested

Posted on 28/09/202628/09/2026 Eric Black By Eric Black No Comments on Your DR Network Is Probably the Part You Haven’t Tested
  1. Home
  2. Network | Monitoring
  3. Your DR Network Is Probably the Part You Haven’t Tested
232323232
A recovered workload is not a recovered service until DNS, routing, firewall policy, load balancing, private connectivity, management access, and application dependencies work from the recovery site.

Most DR tests spend a lot of time proving that a workload can be restored and surprisingly little time proving that anything can reach it afterward. A VM can boot cleanly in the recovery site, pass a guest heartbeat, and still be useless because the route is missing, DNS still resolves to production, the firewall policy never made it over, or the load balancer has no healthy recovery targets.

The network is often treated as something the recovery environment will already have. That assumption is where otherwise solid recovery plans start turning into live troubleshooting exercises.

Recovery test: If the application only works when someone connects directly to the recovered VM’s IP from the recovery console, you have proved compute recovery. You have not proved service recovery.

A Green VM Is Only One Milestone

Powering on the recovered workload matters. So does getting the operating system up, attaching the right disks, starting the application services, and confirming the data is usable. None of those checks proves that a client can reach the service through the path the production application normally uses.

Recovery platforms expose this separation pretty clearly when you look at how they handle networking. Azure Site Recovery maps source networks to target VNets and lets the recovery configuration select target subnets and addressing. AWS Elastic Disaster Recovery uses launch settings and an EC2 launch template to define the recovery VPC, subnet, and security groups. Those are recovery inputs that sit beside the replicated workload, not inside it. [1] [5]

Architecture point: The recovered VM and the recovered network are separate parts of the same service. A DR test that validates only one of them leaves the application contract half tested.

A recovery test should follow the same path the application will use during an incident. Start with the name or VIP the user, API client, or upstream system will actually call. Follow that path through DNS, load balancing or reverse proxying, firewall policy, NAT, routing, the recovery subnet, and finally the application. Then test the application’s outbound and east-west dependencies in the opposite direction.

The DR Network Is a Chain of Dependencies

For most applications, there is no single DR network dependency. There is a chain.

The north-south path can include authoritative and recursive DNS, a global traffic manager, a public or private VIP, a load balancer, a reverse proxy, edge firewall rules, NAT, a VPN or private circuit, internal routing, a recovery VLAN or subnet, and the default gateway used by the recovered host. The application can be running perfectly behind any one of those broken hops.

The east-west path is usually less visible during a basic failover test. The recovered application may need Active Directory, DNS, a database, message queue, cache, file service, API, license server, certificate service, monitoring collector, or another application tier that lives somewhere else. A web tier that returns a local health page is not enough if its database connection still points to an unreachable production address.

There is a third path that gets missed just as often: management. Operators still need access to the hypervisor or cloud control plane, jump hosts, bastions, firewall management, DNS administration, monitoring, logs, consoles, and recovery orchestration. If those paths depend on the site that just failed, the recovery procedure can stall before the application traffic even becomes the problem.

Addressing Is Where the Diagram Starts Lying

IP behavior deserves its own DR test because the recovery address is not always the production address. Microsoft documents that Azure Site Recovery uses the same source IP in the target when source and target subnets use the same address space and the address is available. With different address space, the service uses an available address in the target subnet. Test failover can also receive a different address when the preferred address is already in use. [1]

If the address changes, anything tied to the old address may now be wrong: firewall objects, NAT rules, monitoring targets, reverse proxy backends, load balancer pools, access control lists, allowlists, scripts, or application configuration.

Hard coded IPs are the obvious version of this problem, but they are not the only version. A DNS name can still resolve correctly while a firewall rule references the production subnet. A load balancer can still have the right hostname while its pool members remain production addresses. A monitoring server can report the application down because it is checking the old IP even though the recovered service is healthy.

Address overlap creates a different failure. Microsoft explicitly notes that retaining the same on-premises address space in Azure prevents site-to-site VPN or ExpressRoute connectivity between the overlapping networks, and that routing must change when the subnet moves to Azure. RFC 1918 makes the underlying problem clear: private address space is only useful while address uniqueness is maintained between networks that need to communicate. [2] [10]

If production and recovery both use 10.20.30.0/24, the question is not whether both networks can exist. The question is what happens when they must be reachable at the same time during a test, partial failover, staged migration, or dependency call. The answer has to be designed before the incident.

Re-IP Has to Include Everything Around the Workload

A re-IP plan that stops at the guest NIC is incomplete. The recovered address has to make sense everywhere the service is referenced.

I would document the translation as a small set of linked changes: production subnet to recovery subnet, production IP to recovery IP, production VIP to recovery VIP, DNS record behavior, NAT behavior, load balancer pool membership, firewall object changes, routing changes, and any monitoring or management targets that use the address directly. Keeping those changes together makes it easier to apply the same translation everywhere.

Test networks need the same address discipline. If the test uses an isolated subnet with addresses that do not match the actual recovery plan, the test can prove boot behavior while skipping the real cutover problem. If the test uses the real recovery subnet, it needs enough isolation to prevent duplicate addresses, duplicate services, or unintended connections back into production.

4343222

Example re-IP translation. On narrow screens, scroll the diagram horizontally. The addresses are illustrative; the useful part is keeping every dependent network reference aligned with the recovery plan.

DNS Is Part of the Cutover

DNS changes are often placed near the end of a runbook as if they were a cleanup step. They are part of the recovery path.

RFC 1035 defines TTL as the amount of time a resource record may be cached before the source should be consulted again. Updating the authoritative record does not force every resolver and client to forget the previous answer immediately. [9] AWS makes the same issue operational in Application Recovery Controller, where DNS records are used to shift traffic between application replicas and AWS recommends lower TTL values for records involved in failover. [7] [8]

A DR plan should therefore know which names change, which names remain stable behind a traffic manager, what the TTLs are, and which resolvers clients actually use. Split horizon DNS, internal forwarding, conditional forwarders, private hosted zones, and local DNS caches can all produce a different answer depending on where the test originates.

I would test name resolution from at least three places when they exist: a normal client network, the recovered application network, and an administration network. Query the production service name, dependency names, and any recovery names. Confirm the returned addresses are the ones the recovery design expects, then make the application connection using the name rather than a direct IP.

A direct IP test is useful for diagnosis. It is a bad acceptance test for a service that users normally reach by name.

Firewall Policy Does Not Follow a VM by Magic

Firewall and security policy needs an explicit recovery method. Microsoft states that Azure Site Recovery does not create or replicate Network Security Groups during failover and recommends creating the required target NSGs before failover, then validating them with a test failover. [3]

Outside Azure, recovery still needs an explicit way to recreate the effective policy at the recovery location. It might be infrastructure as code, firewall configuration replication, a recovery script, an SDN policy engine, or a manual change. Whatever the method is, the recovery test should prove the resulting flow, not merely confirm that a rule object exists in a management console.

For each application path, capture source, destination, protocol, port, and direction. Then test the flow from the side that will actually initiate it. A successful inbound client connection does not prove the application can make an outbound database connection. A stateful firewall session in one direction does not prove the return path uses the same device or route.

The evidence I want is simple: connection from the expected source zone, to the expected destination name or address, on the expected port, through the recovery path that will exist during a real event.

Load Balancers and Reverse Proxies Need Recovery State Too

A load balancer can make a recovered application look dead even when every VM behind it is healthy. The VIP may still live only in production. The recovery pool may be empty. Health probes may use the wrong port or host header. The reverse proxy may still resolve backend names to production addresses.

Microsoft’s current Site Recovery networking guidance supports preconfigured internal load balancers, public IPs, secondary IPs, and NSGs for failover VMs. For internal load balancers, the backend pool and frontend configuration must already exist. [4] Regardless of platform, the VIP, listener, health check, and backend pool have recovery state that exists outside the guest.

The test should hit the application through the same front door the real client uses. If production uses app.example.com through a VIP and reverse proxy, testing https://172.20.30.25 from a jump box is only a troubleshooting step. The acceptance test should use https://app.example.com, observe which VIP answers, confirm the recovery backend is selected, and verify the application response.

For TLS services, include the certificate and hostname path in the test. A backend can be reachable while SNI, certificate selection, trust, or reverse proxy routing is wrong.

Routes, Gateways, VPNs, and Private Connectivity Need Their Own Proof

Routing is another place where a diagram can look complete while the forwarding path is not. The recovery subnet can exist, the gateway can be configured, and the VPN can show up while the application prefix is still missing from the route table or advertised to the wrong place.

If the design uses BGP, prove that the recovery prefixes are learned where they need to be learned. If it uses static routes, prove the next hop is correct after failover. If it uses a site-to-site VPN or private circuit, test the application prefixes that matter rather than accepting tunnel state as proof of service reachability.

The overlap case is especially dangerous during partial recovery. A database that stays in production while the application tier fails over may require routed connectivity between sites. If both sides advertise the same subnet, the network cannot infer which copy of the address should receive the traffic. The recovery sequence must account for which copy of each prefix is reachable at every stage.

A full site recovery can sometimes simplify routing because an entire prefix moves together. An application level failover can be harder because one tier moves while another remains behind. The DR runbook needs to match the recovery granularity you actually plan to use.

Management Access Is Part of the Recovery Design

Recovery procedures usually assume an operator can still reach the systems that perform the recovery. That assumption needs a test.

AWS recommends keeping purpose built recovery credentials accessible and using the highly available routing control data plane for failover actions instead of depending on the console. [8] Outside AWS, test the same dependency: can the team still make the network changes if the normal identity path, management network, VPN concentrator, jump host, or admin portal is unavailable?

I would test the management route separately from the application route. Can the recovery team reach the recovery orchestrator? Can it log into the firewall, load balancer, DNS platform, cloud account, hypervisor, and monitoring system? Can it do that from the location the team expects to use during a site outage?

A DR plan that needs production Active Directory to authenticate into the system that recovers production Active Directory has a dependency loop. The same problem can exist with DNS, PKI, MFA, privileged access systems, and remote access gateways.

Test the Recovered Application From Its Own Point of View

AWS Elastic Disaster Recovery drills use the same source server launch settings and point-in-time snapshots as recovery, and AWS provides post-launch checks that can validate network connectivity to specified hosts and ports and verify HTTP or HTTPS responses. [5] [6]

DR network tests should originate where the recovered workload will actually run and where its clients will actually connect. From the recovered application tier, test the database, directory services, DNS, APIs, queues, object stores, file services, license servers, certificate endpoints, and any SaaS dependencies. From the client side, test the service through the production name and front door.

Monitoring should be part of that proof. Confirm the recovered workload can send logs and metrics and that the monitoring platform can reach whatever it needs to reach. A service that works but disappears from monitoring is recoverable only until the next problem occurs.

Do the same for backup and security controls when they matter to the application after failover. If the recovered workload runs for several days, can it still be backed up, scanned, patched, and managed from the recovery site?

Build the Recovery Test Around Flows, Not Boxes

During a recovery test, a communication matrix tied to the runbook is usually more actionable than a giant topology drawing.

For each critical flow, record the initiating system or zone, destination name, destination service, protocol and port, production path, recovery path, whether the address changes, whether DNS changes, which firewall or security policy owns the rule, whether NAT is involved, and who owns the change. Then add one more column: how the team proves it during a test.

A practical validation sequence looks like this:

  1. Map the application path. Start with the real client entry point and trace every network hop to the recovered workload.
  2. Map application dependencies. Trace outbound and east-west calls from the recovered workload to every service it needs.
  3. Define recovery translation. Record subnet, IP, DNS, NAT, load balancer, firewall, and route changes together.
  4. Stage the target network. Build the VLANs, subnets, gateways, security policy, load balancer objects, private connectivity, and management access the plan expects.
  5. Run the recovery. Use the actual recovery network or a test network that faithfully represents the same translation behavior.
  6. Test north-south and east-west flows. Use names and front doors where production uses names and front doors. Test required ports from the initiating side.
  7. Test management and monitoring. Prove the operators and observability systems can reach the recovered environment without depending on the failed site.
  8. Capture evidence and fix the runbook. Record what had to change, who made the change, how long it took, and what should be automated before the next test.

DR networking tends to accumulate manual knowledge. Someone remembers which static route has to be added, which DNS record needs a lower TTL, or which firewall object has to be swapped. If the test discovered it, put it in the runbook while the evidence is still fresh.

What I Would Require Before Signing Off a DR Network Test

AreaWhat to proveTest fromFailure it catches
DNSProduction service names and dependency names resolve to the intended recovery targetsClient, recovery workload, admin networkStale records, wrong forwarding, split DNS mistakes
North-southUsers or upstream systems can reach the service through the normal VIP, proxy, or gatewayRealistic client networkMissing routes, NAT, firewall, VIP, proxy, or load balancer state
East-westApplication tiers can reach databases, directory services, APIs, queues, and other dependenciesRecovered workloadMissing internal routes, rules, hard coded production IPs
EgressRecovery workloads can reach required external services through the intended source NAT and security pathRecovered workloadBroken NAT, egress policy, provider routing, allowlists
ManagementOperators can reach recovery tooling, network controls, consoles, and jump hostsRecovery admin locationDependency on failed management or identity path
MonitoringLogs, metrics, health checks, and alerts work from the recovery siteRecovery workload and monitoring platformSilent recovery environment, old monitoring targets

I would not sign off because every VM powered on. I would sign off when the service name resolves correctly, intended clients can use the application, the application can use its dependencies, the team can manage it, and monitoring sees it through the recovery paths that will exist during a real event.

The Bottom Line

Recovery testing has a natural bias toward the things the recovery platform can show on a status screen: restore completed, failover completed, VM started, guest tools running, application process healthy. Those are necessary milestones, but they stop at the edge of the recovered workload.

The network test begins there. Routing, VLANs and subnets, IP translation, DNS, firewalls, NAT, load balancers, reverse proxies, gateways, VPNs, private connectivity, management access, monitoring, and external dependencies all determine whether the recovered service can actually do its job.

Test the service from the perspective of the systems that need to use it. If the only proven path is from a recovery console to a VM IP, the DR network is still an assumption.

Key Takeaways

  • A VM starting successfully proves workload recovery, not end-to-end service reachability.
  • DR networking includes client entry paths, east-west dependencies, outbound connectivity, and management access.
  • Re-IP affects more than the guest NIC. DNS, NAT, load balancers, firewall objects, routes, monitoring, and application configuration may all need the same translation.
  • DNS failover has cache behavior. Validate records and resolution from the networks that will actually use the recovered service.
  • Security policy and load balancer state must exist in the recovery environment and should be tested by sending real traffic through them.
  • Address overlap can break hybrid or partial failover connectivity even when both networks work independently.
  • Validation should start from the recovered workload and realistic client networks, not from a topology diagram.

Source Notes

Product behavior changes. These sources were checked against public vendor and standards documentation available on September 28, 2026.

1. Microsoft: Set up network mapping and IP addressing for VNets in Azure Site Recovery.
https://learn.microsoft.com/en-us/azure/site-recovery/azure-to-azure-network-mapping

2. Microsoft: Connect to Azure VMs after on-premises failover with Azure Site Recovery.
https://learn.microsoft.com/en-us/azure/site-recovery/concepts-on-premises-to-azure-networking

3. Microsoft: Network Security Groups with Azure Site Recovery.
https://learn.microsoft.com/en-us/azure/site-recovery/concepts-network-security-group-with-site-recovery

4. Microsoft: Customize networking configurations of the target Azure VM.
https://learn.microsoft.com/en-us/azure/site-recovery/azure-to-azure-customize-networking

5. AWS: Preparing for recovery in Elastic Disaster Recovery.
https://docs.aws.amazon.com/drs/latest/userguide/preparing-failover.html

6. AWS: Configuring the default post-launch actions in Elastic Disaster Recovery.
https://docs.aws.amazon.com/drs/latest/userguide/post-launch-action-settings-overview.html

7. AWS: About routing control in Amazon Application Recovery Controller.
https://docs.aws.amazon.com/r53recovery/latest/dg/routing-control.about.html

8. AWS: Best practices for routing control in Amazon Application Recovery Controller.
https://docs.aws.amazon.com/r53recovery/latest/dg/route53-arc-best-practices.regional.html

9. IETF RFC 1035: Domain Names, Implementation and Specification.
https://www.rfc-editor.org/info/rfc1035/

10. IETF RFC 1918: Address Allocation for Private Internets.
https://www.rfc-editor.org/info/rfc1918/

Thank you for reading this post. Kindly share it with others.

  • Share on X (Opens in new window) X
  • Share on Reddit (Opens in new window) Reddit
  • Share on LinkedIn (Opens in new window) LinkedIn
  • Share on Facebook (Opens in new window) Facebook
  • Share on Pinterest (Opens in new window) Pinterest
  • Share on Tumblr (Opens in new window) Tumblr
  • Share on Telegram (Opens in new window) Telegram
  • Share on WhatsApp (Opens in new window) WhatsApp
  • Share on Mastodon (Opens in new window) Mastodon
  • Share on Bluesky (Opens in new window) Bluesky
Network | Monitoring

Post navigation

Previous Post: Your Hypervisor Exit Plan Should Be a Recovery Runbook

Related Posts

  • SSU
    What to know about the servicing stack update and latest cumulative update in Windows Network | Monitoring
  • tape cleaning e1778595685861
    How to perform Tape Drive Cleaning in Practice Network | Monitoring
  • 980239e9 cisco logo 2
    LACP Configuration on Cisco 3650 Switch Network | Monitoring
  • screenshot 2020 03 19 at 19.17.42
    SG300 Firmware Upgrade Copy: Illegal software format Network | Monitoring
  • Add camaeras
    Add additional CC400W Cameras to Synology Surveillance Station Backup
  • elastic ip association error screen
    Fix Elastic IP Address Could not be Associated AWS/Azure/OpenShift

More Related Articles

SSU What to know about the servicing stack update and latest cumulative update in Windows Network | Monitoring
tape cleaning e1778595685861 How to perform Tape Drive Cleaning in Practice Network | Monitoring
980239e9 cisco logo 2 LACP Configuration on Cisco 3650 Switch Network | Monitoring
screenshot 2020 03 19 at 19.17.42 SG300 Firmware Upgrade Copy: Illegal software format Network | Monitoring
Add camaeras Add additional CC400W Cameras to Synology Surveillance Station Backup
elastic ip association error screen Fix Elastic IP Address Could not be Associated AWS/Azure/OpenShift

Leave a Reply Cancel reply

You must be logged in to post a comment.

Microsoft MVP

vexpert-badge-stars-5

Virtual Background

VEEAMLEGEND

Categories

veeaam100

Veeam Vanguard

  • images 4
    How to set up WatchGuard Log Server Network | Monitoring
  • trrdf
    Remote Desktop cannot find the computer this in the specified network: Verify the computer name and domain that you are trying to connect Windows Server
  • log out due to inactivity
    Automatically Log Out After a Period of Inactivity on Mac Mac
  • How to create blue screen using the Not my Fault tool from Sysinternals
    How to create blue screen using the Not my Fault tool from Sysinternals Windows
  • xxxxxx
    How to move the Taskbar to a second screen in Windows Windows
  • WSUS Post deployment Configuration Failed
    The schema version of the database is from a newer version of wsus Windows Server
  • FixThunderboltissue
    Fix the Thunderbolt application is not in use and can be safely uninstalled Windows
  • OxscsIP
    Enable Virtualization in Windows: Fixing VirtualBox’s 32-bit Option Virtualization

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 1,757 other subscribers
  • RSS - Posts
  • RSS - Comments
  • About
  • Authors
  • Write for us
  • Contact
  • Advertise with us
  • General Terms and Conditions
  • Privacy policy
  • Feedly
  • Telegram
  • Youtube
  • Facebook
  • Instagram
  • LinkedIn
  • Tumblr
  • Pinterest
  • Twitter
  • mastodon
  • Bsky

Tags

Active Directory Azure Bitlocker Microsoft Windows PowerShell WDS Windows 10 Windows 11 Windows Deployment Services Windows Server 2016

Copyright © 2026 TechDirectArchive

Loading Comments...

You must be logged in to post a comment.