A problem with the primary server can make a website unavailable and interrupt sales, customer inquiries, or other important online services. Server failover reduces this risk by redirecting traffic to a prepared backup environment when a problem is detected.
Simply having a second server, however, is not enough. Reliable failover requires monitoring of the primary system, up-to-date data in the backup environment, and a mechanism for redirecting traffic.
In this article, we will look at how the process works, which components are required, and what happens after the primary server is restored.
Key Takeaway:
During a failover, the system detects a problem with the primary server and redirects traffic to a prepared backup server or to another healthy environment. To keep the website functioning normally, the backup environment needs up-to-date data, sufficient resources, and access to all critical services, including the database. Failover significantly reduces downtime but does not guarantee completely uninterrupted service.
What Is the Role of a Server in Web Hosting?
In web hosting, a server provides the resources a website needs to operate. It can store website files, run applications, and process requests from visitors when opening pages or using different features.
Depending on the hosting environment, this may be a physical server, a virtual private server (VPS), a cloud instance, or part of a larger infrastructure. More complex websites may use separate servers or services for the application, database, files, and other components.
For failover to be possible, at least one working server is required. A backup environment is needed to take over when a problem occurs, along with a system that monitors the primary server and redirects traffic when necessary.
How Does Server Failover Work?
Failover is triggered when the primary server becomes unavailable or can no longer handle requests normally. Traffic is redirected to a prepared backup environment, such as a VPS, cloud instance, or another healthy server.
Failover should not be confused with a backup. A backup stores a copy of data for later recovery, while failover is intended to keep the website running when the primary server experiences a problem. It can also form part of a broader disaster recovery plan that covers the restoration of data and services after more serious incidents.
Detecting the Problem
Health checks monitor whether the server, application, and connected services are responding normally. Failover is usually triggered after several consecutive failed checks to avoid unnecessary switching due to a brief slowdown or a temporary network issue.
Removing the Problem Server
Once the problem is confirmed, the system stops directing new requests to the affected server. In an active-active configuration, the remaining healthy servers take over the traffic, while in an active-passive configuration, the backup environment is activated.
Activating the Backup Environment
The backup environment must have up-to-date website files and configuration, sufficient resources, and access to the required services. If the database remains unavailable, for example, a working backup server alone will not be enough to keep the website functioning normally.
Redirecting Traffic
When a load balancer is used, new requests can quickly be directed to healthy servers. With DNS failover, the change may take longer because some visitors may still be using cached DNS records.
Checking the Website After Failover
After traffic has been redirected, the website and its key functions are checked to ensure they are functioning normally, including database connectivity, forms, and user sessions. The affected server remains out of service until the cause of the failure has been identified and resolved.
What Does This Look Like in Practice?
Imagine an online store running on a primary server with a prepared backup environment. At 2:00 p.m., the primary server stops responding due to a hardware or network issue. After several failed health checks, the system removes it from service and redirects new visitors to the backup server.
If the files, database, and other required services are synchronized and available, customers can continue browsing products and placing orders while the technical team resolves the problem. Some visitors may experience a brief delay or need to reload the page, but the website does not have to remain unavailable until the primary server is restored.
Will Visitors Notice the Interruption?
With a well-configured system, most visitors may not notice the problem at all. However, they may briefly see an error, experience a slower page load, or need to reload the page.
In some cases, active user sessions may also be interrupted, requiring visitors to sign in again. How noticeable the interruption is depends on how the backup configuration is designed and tested. The goal is therefore to minimize downtime rather than promise completely uninterrupted service.
Two Ways to Configure Server Failover
Active and Standby Servers (Active-Passive)
In this model, the primary server handles visitors while the second server remains on standby. If the primary server fails, the backup server is activated and takes over the traffic. The configuration is relatively easy to manage, but the backup server must be kept up to date, tested regularly, and powerful enough to handle the full traffic load.
Example: An online store runs on a primary VPS, while a second VPS maintains a synchronized backup environment. If the primary server experiences a problem, traffic is redirected to the backup server and the store continues to receive visitors and orders while the issue is resolved.
Multiple Active Servers (Active-Active)
In this model, two or more servers handle visitors simultaneously. If one of them fails, it is removed from traffic distribution while the remaining servers continue processing requests. This distributes the workload across multiple servers but requires more complex configuration and synchronization.
Example: A high-traffic website uses three active servers, with a load balancer distributing requests between them. If one stops responding, it is automatically removed from traffic distribution while the other two continue serving visitors. Once the problem has been resolved, the server can be added back to the system.
What Is Required for Reliable Failover?
A second server is only one part of the system. To keep the website running when a problem occurs, the backup environment must be ready to take over traffic, maintain up-to-date data, and have access to the required services.
Prepared Backup Infrastructure
The backup server must have sufficient resources to handle the required workload. Where possible, the primary and backup environments can be placed in different availability zones or data centers so that a single infrastructure problem does not affect both.
System Monitoring
Health checks monitor not only whether the server is reachable but also whether critical services are operating normally. This allows problems to be detected in time and the switch to the backup environment to be triggered.
At Delta.BG, we monitor the status of server and network equipment as well as the operating environment, ensuring service availability of over 99.9%.
If needed, clients can also utilize an additional system administration service for their servers and cloud infrastructure.
Traffic Redirection
A load balancer or DNS failover directs visitors to the healthy environment. With DNS, the change may take longer to reach some visitors because of cached records.
Up-to-Date Data and Services
The backup environment must have current files and configuration, as well as access to the database and other required services. Without proper synchronization, the website may be online but display outdated information or fail to process requests and transactions correctly.
How Is the Server Returned to Service?
Once the problem has been resolved, the server should not immediately begin handling traffic again. Before returning it to service, the hosting team should:
- identify and resolve the cause of the problem;
- verify that the server is operating reliably;
- synchronize files, configuration, and data with the active environment;
- verify access to required services, including the database;
- test the website and its key functions before the server handles live traffic again.
After these checks are completed successfully, the server can be returned to the active infrastructure or resume its role as the standby server. Returning traffic to the restored server is called failback and can be performed automatically or manually.
Conclusion
The most important part of failover is preparing before a real problem occurs. A backup environment that is not synchronized or has never been tested under real-world load can create additional problems precisely when it is needed most.
When planning hosting infrastructure, it is therefore important to consider not only which servers are required during normal operation, but also how the website will continue to function if the primary system fails. The more critical website availability is to the business, the more important a prepared failover plan and regular testing become.
At Delta.bg, our Cloud VPS platform provides a resilient foundation for building such a hosting architecture. Cloud VPS servers based on OpenStack run on NVMe infrastructure with Ceph triple data replication and redundant network connectivity, helping reduce infrastructure-level risks.
Delta.BG’s infrastructure is continuously monitored, including network connectivity, the status of server and network equipment, and data center conditions. We ensure 99.9% service uptime, and clients can also take advantage of additional system administration services if needed.
For help choosing a Cloud VPS configuration for your website, contact us at support@delta.bg or call +359 2 4 288 288.