One address, many computers

When you open a popular website, you type one address. But one computer could never answer millions of people at once. Behind that single address there are usually many servers (a server is a computer that waits for requests and answers them), and something has to decide which of them handles your request. That something is a load balancer.

Think of a supermarket with three checkouts and one person at the door pointing each customer to a till. The person at the door never scans anything. They only choose where you go. That is the whole job.

Try it: you are the load balancer

Pick a strategy, press Send request, and watch where each request lands. Each request keeps a server busy for a few seconds (one square = one request being handled). Then switch a server off, the way a crash would, and see what changes.

Nothing sent yet.

Each request occupies a server for 3 to 6 seconds. A server that is switched off gets no new requests (that is what a health check achieves).

Try these three experiments. First, keep the strategy on round robin and press Send 6 at once: every server gets two. Second, switch server-B off and send again: the requests are shared by the two that are left, and nobody notices a thing. Third, pick Always the first server: one server drowns while two sit idle, which is exactly the situation a load balancer exists to prevent.

Three common ways to choose

Round robin takes turns in a fixed order: A, B, C, A, B, C. It is simple and fair when every request costs about the same.

Least connections looks at how many requests each server is handling right now and picks the quietest. It copes better when some requests are slow and others are quick, because a server stuck on a slow request stops being chosen. Try it with Auto switched on and compare the squares to round robin.

There are more (sending the same visitor to the same server, weighting bigger machines higher), but these two cover the idea.

The same idea in code

Here is round robin with a health check in a few lines of Python. A health check is a regular question the load balancer asks each server, roughly "are you alive?". A server that fails is skipped until it recovers.

from itertools import cycle

servers = ["server-A", "server-B", "server-C"]
healthy = {"server-A": True, "server-B": True, "server-C": True}
ring = cycle(servers)

def pick():
    for _ in range(len(servers)):
        s = next(ring)
        if healthy[s]:
            return s
    return None   # nobody is up

for n in range(1, 5):
    print("request", n, "->", pick())

healthy["server-B"] = False
print("server-B goes down")
for n in range(5, 9):
    print("request", n, "->", pick())

Output when we ran it:

request 1 -> server-A
request 2 -> server-B
request 3 -> server-C
request 4 -> server-A
server-B goes down
request 5 -> server-C
request 6 -> server-A
request 7 -> server-C
request 8 -> server-A

Once server-B is marked unhealthy, pick() steps past it. If every server were down, pick() would return None, and a real load balancer would answer with an error such as HTTP 503 ("Service Unavailable").

Where the load balancer sits

When you type a web address, DNS (the internet's address book) turns the name into an IP address. For a busy site, that IP address belongs to the load balancer, not to any one of the application servers. Your request arrives there first, the balancer picks a server, passes the request along, and then hands the server's answer back to you. From your side it looks like you talked to a single computer.

That position also explains why the load balancer itself must not become the weak spot. Real setups often run more than one, or use a managed service that does this for you, so that the thing guarding against failure is not a single point of failure itself.

What happens when a server dies mid-request

Notice what happened in the demo when you switched a server off: the squares on it vanished. Those were requests that were being handled when the crash happened, and they are lost. A load balancer cannot rescue work a dead server was halfway through; it can only stop sending new requests there. Good applications cope by letting the visitor's browser or app retry, and by making sure that repeating a request is safe. For example, charging a card twice because of a retry would be a bug, so payment systems are designed to recognise a repeated request.

A health check also takes time. If the balancer checks every 10 seconds, a crashed server may still receive requests for up to 10 seconds before it is marked down. Shorter intervals notice faster but create more checking traffic. It is a trade-off, and real systems tune it.

What this means for your own code

Because any request may land on any server, a server should not keep important things only in its own memory. If you log in on server A and your next click lands on server B, B has never heard of you. This is why web apps keep sessions in a shared place such as a database, or pass a signed token with each request. Programs built this way are called stateless, and they are much easier to run on many servers.

Load balancers also give you a second benefit: you can update servers one at a time while the others keep answering. Tools you may meet later include NGINX, HAProxy, and the managed load balancers offered by cloud providers. All of them do the same basic job you did by hand above.