"The server is down." "It's stored on our servers." "We're moving to the cloud." These words make servers sound like mysterious, special machines humming in a secret room. They're not. A server is just a computer running a program whose whole job is to wait for someone to ask it for something, and then answer. By the end of this page you'll have seen one take a request apart, you'll have the code to run your own in ten lines of Python, and you'll know exactly what "the cloud" is.

Two sides of every conversation

Almost everything you do online is a conversation between two programs:

The question is called a request and the answer is called a response. This pattern, client/server or request/response, is the backbone of the web. The client always speaks first; the server never sends you a page you didn't ask for.

Notice that "server" really names a role, played by a program. People also call the computer that program runs on "a server", which is fine, but the computer itself is nothing magic: it has a processor, memory and a disk, just like your laptop. Most server computers simply have no screen or keyboard plugged in, because nobody sits in front of them.

Here's one trip, start to finish. Step through it.

One request, step by step

browser client DNS internet server program

Two things to take from that. First, a request and a response are just text messages with an agreed format. The format the web uses is called HTTP (HyperText Transfer Protocol; a protocol is just a set of rules for a conversation). Second, the server did very little: it read a message, worked out an answer, and wrote one back. Then it went right back to waiting.

One web page is rarely one request. The HTML arrives first, then your browser notices it needs images, fonts, style sheets and scripts, and sends a separate request for each. A busy news site can easily trigger a hundred requests for a single page load.

A real server in ten lines

Here's the proof that a server is "just a program". This is a complete web server written in Python, using a module called http.server that comes built into Python, so there's nothing to install.

from http.server import BaseHTTPRequestHandler, HTTPServer

class Hello(BaseHTTPRequestHandler):
    def do_GET(self):
        message = "Hello! You asked for " + self.path + "\n"
        self.send_response(200)
        self.send_header("Content-Type", "text/plain")
        self.end_headers()
        self.wfile.write(message.encode())

print("Waiting for requests on http://localhost:8000")
HTTPServer(("localhost", 8000), Hello).serve_forever()

Here's what each part means, in plain words:

The address deserves a word too. localhost is a special name that always means "this same computer". 8000 is the port: one computer can run many server programs at once, and the port number says which one you want, like an apartment number in a building. Normal websites use port 80 (or 443 for secure HTTPS), which is why you never have to type it.

To try it for real, save the code as server.py and run python3 server.py in a terminal (on Windows: py server.py). Then open http://localhost:8000/cats in your browser. Your laptop is now a server. Stop it with Ctrl+C; Python will print a KeyboardInterrupt message, which is normal. Until you do that, you can play with this replica: type a path and see both sides of the conversation.

Try it · browser and server side by side

Your browser (the client)
localhost:8000
Nothing requested yet. Press Go.
Terminal running server.py

The output on both sides is exactly what the real program produces. Each line in the terminal is the server's log, a diary entry Python writes for every request it handles: who asked (127.0.0.1 is the numeric address for "this computer"), when, what they asked for, and the status code it answered with. Real servers keep logs like this too, and reading them is how people find out what went wrong when something breaks.

What happens when everyone asks at once

Our little server handles one request at a time. While it's working on one, any others that arrive wait in a line called a queue. That's fine when requests trickle in. But a server has a limit: if it needs 0.4 seconds per request, it can finish at most 2.5 requests per second. If requests arrive faster than that, the queue only grows. Every person in it waits longer, and once the queue is full the server starts turning people away with the error 503 Service Unavailable. That's what "the site crashed because too many people tried to buy tickets" usually means.

Play with it. Each little screen on the left is a user, and each user sends a request every 2 seconds on average. Blue dots are requests, green dots are answers, red ones are refusals.

Simulator · clients, queue, servers

requestresponse503 refuseduser waiting > 2 s
Load0%
Avg response–
Answered0
Refused (503)0

Some things to try. Slide up to 5 users: the server is now busy 100% of the time on average, and the queue swings wildly because requests don't arrive evenly; they bunch up by chance. Push to 8 or more and the queue fills, response times climb to several seconds and refusals start. Then tick the box.

The load balancer is itself a small server that does nothing but stand at the door and hand each incoming request to whichever real server has the shortest queue. To the outside world there's still one address; behind it there can be two servers, or two thousand. This is called scaling out: instead of buying one giant computer, you add more ordinary ones. It's how big sites survive.

Real servers are smarter than our toy: they handle many requests at the same time instead of strictly one by one, and a fast one can answer thousands per second. But the shape of the problem never changes. Every server has a limit, and when requests arrive faster than it can answer, people wait, and then they get errors.

Why servers live in data centres

You just saw that your laptop can be a server. So why doesn't everyone run their website from home? Because a server has to be there every time someone asks, and a home is a terrible place for that:

A data centre is a building designed to solve exactly those problems. Inside are long rows of metal cabinets called racks, each stacked with flat server computers. It has industrial cooling, several independent high-speed internet connections, batteries that take over the instant the power flickers and diesel generators that start up for longer cuts, plus guards and locked doors. Nobody sits at these machines: engineers control them over the network from anywhere, and many servers run for years without anyone touching them.

So what is "the cloud"?

Running your own data centre is enormously expensive. So a few companies built gigantic ones and started renting out the computers inside. Amazon began doing this in 2006 with Amazon Web Services (AWS); Microsoft Azure and Google Cloud followed. That's the cloud: computers in someone else's data centre, which you rent and control over the internet. There's an old joke among programmers that sums it up: "There is no cloud. It's just someone else's computer."

What makes renting so flexible is a trick called a virtual machine: software splits one powerful physical computer into several pretend computers, each of which looks and behaves like a complete computer of its own. When you rent "a server" in the cloud, you usually get one of those slices. You can have it running in about a minute, pay by the hour or even by the second, and switch it off when you don't need it. That's exactly what the toggle in the simulator does: when the traffic spike comes, rent a second server; when it's over, give it back.

You'll also hear the word serverless. Don't be fooled: there are still servers. It just means the cloud company runs and scales them for you, and you hand over only the code that handles each request, like our do_GET function.

Check yourself

You open a video in the YouTube app on your phone. Which part is the client?

The client is whoever asks. The app sends the requests; YouTube's servers answer them. The router just passes messages along.

A server handles one request at a time and needs 0.5 seconds for each. Three requests arrive at the same moment. How long until the third one is finished?

It waits in the queue while the first two are handled (0.5 + 0.5 = 1 s), then takes its own 0.5 s. Being third in line costs you time even though your own request is just as quick.

With server.py running, you visit http://localhost:8000/dogs. What does the page say?

self.path is everything after the domain and port, and it includes the leading slash. Our server never sends a 404: it answers every GET with 200.

A company says it "moved its website to the cloud". What actually changed?

The cloud is rented computers, usually virtual machines, in a provider's data centre. Still real servers, just not ones the company owns.

The short version

Next time a site says "503 Service Unavailable", you'll know what happened: too many clients, not enough servers, and a queue that ran out of room.