SiteUp: The North Pole Has Never Had an Outage and the Elves Would Like You to Know That
Dear Santa,
My name is Oliver and I am twelve years old. I live in Bristol, England, where my dad works in IT and spends a lot of time on the phone saying things like “have you tried turning it off and on again” and “I’m looking at the ticket now” and “whose idea was it to update on a Friday.” He has been doing this for as long as I can remember. He seems tired.

Dear Oliver, your father is performing one of the most important and least celebrated jobs in modern civilisation. The people who keep systems running are never noticed until something stops running, at which point they are noticed by absolutely everyone all at once, loudly, and usually via a channel that is also not working properly. Please tell him the North Pole sees him. We have our own IT situation and it has given me enormous respect for people in his position.
Santa, my question is this. Does your website ever crash? Because I tried to visit santaclaus.top last Tuesday and it loaded fine, but my friend George said he heard that Santa’s systems go down sometimes in December when too many children write in at once, and I told him that seemed very unlikely because surely the North Pole would have planned for that, and he said maybe Santa’s hosting provider is bad, and I said I didn’t think Santa used a hosting provider in the traditional sense, and then we argued about it for twenty minutes and nobody won and I decided to write to you directly to settle it.
Oliver. Sit down. I am going to tell you about North Pole infrastructure, and it is going to be the most reassuring thing you have read this year.

The North Pole does not go down. Not on December 1st when the first wave of letters arrives. Not on December 20th when the volume peaks. Not on Christmas Eve when the entire operation is running simultaneously across every time zone on Earth with a nine-reindeer delivery vehicle, several thousand active elves, one very focused Mrs. Claus with a clipboard, and a logistics load that would cause any conventional system to submit a formal incident report and go home early. The North Pole stays up. It has always stayed up. This is not luck. It is architecture.
How does the North Pole handle so much traffic? My dad says our home router struggles when more than three people are streaming at the same time and one of them is my sister watching the same episode of the same programme for the seventh time.
The North Pole does not rely on a single router. The infrastructure is distributed across a redundant network that the elves spent several decades building and which has been stress-tested under conditions that no conventional load-testing environment could replicate, because no conventional load-testing environment has ever had to simulate the simultaneous submission of wish lists from two hundred countries in seventeen languages including one regional dialect spoken exclusively in a village in northern Finland that has not appeared on any map since 1932. Your sister streaming the same episode repeatedly is, from an infrastructure perspective, a known and manageable load type. The North Pole has protocols for this.
Has anything ever gone wrong? Technically, I mean. My dad says every system has a failure mode. He says this a lot. Usually at dinner.
Your father is correct that every system has a failure mode, and I will be honest with you because you asked directly and you deserve a direct answer. There was an incident in 2019. I will tell you about it.
In November 2019, an elf named Gerald — who is in the workshop logistics division and has, since this incident, been strongly encouraged to stay in the workshop logistics division — decided to update the correspondence management system eleven days before December 1st. This was done without a change approval. It was done on a Thursday, which is already inadvisable, and Gerald had described the update as “low risk,” which is the two most dangerous words in the English language when spoken by someone holding administrator credentials.
The system went down for four hours and seventeen minutes. During this time, approximately forty thousand letters were queued and unprocessed, the elf correspondence team operated on paper backups that had not been tested since 2003 and turned out to have several gaps, and Prancer filed a support ticket from the reindeer barn on a tablet that nobody knew he had. The ticket read: “Is this related to why my flight schedule hasn’t updated. Asking for a colleague.” It was not related. Prancer filed it anyway. We found it three weeks later.

The system was restored. All letters were processed. No child was affected. Gerald completed a mandatory change management course and has not touched a production system since. The incident is documented internally as Case 2019-NP-001 and is referred to by the IT elves as “the Gerald Situation,” always with a specific tone of voice that Gerald finds unnecessary and everyone else considers proportionate.
What does North Pole IT actually look like? Do the elves do it?
The North Pole IT team consists of eleven elves, one senior elf architect who joined from a previous role she declines to specify but which involved a very large online retailer that delivers globally, and a reindeer named Comet who cannot write code but who sits near the server room and whose presence the team says is “calming,” which Comet accepts without comment because Comet is fundamentally a professional.
The infrastructure runs on a hybrid model combining on-premises North Pole servers — which are cooled by ambient Arctic air, which is one operational advantage of being located at the top of the planet — and a distributed cloud architecture that the elves built themselves because no existing cloud provider had a data centre north of Tromsø and the latency was unacceptable. The system handles inbound letter processing, wish list analysis, naughty-and-nice list management, route optimisation, workshop inventory, reindeer scheduling, and the overnight Christmas Eve delivery coordination that runs for approximately thirty-one hours across all time zones without a single planned maintenance window, because there is no appropriate maintenance window on Christmas Eve and there never will be.
My dad says you should always have a disaster recovery plan. Does the North Pole have one?

The North Pole disaster recovery plan is one hundred and forty-three pages long, was last updated in March, and is reviewed annually by Mrs. Claus, who reads it cover to cover, annotates it in red pen, and returns it with questions that the IT elves describe as “extremely specific” and “occasionally uncomfortable.” The plan covers server failure, network partition, power loss, extreme weather events — which at the North Pole is essentially “weather,” because extreme is the baseline — elf staffing shortages, reindeer medical incidents, and a section titled “Unforeseen Circumstances” that was added after the Gerald Situation and is now twelve pages long.
There is also a laminated one-page summary kept near the main server room door. It says: “Stay calm. Call Mrs. Claus. Do not update anything.” This covers approximately eighty percent of scenarios.
Has Prancer filed any other tickets?
Prancer has filed eleven tickets since 2019. Three were legitimate. Four were about things that were not IT issues but which Prancer felt should be escalated through a formal channel. Two were filed during training flights and described conditions that turned out to be user error. One was filed about the temperature in the reindeer barn, which is definitively a facilities issue and not an IT issue, a distinction that Prancer refuses to accept. The eleventh ticket was filed last February and consisted entirely of the words “you know what you did” with no further context. It has been assigned to Gerald for investigation. Gerald is making very slow progress.
My final question. Can I ask Santa something on behalf of my dad? He wants to know what uptime percentage the North Pole runs at.

Tell your father: excluding the Gerald Situation, which lasted four hours and seventeen minutes across a twelve-year measurement window, the North Pole infrastructure runs at 99.9996 percent uptime. This figure has been reviewed by the elf metrics team and is accurate. Your father will know that this is a number most enterprise systems do not achieve. He is welcome to ask how we do it. The answer is: redundancy, documentation, a disaster recovery plan that Mrs. Claus reads with a red pen, and a strict policy that nobody touches production in November or December under any circumstances whatsoever, not even Gerald, not even if Gerald says it is low risk, especially if Gerald says it is low risk.
Merry Christmas, Oliver. Tell your dad his instincts are good, his dinner-table infrastructure philosophy is correct, and that the North Pole considers people who keep systems running to be among the most important people in any organisation. We know this from experience. We found out in November 2019. Gerald taught us.
Thank you, Santa. I am going to show this letter to my dad. He is going to have a lot of follow-up questions.
Your friend,
Oliver James Whitmore
Bristol, England
P.S. My dad wants to know what monitoring stack the North Pole uses.
The North Pole monitoring stack is proprietary, built in-house by the elf infrastructure team, and runs on servers cooled by actual Arctic air which gives us a power usage effectiveness rating that the elves are extremely smug about at conferences. Your dad is welcome to write in directly. We will send him a postcard from the North Pole and a very brief answer that will raise more questions than it resolves, because that is the nature of monitoring infrastructure and also of Christmas magic, and at a certain point the two are not entirely different things.

by