Maintenance page workaround - can this be done?

Not being an experienced developer (I would say I’m on the basic level, actually), trying the nginx solution to present a maintenance page was a bit of a nightmare. Today, I had this idea and decided to ask Claude if this would be possible, now that I use Railway for other websites and applications:

Since Discourse uses Docker, it seems very complicated to show a maintenance message, because it requires nginx, which is a complicated process for me.
Can I, via Railway or any other service, create a “switch” that quickly directs traffic from one server to the other?
For example, right now Cloudflare redirects example.com to my Hetzner server where Discourse is installed.
Would I be able to send Cloudflare to a Railway project/environment/service where that “switch” would be implemented and maybe where changing a variable value from “maintenance” to “live” would direct the traffic to either another Railway service for the maintenance page, or to Hetzner?

And Claude’s reply:

Yes, technically it works, but running the switch through Railway is the wrong layer for this. You’d be adding a second network hop (Cloudflare → Railway → Hetzner) for every single request when Discourse is live, which means:

- Extra latency on every request, permanently, not just during maintenance
- A new single point of failure (if Railway service goes down, Discourse goes down even though it’s healthy)
- Egress costs on Railway for all your normal traffic, not just maintenance windows

Better approach: do the switch in a Cloudflare Worker. Workers run at the edge (300+ PoPs), so there’s no added round trip when you’re in “live” mode, and toggling is instant.

Architecture:

Cloudflare Worker (bound to example.com)
  → checks KV value "maintenance_mode"
  → if "off": fetch(origin request to Hetzner IP/hostname) and return it
  → if "on": return static maintenance HTML directly from the Worker

Toggle happens via a KV write (through API call or dashboard), not a redeploy. Propagates in seconds, no code push needed.

Minimal worker:

export default {
  async fetch(request, env) {
    const mode = await env.MAINTENANCE_KV.get("mode"); // "on" or "off"

    if (mode === "on") {
      return new Response(MAINTENANCE_HTML, {
        status: 503,
        headers: { "Content-Type": "text/html", "Retry-After": "3600" },
      });
    }

    // pass through to Hetzner untouched
    return fetch(request);
  },
};

const MAINTENANCE_HTML = `<!DOCTYPE html><html><body><h1>Back soon</h1></body></html>`;

Setup:
1. Create a KV namespace, bind it as MAINTENANCE_KV.
2. Set the Worker route to example.com/*.
3. Toggle with wrangler kv:key put --binding=MAINTENANCE_KV mode "on" or via the Cloudflare API from a shell script.

This gives you the exact “switch a variable, traffic redirects” behavior you described, without touching Discourse’s nginx/docker setup at all, and without adding a permanent extra hop for your live traffic.

One thing worth double-checking before you build this: is your Cloudflare proxy currently doing a DNS-only passthrough or already proxied (orange cloud)? If it’s DNS-only right now, Workers won’t intercept anything, you’d need to switch that record to proxied first.


So I would be interested to know if Claude’s reply is valid and can indeed be implemented? Would that be an instant switch and any user would be presented with the maintenance page? Or does it still redirect some users to the Hetzner installation and showing them the broken page, because of some cache being involved?

And replying to Claude: Yes, my example.com is proxied in Cloudflare.


I then decided to open a new Claude window and ask about this interaction, and I got this (so it seems that a bit more work needs to be done:

The architecture is sound. Cloudflare Workers on a proxied (orange-cloud) zone intercept requests at the edge before they hit Hetzner, so the design works. But the “instant, all users see it” claim needs qualification:

What’s actually true:

Worker route matching is fast, effectively immediate once deployed
KV writes propagate globally, but not instantly — Cloudflare’s docs say up to 60 seconds for KV to replicate across all edge locations. Some users could hit stale “off” state for up to a minute after toggling.
Since Discourse is proxied through Cloudflare already, the Worker will intercept 100% of requests to example.com — no bypass path exists (assuming no other DNS records like a subdomain pointing straight at Hetzner IP, or a CDN cache serving stale content in front of it).

What the reply glossed over / got wrong:

Browser cache: if any prior response (from Hetzner) was cached with a long TTL and Cache-Control headers, a returning visitor’s browser could serve a cached page without even hitting Cloudflare. Minor risk, but real — worth setting Cache-Control: no-store on the maintenance response.
Cloudflare cache: if you have caching rules/page rules caching HTML for example.com, a cached copy could be served without ever reaching the Worker logic, depending on rule order. Workers execute before cache lookup by default in most configs, but if there’s an existing Cache Rule set to “Cache Everything” it can interact in ways worth testing, not assuming.
503 status code: using 503 is correct practice (search engines won’t deindex during brief maintenance), but he should add Retry-After properly formatted — the example does this correctly.
The claim “without touching Discourse’s nginx/docker setup at all” is accurate.
No mention of KV read latency inside the Worker itself — trivial, sub-ms, not a real concern.

Bottom line: workable, and better than the nginx approach. “Instant” is optimistic — call it “propagates within ~60 seconds worldwide” and tell them to test with cache disabled first. The reply should have flagged KV propagation delay instead of implying it’s instant.

Hey!

I’m not answering your question, just being curious. Why do you think you need a maintenance page? Seems quite annoying to set up for little benefit.

My instances are down 5 minutes a month for updates and that’s basically all.

Whenever I made significant changes, like installing plugins for example, it would take sometimes 20 minutes.

Not a big deal, if we consider that these are not done every week or even every month, but for a new visitor that ugly page that something is not working, is not a good first impression. Even for returning visitors who are not aware that the website will be down, that can throw them into “panic mode” thinking that the whole community has been shut down or something, maybe some of them promptly emailing me asking what’s happening.

I think it’s just a nice touch to give the visitor, new or old, to let them know what’s happening.

If this approach Claude suggested is something that takes just a few minutes, but it’s not touched every again, I believe it’s time well spent.

It takes seconds if you move to the two container install.

You can even schedule the switch over at night if you really want to, whilst you and/or your main audience sleep.

You don’t need a maintenance page.

i use the dual container build behind Cloudlfare CDN. the dual container minimizes downtime and my two main production forums both have 30 seconds max offline during rebuilds. i like having a maintenance page because the default web server is down page is ugly and gives no clue to how long it will be offline.


documentation to do this is now here:

Wow! I really appreciate that detailed walkthrough! :raising_hands:
I just saved your reply in my notes so when I install Discourse again soon, I will come back to it. I will definitely let you know how it went, even if it takes a while until I install it. I’m finishing some coding stuff at the moment and only then will I focus on Discourse again.

And I agree that the “web serve is down” page is ugly, and for inexperienced users, that message doesn’t mean much. Having a simple maintenance page with a custom message, is always preferable, even if it involves some work setting up.

Again, thank you so much for taking the time to share this. I hope other people will find this useful :slight_smile:

and there’s the rub.

It’s not a simple change and additional infrastructure to worry about and maintain, and imho entirely not worth it for a few seconds of downtime.

Also I believe that people with the Discourse site open will not be redirected to the maintenance page, things will just error? They will only be redirected if they hit refresh, by which time the server may be back up in any case …

… so this will benefit new browsers, but only those who navigate within the short switchover

I understand what you mean, but that’s just one case. What if something really breaks and I need to fix it, taking longer than just a few seconds?

Also, how are you getting a few seconds of downtime and I was seeing like 20 minutes, after installing plugins?

Because in the two container setup (linked in both of our posts and a non-negotiable dependency) you bootstrap the new build using ./launcher bootstrap web_only (during which your site is still 100% functional), then you simply destroy and immediately start the newly, already built container (which takes seconds)

Oh ok, that’s still for the 2 container setup. I thought that you were saying it was not worth the work building the 2 container setup completely. What you are saying is not worth it, is just that maintenance page extra work, right?

Personally yes.

If you want belts and braces for web viewers that appear with fresh navigation during the 30-60 seconds, then a maintenance page is fair.

You have to weigh the benefit of that against the time to maintain the infrastructure around it.

Now I understand. Thanks.

So now I ask: with 2 containers, those 20 minutes or so, are still a thing, the only difference is that I can let it do it on one of the containers, the one that’s not live, and when it’s ready, do the switch. Is that how it works?

Yes, it still takes a while for the container to build and FYI you need to make sure your server is beefy enough to allow that process to take place in parallel whilst serving your community as normal (essentially the most important thing is ensuring enough memory for starters - so very important to make sure you have enough SWAP). The build will also knock out at least one core so think about having a server with 1 or 2 more cores.

Now it all makes sense. Thank you.

I was initially using this, but it seems like it’s not available anymore:

So now I need to go to this one:

Which is a bit jump in price… :confused: and what seems to be a less effective server…?

I would recommend at least 4GB and 3 cores (“vcpu”) for this kind of setup …

For reference as it was asked on another topic, the (edit: wrong) answer: Any cheaper alternatives to Hetzner? - #2 by Canapin

This is the only one that matches that… what a difference in price…

image

I guess I will have to build it with one container now and when the time (:money_bag:) comes, upgrade.

Thanks for the info, Robert!

i disagree and do think it’s worth the maintenance page. in my experience, when people see the Discoure error page the first thing they do is hit refresh and then they will see the brief maintenance page that will then automatically reload the site when the container swap is complete. it isn’t a lot of work to set up and you only have to do it once. during a rebuild with the dual container configuration, the bootstrap part happens in the background and users can still use the site; it is the container switch over that causes the brief downtime (unlike a standard single container configuration)

server cost and selection is a completely different discussion and should be in another separate topic than setting up a maintenance page.

edit: i see there is a topic here and i replied there.

I guess I’m in the middle right now. I understand what you mean, as well as @merefield. It’s a nice touch to have it and if it’s only set up once, not a big deal, I guess?

At the same time, if it does indeed take up to 60 seconds, worst case scenario, not everyone will visit the website exactly at the 0 second mark and have to wait 60 seconds. Some people will see the ugly page for 5 seconds, some 20, some 60. Some people will not even see it (I believe?) for example if they are just reading a reply or writing one, that switch could happen before they hit SEND or before they stop reading the topic and hit REPLY (or visit another page).

My issue with the nginx route, was more related to the 20 minute thing on a one container setup. With the option to have 2 and going from 20 minutes to maybe 60 seconds tops, I wonder if it’s really relevant? Especially, if I am able to add an announcement at the top saying where things will go down just for a few seconds, on a low-traffic time, this won’t be a big issue?

I have to sit and think about it. The 2 container setup will definitely be something to implement, if I find a good server deal, though.

Can you clarify this, in a way that a basic person like me can understand? I’m still not familiar with workers…

Why do you add and delete workers route page, if leaving them is not an issue? So, if I do add the maintenance page, can I really just set it up once, and from that moment on, I always just SSH my server via Terminal and do everything there like I do with a single container, without ever having to go to Cloudflare?

@merefield mentioned that this (maintenance page, workers, etc) is also something that needs to be maintained, so I wonder if it’s really something that you set and forget, or if there’s anything that needs to be done every once in a while?