드디어 문제를 해결했습니다. 정말로 고쳐졌는지 확인하기 위해 꽤 오랜 시간을 기다렸고, 이제 정상적으로 작동하는 것 같습니다.
구성 (Setup)
Discourse는 Docker에서 실행되며 nginx를 Unix 소켓(/var/discourse/shared/standalone/nginx.http.sock)으로 노출합니다. Caddy는 그 앞에 리버스 프록시로 위치하며 X-Real-IP 헤더를 통해 클라이언트의 실제 IP를 전달합니다.
변경해야 할 두 가지가 있었습니다: 컨테이너 내부의 nginx 설정과 재빌드(rebuild) 실행 방식입니다.
1. app.yml
## Plugins go here
## see https://meta.discourse.org/t/19157 for details
hooks:
# This after_code block is just my own plugin list — it has nothing to do with
# the fix. Keep whatever you already have here.
after_code:
- exec:
[...]
# This is the part that matters.
# 1. Writes an http-level map that turns the literal "unix:" into 127.0.0.1.
# The 00- prefix makes nginx load it before discourse.conf.
# 2. Rewrites discourse.conf so X-Forwarded-For uses the mapped variable
# instead of the raw $remote_addr.
# 3. nginx -t fails the build if the result is not valid.
after_web_config:
- exec: >-
printf 'map $remote_addr $safe_remote_addr {\n "unix:" 127.0.0.1;\n default $remote_addr;\n}\nreal_ip_header X-Real-IP;\n'
> /etc/nginx/conf.d/00-safe-remote-addr.conf
- exec: >-
sed -i 's/X-Forwarded-For \$remote_addr;/X-Forwarded-For $safe_remote_addr;/g'
/etc/nginx/conf.d/discourse.conf
- exec: nginx -t
핵심 코드는 after_web_config 블록 내부에 있는 부분입니다.
2. Caddyfile 블록
아래 스크립트는 플래그 파일을 통해 유지보수 모드를 제어하므로, 사이트 블록에서 이를 인식해야 합니다:
forum.example.com {
request_header X-Real-IP {remote_host}
root * /var/caddy/flags
@maintenance file maintenance.flag
handle @maintenance {
respond "Maintenance in progress. We'll be back in a few minutes." 503
}
handle {
reverse_proxy unix//var/discourse/shared/standalone/nginx.http.sock
}
}
스크립트에서 사용하는 플래그 경로와 @maintenance 매처가 찾는 파일은 반드시 동일한 파일이어야 합니다. 하나를 변경하면 다른 하나도 함께 변경해야 하며, 그렇지 않으면 유지보수 모드가 조용히 작동하지 않습니다.
같은 Caddy 인스턴스가 다른 애플리케이션도 서빙하는 경우, 포럼 블록만 매칭되는 플래그를 사용하세요. 그렇지 않으면 Discourse 재빌드 시 다른 애플리케이션들도 함께 다운됩니다.
3. 재빌드 스크립트
이제 더 이상 ./launcher rebuild app을 직접 사용하지 않습니다. 대신 discourse-rebuild.sh를 사용합니다:
#!/bin/bash
#
# discourse-rebuild.sh — Rebuilds the Discourse container without leaving orphaned
# requests behind and without the `invalid input syntax for type inet: "unix:"` error.
#
# CONTEXT
# Discourse runs in Docker and exposes nginx on a Unix socket
# (/var/discourse/shared/standalone/nginx.http.sock). Caddy acts as a reverse
# proxy in front of it and passes the client's real IP through X-Real-IP.
#
# A `rebuild` destroys the container and recreates the socket with a new inode.
# During that transition there are two problematic windows:
#
# 1) Caddy keeps state from the previous socket until it is reloaded.
# 2) nginx starts accepting connections as soon as it boots, but Unicorn
# takes ~15s longer before it can serve them.
#
# A request landing in either window may arrive with no X-Real-IP. $remote_addr
# is then left holding the literal "unix:", which PostgreSQL rejects when
# inserting it into an inet column -> HTTP 500.
#
# WHAT IT DOES
# 1. Raises a flag file that puts the site into 503 (Caddy checks it on every
# request, so it takes effect instantly and with no reload).
# 2. Waits until nothing but nginx itself is holding the socket open.
# 3. Rebuilds the container.
# 4. Polls /srv/status against the socket until Unicorn answers 200.
# 5. Reloads Caddy so it picks up the new socket, and clears the flag.
# 6. Starts the watcher that logs any leftover "unix:" hit.
#
# If the rebuild fails, or Discourse never answers within the polling window,
# the flag is NOT removed: the site stays in maintenance on purpose, so a broken
# container is never exposed. Bring it back up by hand with:
# rm -f /var/caddy/flags/maintenance.flag
#
# REQUIREMENTS
# - lsof installed, and root privileges.
# - FLAG below must point at the exact same file the Caddyfile @maintenance
# matcher looks for.
# - request_header X-Real-IP {remote_host} in that same Caddyfile block.
#
# USAGE
# ./discourse-rebuild.sh
#
# TO SEE WHAT THE WATCHER CAUGHT
# cat /var/log/unixip-hits.log
#
set -e
FLAG=/var/caddy/flags/maintenance.flag
SOCK=/var/discourse/shared/standalone/nginx.http.sock
cleanup() {
local code=$?
systemctl reload caddy
if [ $code -eq 0 ]; then
rm -f "$FLAG"
echo "✅ Rebuild finished. Site is online."
else
echo "⚠️ Rebuild failed (exit code $code). The site is still in maintenance."
echo " Check it, and once it's ready: rm -f $FLAG"
fi
}
trap cleanup EXIT
mkdir -p "$(dirname "$FLAG")"
touch "$FLAG"
echo "🔧 Maintenance is on. Waiting for in-flight requests..."
# Rough heuristic: count the processes holding the socket open and wait until
# only the listener is left. Up to 30s.
for i in $(seq 30); do
n=$(lsof -t "$SOCK" 2>/dev/null | wc -l || echo 0)
[ "$n" -le 1 ] && break
sleep 1
done
/var/discourse/launcher rebuild app
# Up to 60 attempts: about 2 minutes of sleeps, more if any curl hits its own
# 5s timeout.
echo "⏳ Waiting for Discourse to answer..."
status=000
for i in $(seq 60); do
status=$(curl -s -o /dev/null -w '%{http_code}' --max-time 5 \
--unix-socket "$SOCK" http://localhost/srv/status 2>/dev/null || echo 000)
[ "$status" = "200" ] && break
sleep 2
done
if [ "$status" != "200" ]; then
echo "⚠️ Discourse never answered within the polling window (last status code: $status)."
exit 1
fi
echo "✅ Discourse ready after ~$((i*2))s."
systemd-run --unit=unixip-watch --collect \
/bin/bash -c "docker exec app tail -F /var/log/nginx/access.log | grep --line-buffered 'unix:' >> /var/log/unixip-hits.log"
스크립트가 실제로 하는 일
- Caddy의 플래그 파일을 통해 사이트를 유지보수 모드로 전환합니다.
- 소켓이 조용해질 때까지 기다린 후
./launcher rebuild app을 실행합니다. - Unicorn이
200을 응답할 때까지 소켓을 통해/srv/status를 폴링합니다. 이것이 핵심적인 부분입니다: Unicorn이 부팅 중일 때 요청이 nginx에 도달하는 것을 막아줍니다. - Caddy를 재로드합니다(안전장치로 추가했으나, 실제로는 중요하지 않은 것으로 나타남).
- 해당 폴링이 성공한 경우에만 유지보수 플래그를 제거합니다.
수 주가 지나고 여러 번의 재빌드를 거쳤지만, 그 에러를 다시는 보지 못했습니다.
P.S.: 제가 대체 뭘 한 건지 정말로 모르겠습니다.