재구축(2025년 2월 4일) 후 사이트 오프라인

최근 재빌드 후 이 메시지를 확인했습니다. 이후 ./launcher rebuild app을 실행했는데, 그 후 인스턴스에 접근할 수 없게 되었습니다. 표준 설치 환경인데, 무슨 일이 발생했는지 어떻게 파악할 수 있을까요?

./launcher logs app 실행 시 오류 발생

cd /var/discourse
./launcher logs app
x86_64 arch detected.
run-parts: executing /etc/runit/1.d/00-ensure-links
run-parts: executing /etc/runit/1.d/00-fix-var-logs
run-parts: executing /etc/runit/1.d/01-cleanup-web-pids
run-parts: executing /etc/runit/1.d/anacron
run-parts: executing /etc/runit/1.d/cleanup-pids
Cleaning stale PID files
run-parts: executing /etc/runit/1.d/copy-env
run-parts: executing /etc/runit/1.d/letsencrypt
[Tue Feb  4 05:38:16 PM UTC 2025] Domains not changed.
[Tue Feb  4 05:38:16 PM UTC 2025] Skip, Next renewal time is: 2025-03-02T20:15:28Z
[Tue Feb  4 05:38:16 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 05:38:17 PM UTC 2025] Installing key to: /shared/ssl/mydomain.com.key
[Tue Feb  4 05:38:17 PM UTC 2025] Installing full chain to: /shared/ssl/mydomain.com.cer
[Tue Feb  4 05:38:17 PM UTC 2025] Run reload cmd: sv reload nginx
warning: nginx: unable to open supervise/ok: file does not exist
[Tue Feb  4 05:38:17 PM UTC 2025] Reload error for :
[Tue Feb  4 05:38:17 PM UTC 2025] Domains not changed.
[Tue Feb  4 05:38:17 PM UTC 2025] Skip, Next renewal time is: 2025-03-02T20:15:33Z
[Tue Feb  4 05:38:17 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 05:38:18 PM UTC 2025] Installing key to: /shared/ssl/mydomain.com_ecc.key
[Tue Feb  4 05:38:18 PM UTC 2025] Installing full chain to: /shared/ssl/mydomain.com_ecc.cer
[Tue Feb  4 05:38:18 PM UTC 2025] Run reload cmd: sv reload nginx
warning: nginx: unable to open supervise/ok: file does not exist
[Tue Feb  4 05:38:18 PM UTC 2025] Reload error for :
Started runsvdir, PID is 567
ok: run: redis: (pid 577) 0s
ok: run: postgres: (pid 581) 0s
nginx: [warn] duplicate extension "wasm", content type: "application/wasm", previous content type: "application/wasm" in /etc/nginx/conf.d/discourse.conf:4
supervisor pid: 575 unicorn pid: 607
Shutting Down
run-parts: executing /etc/runit/3.d/01-nginx
ok: down: nginx: 1s, normally up
run-parts: executing /etc/runit/3.d/02-unicorn
(575) exiting
ok: down: unicorn: 0s, normally up
run-parts: executing /etc/runit/3.d/10-redis
ok: down: redis: 1s, normally up
run-parts: executing /etc/runit/3.d/99-postgres
ok: down: postgres: 0s, normally up
ok: down: nginx: 5s, normally up
ok: down: postgres: 1s, normally up
ok: down: redis: 3s, normally up
ok: down: cron: 0s, normally up
ok: down: unicorn: 4s, normally up
ok: down: rsyslog: 0s, normally up
run-parts: executing /etc/runit/1.d/00-ensure-links
run-parts: executing /etc/runit/1.d/00-fix-var-logs
run-parts: executing /etc/runit/1.d/01-cleanup-web-pids
run-parts: executing /etc/runit/1.d/anacron
run-parts: executing /etc/runit/1.d/cleanup-pids
Cleaning stale PID files
run-parts: executing /etc/runit/1.d/copy-env
run-parts: executing /etc/runit/1.d/letsencrypt
[Tue Feb  4 05:58:32 PM UTC 2025] Domains not changed.
[Tue Feb  4 05:58:32 PM UTC 2025] Skip, Next renewal time is: 2025-03-02T20:15:28Z
[Tue Feb  4 05:58:32 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 05:58:32 PM UTC 2025] Installing key to: /shared/ssl/mydomain.com.key
[Tue Feb  4 05:58:32 PM UTC 2025] Installing full chain to: /shared/ssl/mydomain.com.cer
[Tue Feb  4 05:58:32 PM UTC 2025] Run reload cmd: sv reload nginx
fail: nginx: runsv not running
[Tue Feb  4 05:58:32 PM UTC 2025] Reload error for :
[Tue Feb  4 05:58:32 PM UTC 2025] Domains not changed.
[Tue Feb  4 05:58:32 PM UTC 2025] Skip, Next renewal time is: 2025-03-02T20:15:33Z
[Tue Feb  4 05:58:32 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 05:58:32 PM UTC 2025] Installing key to: /shared/ssl/mydomain.com_ecc.key
[Tue Feb  4 05:58:32 PM UTC 2025] Installing full chain to: /shared/ssl/mydomain.com_ecc.cer
[Tue Feb  4 05:58:32 PM UTC 2025] Run reload cmd: sv reload nginx
fail: nginx: runsv not running
[Tue Feb  4 05:58:32 PM UTC 2025] Reload error for :
Started runsvdir, PID is 561
ok: run: redis: (pid 575) 0s
nginx: [warn] duplicate extension "wasm", content type: "application/wasm", previous content type: "application/wasm" in /etc/nginx/conf.d/discourse.conf:4
ok: run: postgres: (pid 580) 1s
supervisor pid: 570 unicorn pid: 601
Shutting Down
run-parts: executing /etc/runit/3.d/01-nginx
ok: down: nginx: 0s, normally up
run-parts: executing /etc/runit/3.d/02-unicorn
(570) exiting
ok: down: unicorn: 1s, normally up
run-parts: executing /etc/runit/3.d/10-redis
ok: down: redis: 0s, normally up
run-parts: executing /etc/runit/3.d/99-postgres
ok: down: postgres: 0s, normally up
ok: down: nginx: 3s, normally up
ok: down: postgres: 1s, normally up
ok: down: redis: 1s, normally up
ok: down: cron: 0s, normally up
ok: down: unicorn: 3s, normally up
ok: down: rsyslog: 0s, normally up
run-parts: executing /etc/runit/1.d/00-ensure-links
run-parts: executing /etc/runit/1.d/00-fix-var-logs
run-parts: executing /etc/runit/1.d/01-cleanup-web-pids
run-parts: executing /etc/runit/1.d/anacron
run-parts: executing /etc/runit/1.d/cleanup-pids
Cleaning stale PID files
run-parts: executing /etc/runit/1.d/copy-env
run-parts: executing /etc/runit/1.d/letsencrypt
[Tue Feb  4 06:01:07 PM UTC 2025] Domains not changed.
[Tue Feb  4 06:01:07 PM UTC 2025] Skip, Next renewal time is: 2025-03-02T20:15:28Z
[Tue Feb  4 06:01:07 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 06:01:07 PM UTC 2025] Installing key to: /shared/ssl/mydomain.com.key
[Tue Feb  4 06:01:07 PM UTC 2025] Installing full chain to: /shared/ssl/mydomain.com.cer
[Tue Feb  4 06:01:07 PM UTC 2025] Run reload cmd: sv reload nginx
fail: nginx: runsv not running
[Tue Feb  4 06:01:07 PM UTC 2025] Reload error for :
[Tue Feb  4 06:01:07 PM UTC 2025] Domains not changed.
[Tue Feb  4 06:01:07 PM UTC 2025] Skip, Next renewal time is: 2025-03-02T20:15:33Z
[Tue Feb  4 06:01:07 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 06:01:07 PM UTC 2025] Installing key to: /shared/ssl/mydomain.com_ecc.key
[Tue Feb  4 06:01:07 PM UTC 2025] Installing full chain to: /shared/ssl/mydomain.com_ecc.cer
[Tue Feb  4 06:01:07 PM UTC 2025] Run reload cmd: sv reload nginx
fail: nginx: runsv not running
[Tue Feb  4 06:01:07 PM UTC 2025] Reload error for :
Started runsvdir, PID is 561
ok: run: redis: (pid 575) 0s
ok: run: postgres: (pid 576) 0s
nginx: [warn] duplicate extension "wasm", content type: "application/wasm", previous content type: "application/wasm" in /etc/nginx/conf.d/discourse.conf:4
supervisor pid: 570 unicorn pid: 601
(570) exiting
nginx: [warn] duplicate extension "wasm", content type: "application/wasm", previous content type: "application/wasm" in /etc/nginx/conf.d/discourse.conf:4

모든 것이 순조롭게 진행되고 있었습니다. 아래 내용을 확인하고 다시 빌드했습니다. 빌드는 오류 없이 완료되었지만, 제 사이트가 열리지 않습니다.

-------------------------------------------------------------------------------------
UPGRADE OF POSTGRES COMPLETE

Old 13 database is stored at /shared/postgres_data_old

To complete the upgrade, rebuild again using:

./launcher rebuild app
-------------------------------------------------------------------------------------

이 명령을 실행하면
tail /var/discourse/shared/standalone/log/var-log/postgres/current
출력 결과는 다음과 같습니다.

2025-02-04 18:11:50.943 UTC [573] LOG:  shutting down
2025-02-04 18:11:50.945 UTC [573] LOG:  checkpoint starting: shutdown immediate
2025-02-04 18:11:50.970 UTC [573] LOG:  checkpoint complete: wrote 139 buffers (0.0%); 0 WAL file(s) added, 0 removed, 0 recycled; write=0.017 s, sync=0.005 s, total=0.027 s; sync files=27, longest=0.002 s, average=0.001 s; distance=410 kB, estimate=410 kB
2025-02-04 18:11:51.034 UTC [547] LOG:  database system is shut down
2025-02-04 18:15:04.302 UTC [548] LOG:  starting PostgreSQL 15.10 (Debian 15.10-1.pgdg120+1) on x86_64-pc-linux-gnu, compiled by gcc (Debian 12.2.0-14) 12.2.0, 64-bit
2025-02-04 18:15:04.303 UTC [548] LOG:  listening on IPv4 address "0.0.0.0", port 5432
2025-02-04 18:15:04.303 UTC [548] LOG:  listening on IPv6 address "::", port 5432
2025-02-04 18:15:04.305 UTC [548] LOG:  listening on Unix socket "/var/run/postgresql/.s.PGSQL.5432"
2025-02-04 18:15:04.313 UTC [575] LOG:  database system was shut down at 2025-02-04 18:14:37 UTC
2025-02-04 18:15:04.318 UTC [548] LOG:  database system is ready to accept connections

또한 ./launcher logs app을 실행하면 다음과 같은 출력이 표시됩니다.

x86_64 arch detected.
run-parts: executing /etc/runit/1.d/00-ensure-links
run-parts: executing /etc/runit/1.d/00-fix-var-logs
run-parts: executing /etc/runit/1.d/01-cleanup-web-pids
run-parts: executing /etc/runit/1.d/anacron
run-parts: executing /etc/runit/1.d/cleanup-pids
Cleaning stale PID files
run-parts: executing /etc/runit/1.d/copy-env
run-parts: executing /etc/runit/1.d/letsencrypt
[Tue Feb  4 06:15:03 PM UTC 2025] Domains not changed.
[Tue Feb  4 06:15:03 PM UTC 2025] Skip, Next renewal time is: 2025-02-09T00:30:10Z
[Tue Feb  4 06:15:03 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 06:15:03 PM UTC 2025] Installing key to: /shared/ssl/forum.myforum.com.key
[Tue Feb  4 06:15:03 PM UTC 2025] Installing full chain to: /shared/ssl/forum.myforum.com.cer
[Tue Feb  4 06:15:03 PM UTC 2025] Run reload cmd: sv reload nginx
warning: nginx: unable to open supervise/ok: file does not exist
[Tue Feb  4 06:15:03 PM UTC 2025] Reload error for :
[Tue Feb  4 06:15:03 PM UTC 2025] Domains not changed.
[Tue Feb  4 06:15:03 PM UTC 2025] Skip, Next renewal time is: 2025-02-09T00:30:15Z
[Tue Feb  4 06:15:03 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 06:15:04 PM UTC 2025] Installing key to: /shared/ssl/forum.myforum.com_ecc.key
[Tue Feb  4 06:15:04 PM UTC 2025] Installing full chain to: /shared/ssl/forum.myforum.com_ecc.cer
[Tue Feb  4 06:15:04 PM UTC 2025] Run reload cmd: sv reload nginx
warning: nginx: unable to open supervise/ok: file does not exist
[Tue Feb  4 06:15:04 PM UTC 2025] Reload error for :
Started runsvdir, PID is 537
ok: run: redis: (pid 552) 0s
ok: run: postgres: (pid 548) 0s
nginx: [warn] duplicate extension "wasm", content type: "application/wasm", previous content type: "application/wasm" in /etc/nginx/conf.d/discourse.conf:4
supervisor pid: 546 unicorn pid: 579

오늘 명령줄을 통해 업데이트한 후, 제 자체 호스팅 사이트 두 곳에서도 동일한 문제가 발생하고 있습니다. 이 사이트들은 커스터마이징이나 비공식 플러그인 없이 매우 기본 상태로 설치되어 있으며, 정기적으로 업데이트를 유지해 왔고, 평소에는 어려움 없이 업데이트가 잘 진행되었습니다.

현재는 위에서 @mwaniki 님이 제안하신 방법들을 시도해 보고 있으며, 결과를 확인한 후 다시 이곳에 보고드리겠습니다.

앱을 다시 빌드(rebuild)한 후 업데이트가 성공적으로 완료되었지만, 오류는 보이지 않는 상태인데도 사이트에는 접근할 수 없습니다. 혹시 아이디어가 있으신가요?

./launcher logs app


WARNING: Docker version 20.10.12 deprecated, recommend upgrade to 24.0.7 or newer.
x86_64 arch detected.
WARNING: containers/app.yml file is world-readable. You can secure this file by running: chmod o-rwx containers/app.yml
run-parts: executing /etc/runit/1.d/00-ensure-links
run-parts: executing /etc/runit/1.d/00-fix-var-logs
run-parts: executing /etc/runit/1.d/01-cleanup-web-pids
run-parts: executing /etc/runit/1.d/anacron
run-parts: executing /etc/runit/1.d/cleanup-pids
Cleaning stale PID files
run-parts: executing /etc/runit/1.d/copy-env
run-parts: executing /etc/runit/1.d/letsencrypt
[Tue Feb  4 07:12:15 PM UTC 2025] Domains not changed.
[Tue Feb  4 07:12:15 PM UTC 2025] Skip, Next renewal time is: 2025-03-06T00:39:07Z
[Tue Feb  4 07:12:15 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 07:12:16 PM UTC 2025] Installing key to: /shared/ssl/forum.******.com.key
[Tue Feb  4 07:12:16 PM UTC 2025] Installing full chain to: /shared/ssl/forum.*****.com.cer
[Tue Feb  4 07:12:16 PM UTC 2025] Run reload cmd: sv reload nginx
warning: nginx: unable to open supervise/ok: file does not exist
[Tue Feb  4 07:12:16 PM UTC 2025] Reload error for :
[Tue Feb  4 07:12:16 PM UTC 2025] Domains not changed.
[Tue Feb  4 07:12:16 PM UTC 2025] Skip, Next renewal time is: 2025-03-06T00:39:11Z
[Tue Feb  4 07:12:16 PM UTC 2025] Add '--force' to force to renew.
[Tue Feb  4 07:12:16 PM UTC 2025] Installing key to: /shared/ssl/forum.*****.com_ecc.key
[Tue Feb  4 07:12:16 PM UTC 2025] Installing full chain to: /shared/ssl/forum.ü_ecc.cer
[Tue Feb  4 07:12:16 PM UTC 2025] Run reload cmd: sv reload nginx
warning: nginx: unable to open supervise/ok: file does not exist
[Tue Feb  4 07:12:16 PM UTC 2025] Reload error for :
Started runsvdir, PID is 535
ok: run: redis: (pid 545) 0s
nginx: [warn] duplicate extension "wasm", content type: "application/wasm", previous content type: "application/wasm" in /etc/nginx/conf.d/discourse.conf:4
ok: run: postgres: (pid 548) 0s
supervisor pid: 542 unicorn pid: 575

이 “사이트가 전혀 반응하지 않음” 문제는 postgres 업데이트와 무관하다고 생각합니다. 지금 바로 확인 중입니다 :eyes:

이 문제가 해결되기를 간절히 기다리고 있습니다. 포럼에 접근할 수 없는 이유를 묻는 사용자들의 이메일을 수백 통이나 받고 있습니다 :frowning:

모두에게 불편을 드려 죄송합니다! 수정 사항이 이제 배포되었으므로, ./launcher rebuild app을 한 번 더 실행하면 서비스가 정상적으로 복구되어야 합니다.

그 이후에도 문제가 계속 발생한다면 알려주세요.

직장에서 :slight_smile:

yuppy가 작동 중이고 포럼이 다시 살아나고 있어요. 빠른 해결에 감사드립니다. 그래서 저희는 Discourse를 좋아하는 거예요 :heart:

수정해 주셔서 감사합니다! 저도 지난 2시간 넘게 리빌드 과정에서 무엇이 잘못되었는지 파악하려고 애를 썼습니다. 지금은 모두 정상입니다!

메타적 관찰: 이 포럼에서 해당 문제를 조사할 때, 포럼 검색 결과의 기본 정렬 방식인 '관련성(Relevance)'이 오히려 저에게 불리하게 작용했습니다. 제가 검색한 에러 로그는 여기와 정확히 동일했는데, 이 최근 게시글은 결과 목록에서 여러 페이지 아래에나 나타났습니다(아마도 게시 시점이 최근이라서겠죠). 그래서 우연히 메타(Meta) 첫 페이지를 열었을 때 트렌딩으로 뜨는 것을 보고 나서야 이 게시글을 찾게 되었습니다. 앞으로 리빌드 관련 문제를 조사할 때 첫 페이지나 최신 결과도 함께 확인해야 한다는 점을 저 자신과 다른 분들에게도 알려두는 메모로 남깁니다.

훌륭한 피드백입니다! 처음에는 이 대화가 PostgreSQL 15 update 에서 진행되고 있었습니다. 데이비드가 PostgreSQL 업데이트와 관련이 없다는 것을 깨닫고 관련 게시물들을 새 주제로 옮긴 후부터가 맞습니다. 이 일이 불과 한 시간 전쯤에 일어났기 때문에, 그 시점까지는 찾기가 어려웠을 것입니다!

업데이트 실패 문제를 해결하는 것은 악명 높게 어렵습니다. Discourse 업데이트는 보통 매우 순조롭게 진행되므로, 대부분 자기 호스팅을 하는 우리는 Discourse의 내부 작동 원리나 문제 해결 단계를 배울 필요가 없기 때문이죠!

이 문제를 살펴보고 신속하게 해결책을 찾아주신 @david님께 감사드립니다!

모바일 앱에서 Docker를 업데이트한 후, 늘 문제가 생길 때 나타나는 그 악명 높은 "콘솔을 통해 업데이트하세요"라는 메시지가 떴습니다.

수동 업데이트 단계를 모두 따랐지만, launcher rebuild app을 실행할 때마다 실패합니다.

launcher start app을 실행하면 복구할 수 있으므로 사이트는 정상적으로 작동하고 있습니다.

이 문제가 Postgres 오류와 관련이 있는지, 아니면 제가 어디서 문제를 겪고 있는지는 명확하지 않습니다.

이것은 이 주제에서 논의된 문제와 동일한 문제가 아님을 시사합니다.
해당 문제에서는 재구성이 오류 없이 성공했지만 사이트가 로드되지 않았습니다. ./launcher start 명령도 도움이 되지 않았습니다.

따라서 확인되는 오류에 대한 세부 정보를 포함하여 새로운 Support 주제를 개설하는 것을 권장합니다.

재구축 후 제 사이트가 다시 온라인으로 돌아왔습니다. 감사합니다! :+1:

결국 제 독해력 문제였네요. Postgres 데이터베이스가 제대로 종료되지 않았던 것이 문제였다는 것을 확인했습니다. 올바른 지시를 따랐더니 모든 것이 정상적으로 작동했습니다. 감사합니다. 일이 꼬이고 조금 당황할 때, 더 차분한 분들의 도움을 받을 수 있는 곳이 있어 정말 좋습니다.

감사합니다!!!

제에게는 작동하지 않습니다

그렇다면 다른 문제일 가능성이 있습니다. 세부 정보를 포함하여 새로운 Support 주제를 개설해 주시면 최선을 다해 도와드리겠습니다.

혼동을 피하기 위해, 이 특정 문제는 해결되었으므로 이 주제를 닫겠습니다.