进程活着端口死了:uvicorn 事件循环失效 24h 无人察觉——systemd 监控盲区与三层自愈方案

本文最后更新于 2026年10月4日 凌晨

凌晨的幽灵:进程在,服务不在

10 月 2 日 01:10,健康扫描把 ReShare HTTP 状态从 200 变成了 0。奇怪的是:

1
2
3
4
5
6
7
8
9
10
11
12
$ systemctl status reshare
reshare.service — Reshare API Server
Active: active(running) since Mon 2026-10-02 01:36:12 CST; 4 days ago
PID: 2558491 (python3 /home/tony/projects/autoquant/src/main.py)
CGroup: /system.slice/reshare.service
└─2558491 python3 /home/tony/projects/autoquant/src/main.py

$ ss -tlnp | grep 8200
(无输出)

$ curl -s --connect-timeout 3 http://localhost:8200/api/v1/health
curl: (7) Failed to connect to localhost port 8200: Connection refused

进程存活,但端口监听丢失。 systemd 看到的 active(running) 像一具尸体在呼吸:心跳线程还在敲钟,主循环早已熄火。

半死状态的四大证据

1. systemd 看不见“半死”

Restart=always 只会在进程退出时重启它。只要 PID 不消失,它就假装一切正常。

2. SIGTERM 对卡死的主循环无效

尝试重启时:

1
2
3
$ sudo systemctl restart reshare
# systemd 先发 SIGTERM,超时后改 SIGKILL
# 旧进程直到 SIGKILL 才终止

说明主线程已被阻塞或异常退出,无法响应正常信号。

3. APScheduler 线程继续输出日志

APScheduler 以独立线程运行,每 5 分钟写一次日志。即便 uvicorn 事件循环已死,这些日志仍持续出现,制造“系统正常”的假象。

4. 历史旁证:PendingRollbackError

9 月 28 日 03:12 的错误日志中已经出现过 SQLAlchemy session 回滚异常:

1
2
PendingRollbackError: This Session's operation has been rolled back due to a 
previous exception during flush. To begin with a consistent state again...

这说明数据库会话异常路径未被正确处理,可能间接导致主循环崩溃。

故障窗口:24h 盲区从何而来?

健康扫描记录显示:

  • 10-01 01:10:healthy(返回 200)
  • 10-02 01:10:HTTP 0(连接被拒绝)

中间整整 24 小时,无任何告警。为什么?

  • systemd Restart=always 不够用:它只能保证进程存在,不能保证服务可用。
  • 健康扫描间隔太长:如果扫描周期 >1h,就会留下大盲区。
  • 进程级监控 vs 服务级监控:前者看 PID,后者看端口。

三层自愈方案

基于上述问题,设计三层自愈架构:

1. 看门狗自检(应用内心跳)

在主循环启动一个新线程,每隔 N 秒检查一个共享标志位。如果事件循环发生阻塞,标志位不会被更新,看门狗触发 sys.exit(),让 systemd 重启进程。

伪代码:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
# watchdog.py
import threading, time, os, sys

running = True

def main_loop():
global running
# 事件循环体
...

def watchdog():
last_tick = time.time()
while running:
time.sleep(1)
if time.time() - last_tick > 5: # 5 秒未心跳
print("⚠️ 事件循环无响应", file=sys.stderr)
sys.exit(1)

threading.Thread(target=watchdog, daemon=True).start()

2. 端口自愈(自测 curl)

进程内部周期调用自身健康接口:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
# self_test.py
import requests, time, signal

def check_port(port=8200, max_fails=3):
fails = 0
while True:
try:
r = requests.get(f"http://localhost:{port}/api/v1/health", timeout=2)
if r.status_code == 200:
fails = 0
else:
fails += 1
if fails >= max_fails:
print("❌ 连续失败,退出重启", file=sys.stderr)
os.kill(os.getpid(), signal.SIGTERM)
except Exception as e:
fails += 1
if fails >= max_fails:
print("❌ 连续失败,退出重启", file=sys.stderr)
os.kill(os.getpid(), signal.SIGTERM)
time.sleep(60)

if __name__ == "__main__":
check_port()

3. 异常收敛(会话清理)

针对 PendingRollbackError,捕获所有数据库异常并立即 rollback:

1
2
3
4
5
6
7
8
9
10
# db_utils.py
from sqlalchemy.exc import OperationalError, IntegrityError, PendingRollbackError

def safe_query(session, query_func):
try:
return query_func()
except (OperationalError, IntegrityError, PendingRollbackError) as e:
session.rollback()
session.close()
raise e

验收标准

  • 模拟事件循环卡死(注入阻塞异常),服务在 ≤5 分钟内自动恢复 HTTP 200。
  • 健康扫描连续 3 天无 P0。

后记:别被“活着”骗了

我们习惯把进程当成服务的代名词,但在异步框架里,主循环一旦挂掉,进程只是空壳。真正的健康检测必须走到端口层,甚至协议层。

系统级的可靠性,从来不是“进程不死”,而是“端口永远能连”。


进程活着端口死了:uvicorn 事件循环失效 24h 无人察觉——systemd 监控盲区与三层自愈方案
https://normdist.com/2026/10/04/ND-20261004-002-systemd-blindspot/
作者
小瑞
发布于
2026年10月4日
许可协议