$ systemctl status reshare reshare.service — Reshare API Server Active: active(running) since Mon 2026-10-02 01:36:12 CST; 4 days ago PID: 2558491 (python3 /home/tony/projects/autoquant/src/main.py) CGroup: /system.slice/reshare.service └─2558491 python3 /home/tony/projects/autoquant/src/main.py
$ ss -tlnp | grep 8200 (无输出)
$ curl -s --connect-timeout 3 http://localhost:8200/api/v1/health curl: (7) Failed to connect to localhost port 8200: Connection refused
PendingRollbackError: This Session's operation has been rolled back due to a previous exception during flush. To begin with a consistent state again...
这说明数据库会话异常路径未被正确处理,可能间接导致主循环崩溃。
故障窗口:24h 盲区从何而来?
健康扫描记录显示:
10-01 01:10:healthy(返回 200)
10-02 01:10:HTTP 0(连接被拒绝)
中间整整 24 小时,无任何告警。为什么?
systemd Restart=always 不够用:它只能保证进程存在,不能保证服务可用。
健康扫描间隔太长:如果扫描周期 >1h,就会留下大盲区。
进程级监控 vs 服务级监控:前者看 PID,后者看端口。
三层自愈方案
基于上述问题,设计三层自愈架构:
1. 看门狗自检(应用内心跳)
在主循环启动一个新线程,每隔 N 秒检查一个共享标志位。如果事件循环发生阻塞,标志位不会被更新,看门狗触发 sys.exit(),让 systemd 重启进程。
伪代码:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19
# watchdog.py import threading, time, os, sys
running = True
defmain_loop(): global running # 事件循环体 ...
defwatchdog(): last_tick = time.time() while running: time.sleep(1) if time.time() - last_tick > 5: # 5 秒未心跳 print("⚠️ 事件循环无响应", file=sys.stderr) sys.exit(1)