编号:L-03 严重级:high 工作线:监听与 HTTP(internal/listener、internal/httpx、serve 的 HTTP 装配) 来源:审查 P-06 依赖:B-08 (#15)(Broker.Shutdown)、D-02 (#28)(写队列关闭安全)、C-01 (#32)(可等待退出的消息循环) 被依赖:无
采用审查 P-06 的方案,按目录拆给三条线、由本 issue 负责总装:
Broker.Shutdown(ctx)
Do
ErrQueueClosed
StopAccept()
Wait(ctx)
OnMQTT
<-ctx.Done()
stop()
brk.Shutdown
db.Close()
stop_grace_period: 30s
TimeoutStopSec
断开时还没回 resp 的请求,SDK 会按原消息号重新提交,防重保证不重复。SDK 需要把 DISCONNECT 0x8B 当作可重试,由 SDK 线核对。
cmd/nixmsg/serve.go(runServe 末尾的停机段)、internal/listener/server.go、deploy/docker-compose.yml、docs/OPS.md。
cmd/nixmsg/serve.go
runServe
internal/listener/server.go
deploy/docker-compose.yml
docs/OPS.md
messageLoops
staticFileHandler
admin.Deps
Close
以下是本次复审各区审查报告的原文段落。A、M、I、P、S 开头的是原始发现编号(A 管理后台与网页、M 消息核心、I 身份认证群在线、P 传输平台部署、S SDK)。解决方案以本 issue 上方的"结论与统一方案"为准;原文里的方案与之不一致时,按上方执行。
lnSrv.Close()
wg.Wait()
OnMQTT → AttachTCP
server.go:245-249
270-275
brk.Close()
mochi.Server.Close()
ClientsWg.Wait()
EstablishConnection("tcp"/"ws")
Queue.Do
store/queue.go:78-90
296-303
signal.NotifyContext
-wal
-shm
<-ctx.Done() loopCancel() _ = lnSrv.Close() drainCtx, drainCancel := context.WithTimeout(context.Background(), 10*time.Second) defer drainCancel() if drainErr := db.Queue.Drain(drainCtx); drainErr != nil && !errors.Is(drainErr, context.DeadlineExceeded) { slog.Error("write queue drain", "err", drainErr) } return nil
func (b *Broker) Close() error { if b.closed.Swap(true) { return nil } b.queuesMu.Lock() for _, q := range b.queues { q.close() } b.queuesMu.Unlock() return b.server.Close() }
func (s *Server) Close() error { close(s.done) s.Log.Info("gracefully stopping server") s.Listeners.CloseAll(s.closeListenerClients)
其余位置:mochi listeners/listeners.go:121-135(只关注册过的监听器,然后 ClientsWg.Wait());queue.go:33-39(向可能已关闭的通道发送)。
listeners/listeners.go:121-135
queue.go:33-39
Shutdown(ctx)
OnConnect
PublishDown
b.server.Clients.GetAll()
DisconnectClient(cl, packets.ErrServerShuttingDown)
server.Close()
store.Queue
internal/broker/broker.go
queue.go
internal/store/queue.go
复审基线:main 4059a15(2026-09-30)。编号说明、各工作线的合并顺序、共享文件归属见总览 #7。
4059a15
L-03 已在 feat/fix-l03-shutdown 落地,未合入 main。提交 9a17c6a26b (#22)
feat/fix-l03-shutdown
9a17c6a26b
合成 origin/feat/fix-message-c01-c03 + origin/feat/fix-listener-l01-l07(broker / store 已含于 C-01–C-03)。未改 PublishDown 签名。
origin/feat/fix-message-c01-c03
origin/feat/fix-listener-l01-l07
做法:
Shutdown
Wait
验证(listen :0、临时目录):
:0
TestStopAcceptDoesNotWaitForMQTT
TestServeShutdownWithLiveMQTT
0x8B
go test ./cmd/nixmsg ./internal/listener ./internal/broker ./internal/store ./internal/app/message -count=1
task check 在合成分支上仍有来自 C-01 / L-01 / store 的既有 gofmt/govet,非本提交引入。TestQ2AcceptAndReport 在 C-01–C-03 基线上即因回执路径超时失败,非 L-03 引入。
task check
TestQ2AcceptAndReport
已合入 origin/main 0c9b459。落地提交 b40ef5c fix: 停机先停接受并等待循环再断开 MQTT (#22)。
0c9b459
b40ef5c
No dependencies set.
The note is not visible to the blocked user.
编号:L-03 严重级:high 工作线:监听与 HTTP(internal/listener、internal/httpx、serve 的 HTTP 装配) 来源:审查 P-06
依赖:B-08 (#15)(Broker.Shutdown)、D-02 (#28)(写队列关闭安全)、C-01 (#32)(可等待退出的消息循环) 被依赖:无
结论与统一方案
采用审查 P-06 的方案,按目录拆给三条线、由本 issue 负责总装:
Broker.Shutdown(ctx),在 B-08 实现。Do不向已关闭通道发送,关闭后返回ErrQueueClosed)。StopAccept()(关 TCP 监听、对 HTTP 调 Shutdown)和带超时的Wait(ctx);记录交给OnMQTT的连接,超时后强制关闭。<-ctx.Done()之后立即调用stop(),让第二次 Ctrl+C / SIGTERM 能直接结束进程。brk.Shutdown(5 秒)→ 用剩余时间再 Drain 一次 →db.Close()。stop_grace_period: 30s;OPS 写明 systemdTimeoutStopSec至少 30 秒。断开时还没回 resp 的请求,SDK 会按原消息号重新提交,防重保证不重复。SDK 需要把 DISCONNECT 0x8B 当作可重试,由 SDK 线核对。
改动文件
cmd/nixmsg/serve.go(runServe末尾的停机段)、internal/listener/server.go、deploy/docker-compose.yml、docs/OPS.md。与其他问题的交互 / 冲突说明
messageLoops由 C-01 改,staticFileHandler由 L-05 改,listener 装配段由 L-07 改,admin.Deps由 L-04 与 H-02 改。验收与测试
runServe在 15 秒内返回、没有 panic、客户端收到 0x8B。Do与Close1000 次不 panic。问题明细(各区审查原文,证据含文件与行号)
[P-06] 有 MQTT 连接时停机退不出,broker.Close 还可能 panic
lnSrv.Close()最后会wg.Wait(),而裸 TCP/TLS 的 MQTT 连接一直阻塞在OnMQTT → AttachTCP里,这些 goroutine 都在 wg 中,没有人去关它们(server.go:245-249、270-275)。只要有一台裸 TCP 设备在线,停机就卡在这一步,连 Drain 都不会执行。brk.Close():mochi.Server.Close()。ClientsWg.Wait();我们的连接都是EstablishConnection("tcp"/"ws")接进来的,不属于任何注册监听器,同样没人断开,于是继续卡住。db.Close()会关闭写队列的通道,而 WakePush、重推计时器和没被等待退出的 messageLoops 可能还在Queue.Do里向它发送,也可能 panic(store/queue.go:78-90、296-303)。signal.NotifyContext的 stop 要等 cmdServe 返回才调用,卡住期间再按 Ctrl+C 或再发 SIGTERM 都会被吞掉,只能靠 SIGKILL(Docker 默认 10 秒后,systemd 默认 90 秒后)。SIGKILL 会留下-wal/-shm文件,和 P-19 叠加有损坏风险。其余位置:mochi
listeners/listeners.go:121-135(只关注册过的监听器,然后ClientsWg.Wait());queue.go:33-39(向可能已关闭的通道发送)。Shutdown(ctx):OnConnect直接拒绝新连接,PublishDown返回已关闭。b.server.Clients.GetAll()(跳过内联客户端),调DisconnectClient(cl, packets.ErrServerShuttingDown)(原因码 0x8B)。server.Close(),用 ctx 限时等待。StopAccept()(关闭 TCP 监听、对 HTTP 调 Shutdown)和带超时的Wait(ctx);记录交给 OnMQTT 的连接,超时后强制关闭。<-ctx.Done()之后立刻调stop(),让第二次信号能直接结束进程。brk.Shutdown(5 秒)→ 用剩余时间再 Drain 一次 →db.Close()。store.Queue:用读写锁(或"不关闭 ch、只关闭 stop 通道")保证Do不会向已关闭的通道发送;关闭之后Do返回ErrQueueClosed。stop_grace_period: 30s;OPS 写明 systemd 的TimeoutStopSec至少 30 秒。cmd/nixmsg/serve.go、internal/broker/broker.go、queue.go、internal/listener/server.go、internal/store/queue.go、deploy/docker-compose.yml、docs/OPS.mdDo和Close1000 次。复审基线:main
4059a15(2026-09-30)。编号说明、各工作线的合并顺序、共享文件归属见总览 #7。L-03 已在
feat/fix-l03-shutdown落地,未合入 main。提交9a17c6a26b(#22)合成
origin/feat/fix-message-c01-c03+origin/feat/fix-listener-l01-l07(broker / store 已含于 C-01–C-03)。未改PublishDown签名。做法:
StopAccept()关 TCP 监听并对 HTTP 调Shutdown;Wait(ctx)超时后强关仍阻塞在OnMQTT的连接<-ctx.Done()后立刻stop();顺序 StopAccept → 取消并等待消息循环 → Drain 10s →brk.Shutdown5s → 再 Drain →Wait→db.Close()stop_grace_period: 30s;OPS 写明 systemdTimeoutStopSec至少 30 秒验证(listen
:0、临时目录):TestStopAcceptDoesNotWaitForMQTT通过TestServeShutdownWithLiveMQTT通过:已登录 WS + TCP 客户端,cancel 后runServe15s 内返回,DISCONNECT0x8Bgo test ./cmd/nixmsg ./internal/listener ./internal/broker ./internal/store ./internal/app/message -count=1通过task check在合成分支上仍有来自 C-01 / L-01 / store 的既有 gofmt/govet,非本提交引入。TestQ2AcceptAndReport在 C-01–C-03 基线上即因回执路径超时失败,非 L-03 引入。已合入 origin/main
0c9b459。落地提交b40ef5cfix: 停机先停接受并等待循环再断开 MQTT (#22)。