mirror of
https://github.com/Sagit-chu/flvx.git
synced 2026-10-05 01:26:37 +08:00
fix: add self-healing for forward service name migration
When upgrading from older versions, service names changed from placeholder IDs (forward_user_0) to real user_tunnel IDs, causing service not found errors during control operations. - Add fallback cleanup+rebuild logic on UpdateService when service not found during the upgrade transition period. - Add self-healing retry on Pause/Resume when all variants are missing. - Refactor controlForwardServicesOnNode to support unit testing. - Add tests for shouldSelfHealForwardServiceControl and controlForwardServiceCommand helper functions. Entire-Checkpoint: a7f0c3175d06
This commit is contained in:
@@ -0,0 +1,28 @@
|
||||
# 011 转发服务名升级兼容与节点滚动升级
|
||||
|
||||
## 目标
|
||||
- 修复旧版本升级后编辑转发/隧道出现 `service not found`(service不存在)的问题。
|
||||
- 在后端加入兼容自愈逻辑,允许旧命名与新命名共存过渡。
|
||||
- 给出低风险节点升级顺序,避免一次性全量切换带来的中断。
|
||||
|
||||
## Checklist
|
||||
- [x] 定位回归路径:服务名从 `forward_user_0` 迁移到真实 `user_tunnel_id` 后,与旧运行态不一致导致控制失败。
|
||||
- [x] 在 `UpdateService` 的兼容路径加入旧服务清理后重建逻辑。
|
||||
- [x] 在 `Pause/Resume` 控制路径加入首次 not found 后自愈重试逻辑。
|
||||
- [x] 增加回归测试覆盖兼容行为。
|
||||
- [x] 执行 `go-backend` 相关测试并记录结果。
|
||||
- [x] 输出运维侧“后端先行 + agent 灰度升级 + 批量重部署”操作步骤。
|
||||
|
||||
## 变更说明(实施中)
|
||||
- 后端控制面将在检测到升级期的服务名不一致时进行自动自愈,降低人工干预和手工重建成本。
|
||||
|
||||
## 测试记录
|
||||
- 命令:`cd go-backend && go test ./internal/http/handler/...`
|
||||
- 结果:通过。
|
||||
|
||||
## 运维升级顺序(推荐)
|
||||
1. 先发布本次后端兼容补丁(无需等待所有 agent 同步升级)。
|
||||
2. 按 10%-20% 灰度分批升级 agent(低风险节点 -> 非高峰节点 -> 全量)。
|
||||
3. 每批升级后执行一次“转发批量重部署”,将运行态统一到新服务命名。
|
||||
4. 观察日志中 `service .* not found` 是否清零,再推进下一批。
|
||||
5. 全量稳定后保留兼容逻辑至少一个小版本周期,再评估收敛。
|
||||
Reference in New Issue
Block a user