Files
flvx/plans/011-forward-service-legacy-compat-and-agent-rollout.md
sagitchu f496f58a4d fix: add self-healing for forward service name migration
When upgrading from older versions, service names changed from
placeholder IDs (forward_user_0) to real user_tunnel IDs, causing
service not found errors during control operations.

- Add fallback cleanup+rebuild logic on UpdateService when service
  not found during the upgrade transition period.
- Add self-healing retry on Pause/Resume when all variants are missing.
- Refactor controlForwardServicesOnNode to support unit testing.
- Add tests for shouldSelfHealForwardServiceControl and
  controlForwardServiceCommand helper functions.

Entire-Checkpoint: a7f0c3175d06
2026-03-05 12:40:20 +08:00

1.6 KiB
Raw Permalink Blame History

011 转发服务名升级兼容与节点滚动升级

目标

  • 修复旧版本升级后编辑转发/隧道出现 service not found(service不存在)的问题。
  • 在后端加入兼容自愈逻辑,允许旧命名与新命名共存过渡。
  • 给出低风险节点升级顺序,避免一次性全量切换带来的中断。

Checklist

  • 定位回归路径:服务名从 forward_user_0 迁移到真实 user_tunnel_id 后,与旧运行态不一致导致控制失败。
  • 在 UpdateService 的兼容路径加入旧服务清理后重建逻辑。
  • 在 Pause/Resume 控制路径加入首次 not found 后自愈重试逻辑。
  • 增加回归测试覆盖兼容行为。
  • 执行 go-backend 相关测试并记录结果。
  • 输出运维侧“后端先行 + agent 灰度升级 + 批量重部署”操作步骤。

变更说明(实施中)

  • 后端控制面将在检测到升级期的服务名不一致时进行自动自愈,降低人工干预和手工重建成本。

测试记录

  • 命令:cd go-backend && go test ./internal/http/handler/...
  • 结果:通过。

运维升级顺序(推荐)

  1. 先发布本次后端兼容补丁(无需等待所有 agent 同步升级)。
  2. 按 10%-20% 灰度分批升级 agent(低风险节点 -> 非高峰节点 -> 全量)。
  3. 每批升级后执行一次“转发批量重部署”,将运行态统一到新服务命名。
  4. 观察日志中 service .* not found 是否清零,再推进下一批。
  5. 全量稳定后保留兼容逻辑至少一个小版本周期,再评估收敛。