Files
flvx/plans/038-monitoring-bugfixes-and-optimizations.md
sagitchu 9de240f034 feat(monitoring): add node/tunnel metrics, service monitors, and health checks
- Add NodeMetric/TunnelMetric/ServiceMonitor models and repository methods
- Implement metrics ingestion service with per-minute bucket aggregation
- Add health checker for node connectivity monitoring
- Wire node metrics from WebSocket SystemInfo messages
- Add tunnel metrics ingestion from flow upload endpoint
- Create monitoring REST API endpoints for nodes, tunnels, services
- Implement service monitor CRUD and execution (TCP/ICMP checks)
- Add MonitorPermission for non-admin access control
- Create frontend monitor page with node/tunnel/service views
- Add tunnel metrics ingestion from agent flow reports
- Include schema migration for tunnel_metric unique index
- Fix tunnel entry port conflict validation to use transaction

Entire-Checkpoint: 030821a7c8e3
2026-03-17 14:59:09 +08:00

2.7 KiB

038 - Monitoring Bug Fixes + Optimizations

Context

Monitoring in FLVX currently spans:

  • Agent -> panel WebSocket realtime system metrics (CPU/mem/disk/net/load/conns)
  • Panel-side ingestion + retention pruning (node_metric)
  • Service monitors (TCP/ICMP) with scheduled checks + stored results
  • Frontend monitor page (/monitor) with charts + monitor CRUD/run/results

While the feature set works end-to-end, there are a few correctness footguns and a couple of obvious performance hot spots (agent-side sampling cost and frontend N+1 polling patterns).

Goals

  • Service monitor updates do not accidentally clear nodeId / enabled when fields are omitted.
  • Checker cadence is explicit (intervals below the scan cadence are clamped / best-effort).
  • Reduce frontend requests for service monitor status (avoid per-monitor polling).
  • Reduce agent sampling overhead and DB write volume without breaking UI expectations.
  • Avoid misclassifying arbitrary JSON as a metric message on the WS channel.

Non-goals

  • A full scheduler (per-monitor next-run queue, jitter/backoff, concurrency budgets).
  • Alerting/notifications.
  • Implementing full tunnel-metrics ingestion (connections/errors/latency) beyond current endpoints.

Checklist

Phase 1: Backend Correctness + Hardening

  • Make /api/v1/monitor/services/update treat nodeId and enabled as optional fields (no accidental zeroing).
  • Clamp intervalSec to a minimum that matches the checker scan cadence (and apply the same clamp in the checker).
  • Add GET /api/v1/monitor/services/latest-results returning the latest result per monitor (for frontend list rendering).
  • WS metric parsing: only treat messages as metrics when they look like a system-metric payload.

Phase 2: Frontend UX + Request Reduction

  • Fix “立即检查” toast severity (failure should be an error toast).
  • Use latest-results endpoint to render service monitor status without N+1 polling.
  • Add a small hint when chart data is truncated by backend row limits.

Phase 3: Agent Sampling Optimizations

  • Reduce default WS metric send interval (2s -> 5s).
  • Make CPU sampling non-blocking and cache heavy metrics (e.g. connection counts) to reduce per-sample cost.

Test Plan

Backend:

cd go-backend && go test ./... -count=1

Agent fork:

cd go-gost/x && go test ./... -count=1

Frontend (best-effort in this environment):

cd vite-frontend && npm run build

Rollout Notes

  • Agent sampling interval change reduces metric resolution and DB growth; charts remain usable and realtime UI remains responsive.
  • Existing monitors with very small intervalSec are best-effort; effective cadence remains bounded by the checker scan loop.