mirror of
https://github.com/Sagit-chu/flvx.git
synced 2026-09-28 07:36:38 +08:00
9de240f034
- Add NodeMetric/TunnelMetric/ServiceMonitor models and repository methods - Implement metrics ingestion service with per-minute bucket aggregation - Add health checker for node connectivity monitoring - Wire node metrics from WebSocket SystemInfo messages - Add tunnel metrics ingestion from flow upload endpoint - Create monitoring REST API endpoints for nodes, tunnels, services - Implement service monitor CRUD and execution (TCP/ICMP checks) - Add MonitorPermission for non-admin access control - Create frontend monitor page with node/tunnel/service views - Add tunnel metrics ingestion from agent flow reports - Include schema migration for tunnel_metric unique index - Fix tunnel entry port conflict validation to use transaction Entire-Checkpoint: 030821a7c8e3
2.7 KiB
2.7 KiB
038 - Monitoring Bug Fixes + Optimizations
Context
Monitoring in FLVX currently spans:
- Agent -> panel WebSocket realtime system metrics (CPU/mem/disk/net/load/conns)
- Panel-side ingestion + retention pruning (
node_metric) - Service monitors (TCP/ICMP) with scheduled checks + stored results
- Frontend monitor page (
/monitor) with charts + monitor CRUD/run/results
While the feature set works end-to-end, there are a few correctness footguns and a couple of obvious performance hot spots (agent-side sampling cost and frontend N+1 polling patterns).
Goals
- Service monitor updates do not accidentally clear
nodeId/enabledwhen fields are omitted. - Checker cadence is explicit (intervals below the scan cadence are clamped / best-effort).
- Reduce frontend requests for service monitor status (avoid per-monitor polling).
- Reduce agent sampling overhead and DB write volume without breaking UI expectations.
- Avoid misclassifying arbitrary JSON as a metric message on the WS channel.
Non-goals
- A full scheduler (per-monitor next-run queue, jitter/backoff, concurrency budgets).
- Alerting/notifications.
- Implementing full tunnel-metrics ingestion (connections/errors/latency) beyond current endpoints.
Checklist
Phase 1: Backend Correctness + Hardening
- Make
/api/v1/monitor/services/updatetreatnodeIdandenabledas optional fields (no accidental zeroing). - Clamp
intervalSecto a minimum that matches the checker scan cadence (and apply the same clamp in the checker). - Add
GET /api/v1/monitor/services/latest-resultsreturning the latest result per monitor (for frontend list rendering). - WS metric parsing: only treat messages as metrics when they look like a system-metric payload.
Phase 2: Frontend UX + Request Reduction
- Fix “立即检查” toast severity (failure should be an error toast).
- Use
latest-resultsendpoint to render service monitor status without N+1 polling. - Add a small hint when chart data is truncated by backend row limits.
Phase 3: Agent Sampling Optimizations
- Reduce default WS metric send interval (2s -> 5s).
- Make CPU sampling non-blocking and cache heavy metrics (e.g. connection counts) to reduce per-sample cost.
Test Plan
Backend:
cd go-backend && go test ./... -count=1
Agent fork:
cd go-gost/x && go test ./... -count=1
Frontend (best-effort in this environment):
cd vite-frontend && npm run build
Rollout Notes
- Agent sampling interval change reduces metric resolution and DB growth; charts remain usable and realtime UI remains responsive.
- Existing monitors with very small
intervalSecare best-effort; effective cadence remains bounded by the checker scan loop.