- Add NodeMetric/TunnelMetric/ServiceMonitor models and repository methods
- Implement metrics ingestion service with per-minute bucket aggregation
- Add health checker for node connectivity monitoring
- Wire node metrics from WebSocket SystemInfo messages
- Add tunnel metrics ingestion from flow upload endpoint
- Create monitoring REST API endpoints for nodes, tunnels, services
- Implement service monitor CRUD and execution (TCP/ICMP checks)
- Add MonitorPermission for non-admin access control
- Create frontend monitor page with node/tunnel/service views
- Add tunnel metrics ingestion from agent flow reports
- Include schema migration for tunnel_metric unique index
- Fix tunnel entry port conflict validation to use transaction
Entire-Checkpoint: 030821a7c8e3
- Update root AGENTS.md to commit 21008cc / tag 2.1.5-rc15
- Add CI workflows info (ci-build.yml, docker-build.yml, deploy-docs.yml)
- Add Repository Layer and Contract Tests to WHERE TO LOOK
- Add websocket_reporter to CODE MAP
- Add Go version conventions (1.24/1.23/1.22)
- Add PostgreSQL migration support note
- Update go-backend AGENTS.md with PostgreSQL support and contract tests
- Update go-gost AGENTS.md with CI build conventions
- Update vite-frontend AGENTS.md with component counts
- Update go-gost/x AGENTS.md with file counts and registry reference
- Update handler AGENTS.md with LOC estimates
When multiple TcpPing requests are sent in parallel for diagnosing
multiple remote addresses, the Go agent was processing them serially.
This caused later requests to timeout (10s) while waiting for earlier
requests to complete.
Changes:
- Only TcpPing commands run in goroutines for parallel execution
- TcpPing (read-only diagnostic) no longer triggers saveConfig()
- Other state-mutating commands remain synchronous with config save
- Add mutex to saveConfig() to protect concurrent file writes
When using relay connector with noDelay=false (default), connection
errors to the final target are deferred until first read/write during
Transport(). Previously the Transport() return value was ignored,
causing the marker to never be called for unreachable targets.
Now we capture the Transport() error and mark the node as failed,
enabling failover for subsequent connections.
The single-node bypass in hop.Select() was preventing FailFilter from
being applied when retry excludes reduced available nodes to one.
This caused failed forwarder nodes to keep being selected instead of
failing over to healthy alternatives.
FailFilter's built-in safety guard (len <= 1 returns as-is) ensures
the last remaining node is never permanently blocked.
- Change version.go default to 'dev' for local development
- Use version variable in WebSocket reporter instead of hardcoded '2.0.2'
- Inject version via -ldflags in CI build from tag name
Address reviewer feedback from PR #14 fix:
1. Single-node case: Bypass selector/FailFilter to ensure availability.
This matches upstream go-gost/x behavior - single nodes should always
be attempted regardless of recent failures.
2. Multi-node case: Preserve FailFilter's backoff contract. When all nodes
are marked as failed, return nil to signal 'no healthy nodes' rather
than falling back to a known-bad node. This prevents hammering unhealthy
nodes and respects the failTimeout window.
The handler's retry loop with ExcludeNodes context handles the multi-node
failover properly - this change ensures hop.Select() provides correct
information about node health status.
Fixes intermittent forwarding failures introduced by #14.
When a forwarding rule has multiple backend nodes configured, the first
node failure would cause the entire forward to fail instead of trying
the next available node.
Root causes fixed:
- FailFilter skipped filtering when only 1 node remained
- hop.Select() bypassed selector for single-node hops
- Handlers only attempted one node before giving up
Changes:
- selector/filter.go: Remove len<=1 early return, always filter failed nodes
- hop/hop.go: Remove single-node bypass, add ExcludeNodes context support
- ctx/value.go: Add ContextWithExcludeNodes/ExcludeNodesFromContext helpers
- handler/forward/local: Add maxRetries config, implement retry loop
- handler/forward/remote: Add maxRetries config, implement retry loop
- forwarder/sniffer.go: Add retry logic to dial() and dialTLS()
Closes#12