Commit Graph

118 Commits

Author SHA1 Message Date
sagit c259645227 feat: implement user and forward max connection limit (#461)
* docs: add implementation plan for max conn limit

* feat: add CLimiters support for websocket reporter

* feat: add max_conn field to user and forward models

* feat: implement max conn limiter dispatching

* feat: add maxConn to user and forward CRUD API

* feat: add max conn UI to user management and forward rules

* fix: load MaxConn in get forward record and add e2e contract test for max conn limit
2026-04-23 17:35:39 +08:00
sagitchu 2b2b417f91 fix(gost): ensure JSON floats are correctly converted in GetInt and remove accidental binary 2026-04-22 11:17:59 +08:00
sagitchu 9aff669c0e fix: ensure tunnel protocol and KCP config are correctly processed and cleaned up 2026-04-22 09:41:41 +08:00
sagitchu a03c320b89 fix(gost): prevent panic in service parsing with empty md 2026-04-21 22:11:22 +08:00
sagitchu f3d6366471 fix: increase kcp tunnel bandwidth limit and fix modal backgrounds 2026-04-21 20:16:58 +08:00
sagitchu c1bc795674 fix(kcp): enable congestion control, add FEC, and fix remote node UDP diagnosis
- KCP: switch from fast3 to fast2 mode, enable FEC (10/3), enable
  congestion control (nc=0) for automatic rate adaptation
- go-gost/x: support kcp.nc metadata to override mode-initialized
  NoCongestion value in dialer and listener Init()
- Diagnosis: add udpPingViaRemoteNode, fix pingViaRemoteNode to
  dispatch based on protocol, add Protocol field to federation
  diagnose request structs
2026-04-21 15:41:32 +08:00
sagitchu 30e1473f06 fix(traffic): fix flow counter inflation from TOCTOU race in agent traffic reporter
Replace read-then-subtract pattern in collectAndReport with atomic
swap-to-zero to eliminate race where AddTraffic increments counters
between snapshot and clearReportedTraffic, causing residual traffic
to accumulate indefinitely and inflate user flow counters.

Also add defensive check in processFlowItem to skip AddFlow when
forward no longer exists, and send DeleteService to clean up orphaned
agent services.
2026-04-21 14:32:08 +08:00
sagitchu e995d70be7 feat(diagnosis): add protocol-aware connectivity diagnosis for tunnel chains
- Pass tunnel protocol through diagnosis work items
- Support protocol-specific ping (TCP/UDP/KCP) via remote nodes
- Add KCP probe support in websocket_reporter for chain hop testing
2026-04-21 14:00:02 +08:00
sagitchu 29407c90b6 perf(kcp): optimize tunnel transport defaults for high throughput
- Set KCP mode to fast3 (NoDelay:1, Interval:10ms) for lower latency
- Double SndWnd/RcvWnd from 1024 to 2048 for higher BDP
- Disable FEC (datashard:0, parityshard:0) to eliminate 30% overhead
- Disable compression for tunnel transport
- Increase relay mux MaxStreamBuffer to 2MB for better UDP throughput
- Add kcp.datashard/kcp.parityshard metadata keys support

Before: 210Mbps TCP / 35Mbps UDP (93% loss)
After:  should approach direct-connection speed (~400Mbps+)
2026-04-21 11:33:41 +08:00
sagitchu 37341af2d1 fix(kcp): prevent zero-value override of window/buffer params causing 10x perf drop
The metadata parser unconditionally assigned GetInt results (which
return 0 for missing keys) to all KCP config fields, overwriting
defaults like SndWnd=1024, RcvWnd=1024, StreamBuf=2097152 with 0.

KCP's WndSize only applies positive values, so snd_wnd=0 caused
the write path condition waitsnd < s.kcp.snd_wnd to never succeed,
effectively stalling throughput at ~10Mbps instead of 400Mbps+.

- Only apply metadata values when the key actually exists (IsExists check)
- Use Clone() instead of sharing the global DefaultConfig pointer
- Add Config.Clone() deep-copy method
2026-04-21 10:20:21 +08:00
sagitchu bb48ab00bd fix(udp): mark connection active on WriteQueue to prevent idle timeout race 2026-04-21 09:55:18 +08:00
sagitchu c1f96180f5 docs: simplify AGENTS.md files, remove stale info and redundancy 2026-04-21 00:17:31 +08:00
sagitchu 630e012ec1 fix(tunnel): resolve UDP stream interruption and add KCP protocol support
- Increase UDP listener default TTL from 5s to 30s to prevent idle disconnect
- Add mux keepalive config (15s interval, 45s timeout) to tunnel relay handler
- Add KCP as tunnel chain transport protocol with keepalive and UDP mode default
- Add KCP protocol option to tunnel UI (frontend)
- Remove generic 'tcp' fallback key from KCP metadata to prevent false TCP mode
- Simplify forward service config by removing unused tunnelTLSProtocol parameter
2026-04-20 23:57:43 +08:00
sagitchu 77b7f066f3 fix(security): update vulnerable dependencies in go-backend 2026-04-20 10:46:27 +08:00
sagitchu 1970a74f6a fix(security): update vulnerable dependencies in go-gost/x 2026-04-20 10:46:27 +08:00
sagitchu b66c4966ba fix(security): update vulnerable dependencies in go-gost 2026-04-20 10:46:27 +08:00
sagit eec6cb4298 fix(agent): add port cleanup before resuming paused services (#398)
When user traffic quota is exceeded, services are paused with
ForceClosePortConnections to kill active connections. However,
when admin resets quota and resumes services, the resume logic
was missing this cleanup, causing "address already in use" errors.

Changes:
- Add ForceClosePortConnections call in resumeServices (socket & api)
- Add ForceClosePortConnections call in resumeService (api single)
- Increase wait time from 100ms to 500ms for port release

Fixes #387
2026-03-31 10:28:28 +08:00
sagit 8ebde9dca9 feat: node OS logo and UI fixes (#367)
* chore: update AGENTS.md with next release info

* feat: node OS logo, UI rate fix, and monitor trend updates
2026-03-22 05:03:04 +00:00
sagitchu 6e3d604618 feat: implement tunnel quality polling and service monitor tuning to 1s/30s intervals 2026-03-21 17:28:52 +08:00
sagitchu 27c13d6c47 chore: release 2.1.9-beta7 2026-03-20 22:15:51 +08:00
sagit ff7c91d277 chore: optimize agent-panel metrics communication (#352) 2026-03-20 03:39:49 +00:00
sagitchu 9de240f034 feat(monitoring): add node/tunnel metrics, service monitors, and health checks
- Add NodeMetric/TunnelMetric/ServiceMonitor models and repository methods
- Implement metrics ingestion service with per-minute bucket aggregation
- Add health checker for node connectivity monitoring
- Wire node metrics from WebSocket SystemInfo messages
- Add tunnel metrics ingestion from flow upload endpoint
- Create monitoring REST API endpoints for nodes, tunnels, services
- Implement service monitor CRUD and execution (TCP/ICMP checks)
- Add MonitorPermission for non-admin access control
- Create frontend monitor page with node/tunnel/service views
- Add tunnel metrics ingestion from agent flow reports
- Include schema migration for tunnel_metric unique index
- Fix tunnel entry port conflict validation to use transaction

Entire-Checkpoint: 030821a7c8e3
2026-03-17 14:59:09 +08:00
qimaoww e5339a8072 chore: improve websocket wss->ws fallback diagnostics 2026-03-08 20:07:37 +08:00
qimaoww fbb4d82a44 feat: add wss/https auto-detect with fallback for node-backend comm 2026-03-08 20:07:37 +08:00
sagit 4bdfa50b0c docs: update AGENTS.md files with current project state (#213)
- Update root AGENTS.md to commit 21008cc / tag 2.1.5-rc15
- Add CI workflows info (ci-build.yml, docker-build.yml, deploy-docs.yml)
- Add Repository Layer and Contract Tests to WHERE TO LOOK
- Add websocket_reporter to CODE MAP
- Add Go version conventions (1.24/1.23/1.22)
- Add PostgreSQL migration support note
- Update go-backend AGENTS.md with PostgreSQL support and contract tests
- Update go-gost AGENTS.md with CI build conventions
- Update vite-frontend AGENTS.md with component counts
- Update go-gost/x AGENTS.md with file counts and registry reference
- Update handler AGENTS.md with LOC estimates
2026-02-26 09:52:11 +08:00
sagit 9a9e83dda0 docs(agents): update knowledge base with encryption, API envelope, and build conventions (#127)
Add comprehensive documentation of project conventions including:
- Encryption patterns (AES with node secret PSK)
- API envelope structure (code, msg, data, ts)
- Build peculiarities (minify: false, rolldown-vite, UPX compression)
- Unique styles (flat monorepo, asymmetric Go layout, hybrid frontend mode)
- Module boundaries and anti-patterns
- Large file hotspots and code map references

Updated 7 AGENTS.md files across root and submodules.
2026-02-15 14:33:00 +00:00
sagit f01c0481cd docs: update AGENTS.md hierarchy with new subdirectory docs
- Update root AGENTS.md with expanded anti-patterns and notes
- Add handler/AGENTS.md for high-complexity backend handlers
- Add connector/AGENTS.md for GOST connector protocols
- Add socket/AGENTS.md for GOST socket utilities
2026-02-13 07:52:17 +00:00
sagit 8628c35802 Merge branch 'main' into opencode/tidy-panda 2026-02-13 13:13:33 +08:00
sagit acea5ea76c fix(gost): filter expected net.ErrClosed noise in service handler
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-02-13 05:09:54 +00:00
sagit cdb2914dbf fix(upgrade): stabilize batch node upgrades under long-running operations 2026-02-12 04:57:58 +00:00
sagit dbd5773717 fix(ws): stabilize node connectivity with ping/pong keepalive 2026-02-12 02:27:10 +00:00
sagit 73bf672e62 fix: use systemd-run for agent upgrade/rollback restart to avoid cgroup kill 2026-02-11 08:07:04 +00:00
sagit 311840b29b feat: complete agent upgrade system with batch upgrade, version selection, progress reporting, and rollback 2026-02-11 07:01:00 +00:00
sagit e94aa01213 fix(gost): append 'B' suffix to speed limit values for correct unit parsing 2026-02-09 06:22:47 +00:00
sagit 67d8f7a381 fix(backend): correct speed limit unit conversion from Mbps to Bytes/s 2026-02-09 06:13:56 +00:00
sagit 3a14b22ebc fix: prevent nil pointer dereference in listener config parsing 2026-02-09 04:39:27 +00:00
sagit 634562e56d fix(config): support raw number string for limiter configuration 2026-02-09 03:17:02 +00:00
sagit d06e02998b fix(limiter): fix traffic limiter ScopeClient behavior to allow per-user limits 2026-02-09 03:12:33 +00:00
sagit a1454a3549 fix(agent): handle config save errors and propagate to reporter 2026-02-08 08:23:58 +00:00
sagit 7ab6545594 fix: make chain and limiter updates idempotent
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-02-08 07:04:06 +00:00
sagit 2e2c182a0d fix: make service update idempotent (upsert)
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-02-08 07:04:05 +00:00
root 2a4e7777ab fix(gost): run TcpPing commands concurrently without config save race
When multiple TcpPing requests are sent in parallel for diagnosing
multiple remote addresses, the Go agent was processing them serially.
This caused later requests to timeout (10s) while waiting for earlier
requests to complete.

Changes:
- Only TcpPing commands run in goroutines for parallel execution
- TcpPing (read-only diagnostic) no longer triggers saveConfig()
- Other state-mutating commands remain synchronous with config save
- Add mutex to saveConfig() to protect concurrent file writes
2026-02-05 06:18:56 +00:00
root 1130a55ef5 fix(gost): mark node failed when transport detects relay error
When using relay connector with noDelay=false (default), connection
errors to the final target are deferred until first read/write during
Transport(). Previously the Transport() return value was ignored,
causing the marker to never be called for unreachable targets.

Now we capture the Transport() error and mark the node as failed,
enabling failover for subsequent connections.
2026-02-05 04:09:57 +00:00
root 7c898154b3 fix(gost): remove single-node optimization to enable forwarder failover
The single-node bypass in hop.Select() was preventing FailFilter from
being applied when retry excludes reduced available nodes to one.
This caused failed forwarder nodes to keep being selected instead of
failing over to healthy alternatives.

FailFilter's built-in safety guard (len <= 1 returns as-is) ensures
the last remaining node is never permanently blocked.
2026-02-05 02:14:37 +00:00
root 09c58e2298 feat(gost): add debug logging for failover mechanism analysis
Add debug logs to trace failover behavior:
- FailFilter.Filter(): log node name, fail count, maxFails, timeSince, failTimeout
- hop.Select(): log excludeNodes list, node selection results
- handler retry loop: log maxRetries, selected nodes, dial failures

This helps diagnose issues where failover between multiple target nodes
is not working as expected.
2026-02-05 01:06:56 +00:00
root ec41202b3c fix(gost): use chain.NewNode() to properly initialize marker for failover
When creating temporary Node instances with struct literals like
&chain.Node{Addr: host}, the marker field was not initialized.
Only chain.NewNode() properly initializes marker = selector.NewFailMarker().

Without a valid marker:
- Failed nodes cannot be marked (marker.Mark() is no-op on nil)
- Subsequent selections cannot filter out failed nodes
- Failover mechanism completely fails

Fixed locations:
- sniffer.go dial(): &chain.Node{Addr: host} -> chain.NewNode("", host)
- sniffer.go dialTLS(): &chain.Node{Addr: host} -> chain.NewNode("", host)
- local/handler.go: target := &chain.Node{} -> var target *chain.Node
- remote/handler.go: &chain.Node{Addr: host} -> chain.NewNode("", host)
2026-02-04 23:10:56 +00:00
root 0443cd9ceb fix(gost): sync agent version with release tag
- Change version.go default to 'dev' for local development
- Use version variable in WebSocket reporter instead of hardcoded '2.0.2'
- Inject version via -ldflags in CI build from tag name
2026-02-04 08:22:37 +00:00
root 3e046fc80e fix(gost): restore single-node bypass and preserve FailFilter backoff
Address reviewer feedback from PR #14 fix:

1. Single-node case: Bypass selector/FailFilter to ensure availability.
   This matches upstream go-gost/x behavior - single nodes should always
   be attempted regardless of recent failures.

2. Multi-node case: Preserve FailFilter's backoff contract. When all nodes
   are marked as failed, return nil to signal 'no healthy nodes' rather
   than falling back to a known-bad node. This prevents hammering unhealthy
   nodes and respects the failTimeout window.

The handler's retry loop with ExcludeNodes context handles the multi-node
failover properly - this change ensures hop.Select() provides correct
information about node health status.

Fixes intermittent forwarding failures introduced by #14.
2026-02-04 07:46:40 +00:00
root 0273bc6921 docs: add AGENTS.md for go-gost/x/registry 2026-02-04 06:36:03 +00:00
root a98057d06a fix(gost): implement failover for multi-node forwarding rules (#12)
When a forwarding rule has multiple backend nodes configured, the first
node failure would cause the entire forward to fail instead of trying
the next available node.

Root causes fixed:
- FailFilter skipped filtering when only 1 node remained
- hop.Select() bypassed selector for single-node hops
- Handlers only attempted one node before giving up

Changes:
- selector/filter.go: Remove len<=1 early return, always filter failed nodes
- hop/hop.go: Remove single-node bypass, add ExcludeNodes context support
- ctx/value.go: Add ContextWithExcludeNodes/ExcludeNodesFromContext helpers
- handler/forward/local: Add maxRetries config, implement retry loop
- handler/forward/remote: Add maxRetries config, implement retry loop
- forwarder/sniffer.go: Add retry logic to dial() and dialTLS()

Closes #12
2026-02-04 04:12:38 +00:00