mirror of
https://github.com/Rain-kl/OpenFlare.git
synced 2026-09-28 05:46:36 +08:00
Compare commits
118 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| c3606bc6f6 | |||
| 2b3be6f3a3 | |||
| 11c8e5c7f3 | |||
| 4d52a5b097 | |||
| b1626d068f | |||
| 079fa7ee53 | |||
| 3867841a2f | |||
| 805eea5bd6 | |||
| aad059ab6e | |||
| 96abbf180d | |||
| fff4f425ae | |||
| e1753f868e | |||
| a49d07e2c2 | |||
| 1fe2763034 | |||
| bf1e77a865 | |||
| 65bc7d9b81 | |||
| 1d4533f183 | |||
| a2404a50ef | |||
| 71a715a64b | |||
| 501c754352 | |||
| a0324dc65c | |||
| bbd7f72c2d | |||
| 81739f9cb4 | |||
| 2b69f4d8d7 | |||
| b56f27632a | |||
| fc733d0295 | |||
| 453f7e5d90 | |||
| 63e3b85294 | |||
| e66dea9090 | |||
| 451ce52592 | |||
| bbf79199aa | |||
| 40232d86ba | |||
| 55db1c01ec | |||
| 3528323b50 | |||
| 2cb339258c | |||
| 63007fc8c7 | |||
| 4f8e7e66e3 | |||
| ed1efd3d54 | |||
| efd8268a5d | |||
| 0dd2cf9e80 | |||
| be5d067af9 | |||
| 3e5f3cc562 | |||
| d73aa5f9f0 | |||
| c9fc9c0eea | |||
| cdc7474d7e | |||
| 983cce3e80 | |||
| 76e9d5b0e7 | |||
| 6b75706b8e | |||
| 59fd8cda7b | |||
| cb94081ebd | |||
| 3de0a54d47 | |||
| e7b8fb2f99 | |||
| 454542c1d0 | |||
| 580c51a73a | |||
| 6b7df5a6b5 | |||
| 497edfd564 | |||
| d7d510ec97 | |||
| f4ec58c0e6 | |||
| 2ba28417b3 | |||
| 3f97193280 | |||
| 2aa0a70762 | |||
| 55c1fcd5f9 | |||
| e5d2e4b6d5 | |||
| 4f360cc7a9 | |||
| da43de56ac | |||
| f538670e1f | |||
| 2f60329886 | |||
| c76c5a697b | |||
| 9609bec8b0 | |||
| 9b89d3c630 | |||
| e87f445218 | |||
| 40eee778ac | |||
| 6c128e0be5 | |||
| 7d03154a8a | |||
| 7f8e257d33 | |||
| 4d78bc1c38 | |||
| 1813bbdba3 | |||
| aa4faddade | |||
| ab70633e9f | |||
| e1b439d6a5 | |||
| 4962bf90d1 | |||
| c455be3002 | |||
| ab0e5fecf2 | |||
| f5c9da03f4 | |||
| a16be014d4 | |||
| c85373ff47 | |||
| d7b8f44f90 | |||
| 85321888e0 | |||
| 65c02ef7a5 | |||
| 63a24da9ee | |||
| e5f6b0ad90 | |||
| 111d2900d7 | |||
| 73d8173018 | |||
| 4ecec2cf1b | |||
| 86fad02c41 | |||
| 600a7acdfb | |||
| 288b74d104 | |||
| ce28f63659 | |||
| d0414b402a | |||
| 699e95f12c | |||
| 4e3d79c001 | |||
| 5aaaf8f197 | |||
| b76f707c8b | |||
| f1f6bb858a | |||
| 305d609d0d | |||
| 5a8722ff07 | |||
| 64fbaa7ef1 | |||
| fa689aedbc | |||
| f960511cc0 | |||
| 465440fa5b | |||
| a4dd5ca9e1 | |||
| a9e4237bbf | |||
| 75d1fcf345 | |||
| f1577bf092 | |||
| 01ed2c5e36 | |||
| 80696c12fa | |||
| 3d4d99081e | |||
| 0639855653 |
@@ -7,7 +7,7 @@ description: "Wavelet 项目专用:当新增或修改 ClickHouse 批量写入
|
||||
|
||||
开始前阅读根目录 `AGENTS.md`。ClickHouse 是辅助 OLAP 存储,**厌恶高频单条写入**(过多小 part);写入路径必须优先批量或异步聚合。
|
||||
|
||||
DDL 与表结构变更见 `database-migration` 技能;本技能只覆盖**运行时写入架构**。
|
||||
DDL 与表结构变更见 `database-migration` 技能。日志/分析用途表的判定、三库回落与切换见 `logstore` 技能。本技能只覆盖**运行时写入架构**。
|
||||
|
||||
## 分层职责
|
||||
|
||||
@@ -17,7 +17,7 @@ DDL 与表结构变更见 `database-migration` 技能;本技能只覆盖**运
|
||||
| 批量框架 | `internal/infra/persistence/batchwriter/` | 泛型队列 + 按条数/时间 flush + 非阻塞入队 + 优雅停机;**各业务域独立实例** |
|
||||
| Model | `internal/model/analytics/` | 列定义、`TableName()`、`BatchInsertSQL()`(及可选 `InsertColumns()`) |
|
||||
| Repository | `internal/repository/analytics/` | `BatchInsert*` / `BatchInsertNodeAccessLogs` 等;`PrepareBatch` + 多行 `Append` + 一次 `Send` |
|
||||
| Apps | `internal/apps/<domain>/` | 采集、入队、背压;`FlushFunc` 只调 repository,不写 SQL、不 `PrepareBatch` |
|
||||
| Apps | `internal/apps/<domain>/` | 采集、入队、背压;`FlushFunc` 只调 logstore / repository,不写 SQL、不 `PrepareBatch` |
|
||||
| 装配 | `internal/platform/bootstrap/bootstrap.go` | 进程启动时调用 `Writer.Start`;初始化时需调用 `lifecycle.OnShutdown` 挂载停机钩子 |
|
||||
| 生命周期 | `internal/platform/lifecycle/lifecycle.go` | 统一协调全局并发优雅停机,业务包无需在 `bootstrap.go` 中硬编码 `Stop` 逻辑 |
|
||||
|
||||
@@ -37,6 +37,7 @@ writer.Stop(stopCtx) // close 队列 + drain + 最终 flush
|
||||
|
||||
- `QueueSize`: 10_000
|
||||
- `MaxBatchSize`: 1_000
|
||||
- `MinBatchSize`: 50(未达阈值则跳过按时间 flush,除非设了 `MaxFlushWait`)
|
||||
- `FlushInterval`: 1s
|
||||
|
||||
各域可独立覆盖;可观测低频指标可用更小 `MaxBatchSize`(如 100)与更长 `FlushInterval`(如 2–5s),但**不要**退化为逐条 `Send`。
|
||||
@@ -49,7 +50,8 @@ writer.Stop(stopCtx) // close 队列 + drain + 最终 flush
|
||||
### FlushFunc 规范
|
||||
|
||||
- 签名:`func(ctx context.Context, items []T) error`
|
||||
- 内部调用 `internal/repository/analytics` 的 `BatchInsert*`(传入 `[]analyticsmodel.X`)
|
||||
- **日志/分析用途表**:`logstore.Active(ctx)` 再调对应 `BatchInsert*`。禁止 apps 直连 `analyticsrepo` 或 `db.ChConn`。
|
||||
- 仅 CH、无需主库回落的分析表:才直接调 `repository/analytics` 的 `BatchInsert*`。
|
||||
- 在 flush 边界记录一次错误日志,不要把 DB 驱动错误直接暴露给 HTTP 客户端
|
||||
- `Start` 使用 `context.WithoutCancel(parent)`,避免请求 ctx 取消中断后台 flush
|
||||
|
||||
@@ -57,11 +59,11 @@ writer.Stop(stopCtx) // close 队列 + drain + 最终 flush
|
||||
|
||||
每个业务域拥有自己的 `Writer`、配置与 `FlushFunc`:
|
||||
|
||||
| 域 | 表 | 现状 | 目标形态 |
|
||||
| :--- | :--- | :--- | :--- |
|
||||
| 管理端审计 | `w_user_access_logs` | `risk_control` → `batchwriter` + `analyticsrepo.BatchInsert` | 已接入 |
|
||||
| 边缘访问日志 | `of_node_access_logs` | `openflare/chwriter` 异步 flush | 已接入 |
|
||||
| 可观测时序 | `of_node_metric_snapshots` 等 5 表 | `openflare/chwriter` 五表独立 writer + 进程内短 TTL 去重 | 已接入 |
|
||||
| 域 | 表 | 写入路径 |
|
||||
| :--- | :--- | :--- |
|
||||
| 管理端审计 | `w_user_access_logs` | `risk_control` → `batchwriter` → `logstore.Active` |
|
||||
| 边缘访问日志 | `of_node_access_logs` | `openflare/chwriter` → `logstore.Active` |
|
||||
| 可观测时序 | `of_node_metric_snapshots` 等 | `openflare/chwriter` 分表 writer + 进程内短 TTL 去重 → `logstore.Active` |
|
||||
|
||||
**不要**把 audit、access log、observability 并入同一 channel。
|
||||
|
||||
@@ -73,8 +75,9 @@ writer.Stop(stopCtx) // close 队列 + drain + 最终 flush
|
||||
- `len(items)==0` 直接返回
|
||||
- `db.ChConn == nil` 返回明确错误
|
||||
- 一次 `PrepareBatch` → 循环 `Append` → 一次 `Send`
|
||||
4. **Writer 胶水**(`internal/apps/<domain>/` 或 `internal/repository/analytics/<domain>_writer.go`):
|
||||
4. **Writer 胶水**(`internal/apps/<domain>/`):
|
||||
- `New` + `Start`,并在初始化逻辑内通过 `lifecycle.OnShutdown("your_writer_name", Stop)` 注册停机回调
|
||||
- 日志表的 `FlushFunc` 调 `logstore.Active`(见 `logstore` skill)
|
||||
- 业务路径 `TryEnqueue`;HTTP 背压用 `IsFull()`
|
||||
5. **测试**:
|
||||
- repository:mock `ChConn` 验证 `BatchInsertSQL` 与 append 列数
|
||||
@@ -123,14 +126,9 @@ var globalChan chan any
|
||||
|
||||
```go
|
||||
// internal/platform/bootstrap/bootstrap.go(示意)
|
||||
var userAccessLogWriter *batchwriter.Writer[*analytics.UserAccessLog]
|
||||
|
||||
func RegisterAPI(ctx context.Context) {
|
||||
// ...
|
||||
if config.Config.ClickHouse.Enabled {
|
||||
initUserAccessLogWriter(ctx) // Start writer
|
||||
risk_control.BindWriter(userAccessLogWriter) // 或逐步替换 InitLogWriter
|
||||
}
|
||||
// 日志 writer 不依赖 clickhouse.enabled:flush 时由 logstore 选库
|
||||
risk_control.InitLogWriter(ctx)
|
||||
}
|
||||
```
|
||||
|
||||
@@ -149,7 +147,8 @@ make code-check
|
||||
- flush 按 `MaxBatchSize` 与 `FlushInterval` 触发
|
||||
- `Stop` 能 drain 队列内剩余项
|
||||
- repository 层无 goroutine、无 channel
|
||||
- `clickhouse.enabled: false` 时不 `Start` writer、不入队
|
||||
- 日志表:`clickhouse.enabled: false` 时 writer 仍 `Start`,flush 走主库 logstore
|
||||
- 仅 CH 的分析表:未启用 CH 时不要 `Start`、不要入队
|
||||
|
||||
## 相关文件速查
|
||||
|
||||
@@ -157,7 +156,8 @@ make code-check
|
||||
- 连接:`internal/infra/persistence/clickhouse.go`
|
||||
- 审计写入:`internal/apps/risk_control/logics.go`
|
||||
- OpenFlare 写入胶水:`internal/apps/openflare/chwriter/writer.go`
|
||||
- 节点访问日志 repository:`internal/repository/analytics/node_access_log_writer.go`
|
||||
- 可观测 repository:`internal/repository/analytics/node_observability_writer.go`
|
||||
- 日志抽象:`internal/repository/logstore`
|
||||
- 节点访问日志 CH 实现:`internal/repository/analytics/node_access_log_writer.go`
|
||||
- 可观测 CH 实现:`internal/repository/analytics/node_observability_writer.go`
|
||||
- 生命周期管理器:`internal/platform/lifecycle/lifecycle.go`
|
||||
- Bootstrap:`internal/platform/bootstrap/bootstrap.go`
|
||||
@@ -81,7 +81,7 @@ make code-check
|
||||
ClickHouse 是**辅助 OLAP 存储**,与 PostgreSQL/SQLite 主库**完全独立**的迁移与访问管线:
|
||||
|
||||
- 主库(PG/SQLite):业务事务数据、`goose_db_version`、双方言 SQL。
|
||||
- 分析库(ClickHouse):访问日志、统计聚合等分析型数据、`goose_clickhouse_version`、单方言 SQL。
|
||||
- 分析库(ClickHouse):分析型数据、`goose_clickhouse_version`、单方言 SQL。日志用途表还必须在主库建回落并走 `logstore`(见该 skill);CH 目录仍只放 CH DDL。
|
||||
|
||||
**不要**把 ClickHouse 表结构混入 PG/SQLite 迁移目录,也**不要**在 `support-files/`、`internal/apps/` 或 `internal/repository/` 中手写 DDL。
|
||||
|
||||
@@ -118,7 +118,7 @@ ClickHouse 是**辅助 OLAP 存储**,与 PostgreSQL/SQLite 主库**完全独
|
||||
1. **Model**:在 `internal/model/analytics/` 定义 struct,`gorm:"column:..."` 与 DDL 列名一一对应;实现 `TableName()`,批量写入表可提供 `InsertColumns()` / `BatchInsertSQL()`。
|
||||
2. **Goose SQL**:在 `internal/infra/persistence/migrator/goose/clickhouse/` 新增递增版本文件(格式同主库,如 `YYYYMMDDNNNN_create_xxx.sql`),编写 `-- +goose Up` / `-- +goose Down`。
|
||||
3. **Repository**:在 `internal/repository/analytics/` 实现 `BatchInsert*`(`db.ChConn` 一次 `PrepareBatch` + 多行 `Append` + 一次 `Send`)与查询(`db.ChDB`);连接未初始化时返回明确错误,**不要**在 handler 写 SQL,**不要**在 repository 内维护 channel/goroutine。
|
||||
4. **Apps**:在 `internal/apps/<domain>/` 编排采集与入队;高频写入通过 `internal/infra/persistence/batchwriter` 各域独立实例异步 flush(详见 `clickhouse-batchwriter` 技能),`FlushFunc` 只调 repository `BatchInsert*`;管理端统计 API 只读 repository,不触达 DDL。
|
||||
4. **Apps**:在 `internal/apps/<domain>/` 编排采集与入队;高频写入通过 `internal/infra/persistence/batchwriter` 各域独立实例异步 flush(详见 `clickhouse-batchwriter` 技能)。**日志/分析用途表**还要同时建 PG/SQLite 回落并接入 `logstore`(见 `logstore` 技能),`FlushFunc` 调 `logstore.Active` 而不是 `analyticsrepo`;普通业务分析表仍只读 repository。
|
||||
|
||||
### ClickHouse 验证
|
||||
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
---
|
||||
name: "logstore"
|
||||
description: "OpenFlare / Wavelet:当新增或修改日志/分析用途表(节点访问日志、用户访问日志、可观测时序)、接入 internal/repository/logstore、切换日志主库、实现 PG/SQLite 回落,或判断一张表该走业务主库还是日志库时必须使用。"
|
||||
---
|
||||
|
||||
# 日志用途表开发
|
||||
|
||||
开始前阅读根目录 `AGENTS.md`。DDL 用 `database-migration`;高频写入队列用 `clickhouse-batchwriter`;切换任务用 `new-async-task`。本技能只回答:**这张表是不是日志表,以及如何接入可切换的日志主库。**
|
||||
|
||||
设计背景见 [日志存储解耦](../../../docs/design/logstore.md)。
|
||||
|
||||
## 先判定
|
||||
|
||||
日志表同时满足:
|
||||
|
||||
- 追加写入、几乎不更新单行
|
||||
- 按时间查询/聚合,允许按保留天数删除
|
||||
- 关闭 ClickHouse 后仍要能写、能查
|
||||
- 不参与网站/节点/证书等事务一致性
|
||||
|
||||
**不要**做成日志表:Zone、节点、配置版本、任务执行、上传元数据。这些走主库 `repository`。
|
||||
|
||||
当前日志域:
|
||||
|
||||
| 域 | 接口 | 表 |
|
||||
| :--- | :--- | :--- |
|
||||
| 节点访问日志 | `AccessLogStore` | `of_node_access_logs` |
|
||||
| 可观测 | `ObservabilityStore` | `of_node_metric_snapshots` / `of_node_edge_health` / `of_node_obs_frps` / `of_node_obs_frpc` |
|
||||
| 用户访问审计 | `UserAccessLogStore` | `w_user_access_logs` |
|
||||
|
||||
## 分层
|
||||
|
||||
| 层级 | 路径 | 职责 |
|
||||
| :--- | :--- | :--- |
|
||||
| 抽象 | `internal/repository/logstore` | 接口 + `Active`/`BuildForMigration`;apps **只**面向这里或 `repository` 门面 |
|
||||
| CH 实现 | `logstore/clickhouse_store.go` 委托 `analytics` | 原生批量 + 现有聚合 SQL |
|
||||
| 主库实现 | `logstore/postgres_store.go` | PG(按月分区)与 SQLite(普通表)共用 GORM |
|
||||
| Model | `internal/model/analytics` | 实体与批量 SQL,无 IO |
|
||||
| 入队 | `chwriter` / `risk_control` + `batchwriter` | flush 调 logstore `BatchInsert*`;CH 入队经 hooks |
|
||||
| 切换 | `of_log_db_switch` | 冻结 → `chwriter.Drain` → 逐表复制 → 翻转 |
|
||||
| 约束 | `logstore/imports_test.go` | apps 禁止 import `repository/analytics` |
|
||||
|
||||
`log_database` 只能是「随主库」或 `clickhouse`。`log_database` / `log_db_migration` 受保护。
|
||||
|
||||
## 新增一张日志表
|
||||
|
||||
1. **Model**(`internal/model/analytics`):`TableName` + `InsertColumns` / `BatchInsertSQL`。
|
||||
2. **三套 DDL**:CH `MergeTree` + `toYYYYMM`;PG `PARTITION BY RANGE(时间列)`(主键含分区键);SQLite 普通表。不要在主库建 CH 物化视图,聚合实时算。
|
||||
3. **挂到已有域或新接口**:能进 `AccessLogStore` / `ObservabilityStore` / `UserAccessLogStore` 就不要再拆包。新域才新增接口并放进 `Store`。
|
||||
4. **方法最少集**:`BatchInsert`(含 `ensureWritable`)、业务查询、`ListForMigration`、`MigrationRange`、`DeleteAll`、`DeleteBefore`、`EnsurePartitions`(仅 PG 预建)。
|
||||
5. **双实现**:CH 委托 `analyticsrepo`;GORM 共用一套,方言 SQL 放 `dialect_*.go`。零值 id 用 `idgen.NextUint64ID()`。
|
||||
6. **`buildStore`**:CH / GORM 两分支都挂上。
|
||||
7. **写入**:独立 `batchwriter`;`FlushFunc` → `logstore.Active`。节点日志/可观测走 `SetAccessLogHooks` / `SetObservabilityHooks`,不要让 apps 碰 `ChConn`。
|
||||
8. **切换任务**:`clearTarget` + `copy*` 增加该表;源数据不删,失败不翻转。
|
||||
9. **清理**:访问类走 `log_retention_days_*`;性能指标走 `metric_retention_days`。不要擅自共用错误的 TTL。
|
||||
10. **import-lint**:apps 新增对 `analytics` 或 `infra/persistence`(`batchwriter`/`idgen` 除外)的 import 必须失败。
|
||||
|
||||
## 禁止
|
||||
|
||||
- apps 直连 `analyticsrepo` / `db.ChConn` / `db.ChDB` 做日志读写
|
||||
- 只建 CH、不建主库回落
|
||||
- Handler 内逐条 `PrepareBatch`
|
||||
- 业务表塞进 logstore
|
||||
- 管理端改 `log_database` / `log_db_migration`
|
||||
|
||||
## 验证
|
||||
|
||||
```bash
|
||||
go test ./internal/repository/logstore ./internal/repository/analytics
|
||||
go test ./internal/apps/openflare/... ./internal/apps/admin/logs ./internal/apps/admin/status
|
||||
make swagger
|
||||
make code-check
|
||||
```
|
||||
|
||||
对照:`of_node_access_logs` 或 `w_user_access_logs` 的 model、三库 goose、`logstore` 双实现、`chwriter`/`risk_control` flush、`LogDBSwitchHandler`。
|
||||
@@ -28,17 +28,14 @@ description: "Wavelet 项目专用:根据自上一个正式版本 Tag 以来
|
||||
|
||||
1. 合并重复或相近提交。
|
||||
2. 删除无意义提交,例如格式化、临时调试、无关重构。
|
||||
3. 将内部实现描述改写为用户可理解的变更。
|
||||
4. 每条使用完整中文句子。
|
||||
5. 尽量说明“修复/优化了什么”以及“带来的效果”。
|
||||
6. 不要编造 commit log 中没有的信息。
|
||||
7. 不要加入 token、密钥、私有地址等敏感信息。
|
||||
8. 如果某个分类没有内容,可以省略。
|
||||
3. 将内部实现描述改写为用户可理解的变更, 说明“修复/优化了什么”以及“带来的效果”。
|
||||
4. 不要写技术细节:只描述用户可感知的行为与效果,禁止内部实现描述,例如字段名/表名/SQL(`node_id = ''`)、框架或库名称(shadcn、GORM、OpenResty)、配置或协议细节(RFC3339、ClickHouse/PostgreSQL 差异)、代码机制(`proxy_intercept_errors`、Lua 过滤器、雪花 ID)。数据库名称仅在说明受影响用户范围时使用(如「PostgreSQL 日志库下无数据」)。
|
||||
5. 如果某个分类没有内容,则省略。
|
||||
|
||||
固定使用以下分类:
|
||||
|
||||
```text
|
||||
### 新增
|
||||
### ✨ 新功能
|
||||
### 🛠 修复
|
||||
### ⚡️ 优化与改进
|
||||
### 💄 其他/体验
|
||||
@@ -46,7 +43,7 @@ description: "Wavelet 项目专用:根据自上一个正式版本 Tag 以来
|
||||
|
||||
分类规则:
|
||||
|
||||
- 新功能、新能力、新配置、新任务:放入 ### 新增
|
||||
- 新功能、新能力、新配置、新任务:放入 ### ✨ 新功能
|
||||
- Bug、异常行为、错误逻辑:放入 ### 🛠 修复
|
||||
- 性能、稳定性、接口、架构、兼容性:放入 ### ⚡️ 优化与改进
|
||||
- 日志、文案、UI、文档、开发体验:放入 ### 💄 其他/体验
|
||||
@@ -67,9 +64,8 @@ chore(release): v3.3.0
|
||||
- 新增笔记库快照备份功能,支持定时备份与手动一键恢复(仅说明新增的功能, 禁止提及新功能开发时期的优化修复等内容)。
|
||||
|
||||
### 🛠 修复
|
||||
- 修复了通过 MCP 接口操作时笔记库范围限制未正确生效的问题。
|
||||
- 修复了 MCP 接口返回数据格式不一致的问题。
|
||||
- 修复了 WebSocket 客户端异常断开后僵尸连接未及时清理的问题。
|
||||
- 修复首页「来源分布」卡片在 PostgreSQL/SQLite 日志库下无数据的问题。
|
||||
- 修复源站错误页「仅针对 GET 请求」未真正透传非 GET 响应的问题:POST/PUT 等非 GET 请求现可完整看到源站原始报错内容。
|
||||
|
||||
### ⚡️ 优化与改进
|
||||
- 优化了 WebGUI 登录机制,引入设备令牌自动轮转,减少因 IP 变化产生的冗余令牌。
|
||||
|
||||
Executable
+53
@@ -0,0 +1,53 @@
|
||||
#!/bin/bash
|
||||
# Correctness gate: must pass after every edit. Fails fast on real breakage.
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
|
||||
echo "==> go vet ./..."
|
||||
go vet ./... 2>&1 | tail -20
|
||||
|
||||
echo "==> go build ./..."
|
||||
go build ./... 2>&1 | tail -20
|
||||
|
||||
echo "==> golangci-lint run (repo config)"
|
||||
golangci-lint run 2>&1 | tail -20
|
||||
|
||||
# 全量单测(sqlite + miniredis,纯本地无需外部服务;2026-08-16 起全绿)
|
||||
echo "==> go test ./internal/... ./pkg/..."
|
||||
go test ./internal/... ./pkg/... 2>&1 | grep -E "^--- FAIL|^FAIL" | head -20 || true
|
||||
if go test ./internal/... ./pkg/... > /tmp/auto_gotest.log 2>&1; then
|
||||
:
|
||||
else
|
||||
tail -30 /tmp/auto_gotest.log
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# 前端测试(vitest;2026-08-16 起全绿)
|
||||
echo "==> pnpm exec vitest run (frontend)"
|
||||
(cd frontend && node scripts/merge-i18n-fragments.mjs && pnpm exec vitest run --reporter=dot > /tmp/auto_vitest.log 2>&1) || {
|
||||
tail -30 /tmp/auto_vitest.log
|
||||
exit 1
|
||||
}
|
||||
|
||||
# SPDX license 头门禁(repo 自带约定)
|
||||
echo "==> make license-check"
|
||||
make license-check 2>&1 | grep "needs license" | head -10 || true
|
||||
if make license-check > /tmp/auto_license.log 2>&1; then
|
||||
:
|
||||
else
|
||||
tail -15 /tmp/auto_license.log
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# 并发密集包 -race 门禁(2026-08-16 全仓 -race 清零后纳入,防回归;
|
||||
# frpc/frps 慢套件不含在此,另做全量周期验证)
|
||||
echo "==> go test -race (concurrency packages)"
|
||||
RACE_PKGS="./internal/apps/oauth/ ./internal/apps/openflare/tls/ ./internal/apps/openflare/uptimekuma/ ./internal/apps/upload/cache/ ./internal/repository/ ./pkg/cache/disk/ ./pkg/logger/ ./internal/infra/persistence/batchwriter/"
|
||||
if go test -race -count=1 $RACE_PKGS > /tmp/auto_race.log 2>&1; then
|
||||
:
|
||||
else
|
||||
grep -E "WARNING: DATA RACE|^--- FAIL|^FAIL" /tmp/auto_race.log | head -20
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "OK: checks passed"
|
||||
+169
@@ -0,0 +1,169 @@
|
||||
# Ideas backlog (代码质量)
|
||||
|
||||
## 已尝试并收尾(2026-08-16 会话,14 个实验,108→8)
|
||||
|
||||
- 生产代码 golangci 扩展集 13 类 linter 全量清理(modernize/perfsprint/
|
||||
errorlint/canonicalheader/usestdlibvars/intrange/wastedassign/errname/
|
||||
forcetypeassert/prealloc/gosec/recvcheck/exhaustive),剩余 8 处全部为
|
||||
有据可查的刻意保留项(telegram %v、3 处嵌套 struct omitempty、
|
||||
3 处 not-found 惯例、1 处 encoding/json 接收者混合)。
|
||||
- 测试代码质量维度(testifylint/usetesting/thelper)25→0。
|
||||
- 前端 eslint/tsc 0。
|
||||
- 修复中积累的工具经验:golangci-lint v2 `--fix` 的 import 管理不可靠,
|
||||
跑完必须 `goimports -w`;`--max-issues-per-linter=0` 才能拿到全量清单
|
||||
(默认 50 + max-same-issues=3 会掩盖重复模式);cyclop 与 exhaustive
|
||||
有张力(显式 case 计入复杂度)。
|
||||
|
||||
## 未来可深化方向(均经评估)
|
||||
|
||||
- 测试可运行性修复:`go test ./internal/...` 目前在 main 上就有失败
|
||||
(无本地 redis、frpc 进程测试 flaky)。修复这些环境问题后,可以把
|
||||
`go test` 加入 checks.sh,解锁 paralleltest/tparallel 维度
|
||||
(t.Parallel 提速 + 正确性,目前因共享状态+不可运行而放弃)。
|
||||
- frontend biome 格式漂移(76 文件):一次性 `make format` 提交,
|
||||
与质量修复分开做,不进基准。
|
||||
- fieldalignment:结构体内存布局优化,但会改变 JSON key 顺序且有
|
||||
位置字面量风险 —— 若做,需按文件人工核对,不进自动基准。
|
||||
- Go 1.26 新特性扫描:`go vet` 新分析器、golangci-lint 新 linter
|
||||
(如 recvcheck 之后的 new receivers 检查)随版本跟进。
|
||||
- 文档/示例代码(docs/、scripts/)质量:目前不在 golangci 范围(tests:false
|
||||
之外还有 scripts 目录),可用同一扩展集扫 scripts/ 下的 main.go。
|
||||
|
||||
## 会话收尾(2026-08-16,run #23 后)
|
||||
|
||||
- 已确认收敛:基准 5 维全下限、-race 全仓清零、双端测试全绿、发布构建可复现、
|
||||
config.example.yaml ↔ model.go 同步无漂移、无 flaky 测试。
|
||||
- 明确评估为不值得做的方向:paralleltest/tparallel(共享全局状态风险)、
|
||||
fieldalignment(JSON key 顺序变化)、biome 格式漂移(纯噪声)、
|
||||
frpc/frps 慢测试注入 backoff(为省 ~40s 改生产时序逻辑,不值)。
|
||||
- 未来如继续:可周期跑 `go test -race ./...` 全量(frpc/frps 慢套件);
|
||||
或前端 a11y 用 axe 做浏览器级审计(超出 eslint 静态规则)。
|
||||
|
||||
## 本会话新增(runs #39-#43)
|
||||
|
||||
已修复:
|
||||
- agent auth_cache negative 缓存无上限 → 10k 上限+过期清理(DoS 防护)
|
||||
- relay/flared 与 agent 三份重复 authenticateAccessToken → 共享 agent 版(负缓存共享,DB 压力下降)
|
||||
- websocket 三 hub:runWritePump 抽取、wsClientCore 嵌入(close/enqueue 单份)、broadcastAgent 合并
|
||||
- frps/frpc TOML 注入 → pkg/protocol/toml.go TOMLQuote 转义全部插值
|
||||
|
||||
评估后不修/暂缓:
|
||||
- cloudflare listMemberItems、config_version snapshot 证书循环的 N+1:管理端小 N 低频,
|
||||
加批量 repo API 属投机优化;若未来组员数量变大再做 ListZoneDomainsByIDs。
|
||||
- fatcontext ×3(oauth/upload/auth_source cache listener):别名赋值误报,非嵌套包装。
|
||||
- objectstore newOSSBackend/newWebDAVBackend 恒 nil error:跨后端工厂签名统一,刻意设计。
|
||||
- edge/updater assetNameForGOOSGOARCH 恒 "linux":跨平台预留参数,刻意泛化。
|
||||
- agent ResolverDirective explicitResolvers 原样插入 nginx conf:管理员配置属可信输入;
|
||||
若未来开放给低权限角色需加格式校验(IP 解析)。
|
||||
- pkg/render/openresty 管理端旋钮(ClientMaxBodySize 等)原样插值:管理员权限范围内。
|
||||
- frontend/settings/profile.tsx(858 行)超 AGENTS.md ~600 行指引:存量组件,拆分属
|
||||
纯重构无质量增益,暂缓;若后续要改该页面功能时顺手拆 components/。
|
||||
|
||||
## Run #44(全仓 -race 扫描)
|
||||
|
||||
- 发现并修复 upload/cache 监听器 DATA RACE:goroutine 读可变全局 db.Redis vs
|
||||
testhelper 清理置 nil。根因修复=启动时捕获 redisClient(oauth×2/repository×2
|
||||
同型监听器一并加固),StopUploadMetaCacheListener 补 done 等待。
|
||||
- 教训:testhelper 不能 import upload/cache(循环依赖);"捕获替代全局读"是
|
||||
无环的根因修法。
|
||||
- 全仓 -race 现为 0 竞争(internal/... + pkg/...);建议周期性重跑。
|
||||
|
||||
## LIKE 转义(本轮已修日志搜索 4 站点;同类遗留)
|
||||
|
||||
- 已修:analytics/node_access_log_filter.go、analytics/access_log_filter.go、
|
||||
logstore/postgres_store.go×2(PG/SQLite 加 ESCAPE '\',CH 用默认反斜杠转义)。
|
||||
新助手 pkg/util/like.go EscapeLike + 单测。
|
||||
- Run #47 已收尾全部 GORM 站点:upload.go keyword、user.go:73/76/188/229
|
||||
(含 OAuth uniqueUsername base 转义——外部输入含 _ 曾误报用户名冲突)、
|
||||
task_execution.go task_type 前缀。均加显式 ESCAPE '\'。
|
||||
- 刻意保留:upload.go:199 `image/%`(系统常量)、config_version.go:65(系统生成)。
|
||||
|
||||
## Run #48(后台 goroutine panic 防护,55db1c01)
|
||||
|
||||
- 全仓 20 处裸 go func() 零 recover → 新增 pkg/util/goroutine.go `Go(fn)`(recover +
|
||||
slog + debug.Stack,runtime.Caller 自动记录调用点无需手写名字),22 个站点全部收口
|
||||
(oauth/upload/system_config/auth_source 的嵌套 ctx-done watcher 也含)。
|
||||
- 教训:脚本括号深度匹配首轮会跳过嵌套内层 goroutine,需跑两轮;新 Go 文件必须先跑
|
||||
scripts/update_go_license.sh(license-check 会拦)。
|
||||
- 已过期记录:go test ./internal/... ./pkg/... 现全过(94 ok)——"main 上测试失败"
|
||||
不再成立。scripts/、docs/ 下 Go 文件用扩展 linter 扫过:0 issues。
|
||||
|
||||
## Run #50(发现型 linter 扫描,全证伪——勿重跑这些维度)
|
||||
|
||||
- errchkjson 12 处:全部为不可能失败的 json.Marshal(纯 string/int/[]string
|
||||
结构体;admin/logs/routers.go:131 与 waf/ip_group_sync.go:255 的 "unsafe type"
|
||||
是传递性保守标记,RawMessage/time.Time 内容来自必然成功的 marshal)。
|
||||
- spancheck 1 处(pkg/trace/trace.go:61):误报,helper 正常返回 span,
|
||||
唯一调用方 internal/infra/task/executor.go:242 有 defer span.End()。
|
||||
- unparam ×2(objectstore oss/webdav 恒 nil error):已在 #43 前评估为跨后端工厂签名统一。
|
||||
- 性能排查:正则全部包级编译(无函数内 MustCompile);包级 map 全为有界静态注册表;
|
||||
task AppendLog 走 DB 非内存累积;push escapeJSONString 用法正确。
|
||||
- 结论:Go 静态可发现的低垂果实已穷尽。剩余方向:frontend axe a11y 浏览器级审计、
|
||||
周期性 -race 重跑(上次 #49 干净)、运维类增长审查。
|
||||
|
||||
## Run #54(认证页 axe a11y 审计+修复,451ce525)
|
||||
|
||||
已修(复扫验证生效):
|
||||
- 布局级全局:sidebar 折叠按钮 aria-label、Sidebar role=navigation(region 18 节点/页清零)、
|
||||
header Kbd 对比度 text-foreground/70、空态/错误/加载 h3→p(heading-order 清零)。
|
||||
- 页面级:dashboard 4 个 Progress aria-label、users 分页 prev/next aria-label、
|
||||
admin/system 无内容 Tabs→aria-pressed 按钮组(aria-valid-attr-value critical 清零)。
|
||||
- / 与 /admin/system 现 axe 0 违规。
|
||||
|
||||
后续可做(页面级批量,工作量大):
|
||||
- admin 数据表格行内操作图标按钮(编辑/删除)与 Switch 开关无 aria-label —— 每张管理表逐个补;
|
||||
- muted 文本对比度(card description、radix tabs trigger、primary 按钮文字)—— shadcn 默认色在浅色主题下 axe 判 fail,改主题变量影响面大需设计确认。
|
||||
- 审计环境复用:后端 :3100 + CONFIG_PATH=/tmp/of-audit/config.yaml(sqlite)、docker redis --network host、
|
||||
pnpm dev --port 3002 WAVELET_BACKEND_URL=:3100;admin 密码 reset-passwd 重置。注意 :3000 是生产实例勿动。
|
||||
|
||||
## Run #54-#55(认证页 a11y 审计,两轮 keep)
|
||||
|
||||
已修复(浏览器 axe 复扫验证):
|
||||
- 全局布局:sidebar 折叠按钮 aria-label、Sidebar role=navigation、header Kbd 对比度、
|
||||
dashboard Progress aria-label、分页 prev/next、空态/加载 h3→p、admin/system Tabs→aria-pressed。
|
||||
- 主题级根因:--primary indigo-500(#6366f1) 白字对比度仅 4.27(AA 需 4.5) → indigo-600
|
||||
oklch(51.1% 0.262 276.966) ≈6.8,一处修复全站 contrast 清零。
|
||||
- 控件名:access-analytics 刷新、events-tab Switch/编辑/删除、openflare-ops Switch/Select/
|
||||
Input(htmlFor)/Textarea、table-browser/sql-console SelectTrigger;heading-order:眉题
|
||||
h4→p(cache-manager/user-detail-sheet)、卡片题 h3→p(task-manager/file-manager)。
|
||||
- 结果:dashboard、admin/system、admin/settings、admin/logs、admin/push、admin/tasks、
|
||||
admin/database、files 共 8 页 axe 0 违规。
|
||||
|
||||
审计方法(可复用):后端 :3100(CONFIG_PATH=/tmp/of-audit/config.yaml,sqlite,
|
||||
api_prefix 必须显式 /api)+ docker redis --network host(本机 bridge NAT 坏)+
|
||||
pnpm dev --port 3002 WAVELET_BACKEND_URL=:3100 + admin 密码经 reset-passwd 重置。
|
||||
axe 注入:eval 建 CDN script → Promise 轮询 window.axe → axe.run。
|
||||
教训:表单页异步渲染,须 wait≥5s 再扫否则漏报 label 规则;Radix SelectValue
|
||||
value='' 时 placeholder 不显示,combobox 无名需 aria-label 兜底。
|
||||
|
||||
## 剩余可做
|
||||
|
||||
- 抽查其余页面(websites/[zoneId]、origins/detail、responses 编辑器等富交互页)
|
||||
——contrast 已由主题修复覆盖,预期只剩个别控件名。
|
||||
- 周期性 go test -race ./... 全量重跑(上次干净为 run #49 后)。
|
||||
|
||||
## Run #56(富交互页抽查,keep,63e3b852)
|
||||
|
||||
- 扫描 11 页:websites/origins/proxy-routes/certificates/dns-accounts 直接 0 违规
|
||||
(indigo-600 主题修复已覆盖全站 contrast)。
|
||||
- 修复 3 处并复扫归零:
|
||||
1. cloudflare/components/sync-tasks-panel.tsx 状态筛选 SelectTrigger 加 aria-label
|
||||
(Radix SelectValue value='' 时 placeholder 不渲染,combobox 无名)。
|
||||
2. components/common/settings/access-token.tsx 安全提示 text-amber-600→amber-700
|
||||
(12px 小字对比度不足)。
|
||||
3. settings/notifications 面包屑页缺 h1 → sr-only h1。教训:h1 不能作为
|
||||
BreadcrumbList 子元素(axe list 规则报 list 语义破坏),须放 <Breadcrumb> 外;
|
||||
BreadcrumbPage 无 asChild 支持。
|
||||
- a11y 维度至此穷尽:累计 14 页 axe 全部 0 违规。
|
||||
|
||||
## Run #59(-shuffle=on 测试顺序随机化扫描,keep,b56f2763)
|
||||
|
||||
- 新维度:`go test -shuffle=on` 抓到 config_version 包测试顺序依赖——
|
||||
TestBuildOpenRestyConfigSnapshotOriginErrorPageDefaults 在 shuffle 下命中
|
||||
Custom 用例留在进程级 RAM 配置缓存的值(GetSystemConfigByGroup 未命中时
|
||||
ram.Set 回填,TTL 跨测试存活;:memory: DB + SetDB 换库不使缓存失效)。
|
||||
- 修复:setupOriginErrorPageSnapshotDB / setupConfigVersionTestDB 换 DB 前后
|
||||
接入既有 ram.ResetForTest()。包内 shuffle×8 + 全仓 shuffle 复扫全过。
|
||||
- 教训:默认源码顺序掩盖顺序依赖;-shuffle=on 是低成本周期扫描手段。
|
||||
全仓 -race(#58 后)同样干净。其余用 SetDB 的测试包如后续 shuffle 复发,
|
||||
同法接入 ResetForTest 即可。
|
||||
@@ -0,0 +1,60 @@
|
||||
{"type":"config","name":"前后端代码质量优化(符合最佳实践)","metricName":"total_issues","metricUnit":"","bestDirection":"lower"}
|
||||
{"run":1,"commit":"305d609","metric":108,"metrics":{"golint_canonicalheader":8,"golint_errname":1,"golint_errorlint":12,"golint_forcetypeassert":3,"golint_gosec":2,"golint_intrange":3,"golint_modernize":37,"golint_nilnil":3,"golint_perfsprint":18,"golint_prealloc":3,"golint_recvcheck":7,"golint_usestdlibvars":3,"golint_wastedassign":7,"golint_total":107,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":36},"status":"checks_failed","description":"基线:总问题 108(golangci 107 + eslint 1)。checks 失败的唯一原因:repo 自带 golangci gate 有 2 个既有 gosec G115 问题(预期内,首次修复后即绿)。","timestamp":1786871594292,"segment":0,"confidence":null,"asi":{"hypothesis":"baseline","next_action_hint":"修复 internal/apps/edge/observability/linux.go 的 2 个 G115 gosec 问题后 checks.sh 才能通过;之后每次迭代即可正常 keep/discard"}}
|
||||
{"run":2,"commit":"f1f6bb8","metric":106,"metrics":{"golint_canonicalheader":8,"golint_errname":1,"golint_errorlint":12,"golint_forcetypeassert":3,"golint_gosec":0,"golint_intrange":3,"golint_modernize":37,"golint_nilnil":3,"golint_perfsprint":18,"golint_prealloc":3,"golint_recvcheck":7,"golint_usestdlibvars":3,"golint_wastedassign":7,"golint_total":105,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":38},"status":"keep","description":"修复 internal/apps/edge/observability/linux.go 的 2 个 gosec G115 整数溢出转换:helper 改为接收 int64 b,用 gosec 认可的饱和乘法模式(uint64 域乘积 + 上界比较),去掉原 //nolint:gosec,语义不变(Bsize 恒为正)。repo 自带 gate 首次全绿。","timestamp":1786872064145,"segment":0,"confidence":null,"asi":{"hypothesis":"修复 gosec G115:multiplyUint64ToInt64 改为 accept int64 b 并采用 gosec 认可的饱和乘法模式","insight":"gosec G115 不接受分支上界证明(a > MaxInt64/b),但接受先算 uint64 乘积再 if v > MaxInt64 饱和的模式,无需 nolint","next_action_hint":"下一步批量清理 modernize(37)/perfsprint(18) 等自动可修复类别,用 golangci-lint --fix 后人工核对 diff"}}
|
||||
{"run":3,"commit":"b76f707","metric":74,"metrics":{"golint_canonicalheader":8,"golint_errname":1,"golint_errorlint":12,"golint_forcetypeassert":3,"golint_gosec":0,"golint_intrange":3,"golint_modernize":5,"golint_nilnil":3,"golint_perfsprint":18,"golint_prealloc":3,"golint_recvcheck":7,"golint_usestdlibvars":3,"golint_wastedassign":7,"golint_total":73,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":38},"status":"keep","description":"modernize 37→5(-32):interface{}→any、内置 max/min、slices/maps 辅助、strings.Cut/SplitSeq、strings.Builder(修复 mail.go O(n²) 拼接)。逐 hunk 核对语义等价;omitzero 冲突修复被自动跳过(wire 格式不变);手动清 4 处遗留 sort import + 2 处 QF1012。","timestamp":1786872502383,"segment":0,"confidence":17,"asi":{"hypothesis":"批量应用 modernize 自动修复(interface{}→any、max/min、slices.Sort/Contains、strings.Cut/SplitSeq、maps.Copy、strings.Builder)","insight":"golangci-lint --fix 会把 omitempty→omitzero 的冲突修复跳过(2个文件保留原 tag,wire 格式不变,好);但 fixer 会遗留未使用的 sort import,需手动清 4 处;mail.go 的 Builder 迁移附带暴露 2 个 QF1012,顺手用 fmt.Fprintf 修复。全部修复语义等价,已逐项核对 diff","next_action_hint":"剩余 modernize=5 应为 omitzero 冲突文件;下轮先处理 perfsprint(18):先看 --fix 是否安全再决定"}}
|
||||
{"run":4,"commit":"699e95f","metric":56,"metrics":{"golint_canonicalheader":8,"golint_errname":1,"golint_errorlint":12,"golint_forcetypeassert":3,"golint_gosec":0,"golint_intrange":3,"golint_modernize":5,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":3,"golint_recvcheck":7,"golint_usestdlibvars":3,"golint_wastedassign":7,"golint_total":55,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":47},"status":"keep","description":"perfsprint 18→0:strconv.Itoa/FormatInt/FormatUint/FormatBool 替代 fmt.Sprintf、无动词 fmt.Errorf→errors.New、纯字符串拼接。全部语义等价(已核对 diff)。修正 fixer 遗留的 import 问题(引入 goimports 统一整理)。","timestamp":1786872884713,"segment":0,"confidence":3.0588235294117645,"asi":{"hypothesis":"perfsprint --fix:%d→strconv.Itoa/FormatInt、%t→FormatBool、%s+const→拼接、无动词 Errorf→errors.New","insight":"重要:golangci-lint v2 fixer 的 import 管理不可靠(删除/添加 import 会出错,53 个文件中 5 处报 undefined)+ 遗留未用 import。已安装 goimports(repo make format 本来就需要它),对改动文件统一 goimports -w 修复。后续只要用 --fix 就要记得跑 goimports -w","next_action_hint":"剩余大头:errorlint(12)、canonicalheader(8)(usestdlibvars 同类)、recvcheck(7)、wastedassign(7)。errorlint 需手工逐处判断;先做 canonicalheader+usestdlibvars(自动可修复但要核对)"}}
|
||||
{"run":5,"commit":"d0414b4","metric":45,"metrics":{"golint_canonicalheader":0,"golint_errname":1,"golint_errorlint":12,"golint_forcetypeassert":3,"golint_gosec":0,"golint_intrange":3,"golint_modernize":5,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":3,"golint_recvcheck":7,"golint_usestdlibvars":0,"golint_wastedassign":7,"golint_total":44,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":38},"status":"keep","description":"canonicalheader 8→0 + usestdlibvars 3→0:header key 改为 Go 规范大小写(wire 格式本就如此,纯代码修正)、HTTP 方法常量替代字符串字面量。","timestamp":1786873098921,"segment":0,"confidence":2.1724137931034484,"asi":{"hypothesis":"canonicalheader+usestdlibvars --fix:Header key 统一规范大小写、GET/OPTIONS 等方法常量","insight":"GitHub header 修正前后的 wire 格式完全一致(Go 在 Set 时本来就会规范化),纯代码层面修正,零行为风险;下次遇到同类 100% 安全","next_action_hint":"剩余:errorlint(12) 需逐处人工判断(其中 3 处 err != context.Canceled、2 处 %v wrap、若干 ==/类型断言);recvcheck(7) 是模型接收者一致性;wastedassign(7) 删 TODO 赋值;intrange(3)/modernize(5)/nilnil(3)/prealloc(3)/forcetypeassert(3)/errname(1)/eslint(1)"}}
|
||||
{"run":6,"commit":"ce28f63","metric":38,"metrics":{"golint_canonicalheader":0,"golint_errname":1,"golint_errorlint":12,"golint_forcetypeassert":3,"golint_gosec":0,"golint_intrange":3,"golint_modernize":5,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":3,"golint_recvcheck":7,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":37,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":47},"status":"keep","description":"wastedassign 7→0:删除 7 处死初始化(snapshot.go 三连、push 三件套 content、format.go numStr),改 var 声明,零行为变化。","timestamp":1786873485497,"segment":0,"confidence":2.978723404255319,"asi":{"hypothesis":"wastedassign 7→0:删除 7 处死初始化(x := \"\" 后所有分支都赋值)改为 var 声明","insight":"replace 工具会归一化 replacement_text 的前导空白;对需要缩进的编辑直接用 sed/gofmt -w 处理更稳","next_action_hint":"剩余:errorlint(12)、recvcheck(7)、modernize(5)、intrange(3)、nilnil(3)、prealloc(3)、forcetypeassert(3)、errname(1)、eslint(1)"}}
|
||||
{"run":7,"commit":"288b74d","metric":33,"metrics":{"golint_canonicalheader":0,"golint_errname":1,"golint_errorlint":12,"golint_forcetypeassert":3,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":3,"golint_recvcheck":7,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":32,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":45},"status":"keep","description":"intrange 3→0 + modernize 5→3:for i:=0;i<len/N;i++ → range len/N(8 处);time.Time 字段 omitempty→omitzero(wire 输出一致);SplitSeq;min() 简化。刻意保留 lark.go omitzero(会改变 wire 行为)。","timestamp":1786873629461,"segment":0,"confidence":4.166666666666667,"asi":{"hypothesis":"intrange(3) + modernize 剩余(2 个 time.Time omitempty→omitzero + SplitSeq + min)","insight":"lark.go larkTextContent omitempty→omitzero 会改变 wire(普通 struct 无 IsZero,当前恒序列化,改后零值省略)—— 判定为行为变化,故意保留;time.Time 字段 omitempty/omitzero 输出一致,可安全替换","next_action_hint":"剩余:errorlint(12) 大头(3 处 != context.Canceled 需确认 runner 是否 wrap;%v→%w 2 处;若干 ==err / 类型断言);recvcheck(7);forcetypeassert(3);nilnil(3);prealloc(3);errname(1);eslint(1)"}}
|
||||
{"run":8,"commit":"86fad02","metric":22,"metrics":{"golint_canonicalheader":0,"golint_errname":1,"golint_errorlint":1,"golint_forcetypeassert":3,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":3,"golint_recvcheck":7,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":21,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":46},"status":"keep","description":"errorlint 12→1:3 处 cmd 入口 err!=context.Canceled→errors.Is(防御性,当前 runner 不 wrap 语义不变);2 处 strconv.NumError 断言、1 处 viper 断言、2 处 ==io.EOF、2 处 ==redis.Nil、1 处 ==gorm.ErrRecordNotFound→errors.As/Is;8 处 %v→%w 保留错误链。刻意保留 telegram.go 单处 %v(原始错误仅作上下文文本,wrap 会改变 errors.Is 匹配语义)。","timestamp":1786873923775,"segment":0,"confidence":4.195121951219512,"asi":{"hypothesis":"errorlint 12→1:errors.Is/As 替代 ==/类型断言(防御 wrap),%v→%w 保留错误链","insight":"errorlint 结果在并行分析时一度不稳定(可能文件缓存竞争),多跑一次确认;telegram.go 的 %v 是刻意保留原始 HTML 错误为文本(只 wrap fallbackErr),判定为合理例外,不计为负债。错误链保留(%w)对多错误组合消息(manager.go、restart_unix.go、service.go)是净收益,调用方无 Is 匹配这些次要错误","next_action_hint":"剩余:recvcheck(7)、forcetypeassert(3)、nilnil(3)、prealloc(3)、modernize(3=lark omitzero 刻意保留)、errname(1)、eslint(1)"}}
|
||||
{"run":9,"commit":"4ecec2c","metric":15,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":7,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":14,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":43},"status":"keep","description":"forcetypeassert 6→0(缓存 list 断言、relay/flared 中间件契约断言、图片压缩 flight 断言,全部带检查+安全失败路径);errname 1→0;prealloc 2 处(另 1 处与 repo mnd 冲突,用命名常量解决)。nilnil 保留(not-found/可选结果惯例,含接口契约注释)。","timestamp":1786874283774,"segment":0,"confidence":4.043478260869565,"asi":{"hypothesis":"forcetypeassert(6处) → 带检查断言(middleware 契约破坏时 Abort 401/返回错误);errname runtimeInitErr→errRuntimeInit;prealloc 2 处(uptimekuma、postgres replicas)","insight":"prealloc 与 repo mnd 门禁冲突(magic number 3):用命名常量 baseTracingOptionCount 同时满足两者;nilnil 5 处判定为合法 not-found/可选结果惯例(含接口注释契约 + 测试断言),全部保留;用 --max-issues-per-linter=0 拿全量清单避免被默认 50 截断误导","next_action_hint":"剩余:recvcheck(7) 接收者一致性(需逐模型判断)、eslint(1) exhaustive-deps、modernize(3=lark omitzero 刻意保留+2 处待查)、nilnil(3 刻意保留)、errorlint(1 刻意保留)"}}
|
||||
{"run":10,"commit":"73d8173","metric":9,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":1,"eslint_errors":0,"eslint_warnings":1,"tsc_errors":0,"measure_s":45},"status":"keep","description":"recvcheck 7→1:6 个 GORM 模型 TableName 改为指针接收者(GORM 源码确认 reflect.New 判定 Tabler,兼容;模型单测通过)。MillisecondDuration 刻意保留(encoding/json 要求 Marshal 值/Unmarshal 指针的混合)。","timestamp":1786874445733,"segment":0,"confidence":4.304347826086956,"asi":{"hypothesis":"recvcheck 7→1:GORM 模型 TableName 值接收者→指针接收者,与其它方法一致","insight":"GORM schema.Parse 用 reflect.New(modelType) 判定 Tabler,指针接收者 TableName 完全兼容(已读 gorm 源码确认 + 模型单测通过);仓库中 (Model{}).TableName() 字面量调用都在未改的类型上,无破坏。MillisecondDuration 保留:MarshalJSON 值接收者是 json 对不可寻址值的行为保障,UnmarshalJSON 必须指针 —— 混合是 encoding/json 硬性要求","next_action_hint":"剩余:modernize(3,含 lark omitzero 刻意保留 + 2 处待查)、nilnil(3 刻意保留)、eslint(1 exhaustive-deps)、errorlint(1 刻意保留)。下一步查 modernize 剩余 2 处并修 eslint 的 hook 依赖"}}
|
||||
{"run":11,"commit":"111d290","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":38},"status":"keep","description":"eslint 1→0:pages-source-card useEffect 补 t 依赖(next-intl 稳定引用)。modernize 补 1 处 time.Time omitzero。剩余 8 全部为刻意保留项。","timestamp":1786874578893,"segment":0,"confidence":4.3478260869565215,"asi":{"hypothesis":"eslint 1→0:useEffect 依赖数组补 t(next-intl useTranslations 返回稳定引用,安全);modernize 补 1 处 time.Time omitempty→omitzero(输出一致)","insight":"modernize 剩余 3 处全部是嵌套 struct omitempty(client.go Release/Asset、lark.go Content)→ omitzero 会改变 wire,全部刻意保留。至此所有可安全修复的类别清零,剩余 8 个全部是有据可查的刻意保留项","next_action_hint":"剩余 8 全部刻意保留(errorlint 1 telegram、modernize 3 嵌套struct、nilnil 3 not-found、recvcheck 1 json)。下一轮做深化方向:测试代码质量(tests:false 之外)、或 golangci 附加 linter(gocritic 更多检查)作为新基准段"}}
|
||||
{"run":12,"commit":"e5f6b0a","metric":33,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":20,"golint_test_thelper":3,"golint_test_usetesting":2,"golint_test_total":25,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":37},"status":"keep","description":"基准扩展(文档化):新增测试代码质量维度 25 处(testifylint 20 + thelper 3 + usetesting 2),生产代码 8 处刻意保留不变。新基线 total=33。","timestamp":1786874744438,"segment":0,"confidence":4.878048780487805,"asi":{"hypothesis":"扩展基准到测试代码质量维度(testifylint 20 + thelper 3 + usetesting 2 = 25)","insight":"刻意排除 paralleltest/tparallel(共享 DB/redis 状态 + 本环境无法跑测试,t.Parallel 有风险)—— 这是范围扩展(抬高门槛),不是 gaming;基准定义已写入 prompt.md","next_action_hint":"修 25 处测试问题:float-compare 3(InDelta)、require-error 3、encoded-compare 1(JSONEq)、empty 3、contains 1、error-is-as 3、len 3、go-require-in-handler 2、t.Helper 3、os.MkdirTemp→t.TempDir 2"}}
|
||||
{"run":13,"commit":"63a24da","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":39},"status":"keep","description":"测试代码质量 25→0:assert↔require 一致性(fail-fast)、float 精确比较→InDelta、Equal(\"\",x)→Empty、Equal(len)→Len、errors.Is/As→ErrorIs/ErrorAs、JSON 字符串→JSONEq、handler goroutine 内 require→assert(真健壮性修复)、t.Helper()、os.MkdirTemp→t.TempDir()(符合 repo AGENTS 约束)。","timestamp":1786875177918,"segment":0,"confidence":4.3478260869565215,"asi":{"hypothesis":"修完测试代码质量维度 25 处(testifylint 20 + thelper 3 + usetesting 2)","insight":"批量修复过程揭示 testifylint 默认 max-same-issues=3 会掩盖重复模式(len(entries) 出现 4+ 次、float64(3) 4 次),需 --max-issues-per-linter=0 反复收敛;全部修复语义中性(assert↔require 仅 fail-fast 差异,InDelta/JSONEq/Empty/Len/ErrorIs 等价断言,t.Helper/t.TempDir 纯改善)。go-require 类(handler 内 require→assert)是真正的健壮性修复","next_action_hint":"测试维度清零。生产代码剩余 8 全部刻意保留。可选深化:gocritic 更多子检查/staticcheck 更多(SA 系列)扫描、或 biome check 格式一次性提交、或前端 a11y(eslint jsx-a11y 已含于 next core-web-vitals 默认关闭项)"}}
|
||||
{"run":14,"commit":"65c02ef","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":36},"status":"keep","description":"基准扩展 exhaustive(文档化)+ 12→0:枚举 switch 补显式 case(全部与现有 default 行为等价,fail-explicit 防未来枚举静默落入 default);source_tasks.go 为控制复杂度合并两个等价校验条件。","timestamp":1786875548060,"segment":0,"confidence":4.25531914893617,"asi":{"hypothesis":"基准扩展 exhaustive(12 处枚举 switch 显式化)+ 全量修复","insight":"12 处全部是 default 已正确处理、缺显式 case 的类型;补显式 case 仅为 fail-explicit(未来枚举新增不会静默落入 default)。source_tasks 补 case 后 Execute 复杂度 20→21 触发 cyclop,合并两个 ActionInvalid 条件(逻辑等价)降回 19。cyclop 与 exhaustive 的张力:显式 case 也计入复杂度","next_action_hint":"剩余 8 全为刻意保留。可再深化:sloglint 全量、govet 附加分析器、或前端 jsx-a11y/next 规则已有覆盖。也可将剩余 8 处文档化后收尾总结"}}
|
||||
{"run":15,"commit":"d7b8f44","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":37},"status":"keep","description":"修复 geoip/runtime.go 真死代码:ensureServerMMDB 的 os.Stat 错误被 if-init 遮蔽,`err != nil && !os.IsNotExist(err)` 恒为 false(外层 err 恒 nil),防御检查从未生效;改为显式捕获 statErr,stat 非 not-exist 错误现在正确返回。基准新增第 4 维度 govet nilness+unusedwrite(文档化扩展),当前 0。","timestamp":1786875949461,"segment":0,"confidence":4.166666666666667,"asi":{"hypothesis":"govet nilness 真实死代码 bug:ensureServerMMDB 的 stat 错误被 if-init 遮蔽,!os.IsNotExist(err) 恒为死条件(外层 err 恒 nil)","insight":"修复:显式捕获 statErr,使防御检查生效(stat 权限错误现在立即返回,不再静默吞掉后走 WriteFile 失败)。顺带基准扩展第 4 维度 govet nilness+unusedwrite(文档化,survey 过 fatcontext/containedctx/unparam/gocritic+29 检查:unparam 有 6+ 处真实死结果但需签名改动,留待下轮)","next_action_hint":"下轮候选:unparam(6+ 处 always-nil/never-used 结果,含 getSQLiteOverview/getPostgresOverview/getStatus 等,需改签名+调用方,churn 中等但都是真实死代码);或 fatcontext/containedctx(3+3 处,需逐处判断是否真反模式)"}}
|
||||
{"run":16,"commit":"c85373f","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":43},"status":"keep","description":"unparam 死代码清理 12→2(保留 2 处 objectstore 构造函数统一签名):移除 10 处恒 nil error / 从未使用的结果(getPoWConfigForRoute 的恒 nil *PoWConfig、getSQLiteOverview/getPostgresOverview/getStatus/loadKumaConfig/filterExpectedRoutes 的恒 nil error、rawJSONString/parsePositiveInt 的弃用 bool、buildProxyRoute 的弃用 []ZoneDomain、getLocked 的恒 nil error),同步简化 12+ 处调用方与死错误检查。9 个受影响包测试通过。metric 持平 8(改进在基准之外)。","timestamp":1786876191447,"segment":0,"confidence":5.128205128205129,"asi":{"hypothesis":"unparam 死代码清理:10 处 always-nil error / never-used 结果从签名移除","insight":"移除后调用方同步简化(db_manage 的 err 检查、option routers 的 AbortBadRequestOnError 成为死代码一并删)。getPoWConfigForRoute 的 *PoWConfig 结果恒 nil 且从未被用 —— 真死代码。保留 2 处 objectstore 构造函数 (X, error):factory switch 统一签名(newS3Backend/newLocalBackend 等可能真实报错),unparam 在此为接口一致性误报。全部 9 个受影响包测试通过。metric 持平 8(改进在基准之外,诚实记录)","next_action_hint":"下一候选:fatcontext(3 处嵌套 context 闭包,多为 slog/otel ctx 传递,需逐处判断是否真反模式) 或 containedctx(3 处 struct 含 ctx 字段,含 webdav/uptimekuma client —— 重构风险中等);或收尾把 unparam 加入基准(2 处已知保留)"}}
|
||||
{"run":17,"commit":"a16be01","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":40},"status":"keep","description":"修复 frpc 进程生命周期真 bug(agent 生产代码):exec.CommandContext 默认只杀直接子进程,被杀 shell 的孤儿 sleep 继续持有 stderr 管道,cmd.Wait() 阻塞到其自然退出(Stop/重启可挂起秒级)。改 Setpgid 进程组 + Kill(-pid) 整组击杀。连带修复两个测试 bug(Manager 拥有 Cmd 的并发 Wait 竞态 → Signal(0) 探测;ssl_renew 用 miniredis 替代 init() 创建的真实 redis 客户端)。go test ./internal/... ./pkg/... 全绿,checks.sh 升级为真实测试门禁。","timestamp":1786877266517,"segment":0,"confidence":7.142857142857143,"asi":{"hypothesis":"frpc 进程生命周期真 bug:exec.CommandContext 只杀直接子进程,孤儿孙进程持有 stderr 管道导致 cmd.Wait 阻塞到其自然退出(实测脚本 sleep 5 时 Stop 挂起 5s)","insight":"修复:Setpgid 独立进程组 + cmd.Cancel 覆盖为 Kill(-pid,SIGKILL) 整组击杀(经隔离复现 + 临时插桩定位,4 次假设检验收敛)。连带修复两个测试 bug:TestStopCancelsRunningProcesses 对 Manager 拥有的 Cmd 并发 Wait(与 os/exec ctxResult 通道竞争永久挂起)改为 Signal(0) 探测;ssl_renew 测试改用 miniredis(task 包 init() 创建真实 redis 客户端,违反 repo 无 init 装配约束)。成果:go test ./internal/... ./pkg/... 从 3 个失败→全绿(81+13 包),checks.sh 升级为真实测试门禁。metric 持平 8(改进在基准之外,但价值最高的一轮)","next_action_hint":"测试全绿后可解锁:paralleltest/tparallel 维度(t.Parallel 提速)——需先评估共享状态(miniredis/sqlite 每测试独立,风险低);或探索 relay/frps 同构代码是否有同样的 group-kill 问题(frps/manager 结构相同,值得检查)"}}
|
||||
{"run":18,"commit":"f5c9da0","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":39},"status":"keep","description":"前端测试套件 44 失败→全绿:10 个测试文件补 NextIntlClientProvider 包装(含 React19 createElement 类型修复、.ts→.tsx 重命名);修复真实 i18n ICU bug(githubUrlInvalid 的 {owner}/{repo} 未转义导致生产渲染成 key,zh/en + fragment 4 文件同步转义);更新 2 处过期测试期望。vitest 116/116 + tsc + eslint 全绿,checks.sh 增加前端测试门禁。","timestamp":1786878501539,"segment":0,"confidence":9.523809523809524,"asi":{"hypothesis":"前端测试可运行性:next-intl 迁移后 44/116 测试失败(缺 NextIntlClientProvider + 3 处真实断言问题)","insight":"修复三类:(1) 10 个测试文件的 render 助手缺 NextIntlClientProvider(createElement 与 JSX 混用踩 React19 类型坑,.ts 文件不能写 JSX → 重命名为 .tsx);(2) 真实 i18n bug:githubUrlInvalid 消息的 {owner}/{repo} 被 ICU 当占位符,t() 无参调用渲染成 key —— 需 '{' 单引号转义('{}' 内层转义不够,必须整体引号包裹 '{owner}'),4 个消息文件(zh/en + fragment 源)同步修复,check:i18n 通过;(3) 2 处测试期望过期(唯一访问者→查询窗口独立访客、检查间隔→检查间隔(分钟),以消息文件为准)。成果:116/116 vitest + tsc/eslint 全绿,checks.sh 增加前端测试门禁","next_action_hint":"前端测试全绿后可把 vitest 失败数纳入基准(当前不在基准内);或检查 app/(main) 目录下 3 个自带 .test.tsx(waf editor 系列)是否也符合新约定"}}
|
||||
{"run":19,"commit":"c455be3","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":116,"measure_s":62},"status":"keep","description":"基准扩展第 5 维度(文档化):前端 vitest 失败数纳入 total_issues(vitest_failed=0, total=116)。5 维全部处于下限,total=8 不变。","timestamp":1786878719509,"segment":0,"confidence":14.285714285714286,"asi":{"hypothesis":"基准扩展第 5 维度:前端 vitest 失败数(全绿后纳入防回归,文档化范围扩展非作弊)","insight":"measure_s 从 39s 升到 62s(vitest ~20s + eslint 冷启动),可接受。5 个维度全部在其下限:生产 8(全刻意保留)+ 测试 0 + govet 0 + eslint/tsc 0 + vitest 0","next_action_hint":"基准已 5 维全下限。后续可深化:paralleltest(现在测试可跑,但共享全局状态风险仍在,低优先);或 frontend biome 格式一次性提交(不进基准);或前端组件更深规则(jsx-a11y 已在 next core-web-vitals 覆盖)。也可认为会话到达稳定收尾点,更新 prompt/ideas 后总结"}}
|
||||
{"run":20,"commit":"4962bf9","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":116,"measure_s":86},"status":"keep","description":"两处真实质量修复:(1) 过期 swagger 文档重新生成(status_2xx/4xx/5xx_count 字段随 a4dd5ca9 加入后未同步 docs,违反 repo 约定,swag init 后差异仅真实新增字段);(2) generate-themes.js 输出补尾换行,themes.json 构建可复现(此前每次 build 弄脏工作树)。验证 next build 成功、musttag/tagalign 调查无真实问题。","timestamp":1786879144888,"segment":0,"confidence":25,"asi":{"hypothesis":"验证生产构建 + 修两处真实质量问题:swagger 文档过期(status_2xx/4xx/5xx_count 新增字段未重新生成)与 themes.json 构建不可复现(generate-themes.js 缺尾换行,每次 build 弄脏工作树)","insight":"next build 成功(无构建问题);musttag 3 处与 tagalign 均判定为非问题(持久化 round-trip 自洽/调试日志/纯格式)。swagger 差异仅 27 行且全部真实(a4dd5ca9 状态码拆分字段)。generate-themes.js 补 '\\n' 后 themes.json 再生与提交版完全一致,构建可复现。metric 持平 8(改进在基准之外)","next_action_hint":"会话已 5 维全下限 + 构建可复现 + 双端测试全绿。收尾候选:更新 prompt/ideas 记录本轮成果后总结;或继续验证 swag 生成的 docs.go 在 CI 中的可复现性"}}
|
||||
{"run":21,"commit":"e1b439d","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":116,"measure_s":74},"status":"keep","description":"全仓 go test -race 扫描(93 包)→ 全绿。修复 6 类数据竞争:frpc/frps 测试的锁外读与并发 Wait;oauth/repository 4 个 Pub/Sub 监听器 goroutine 读可变包变量(局部捕获 + done 通道等待);oauth 测试换 db.Redis 前停监听器;【真实生产 bug】tls 响应快照与异步续签 goroutine 并发写 cert 竞争(先快照再起 goroutine);upload/cache 监听器 goroutine 内读 db.Redis(调用方捕获)。","timestamp":1786881379399,"segment":0,"confidence":100,"asi":{"hypothesis":"全仓 -race 扫描:发现并修复 6 类数据竞争,含 1 个真实生产竞争","insight":"修复清单:(1) frpc 测试助手锁外读 proc.Status;(2) frps 测试对 Manager 拥有的 Cmd 并发 Wait(与 frpc 同类);(3) oauth/repository 4 个 Pub/Sub 监听器 goroutine 内读可变包变量 userListenerCtx 等 → 局部捕获 + stop 增加 done 通道等待;(4) oauth 测试 setupTestRouter 换 db.Redis 前先停各层监听器;(5)【真实生产 bug】tls logics 的 sanitizeCertificateForResponse 整体拷贝 cert 与异步续签 goroutine 字段写入并发 → 先快照再起 goroutine;(6) upload/cache 监听器 goroutine 内读 db.Redis → 调用方捕获。成果:93 包 go test -race 全绿。基准 5 维全下限不变(改进在基准外,但 tls 竞争是生产级真实问题)","next_action_hint":"可考虑把 -race 纳入周期性验证(不进每次 checks,全仓 ~3 分钟);或在 prompt/ideas 记录本轮成果"}}
|
||||
{"run":22,"commit":"ab70633","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":116,"measure_s":68},"status":"keep","description":"checks.sh 新增并发密集包 -race 门禁(8 个快速包,全仓 -race 清零后纳入防回归;frpc/frps 慢套件留作周期全量验证)。核查 7 处 t.Skip 均为合法环境门控。","timestamp":1786881700477,"segment":0,"confidence":200,"asi":{"hypothesis":"把 -race 门禁纳入 checks.sh(并发密集包子集,全仓 -race 清零后防回归)","insight":"选 8 个快速并发包(oauth/tls/uptimekuma/upload-cache/repository/disk-cache/logger/batchwriter,约 20s),排除 frpc/frps 慢套件(15s backoff + 25s supervisor,另做周期全量验证)。checks 总耗时 ~60s 可接受。顺带核查 7 处 t.Skip 均为合法环境门控(PG/CH 集成、symlink 平台差异),无隐藏损坏。metric 持平 8","next_action_hint":"会话已达全面收敛:5 维基准全下限 + 全仓 -race 清零 + 双端测试全绿 + 构建可复现 + -race 门禁入 checks。可收尾:更新 prompt/ideas 后总结"}}
|
||||
{"run":23,"commit":"aa4fadd","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":116,"measure_s":81},"status":"keep","description":"补齐 131 个 .go 文件的 SPDX license 头(repo 自带 make license 约定,早于约定新增的文件含 2 个生产文件;纯注释插入零行为影响),make license-check 转绿。go mod tidy -diff 确认干净。","timestamp":1786882164158,"segment":0,"confidence":null,"asi":{"hypothesis":"跑 repo 自带门禁 make license-check:发现 131 个 .go 文件缺 SPDX 头(早于 license 约定,含 2 个生产文件与多份本会话早期改过的文件)","insight":"用 repo 自带 scripts/update_go_license.sh(make license)自动补头(391 行纯注释插入),license-check 转绿;go mod tidy -diff 干净。gofmt/build/全测试/-race 均通过。metric 持平 8","next_action_hint":"会话全维度收敛。收尾:把 license-check 加入 checks.sh 防再犯;更新 prompt/ideas 最终状态"}}
|
||||
{"run":24,"commit":"4d78bc1","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":116,"measure_s":83},"status":"keep","description":"收敛验证轮:3× 时序敏感包测试无 flake;config.example.yaml ↔ model.go 同步确认无漂移(scheduler 空结构、replicas 注释化、queues[].name 嵌套均为误报);make build-embedded 发布路径构建成功且工作树干净(bin/、frontend/out 已 gitignore)。更新 prompt/ideas 记录会话终态。metric 持平 8。","timestamp":1786882694298,"segment":0,"confidence":null,"asi":{"hypothesis":"会话收敛验证:跑 3× 时序敏感包(无 flake)、config.example.yaml↔model.go 同步(3 个疑似缺失均为误报:scheduler 空结构/replicas 注释化/queues[].name 嵌套)、make build-embedded 发布路径(成功且树干净)","insight":"全部为确认性结果,无新问题。诚实记录:本轮零代码改动(上一轮已修 license),metric 持平 8。会话正式收敛:基准 5 维全下限 + 全仓 -race 清零 + 双端测试全绿 + 发布可复现 + config/文档同步。prompt/ideas 已更新终态","next_action_hint":"会话已收敛。若继续:周期全量 -race、前端 axe 浏览器级 a11y 审计(ideas.md 已记录)"}}
|
||||
{"run":25,"commit":"7f8e257","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":118,"measure_s":71},"status":"keep","description":"前端真实 a11y 审计:新增 axe-core(devDep)+ tests/a11y.test.tsx,对登录页与注册页渲染完整表单后运行 axe 结构性规则(label/button-name/heading-order/landmark/aria),两页均零违规。摸清并处理了渲染依赖(UserProvider 会话检查、publicConfigQuery 门控、configBool 字符串语义)。vitest 116→118 全绿。","timestamp":1786884595141,"segment":0,"confidence":null,"asi":{"hypothesis":"前端真实 a11y 审计:axe-core(jsdom 结构性规则)覆盖登录/注册页,超出 eslint 静态 jsx-a11y 的动态可访问性验证","insight":"新增 tests/a11y.test.tsx(2 测试)+ axe-core devDependency。调试中摸清登录/注册页渲染依赖链(UserProvider 挂载跳查 getUserInfo、LoginForm/RegisterForm 门控 publicConfigQuery、configBool 期望字符串 'true' 而非布尔 —— mock 需给字符串)。两页均零 axe 违规(color-contrast 因 jsdom 无布局引擎禁用,文档化)。vitest 116→118,checks 全绿。metric 持平 8","next_action_hint":"可扩展 axe 到更多页面(如登录 OTP 态、设置页),或收尾。axe 依赖仅 devDependency,不进基准计数(vitest_failed 已含新测试)"}}
|
||||
{"run":26,"commit":"7d03154","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":120,"measure_s":71},"status":"keep","description":"axe a11y 审计扩展到登录 OTP 验证表单(input-otp 分段输入,FieldLabel htmlFor 正确关联,零违规)与人机验证小部件手动模式(零违规)。环境修复:tests/setup.ts 加 ResizeObserver mock(input-otp 依赖,jsdom 未内置)。vitest 118→120 全绿。","timestamp":1786885103995,"segment":0,"confidence":null,"asi":{"hypothesis":"把 axe a11y 审计扩展到登录/注册外的关键表单状态:登录 OTP 验证(input-otp 分段输入是经典 a11y 雷区)与人机验证小部件(CapWidget 手动模式)","insight":"新增 2 个 axe 测试(vitest 118→120)。OTPForm 验证通过 —— FieldLabel htmlFor='otp-verification' 与 InputOTP id 正确关联,分段输入零违规;CapWidget 手动模式零违规。环境修复:input-otp 依赖 ResizeObserver,jsdom 未内置 → tests/setup.ts 加 ResizeObserver mock(全局,其它测试无害)。渲染上下文修正:裸组件需 <main> landmark 包裹(region 规则,真实页面由 AuthShell 提供)。scripts/ Go 代码用扩展 linter 集扫描为 0 问题(死路,未纳入基准)。metric 持平 8","next_action_hint":"a11y 覆盖已达:登录/注册页/OTP/CapWidget 全零违规。剩余页面(设置页、admin)需较重 mock。会话可收尾,或周期跑全量 -race 验证无 flake"}}
|
||||
{"run":27,"commit":"6c128e0","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":121,"measure_s":71},"status":"keep","description":"axe a11y 审计扩展到最复杂认证路径:注册页开启人机验证(CapWidget 自动求解→已通过状态 + 完整表单),mock getCapToken 避免 jsdom 无 Worker 环境限制。零违规。vitest 120→121 全绿。","timestamp":1786885536618,"segment":0,"confidence":null,"asi":{"hypothesis":"axe 审计覆盖最复杂的认证渲染路径:注册页开启人机验证(CapWidget 自动求解→已通过状态 + 完整表单)","insight":"新增第 5 个 axe 测试(vitest 120→121):mock @/lib/cap-solver 的 getCapToken 直接 resolve token(避免 jsdom 无 Web Worker),config 开 cap_login_enabled/cap_auto_solve,注册页渲染出 CAPTCHA 已通过态 + 表单全字段 → 零违规。vi.mock('@/lib/cap-solver') 对其它测试无害(仅 capEnabled 时渲染 CapWidget)。metric 持平 8","next_action_hint":"axe 覆盖已达 5 个认证表单态(登录/注册/OTP/验证小部件手动/注册+验证)。剩余:设置页与 admin 页需较重 mock。可收尾,或周期跑全量 -race 验证无 flake"}}
|
||||
{"run":28,"commit":"40eee77","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_vetx_total":0,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"vitest_failed":0,"vitest_total":121,"measure_s":0},"status":"keep","description":"前端显式 any 类型清理 2→0:Slot children?: any → ReactNode | MotionValue 联合(motion 真实类型),顺带修复潜在崩溃(原代码在 isValidElement 前访问 children.type,缺失时 TypeError,现无效 children 返回 null,hooks 无条件合规);useControlledState Rest extends any[] → unknown[]。两处 eslint-disable 注释删除。tsc/eslint/vitest 121 全绿。","timestamp":1786886086713,"segment":0,"confidence":null,"asi":{"hypothesis":"前端显式 any 类型清理:全仓 grep 仅 2 处 any —— Slot children?: any 与 useControlledState 的 Rest extends any[],均为真实类型缺陷","insight":"全前端 any 计数 2→0。slot.tsx:children?: any → React.ReactNode | MotionValue<string> | MotionValue<number>(motion HTMLMotionProps 的真实 children 类型);顺带修复潜在崩溃 —— 原代码在 isValidElement 检查前就访问 children.type,children 缺失时 TypeError,改为 isValidChild/childrenType 先计算(hooks 无条件,rules-of-hooks 合规),无效 children 返回 null。use-controlled-state.tsx:Rest extends any[] → unknown[]。两处 eslint-disable no-explicit-any 注释随之删除(无抑制注释)。tsc/eslint/vitest 121/checks.sh 全绿。benchmark 无关(metric 持平 8)。注意:run #28 的 run_experiment 被用户中断(aborted),但代码修复已通过全部门禁验证","next_action_hint":"用户要求合并到 main 并推送"}}
|
||||
{"run":29,"commit":"511bed8","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":63,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":124},"status":"keep","description":"修复 2 个新增 unconvert 问题(linux.go 中 int64(stat.Bsize) 恒等转换,Statfs_t.Bsize 在 Linux 上本就是 int64),删除多余转换零行为变化;total 10→8 回到 5 维全下限。","timestamp":1786894372432,"segment":0,"confidence":null,"asi":{"category":"unconvert","hypothesis":"会话恢复后 measure 显示 total=10,出现 2 个新的 unconvert 问题(internal/apps/edge/observability/linux.go:261-262 的 int64(stat.Bsize) 恒等转换,Linux Statfs_t.Bsize 本就是 int64)。删除多余转换,零行为变化","finding":"unconvert 是 repo 自带配置启用的 linter,此前 baseline 无此问题,最近用户提交/Go 版本变化后新增;修复后 5 维回到全下限 8","next_action_hint":"会话恢复点确认:total=8(5 维全下限,8 项均为有据可查的刻意保留)。下一轮候选:静态检查新维度(staticcheck SA 系列在 repo 配置中已启用且为 0)、或把 docs/ 下 vitepress 站点的构建纳入 measure 防回归(docs build 不属质量计数,不进基准)"}}
|
||||
{"run":30,"commit":"d49c7e1","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":183,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":126},"status":"checks_failed","description":"Agent 发现 Token 比较改为 SHA-256 后恒定时间 Compare,堵住未授权节点注册口的计时侧信道。checks 在 -race 阶段超时(包本身已单独跑绿)。","timestamp":1787667218065,"segment":0,"confidence":null,"asi":{"hypothesis":"discovery token 用 != 比较,未授权 /agent/nodes/register 可被计时;改 SHA-256 + ConstantTimeCompare","rollback_reason":"checks.sh 在 go test -race 阶段 300s 超时(包单独跑全绿,预算不够)","next_action_hint":"同一修复用 checks_timeout_seconds=600 重跑"}}
|
||||
{"run":31,"commit":"69055a9","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":71,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"未授权 Agent 注册口的 discovery token 改为 SHA-256 后恒定时间比较,堵住计时侧信道;空 token / 末字节翻转用例同步补上。metric 持平 8。","timestamp":1787667401636,"segment":0,"confidence":null,"asi":{"hypothesis":"discovery token 用 != 比较,未授权 /agent/nodes/register 可被计时;改 SHA-256 + ConstantTimeCompare","finding":"公开面注册口 ValidateDiscoveryToken 是入侵入口;管理员已登录操作不在范围内。checks 全绿。","next_action_hint":"下一轮可查边缘 Token 比较(agent/relay/flared 走 DB 查找,计时面更弱)或登录口限流"}}
|
||||
{"run":32,"commit":"fb62802","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":81,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"公开登录/注册邮箱验证码比较改为 SHA-256 后恒定时间 Compare,堵住未授权口的计时侧信道。metric 持平 8。","timestamp":1787667660993,"segment":0,"confidence":null,"asi":{"hypothesis":"verifyEmailCode 用 != 比较 6 位码,公开登录/注册口可被计时","finding":"公开面验证码比较已改恒定时间;冷却仍在,不改限流策略。","next_action_hint":"下一轮可查边缘节点 access_token 比较(DB 查找,计时面更弱)或登录失败锁定"}}
|
||||
{"run":33,"commit":"b8bf82b","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":99,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"未授权登录口补哑 bcrypt 比较,用户不存在与密码错误耗时对齐;禁用账号不再返回不同文案,堵住用户枚举。metric 持平 8。","timestamp":1787668237322,"segment":0,"confidence":null,"asi":{"hypothesis":"未授权 /user/login 在用户不存在时跳过 bcrypt,且禁用账号返回不同文案,可枚举用户","finding":"DummyCheckPassword 启动时生成哑哈希,gosec 不报警;禁用账号改统一错误文案。管理员已登录不在范围内。","next_action_hint":"下一轮可查边缘节点 access_token 明文比较,或公开 CAP challenge 滥用"}}
|
||||
{"run":34,"commit":"380a42a","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":90,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"登录/注册/OAuth 回调统一走 SetLoginSession,保存前清空 Redis 会话 ID,堵住未授权会话固定。metric 持平 8。","timestamp":1787669059138,"segment":0,"confidence":null,"asi":{"hypothesis":"生产 Redis 会话在登录时复用同一 ID,未授权方可固定会话 cookie","finding":"SetLoginSession 先 Clear 再把 gorilla session.ID 置空,Save 时 redistore 生成新 ID;明文改密标记经 extras 写回。","next_action_hint":"下一轮可查边缘节点 access_token 明文比较,或公开 CAP challenge 滥用"}}
|
||||
{"run":35,"commit":"dfda2d3","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":75,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"去掉公开 CAP 口硬编码默认密钥;SessionSecret 为空时拒绝签发/核销,防止未授权伪造 PoW。metric 持平 8。","timestamp":1787669542055,"segment":0,"confidence":null,"asi":{"hypothesis":"公开 /api/cap/challenge 在 SessionSecret 为空时用硬编码默认密钥,未授权方可伪造 PoW","finding":"GetDefaultManager 无密钥时返回 nil;Challenge/Redeem 拒绝,VerifyMiddleware 在 CAP 开启时同样拒绝。测试自行设置密钥。","next_action_hint":"下一轮可查公开 OAuth state 洪水或边缘节点 access_token 明文比较"}}
|
||||
{"run":36,"commit":"7fa9e46","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":86,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"注册开关读取失败时改为关闭,堵住配置缺失时未授权开注册;OAuth 自动注册同样 fail-closed。metric 持平 8。","timestamp":1787669960693,"segment":0,"confidence":null,"asi":{"hypothesis":"registration_enabled/password_register_enabled 读取失败默认 true,和种子 false 相反,配置缺失时未授权开注册","finding":"密码注册与 OAuth 自动注册均 fail-closed;测试改为显式开启注册并正确失效缓存。","next_action_hint":"下一轮可查 OIDC 开关 fail-open(种子默认 true,风险较低)或公开 OAuth state 洪水"}}
|
||||
{"run":37,"commit":"0290c93","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":95,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"公开 OAuth 登录/授权入口按会话限制 10 分钟内最多 20 个 state,堵住未授权 Redis 洪水。metric 持平 8。","timestamp":1787670327304,"segment":0,"confidence":null,"asi":{"hypothesis":"公开 /oauth/login 与 /oauth/{source}/authorize 每次请求都往 Redis 写 10 分钟 state,无上限","finding":"按 sessionHash 计数,10 分钟内最多 20 个;超出返回业务错误。mock Redis 补 Incr/Expire。","next_action_hint":"下一轮可查边缘节点 access_token 明文比较,或公开 CAP challenge 洪水"}}
|
||||
{"run":38,"commit":"c0a82f8","metric":8,"metrics":{"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_usestdlibvars":0,"golint_wastedassign":0,"golint_total":8,"eslint_problems":0,"eslint_errors":0,"eslint_warnings":0,"tsc_errors":0,"measure_s":85,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_usetesting":0,"golint_test_total":0,"golint_exhaustive":0,"golint_vetx_total":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"公开密码登录口按 IP 限制 10 分钟内最多 20 次失败,堵住未授权爆破。metric 持平 8。","timestamp":1787670665553,"segment":0,"confidence":null,"asi":{"hypothesis":"公开 /user/login 失败无 IP 限流,未授权方可无限爆破","finding":"按 ClientIP 计数,10 分钟 20 次失败后拒绝;成功清零。管理员已登录不在范围内。","next_action_hint":"下一轮可查公开 CAP challenge 洪水或边缘节点 access_token 明文比较"}}
|
||||
{"run":39,"commit":"be5d067","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":76,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"auth_cache negative 缓存加上限防 DoS + relay/flared 删除重复 authenticateAccessToken 改用 agent 共享缓存版","timestamp":1787708052241,"segment":0,"confidence":null,"asi":{"hypothesis":"negative cache 无上限可被伪造 token 撑爆内存;relay/flared 与 agent 三份重复的 authenticateAccessToken","next_action_hint":"继续扫其他无界缓存/限流缺口","result":"metric 持平 8(8 个均为 deliberate keeper),安全修复不计入 metric","security":"negative cache 加 10k 上限+过期清理;relay/flared 复用 agent.AuthenticateAccessToken(共享 2min 正/10min 负缓存,DB 压力下降)"}}
|
||||
{"run":40,"commit":"0dd2cf9","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":77,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"websocket 三 hub 去重:抽 runWritePump 共享写泵 + 合并 agent 广播函数为 broadcastAgent","timestamp":1787708370650,"segment":0,"confidence":null,"asi":{"hypothesis":"三份 hub 的 writePump 完全重复(仅日志前缀不同),readPump 已有 runReadPump 抽取先例;BroadcastWAFIPGroups/BroadcastActiveConfig 复制粘贴","next_action_hint":"close() 3 份小重复可再合并但收益低;继续找其他模块的重复/无界增长","result":"metric 持平 8,全测试绿","refactor":"新增 websocket/write_pump.go runWritePump(对齐 runReadPump 模式),agent/relay/flared writePump 改委托;agent_hub 抽 broadcastAgent 合并两个广播函数"}}
|
||||
{"run":41,"commit":"efd8268","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":75,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"websocket 三 client 结构体去重:嵌入共享 wsClientCore(close/enqueue 单份实现)","timestamp":1787708975609,"segment":0,"confidence":null,"asi":{"hypothesis":"agentClient/relayClient/flaredClient 字段与 close/enqueue 完全相同,用组合(嵌入 wsClientCore)消除三份重复","next_action_hint":"代码库经 40 轮已高度收敛;后续可周期性跑 go test -race 全量","result":"metric 持平 8,全测试绿;净减 ~60 行重复代码","refactor":"新增 websocket/client_core.go:wsClientCore(nodeID/conn/send/done/once) + 共享 close/enqueue;三个 client 结构体改为嵌入"}}
|
||||
{"run":42,"commit":"ed1efd3","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":77,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"补 wsClientCore 并发测试 + close() 防 nil conn 守卫","timestamp":1787709222794,"segment":0,"confidence":null,"asi":{"hypothesis":"wsClientCore 并发语义(close 幂等、enqueue 不阻塞/关后拒绝)无测试覆盖","next_action_hint":"websocket 包已有基础并发测试;继续其他模块扫描","result":"metric 持平 8;测试还暴露 close 未防 nil conn 的防御缺口,已补守卫","refactor":"新增 websocket/client_core_test.go 3 个 -race 测试;client_core.go close() 增加 nil conn 守卫"}}
|
||||
{"run":43,"commit":"4f8e7e6","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":77,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"修复 frps/frpc TOML 配置注入:新增 protocol.TOMLQuote 并在两处配置渲染全部使用","timestamp":1787709693698,"segment":0,"confidence":null,"asi":{"hypothesis":"frps/frpc TOML 配置用裸 Fprintf 拼接,token/password/域名含引号、反斜杠、换行时会破坏配置或注入键","next_action_hint":"检查其他配置生成点是否有同类注入面(nginx/openresty 配置)","result":"metric 回到 8;frpc 慢套件 16.8s 全绿;mnd 曾短暂+1(Grow 魔法数),删除微优化后消除","security":"新增 pkg/protocol/toml.go TOMLQuote 转义助手 + toml_test.go;relay/frps renderConfig 与 flared/frpc buildFrpcToml 全部插值改为转义输出"}}
|
||||
{"run":44,"commit":"63007fc","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":92,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"全仓 race 扫描发现 upload/cache 监听器 DATA RACE:捕获 redis 客户端消除全局读竞争 + Stop 等待 done + 同型监听器(oauth×2/repository×2)加固","timestamp":1787711092906,"segment":0,"confidence":null,"asi":{"hypothesis":"全仓 go test -race 可能暴露并发 bug(此前仅局部验证)","next_action_hint":"继续扫其他模块;可考虑把 -race 纳入周期性检查","result":"发现并修复 1 个真实 DATA RACE;修复后全仓 -race 0 竞争,metric 持平 8","root_cause":"upload/cache 监听器 goroutine 读可变全局 db.Redis,与 testhelper 清理置 nil 竞争;testhelper 导入 upload/cache 有循环依赖,故用启动时捕获客户端的根因修复(oauth/repository 同型监听器一并加固),并补 StopUploadMetaCacheListener 同步等待 done"}}
|
||||
{"run":45,"commit":"63007fc","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":70,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"discard","description":"探索轮:索引对齐/前端请求瀑布/BasicAuth 注入面三假设均证伪,无代码变更","timestamp":1787711474404,"segment":0,"confidence":null,"asi":{"hypothesis":"SQLite 迁移缺 PG 同款索引;前端存在串行请求瀑布;nginx BasicAuth 密码有注入面","next_action_hint":"代码库已高度收敛;下轮可考虑 observability 查询构造器审计或周期性重跑 -race","rollback_reason":"纯探索无代码变更,无需回滚","result":"三个假设均无产出:①索引对比(修正提取正则后)PG/SQLite 完全对齐,SQLite 仅多 legacy w_* 冗余索引;②前端 await Service 均在事件处理器非渲染期;③BasicAuth 密码经 base64 编码(字母表无元字符)无注入面","lessons":"grep 提取 SQL 时注意 IF NOT EXISTS 变体,否则产生假缺口"}}
|
||||
{"run":46,"commit":"2cb3392","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":106,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"LIKE 过滤器转义修复:日志搜索含 %/_ 的输入不再被当通配符;pkg/util 新增 EscapeLike 共享助手 + 单测","timestamp":1787712116152,"segment":0,"confidence":null,"asi":{"hypothesis":"日志搜索 LIKE 过滤器不转义 %/_/\\,含下划线的路径/主机名搜索结果错误","next_action_hint":"同类遗留站点(upload/user/task_execution GORM 搜索)已记 ideas.md,可作后续轮次","result":"修复 4 个站点:analytics 两处 CH 过滤器 + logstore postgres_store 两处(PG/SQLite 加 ESCAPE '\\')。新增 pkg/util/like.go EscapeLike + 单测。metric 持平 8,全部测试通过","scope_decision":"GORM 实体搜索站(upload keyword、user username/email)同 bug 类但低风险且可能依赖现有通配语义,本轮不动"}}
|
||||
{"run":47,"commit":"3528323","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":102,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"GORM 实体搜索 LIKE 转义收尾:6 站点复用 EscapeLike + 显式 ESCAPE 子句,含 OAuth 用户名冲突误报修复","timestamp":1787712555794,"segment":0,"confidence":null,"asi":{"hypothesis":"GORM 实体搜索站与 #46 日志搜索同 bug 类:LIKE 模式不转义通配符","next_action_hint":"LIKE 类已全部收尾;下轮可考虑 ideas.md 的测试可运行性方向或周期性全仓 -race 重跑","result":"6 站点修复(upload keyword、user username/email 前缀+contains、OAuth uniqueUsername base、task_type 前缀),PG/SQLite 加显式 ESCAPE。系统常量模式刻意保留(upload.go:199 image/%)。metric 持平 8,测试全绿","scope_decision":"uniqueUsername 的 base 来自 OAuth 用户信息属外部输入,含 _ 会误报用户名冲突——虽是系统生成后缀模式也需转义 base 本身"}}
|
||||
{"run":48,"commit":"55db1c0","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":112,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"后台 goroutine panic 防护:新增 pkg/util.Go 共享助手(recover+调用点日志),全仓 22 个裸 go func() 站点统一收口","timestamp":1787713583118,"segment":0,"confidence":null,"asi":{"hypothesis":"全仓 20 处后台 goroutine 裸跑零 recover,任一 panic 击穿 gin handler 级恢复直接崩溃进程","next_action_hint":"goroutine 收口完成;下轮可周期性 go test -race ./... 全量重跑(上次 #44)","result":"pkg/util.Go(fn) 共享助手(runtime.Caller 自动记录调用点 + slog + debug.Stack),22 个站点全部收口(含嵌套 watcher)。脚本转换两轮(首轮漏嵌套内层)。首次 checks_failed 因新文件缺 SPDX 头,update_go_license.sh 修复后全绿。metric 持平 8"}}
|
||||
{"run":49,"commit":"40232d8","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":75,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"修复 frpc restartProcess 发布未初始化 exec.Cmd 的数据竞争:proc.Cmd/Status 改为 Start 成功后加锁发布","timestamp":1787714358791,"segment":0,"confidence":null,"asi":{"hypothesis":"周期性全仓 go test -race ./... 重跑(上次 #44 后又改了 repository/logstore/goroutine 站点)能抓出新数据竞争","next_action_hint":"-race 全仓清零;下轮候选:frontend axe a11y 审计,或 Go 1.26 新 linter 扫描","result":"全仓 -race 抓到 1 个真实 race:frpc/manager.go restartProcess 在 cmd.Start() 前就发布 proc.Cmd+Status=running(Start 中 cmd.Process 未赋值),测试读句柄与之竞争。修复=Start 成功后再加锁发布(manager.go:219-220 移入 err==nil 分支)。frpc 包 -race 连续 3 次通过。其余全仓 -race 干净"}}
|
||||
{"run":50,"commit":"40232d8","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":70,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"discard","description":"扩展 linter 发现扫描 + 热路径性能排查:errchkjson/unparam/spancheck 等 9 个新维度,全部核实为不可失败/刻意设计/误报","timestamp":1787714798689,"segment":0,"confidence":null,"asi":{"hypothesis":"基准外发现型 linter(errchkjson/unparam/spancheck/exptostd/durationcheck/makezero/reassign/asasalint/bidichk)+ 热路径性能 grep 能找到真实缺陷","next_action_hint":"发现型 linter 已穷尽;下轮候选:frontend axe a11y 浏览器级审计,或任务执行日志/DB 增长类运维审查","result":"全部证伪:errchkjson 12 处均核实为不可能失败的 marshal(纯 string/int/[]string 结构体;2 处 unsafe 标记是传递性保守);spancheck 1 处误报(唯一调用方 executor.go:242 有 defer span.End());unparam×2 为已评估的工厂签名设计;正则全在包级编译无热路径重编译;包级 map 全为有界静态注册表;AppendLog 走 DB 无内存累积。escapeJSONString 用法正确。无代码变更"}}
|
||||
{"run":51,"commit":"bbf7919","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":72,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"discard","description":"运行时资源审计:HTTP 客户端超时覆盖 + 查询热路径索引覆盖,两项全部干净无缺陷","timestamp":1787715135724,"segment":0,"confidence":null,"asi":{"hypothesis":"运行时资源审计:出站 HTTP 客户端超时覆盖 + LIKE/精确匹配热路径的 DB 索引支撑","next_action_hint":"两项审计干净。剩余:frontend axe a11y(需起前端+浏览器)、周期性 -race 重跑、uploads LOWER(file_name) contains 若成为性能痛点需改前缀语义+表达式索引","result":"全部干净:15 个 http.Client 中 14 个显式 Timeout,唯一无 Timeout 的 agent/nginx checkStubStatus 走 NewRequestWithContext+WithTimeout 边界;users.username 全部精确匹配热路径由 UNIQUE 内联索引覆盖(PG+SQLite 均确认),email/task_type/logstore 过滤列均已有索引;uploads LOWER(file_name) contains 不可用 b-tree 但属管理端低频,改语义才有收益故不动"}}
|
||||
{"run":52,"commit":"bbf7919","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":71,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"discard","description":"SQL 注入面 + Go 运行时陷阱模式 + react-hooks 依赖三重审计,全部干净无缺陷","timestamp":1787715503278,"segment":0,"confidence":null,"asi":{"hypothesis":"原始 SQL 拼接注入面 + 经典 Go 运行时陷阱(time.After 循环泄漏/defer-in-loop/context.Background 丢失取消)+ 前端 react-hooks 依赖正确性","next_action_hint":"静态+运行时审计维度已穷尽。剩余唯一大项:frontend axe a11y 浏览器级审计(需起前端 dev server + agent_browser)","result":"全部干净:db_manage SQL 控制台为管理端允许例外且表名双引号转义正确、analytics Sprintf 均内部常量表名+参数化占位符;time.After 仅 3 处且均为 select 单次等待/有界重试;defer 均在函数级非循环内;19 处 context.Background() 全部为后台监听器(WithCancel)/重启路径/自带超时的清理任务,无请求 ctx 丢弃;react-hooks/exhaustive-deps 全仓零违规(CLI 临时规则,未改配置)"}}
|
||||
{"run":53,"commit":"bbf7919","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":72,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"discard","description":"前端 axe a11y 浏览器审计:唯一违规为无后端环境产物,无代码缺陷","timestamp":1787716027952,"segment":0,"confidence":null,"asi":{"hypothesis":"前端 axe-core 浏览器级 a11y 审计(最后一个未探索大维度)","next_action_hint":"a11y 维度已探索但受登录墙限制:完整审计需起后端+种子账号登录。若未来重跑:起 Go 后端 + admin 登录后逐页 axe.run","result":"agent-browser 0.34.0 已装好可复用。axe 审计覆盖所有无认证可达页面(/login、/register、/docs/* 全被登录墙拦截):唯一违规 page-has-heading-one 是环境产物——后端未启动时页面卡在 session-check/publicConfig-pending 态只渲染 Spinner,真实表单的 AuthHeading h1 未渲染;瞬态态用 h3 属可接受的瞬态层级。无代码缺陷。已认证页面需后端才能审计"}}
|
||||
{"run":54,"commit":"451ce52","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":93,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"认证页 axe a11y 审计+修复:7 处布局级真实违规全修,复扫验证 dashboard/admin/system 归零;基准 total_issues 保持 8 不变(纯质量收益)","timestamp":1787718397798,"segment":0,"confidence":null,"asi":{"hypothesis":"认证页 axe a11y 审计(起后端+登录突破登录墙):修复布局级真实违规","next_action_hint":"已验证 / 与 /admin/system 归零。剩余页面级:admin 表格行内操作按钮/Switch 无 aria-label、muted 文本对比度——需逐表补标签,工作量大已归档 ideas.md","result":"修复 7 处全局问题并复扫验证:sidebar 折叠按钮 aria-label、Sidebar role=navigation(region 违规 18 节点/页清零)、header Kbd 对比度 text-foreground/70(每页 1 处)、dashboard 4 个 Progress aria-label、分页按钮 aria-label、空态/错误/加载 h3→p(heading-order 清零)、admin/system 无内容 Tabs 改 aria-pressed 按钮组(aria-valid-attr-value critical 清零)。dashboard 与 admin/system 现 0 违规","setup":"审计环境:后端 go run . api @:3100(CONFIG_PATH=/tmp/of-audit/config.yaml,sqlite+redis host 网络 docker)、前端 pnpm dev --port 3002(WAVELET_BACKEND_URL=:3100)、admin 密码经 reset-passwd 重置"}}
|
||||
{"run":55,"commit":"e66dea9","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":85,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"a11y 收尾:主题级对比度根因修复(indigo-500→600)+12 处控件 accessible name+4 处 heading-order,7 页复扫全 0 违规;基准 total_issues 保持 8","timestamp":1787719908229,"segment":0,"confidence":null,"asi":{"hypothesis":"页面级 a11y 批量收尾:主题级 color-contrast 根因 + 表格/表单控件 accessible name","next_action_hint":"7 页复扫全 0 违规。剩余:其余页面(websites/origins/cloudflare 等仅扫过 contrast 已由主题修复覆盖)可抽查;-race 周期重跑","result":"根因1:--primary indigo-500(#6366f1) 对 #fafafa 仅 4.27 → 改 indigo-600 oklch(51.1% 0.262 276.966)(~6.8 AA),全站 contrast 清零(一处主题修复覆盖所有页面)。修复 12 处控件名:access-analytics 刷新按钮、events-tab Switch/edit/delete、openflare-ops ToggleRow Switch+geoip/kuma Select+FieldInput Input htmlFor+discovery Textarea、table-browser/sql-console SelectTrigger;heading-order:cache-manager/user-detail-sheet h4→p、task-manager h3→p、file-manager noFiles h3→p;新增 admin.logs.analytics.refresh i18n 键(en/zh)+merge-i18n-fragments。教训:settings 表单异步渲染,早前扫描漏报 label 违规需 wait 5s 后再 axe.run;Radix SelectValue value='' 时 placeholder 不显示致 combobox 无名,须 aria-label 兜底","setup":"审计环境同 run#54:后端:3100(sqlite) + docker redis host 网络 + pnpm dev --port 3002"}}
|
||||
{"run":56,"commit":"63e3b85","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":85,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"富交互页 a11y 抽查收尾:8+3 页扫描,修复 cloudflare 筛选器无名/access-token amber 对比度/notifications 缺 h1 共 3 处,全部复扫归零;基准 total_issues 保持 8","timestamp":1787720716912,"segment":0,"confidence":null,"asi":{"hypothesis":"富交互页抽查(websites/origins/proxy-routes/certificates/cloudflare/dns-accounts/settings 子页)","next_action_hint":"11 页扫描全部归零,a11y 维度已穷尽。剩余:周期性 -race 重跑;审计环境复用法在 ideas.md","result":"websites/origins/proxy-routes/certificates/dns-accounts 5 页直接 0 违规(主题修复覆盖);3 处新发现全修复并复扫验证:cloudflare 同步面板状态筛选 SelectTrigger 加 aria-label(statusPlaceholder);access-token 安全提示 amber-600→amber-700(12px 小字对比度 4.5 不达标);notifications 面包屑页加 sr-only h1——教训:h1 不能放 BreadcrumbList 内(破坏 list 语义 axe list 规则),BreadcrumbPage 无 asChild 需放 Breadcrumb 外","setup":"审计环境同前:后端:3100 + docker redis host 网络 + pnpm dev --port 3002"}}
|
||||
{"run":57,"commit":"453f7e5","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":95,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"周期性 -race 重跑抓到真实 bug:wsClientCore.enqueue close 后 select 随机选择致契约违反;确定性先查 done 修复+测试循环加固+gofmt 存量漂移清理","timestamp":1787721299485,"segment":0,"confidence":null,"asi":{"hypothesis":"周期性全仓 -race 重跑(上次干净为 run #49)","next_action_hint":"websocket 包 -race 10×count=1 全过。教训已记录:select 多 case 同时就绪时随机选择,closed 检查须独立 select 先行;replace 工具锚点选错会级联破坏文件,小文件直接 write 重写更安全","root_cause":"enqueue 把 closed 检查与发送合并在同一个 select,两 case 同时就绪时 Go 随机选择,close 后约 50% 概率仍投递成功——违反 fail-fast 契约且测试 flaky。修复=独立 select 确定性先查 done;测试加固为循环 50 次","result":"抓到真实 bug:wsClientCore.enqueue close 后非确定返回 true(TestWSClientCoreEnqueueFailsAfterClose 必失败)。调用方 agent_hub×3 语义无影响(false=丢弃本就正确)。顺带修 3 个 hub 文件存量 gofmt 漂移","scope_note":"-race 重跑仅 websocket 包 1 个 FAIL,其余 internal/... pkg/... 全部通过"}}
|
||||
{"run":58,"commit":"fc733d0","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":77,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"#57 enqueue 修复的同型残留收口:SendFlaredPong/SendRelayPong 合并 select 随机选择 bug,委托 client.enqueue 去重修复","timestamp":1787721635717,"segment":0,"confidence":null,"asi":{"hypothesis":"#57 修复 enqueue 后,grep 全 hub 同型合并 select——发现 SendFlaredPong/SendRelayPong 残留相同 bug","lesson":"修一个 bug 后应 grep 所有同型调用点(本会话 run #44/#46/#57 三次都是同型残留收口模式);委托共享 enqueue 是去重+根因一步到位","next_action_hint":"websocket 并发面已全清。下轮可做:周期性全仓 -race 或 go test -count=10 稳定性抽查","root_cause":"SendFlaredPong (flared_hub.go) 与 SendRelayPong (relay_hub.go) 把 case <-client.done 与 case client.send <- 合并同一 select,两 case 同时就绪时 Go 随机选择,close 后仍可能投递成功。修复=委托 client.enqueue(内含确定性先查 done),同时消除重复代码"}}
|
||||
{"run":59,"commit":"b56f276","metric":8,"metrics":{"eslint_errors":0,"eslint_problems":0,"eslint_warnings":0,"golint_canonicalheader":0,"golint_errname":0,"golint_errorlint":1,"golint_exhaustive":0,"golint_forcetypeassert":0,"golint_gosec":0,"golint_intrange":0,"golint_modernize":3,"golint_nilnil":3,"golint_perfsprint":0,"golint_prealloc":0,"golint_recvcheck":1,"golint_test_testifylint":0,"golint_test_thelper":0,"golint_test_total":0,"golint_test_usetesting":0,"golint_total":8,"golint_usestdlibvars":0,"golint_vetx_total":0,"golint_wastedassign":0,"measure_s":66,"tsc_errors":0,"vitest_failed":0,"vitest_total":126},"status":"keep","description":"#59 -shuffle=on 扫描抓到测试顺序依赖:config_version RAM 配置缓存跨测试污染,setup/cleanup 接入 ram.ResetForTest() 修复","timestamp":1787722520315,"segment":0,"confidence":null,"asi":{"hypothesis":"-shuffle=on 测试顺序随机化扫描(未查过的维度),暴露测试间共享状态依赖","lesson":"repository 读配置会写进程级 RAM 缓存(ram.Set,TTL 跨测试存活);测试用 :memory: DB + SetDB 换库时缓存不随之失效。默认源码顺序下 Defaults 先跑掩盖了问题。-shuffle=on 是暴露此类顺序依赖的低成本手段,可周期重跑","next_action_hint":"全仓 shuffle 已干净。下轮候选:-count 多轮稳定性、或从 ideas.md 剩余条目挑;明确不做清单见 ideas.md","root_cause":"TestBuildOpenRestyConfigSnapshotOriginErrorPageDefaults 在 shuffle 下命中 Custom 用例留在进程级 RAM 配置缓存的 enabled=false/[\"522\",\"500-502\"](GetSystemConfigByGroup 未命中时 ram.Set 回填)。修复=两个测试 setup(setupOriginErrorPageSnapshotDB/setupConfigVersionTestDB)接入既有 ram.ResetForTest():换 DB 前后各清一次"}}
|
||||
Executable
+90
@@ -0,0 +1,90 @@
|
||||
#!/bin/bash
|
||||
# Benchmark: total code-quality issues across backend + frontend (lower is better).
|
||||
# Fixed linter set — see .auto/prompt.md. Never tune this file to game counts.
|
||||
set -euo pipefail
|
||||
cd "$(dirname "$0")/.."
|
||||
start=$(date +%s)
|
||||
|
||||
# ---------- Backend: golangci-lint, repo config + fixed best-practice extras ----------
|
||||
EXTRA_LINTERS="errorlint,errname,nilnil,forcetypeassert,copyloopvar,intrange,mirror,perfsprint,prealloc,usestdlibvars,modernize,sloglint,canonicalheader,nosprintfhostport,recvcheck,wastedassign,exhaustive"
|
||||
golang_out=$(golangci-lint run --enable="$EXTRA_LINTERS" 2>&1 || true)
|
||||
|
||||
golang_total=0
|
||||
while IFS= read -r line; do
|
||||
if [[ "$line" =~ ^\*\ ([a-zA-Z0-9_]+):\ ([0-9]+)$ ]]; then
|
||||
name="${BASH_REMATCH[1]}"
|
||||
n="${BASH_REMATCH[2]}"
|
||||
golang_total=$((golang_total + n))
|
||||
echo "METRIC golint_${name}=$n"
|
||||
fi
|
||||
done <<< "$golang_out"
|
||||
echo "METRIC golint_total=$golang_total"
|
||||
|
||||
# ---------- Backend: test-code quality (tests excluded from repo config; safe linters only) ----------
|
||||
test_out=$(golangci-lint run --tests=true --enable=testifylint,usetesting,thelper --enable-only=testifylint,usetesting,thelper 2>&1 || true)
|
||||
golang_test_total=0
|
||||
while IFS= read -r line; do
|
||||
if [[ "$line" =~ ^\*\ ([a-zA-Z0-9_]+):\ ([0-9]+)$ ]]; then
|
||||
name="${BASH_REMATCH[1]}"
|
||||
n="${BASH_REMATCH[2]}"
|
||||
golang_test_total=$((golang_test_total + n))
|
||||
echo "METRIC golint_test_${name}=$n"
|
||||
fi
|
||||
done <<< "$test_out"
|
||||
echo "METRIC golint_test_total=$golang_test_total"
|
||||
|
||||
# ---------- Backend: govet extra analyzers (dead code / nil deref — real-bug finders) ----------
|
||||
cat > /tmp/govetx.yml <<'EOF'
|
||||
version: "2"
|
||||
linters:
|
||||
default: none
|
||||
enable:
|
||||
- govet
|
||||
settings:
|
||||
govet:
|
||||
enable:
|
||||
- nilness
|
||||
- unusedwrite
|
||||
EOF
|
||||
vetx_out=$(golangci-lint run --config /tmp/govetx.yml --max-issues-per-linter=0 2>&1 || true)
|
||||
rm -f /tmp/govetx.yml
|
||||
golang_vetx_total=0
|
||||
while IFS= read -r line; do
|
||||
if [[ "$line" =~ ^\*\ ([a-zA-Z0-9_]+):\ ([0-9]+)$ ]]; then
|
||||
name="${BASH_REMATCH[1]}"
|
||||
n="${BASH_REMATCH[2]}"
|
||||
golang_vetx_total=$((golang_vetx_total + n))
|
||||
echo "METRIC golint_vetx_${name}=$n"
|
||||
fi
|
||||
done <<< "$vetx_out"
|
||||
echo "METRIC golint_vetx_total=$golang_vetx_total"
|
||||
|
||||
# ---------- Frontend: eslint (repo gate) ----------
|
||||
cd frontend
|
||||
eslint_out=$(pnpm exec eslint . --max-warnings 0 2>&1 || true)
|
||||
eslint_problems=0; eslint_errors=0; eslint_warnings=0
|
||||
if [[ "$eslint_out" =~ ([0-9]+)\ problems? ]]; then eslint_problems="${BASH_REMATCH[1]}"; fi
|
||||
if [[ "$eslint_out" =~ \(([0-9]+)\ errors?, ]]; then eslint_errors="${BASH_REMATCH[1]}"; fi
|
||||
if [[ "$eslint_out" =~ ,\ ([0-9]+)\ warnings? ]]; then eslint_warnings="${BASH_REMATCH[1]}"; fi
|
||||
echo "METRIC eslint_problems=$eslint_problems"
|
||||
echo "METRIC eslint_errors=$eslint_errors"
|
||||
echo "METRIC eslint_warnings=$eslint_warnings"
|
||||
|
||||
# ---------- Frontend: tsc (repo gate) ----------
|
||||
tsc_out=$(pnpm exec tsc --noEmit --jsx preserve 2>&1 || true)
|
||||
tsc_errors=$(grep -cE "error TS" <<< "$tsc_out" || true)
|
||||
echo "METRIC tsc_errors=$tsc_errors"
|
||||
|
||||
# ---------- Frontend: vitest (2026-08-16 起全绿,纳入基准防回归) ----------
|
||||
vitest_out=$(pnpm exec vitest run --reporter=dot 2>&1 || true)
|
||||
vitest_failed=0; vitest_total=0
|
||||
if [[ "$vitest_out" =~ ([0-9]+)\ failed ]]; then vitest_failed="${BASH_REMATCH[1]}"; fi
|
||||
if [[ "$vitest_out" =~ Tests[[:space:]]+([0-9]+)\ passed ]]; then vitest_total="${BASH_REMATCH[1]}"; fi
|
||||
if [[ "$vitest_out" =~ Tests[[:space:]]+([0-9]+) ]]; then vitest_total="${BASH_REMATCH[1]}"; fi
|
||||
echo "METRIC vitest_failed=$vitest_failed"
|
||||
echo "METRIC vitest_total=$vitest_total"
|
||||
|
||||
end=$(date +%s)
|
||||
total=$((golang_total + golang_test_total + golang_vetx_total + eslint_problems + tsc_errors + vitest_failed))
|
||||
echo "METRIC total_issues=$total"
|
||||
echo "METRIC measure_s=$((end - start))"
|
||||
+169
@@ -0,0 +1,169 @@
|
||||
# Autoresearch: 前后端代码质量符合最佳代码实践
|
||||
|
||||
## Objective
|
||||
|
||||
Improve backend (Go) and frontend (Next.js/TS) code quality so the codebase
|
||||
conforms to best practices. NOT a performance task. Each experiment is a code
|
||||
change that removes real, lint-diagnosed code-quality issues (dead assignments,
|
||||
error-wrapping bugs, non-idiomatic loops, mixed receivers, unsafe error
|
||||
comparisons, unnecessary string fmt, etc.) without changing behavior.
|
||||
|
||||
Genuine quality work only: fix code, never weaken the checks. Do NOT edit
|
||||
`.golangci.yml`, eslint/biome config, or add `nolint`/`eslint-disable`
|
||||
comments to reduce counts. Do NOT reformat code that isn't part of a fix
|
||||
(no formatted-only churn).
|
||||
|
||||
## Metrics
|
||||
|
||||
- **Primary**: `total_issues` (unitless, lower is better) = backend golangci
|
||||
issues (extended linter set below) + frontend eslint problems + tsc errors.
|
||||
- **Secondary**: per-linter counts (`golint_modernize`, `golint_perfsprint`,
|
||||
`golint_errorlint`, `golint_gosec`, `golint_canonicalheader`,
|
||||
`golint_recvcheck`, `golint_wastedassign`, `golint_usestdlibvars`,
|
||||
`golint_intrange`, `golint_forcetypeassert`, `golint_nilnil`,
|
||||
`golint_prealloc`, `golint_errname`, `golint_sloglint`,
|
||||
`golint_copyloopvar`, `golint_mirror`, `golint_nosprintfhostport`),
|
||||
`eslint_problems`, `eslint_errors`, `eslint_warnings`, `tsc_errors`,
|
||||
`measure_s` (benchmark wall time).
|
||||
|
||||
## How to Run
|
||||
|
||||
`./.auto/measure.sh` — outputs `METRIC name=value` lines. Parsed by
|
||||
run_experiment automatically.
|
||||
|
||||
Correctness gate: `./.auto/checks.sh` runs `go vet ./...`, `go build ./...`,
|
||||
and the repo's own `golangci-lint run` (repo config, tests excluded) — all
|
||||
must pass. Note: `go test ./...` is NOT in checks.sh — several tests fail on
|
||||
main today for environmental reasons (no local redis; flaky frpc process
|
||||
tests). Don't "fix" those unless cheap and clearly unrelated to redis/flaky.
|
||||
|
||||
## Benchmark Definition (fixed — never change mid-session)
|
||||
|
||||
Backend: `golangci-lint run --enable=errorlint,errname,nilnil,forcetypeassert,
|
||||
copyloopvar,intrange,mirror,perfsprint,prealloc,usestdlibvars,modernize,
|
||||
sloglint,canonicalheader,nosprintfhostport,recvcheck,wastedassign`
|
||||
(repo `.golangci.yml` linters stay active too; `tests: false` as configured).
|
||||
|
||||
Frontend: `pnpm exec eslint . --max-warnings 0` (repo gate) +
|
||||
`pnpm exec tsc --noEmit --jsx preserve` (repo gate).
|
||||
|
||||
Test-code dimension (added 2026-08-16, run #12+, documented scope extension —
|
||||
raising the bar, not gaming): `golangci-lint run --tests=true
|
||||
--enable=testifylint,usetesting,thelper --enable-only=testifylint,usetesting,thelper`
|
||||
counts test-file quality. DELIBERATELY excludes paralleltest/tparallel
|
||||
(t.Parallel advice is unsafe here: many suites share DB/redis state and tests
|
||||
cannot be run in this env) and gocritic extras (noise). Fix test issues only
|
||||
when compile-safe (go vet compiles tests) and semantically neutral.
|
||||
|
||||
Frontend vitest dimension (added run #19, after suite went green in run #18):
|
||||
`pnpm exec vitest run --reporter=dot` — `vitest_failed` counts into total.
|
||||
The suite is fully runnable locally (jsdom + mocks; no external services).
|
||||
Do not add/remove linters or change settings to make the number go down.
|
||||
|
||||
## Files in Scope
|
||||
|
||||
Backend (Go): `cmd/`, `internal/`, `pkg/`. Anything lint-flagged in the
|
||||
extended set above. Note: module name in go.mod is `github.com/Rain-kl/Wavelet`.
|
||||
|
||||
Frontend (TS/React): `frontend/app/`, `frontend/components/`, `frontend/lib/`,
|
||||
`frontend/contexts/`, `frontend/hooks/`, `frontend/types/`, frontend scripts.
|
||||
|
||||
Infra: `frontend/pnpm-workspace.yaml` — approved @parcel/watcher + @swc/core
|
||||
builds (fixes `make code-check` under pnpm 11; ERR_PNPM_IGNORED_BUILDS
|
||||
otherwise). Already committed in setup.
|
||||
|
||||
## Off Limits
|
||||
|
||||
- `.golangci.yml`, `eslint.config.mjs`, `biome.json` — never touch to reduce counts.
|
||||
- No `//nolint` / `eslint-disable` comments to silence checks.
|
||||
- No reformat-only commits (biome/gofmt churn without a fix).
|
||||
- No behavior changes: refactors must compile (checks.sh gate) and keep tests
|
||||
semantics identical. Re-run checks.sh after every edit.
|
||||
- `frontend/node_modules`, `frontend/bun.lock` (untracked, not ours).
|
||||
- Do not run `go test` suites that need redis/network to declare success.
|
||||
|
||||
## Constraints
|
||||
|
||||
- Backend conventions (AGENTS.md): apps → repository → model layering;
|
||||
`pkg/util/` must not import Gin/GORM/sessions; no `db.DB` in model;
|
||||
response.Abort* for API errors; Chinese docs for content changes
|
||||
(code-quality fixes are not content changes — no doc sync needed unless
|
||||
behavior/UX changes; changelog only for user-visible changes, typically
|
||||
none here).
|
||||
- Frontend: run `pnpm exec biome format --write` only on files you edit
|
||||
(repo `make format` uses biome); keep component placement rules.
|
||||
- `golangci-lint --fix` is allowed and preferred for safe fixes
|
||||
(modernize/intrange/perfsprint/usestdlibvars/canonicalheader/mirror/
|
||||
copyloopvar/sloglint/errname) — review the resulting diff before keeping.
|
||||
For no-fix linters (errorlint wrapping, wastedassign, recvcheck, nilnil,
|
||||
prealloc, forcetypeassert) edit by hand.
|
||||
|
||||
## Workflow per iteration
|
||||
|
||||
1. Read current measure output: which categories remain, where.
|
||||
2. Pick ONE category (or a coherent set of similar fixes), locate files, fix
|
||||
by hand or with golangci-lint --fix scoped to that category.
|
||||
3. `./.auto/measure.sh` → if total dropped → `./.auto/checks.sh` → log keep.
|
||||
If flat/worse → discard or adjust.
|
||||
|
||||
## What's Been Tried
|
||||
|
||||
- Setup commit `ee6974d` (autoresearch/code-quality-2026-08-16): branch,
|
||||
.auto/ session files, frontend/pnpm-workspace.yaml build approvals.
|
||||
- Baseline (before any code fix): total_issues = 108
|
||||
(golangci 107 = modernize 37, perfsprint 18, errorlint 12, canonicalheader 8,
|
||||
recvcheck 7, wastedassign 7, usestdlibvars 3, intrange 3, forcetypeassert 3,
|
||||
nilnil 3, prealloc 3, errname 1, gosec 2; eslint 1 warning
|
||||
[react-hooks/exhaustive-deps in
|
||||
app/(main)/pages/detail/components/pages-source-card.tsx:275]; tsc 0).
|
||||
- Environment notes: golangci-lint 2.12.2 warm cache ~3s; eslint cold ~27s
|
||||
(ignore stderr pnpm noise); go vet+go build ~15-30s after edits.
|
||||
|
||||
### 最终状态(run #23,提交 aa4fadda,本会话收敛点)
|
||||
|
||||
基准 5 维全下限 total=8(全为刻意保留);后端 94 包 + 前端 vitest 116 全绿;
|
||||
`go test -race ./internal/... ./pkg/...` 93 包零警告;`make build-embedded`
|
||||
(发布路径)成功且工作树干净;`make license-check` / `go mod tidy -diff` /
|
||||
`go test -count=3`(时序敏感包)全部通过。checks.sh 门禁:vet + build +
|
||||
golangci + 单测 + vitest + 并发包 -race + license-check。
|
||||
|
||||
### Session result (14 experiments, commits f1f6bb85→65c02ef7)
|
||||
|
||||
108 → **8** (-92.6%) across 3 benchmark dimensions, all remaining 8 are
|
||||
deliberate, documented keepers (see below). Never weakened a check; never
|
||||
added nolint/eslint-disable; benchmark extensions were transparently
|
||||
documented (test-code dimension run #12, exhaustive run #14).
|
||||
|
||||
Fixed (zero behavior change, each reviewed):
|
||||
- gosec 2→0 (saturating multiply pattern gosec accepts without nolint)
|
||||
- modernize 37→5→3 (any, max/min, slices/maps, strings.Cut/SplitSeq,
|
||||
strings.Builder; omitted omitted-lark: nested struct omitzero = wire change)
|
||||
- perfsprint 18→0, canonicalheader 8→0, usestdlibvars 3→0, intrange 3→0,
|
||||
wastedassign 7→0, errname 1→0, forcetypeassert 6→0, prealloc 2→0
|
||||
- errorlint 12→1 (errors.Is/As, %v→%w chains)
|
||||
- recvcheck 7→1 (GORM TableName → pointer receiver; verified gorm source uses
|
||||
reflect.New, tests pass)
|
||||
- eslint 1→0 (exhaustive-deps: add stable `t` to dep array)
|
||||
- test dimension 25→0 (testifylint 20, thelper 3, usetesting 2)
|
||||
- exhaustive 12→0 (explicit enum cases = fail-explicit)
|
||||
|
||||
Deliberate keepers (8) — do NOT "fix" without new evidence:
|
||||
- errorlint 1: pkg/push/telegram.go %v — wrapping the original error would
|
||||
change errors.Is matching semantics; it's intentionally textual context.
|
||||
- modernize 3: nested-struct omitempty (client.go Release/Asset,
|
||||
lark.go Content) — omitzero would CHANGE wire output (plain structs
|
||||
serialize always today).
|
||||
- nilnil 3: not-found/optional-result conventions — postgres_store.go
|
||||
ClickHouseOperationalStats (interface contract, documented in comment),
|
||||
openflare_apply_log.go GetLatestOpenFlareApplyLogByNodeID (tested),
|
||||
github_source_action.go guarded outcome (callers check != nil).
|
||||
- recvcheck 1: MillisecondDuration — encoding/json requires Marshal value
|
||||
receiver + Unmarshal pointer receiver.
|
||||
|
||||
Surveyed and rejected (noise/risk, do not add):
|
||||
- fieldalignment (~100+): JSON key order change + positional literal risk.
|
||||
- sloglint full / gocritic extras: 0 findings.
|
||||
- paralleltest/tparallel: t.Parallel advice unsafe (shared DB/redis state;
|
||||
tests not runnable in this env).
|
||||
- biome format drift (76 files): pure formatting noise; repo's make format
|
||||
covers it.
|
||||
@@ -1,19 +0,0 @@
|
||||
root = true
|
||||
|
||||
[*]
|
||||
indent_style = space
|
||||
indent_size = 4
|
||||
charset = utf-8
|
||||
end_of_line = lf
|
||||
trim_trailing_whitespace = true
|
||||
insert_final_newline = true
|
||||
|
||||
[*.{json,yml,yaml}]
|
||||
indent_size = 2
|
||||
|
||||
[*.md]
|
||||
insert_final_newline = false
|
||||
trim_trailing_whitespace = false
|
||||
|
||||
[*.{js,ts,css,html,jsx,tsx,vue}]
|
||||
indent_size = 2
|
||||
@@ -82,3 +82,8 @@ profile.cov
|
||||
|
||||
/.superpowers/
|
||||
/.worktrees/
|
||||
/.pi-subagents/
|
||||
|
||||
# i18n 生成物(由 scripts/merge-i18n-fragments.mjs 从 fragments 生成)
|
||||
frontend/messages/zh-CN.json
|
||||
frontend/messages/en.json
|
||||
|
||||
@@ -72,6 +72,7 @@ Strong success criteria let you loop independently. Weak criteria ("make it work
|
||||
| `new-async-task` | Asynq 任务、定时任务、TaskHandler、任务元数据 |
|
||||
| `new-setting` | 系统/业务/公开设置、`/admin/system`、`/admin/settings` |
|
||||
| `database-migration` | 表结构、goose 迁移(PG/SQLite/ClickHouse)、seed |
|
||||
| `logstore` | 日志/分析用途表、`internal/repository/logstore`、切换日志主库、PG/SQLite 回落 |
|
||||
| `clickhouse-batchwriter` | CH 批量写入、batchwriter、分析表 flush/背压 |
|
||||
| `file-upload` | 上传/摄取、`upload.Ingest`、文件访问、`w_uploads` |
|
||||
| `cache-framework` | 业务缓存(RAM/Redis/DB)、失效、多节点同步 |
|
||||
@@ -91,6 +92,7 @@ Strong success criteria let you loop independently. Weak criteria ("make it work
|
||||
- **分层**:`apps → repository → model`,`repository → infra/persistence`;禁止 `model → repository`。
|
||||
- `model`:实体、表名、配置 key、查询 DTO、无 IO 规则。禁止 `db.DB` / Redis / CH;禁止 `import repository`。GORM hook 仅可 mutate 自身字段,禁止在 hook 内再查 DB/缓存。
|
||||
- `repository`:唯一持久化入口。apps/logics 禁止为业务 CRUD 直调 `db.DB`(管理端 SQL 控制台、infra 内部等例外保留)。禁止新增 `model.Get/List/Create/...` 类数据访问 API。
|
||||
- 日志/分析表(节点访问日志、用户访问日志、可观测时序)走 `internal/repository/logstore`,禁止 apps 直连 `repository/analytics` 或 `db.ChConn`/`db.ChDB`。判定与接入步骤见 `logstore` skill。
|
||||
- 跨模块集成(任务 Handler、推送事件、域监听、完成钩子)禁止 `init()` 注册;经 `internal/platform/bootstrap` 在 `internal/cmd` 入口显式装配。
|
||||
- 核心业务(如 `oauth`、`user`)禁止直接 import push/custom_events;经 `internal/listener` 发域事件,push 在 bootstrap 订阅。
|
||||
- 依赖任务/推送注册的测试须显式 `bootstrap.RegisterTasks()` / `RegisterPushDomainEvents()` 等,不依赖 `init()`。
|
||||
@@ -226,3 +228,10 @@ frontend/lib/services/<name>/
|
||||
|
||||
- 继承 `BaseService`,定义 `basePath`,有类型静态方法;在 `frontend/lib/services/index.ts` 注册。
|
||||
- 回调/`mutationFn`/`queryFn` **禁止**直接传静态方法引用(丢 `this`);用箭头:`(p) => XxxService.create(p)`。
|
||||
|
||||
### 国际化 (i18n)
|
||||
|
||||
- 使用 `next-intl`(无 URL locale 前缀 / provider 模式),兼容 `NEXT_STANDALONE_EXPORT`。
|
||||
- 语言:`zh-CN`、`en`;默认 `zh-CN`。优先级:cookie `NEXT_LOCALE` → 浏览器语言 → 默认。
|
||||
- 文案放在 `frontend/messages/fragments`。参考已有代码,按模块拆文件夹,en.json 和 zh-CN.json 是 ci 生成的(node scripts/merge-i18n-fragments.mjs),禁止手动修改。
|
||||
- 禁止在页面/组件里直接写文案,文案必须支持 i18
|
||||
|
||||
@@ -45,7 +45,7 @@ code-check:
|
||||
exit 1; \
|
||||
fi
|
||||
golangci-lint run
|
||||
cd frontend && pnpm tsc --noEmit --jsx preserve && npx eslint . --max-warnings 0
|
||||
cd frontend && node scripts/merge-i18n-fragments.mjs && pnpm tsc --noEmit --jsx preserve && npx eslint . --max-warnings 0
|
||||
|
||||
build-backend:
|
||||
@echo "==> Building backend version=$(VERSION) build_date=$(BUILD_DATE)..."
|
||||
|
||||
-225
@@ -1,225 +0,0 @@
|
||||
<div align="center">
|
||||
|
||||
# OpenFlare
|
||||
|
||||
**[English](./README.en.md) | [📖 中文](./README.md)**
|
||||
|
||||
OpenFlare is an open-source CDN orchestration and edge security platform. It supports reverse proxies, centralized configuration synchronization, secure intranet penetration (Tunnels), dynamic WAF protection, and anti-CC challenges.
|
||||
|
||||
</div>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/LICENSE">
|
||||
<img src="https://img.shields.io/github/license/Rain-kl/OpenFlare?color=brightgreen" alt="license">
|
||||
</a>
|
||||
<a href="https://github.com/Rain-kl/OpenFlare/releases/latest">
|
||||
<img src="https://img.shields.io/github/v/release/Rain-kl/OpenFlare?color=brightgreen&include_prereleases" alt="release">
|
||||
</a>
|
||||
<a href="https://github.com/Rain-kl/OpenFlare/pkgs/container/openflare">
|
||||
<img src="https://img.shields.io/badge/GHCR-ghcr.io%2Frain--kl%2Fopenflare-brightgreen" alt="ghcr">
|
||||
</a>
|
||||
</p>
|
||||
|
||||
> [!WARNING]
|
||||
> After logging in for the first time with the `root` user, make sure to change the default password `123456`.
|
||||
>
|
||||
> The BETA version is a temporary product for the development and testing phase. It may contain unknown issues and should not be used in production environments.
|
||||
|
||||
## Documentation
|
||||
|
||||
**https://open-flare.pages.dev**
|
||||
|
||||
Quick links:
|
||||
|
||||
* [Quick Start](https://open-flare.pages.dev/en/guide/quick-start)
|
||||
* [Deployment Guide](https://open-flare.pages.dev/en/deployment/deployment)
|
||||
* [Configuration Reference](https://open-flare.pages.dev/reference/configuration)
|
||||
* [System Design](https://open-flare.pages.dev/design/)
|
||||
|
||||
## Core Features
|
||||
|
||||
* **Reverse Proxy Management**: Website rules as the aggregation boundary, supporting multi-domain binding and multi-upstream load balancing with unified management of all OpenResty node configurations.
|
||||
* **Immutable Config Version Control**: Full-snapshot publish model based on version numbers (`YYYYMMDD-NNN`), with pre-publish diff preview, a single globally active version, and one-click sub-second rollback.
|
||||
* **Secure Intranet Penetration (Tunnels)**: An open-source alternative to Cloudflare Tunnels. Securely expose local intranet Web services to the public network via Relay and OpenFlared clients — no public IP or open inbound ports required.
|
||||
* **Edge WAF Safety Protection**: Provides global and custom rule groups, supporting manual/automatic/subscription IP groups, MaxMind GeoIP country-level access control, Checksum-based differential IP group sync (no Nginx reload), and custom block responses.
|
||||
* **Anti-CC & Human-Machine Challenge (PoW)**: Built-in high-performance client-side cryptographic Proof of Work challenges (similar to Turnstile) to block and intercept botnets and scrapers at the gateway edge in seconds.
|
||||
* **Pages Static Hosting**: Upload pre-built ZIP packages directly; edge Agents pull and serve them via local OpenResty, with SPA Fallback and built-in API reverse proxy configuration.
|
||||
* **Automated TLS Certificate Management**: Supports dynamic certificate upload, automatic multi-domain certificate matching and binding, and ACME-based automatic issuance and renewal via Let's Encrypt.
|
||||
* **Uptime Kuma Monitoring Sync**: Integrates with Uptime Kuma to automatically sync the monitoring site list using differential updates, providing real-time awareness of node availability and service health.
|
||||
* **SSO Single Sign-On**: Supports GitHub OAuth and standard OIDC protocol for seamless integration with enterprise identity providers.
|
||||
* **Unified Observability**: Aggregates node request metrics, real-time access log details, host/Nginx resource snapshots, health events, and a re-upload buffer for network fluctuations.
|
||||
|
||||
## Quick Start
|
||||
|
||||
### 1. Launch Server
|
||||
|
||||
```yaml
|
||||
services:
|
||||
openflare:
|
||||
image: ghcr.io/rain-kl/openflare:latest
|
||||
restart: unless-stopped
|
||||
env_file: .env
|
||||
environment:
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
ports:
|
||||
- "3000:3000"
|
||||
volumes:
|
||||
- openflare_uploads:/app/uploads
|
||||
depends_on:
|
||||
postgres:
|
||||
condition: service_healthy
|
||||
redis:
|
||||
condition: service_healthy
|
||||
clickhouse:
|
||||
condition: service_healthy
|
||||
|
||||
postgres:
|
||||
image: postgres:17-alpine
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
POSTGRES_DB: ${DB_NAME:-openflare}
|
||||
POSTGRES_USER: ${DB_USERNAME:-openflare}
|
||||
POSTGRES_PASSWORD: ${DB_PASSWORD:-replace-with-strong-password}
|
||||
volumes:
|
||||
- openflare_postgres_data:/var/lib/postgresql/data
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U ${DB_USERNAME:-openflare} -d ${DB_NAME:-openflare}"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
|
||||
redis:
|
||||
image: valkey/valkey:8.0-alpine
|
||||
restart: unless-stopped
|
||||
command: ["valkey-server", "--appendonly", "yes"]
|
||||
volumes:
|
||||
- openflare_redis_data:/data
|
||||
healthcheck:
|
||||
test: ["CMD", "valkey-cli", "ping"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
start_period: 5s
|
||||
|
||||
clickhouse:
|
||||
image: clickhouse/clickhouse-server:25.3-alpine
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
CLICKHOUSE_DB: ${CLICKHOUSE_NAME:-openflare}
|
||||
CLICKHOUSE_USER: ${CLICKHOUSE_USERNAME:-default}
|
||||
CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}
|
||||
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
volumes:
|
||||
- openflare_clickhouse_data:/var/lib/clickhouse
|
||||
healthcheck:
|
||||
test: ["CMD", "clickhouse-client", "--user", "${CLICKHOUSE_USERNAME:-default}", "--password", "${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}", "--query", "SELECT 1"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
start_period: 15s
|
||||
|
||||
volumes:
|
||||
openflare_uploads:
|
||||
openflare_postgres_data:
|
||||
openflare_redis_data:
|
||||
openflare_clickhouse_data:
|
||||
```
|
||||
|
||||
```bash
|
||||
docker compose up -d
|
||||
```
|
||||
|
||||
Access at: `http://localhost:3000`
|
||||
|
||||
Default credentials:
|
||||
|
||||
* Username: `root`
|
||||
* Password: `123456`
|
||||
|
||||
### 2. Install Agent
|
||||
|
||||
Before installing an Agent, please install OpenResty on the target node first, or use the Agent Docker image with OpenResty built-in.
|
||||
|
||||
You can copy the installation command from **Node Management -> Details -> Node Info -> Node Token & Deployment** in the control panel, or directly use the scripts below:
|
||||
|
||||
#### Docker Deployment
|
||||
|
||||
For Docker deployment, you can directly run the Agent image:
|
||||
|
||||
```bash
|
||||
docker pull ghcr.io/rain-kl/openflare-agent:latest
|
||||
docker rm -f openflare-agent 2>/dev/null || true
|
||||
docker run -d --name openflare-agent --restart unless-stopped \
|
||||
-p 80:80 -p 443:443/tcp -p 443:443/udp \
|
||||
-e OPENFLARE_SERVER_URL=http://your-server:3000 \
|
||||
-e OPENFLARE_AGENT_TOKEN=YOUR_AGENT_TOKEN \
|
||||
ghcr.io/rain-kl/openflare-agent:latest
|
||||
```
|
||||
|
||||
#### Local Installation
|
||||
|
||||
Using `discovery_token` to register:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--discovery-token YOUR_DISCOVERY_TOKEN
|
||||
```
|
||||
|
||||
Using node-specific `agent_token`:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--agent-token YOUR_AGENT_TOKEN
|
||||
```
|
||||
|
||||
The installation script defaults to `/opt/openflare-agent`, creates a `openflare-agent.service`, automatically searches for `openresty`, and can be executed repeatedly to reinstall or upgrade the Agent.
|
||||
|
||||
### 3. Uninstall Agent
|
||||
|
||||
To completely uninstall the Agent and clear local data, run:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/uninstall-agent.sh | bash
|
||||
```
|
||||
|
||||
The uninstallation script will stop and remove the `openflare-agent.service`, and delete the entire `/opt/openflare-agent` directory. It will not delete the local OpenResty installation.
|
||||
|
||||
### 4. Publish Your First Configuration
|
||||
|
||||
1. Log in to the management panel and add a reverse proxy rule.
|
||||
2. View the preview or change summary before publishing.
|
||||
3. Activate the new version.
|
||||
4. Agents will receive the configuration and apply it via WebSocket notification or subsequent heartbeats.
|
||||
|
||||
The version number format is fixed as `YYYYMMDD-NNN`. Historical versions are immutable, and rollback is achieved by reactivating an older version.
|
||||
|
||||
## UI Preview
|
||||
|
||||
### Dashboard Overview
|
||||
|
||||

|
||||
|
||||
### Node Details
|
||||
|
||||

|
||||
|
||||
### Proxy Configuration
|
||||
|
||||

|
||||
|
||||
## License
|
||||
|
||||
This project is licensed under [Apache License 2.0](./LICENSE).
|
||||
|
||||
## Star History
|
||||
|
||||
<a href="https://www.star-history.com/?repos=Rain-kl%2FOpenFlare&type=date&legend=bottom-right">
|
||||
<picture>
|
||||
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/chart?repos=Rain-kl/OpenFlare&type=date&theme=dark&legend=top-left" />
|
||||
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/chart?repos=Rain-kl/OpenFlare&type=date&legend=top-left" />
|
||||
<img alt="Star History Chart" src="https://api.star-history.com/chart?repos=Rain-kl/OpenFlare&type=date&legend=top-left" />
|
||||
</picture>
|
||||
</a>
|
||||
@@ -2,9 +2,9 @@
|
||||
|
||||
# OpenFlare
|
||||
|
||||
**[📖 中文](./README.md) | [English](./README.en.md)**
|
||||
**[English](./README.md) | [简体中文](./README.zh-CN.md)**
|
||||
|
||||
OpenFlare 是开源 CDN 编排与边缘安全平台。它支持反向代理、集中式配置同步、内网穿透(Tunnels)、动态 WAF 防护以及防 CC 挑战。
|
||||
OpenFlare is an open-source CDN orchestration and edge security platform. It supports reverse proxy, centralized configuration synchronization, in-network tunneling (Tunnels), dynamic WAF protection, and CC defense challenges.
|
||||
|
||||
</div>
|
||||
|
||||
@@ -21,64 +21,64 @@ OpenFlare 是开源 CDN 编排与边缘安全平台。它支持反向代理、
|
||||
</p>
|
||||
|
||||
> [!WARNING]
|
||||
> 使用 `admin` 用户初次登录系统后,务必修改默认密码 `12345678`。
|
||||
> After the first login with the `admin` user, you must change the default password `12345678`.
|
||||
>
|
||||
> BETA 版本为开发测试阶段的临时产物,可能存在未知问题,请勿在生产环境使用。
|
||||
> The BETA version is a temporary product in the development and testing stage and may have unknown issues. It should not be used in production environments.
|
||||
|
||||
## 文档
|
||||
## Documentation
|
||||
|
||||
**https://open-flare.pages.dev**
|
||||
**https://openflare.fyrn.link**
|
||||
|
||||
常用入口:
|
||||
Common entry points:
|
||||
|
||||
* [快速开始](https://open-flare.pages.dev/guide/quick-start)
|
||||
* [部署说明](https://open-flare.pages.dev/deployment/deployment)
|
||||
* [配置项参考](https://open-flare.pages.dev/reference/configuration)
|
||||
* [系统设计](https://open-flare.pages.dev/design/)
|
||||
* [Quick Start](https://openflare.fyrn.link/guide/quick-start)
|
||||
* [Deployment Guide](https://openflare.fyrn.link/deployment/deployment)
|
||||
* [Configuration Reference](https://openflare.fyrn.link/reference/configuration)
|
||||
* [System Design](https://openflare.fyrn.link/design/)
|
||||
|
||||
## 核心能力
|
||||
## Core Capabilities
|
||||
|
||||
* **反代配置管理**:以网站规则为聚合边界,支持多域名绑定与多上游负载均衡,统一管理所有 OpenResty 节点的反代配置。
|
||||
* **安全内网穿透(Tunnels)**:开源版的 Cloudflare Tunnels。无须公网 IP 或暴露入向端口,通过 Relay 中继节点与 OpenFlared 客户端安全反向穿透内网 Web 服务至公网。
|
||||
* **边缘 WAF 安全防护**:提供全局与自定义规则组,支持手动/自动/订阅型 IP 组、MaxMind GeoIP 国家级地域准入、IP 组成员 Checksum 差分同步(无需 Nginx 重载)以及自定义拦截响应。
|
||||
* **防 CC 与人机挑战(PoW)**:内置高性能客户端密码学 Proof of Work 挑战(类似 Turnstile),在网关边缘秒级拦截并阻断僵尸网络与爬虫。
|
||||
* **Pages 静态托管**:支持上传或从受限 Remote URL、公开 GitHub Release asset 同步预构建产物;GitHub latest 可定时检查并可选自动发布。所有来源统一生成不可变部署,由边缘 Agent 拉取并通过 OpenResty 本地提供服务,支持回滚、SPA Fallback 与 API 反向代理。
|
||||
* **TLS 证书自动化**:支持证书动态上传、多域名证书自动匹配绑定,以及通过 ACME 协议向 Let's Encrypt 自动申请与续期证书。
|
||||
* **Uptime Kuma 监控同步**:与 Uptime Kuma 集成,自动差分同步监控站点列表,实时感知节点存活与服务可用状态。
|
||||
* **SSO 单点登录**:支持 GitHub OAuth 与标准 OIDC 协议,无缝接入企业身份提供商实现统一登录。
|
||||
* **统一观测**:聚合节点请求指标、实时访问日志明细、宿主机与 Nginx 资源快照、健康事件以及网络波动补传缓冲。
|
||||
* **Reverse Proxy Configuration Management**: Uses website rules as the aggregation boundary, supports multi-domain binding and multi-upstream load balancing, and centrally manages reverse proxy configurations for all OpenResty nodes.
|
||||
* **Secure In-Network Tunneling (Tunnels)**: Open-source version of Cloudflare Tunnels. No public IP or exposed inbound ports are required. Securely reverse-proxy internal web services to the public internet through Relay relay nodes and OpenFlared clients.
|
||||
* **Edge WAF Security Protection**: Provides global and custom rule groups, supports manual/auto/subscription-type IP groups, MaxMind GeoIP national-level geographic access control, IP group member Checksum differential synchronization (no Nginx reload required), and custom blocking responses.
|
||||
* **CC Defense and Human-Computer Challenge (PoW)**: Built-in high-performance client-side cryptography Proof of Work challenge (similar to Turnstile). Secures high-speed interception and blocking of zombie networks and crawlers at the gateway edge.
|
||||
* **Pages Static Hosting**: Supports uploading or synchronizing pre-built artifacts from restricted Remote URLs or public GitHub Release assets. GitHub latest can be checked periodically and optionally auto-published. All sources are unified to generate immutable deployments, pulled by the edge Agent and served locally by OpenResty, supporting rollbacks, SPA Fallback, and API reverse proxy.
|
||||
* **TLS Certificate Automation**: Supports dynamic certificate uploads, automatic multi-domain certificate matching and binding, and automatic issuance and renewal of certificates from Let's Encrypt via the ACME protocol.
|
||||
* **Uptime Kuma Monitoring Synchronization**: Integrated with Uptime Kuma to automatically perform differential synchronization of monitoring site lists, real-time awareness of node availability and service status.
|
||||
* **SSO Single Sign-On**: Supports GitHub OAuth and standard OIDC protocol for seamless integration with enterprise identity providers to achieve unified login.
|
||||
* **Unified Observability**: Aggregates node request metrics, real-time access log details, host and Nginx resource snapshots, health events, and network fluctuation replenishment buffers.
|
||||
|
||||
## 界面预览
|
||||
## Interface Preview
|
||||
|
||||
### 仪表盘总览
|
||||
### Dashboard Overview
|
||||
|
||||

|
||||
|
||||
### 访问日志
|
||||
### Access Logs
|
||||
|
||||

|
||||
|
||||
### WAF 防护
|
||||
### WAF Protection
|
||||
|
||||

|
||||
|
||||
## 快速开始
|
||||
## Quick Start
|
||||
|
||||
### 硬件配置推荐
|
||||
### Hardware Configuration Recommendations
|
||||
|
||||
| 组件 | 最低硬件配额 | 推荐硬件配额 | 说明 |
|
||||
| --- |-------------------------------| --- | --- |
|
||||
| **Server 控制面** | 1 核 CPU / 2 GB 内存 / 20 GB 磁盘 | 2 核 CPU / 4 GB 内存 / 50 GB+ 磁盘 | 磁盘用量需根据访问日志留存时长与并发流量合理扩容 |
|
||||
| **Agent 数据面** | 1 核 CPU / 512 MB 内存 / 2 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 10 GB+ 磁盘 | 根据 OpenResty 的并发代理连接量与 WAF 拦截处理扩容 |
|
||||
| **Relay 中继节点**| 1 核 CPU / 1 GB 内存 / 5 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 20 GB 磁盘 | frps 传输中继吞吐量主要受带宽与 CPU 吞吐能力限制 |
|
||||
| **OpenFlared 客户端**| 1 核 CPU / 256 MB 内存 / 1 GB 磁盘 | 1 核 CPU / 512 MB 内存 / 5 GB 磁盘 | 独立运行于内网,自身资源占用极小,保障网络吞吐即可 |
|
||||
| Component | Minimum Hardware Requirements | Recommended Hardware Requirements | Notes |
|
||||
|------------------------|-----------------------------------|-----------------------------------|-------|
|
||||
| **Server Control Plane** | 1 CPU core / 2 GB RAM / 20 GB disk | 2 CPU cores / 4 GB RAM / 50 GB+ disk | Disk usage should be expanded reasonably based on access log retention duration and concurrent traffic |
|
||||
| **Agent Data Plane** | 1 CPU core / 512 MB RAM / 2 GB disk | 2 CPU cores / 2 GB RAM / 10 GB+ disk | Expanded based on OpenResty concurrent proxy connections and WAF interception processing |
|
||||
| **Relay Relay Node** | 1 CPU core / 1 GB RAM / 5 GB disk | 2 CPU cores / 2 GB RAM / 20 GB disk | frps transmission relay throughput is mainly limited by bandwidth and CPU throughput |
|
||||
| **OpenFlared Client** | 1 CPU core / 256 MB RAM / 1 GB disk | 1 CPU core / 512 MB RAM / 5 GB disk | Runs independently on the internal network with extremely low resource consumption; only network throughput needs to be guaranteed |
|
||||
|
||||
### 1. 启动 Server
|
||||
### 1. Start the Server
|
||||
|
||||
使用 docker-compose
|
||||
Use `docker-compose`:
|
||||
|
||||
```bash
|
||||
# 下载环境变量模板并创建 .env 文件
|
||||
# Download environment variable template and create .env file
|
||||
curl -o .env.example https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/.env.example
|
||||
cp .env.example .env
|
||||
```
|
||||
@@ -135,24 +135,24 @@ volumes:
|
||||
openflare_redis_data:
|
||||
```
|
||||
|
||||
详细部署说明见 [部署文档](https://open-flare.pages.dev/deployment/deployment)。
|
||||
See the [deployment documentation](https://openflare.fyrn.link/deployment/deployment) for details.
|
||||
|
||||
访问地址:`http://localhost:3000`
|
||||
Access address: `http://localhost:3000`
|
||||
|
||||
默认账号:
|
||||
Default account:
|
||||
|
||||
* 用户名:`admin`
|
||||
* 密码:`12345678`
|
||||
* Username: `admin`
|
||||
* Password: `12345678`
|
||||
|
||||
### 2. 安装 Agent
|
||||
### 2. Install Agent
|
||||
|
||||
安装 Agent 前请先在节点上安装 OpenResty,或改用内置 OpenResty 的 Agent Docker 镜像。
|
||||
Before installing the Agent, first install OpenResty on the node or use the built-in OpenResty Agent Docker image.
|
||||
|
||||
你可以在控制面板的节点管理->详情->节点信息->节点标识与部署复制安装命令,或直接使用下面的脚本:
|
||||
You can copy the installation command from the control panel's **Nodes Management -> Details -> Node Information -> Node ID and Deployment**, or use the script below:
|
||||
|
||||
#### Docker 部署
|
||||
#### Docker Deployment
|
||||
|
||||
Docker 部署可直接运行 Agent 镜像:
|
||||
Docker deployment can directly run the Agent image:
|
||||
|
||||
```bash
|
||||
docker pull ghcr.io/rain-kl/openflare-agent:latest
|
||||
@@ -165,9 +165,9 @@ docker run -d --name openflare-agent --restart unless-stopped \
|
||||
ghcr.io/rain-kl/openflare-agent:latest
|
||||
```
|
||||
|
||||
## 开源协议
|
||||
## Open Source License
|
||||
|
||||
本项目采用 [Apache License 2.0](./LICENSE) 开源。
|
||||
This project is licensed under the [Apache License 2.0](./LICENSE).
|
||||
|
||||
## Star History
|
||||
|
||||
|
||||
+180
@@ -0,0 +1,180 @@
|
||||
<div align="center">
|
||||
|
||||
# OpenFlare
|
||||
|
||||
**[English](./README.md) | [简体中文](./README.zh-CN.md)**
|
||||
|
||||
OpenFlare 是开源 CDN 编排与边缘安全平台。它支持反向代理、集中式配置同步、内网穿透(Tunnels)、动态 WAF 防护以及防 CC 挑战。
|
||||
|
||||
</div>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/LICENSE">
|
||||
<img src="https://img.shields.io/github/license/Rain-kl/OpenFlare?color=brightgreen" alt="license">
|
||||
</a>
|
||||
<a href="https://github.com/Rain-kl/OpenFlare/releases/latest">
|
||||
<img src="https://img.shields.io/github/v/release/Rain-kl/OpenFlare?color=brightgreen&include_prereleases" alt="release">
|
||||
</a>
|
||||
<a href="https://github.com/Rain-kl/OpenFlare/pkgs/container/openflare">
|
||||
<img src="https://img.shields.io/badge/GHCR-ghcr.io%2Frain--kl%2Fopenflare-brightgreen" alt="ghcr">
|
||||
</a>
|
||||
</p>
|
||||
|
||||
> [!WARNING]
|
||||
> 使用 `admin` 用户初次登录系统后,务必修改默认密码 `12345678`。
|
||||
>
|
||||
> BETA 版本为开发测试阶段的临时产物,可能存在未知问题,请勿在生产环境使用。
|
||||
|
||||
## 文档
|
||||
|
||||
**https://openflare.fyrn.link**
|
||||
|
||||
常用入口:
|
||||
|
||||
* [快速开始](https://openflare.fyrn.link/guide/quick-start)
|
||||
* [部署说明](https://openflare.fyrn.link/deployment/deployment)
|
||||
* [配置项参考](https://openflare.fyrn.link/reference/configuration)
|
||||
* [系统设计](https://openflare.fyrn.link/design/)
|
||||
|
||||
## 核心能力
|
||||
|
||||
* **反代配置管理**:以网站规则为聚合边界,支持多域名绑定与多上游负载均衡,统一管理所有 OpenResty 节点的反代配置。
|
||||
* **安全内网穿透(Tunnels)**:开源版的 Cloudflare Tunnels。无须公网 IP 或暴露入向端口,通过 Relay 中继节点与 OpenFlared 客户端安全反向穿透内网 Web 服务至公网。
|
||||
* **边缘 WAF 安全防护**:提供全局与自定义规则组,支持手动/自动/订阅型 IP 组、MaxMind GeoIP 国家级地域准入、IP 组成员 Checksum 差分同步(无需 Nginx 重载)以及自定义拦截响应。
|
||||
* **防 CC 与人机挑战(PoW)**:内置高性能客户端密码学 Proof of Work 挑战(类似 Turnstile),在网关边缘秒级拦截并阻断僵尸网络与爬虫。
|
||||
* **Pages 静态托管**:支持上传或从受限 Remote URL、公开 GitHub Release asset 同步预构建产物;GitHub latest 可定时检查并可选自动发布。所有来源统一生成不可变部署,由边缘 Agent 拉取并通过 OpenResty 本地提供服务,支持回滚、SPA Fallback 与 API 反向代理。
|
||||
* **TLS 证书自动化**:支持证书动态上传、多域名证书自动匹配绑定,以及通过 ACME 协议向 Let's Encrypt 自动申请与续期证书。
|
||||
* **Uptime Kuma 监控同步**:与 Uptime Kuma 集成,自动差分同步监控站点列表,实时感知节点存活与服务可用状态。
|
||||
* **SSO 单点登录**:支持 GitHub OAuth 与标准 OIDC 协议,无缝接入企业身份提供商实现统一登录。
|
||||
* **统一观测**:聚合节点请求指标、实时访问日志明细、宿主机与 Nginx 资源快照、健康事件以及网络波动补传缓冲。
|
||||
|
||||
## 界面预览
|
||||
|
||||
### 仪表盘总览
|
||||
|
||||

|
||||
|
||||
### 访问日志
|
||||
|
||||

|
||||
|
||||
### WAF 防护
|
||||
|
||||

|
||||
|
||||
## 快速开始
|
||||
|
||||
### 硬件配置推荐
|
||||
|
||||
| 组件 | 最低硬件配额 | 推荐硬件配额 | 说明 |
|
||||
| --- |-------------------------------| --- | --- |
|
||||
| **Server 控制面** | 1 核 CPU / 2 GB 内存 / 20 GB 磁盘 | 2 核 CPU / 4 GB 内存 / 50 GB+ 磁盘 | 磁盘用量需根据访问日志留存时长与并发流量合理扩容 |
|
||||
| **Agent 数据面** | 1 核 CPU / 512 MB 内存 / 2 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 10 GB+ 磁盘 | 根据 OpenResty 的并发代理连接量与 WAF 拦截处理扩容 |
|
||||
| **Relay 中继节点**| 1 核 CPU / 1 GB 内存 / 5 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 20 GB 磁盘 | frps 传输中继吞吐量主要受带宽与 CPU 吞吐能力限制 |
|
||||
| **OpenFlared 客户端**| 1 核 CPU / 256 MB 内存 / 1 GB 磁盘 | 1 核 CPU / 512 MB 内存 / 5 GB 磁盘 | 独立运行于内网,自身资源占用极小,保障网络吞吐即可 |
|
||||
|
||||
### 1. 启动 Server
|
||||
|
||||
使用 docker-compose
|
||||
|
||||
```bash
|
||||
# 下载环境变量模板并创建 .env 文件
|
||||
curl -o .env.example https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/.env.example
|
||||
cp .env.example .env
|
||||
```
|
||||
|
||||
```yaml
|
||||
services:
|
||||
openflare:
|
||||
image: ghcr.io/rain-kl/openflare:latest
|
||||
restart: unless-stopped
|
||||
env_file: .env
|
||||
environment:
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
ports:
|
||||
- "3000:3000"
|
||||
volumes:
|
||||
- openflare_uploads:/app/uploads
|
||||
depends_on:
|
||||
postgres:
|
||||
condition: service_healthy
|
||||
redis:
|
||||
condition: service_healthy
|
||||
|
||||
postgres:
|
||||
image: postgres:17-alpine
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
POSTGRES_DB: ${DB_NAME:-openflare}
|
||||
POSTGRES_USER: ${DB_USERNAME:-openflare}
|
||||
POSTGRES_PASSWORD: ${DB_PASSWORD:-replace-with-strong-password}
|
||||
volumes:
|
||||
- openflare_postgres_data:/var/lib/postgresql/data
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U ${DB_USERNAME:-openflare} -d ${DB_NAME:-openflare}"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
|
||||
redis:
|
||||
image: valkey/valkey:8.0-alpine
|
||||
restart: unless-stopped
|
||||
command: ["valkey-server", "--appendonly", "yes"]
|
||||
volumes:
|
||||
- openflare_redis_data:/data
|
||||
healthcheck:
|
||||
test: ["CMD", "valkey-cli", "ping"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
start_period: 5s
|
||||
|
||||
volumes:
|
||||
openflare_uploads:
|
||||
openflare_postgres_data:
|
||||
openflare_redis_data:
|
||||
```
|
||||
|
||||
详细部署说明见 [部署文档](https://openflare.fyrn.link/deployment/deployment)。
|
||||
|
||||
访问地址:`http://localhost:3000`
|
||||
|
||||
默认账号:
|
||||
|
||||
* 用户名:`admin`
|
||||
* 密码:`12345678`
|
||||
|
||||
### 2. 安装 Agent
|
||||
|
||||
安装 Agent 前请先在节点上安装 OpenResty,或改用内置 OpenResty 的 Agent Docker 镜像。
|
||||
|
||||
你可以在控制面板的节点管理->详情->节点信息->节点标识与部署复制安装命令,或直接使用下面的脚本:
|
||||
|
||||
#### Docker 部署
|
||||
|
||||
Docker 部署可直接运行 Agent 镜像:
|
||||
|
||||
```bash
|
||||
docker pull ghcr.io/rain-kl/openflare-agent:latest
|
||||
docker rm -f openflare-agent 2>/dev/null || true
|
||||
docker run -d --name openflare-agent --restart unless-stopped \
|
||||
-p 80:80 -p 443:443/tcp -p 443:443/udp \
|
||||
-v openflare-agent-pages:/data/var/lib/openflare/pages \
|
||||
-e OPENFLARE_SERVER_URL=http://your-server:3000 \
|
||||
-e OPENFLARE_AGENT_TOKEN=YOUR_AGENT_TOKEN \
|
||||
ghcr.io/rain-kl/openflare-agent:latest
|
||||
```
|
||||
|
||||
## 开源协议
|
||||
|
||||
本项目采用 [Apache License 2.0](./LICENSE) 开源。
|
||||
|
||||
## Star History
|
||||
|
||||
<a href="https://www.star-history.com/?repos=Rain-kl%2FOpenFlare&type=date&legend=bottom-right">
|
||||
<picture>
|
||||
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/chart?repos=Rain-kl/OpenFlare&type=date&theme=dark&legend=top-left" />
|
||||
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/chart?repos=Rain-kl/OpenFlare&type=date&legend=top-left" />
|
||||
<img alt="Star History Chart" src="https://api.star-history.com/chart?repos=Rain-kl/OpenFlare&type=date&legend=top-left" />
|
||||
</picture>
|
||||
</a>
|
||||
+5
-1
@@ -1,8 +1,12 @@
|
||||
// Copyright 2026 Arctel.net
|
||||
// SPDX-License-Identifier: Apache-2.0
|
||||
|
||||
// Command agent runs the OpenFlare edge agent daemon.
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"log/slog"
|
||||
"os"
|
||||
@@ -132,7 +136,7 @@ func main() {
|
||||
go geoIPUpdater.Run(ctx)
|
||||
slog.Info("agent process started")
|
||||
|
||||
if err = runner.Run(ctx); err != nil && err != context.Canceled {
|
||||
if err = runner.Run(ctx); err != nil && !errors.Is(err, context.Canceled) {
|
||||
slog.Error("agent process exited with error", "error", err)
|
||||
stop()
|
||||
os.Exit(1)
|
||||
|
||||
@@ -1,3 +1,6 @@
|
||||
// Copyright 2026 Arctel.net
|
||||
// SPDX-License-Identifier: Apache-2.0
|
||||
|
||||
package main
|
||||
|
||||
import (
|
||||
|
||||
+5
-1
@@ -1,8 +1,12 @@
|
||||
// Copyright 2026 Arctel.net
|
||||
// SPDX-License-Identifier: Apache-2.0
|
||||
|
||||
// Command flared runs the OpenFlare tunnel client daemon.
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"log/slog"
|
||||
"os"
|
||||
@@ -63,7 +67,7 @@ func main() {
|
||||
|
||||
slog.Info("flared process started")
|
||||
|
||||
if err := runner.Run(ctx); err != nil && err != context.Canceled {
|
||||
if err := runner.Run(ctx); err != nil && !errors.Is(err, context.Canceled) {
|
||||
slog.Error("flared process exited with error", "error", err)
|
||||
stop()
|
||||
os.Exit(1)
|
||||
|
||||
+5
-1
@@ -1,8 +1,12 @@
|
||||
// Copyright 2026 Arctel.net
|
||||
// SPDX-License-Identifier: Apache-2.0
|
||||
|
||||
// Command relay runs the OpenFlare relay node daemon.
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"log/slog"
|
||||
"os"
|
||||
@@ -63,7 +67,7 @@ func main() {
|
||||
|
||||
slog.Info("relay process started")
|
||||
|
||||
if err := runner.Run(ctx); err != nil && err != context.Canceled {
|
||||
if err := runner.Run(ctx); err != nil && !errors.Is(err, context.Canceled) {
|
||||
slog.Error("relay process exited with error", "error", err)
|
||||
stop()
|
||||
os.Exit(1)
|
||||
|
||||
@@ -26,11 +26,14 @@ RUN --mount=type=cache,target=/go/pkg/mod \
|
||||
RUN apk add --no-cache bash curl \
|
||||
&& bash scripts/fetch-agent-geoip-mmdb.sh
|
||||
|
||||
FROM openresty/openresty:alpine
|
||||
FROM openresty/openresty:alpine-slim
|
||||
|
||||
RUN apk add --no-cache ca-certificates tzdata perl libmaxminddb su-exec libcap \
|
||||
RUN apk add --no-cache ca-certificates tzdata libmaxminddb su-exec libcap \
|
||||
&& ln -sf /usr/lib/libmaxminddb.so.0 /usr/lib/libmaxminddb.so \
|
||||
&& apk add --no-cache --virtual .build-deps perl curl \
|
||||
&& opm get anjia0532/lua-resty-maxminddb \
|
||||
&& apk del .build-deps \
|
||||
&& rm -rf /root/.opm \
|
||||
&& addgroup -S openflare \
|
||||
&& adduser -S -G openflare -H -h /data -s /sbin/nologin openflare \
|
||||
&& mkdir -p /etc/openflare /data/etc/openflare \
|
||||
@@ -42,13 +45,10 @@ ENV OPENFLARE_OPENRESTY_PATH=openresty \
|
||||
|
||||
COPY --from=builder /build/bin/openflare-agent /usr/local/bin/openflare-agent
|
||||
# Default agent paths: data_dir/etc/openflare/GeoLite2-*.mmdb
|
||||
COPY --from=builder /build/dist/geoip/GeoLite2-Country.mmdb /data/etc/openflare/GeoLite2-Country.mmdb
|
||||
COPY --from=builder /build/dist/geoip/GeoLite2-City.mmdb /data/etc/openflare/GeoLite2-City.mmdb
|
||||
RUN chown openflare:openflare /data/etc/openflare/GeoLite2-Country.mmdb /data/etc/openflare/GeoLite2-City.mmdb \
|
||||
&& chmod 644 /data/etc/openflare/GeoLite2-Country.mmdb /data/etc/openflare/GeoLite2-City.mmdb
|
||||
COPY --chown=openflare:openflare --chmod=644 --from=builder /build/dist/geoip/GeoLite2-Country.mmdb /data/etc/openflare/GeoLite2-Country.mmdb
|
||||
COPY --chown=openflare:openflare --chmod=644 --from=builder /build/dist/geoip/GeoLite2-City.mmdb /data/etc/openflare/GeoLite2-City.mmdb
|
||||
|
||||
COPY scripts/agent-entrypoint.sh /usr/local/bin/openflare-agent-entrypoint.sh
|
||||
RUN chmod +x /usr/local/bin/openflare-agent-entrypoint.sh
|
||||
COPY --chmod=755 scripts/agent-entrypoint.sh /usr/local/bin/openflare-agent-entrypoint.sh
|
||||
|
||||
EXPOSE 80 443 18081
|
||||
ENTRYPOINT ["/usr/local/bin/openflare-agent-entrypoint.sh"]
|
||||
|
||||
+46
-1
@@ -8,6 +8,37 @@ sidebar: false
|
||||
|
||||
格式基于 [Keep a Changelog](http://keepachangelog.com/),版本号遵循 [语义化版本](http://semver.org/)。
|
||||
|
||||
## [Unreleased]
|
||||
|
||||
### 新增
|
||||
- Cloudflare 指向分组:支持将已添加的域名在不同分组之间移动,自动排队更新远程 DNS 指向。
|
||||
- Cloudflare 指向分组:详情页支持批量勾选域名进行批量移动与批量移出操作。
|
||||
- WAF IP 组:在查看 IP 组弹窗中新增即时搜索功能,支持快速过滤和定位 IP 地址。
|
||||
|
||||
## [v3.5.5] - 2026-09-19
|
||||
|
||||
### 🛠 修复
|
||||
- 修复 Cloudflare 指向分组引用的节点已被删除时,分组列表/详情接口整体返回「Cloudflare 资源不存在」的问题;现会跳过缺失节点并继续返回其余分组。
|
||||
- 修复静态导出部署下访问 Cloudflare 指向分组详情(`/cloudflare/groups/{id}`,id 不为 1)会跳回首页并触发 React hydration 报错的问题。
|
||||
|
||||
### ⚡️ 优化与改进
|
||||
- 优化 openflare-agent Docker 镜像体积:精简运行时依赖并消除离线 IP 库冗余层,镜像总体积从 500MB+ 缩减至约 150MB。
|
||||
|
||||
## [v3.5.4] - 2026-08-29
|
||||
|
||||
### ✨ 新功能
|
||||
- 控制台接入中英双语(next-intl,无 URL 语言前缀):默认中文,可在顶栏或「外观设置」切换;选择写入 cookie 后刷新生效。
|
||||
|
||||
### 🛠 修复
|
||||
- 修复在网站列表中删除已加入 Cloudflare 指向分组的域名后,访问 Cloudflare 指向分组详情报错「Cloudflare 资源不存在」的问题。
|
||||
- 修复自定义 Webhook 推送在企业微信/钉钉返回 HTTP 200 但 `errcode` 非零时仍记为成功的问题;任务日志会记录上游响应体。
|
||||
- 修复 OpenTelemetry Resource 绑定 semconv schema 版本导致 SDK 升级后可能无法启动的问题。
|
||||
- 修复静态导出(build:embed)部署下切换语言无效的问题:此前页面在构建时固定为默认中文,运行时不再读取 `NEXT_LOCALE`;现在客户端会按 cookie/浏览器语言重新解析并切换界面语言与 `html lang`。
|
||||
- 修复 frpc 子进程在被杀后孤儿进程继续持有管道导致退出阻塞的问题。
|
||||
|
||||
### 💄 其他/体验
|
||||
- 前端使用 `next/font` 自托管 Inter 字体,并忽略浏览器扩展改写 `body` 属性引起的 hydration 警告。
|
||||
|
||||
## 重大变更
|
||||
|
||||
> [!IMPORTANT]
|
||||
@@ -16,7 +47,21 @@ sidebar: false
|
||||
>
|
||||
|
||||
|
||||
## [Unreleased]
|
||||
## [v3.5.3] - 2026-08-13
|
||||
|
||||
### 新增
|
||||
- 访问日志「日志明细」支持按 HTTP 状态码筛选,可直接输入任意状态码。
|
||||
- 访问日志「日志明细」支持自定义时间范围筛选,可按起止时间检索日志。
|
||||
- 首页看板改版:24 小时请求趋势拆分展示请求总量与 2xx/4xx/5xx 状态码类请求量并独占一行;移除宿主机磁盘指标,24 小时容量趋势(CPU/内存)并入业务流量卡片展示。
|
||||
|
||||
### 🛠 修复
|
||||
- 修复首页「来源分布」卡片在 PostgreSQL/SQLite 日志库下无数据的问题。
|
||||
- 修复源站错误页「仅针对 GET 请求」未真正透传非 GET 响应的问题:POST/PUT 等非 GET 请求现可完整看到源站原始报错内容。
|
||||
|
||||
## [v3.5.2] - 2026-08-09
|
||||
|
||||
### 🛠 修复
|
||||
- 修复 PostgreSQL 作为日志库时节点访问日志/可观测指标/用户访问日志批量写入失败的问题,现可正常写入。
|
||||
|
||||
## [v3.5.1] - 2026-08-09
|
||||
|
||||
|
||||
@@ -138,6 +138,7 @@ function sidebarDesign(): DefaultTheme.SidebarItem[] {
|
||||
{ text: '边缘可观测与业务流量统计', link: 'observability-design' },
|
||||
{ text: '观测数据传输模型', link: 'observability-transport-model' },
|
||||
{ text: '观测上报协议与表结构', link: 'observability-data-model' },
|
||||
{ text: '日志存储解耦', link: 'logstore' },
|
||||
{ text: 'Uptime Kuma 监控同步设计', link: 'kuma-design' },
|
||||
{ text: '登录验证码设计', link: 'login-captcha' }
|
||||
]
|
||||
|
||||
@@ -20,7 +20,7 @@ OpenFlare Agent 运行在代理节点侧。它不会接收远程 shell 指令,
|
||||
|
||||
## 一键安装
|
||||
|
||||
### 交互式安装 (推荐)
|
||||
### 交互式安装(推荐)
|
||||
|
||||
如果在不传递任何参数的情况下运行安装脚本,脚本将进入交互模式。您将可以通过向导选择安装方式(本地运行 / Docker 容器运行),并配置 Server 地址与认证 Token(若选择 Docker 方式且本地没有 Docker,脚本还会询问并智能安装 Docker):
|
||||
|
||||
@@ -28,7 +28,7 @@ OpenFlare Agent 运行在代理节点侧。它不会接收远程 shell 指令,
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash
|
||||
```
|
||||
|
||||
### 自动化 (非交互式) 安装
|
||||
### 自动化(非交互式)安装
|
||||
|
||||
如果在执行脚本时附加了任何参数,脚本将进入自动化安装模式,不需要任何交互。
|
||||
|
||||
@@ -115,7 +115,7 @@ curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/inst
|
||||
}
|
||||
```
|
||||
|
||||
如果不配置 `openresty_path`,Agent 默认调用 `openresty`。完整字段见 [配置项参考](../reference/configuration.md#agent-配置字段)。
|
||||
如果不配置 `openresty_path`,Agent 默认调用 `openresty`。完整字段见 [配置项参考](../reference/configuration.md#agent-命令行参数与配置字段)。
|
||||
|
||||
## Docker 运行
|
||||
|
||||
@@ -138,7 +138,7 @@ docker run -d --name openflare-agent --restart unless-stopped \
|
||||
|
||||
## 卸载
|
||||
|
||||
### 交互式卸载 (推荐)
|
||||
### 交互式卸载(推荐)
|
||||
|
||||
如果在不传递任何参数的情况下运行卸载脚本,脚本将进入交互模式。您可以通过提示菜单选择卸载方式(本地卸载 / Docker 容器卸载):
|
||||
|
||||
@@ -146,7 +146,7 @@ docker run -d --name openflare-agent --restart unless-stopped \
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/uninstall-agent.sh | bash
|
||||
```
|
||||
|
||||
### 卸载
|
||||
### Docker 容器卸载
|
||||
|
||||
停止并删除 `openflare-agent` 容器即可
|
||||
|
||||
@@ -156,4 +156,4 @@ curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/unin
|
||||
| --- |---------------------------------------------------------------------------------------------------------|
|
||||
| `agent_token 和 discovery_token 不能同时为空` | 检查 `agent.json` 至少配置了一个 Token |
|
||||
| 节点一直离线 | 在 Agent 节点执行 `curl -I http://your-server:3000`,确认 Server 地址可达 |
|
||||
| 发布后重复失败 | Agent 会阻断同一 `version + checksum` 的重复应用;在节点尝试强制同步,或者重新发布版本 |
|
||||
| 发布后重复失败 | Agent 会阻断同一 `version + checksum` 的重复应用;在节点详情页点击「强制同步」,或重新发布新版本 |
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
你会学到:OpenFlare 的推荐部署方式、Server 与 Agent 的运行要求、源码启动方式、联调步骤、升级与卸载入口。
|
||||
|
||||
生产环境建议使用 PostgreSQL 作为 Server 数据库,并通过 `config.yaml` 或环境变量配置 `APP_SESSION_SECRET` 等参数。完整 Docker Compose 部署还需 Redis 与 ClickHouse(见仓库根目录 `docker-compose.yaml`)。Agent 部署方式推荐为 Docker 部署(即直接使用内置 OpenResty 的 Agent 镜像);亦支持通过安装脚本或手动本地运行。
|
||||
生产环境建议使用 PostgreSQL 作为 Server 数据库,并通过 `config.yaml` 或环境变量配置 `APP_SESSION_SECRET` 等参数。完整 Docker Compose 部署需要 Redis;ClickHouse 可选,用于海量访问日志与观测时序(见仓库根目录 `docker-compose.yaml`)。Agent 支持 Docker 部署与本地安装脚本两种方式,Docker 镜像已内置 OpenResty 二进制。日志库判定与切换见 [日志存储解耦](../design/logstore.md)。
|
||||
|
||||
## 部署拓扑
|
||||
|
||||
@@ -49,7 +49,7 @@ Internal Service (192.168.x.x)
|
||||
|
||||
### 硬件配置推荐
|
||||
|
||||
| 组件 | 最低硬件配额 | 推荐硬件配额 | 说明 |
|
||||
| 组件 | 参考配置(入门) | 参考配置(生产) | 说明 |
|
||||
| --- |-------------------------------| --- | --- |
|
||||
| **Server 控制面** | 1 核 CPU / 2 GB 内存 / 20 GB 磁盘 | 2 核 CPU / 4 GB 内存 / 50 GB+ 磁盘 | 磁盘用量需根据访问日志留存时长与并发流量合理扩容 |
|
||||
| **Agent 数据面** | 1 核 CPU / 512 MB 内存 / 2 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 10 GB+ 磁盘 | 根据 OpenResty 的并发代理连接量与 WAF 拦截处理扩容 |
|
||||
|
||||
@@ -8,10 +8,10 @@
|
||||
|
||||
## 前置条件
|
||||
|
||||
1. **获取 Tunnel Token**:在 OpenFlare 管理端的「内网穿透」或「隧道管理」页面中,创建一个新的隧道实例,系统会自动生成唯一的 `tunnel_id` 与 `tunnel_token`(形如 `tun-<32hex>`)。
|
||||
1. **获取 Tunnel Token**:在管理端「节点管理」中新增一个类型为 **Tunnel** 的节点,保存后进入节点详情页即可查看该节点专属的接入 Token。
|
||||
2. **网络出方向权限**:内网服务器无需任何公网入方向 IP 或端口映射,但必须能够通过网络访问公网上的 **OpenFlare Server 地址** 以及对应的 **TunnelRelay 节点中继端口 (默认 7000)**。
|
||||
3. **软件依赖**(仅限宿主机直接部署):
|
||||
- 本地需有可执行的 `frpc` 二进制文件(建议版本为 `v0.61.0+` 或最新稳定版 `v0.69.0`),或通过参数显式指定路径。
|
||||
- 本地需有可执行的 `frpc` 二进制文件,或通过参数显式指定路径。
|
||||
|
||||
---
|
||||
|
||||
@@ -58,8 +58,8 @@ docker run -d --name openflared --restart unless-stopped \
|
||||
启动成功后,OpenFlared 将执行以下工作流:
|
||||
- **心跳与配置获取**:周期性向 Server 的 `/api/v1/tunnel/heartbeat` 和 `/api/v1/tunnel/config/active` 接口发起同步,验证 Token 并检测配置版本。
|
||||
- **文件渲染**:当检测到配置版本(或校验和 Checksum)变化时,会自动拉取该隧道的完整路由规则。如果绑定了多个中继 Relay,将为每个 Relay 分别在 `data_dir` 下渲染出 `frpc_{relayNodeID}.toml`。
|
||||
- **热重载或重启**:拉起对应的 `frpc` 子进程,或在配置文件发生改变时执行 `frpc reload` / 重启动作,以确保流量映射保持最新。
|
||||
- **异常自恢复**:如果本地 `frpc` 隧道进程异常退出,主控程序会在 5 秒的退避惩罚后自动尝试重新启动。
|
||||
- **配置变更重启**:当配置或校验和变化时,重新拉起对应的 `frpc` 子进程,以确保流量映射保持最新。
|
||||
- **异常自恢复**:如果本地 `frpc` 隧道进程异常退出,主控程序会按指数退避(初始 1 秒,上限 60 秒)自动重启。
|
||||
|
||||
### 2. 查看日志与连接状态
|
||||
|
||||
@@ -79,6 +79,6 @@ frpc process missing, starting {"relay_id": "..."}
|
||||
|
||||
### 3. 管理端确认
|
||||
|
||||
打开管理后台的 **「内网穿透」** 页面:
|
||||
- 查看对应隧道的在线状态,此时应当绿灯显示 **「在线」**。
|
||||
- 您可以清晰地看到该隧道目前连接了哪些中继节点,以及各内网服务的穿透路由详情。
|
||||
打开管理后台的 **「节点管理」**,进入对应 Tunnel 节点的详情页:
|
||||
- 查看节点在线状态与 flared 运行状态(WebSocket 已连接 / 运行中 / 离线)。
|
||||
- 查看当前应用版本与最近一次应用记录。
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
# 部署 Relay (Tunnel 中继)
|
||||
# 部署 Relay(Tunnel 中继)
|
||||
|
||||
你会学到:TunnelRelay 节点的职责、`openflare-relay` 的配置项与环境变量、使用 Docker 运行 Relay 的方法,以及如何通过源码手动构建并部署 Relay。
|
||||
|
||||
@@ -40,7 +40,7 @@
|
||||
|
||||
---
|
||||
|
||||
## Docker 运行)
|
||||
## Docker 运行
|
||||
|
||||
Docker 运行是 TunnelRelay 节点最便捷的部署方案。官方镜像内置了 `openflare-relay` 控制器与 `frps` 运行时,开箱即用。
|
||||
|
||||
@@ -84,7 +84,7 @@ docker logs -f openflare-relay
|
||||
- 从控制面获取最新的 frps 基础配置(包括 `bindPort`、`vhostHTTPPort` 与自动生成的隧道认证凭证 `auth_token`)。
|
||||
- 在本地自动渲染出 `data/frps.toml` 配置文件。
|
||||
- 自动拉起子进程 `frps -c data/frps.toml`。
|
||||
- 如果进程意外崩溃,Relay 将在 2 秒后自动拉起它。
|
||||
- 如果进程意外退出,Relay 会按指数退避(初始 1 秒,上限 60 秒)自动重启 frps。
|
||||
|
||||
### 3. 管理端确认
|
||||
|
||||
|
||||
@@ -7,7 +7,7 @@ OpenFlare Server 是 Gin + GORM 单体控制面,负责管理端 UI、管理 AP
|
||||
> [!IMPORTANT]
|
||||
> **关于外部依赖**:
|
||||
> OpenFlare 系统内建了对后台异步任务(Asynq 框架)的支持。因此,**无论采用何种部署模式,系统都必须依赖 Redis(或 Valkey)**。各个部署方案的主要差异在于主关系型数据库的选择(SQLite vs PostgreSQL)以及是否启用链路追踪服务(Jaeger)。
|
||||
> 若业务流量过大, 建议使用 ClickHouse 存储日志。
|
||||
> 若业务流量过大,建议使用 ClickHouse 存储日志。
|
||||
|
||||
> [!TIP]
|
||||
> **ClickHouse 服务端性能配置(推荐挂载)**
|
||||
@@ -34,11 +34,11 @@ volumes:
|
||||
|
||||
---
|
||||
|
||||
## 方式一:Docker 部署 (推荐)
|
||||
## 方式一:Docker 部署(推荐)
|
||||
|
||||
使用 Docker 部署可以免去本地配置 Go 与 Node.js 前端构建环境的麻烦。根据你的服务器硬件配置及业务需求,你可以选择以下三种方案之一:
|
||||
|
||||
### 1. 快速启动 (SQLite + Redis)
|
||||
### 1. 快速启动(SQLite + Redis)
|
||||
|
||||
> **适用场景**:测试体验、轻量化单机部署。
|
||||
>
|
||||
@@ -66,8 +66,6 @@ services:
|
||||
SQLITE_PATH: "/data/openflare.db"
|
||||
REDIS_ENABLED: "true"
|
||||
REDIS_ADDR: "redis:6379"
|
||||
CLICKHOUSE_ENABLED: "true"
|
||||
CLICKHOUSE_HOST: "clickhouse:9000"
|
||||
depends_on:
|
||||
redis:
|
||||
condition: service_healthy
|
||||
@@ -87,9 +85,9 @@ services:
|
||||
|
||||
---
|
||||
|
||||
### 2. 小流量业务场景 (PostgreSQL + Redis)
|
||||
### 2. 小流量业务场景(PostgreSQL + Redis)
|
||||
|
||||
> **适用场景**:生产环境、业务流量中小, PostgreSQL 不会成为日志记录的瓶颈。
|
||||
> **适用场景**:生产环境、业务流量中小,PostgreSQL 不会成为日志记录的瓶颈。
|
||||
|
||||
创建 `docker-compose.yaml` 文件:
|
||||
|
||||
@@ -157,11 +155,11 @@ docker compose up -d
|
||||
|
||||
---
|
||||
|
||||
### 3. 进阶版 (含 Jaeger 链路追踪的完整编排)
|
||||
### 3. 进阶版(含 Jaeger 链路追踪的完整编排)
|
||||
|
||||
> **适用场景**:大流量场景, 需要进行链路性能指标追踪。
|
||||
> **适用场景**:大流量场景,需要进行链路性能指标追踪。
|
||||
>
|
||||
> **特点**:在“生产推荐”全家桶的基础上,使用 ClickHouse 存储日志, 联动 Jaeger 作为 OpenTelemetry (OTel) 链路追踪的后端。
|
||||
> **特点**:在“生产推荐”全家桶的基础上,使用 ClickHouse 存储日志,联动 Jaeger 作为 OpenTelemetry (OTel) 链路追踪的后端。
|
||||
|
||||
创建 `docker-compose.yaml` 文件:
|
||||
|
||||
@@ -177,7 +175,7 @@ services:
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
OTEL_EXPORTER_OTLP_ENDPOINT: "http://jaeger:4317"
|
||||
OTEL_EXPORTER_OTLP_INSECURE: "true"
|
||||
OTEL_SAMPLING_RATE: "1.0" # 本地调试建议设为 1.0 以采样所有 Trace
|
||||
OTEL_SAMPLING_RATE: "1.0" # 采样率,1.0 表示采样全部 Trace
|
||||
ports:
|
||||
- "3000:3000"
|
||||
volumes:
|
||||
|
||||
@@ -17,4 +17,4 @@ docker compose up
|
||||
|
||||
## Agent 升级
|
||||
|
||||
Agent 是完全无状态的,升级时直接拉取最新镜像重建容器即可。具体部署命令与安装方式请参考 **[接入 Agent](./agent.md)**。
|
||||
Agent 本地仅缓存运行配置与状态文件,不保存业务数据;升级时直接拉取最新镜像重建容器即可。具体部署命令与安装方式请参考 **[接入 Agent](./agent.md)**。
|
||||
|
||||
@@ -92,12 +92,12 @@ sequenceDiagram
|
||||
Agent 对数据面 OpenResty 的管控实现了端到端的闭环,包含配置落地、语法验证、平滑重载和异常状态捕获:
|
||||
|
||||
### 1. 配置文件的落地组织
|
||||
同步成功后,Agent 会将配置按照特定的物理结构写入到本地 `/etc/nginx/openflare-lua/` 目录下(或配置指定的 `LuaDir`):
|
||||
同步成功后,Agent 将配置写入 `data_dir` 下(默认相对路径 `etc/nginx/`、`etc/openflare/`、`var/lib/openflare/`,具体以 `agent.json` 中 `main_config_path`、`route_config_path`、`cert_dir`、`lua_dir`、`runtime_config_dir`、`pages_dir` 等字段为准):
|
||||
* `nginx.conf`:主配置文件(替换相关占位符,配置性能参数、Shared Dictionaries 及全局 Server)。
|
||||
* `routes.conf`:路由配置文件(由 Agent 生成,包含所有代理网站的 Server 块、证书路径、缓存及速率限制指令)。
|
||||
* `conf.d/openflare_routes.conf`:路由配置文件(由 Agent 生成,包含所有代理网站的 Server 块、证书路径、缓存及速率限制指令)。
|
||||
* `certs/`:证书存放目录(文件命名为 `{cert_id}.crt` 和 `{cert_id}.key`)。
|
||||
* `waf/` 与 `pow/`:WAF 及防 CC 挑战所需的专用 Lua 运行时脚本。
|
||||
* `waf_config.json` 与 `waf_ip_groups.json`:WAF 过滤引擎所需的结构化规则配置文件。
|
||||
* `lua/waf/` 与 `lua/pow/`:WAF 及防 CC 挑战所需的专用 Lua 运行时脚本。
|
||||
* `etc/openflare/waf_config.json` 与 `waf_ip_groups.json`:WAF 过滤引擎所需的结构化规则配置文件。
|
||||
* `pages_dir`:Pages 静态站点部署目录,默认位于 `data_dir/var/lib/openflare/pages`。当激活配置引用 Pages **项目**时,Agent 按 `project_id` 请求控制面「最新激活包」(hash + package),以流式方式写入临时文件并执行实际响应上限与 SHA-256 校验,再安全解压到 `projects/{project_id}/releases/{hash}`。解压后会复核文件数与总字节,绝对防御上限为 2 GiB 包、1,000 个文件、单文件及总量 8 GiB;随后原子切换 `current` 并**立即删除同项目其它历史 release**(仅保留最新)。项目内切换激活无需重发主配置;多项目对账时单项目失败不阻塞其它项目。
|
||||
|
||||
### 2. 精细化的重载动作
|
||||
@@ -105,13 +105,13 @@ Agent 对数据面 OpenResty 的管控实现了端到端的闭环,包含配置
|
||||
2. **写入并替换占位符**:将最新拉取的模板写入,自动将模板中的绝对路径占位符(如 `__OPENFLARE_LUA_DIR__`、`__OPENFLARE_PAGES_DIR__`)替换为本地实际运行路径。
|
||||
3. **语法校验**:调用 `openresty -t -c <temp_nginx.conf>` 进行严格的语法测试。
|
||||
4. **平滑重载**:若校验通过,将新配置移至正式路径,执行 `openresty -s reload`。若 OpenResty 处于未启动状态,则使用当前配置拉起进程。
|
||||
5. **捕获异常**:校验或重载失败时,Agent 会截获标准错误输出(stderr),提取前 2000 个字符的详细报错信息。
|
||||
5. **捕获异常**:校验或重载失败时,Agent 截获命令标准输出(stderr/stdout)作为失败详情上报。
|
||||
|
||||
---
|
||||
|
||||
## 发布与配置应用模型
|
||||
|
||||
OpenFlare 摒弃了动态 Patch 节点配置的落后方式,采用 **不可变配置版本发布模型**。
|
||||
OpenFlare 采用 **不可变配置版本发布模型**,而非对节点配置进行在线动态 Patch。
|
||||
|
||||
```text
|
||||
修改规则 -> 预览 / 查看 diff -> 发布 -> 生成完整配置版本 -> 激活版本 -> Agent 拉取 -> 本地应用 -> 上报结果
|
||||
|
||||
@@ -145,7 +145,7 @@ OpenResty access.log(业务事实)
|
||||
|
|
||||
| Agent tail 增量明细(不 sum/count/uniq)
|
||||
v
|
||||
Server 入库 ClickHouse
|
||||
Server 经 logstore 入库(当前日志主库:PostgreSQL / SQLite / ClickHouse)
|
||||
|
|
||||
+---> 全局聚合 --> 看板「已提供数据 / 请求 / UV」
|
||||
+---> host∈Zone --> Zone「已提供数据」等(同一套语义)
|
||||
@@ -206,19 +206,3 @@ OpenResty 健康与连接数 --> 边缘健康(瞬时,不作 24h 业务总量
|
||||
| Pages artifact 与仓库构建分离 | 现有来源只导入预构建产物;未来 checkout/build 由 Server 隔离 executor 完成并复用 artifact pipeline,Agent 不执行第三方构建 |
|
||||
|
||||
---
|
||||
|
||||
## 贡献者阅读建议
|
||||
|
||||
修改系统架构或开发新功能前,请按以下顺序阅读:
|
||||
|
||||
1. **[产品边界](./index.md)**:了解 OpenFlare 核心定位与不允许逾越的设计边界。
|
||||
2. **[Agent 与发布模型](./agent-design.md)**:理解版本快照同步及失败回滚的安全兜底逻辑。
|
||||
3. **细分领域设计**:
|
||||
* Zone 与域名相关开发:阅读 [Zone 与域名资源设计](./zone-design.md)。
|
||||
* Cloudflare DNS 指向开发:阅读 [Cloudflare DNS 指向设计](./cloudflare-pointing.md)。
|
||||
* 穿透相关开发:阅读 [内网穿透隧道设计](./tunnel-design.md)。
|
||||
* WAF 相关开发:阅读 [WAF 设计](./waf-design.md) 与 [WAF 可编排规则设计](./waf-orchestration-design.md)。
|
||||
* Pages 托管开发:阅读 [Pages 静态托管设计](./pages-design.md)。
|
||||
* 监控同步开发:阅读 [Uptime Kuma 监控同步设计](./kuma-design.md)。
|
||||
* 看板/访问日志/节点指标开发:阅读 [观测数据传输模型](./observability-transport-model.md) 与 [边缘可观测与业务流量统计](./observability-design.md)。
|
||||
4. **[仓库结构](./index.md#仓库结构)**:明确各个物理目录分层职责,避免堆砌和重复开发。
|
||||
|
||||
@@ -187,8 +187,7 @@ OpenFlare 库表为 Source of Truth。每个成员期望:
|
||||
| 可选域名 | `GET /domains/available` |
|
||||
|
||||
* 成功 `response.OK`;失败 `response.Abort*`;**永不**在 JSON 中返回 Token。
|
||||
* Handler 与 `logics.go` 分离;CF 客户端可 mock 接口。
|
||||
* 变更后维护 Swagger(`make swagger`)。
|
||||
* Handler 与 `logics.go` 分离;CF 客户端以接口抽象便于替换。
|
||||
|
||||
## 前端
|
||||
|
||||
@@ -207,18 +206,9 @@ OpenFlare 库表为 Source of Truth。每个成员期望:
|
||||
* 典型:未配置 Token、Token 无效、节点无 IP、CF 无 Zone、同名多 A、限流。
|
||||
* Token 仅服务端解密使用;响应与日志禁止明文 Token。
|
||||
|
||||
## 数据迁移与测试
|
||||
## 数据迁移
|
||||
|
||||
* goose 双方言(PG/SQLite)新建三张表;默认值与 Go 零值一致。
|
||||
* 单测:Token 解析、reconcile 0/1/多条、橙云只初始化新成员、移出删远端(mock)、节点 IP 变更入队。
|
||||
* 禁止单测打真实 Cloudflare。
|
||||
|
||||
## 文档与边界同步
|
||||
|
||||
* 更新 [Zone 与域名资源设计](./zone-design.md):Zone 仍不内建权威 DNS;可选本模块负责 CF A 指向。
|
||||
* 更新 [系统架构](./architecture.md) 核心对象与阅读建议。
|
||||
* 更新 [产品边界](./index.md) 能力表。
|
||||
* 实现完成后写入 `docs/changelog/index.md` 的 `[Unreleased]`(纯设计文档变更不写 changelog)。
|
||||
|
||||
## 关键决策摘要
|
||||
|
||||
|
||||
@@ -204,7 +204,6 @@ access.log cache_status=$upstream_cache_status
|
||||
| 模型/默认 | 创建路由默认 `cache_policy=static`;读写时 `url`→`all` |
|
||||
| 快照 | `config_version` 快照规范化 |
|
||||
| UI | `proxy-routes/detail/components/cache-section.tsx` |
|
||||
| 测试 | `pkg/render/openresty/render_test.go` 等 |
|
||||
|
||||
---
|
||||
|
||||
@@ -235,21 +234,7 @@ access.log cache_status=$upstream_cache_status
|
||||
|
||||
---
|
||||
|
||||
## 7. 验证要点
|
||||
|
||||
* 渲染:无 Cookie/Auth/请求 Cache-Control 旁路;含 `proxy_cache_valid` 三行;`proxy_no_cache` 含 `$upstream_http_set_cookie`。
|
||||
* 单测:内置表含 `css`/`js`/`map`/`mjs`,**不含** `html`/`json`。
|
||||
* 手动:
|
||||
* 带 session Cookie 请求 `/a.js` → 第二次 `HIT`;
|
||||
* `/index.html` + `static` → 未缓存;
|
||||
* 源站对 eligible 路径返回 `Set-Cookie` → 不入库(持续 MISS/不 HIT);
|
||||
* 源站 `Cache-Control: private` → 不入库。
|
||||
* 观测:access log 三态与原始 `cache_status` 一致。
|
||||
* 生效:配置版本发布并节点应用后验证。
|
||||
|
||||
---
|
||||
|
||||
## 8. 决策矩阵(防漏判)
|
||||
## 7. 决策矩阵(防漏判)
|
||||
|
||||
| 场景 | CF | OpenFlare(本设计) |
|
||||
| --- | --- | --- |
|
||||
@@ -264,18 +249,7 @@ access.log cache_status=$upstream_cache_status
|
||||
|
||||
---
|
||||
|
||||
## 9. 后续路线图
|
||||
|
||||
1. Auth 完整 RFC/CF 条件缓存(Lua)
|
||||
2. 强制 Edge TTL / `proxy_ignore_headers`(Cache Rules 级)
|
||||
3. Purge API
|
||||
4. Cache Rules(有序规则 + 动作)
|
||||
5. 全局默认可缓存扩展名可配置;可选对齐 CF 更长扩展名表
|
||||
6. HEAD→GET
|
||||
|
||||
---
|
||||
|
||||
## 10. 决策记录
|
||||
## 8. 决策记录
|
||||
|
||||
| 决策 | 选择 | 原因 |
|
||||
| --- | --- | --- |
|
||||
|
||||
@@ -32,6 +32,8 @@ OpenFlare 适合需要统一管理多台 OpenResty 代理节点的团队,具
|
||||
| **Pages 静态托管** | 支持上传或从 Remote URL、公开 GitHub Release 同步预构建产物;GitHub latest 可定时检查并可选自动发布。不可变部署由边缘节点拉取并由 OpenResty 本地服务,支持回滚、API 反代与 SPA Fallback | [Pages 静态托管设计](./pages-design.md) / [Pages 使用指南](../guide/pages-usage.md) |
|
||||
| **TLS 证书自动续期** | 将证书显式绑定到 Zone 域名,并通过 ACME 协议向 Let's Encrypt 申请/续期证书 | [Zone 与域名资源设计](./zone-design.md) |
|
||||
| **多节点监控与观测** | 访问日志为业务流量唯一真相;Agent 只上报明细与主机读数,Server 统一聚合;与 Zone/看板对账 | [观测数据传输模型](./observability-transport-model.md) / [边缘可观测与业务流量统计](./observability-design.md) / [上报协议与表结构](./observability-data-model.md) / [系统架构](./architecture.md) |
|
||||
| **日志存储** | 访问日志与可观测时序走可切换日志主库(随业务主库或 ClickHouse);关闭 ClickHouse 后仍可写可查 | [日志存储解耦](./logstore.md) |
|
||||
| **控制台双语** | 无 URL 前缀的 zh-CN / en,cookie `NEXT_LOCALE` 优先,兼容静态导出 | [前端 i18n 设计](../superpowers/specs/2026-07-24-frontend-i18n-design.md) |
|
||||
|
||||
---
|
||||
|
||||
@@ -62,7 +64,7 @@ OpenFlare 适合需要统一管理多台 OpenResty 代理节点的团队,具
|
||||
### 5. 系统与版本边界
|
||||
* **全局单一激活版本**:所有节点拉取并消费同一份全局激活配置。不进行按节点分组的差异化配置发布。
|
||||
* **单租户架构**:OpenFlare 仅供单团队在受信任的内部网络部署使用。采用单租户设计,不支持细粒度的多用户角色或多租户资源隔离。
|
||||
* **外部基础设施依赖性**:Server 虽支持 SQLite 作为本地轻量关系数据库,但**系统必须强制依赖外部 Redis(或 Valkey)及 ClickHouse 实例**。Redis 用于处理分布式协调、后台异步队列(Asynq 框架)及系统级全局缓存;ClickHouse 用于接收海量节点访问日志与基础观测的异步 Flush。系统不支持完全脱离这两个组件运行。
|
||||
* **外部基础设施依赖性**:Server **必须依赖**外部 Redis(或 Valkey),用于分布式协调、Asynq 队列与系统缓存。关系库为 PostgreSQL,或关闭 `database.enabled` 时使用 SQLite。ClickHouse **可选**:不启用时,访问日志与可观测时序由当前日志主库(随业务主库)承接;启用后可通过「切换日志数据库」任务迁到 ClickHouse。系统不支持脱离 Redis 运行。详情见 [日志存储解耦](./logstore.md)。
|
||||
|
||||
---
|
||||
|
||||
@@ -98,7 +100,7 @@ OpenFlare 已收敛为**单 monorepo**(Go 模块 `github.com/Rain-kl/Wavelet`
|
||||
| `internal/apps/openflare/{agent,relay,flared}/` | **Server 侧**边缘协议处理器(鉴权、心跳、WS) |
|
||||
| `internal/model/` | GORM 实体 / DTO / 无 IO 领域规则(`openflare_*.go` + 平台模型);**不含** DB 访问 |
|
||||
| `internal/infra/persistence/migrator/goose/` | goose SQL 迁移(PostgreSQL / SQLite / ClickHouse) |
|
||||
| `internal/repository/` | 数据访问层(平台 + OpenFlare 业务 CRUD、缓存、ClickHouse 分析读写);**唯一**持久化入口 |
|
||||
| `internal/repository/` | 数据访问层(平台 + OpenFlare 业务 CRUD、缓存、`logstore` 日志读写);**唯一**持久化入口 |
|
||||
| `internal/infra/task/` | Asynq 异步任务(Worker + Scheduler) |
|
||||
| `internal/infra/config/` | Viper 配置加载 |
|
||||
| `internal/shared/` | 统一 API 响应封装(`response/`) |
|
||||
@@ -191,6 +193,7 @@ OpenFlare 已收敛为**单 monorepo**(Go 模块 `github.com/Rain-kl/Wavelet`
|
||||
## 文档维护原则
|
||||
|
||||
* 产品范围或系统边界变化:更新本文档([产品边界](./index.md))。
|
||||
* 日志存储、日志表判定或切换协议变化:更新 [日志存储解耦](./logstore.md)。
|
||||
* 系统结构、组件分工变化:更新 [系统架构](./architecture.md)。
|
||||
* 发布、同步、回滚与 Agent 模型变化:更新 [Agent 与发布模型](./agent-design.md)。
|
||||
* 部署方式变化:更新 [部署说明](../deployment/deployment.md) 与 README。
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
在多节点的网关架构中,监控系统的状态与反向代理路由的状态通常是相互脱节的:
|
||||
1. **录入开销大**:每当网关控制面新增或下线一个站点,管理员都必须在监控系统(如 Uptime Kuma)中重复配置对应的探测地址与告警策略。
|
||||
2. **数据不一致**:当代理路由域名发生变更或切换 HTTPS 时,容易遗漏修改监控参数,导致监控系统误报或漏报。
|
||||
3. **环境污染隐患**:如果简单的在监控中执行全量“删除-重建”同步,不仅会清空监控系统中的历史统计指标和 SLA 曲线,还会影响到用户在此监控实例上自行配置的、与网关无关的其他监控任务。
|
||||
3. **环境污染隐患**:若在监控中执行全量“删除-重建”同步,会清空监控系统中的历史统计指标与 SLA 曲线,还会影响用户在此监控实例上自行配置的、与网关无关的其他监控任务。
|
||||
|
||||
为了解决这些痛点,OpenFlare 引入了基于客户端/服务器模式的 **Uptime Kuma 自动监控同步机制**,实现网关站点路由定义与可用性监测系统的强一致、低开销以及零污染同步。
|
||||
|
||||
@@ -106,4 +106,4 @@ stateDiagram-v2
|
||||
* Server 周期性(每 1 分钟)通过后台的 Cron Job 探测是否达到配置的同步间隔(`UptimeKumaSyncInterval`)。
|
||||
* 任务内部设计了互斥锁(Mutex Locking)。如果前一次同步请求因为网络延迟等原因尚未结束,下一次调度将自动跳过,防止并发多个 Socket.IO 连接对 Uptime Kuma 实例造成 DDOS 冲击。
|
||||
2. **WebSocket 状态监听**:
|
||||
* 同步程序利用 Socket.IO 的事件监听机制,在连接建立后,必须等到监听到 `monitorList` 事件的完整列表推送后,才允许向下执行差分算法,以规避因为数据加载不完整导致误删监控项的边界情况。
|
||||
* 同步程序利用 Socket.IO 的事件监听机制,在连接建立后,必须等到监听到 `monitorList` 事件的完整列表推送后,才允许向下执行差分算法,避免因数据加载不完整导致误删监控项。
|
||||
|
||||
@@ -7,14 +7,14 @@
|
||||
## 1. 业务背景与产品范围
|
||||
|
||||
### 背景与痛点
|
||||
根据我们的系统安全分析,OpenFlare 的登录端点 `/api/user/login` 虽然配置了基于 IP 的限流限制,但由于缺少用户维度的防护机制,攻击者可使用代理池绕过 IP 限制对高权限账户(如 `root`)实施撞库和暴力破解。同时,对于系统登录页面,标准的视觉验证码对用户体验和无障碍不够友好。
|
||||
OpenFlare 的登录端点 `/api/v1/user/login` 缺少用户维度的防护机制,攻击者可使用代理池对高权限账户(如 `root`)实施撞库和暴力破解。同时,标准的视觉验证码对登录页用户体验和无障碍不够友好。
|
||||
|
||||
### 产品范围与技术选型
|
||||
* **技术选型**:Cap (Proof-of-Work 驱动的无感无图像验证码解决方案)。
|
||||
- **核心原理**:客户端(Widget/网页)从服务器获取工作量证明 (PoW) 的难题,使用浏览器后台计算求解并将答案回传。服务器验证答案的正确性,完成人机识别。
|
||||
- **优势**:无感、无图像验证、不依赖任何外部第三方 API 节点(私密)、包极小。
|
||||
* **接入范围**:控制面 Server 登录 API(`/api/user/login`)以及前端登录页面。
|
||||
* **配置粒度**:支持管理员通过控制台 Option 表随时开启/关闭验证码(`CapLoginEnabled`)。
|
||||
* **接入范围**:控制面 Server 登录 API(`/api/v1/user/login`)以及前端登录页面。
|
||||
* **配置粒度**:支持管理员通过控制台 Option 表随时开启/关闭验证码(`cap_login_enabled`)。
|
||||
|
||||
---
|
||||
|
||||
@@ -28,7 +28,7 @@
|
||||
* 暴露 `POST /api/cap/challenge` 接口,为客户端分发 PoW 难题和签名的 JWT Token。
|
||||
* 暴露 `POST /api/cap/redeem` 接口,校验客户端提交的 PoW 解答并核发带有失效时间的登录凭证(Redeem Token)。
|
||||
* 将 Redeem Token 与对应过期时间保存在内存缓存/Redis 缓存中。
|
||||
* 在 `POST /api/user/login` 接口中,若启用了验证码保护,先校验并消耗(单次失效)对应的 `cap-token`。
|
||||
* 在 `POST /api/v1/user/login` 接口中,若启用了验证码保护,先校验并消耗(单次失效)对应的 `cap-token`。
|
||||
|
||||
### 2.2 验证流时序图
|
||||
```mermaid
|
||||
@@ -51,7 +51,7 @@ sequenceDiagram
|
||||
Server->>Browser: 返回 {success: false, reason}
|
||||
end
|
||||
User->>Browser: 输入账号密码,点击登录
|
||||
Browser->>Server: POST /api/user/login (在 HTTP 请求头中携带 X-Cap-Token)
|
||||
Browser->>Server: POST /api/v1/user/login (在 HTTP 请求头中携带 X-Cap-Token)
|
||||
alt CapLoginEnabled = true
|
||||
Server->>Server: Middleware (CapAuth) 校验并消费 X-Cap-Token
|
||||
alt token 合法且未过期且未被消费
|
||||
@@ -80,7 +80,7 @@ sequenceDiagram
|
||||
"error_msg": "",
|
||||
"data": {
|
||||
"challenge": {
|
||||
"c": 50,
|
||||
"c": 1,
|
||||
"s": 32,
|
||||
"d": 4
|
||||
},
|
||||
@@ -108,7 +108,7 @@ sequenceDiagram
|
||||
}
|
||||
```
|
||||
|
||||
#### 3. 登录接口 (POST /api/user/login)
|
||||
#### 3. 登录接口 (POST /api/v1/user/login)
|
||||
* **请求负载保持不变**:
|
||||
```json
|
||||
{
|
||||
@@ -124,4 +124,4 @@ sequenceDiagram
|
||||
1. **JWT 临时状态绑定**:难题在生成时就被签入 JWT payload,包含过期时间限制(10 分钟)。
|
||||
2. **Replay 拦截(Nonce 消耗)**:当客户端调用 `/redeem` 提交解答时,后端在缓存中标记该 JWT Signature 已使用。重复提交相同的解密包将返回 `already_redeemed`。
|
||||
3. **Redeem 一次性核销(单次失效)**:当客户端登录并提交 `cap-token` 时,后端在检验到合法性后立即从缓存中删除该 Key,防止黑客提取历史正确的 `cap-token` 进行重放登录。
|
||||
4. **验证机制无感化**:通过调整 `c (难题数)=50`,`d (难度)=4`,普通用户在桌面端和移动端只需 0.5 秒至 1.5 秒即可静默解出,极大地兼顾了用户体验和反爬效果。
|
||||
4. **验证机制无感化**:通过调整 `c (难题数)`、`d (难度)` 等参数平衡求解耗时与反爬强度,用户在后台静默解出,不打断登录流程。
|
||||
|
||||
@@ -0,0 +1,86 @@
|
||||
# 日志存储解耦
|
||||
|
||||
你会学到:哪些表属于日志用途、为什么不能绑死 ClickHouse,以及新增一张日志表时必须走哪条代码路径。
|
||||
|
||||
观测字段与上报协议仍以 [观测上报协议与表结构](./observability-data-model.md) 为准;本文只约定**存到哪、怎么切库**。
|
||||
|
||||
---
|
||||
|
||||
## 1. 目标
|
||||
|
||||
* **ClickHouse 可选**:不启用时,PostgreSQL(或关闭主库时的 SQLite)完整承接写入、查询、聚合与清理。
|
||||
* **上层不碰底层库**:apps 只面向 `internal/repository/logstore`(或 `repository` 门面)。`repository/analytics` 与 `db.ChConn` / `db.ChDB` 仅供 logstore 的 ClickHouse 实现使用。
|
||||
* **可切换**:任务管理里的「切换日志数据库」在 PostgreSQL/SQLite 与 ClickHouse 之间复制数据并翻转主库;迁移期间冻结写入,成功才切换,源数据不删。
|
||||
|
||||
---
|
||||
|
||||
## 2. 什么算日志表
|
||||
|
||||
同时满足才进 logstore:
|
||||
|
||||
* 追加写入,几乎不更新单行
|
||||
* 按时间查询或聚合,允许按保留天数删除
|
||||
* 关闭 ClickHouse 后仍要能写、能查
|
||||
* 不参与网站 / 节点 / 证书等事务一致性
|
||||
|
||||
**不要**做成日志表:Zone、节点、配置版本、任务执行、上传元数据。这些走业务主库 `repository`。
|
||||
|
||||
当前日志域:
|
||||
|
||||
| 域 | 接口 | 表 |
|
||||
| --- | --- | --- |
|
||||
| 节点访问日志 | `AccessLogStore` | `of_node_access_logs` |
|
||||
| 可观测时序 | `ObservabilityStore` | `of_node_metric_snapshots` / `of_node_edge_health` / `of_node_obs_frps` / `of_node_obs_frpc` |
|
||||
| 用户访问审计 | `UserAccessLogStore` | `w_user_access_logs` |
|
||||
|
||||
ClickHouse 上的小时级物化视图(如 `of_access_log_hourly`)只服务 CH 查询加速。PostgreSQL / SQLite **不建**同构聚合表,查询时从原始日志实时聚合。
|
||||
|
||||
---
|
||||
|
||||
## 3. 分层
|
||||
|
||||
| 层级 | 路径 | 职责 |
|
||||
| --- | --- | --- |
|
||||
| 抽象 | `internal/repository/logstore` | 接口 + `Active` / `BuildForMigration`;按 `log_database` 选实现 |
|
||||
| CH 实现 | `logstore/clickhouse_store.go` | 委托 `repository/analytics`(原生批量 + 现有聚合 SQL) |
|
||||
| 主库实现 | `logstore/postgres_store.go` | PostgreSQL(高频表按月分区)与 SQLite(普通表)共用 GORM |
|
||||
| Model | `internal/model/analytics` | 实体与批量 SQL,无 IO |
|
||||
| 入队 | `chwriter` / `risk_control` + `batchwriter` | `FlushFunc` 调 `logstore.Active`;节点日志 / 可观测经 hooks 入队 |
|
||||
| 约束 | `logstore/imports_test.go` | apps 禁止 import `repository/analytics` |
|
||||
|
||||
`log_database` 只有两种合法状态:**随业务主库**(`postgres` 或 `sqlite`)或 **`clickhouse`**。不存在「主库 PostgreSQL + 日志 SQLite」。`log_database` / `log_db_migration` 受保护,管理端不可改。
|
||||
|
||||
启动时:`log_database=clickhouse` 但 ClickHouse 未启用会拒绝启动,须先重新启用 ClickHouse 并切回主库后再关掉。
|
||||
|
||||
---
|
||||
|
||||
## 4. 切换协议
|
||||
|
||||
任务类型 `of_log_db_switch`(管理端名称「切换日志数据库」),参数 `target`。
|
||||
|
||||
1. 校验目标合法且不等于当前库。
|
||||
2. 写 `log_db_migration=migrating`,排空在途 batchwriter(`Drain`,不要 `Stop` writer)。此后写入返回明确错误(HTTP 503),不排队积压。
|
||||
3. 清空目标日志表后按 id 分页复制;复制前对 PostgreSQL 目标 `EnsurePartitions`。
|
||||
4. 全部成功才写 `log_database=target` 并清除迁移标记;失败清除标记,写入继续走源库。
|
||||
5. 源数据不删;重试前重新清空目标以保证幂等。
|
||||
|
||||
不要另起切换协议,也不要在任务里直连 `analyticsrepo`。
|
||||
|
||||
---
|
||||
|
||||
## 5. 新增一张日志表
|
||||
|
||||
列名必须在 ClickHouse / PostgreSQL / SQLite 三套 goose 迁移中一致。要点:
|
||||
|
||||
* 高频表:CH 用 `MergeTree` + `toYYYYMM`;PG 用 `PARTITION BY RANGE(时间列)`,主键含分区键;SQLite 普通表 + 索引。
|
||||
* ID 用 snowflake `uint64`,迁移时原样保留。
|
||||
* 写入走独立 `batchwriter`;flush 调 `logstore.Active`,不要 `analyticsrepo.BatchInsert`。
|
||||
* 切换任务的 `copy*` 必须覆盖新表;清理走已有 `log_retention_days_*` 或 `metric_retention_days`,不要用错 TTL。
|
||||
|
||||
运行时配置见 [配置项参考 · 日志存储](../reference/configuration.md#8-日志存储log-database)。
|
||||
|
||||
---
|
||||
|
||||
## 6. 相关文档
|
||||
|
||||
* 观测字段与上报协议:[观测上报协议与表结构](./observability-data-model.md)
|
||||
@@ -264,7 +264,7 @@ Agent:tail access.log → 解析 JSON 行 → 原样字段上报(可截断 p
|
||||
#### 边界
|
||||
|
||||
* Pages 静态 / 无 `proxy_cache` 的 location:多为空或 `-` → **未使用缓存**,不得标成「命中」。
|
||||
* 第一期只做明细可见;命中率看板、hourly 维度可后续用同一列聚合。
|
||||
* 明细详情展示缓存状态;命中率看板与 hourly 维度可基于同一列扩展。
|
||||
|
||||
**单次心跳条数建议:**
|
||||
|
||||
@@ -302,10 +302,10 @@ Agent:tail access.log → 解析 JSON 行 → 原样字段上报(可截断 p
|
||||
|
||||
写入关系库健康事件表(现有模型即可),不进访问日志湖。
|
||||
|
||||
### 3.8 Go 协议草图(目标)
|
||||
### 3.8 Go 协议结构
|
||||
|
||||
```go
|
||||
// pkg/protocol/agent.go(目标形态,实现时替换旧类型)
|
||||
// pkg/protocol/agent.go(当前实现)
|
||||
|
||||
type NodePayload struct {
|
||||
SchemaVersion int `json:"schema_version,omitempty"`
|
||||
@@ -447,7 +447,7 @@ type BufferedFacts struct {
|
||||
|
||||
---
|
||||
|
||||
## 5. 表结构(目标 DDL)
|
||||
## 5. 表结构(DDL)
|
||||
|
||||
> 引擎与 TTL 与现网一致倾向:访问日志 90 天,指标 30 天。
|
||||
> `id` 使用控制面 Snowflake/唯一 UInt64。
|
||||
@@ -556,7 +556,6 @@ GROUP BY node_id, hour, host;
|
||||
2. 即便存每小时 UV,对多小时窗口 **相加会严重高估**(同一 IP 跨小时重复计)。
|
||||
3. 产品「24h 独立访客」只认整窗 `uniqExact`;趋势图主序列是请求量/错误/字节,分时 UV 非主指标。
|
||||
|
||||
可选未来:若需要分时 UV 曲线,再单独加 `AggregatingMergeTree` 状态表或查询时对明细做 `uniqExact` 按小时 group(成本更高,不阻塞当前看板)。
|
||||
### 5.3 L3 事实表:`of_node_metric_snapshots`(保留,语义明确)
|
||||
|
||||
```sql
|
||||
@@ -761,18 +760,7 @@ Agent 解析:
|
||||
|
||||
---
|
||||
|
||||
## 10. 实现检查清单
|
||||
|
||||
- [x] `pkg/protocol`:仅 v2 字段,无兼容别名
|
||||
- [x] Agent:只组 `host_metrics` / `edge_health` / `access_logs` / `buffered`
|
||||
- [x] Server:无 request_reports / openresty 吞吐;健康当前态 PG、时序 CH
|
||||
- [x] CH migration:`request_length`、`request_time_ms`、`of_node_edge_health`、`of_access_log_hourly`、hourly 回填
|
||||
- [x] 看板/Zone API 统一读 access log 聚合
|
||||
- [x] UV:整窗 uniqExact;Zone 曲线标明分桶 UV;小时趋势不绘 UV
|
||||
|
||||
---
|
||||
|
||||
## 11. 修订记录
|
||||
## 10. 修订记录
|
||||
|
||||
| 日期 | 说明 |
|
||||
| --- | --- |
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# 边缘可观测与业务流量统计重构设计
|
||||
|
||||
你会学到:当前观测链路为何出现「看板 OpenResty 出站」与「Zone 已提供数据」不一致、字段与聚合为何冗余,以及目标架构如何让 **Agent 只上报事实、Server 只解释事实**,业务流量以访问日志为唯一真相源。
|
||||
你会学到:本次重构要解决的问题(「看板 OpenResty 出站」与「Zone 已提供数据」不一致、字段与聚合冗余),以及目标架构如何让 **Agent 只上报事实、Server 只解释事实**,业务流量以访问日志为唯一真相源。
|
||||
|
||||
---
|
||||
|
||||
@@ -38,7 +38,7 @@
|
||||
### 2.1 产品约束(继承)
|
||||
|
||||
* 单租户、全局单激活配置;观测不引入多租户计费隔离。
|
||||
* ClickHouse 为访问日志与时序观测的强制分析存储。
|
||||
* 访问日志与时序观测走可切换日志主库(默认 ClickHouse,可切换 PostgreSQL/SQLite),见 [日志存储解耦](./logstore.md)。
|
||||
* Agent 无入向控制、Pull 模型;离线期间本地 OpenResty 继续服务,观测可本地缓冲后补传。
|
||||
|
||||
### 2.2 工程约束
|
||||
@@ -98,9 +98,9 @@ Server = 入库 + 聚合 + 归属 + 趋势 + 对账
|
||||
|
||||
---
|
||||
|
||||
## 4. 现状问题(基线)
|
||||
## 4. 重构前的问题(基线)
|
||||
|
||||
### 4.1 当前数据流(冗余)
|
||||
### 4.1 重构前数据流(冗余)
|
||||
|
||||
```text
|
||||
一次 HTTP 请求
|
||||
@@ -307,11 +307,10 @@ Agent 职责:
|
||||
|
||||
### 7.4 OpenResty 本地观测
|
||||
|
||||
**收敛后建议:**
|
||||
收敛后的状态:
|
||||
|
||||
* 保留:健康检查、`stub_status` 当前连接。
|
||||
* 删除主路径依赖:`log.lua` 中对 request/status/domain/rx/tx 的 shared dict 业务计数,以及 `/openflare/observability` 作为 TrafficReport 来源。
|
||||
* 若短期内保留 endpoint 供调试,不得再写入 Server 权威分析表。
|
||||
* 主路径不再依赖 `log.lua` 的 shared dict 业务计数;`/openflare/observability` 只返回健康与连接快照,不作为业务报表来源。
|
||||
|
||||
### 7.5 与 Agent 设计文档的关系
|
||||
|
||||
@@ -479,7 +478,7 @@ bytes_sent (= $body_bytes_sent), request_length
|
||||
### 11.4 健康状态权威
|
||||
|
||||
* **当前态**:PG `openresty_status` / `openresty_message`。
|
||||
* **时序**:CH `of_node_edge_health`(status + connections;无 message)。
|
||||
* **时序**:日志主库 `of_node_edge_health`(status + connections;无 message)。
|
||||
|
||||
### 11.5 UV
|
||||
|
||||
@@ -497,37 +496,7 @@ bytes_sent (= $body_bytes_sent), request_length
|
||||
|
||||
---
|
||||
|
||||
## 13. 验证标准
|
||||
|
||||
### 13.1 对账
|
||||
|
||||
在仅有单一 Zone 产生流量的环境:
|
||||
|
||||
```text
|
||||
看板「已提供数据」(24h) ≈ Zone「已提供的数据总计」(24h)
|
||||
误差仅来自时间窗对齐(整点截断)与未计入 Host
|
||||
```
|
||||
|
||||
多 Zone 时:
|
||||
|
||||
```text
|
||||
sum(各 Zone 已提供) + sum(未归属 Host) = 全局已提供
|
||||
```
|
||||
|
||||
### 13.2 回归
|
||||
|
||||
* Agent 单测:只解析与 offset,不出现业务 sum 断言为「上报契约」。
|
||||
* Server:Zone stats 与 dashboard business traffic 共用聚合测例。
|
||||
* 前端:文案快照/测试中不再出现业务含义的「OpenResty 出站」与「已提供数据」双卡片。
|
||||
|
||||
### 13.3 性能
|
||||
|
||||
* 24h 看板聚合 P95 可接受(必要时 hourly MV)。
|
||||
* 心跳 payload 体积:明细批量有上限;超限拆缓冲,不在 Agent 做摘要替代。
|
||||
|
||||
---
|
||||
|
||||
## 14. 风险与权衡
|
||||
## 13. 风险与权衡
|
||||
|
||||
| 风险 | 缓解 |
|
||||
| --- | --- |
|
||||
@@ -543,7 +512,7 @@ sum(各 Zone 已提供) + sum(未归属 Host) = 全局已提供
|
||||
|
||||
---
|
||||
|
||||
## 15. 关键决策摘要
|
||||
## 14. 关键决策摘要
|
||||
|
||||
| 决策 | 选择 | 否决方案 |
|
||||
| --- | --- | --- |
|
||||
@@ -556,7 +525,7 @@ sum(各 Zone 已提供) + sum(未归属 Host) = 全局已提供
|
||||
|
||||
---
|
||||
|
||||
## 16. 文档与代码映射(落地时)
|
||||
## 15. 文档与代码映射
|
||||
|
||||
| 区域 | 主要路径 |
|
||||
| --- | --- |
|
||||
@@ -568,8 +537,6 @@ sum(各 Zone 已提供) + sum(未归属 Host) = 全局已提供
|
||||
| 看板 | `internal/apps/openflare/dashboard/`、`internal/apps/openflare/observability/analytics.go` |
|
||||
| 前端 | `frontend/app/(main)/page.tsx`、`components/dashboard/*`、`websites/.../zone-overview.tsx` |
|
||||
|
||||
实现计划见:`docs/plan/20260717-observability-redesign.md`。
|
||||
|
||||
**推荐阅读顺序:**
|
||||
|
||||
1. **[观测数据传输模型](./observability-transport-model.md)**(最新:传什么、从哪采、频率、示例 JSON)
|
||||
@@ -577,7 +544,7 @@ sum(各 Zone 已提供) + sum(未归属 Host) = 全局已提供
|
||||
|
||||
---
|
||||
|
||||
## 17. 修订记录
|
||||
## 16. 修订记录
|
||||
|
||||
| 日期 | 说明 |
|
||||
| --- | --- |
|
||||
|
||||
@@ -6,7 +6,7 @@
|
||||
|
||||
---
|
||||
|
||||
## 0. 先记住三层(不要混)
|
||||
## 0. 先记住三层
|
||||
|
||||
| 层 | 回答的问题 | 唯一数据来源 | 产品例子 |
|
||||
| --- | --- | --- | --- |
|
||||
@@ -215,7 +215,7 @@ cache_status ← $upstream_cache_status 【缓存状态;UI 可推导命中/
|
||||
|
||||
落库表:`of_node_access_logs`(可选 Server 侧 `of_access_log_hourly` 加速,**Agent 不写**)。
|
||||
|
||||
### 4.4 频率再强调
|
||||
### 4.4 上报频率
|
||||
|
||||
```text
|
||||
请求发生 ──立即──► 写 access.log
|
||||
@@ -229,9 +229,9 @@ Server ──立即/批量──► CH
|
||||
|
||||
## 5. L2 健康:edge_health 与 `/openflare/observability`
|
||||
|
||||
### 5.1 本机监测口(合并后目标)
|
||||
### 5.1 本机监测口
|
||||
|
||||
**只保留一个接口:**
|
||||
**数据采集接口:**
|
||||
|
||||
```http
|
||||
GET http://127.0.0.1:{openresty_observability_port}/openflare/observability
|
||||
@@ -241,7 +241,7 @@ GET http://127.0.0.1:{openresty_observability_port}/openflare/observability
|
||||
|
||||
**职责:** 回答「OpenResty 此刻怎样」,**不**回答业务已提供多少数据。
|
||||
|
||||
#### 返回示例(目标 JSON)
|
||||
#### 返回示例
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -263,7 +263,7 @@ GET http://127.0.0.1:{openresty_observability_port}/openflare/observability
|
||||
| `connections.active` | **瞬时** | Nginx 连接状态(原 stub_status Active) | 当前活跃连接 |
|
||||
| `reading` / `writing` / `waiting` | **瞬时** | 同上细分 | 可选但建议带 |
|
||||
|
||||
**不返回(已从目标模型删除):**
|
||||
**不返回(已删除):**
|
||||
|
||||
| 旧字段 | 原因 |
|
||||
| --- | --- |
|
||||
@@ -272,9 +272,9 @@ GET http://127.0.0.1:{openresty_observability_port}/openflare/observability
|
||||
| `source_countries` | 从未实现;国家走 Server GeoIP |
|
||||
| `server.accepts/handled/requests` | 进程累计 counter,易与业务请求混淆;主路径不收录 |
|
||||
|
||||
**`/openflare/stub_status`:** 合并进上述 JSON 后 **删除**(过渡期可双挂,Agent 只打合并口)。
|
||||
**`/openflare/stub_status`:** 保留;`/openflare/observability` 内部读取该口组装连接数 JSON,Agent 健康检查也直接探测该口。
|
||||
|
||||
### 5.2 采集机制(读快照,不是「调用才开始统计业务」)
|
||||
### 5.2 采集机制(读快照)
|
||||
|
||||
```text
|
||||
Nginx 在连接建立/释放时维护 Active connections 等
|
||||
@@ -284,9 +284,8 @@ Agent GET /openflare/observability
|
||||
只读取「当前值」拼 JSON 返回
|
||||
```
|
||||
|
||||
- **不是** GET 一次才去扫 access.log。
|
||||
- **不是** 60 秒业务均值。
|
||||
- 是 **瞬时 gauge 快照**。
|
||||
- 不扫 access.log、不算 60 秒业务均值。
|
||||
- 返回 **瞬时 gauge 快照**。
|
||||
|
||||
### 5.3 上报示例(装进 NodePayload)
|
||||
|
||||
@@ -361,7 +360,7 @@ Agent 读本机(如 `/proc`、磁盘统计等),**每次组包时读一次*
|
||||
|
||||
---
|
||||
|
||||
## 7. 一次完整上报示例(拼起来)
|
||||
## 7. 一次完整上报示例
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -466,13 +465,13 @@ t=6s 下一轮…
|
||||
|
||||
---
|
||||
|
||||
## 10. 旧模型对照(帮助消歧)
|
||||
## 10. 旧模型对照
|
||||
|
||||
| 旧做法 | 新模型 |
|
||||
| --- | --- |
|
||||
| Lua dict 60s 窗 request_count + Agent 10s 拉 + Server sum | **删除**;请求数 = 日志 count |
|
||||
| openresty_tx 当「出站」 | **删除**;已提供数据 = `sum(bytes_sent)` |
|
||||
| 两个口 observability + stub_status | **合并为一个** observability,只返回连接/探活 |
|
||||
| 两个口 observability + stub_status | 数据采集统一走 observability;stub_status 保留为探活与内部读取口 |
|
||||
| TrafficReport 预聚合 | **删除**;协议与 API 均无此路径 |
|
||||
| 业务与网卡混称「流量」 | **分文案、分 API、分表** |
|
||||
| 健康 status/message | **PG 最新态权威**;CH 仅 status+连接时序 |
|
||||
|
||||
@@ -15,7 +15,7 @@
|
||||
* **默认可视**:默认启用,默认状态码标签 `500-599`,默认 OpenFlare 极简错误页。
|
||||
* **可自定义**:管理员可在线编辑完整 HTML;空 HTML 表示使用内置默认模板。
|
||||
* **状态码透传**:HTTP 响应 `status` 保持原错误码(如 502、522);页面正文通过 `{{status}}` 展示同一数值。
|
||||
* **全局统一**:侧栏「网站管理 → 错误页」单一配置,全站反代路由共用。
|
||||
* **全局统一**:侧栏「网站管理 → 响应页面」单一配置,全站反代路由共用。
|
||||
* **与发布一致**:配置经 Option 持久化,进入配置版本快照后随发布/回滚下发。
|
||||
|
||||
### 1.2 非目标
|
||||
@@ -35,12 +35,13 @@
|
||||
| 条件 | 行为 |
|
||||
| --- | --- |
|
||||
| 开关开启,且响应状态码落在展开后的集合内 | 返回自定义/默认 HTML,**status 不变** |
|
||||
| 开关开启且启用 GET-only,非 GET 请求返回匹配状态码 | 透传源站原始响应,不替换 |
|
||||
| 开关关闭 | 不生成 `error_page` 相关指令,透传 |
|
||||
| 状态码不在集合内 | 不替换 |
|
||||
| Pages 上游路由 | 不应用本功能 |
|
||||
| 源站成功返回 2xx/3xx/4xx(未配置时) | 不替换 |
|
||||
|
||||
实现上对反代 `location` 启用 `proxy_intercept_errors on`,因此**源站返回的**匹配 5xx 等也会被拦截,而不仅是网关本地生成的 502。
|
||||
全方法模式下对反代 `location` 启用 `proxy_intercept_errors on`,因此**源站返回的**匹配 5xx 等也会被拦截,而不仅是网关本地生成的 502;GET-only 模式改用 Lua header/body 过滤器仅替换 GET 响应正文。
|
||||
|
||||
### 2.2 状态码标签语法
|
||||
|
||||
@@ -81,6 +82,7 @@ Tags Input 每条标签:
|
||||
| `origin_error_page_enabled` | bool 字符串 | `true` | 总开关 |
|
||||
| `origin_error_page_status_codes` | JSON 字符串数组 | `["500-599"]` | 原始标签 |
|
||||
| `origin_error_page_html` | 文本 | `""` | 空 = 内置默认;最大 **256 KiB** |
|
||||
| `origin_error_page_get_only` | bool 字符串 | `false` | 仅对 GET 请求替换错误页,其它方法透传 |
|
||||
|
||||
API 复用:
|
||||
|
||||
@@ -106,6 +108,7 @@ API 复用:
|
||||
OriginErrorPageEnabled bool
|
||||
OriginErrorPageStatusCodes []string // 原始标签
|
||||
OriginErrorPageHTML string // 空则渲染器用内置默认
|
||||
OriginErrorPageGetOnly bool
|
||||
```
|
||||
|
||||
构建快照时从 Option 读取;Agent 只消费快照,不直读控制面 DB。
|
||||
@@ -121,26 +124,27 @@ OriginErrorPageHTML string // 空则渲染器用内置默认
|
||||
|
||||
```nginx
|
||||
proxy_intercept_errors on;
|
||||
error_page <expanded codes...> = /__openflare_origin_error;
|
||||
error_page <expanded codes...> @__openflare_origin_error;
|
||||
|
||||
location = /__openflare_origin_error {
|
||||
internal;
|
||||
location @__openflare_origin_error {
|
||||
default_type text/html;
|
||||
charset utf-8;
|
||||
# 保持 ngx.status 为原错误码
|
||||
content_by_lua_block {
|
||||
# 读取模板,替换 {{status}} / {{host}} 后输出 body
|
||||
# ngx.status 保持原错误码
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 4.2 运行时替换
|
||||
|
||||
采用 **internal location 内轻量 Lua(或现有 resty 能力)** 读模板并 `string.gsub` 替换占位符,**不**把 status 固化进静态文件(请求间状态码不同)。
|
||||
采用 **命名 location 内 `content_by_lua_block`** 读模板并替换占位符,**不**把 status 固化进静态文件(请求间状态码不同)。GET-only 模式在反代 location 内用 `header_filter_by_lua_block` + `body_filter_by_lua_block` 仅替换 GET 响应正文,非 GET 请求透传。
|
||||
|
||||
禁止将错误页统一改为 HTTP 200。
|
||||
|
||||
### 4.3 关闭时
|
||||
|
||||
不输出 `proxy_intercept_errors`、`error_page`、内部 location 与对应 SupportFile(或文件可写但不被引用)。
|
||||
不输出 `proxy_intercept_errors`、`error_page`、内部 location 与对应 SupportFile(或文件可写但不被引用)。GET-only 模式同时不输出 Lua 过滤器。
|
||||
|
||||
### 4.4 与缓存 / stale
|
||||
|
||||
@@ -152,8 +156,7 @@ location = /__openflare_origin_error {
|
||||
|
||||
### 5.1 入口
|
||||
|
||||
* 侧栏「网站管理」新增:**错误页** → `/error-pages`
|
||||
* 更新 `openflareWebsiteNavGroup`、`openflareWebsiteSubNav`(若使用)、全局搜索关键词
|
||||
* 侧栏「网站管理 → 响应页面」:错误页 Tab(`/responses`),编辑页 `/responses/error-page/edit`、预览页 `/responses/error-page/preview`。
|
||||
|
||||
### 5.2 页面结构
|
||||
|
||||
@@ -165,14 +168,14 @@ location = /__openflare_origin_error {
|
||||
|
||||
### 5.3 组件依赖
|
||||
|
||||
若仓库尚无 Tags Input,按项目 shadcn 流程添加;样式与现有 UI 一致。
|
||||
Tags Input 与 HTML 编辑器复用现有 shadcn/ui 组件,样式与现有 UI 一致。
|
||||
|
||||
---
|
||||
|
||||
## 6. 数据流
|
||||
|
||||
```text
|
||||
管理员 /error-pages
|
||||
管理员 /responses(错误页 Tab)
|
||||
→ Option update-batch(校验标签与 HTML)
|
||||
→ w_system_configs
|
||||
|
||||
@@ -183,50 +186,13 @@ location = /__openflare_origin_error {
|
||||
|
||||
访客请求反代域名
|
||||
→ 源站/网关产生匹配状态码
|
||||
→ error_page → internal location
|
||||
→ error_page → 命名 location
|
||||
→ 替换占位符,status 保持原码,返回 HTML
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. 测试与验收
|
||||
|
||||
### 7.1 自动化
|
||||
|
||||
* 状态码解析:单码、区间、去重、越界、反序、默认 `500-599`
|
||||
* 渲染:enabled/disabled conf 片段;空 HTML 用默认;自定义进 SupportFile
|
||||
* Option 校验:非法标签 / 超大 HTML → 4xx
|
||||
|
||||
### 7.2 手动
|
||||
|
||||
1. 默认配置:源站不可达 → CF 风格页,真实 502/504,页内数字一致
|
||||
2. 源站返回 503 → 替换页,status 503
|
||||
3. 仅标签 `522` → 仅 522 替换
|
||||
4. 关闭开关并发布 → 透传恢复
|
||||
5. 自定义 HTML 占位符预览与线上一致
|
||||
6. Pages 路由不受影响
|
||||
|
||||
### 7.3 文档
|
||||
|
||||
* 本设计文档;`docs/design/index.md` 能力表;`docs/config.ts` 侧栏
|
||||
* changelog `[Unreleased]` 用户可读改进条
|
||||
|
||||
---
|
||||
|
||||
## 8. 实现要点清单(供计划拆分)
|
||||
|
||||
1. goose seed 三个 Option key + model 常量
|
||||
2. 状态码解析/校验纯函数 + 单测
|
||||
3. Option update 路径挂接校验
|
||||
4. 快照填充 `ConfigSnapshot` 新字段
|
||||
5. `pkg/render/openresty`:error_page 块、SupportFile、默认 HTML、单测
|
||||
6. Agent 侧若需 Lua 辅助文件,随现有 nginx lua 目录同步
|
||||
7. 前端 Tags Input + `/error-pages` 页 + 导航
|
||||
8. changelog 与设计索引
|
||||
|
||||
---
|
||||
|
||||
## 9. 决策记录
|
||||
## 7. 决策记录
|
||||
|
||||
| 决策 | 选择 | 原因 |
|
||||
| --- | --- | --- |
|
||||
@@ -236,4 +202,3 @@ location = /__openflare_origin_error {
|
||||
| 响应 status | 保持原码 | 监控/SEO/客户端语义正确 |
|
||||
| 运行时替换 | internal + 轻量模板替换 | 每请求 status 不同 |
|
||||
| 自定义方式 | 在线 HTML | 灵活且无需文件上传链路 |
|
||||
`}
|
||||
@@ -27,15 +27,13 @@ Pages 静态托管子系统包含以下核心能力:
|
||||
* **安全包校验与解压缩**:内置路径逃逸防御、防软链接劫持、文件大小/数量上限与可配置上传包体积控制,保障节点物理安全。
|
||||
* **可配置限额**:管理员可在运维设置中调整「部署包大小上限」与「历史部署保留数」。
|
||||
|
||||
### 部署源与未来构建边界
|
||||
### 部署源
|
||||
|
||||
项目当前支持 manual、Remote URL、GitHub Release 三种来源视图。无 source 记录即 manual;切换或删除 source 不删除历史 deployment,也不改变当前 active deployment。Remote URL 只允许手动“同步并发布”;GitHub Release 支持 latest/tag 手动检查与同步,只有 latest 可选择定时检查和自动更新。
|
||||
|
||||
source 是可变配置,deployment 是不可变事实。source 配置与运行态游标、状态、租约分别存储;deployment 只保存创建时的安全 provenance 快照。所有产物都复用“下载或接收产物 → 真实字节与入口校验 → `upload.Ingest` → deployment”的 artifact pipeline:manual 上传停在 candidate,等待管理员显式激活;持久来源 sync 才在同一业务事务中 create-or-load 并原子激活。Agent 只消费 active deployment,不感知来源类型。
|
||||
source 是可变配置,deployment 是不可变事实。source 配置与运行态游标、状态、租约分别存储;deployment 只保存创建时的安全 provenance 快照。所有产物都复用“下载或接收产物 → 真实字节与入口校验 → `upload.Ingest` → deployment”的 artifact pipeline:manual 上传停在 candidate,等待管理员显式激活;持久来源 sync 才在同一业务事务中 create-or-load 并原子激活。
|
||||
|
||||
后续从 Git 仓库拉取源码并自动构建时,将新增独立 `git_repository` provider 与隔离的 build executor。它输出受限的预构建产物后继续复用上述导入管线;不得把 clone、依赖安装或任意构建命令下发给 Agent,也不得把 branch/build/env 字段塞入现有 `github_release` source。当前 V2 不增加这些未来字段或空任务,只稳定 provider 输出、source discriminated view 与 deployment provenance 三个扩展边界。
|
||||
|
||||
管理端信息架构参考 Cloudflare Pages 当前把 [Git integration](https://developers.cloudflare.com/pages/configuration/git-integration/) 与 [Direct Upload](https://developers.cloudflare.com/pages/get-started/direct-upload/) 分离、并统一展示生产状态与历史部署的方式:OpenFlare 项目详情按“当前生产部署 → 部署源 → 部署历史”组织。OpenFlare 仍允许切换来源并保留历史部署,不采用 Cloudflare 项目创建后来源不可切换的限制。
|
||||
管理端项目详情按“当前生产部署 → 部署源 → 部署历史”组织。
|
||||
|
||||
---
|
||||
|
||||
@@ -67,7 +65,7 @@ graph TD
|
||||
```
|
||||
|
||||
* **控制面(Control Plane)**:Server 接收本地上传,或通过受限 Provider 获取 Remote/GitHub 预构建产物;action task 与内部 scanner 负责检查、同步和自动更新。所有产物经统一 inspect 与 `upload.Ingest` 写入平台存储后端;manual 上传创建新的 candidate,持久来源 sync 则 create-or-load deployment 并原子激活。配置发布时只编译稳定的项目锚点与静态服务元数据。
|
||||
* **数据面(Data Plane)**:Agent 在心跳/WS 对账中发现配置引用的 Pages 项目,通过专属 API 拉取该项目当前激活包并执行校验解压缩。OpenResty 在本地提供静态文件服务;Agent 不感知产物来自上传、Remote、GitHub 或未来 build executor。
|
||||
* **数据面(Data Plane)**:Agent 在心跳/WS 对账中发现配置引用的 Pages 项目,通过专属 API 拉取该项目当前激活包并执行校验解压缩。OpenResty 在本地提供静态文件服务;Agent 不感知产物来自上传、Remote 或 GitHub。
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -116,14 +116,14 @@ Kill 并重新拉起 frps 进程
|
||||
server_name intranet.example.com;
|
||||
# ... TLS 证书与 WAF 过滤逻辑 ...
|
||||
location / {
|
||||
proxy_pass http://127.0.0.1:18080; # 指向本地 frps 的虚拟主机端口
|
||||
proxy_pass http://127.0.0.1:8080; # 指向本地 frps 的虚拟主机端口
|
||||
proxy_set_header Host $host; # 必须保留原 Host,因为 frps 依靠 Host 进行内部路由分发
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
}
|
||||
}
|
||||
```
|
||||
2. **中继节点 (frps)**:
|
||||
`frps` 监听到 `18080` 端口有 HTTP 请求进来,读取 HTTP 请求头中的 `Host: intranet.example.com`,在其已注册的活跃隧道表中检索该域名对应的加密 TCP 连接(由内网 frpc 建立)。
|
||||
`frps` 在虚拟主机端口(默认 `8080`)收到 HTTP 请求,读取 HTTP 请求头中的 `Host: intranet.example.com`,在其已注册的活跃隧道表中检索该域名对应的加密 TCP 连接(由内网 frpc 建立)。
|
||||
3. **加密隧道传输 (TCP)**:
|
||||
`frps` 将 HTTP 请求封装进内部 TCP 隧道协议,发送给内网的 `frpc` 客户端。
|
||||
4. **内网客户端分发 (frpc)**:
|
||||
|
||||
@@ -8,7 +8,7 @@ Server 保存带坐标和修订号的编辑图,发布时再次校验并编译
|
||||
|
||||
IP 组独立于规则拓扑更新。手动、订阅和自动 IP 组由控制面维护,Agent 先原子替换 JSON、最后更新 checksum。协调 Worker 每 5 秒检查 checksum,仅变化时读取完整快照并分发给其它 Worker;失败时保留上一份有效数据。完整运行时快照上限为 20 MiB,Server 发布/同步与 Agent 落盘使用同一序列化校验;OpenResty 使用独立的 64 MiB 共享字典和非淘汰写入,容量不足时拒绝新版本而不破坏已提交快照。
|
||||
|
||||
地域节点使用 Country 与 City MMDB。Agent 首次启动时从程序内嵌数据库初始化缺失文件,后续按配置周期下载更新,请求处理始终读取 OpenResty 已加载的数据库。数据库不可用时地域匹配返回 `false` 并限频告警,不允许因数据损坏意外放行其它执行错误。
|
||||
地域节点使用 Country 与 City MMDB。Docker 镜像内置数据库文件,裸二进制安装由 Agent 首次启动时下载缺失文件,后续按配置周期更新,请求处理始终读取 OpenResty 已加载的数据库。数据库不可用时地域匹配返回 `false` 并限频告警,不允许因数据损坏意外放行其它执行错误。
|
||||
|
||||
## 安全顺序
|
||||
|
||||
|
||||
@@ -6,7 +6,7 @@
|
||||
|
||||
用户新增 WAF 规则时只输入名称。Server 随即创建一张合法的默认图 `开始 → 通过`,前端进入基于 React Flow 的独立编排页面。用户通过添加处理单元、配置节点并连接分支构建策略,不再填写固定顺序的黑白名单与 PoW 表单。
|
||||
|
||||
第一阶段支持以下节点:
|
||||
支持的节点:
|
||||
|
||||
| 节点 | 数量约束 | 输入 | 输出 | 配置 |
|
||||
| --- | --- | --- | --- | --- |
|
||||
@@ -21,7 +21,7 @@
|
||||
|
||||
IP 匹配、地域匹配、UA 检查与安全防护不区分黑名单或白名单。`true` 只表示请求通过该节点判定,`false` 只表示未通过;放行或阻止的业务含义完全由连线决定。UA 检查的求值顺序为:要求携带 UA → 屏蔽爬虫/非正常 UA → 白名单匹配。安全防护在请求 Path/Query/Header/Cookie/Body(有限)上做特征匹配。PoW 验证完成后沿 `next` 继续,未完成时由挑战页面接管当前请求,不产生 `false` 分支。
|
||||
|
||||
不在第一阶段实现循环、脚本节点、任意表达式节点、子图调用和跨规则跳转。
|
||||
不实现循环、脚本节点、任意表达式节点、子图调用和跨规则跳转。
|
||||
|
||||
## 控制面架构
|
||||
|
||||
@@ -116,13 +116,3 @@ WAF 列表展示规则名称、启用状态、节点数量、应用路由数量
|
||||
新 Worker 只接受完整且可解析的规则运行态配置。旧 Worker 在 OpenResty 优雅 reload 期间继续使用旧内存图,新 Worker 使用新图,因此请求不会观察到半更新状态。
|
||||
|
||||
地域数据库不可用时,地域匹配返回 `false` 并限频告警,保持现有行为。IP 组刷新失败时保留旧内存快照。PoW 未完成由挑战模块接管请求,不视为执行错误;PoW 节点配置先以短期键写入 OpenResty 共享内存,再通过 `ngx.exec` 的显式参数传给内部挑战处理器,不能依赖内部重定向保留 `ngx.ctx` 或隐式继承请求参数。发布快照中的空规则绑定必须编码为 JSON 空数组;运行时将旧快照中的 `null` 可选数组按空数组处理,禁止因 `cjson` 的 `ngx.null` userdata 中断请求。
|
||||
|
||||
## 测试与验收
|
||||
|
||||
* Go 单元测试覆盖图结构、端口、可达性、终止性、节点配置、体积限制、编译结果、修订冲突和绑定顺序。
|
||||
* 数据库测试覆盖 PostgreSQL/SQLite 迁移、默认图、旧绑定稳定排序和回滚。
|
||||
* Lua 测试覆盖所有节点出口、多规则顺序、全局规则前置、多个阻止响应、PoW 接管和损坏运行时图保护。
|
||||
* Agent/OpenResty 测试覆盖发布 reload、加载一次、失败回滚、IP 组五秒 checksum 刷新和旧快照保留。
|
||||
* 前端测试覆盖创建后导航、特殊节点唯一性、连线限制、属性编辑、即时校验、未保存提示和并发冲突。
|
||||
* 集成测试从控制面创建并编排规则,发布后用真实请求验证放行、阻止、PoW 和 IP 组热刷新。
|
||||
* API 变更后运行 `make swagger`;完成实现后运行前端检查与构建以及 `make code-check`。
|
||||
|
||||
@@ -54,7 +54,7 @@ erDiagram
|
||||
* `GET/POST /api/v1/d/zones`
|
||||
* `GET/POST /api/v1/d/zones/:id/update`
|
||||
* `POST /api/v1/d/zones/:id/delete`
|
||||
* `GET/POST /api/v1/d/zones/:id/domains`
|
||||
* `POST /api/v1/d/zones/:id/domains`(列表经 overview 返回)
|
||||
* `POST /api/v1/d/zones/:id/domains/:domainID/update`
|
||||
* `POST /api/v1/d/zones/:id/domains/:domainID/delete`
|
||||
* `GET /api/v1/d/zones/:id/overview`
|
||||
@@ -91,10 +91,3 @@ WAF、Pages、上游与发布版本仍属于 `proxy_routes`。Zone 概览只聚
|
||||
* 持久化:域名与证书只存在于 `of_zone_domains`;`of_proxy_routes` 仅保存路由策略(上游、缓存、限流、WAF 绑定键等)。
|
||||
* 渲染:配置快照在内存中组装临时 `Domains` / `DomainCertIDs` 供 OpenResty 渲染,不写回数据库。
|
||||
* 结构迁移仅使用 `internal/infra/persistence/migrator/goose/{postgres,sqlite}/*.sql`;启动时自动导入历史域名,第二阶段后旧列不存在则为空操作。
|
||||
|
||||
## 验证
|
||||
|
||||
* 单元测试:Public Suffix List 分组、FQDN / 通配符拒绝、跨 Zone 路由、证书 SAN 覆盖、删除保护及迁移幂等性;清理后断言旧列/旧表不存在。
|
||||
* 集成测试:Zone、Zone 域名与路由 API 的成功与失败响应;现有路由迁移后生成相同 OpenResty 域名与证书配置。
|
||||
* 前端测试:Zone 列表、ID 路由、详情加载 / 错误 / 空状态、域名选择器与 API 负载。
|
||||
* 手动验证:迁移前后比较激活配置快照中的 `server_name` 和证书路径,发布后使用根域及各子域请求验证路由。
|
||||
|
||||
+247
-3
@@ -4809,7 +4809,7 @@ const docTemplate = `{
|
||||
"SessionCookie": []
|
||||
}
|
||||
],
|
||||
"description": "分页返回 OpenFlare 访问日志,支持按节点、IP、主机与路径筛选,需要管理员权限",
|
||||
"description": "分页返回 OpenFlare 访问日志,支持按节点、IP、主机、路径与状态码筛选,需要管理员权限",
|
||||
"produces": [
|
||||
"application/json"
|
||||
],
|
||||
@@ -4842,6 +4842,24 @@ const docTemplate = `{
|
||||
"name": "path",
|
||||
"in": "query"
|
||||
},
|
||||
{
|
||||
"type": "integer",
|
||||
"description": "HTTP 状态码(100-599)",
|
||||
"name": "status_code",
|
||||
"in": "query"
|
||||
},
|
||||
{
|
||||
"type": "string",
|
||||
"description": "起始时间(RFC3339,需与 until 成对提供)",
|
||||
"name": "since",
|
||||
"in": "query"
|
||||
},
|
||||
{
|
||||
"type": "string",
|
||||
"description": "结束时间(RFC3339,需与 since 成对提供)",
|
||||
"name": "until",
|
||||
"in": "query"
|
||||
},
|
||||
{
|
||||
"type": "integer",
|
||||
"description": "页码",
|
||||
@@ -6344,6 +6362,190 @@ const docTemplate = `{
|
||||
}
|
||||
}
|
||||
},
|
||||
"/api/v1/d/cloudflare/groups/{id}/members/batch-move": {
|
||||
"post": {
|
||||
"security": [
|
||||
{
|
||||
"SessionCookie": []
|
||||
}
|
||||
],
|
||||
"consumes": [
|
||||
"application/json"
|
||||
],
|
||||
"produces": [
|
||||
"application/json"
|
||||
],
|
||||
"tags": [
|
||||
"openflare-cloudflare"
|
||||
],
|
||||
"summary": "批量移动 Cloudflare 指向成员",
|
||||
"parameters": [
|
||||
{
|
||||
"type": "integer",
|
||||
"description": "原分组 ID",
|
||||
"name": "id",
|
||||
"in": "path",
|
||||
"required": true
|
||||
},
|
||||
{
|
||||
"description": "批量移动参数",
|
||||
"name": "body",
|
||||
"in": "body",
|
||||
"required": true,
|
||||
"schema": {
|
||||
"$ref": "#/definitions/cloudflare.MemberBatchMoveInput"
|
||||
}
|
||||
}
|
||||
],
|
||||
"responses": {
|
||||
"200": {
|
||||
"description": "OK",
|
||||
"schema": {
|
||||
"$ref": "#/definitions/response.Any"
|
||||
}
|
||||
},
|
||||
"400": {
|
||||
"description": "Bad Request",
|
||||
"schema": {
|
||||
"$ref": "#/definitions/response.Any"
|
||||
}
|
||||
},
|
||||
"404": {
|
||||
"description": "Not Found",
|
||||
"schema": {
|
||||
"$ref": "#/definitions/response.Any"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"/api/v1/d/cloudflare/groups/{id}/members/batch-remove": {
|
||||
"post": {
|
||||
"security": [
|
||||
{
|
||||
"SessionCookie": []
|
||||
}
|
||||
],
|
||||
"consumes": [
|
||||
"application/json"
|
||||
],
|
||||
"produces": [
|
||||
"application/json"
|
||||
],
|
||||
"tags": [
|
||||
"openflare-cloudflare"
|
||||
],
|
||||
"summary": "批量移出 Cloudflare 指向成员",
|
||||
"parameters": [
|
||||
{
|
||||
"type": "integer",
|
||||
"description": "分组 ID",
|
||||
"name": "id",
|
||||
"in": "path",
|
||||
"required": true
|
||||
},
|
||||
{
|
||||
"description": "批量移出参数",
|
||||
"name": "body",
|
||||
"in": "body",
|
||||
"required": true,
|
||||
"schema": {
|
||||
"$ref": "#/definitions/cloudflare.MemberBatchRemoveInput"
|
||||
}
|
||||
}
|
||||
],
|
||||
"responses": {
|
||||
"200": {
|
||||
"description": "OK",
|
||||
"schema": {
|
||||
"$ref": "#/definitions/response.Any"
|
||||
}
|
||||
},
|
||||
"400": {
|
||||
"description": "Bad Request",
|
||||
"schema": {
|
||||
"$ref": "#/definitions/response.Any"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"/api/v1/d/cloudflare/groups/{id}/members/{memberId}/move": {
|
||||
"post": {
|
||||
"security": [
|
||||
{
|
||||
"SessionCookie": []
|
||||
}
|
||||
],
|
||||
"consumes": [
|
||||
"application/json"
|
||||
],
|
||||
"produces": [
|
||||
"application/json"
|
||||
],
|
||||
"tags": [
|
||||
"openflare-cloudflare"
|
||||
],
|
||||
"summary": "移动 Cloudflare 指向成员到其他分组",
|
||||
"parameters": [
|
||||
{
|
||||
"type": "integer",
|
||||
"description": "原分组 ID",
|
||||
"name": "id",
|
||||
"in": "path",
|
||||
"required": true
|
||||
},
|
||||
{
|
||||
"type": "integer",
|
||||
"description": "成员 ID",
|
||||
"name": "memberId",
|
||||
"in": "path",
|
||||
"required": true
|
||||
},
|
||||
{
|
||||
"description": "目标分组参数",
|
||||
"name": "body",
|
||||
"in": "body",
|
||||
"required": true,
|
||||
"schema": {
|
||||
"$ref": "#/definitions/cloudflare.MemberMoveInput"
|
||||
}
|
||||
}
|
||||
],
|
||||
"responses": {
|
||||
"200": {
|
||||
"description": "OK",
|
||||
"schema": {
|
||||
"allOf": [
|
||||
{
|
||||
"$ref": "#/definitions/response.Any"
|
||||
},
|
||||
{
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"data": {
|
||||
"$ref": "#/definitions/cloudflare.MemberItem"
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
},
|
||||
"400": {
|
||||
"description": "Bad Request",
|
||||
"schema": {
|
||||
"$ref": "#/definitions/response.Any"
|
||||
}
|
||||
},
|
||||
"404": {
|
||||
"description": "Not Found",
|
||||
"schema": {
|
||||
"$ref": "#/definitions/response.Any"
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"/api/v1/d/cloudflare/groups/{id}/members/{memberId}/remove": {
|
||||
"post": {
|
||||
"security": [
|
||||
@@ -14242,7 +14444,7 @@ const docTemplate = `{
|
||||
}
|
||||
},
|
||||
"400": {
|
||||
"description": "用户名或密码错误、帐号已禁用等",
|
||||
"description": "用户名或密码错误",
|
||||
"schema": {
|
||||
"$ref": "#/definitions/response.Any"
|
||||
}
|
||||
@@ -15121,6 +15323,31 @@ const docTemplate = `{
|
||||
}
|
||||
}
|
||||
},
|
||||
"cloudflare.MemberBatchMoveInput": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"member_ids": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "integer"
|
||||
}
|
||||
},
|
||||
"target_group_id": {
|
||||
"type": "integer"
|
||||
}
|
||||
}
|
||||
},
|
||||
"cloudflare.MemberBatchRemoveInput": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"member_ids": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "integer"
|
||||
}
|
||||
}
|
||||
}
|
||||
},
|
||||
"cloudflare.MemberCreateInput": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
@@ -15167,6 +15394,14 @@ const docTemplate = `{
|
||||
}
|
||||
}
|
||||
},
|
||||
"cloudflare.MemberMoveInput": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"target_group_id": {
|
||||
"type": "integer"
|
||||
}
|
||||
}
|
||||
},
|
||||
"cloudflare.MemberUpdateInput": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
@@ -15564,7 +15799,7 @@ const docTemplate = `{
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"additionalProperties": true
|
||||
"additionalProperties": {}
|
||||
}
|
||||
},
|
||||
"type": {
|
||||
@@ -18385,6 +18620,15 @@ const docTemplate = `{
|
||||
"request_count": {
|
||||
"type": "integer"
|
||||
},
|
||||
"status_2xx_count": {
|
||||
"type": "integer"
|
||||
},
|
||||
"status_4xx_count": {
|
||||
"type": "integer"
|
||||
},
|
||||
"status_5xx_count": {
|
||||
"type": "integer"
|
||||
},
|
||||
"unique_visitor_count": {
|
||||
"type": "integer"
|
||||
}
|
||||
|
||||
@@ -0,0 +1,5 @@
|
||||
# Changelog
|
||||
|
||||
The changelog is maintained in Simplified Chinese in this repository. See the Chinese version:
|
||||
|
||||
* [更新日志(简体中文)](../changelog/)
|
||||
@@ -0,0 +1,147 @@
|
||||
import { defineAdditionalConfig, type DefaultTheme } from 'vitepress'
|
||||
|
||||
export default defineAdditionalConfig({
|
||||
description:
|
||||
'OpenFlare is a lightweight, self-hosted OpenResty control plane for managing reverse proxy rules, configuration publishing, node synchronization, TLS certificates, and basic observability.',
|
||||
|
||||
themeConfig: {
|
||||
nav: nav(),
|
||||
|
||||
sidebar: {
|
||||
'/en/guide/': { base: '/en/guide/', items: sidebarGuide() },
|
||||
'/en/reference/': { base: '/en/reference/', items: sidebarReference() },
|
||||
'/en/deployment/': { base: '/en/deployment/', items: sidebarDeployment() },
|
||||
'/en/design/': { base: '/en/design/', items: sidebarDesign() },
|
||||
'/en/changelog/': { base: '/en/changelog/', items: [] }
|
||||
},
|
||||
|
||||
editLink: {
|
||||
pattern: 'https://github.com/Rain-kl/OpenFlare/edit/main/docs/:path',
|
||||
text: 'Edit this page on GitHub'
|
||||
},
|
||||
|
||||
footer: {
|
||||
message: 'Released under the Apache License 2.0',
|
||||
copyright: 'Copyright © OpenFlare contributors'
|
||||
},
|
||||
|
||||
docFooter: {
|
||||
prev: 'Previous Page',
|
||||
next: 'Next Page'
|
||||
},
|
||||
|
||||
outline: {
|
||||
label: 'On this page'
|
||||
},
|
||||
|
||||
lastUpdated: {
|
||||
text: 'Last updated at'
|
||||
},
|
||||
|
||||
notFound: {
|
||||
title: 'Page Not Found',
|
||||
quote: 'This document does not have a corresponding page yet.',
|
||||
linkLabel: 'Go to Home',
|
||||
linkText: 'Back to OpenFlare Docs'
|
||||
},
|
||||
|
||||
langMenuLabel: 'Language',
|
||||
returnToTopLabel: 'Back to top',
|
||||
sidebarMenuLabel: 'Menu',
|
||||
darkModeSwitchLabel: 'Theme',
|
||||
lightModeSwitchTitle: 'Switch to light theme',
|
||||
darkModeSwitchTitle: 'Switch to dark theme',
|
||||
skipToContentLabel: 'Skip to content'
|
||||
}
|
||||
})
|
||||
|
||||
function nav(): DefaultTheme.NavItem[] {
|
||||
return [
|
||||
{ text: 'Guide', link: '/en/guide/', activeMatch: '/en/guide/' },
|
||||
{ text: 'Deployment', link: '/en/deployment/', activeMatch: '/en/deployment/' },
|
||||
{ text: 'Reference', link: '/en/reference/', activeMatch: '/en/reference/' },
|
||||
{ text: 'Design', link: '/en/design/', activeMatch: '/en/design/' },
|
||||
{ text: 'Changelog', link: '/en/changelog/', activeMatch: '/en/changelog/' }
|
||||
]
|
||||
}
|
||||
|
||||
function sidebarGuide(): DefaultTheme.SidebarItem[] {
|
||||
return [
|
||||
{
|
||||
text: 'Guide',
|
||||
items: [
|
||||
{ text: 'Overview', link: '' },
|
||||
{ text: 'Quick Start', link: 'quick-start' },
|
||||
{ text: 'TLS Certificates & Auto-Renewal', link: 'certificates' },
|
||||
{ text: 'Zone Domain Migration', link: 'zone-domain-migration' },
|
||||
{ text: 'Create a Reverse Proxy Config', link: 'proxy-config' },
|
||||
{ text: 'Pages Static Hosting Usage', link: 'pages-usage' },
|
||||
{ text: 'Tunnel & Intranet Penetration', link: 'tunnel-usage' },
|
||||
{ text: 'WAF Security Protection', link: 'waf-usage' },
|
||||
{ text: 'WAF Auto IP Group Expressions', link: 'waf-ip-group-expr' },
|
||||
{ text: 'Uptime Kuma Monitoring Sync', link: 'uptime-kuma' },
|
||||
{ text: 'SSO Login Configuration', link: 'sso' },
|
||||
{ text: 'Publish First Configuration', link: 'first-site' },
|
||||
{ text: 'Troubleshooting', link: 'troubleshooting' },
|
||||
{ text: 'Credits', link: 'credits' }
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
function sidebarReference(): DefaultTheme.SidebarItem[] {
|
||||
return [
|
||||
{
|
||||
text: 'Reference',
|
||||
items: [
|
||||
{ text: 'Overview', link: '' },
|
||||
{ text: 'Configuration Options', link: 'configuration' },
|
||||
{ text: 'CLI Commands', link: 'cli' }
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
function sidebarDeployment(): DefaultTheme.SidebarItem[] {
|
||||
return [
|
||||
{
|
||||
text: 'Deployment',
|
||||
items: [
|
||||
{ text: 'Overview', link: '' },
|
||||
{ text: 'Deployment Guide', link: 'deployment' },
|
||||
{ text: 'Start the Server', link: 'server' },
|
||||
{ text: 'Access Agent', link: 'agent' },
|
||||
{ text: 'Deploy Relay (Tunnel)', link: 'relay' },
|
||||
{ text: 'Deploy OpenFlared', link: 'openflared' },
|
||||
{ text: 'Upgrade & Maintenance', link: 'upgrade' }
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
|
||||
function sidebarDesign(): DefaultTheme.SidebarItem[] {
|
||||
return [
|
||||
{
|
||||
text: 'Design',
|
||||
items: [
|
||||
{ text: 'Product Boundaries', link: '' },
|
||||
{ text: 'System Architecture', link: 'architecture' },
|
||||
{ text: 'Zone & Domain Resource Design', link: 'zone-design' },
|
||||
{ text: 'Cloudflare DNS Pointing Design', link: 'cloudflare-pointing' },
|
||||
{ text: 'Agent & Publish Model', link: 'agent-design' },
|
||||
{ text: 'Tunnel & Intranet Penetration', link: 'tunnel-design' },
|
||||
{ text: 'WAF Design', link: 'waf-design' },
|
||||
{ text: 'WAF Orchestration Rule Design', link: 'waf-orchestration-design' },
|
||||
{ text: 'Pages Static Hosting Design', link: 'pages-design' },
|
||||
{ text: 'Edge Cache Strategy Design', link: 'edge-cache-design' },
|
||||
{ text: 'Origin Error Page Design', link: 'origin-error-page' },
|
||||
{ text: 'Edge Observability & Traffic Stats', link: 'observability-design' },
|
||||
{ text: 'Observability Transport Model', link: 'observability-transport-model' },
|
||||
{ text: 'Observability Protocol & Tables', link: 'observability-data-model' },
|
||||
{ text: 'Log Store Decoupling', link: 'logstore' },
|
||||
{ text: 'Uptime Kuma Sync Design', link: 'kuma-design' },
|
||||
{ text: 'Login CAPTCHA Design', link: 'login-captcha' }
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,159 @@
|
||||
# Access Agent
|
||||
|
||||
You will learn: the Agent's responsibilities, the difference between the two access Tokens, install script parameters, `agent.json` config, and how to confirm the node is online.
|
||||
|
||||
The OpenFlare Agent runs on the proxy node side. It doesn't accept remote shell commands; instead it pulls released config versions from the control plane via the Agent API, writes OpenResty files locally, runs config validation, reloads, and attempts to roll back to a runnable config on failure.
|
||||
|
||||
## Connection Methods
|
||||
|
||||
| Method | Use Case |
|
||||
| --- | --- |
|
||||
| `discovery_token` | first-time auto-registration; the Server exchanges it for a node-specific credential |
|
||||
| `agent_token` | node already created/assigned in the admin panel; connect with the node-specific credential |
|
||||
|
||||
At least one of `agent_token` / `discovery_token` is required.
|
||||
|
||||
### Credential Paths
|
||||
|
||||
- **`discovery_token` (auto-registration credential)**: log in to the admin panel, navigate to「System Settings」->「Auto Registration」; generate, view, and copy the global auto-registration credential there.
|
||||
- **`agent_token` (node-specific credential)**: log in to the admin panel, navigate to「Node Management」->「Add Node」; after filling in basic info and saving, copy the node-specific access Token on the node detail page.
|
||||
|
||||
## One-Click Install
|
||||
|
||||
### Interactive Install (recommended)
|
||||
|
||||
Running the install script without any arguments enters interactive mode, with a wizard choosing the install method (local / Docker container) and configuring the Server address and auth Token (if Docker is chosen and Docker isn't installed locally, the script asks and intelligently installs Docker):
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash
|
||||
```
|
||||
|
||||
### Automated (non-interactive) Install
|
||||
|
||||
Adding any arguments enters automated install mode with no interaction.
|
||||
|
||||
Local install with `discovery_token`:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--discovery-token YOUR_DISCOVERY_TOKEN
|
||||
```
|
||||
|
||||
Local install with node-specific `agent_token`:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--agent-token YOUR_AGENT_TOKEN
|
||||
```
|
||||
|
||||
Automated Docker container install:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--discovery-token YOUR_DISCOVERY_TOKEN \
|
||||
--docker
|
||||
```
|
||||
|
||||
In local install mode, the script downloads the latest Agent, writes to `/opt/openflare-agent` by default, generates `agent.json`, auto-detects and creates the low-privilege system account `openflare` (granting the whole install dir to it), and creates the `openflare-agent.service` systemd service on Linux + systemd. The service runs as the `openflare` unprivileged user, with Linux Capabilities (`CAP_NET_BIND_SERVICE`) enabling privileged ports (e.g. 80, 443).
|
||||
|
||||
Supported parameters:
|
||||
|
||||
| Parameter | Description |
|
||||
| --- | --- |
|
||||
| `--server-url` | Server address |
|
||||
| `--discovery-token` | first-time auto-registration Token |
|
||||
| `--agent-token` | node-specific Token |
|
||||
| `--install-dir` | install dir, default `/opt/openflare-agent` (local install only) |
|
||||
| `--openresty-path` | OpenResty binary path; auto-finds `openresty` when omitted (local install only) |
|
||||
| `--repo` | GitHub repo for downloading the Agent, default `Rain-kl/OpenFlare` |
|
||||
| `--no-service` | don't create the systemd service (local install only) |
|
||||
| `--docker` | install via Docker container |
|
||||
| `--method` | install method: `local` or `docker` (default `local`) |
|
||||
|
||||
## Config File
|
||||
|
||||
Default config file path:
|
||||
|
||||
```text
|
||||
/opt/openflare-agent/agent.json
|
||||
```
|
||||
|
||||
Local config example:
|
||||
|
||||
```json
|
||||
{
|
||||
"server_url": "http://127.0.0.1:3000",
|
||||
"agent_token": "replace-with-node-auth-token",
|
||||
"data_dir": "./data",
|
||||
"openresty_path": "openresty",
|
||||
"openresty_observability_port": 18081,
|
||||
"observability_replay_minutes": 60,
|
||||
"heartbeat_interval": 3000,
|
||||
"request_timeout": 10000
|
||||
}
|
||||
```
|
||||
|
||||
Custom OpenResty path example:
|
||||
|
||||
```json
|
||||
{
|
||||
"server_url": "http://127.0.0.1:3000",
|
||||
"agent_token": "replace-with-node-auth-token",
|
||||
"data_dir": "/var/lib/openflare-agent",
|
||||
"openresty_path": "/usr/local/openresty/nginx/sbin/openresty",
|
||||
"main_config_path": "/var/lib/openflare-agent/etc/nginx/nginx.conf",
|
||||
"route_config_path": "/var/lib/openflare-agent/etc/nginx/conf.d/openflare_routes.conf",
|
||||
"access_log_path": "/var/lib/openflare-agent/var/log/openflare/access.log",
|
||||
"cert_dir": "/var/lib/openflare-agent/etc/nginx/certs",
|
||||
"lua_dir": "/var/lib/openflare-agent/etc/nginx/lua",
|
||||
"runtime_config_dir": "/var/lib/openflare-agent/etc/openflare",
|
||||
"heartbeat_interval": 3000,
|
||||
"request_timeout": 10000
|
||||
}
|
||||
```
|
||||
|
||||
Without `openresty_path`, the Agent calls `openresty` by default. Full fields: [Configuration Reference](../reference/configuration.md#agent-命令行参数与配置字段).
|
||||
|
||||
## Running with Docker
|
||||
|
||||
For Docker deployment, directly run the Agent image with a built-in OpenResty:
|
||||
|
||||
```bash
|
||||
docker pull ghcr.io/rain-kl/openflare-agent:latest
|
||||
docker rm -f openflare-agent 2>/dev/null || true
|
||||
docker run -d --name openflare-agent --restart unless-stopped \
|
||||
-p 80:80 -p 443:443/tcp -p 443:443/udp \
|
||||
-v openflare-agent-pages:/data/var/lib/openflare/pages \
|
||||
-e OPENFLARE_SERVER_URL=http://your-server:3000 \
|
||||
-e OPENFLARE_AGENT_TOKEN=YOUR_AGENT_TOKEN \
|
||||
ghcr.io/rain-kl/openflare-agent:latest
|
||||
```
|
||||
|
||||
> [!NOTE]
|
||||
> **Pages persistence**
|
||||
> By default the Pages deployment dir is mounted to the Docker named volume `openflare-agent-pages` (container path `/data/var/lib/openflare/pages`). Rebuilding or upgrading the Agent container doesn't require re-pulling static site packages.
|
||||
|
||||
## Uninstall
|
||||
|
||||
### Interactive Uninstall (recommended)
|
||||
|
||||
Running the uninstall script without any arguments enters interactive mode with a menu choosing the method (local uninstall / Docker container uninstall):
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/uninstall-agent.sh | bash
|
||||
```
|
||||
|
||||
### Docker Container Uninstall
|
||||
|
||||
Stop and remove the `openflare-agent` container.
|
||||
|
||||
## FAQ
|
||||
|
||||
| Symptom | Handling |
|
||||
| --- | --- |
|
||||
| `agent_token and discovery_token cannot both be empty` | check that `agent.json` has at least one Token |
|
||||
| Node stays offline | run `curl -I http://your-server:3000` on the Agent node to confirm the Server address is reachable |
|
||||
| Repeated failure after release | the Agent blocks re-applying the same `version + checksum`; click「Force Sync」in the node detail page, or republish a new version |
|
||||
@@ -0,0 +1,151 @@
|
||||
# Deployment Guide
|
||||
|
||||
You will learn: OpenFlare's recommended deployment approaches, Server and Agent runtime requirements, source startup, integration steps, and upgrade/uninstall entries.
|
||||
|
||||
Production recommends PostgreSQL as the Server DB, with `APP_SESSION_SECRET` etc. configured via `config.yaml` or env vars. The full Docker Compose deployment requires Redis; ClickHouse is optional for massive access logs and observability time series (see the repo root `docker-compose.yaml`). The Agent supports both Docker deployment and a local install script; the Docker image bundles the OpenResty binary. Log-DB determination and switching: [Log Store Decoupling](../design/logstore.md).
|
||||
|
||||
## Deployment Topology
|
||||
|
||||
### Standard Reverse Proxy Traffic Path
|
||||
|
||||
```text
|
||||
Browser
|
||||
|
|
||||
v
|
||||
OpenFlare Server :3000
|
||||
|
|
||||
| Agent API / heartbeat / config pull
|
||||
v
|
||||
OpenFlare Agent
|
||||
|
|
||||
v
|
||||
OpenResty binary
|
||||
|
|
||||
v
|
||||
Origin service
|
||||
```
|
||||
|
||||
### Intranet Penetration Traffic Path
|
||||
|
||||
```text
|
||||
Browser
|
||||
|
|
||||
v
|
||||
OpenResty (Agent, WAF/HTTPS termination) <-- TunnelRelay node
|
||||
|
|
||||
| proxy_pass (127.0.0.1:{vhost_port})
|
||||
v
|
||||
OpenFlareRelay (frps process) <-- TunnelRelay node
|
||||
|
|
||||
| frp tunnel protocol
|
||||
v
|
||||
OpenFlared (frpc client) <-- intranet server
|
||||
|
|
||||
v
|
||||
Internal Service (192.168.x.x)
|
||||
```
|
||||
|
||||
## Prerequisites
|
||||
|
||||
### Hardware Recommendations
|
||||
|
||||
| Component | Reference (entry) | Reference (production) | Notes |
|
||||
| --- | --- | --- | --- |
|
||||
| **Server control plane** | 1 core / 2 GB RAM / 20 GB disk | 2 cores / 4 GB RAM / 50 GB+ disk | expand disk by access-log retention and concurrent traffic |
|
||||
| **Agent data plane** | 1 core / 512 MB RAM / 2 GB disk | 2 cores / 2 GB RAM / 10 GB+ disk | expand by OpenResty concurrent proxy connections and WAF interception |
|
||||
| **Relay node** | 1 core / 1 GB RAM / 5 GB disk | 2 cores / 2 GB RAM / 20 GB disk | frps relay throughput limited by bandwidth and CPU |
|
||||
| **OpenFlared client** | 1 core / 256 MB RAM / 1 GB disk | 1 core / 512 MB RAM / 5 GB disk | runs independently in the intranet, tiny footprint |
|
||||
|
||||
## Docker Compose Server Deployment
|
||||
|
||||
The repo root provides a full `docker-compose.yaml` (PostgreSQL, Redis, ClickHouse, Jaeger).
|
||||
|
||||
```bash
|
||||
curl -o .env.example https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/.env.example
|
||||
cp .env.example .env
|
||||
# edit .env; at minimum change APP_SESSION_SECRET and the DB passwords
|
||||
docker compose up -d
|
||||
docker compose ps
|
||||
docker compose logs -f openflare
|
||||
```
|
||||
|
||||
First visit `http://localhost:3000`; default account `admin` / `12345678`. Change the default password immediately after login.
|
||||
|
||||
## Source Startup
|
||||
|
||||
First build the admin frontend:
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
corepack enable
|
||||
pnpm install
|
||||
pnpm build:embed
|
||||
```
|
||||
|
||||
Then start the Server (repo root):
|
||||
|
||||
```bash
|
||||
cp config.example.yaml config.yaml
|
||||
export APP_SESSION_SECRET='replace-with-a-long-random-string'
|
||||
# optional: use PostgreSQL
|
||||
# export DB_HOST=127.0.0.1 DB_USERNAME=postgres DB_PASSWORD=postgres DB_NAME=openflare
|
||||
go run main.go all
|
||||
```
|
||||
|
||||
Listens on `:3000` by default (controlled by `app.addr` in `config.yaml` or `APP_ADDR`).
|
||||
|
||||
## Run the Agent with Docker (recommended)
|
||||
|
||||
Docker is the recommended Agent deployment. The Agent image is built on the OpenResty image, bundling the Agent controller and the OpenResty binary. Without an explicit `node_ip`, the Agent prefers fetching the real egress IP via a third-party API, avoiding registering the Docker bridge address as the node IP.
|
||||
|
||||
```bash
|
||||
docker pull ghcr.io/rain-kl/openflare-agent:latest
|
||||
docker rm -f openflare-agent 2>/dev/null || true
|
||||
docker run -d --name openflare-agent --restart unless-stopped \
|
||||
-p 80:80 -p 443:443/tcp -p 443:443/udp \
|
||||
-v openflare-agent-pages:/data/var/lib/openflare/pages \
|
||||
-e OPENFLARE_SERVER_URL=http://your-server:3000 \
|
||||
-e OPENFLARE_AGENT_TOKEN=YOUR_AGENT_TOKEN \
|
||||
ghcr.io/rain-kl/openflare-agent:latest
|
||||
```
|
||||
|
||||
The named volume `openflare-agent-pages` persists the Pages deployment dir; rebuilding the container doesn't require re-pulling static site packages.
|
||||
|
||||
## Agent Access (script install)
|
||||
|
||||
Besides Docker, the install script can deploy the Agent to the local host. The script registers the low-privilege `openflare` service account and runs the systemd service as that user, using Linux Capabilities to safely listen on privileged ports 80/443.
|
||||
|
||||
Auto-register with `discovery_token`:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--discovery-token YOUR_DISCOVERY_TOKEN
|
||||
```
|
||||
|
||||
With node-specific `agent_token`:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--agent-token YOUR_AGENT_TOKEN
|
||||
```
|
||||
|
||||
Install script parameters:
|
||||
|
||||
| Parameter | Description |
|
||||
| --- | --- |
|
||||
| `--server-url` | Server address, required |
|
||||
| `--discovery-token` | first-time auto-registration Token, one of two with `--agent-token` |
|
||||
| `--agent-token` | node-specific Token, one of two with `--discovery-token` |
|
||||
| `--install-dir` | install dir, default `/opt/openflare-agent` |
|
||||
| `--openresty-path` | OpenResty binary path; auto-finds `openresty` when omitted |
|
||||
| `--repo` | GitHub repo for downloading the Agent, default `Rain-kl/OpenFlare` |
|
||||
| `--no-service` | don't create the systemd service |
|
||||
|
||||
Confirm state:
|
||||
|
||||
```bash
|
||||
systemctl status openflare-agent
|
||||
journalctl -u openflare-agent -f
|
||||
```
|
||||
@@ -0,0 +1,24 @@
|
||||
# Deployment & Upgrades
|
||||
|
||||
This section provides detailed deployment guides, configuration notes, and upgrade/maintenance steps for the OpenFlare Server, Agent, Relay, and the OpenFlared intranet penetration client.
|
||||
|
||||
## Content Navigation
|
||||
|
||||
### Quick Start
|
||||
* **[Quick Start](../guide/quick-start.md)**: start the Server and your first Agent with Docker Compose in 5 minutes (recommended for new users)
|
||||
|
||||
### Server Deployment
|
||||
* **[Start the Server](./server.md)**: build the frontend from source, start the Server, choose SQLite or PostgreSQL
|
||||
|
||||
### Agent Deployment
|
||||
* **[Access Agent](./agent.md)**: Agent connection methods, Docker deployment, script install, config file, and troubleshooting
|
||||
|
||||
### Tunnel Intranet Penetration Deployment
|
||||
* **[Deploy Relay](./relay.md)**: TunnelRelay node config notes, Docker deployment, and host running guide
|
||||
* **[Deploy OpenFlared](./openflared.md)**: intranet penetration client config notes, Docker running, and self-sync mechanism
|
||||
|
||||
### Upgrades & Maintenance
|
||||
* **[Upgrade & Maintenance](./upgrade.md)**: Server and Agent upgrade steps, data cleanup policy, verification commands
|
||||
|
||||
### Reference
|
||||
* **[Deployment Guide](./deployment.md)**: deployment topology, prerequisites, Docker Compose config examples, and an overview of multiple deployment approaches
|
||||
@@ -0,0 +1,84 @@
|
||||
# Deploy the OpenFlared Client
|
||||
|
||||
You will learn: OpenFlared's responsibilities, config parameters and env vars, running the client with Docker, and deploying it standalone on an intranet server via binary.
|
||||
|
||||
**OpenFlared** is the tunnel client deployed in your intranet (LAN, private cloud, or any environment not directly reachable from the public internet). Its core responsibility is communicating with the control plane (OpenFlare Server) via `X-Tunnel-Token`, and locally spawning and managing one or more **frpc (fast reverse proxy client)** processes to securely and stably tunnel intranet HTTP traffic to public relay nodes.
|
||||
|
||||
---
|
||||
|
||||
## Prerequisites
|
||||
|
||||
1. **Get a Tunnel Token**: add a node of type **Tunnel** in admin「Node Management」, save it, then open the node detail page to view the dedicated access Token.
|
||||
2. **Outbound network access**: the intranet server needs no public inbound IP or port mapping, but must reach the public **OpenFlare Server address** and the corresponding **TunnelRelay node relay port (default 7000)** over the network.
|
||||
3. **Software dependency** (host deployment only):
|
||||
- an executable `frpc` binary locally, or an explicitly specified path via parameter.
|
||||
|
||||
---
|
||||
|
||||
## Config File and Env Vars
|
||||
|
||||
`openflared` reads `flared.json` in the current directory by default at startup, fully overridable via env vars.
|
||||
|
||||
### Config Field Details
|
||||
|
||||
| JSON field | Env var | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `server_url` | `OPENFLARE_SERVER_URL` | OpenFlare Server API service address | **none (required)** |
|
||||
| `tunnel_token` | `OPENFLARE_TUNNEL_TOKEN` | tunnel client's dedicated auth Token | **none (required)** |
|
||||
| `frpc_path` | `OPENFLARE_FRPC_PATH` | frpc executable binary path | `"frpc"` |
|
||||
| `data_dir` | `OPENFLARE_DATA_DIR` | local data and generated `frpc_{relayNodeID}.toml` directory | `"./data"` |
|
||||
| `state_path` | - | local state record file path (last applied config version) | `"{data_dir}/flared-state.json"` |
|
||||
| `heartbeat_interval`| - | state heartbeat report period (ms int or Go Duration string) | `10000` (10s) |
|
||||
| `sync_interval` | - | tunnel config pull/sync period (ms int or Go Duration string) | `30000` (30s) |
|
||||
| `request_timeout` | - | API network request timeout | `10000` (10s) |
|
||||
|
||||
---
|
||||
|
||||
## Running with Docker
|
||||
|
||||
Docker deployment is the simplest and safest way to run in the intranet. The official `openflared` image bundles the client controller and the `frpc v0.69.0` binary runtime — no extra environment needed.
|
||||
|
||||
```bash
|
||||
docker pull ghcr.io/rain-kl/openflared:latest
|
||||
docker rm -f openflared 2>/dev/null || true
|
||||
|
||||
docker run -d --name openflared --restart unless-stopped \
|
||||
-e OPENFLARE_SERVER_URL=http://your-server:3000 \
|
||||
-e OPENFLARE_TUNNEL_TOKEN=YOUR_TUNNEL_TOKEN \
|
||||
-v openflared-data:/app/data \
|
||||
ghcr.io/rain-kl/openflared:latest
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Startup and Verification
|
||||
|
||||
### 1. Auto-Sync Logic
|
||||
|
||||
After starting successfully, OpenFlared runs this workflow:
|
||||
- **Heartbeat & config fetch**: periodically syncs with the Server's `/api/v1/tunnel/heartbeat` and `/api/v1/tunnel/config/active` endpoints, validating the Token and detecting config versions.
|
||||
- **File rendering**: when a config version (or checksum) changes, it auto-pulls the tunnel's full route rules. If multiple Relay nodes are bound, it renders `frpc_{relayNodeID}.toml` per Relay under `data_dir`.
|
||||
- **Config-change restart**: when config or checksum changes, it re-spawns the corresponding `frpc` child processes to keep traffic mappings current.
|
||||
- **Abnormal self-recovery**: if a local `frpc` tunnel process exits abnormally, the supervisor restarts it with exponential backoff (initial 1s, cap 60s).
|
||||
|
||||
### 2. View Logs and Connection State
|
||||
|
||||
```bash
|
||||
# Docker container logs
|
||||
docker logs -f openflared
|
||||
```
|
||||
|
||||
If the process runs correctly, you'll see output like:
|
||||
```text
|
||||
flared config loaded ...
|
||||
detected frpc version v0.69.0
|
||||
flared process started
|
||||
applying new tunnel config {"version": "...", "checksum": "..."}
|
||||
frpc process missing, starting {"relay_id": "..."}
|
||||
```
|
||||
|
||||
### 3. Confirm in the Admin Panel
|
||||
|
||||
Open **「Node Management」** in the admin panel and enter the Tunnel node's detail page:
|
||||
- View the node online state and flared runtime state (WebSocket connected / running / offline).
|
||||
- View the current applied version and the latest apply record.
|
||||
@@ -0,0 +1,92 @@
|
||||
# Deploy Relay (Tunnel Relay)
|
||||
|
||||
You will learn: TunnelRelay node responsibilities, `openflare-relay` config items and env vars, running Relay with Docker, and building/deploying Relay manually from source.
|
||||
|
||||
In OpenFlare's intranet penetration system, the **TunnelRelay node** plays a key role. Unlike regular edge nodes, besides running the traditional Agent (hosting OpenResty for HTTPS/WAF processing), it also runs the **Relay (frps tunnel manager)** service on the same machine, listening for tunnel connections from intranet clients (OpenFlared) and relaying traffic.
|
||||
|
||||
---
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Before deploying a TunnelRelay node, make sure:
|
||||
|
||||
1. **Registered as a TunnelRelay-type node**: in OpenFlare admin「Node Management」, add a node of type `tunnel_relay` and get its dedicated `agent_token`, or use the global `discovery_token`.
|
||||
2. **Network ports**:
|
||||
- `bindPort` (frpc connection port, default `7000`) must be reachable by public/intranet clients.
|
||||
- `vhostHTTPPort` (HTTP Vhost port, default `8080`) must be free; the Agent exchanges traffic with frps on this port.
|
||||
3. **Software dependency** (host deployment only):
|
||||
- an executable `frps` binary locally, or an explicitly specified path via parameter.
|
||||
|
||||
---
|
||||
|
||||
## Config File and Env Vars
|
||||
|
||||
`openflare-relay` reads `relay.json` in the current directory by default at startup, fully overridable via env vars.
|
||||
|
||||
### Config Field Details
|
||||
|
||||
| JSON field | Env var | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `server_url` | `OPENFLARE_SERVER_URL` | OpenFlare Server API service address | **none (required)** |
|
||||
| `agent_token` | `OPENFLARE_AGENT_TOKEN` | node-specific Token | one of these two |
|
||||
| `discovery_token` | `OPENFLARE_DISCOVERY_TOKEN` | auto-registration Token | one of these two |
|
||||
| `node_name` | `OPENFLARE_NODE_NAME` | node identifier name | local hostname by default |
|
||||
| `node_ip` | `OPENFLARE_NODE_IP` | node egress/listen IP | auto-detected real egress IP |
|
||||
| `frps_path` | `OPENFLARE_FRPS_PATH` | frps executable binary path | `"frps"` |
|
||||
| `data_dir` | `OPENFLARE_DATA_DIR` | local data and generated `frps.toml` directory | `"./data"` |
|
||||
| `state_path` | - | local state JSON record file path | `"{data_dir}/relay-state.json"` |
|
||||
| `heartbeat_interval`| - | heartbeat period (ms int or Go Duration string) | `10000` (10s) |
|
||||
| `request_timeout` | - | API request timeout | `10000` (10s) |
|
||||
|
||||
---
|
||||
|
||||
## Running with Docker
|
||||
|
||||
Docker is the most convenient deployment for a TunnelRelay node. The official image bundles the `openflare-relay` controller and the `frps` runtime — out of the box.
|
||||
|
||||
```bash
|
||||
docker pull ghcr.io/rain-kl/openflare-relay:latest
|
||||
docker rm -f openflare-relay 2>/dev/null || true
|
||||
|
||||
docker run -d --name openflare-relay --restart unless-stopped \
|
||||
-p 7000:7000 \
|
||||
-p 17500:17500 \
|
||||
-e OPENFLARE_SERVER_URL=http://your-server:3000 \
|
||||
-e OPENFLARE_AGENT_TOKEN=YOUR_AGENT_TOKEN \
|
||||
-v openflare-relay-data:/app/data \
|
||||
ghcr.io/rain-kl/openflare-relay:latest
|
||||
```
|
||||
|
||||
> [!TIP]
|
||||
> The `-p 7000:7000` mapping is the port frpc clients connect to for relaying. If the admin panel configures a custom `relay_bind_port`, adjust the host port mapping accordingly.
|
||||
|
||||
> [!NOTE]
|
||||
> **Enable the embedded frps Web UI**:
|
||||
> If the Server control panel enables the relay traffic monitoring panel (i.e. `relay_frps_web_ui_enabled` set to `true` in DB/system settings), you need to map the Web port (default `17500`, controlled by `relay_frps_web_ui_port` in system settings) to the host via `-p 17500:17500`.
|
||||
> The Web UI username is fixed to `admin`, and the password is the relay node's `agent_token`.
|
||||
|
||||
---
|
||||
|
||||
## Startup and Verification
|
||||
|
||||
### 1. View Process Logs
|
||||
|
||||
```bash
|
||||
# Docker container logs
|
||||
docker logs -f openflare-relay
|
||||
```
|
||||
|
||||
### 2. Verify Runtime State
|
||||
|
||||
After starting successfully, the Relay will:
|
||||
- Send HTTP heartbeats to the control plane to register/go online.
|
||||
- Fetch the latest frps base config from the control plane (including `bindPort`, `vhostHTTPPort`, and the auto-generated tunnel auth credential `auth_token`).
|
||||
- Render the local `data/frps.toml` config file.
|
||||
- Spawn the child process `frps -c data/frps.toml`.
|
||||
- If the process exits unexpectedly, the Relay auto-restarts frps with exponential backoff (initial 1s, cap 60s).
|
||||
|
||||
### 3. Confirm in the Admin Panel
|
||||
|
||||
Log in to the admin panel, navigate to **「Node Management」**, and confirm:
|
||||
- The TunnelRelay node status is marked **「Online」**.
|
||||
- The node type is correctly marked as **Relay node** and the frps runtime state is **Healthy**.
|
||||
@@ -0,0 +1,300 @@
|
||||
# Start the Server
|
||||
|
||||
You will learn: how to deploy with Docker (quick start, production-recommended, advanced) and how to deploy the OpenFlare Server locally from source.
|
||||
|
||||
The OpenFlare Server is a Gin + GORM monolithic control plane responsible for the admin UI, admin API, Agent API, config rendering, version release, data storage, and aggregation queries.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> **About external dependencies**:
|
||||
> OpenFlare has built-in support for background async tasks (Asynq framework). Therefore, **regardless of deployment mode, Redis (or Valkey) is required**. The main difference between deployment options is the primary relational DB choice (SQLite vs PostgreSQL) and whether tracing (Jaeger) is enabled.
|
||||
> For high business traffic, ClickHouse is recommended for log storage.
|
||||
|
||||
> [!TIP]
|
||||
> **ClickHouse server performance config (recommended mount)**
|
||||
> The control plane is typically a small host (e.g. 3c6g). The `performance.xml` provided in the repo tightens the background merge/mutation thread pools, avoiding high idle CPU or ClickHouse 25.x startup validation failures on small machines.
|
||||
> Mount the local `./config/clickhouse/performance.xml` as a single file at `/etc/clickhouse-server/config.d/performance.xml` to keep the official image's built-in Docker network listening config.
|
||||
|
||||
Pull the config locally before deploying:
|
||||
|
||||
```bash
|
||||
mkdir -p ./config/clickhouse
|
||||
curl -fsSL -o ./config/clickhouse/performance.xml \
|
||||
https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/config/clickhouse/performance.xml
|
||||
```
|
||||
|
||||
Add it to the ClickHouse service `volumes` (alongside the data volume):
|
||||
|
||||
```yaml
|
||||
volumes:
|
||||
- ./data/clickhouse_data:/var/lib/clickhouse # or named volume
|
||||
- ./config/clickhouse/performance.xml:/etc/clickhouse-server/config.d/performance.xml:ro
|
||||
```
|
||||
|
||||
After modifying `performance.xml`, run `docker compose restart clickhouse` for it to take effect.
|
||||
|
||||
---
|
||||
|
||||
## Method 1: Docker Deployment (recommended)
|
||||
|
||||
Docker deployment avoids configuring Go and Node.js frontend build environments locally. Choose one of the three options based on your hardware and needs:
|
||||
|
||||
### 1. Quick Start (SQLite + Redis)
|
||||
|
||||
> **Use case**: testing/experience, lightweight single-machine deployment.
|
||||
>
|
||||
> **Features**: primary relational DB is SQLite.
|
||||
|
||||
Create a `docker-compose.yaml`:
|
||||
|
||||
```yaml
|
||||
version: '3.8'
|
||||
|
||||
services:
|
||||
openflare:
|
||||
image: ghcr.io/rain-kl/openflare:latest
|
||||
container_name: openflare-server
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "3000:3000"
|
||||
volumes:
|
||||
- ./openflare-data:/data
|
||||
- ./uploads:/app/uploads
|
||||
environment:
|
||||
TZ: Asia/Shanghai
|
||||
APP_SESSION_SECRET: 'replace-with-a-long-random-string' # replace with a long random string in production
|
||||
DB_ENABLED: "false" # disables PostgreSQL, auto-enables the built-in SQLite fallback
|
||||
SQLITE_PATH: "/data/openflare.db"
|
||||
REDIS_ENABLED: "true"
|
||||
REDIS_ADDR: "redis:6379"
|
||||
depends_on:
|
||||
redis:
|
||||
condition: service_healthy
|
||||
|
||||
redis:
|
||||
image: valkey/valkey:8.0-alpine
|
||||
restart: unless-stopped
|
||||
command: ["valkey-server", "--appendonly", "yes"]
|
||||
volumes:
|
||||
- ./data/valkey:/data
|
||||
healthcheck:
|
||||
test: ["CMD", "valkey-cli", "ping"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 2. Small-Traffic Business (PostgreSQL + Redis)
|
||||
|
||||
> **Use case**: production, small-to-medium traffic; PostgreSQL won't be the log-write bottleneck.
|
||||
|
||||
Create a `docker-compose.yaml`:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
openflare:
|
||||
image: ghcr.io/rain-kl/openflare:latest
|
||||
restart: unless-stopped
|
||||
env_file: .env
|
||||
environment:
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
ports:
|
||||
- "3000:3000"
|
||||
volumes:
|
||||
- openflare_uploads:/app/uploads
|
||||
depends_on:
|
||||
postgres:
|
||||
condition: service_healthy
|
||||
redis:
|
||||
condition: service_healthy
|
||||
|
||||
postgres:
|
||||
image: postgres:17-alpine
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
POSTGRES_DB: ${DB_NAME:-openflare}
|
||||
POSTGRES_USER: ${DB_USERNAME:-openflare}
|
||||
POSTGRES_PASSWORD: ${DB_PASSWORD:-replace-with-strong-password}
|
||||
volumes:
|
||||
- openflare_postgres_data:/var/lib/postgresql/data
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U ${DB_USERNAME:-openflare} -d ${DB_NAME:-openflare}"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
|
||||
redis:
|
||||
image: valkey/valkey:8.0-alpine
|
||||
restart: unless-stopped
|
||||
command: ["valkey-server", "--appendonly", "yes"]
|
||||
volumes:
|
||||
- openflare_redis_data:/data
|
||||
healthcheck:
|
||||
test: ["CMD", "valkey-cli", "ping"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
start_period: 5s
|
||||
|
||||
volumes:
|
||||
openflare_uploads:
|
||||
openflare_postgres_data:
|
||||
openflare_redis_data:
|
||||
```
|
||||
|
||||
Create a matching `.env` file for system env vars (copy and modify the root `.env.example`):
|
||||
|
||||
```bash
|
||||
curl -o .env.example https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/.env.example
|
||||
cp .env.example .env
|
||||
# edit .env: fill in DB, Redis, passwords, and APP_SESSION_SECRET
|
||||
|
||||
docker compose up -d
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### 3. Advanced (full orchestration with Jaeger tracing)
|
||||
|
||||
> **Use case**: high traffic; needs trace performance metrics.
|
||||
>
|
||||
> **Features**: on top of the "production-recommended" bundle, stores logs with ClickHouse and uses Jaeger as the OpenTelemetry (OTel) tracing backend.
|
||||
|
||||
Create a `docker-compose.yaml`:
|
||||
|
||||
```yaml
|
||||
version: '3.8'
|
||||
|
||||
services:
|
||||
openflare:
|
||||
image: ghcr.io/rain-kl/openflare:latest
|
||||
restart: unless-stopped
|
||||
env_file: .env
|
||||
environment:
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
OTEL_EXPORTER_OTLP_ENDPOINT: "http://jaeger:4317"
|
||||
OTEL_EXPORTER_OTLP_INSECURE: "true"
|
||||
OTEL_SAMPLING_RATE: "1.0" # sampling rate; 1.0 samples all traces
|
||||
ports:
|
||||
- "3000:3000"
|
||||
volumes:
|
||||
- openflare_uploads:/app/uploads
|
||||
depends_on:
|
||||
postgres:
|
||||
condition: service_healthy
|
||||
redis:
|
||||
condition: service_healthy
|
||||
clickhouse:
|
||||
condition: service_healthy
|
||||
jaeger:
|
||||
condition: service_started
|
||||
|
||||
postgres:
|
||||
image: postgres:17-alpine
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
POSTGRES_DB: ${DB_NAME:-openflare}
|
||||
POSTGRES_USER: ${DB_USERNAME:-openflare}
|
||||
POSTGRES_PASSWORD: ${DB_PASSWORD:-replace-with-strong-password}
|
||||
volumes:
|
||||
- openflare_postgres_data:/var/lib/postgresql/data
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U ${DB_USERNAME:-openflare} -d ${DB_NAME:-openflare}"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
|
||||
redis:
|
||||
image: valkey/valkey:8.0-alpine
|
||||
restart: unless-stopped
|
||||
command: ["valkey-server", "--appendonly", "yes"]
|
||||
volumes:
|
||||
- openflare_redis_data:/data
|
||||
healthcheck:
|
||||
test: ["CMD", "valkey-cli", "ping"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
start_period: 5s
|
||||
|
||||
jaeger:
|
||||
image: jaegertracing/jaeger:2.19.0
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
ports:
|
||||
- "16686:16686" # Web UI port
|
||||
- "4317:4317" # OTLP gRPC receive port
|
||||
- "4318:4318" # OTLP HTTP receive port
|
||||
|
||||
clickhouse:
|
||||
image: clickhouse/clickhouse-server:25.3-alpine
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
CLICKHOUSE_DB: ${CLICKHOUSE_NAME:-openflare}
|
||||
CLICKHOUSE_USER: ${CLICKHOUSE_USERNAME:-default}
|
||||
CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}
|
||||
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1
|
||||
TZ: ${TZ:-Asia/Shanghai}
|
||||
ulimits:
|
||||
nofile:
|
||||
soft: 262144
|
||||
hard: 262144
|
||||
volumes:
|
||||
- openflare_clickhouse_data:/var/lib/clickhouse
|
||||
- ./config/clickhouse/performance.xml:/etc/clickhouse-server/config.d/performance.xml:ro
|
||||
healthcheck:
|
||||
test: ["CMD", "clickhouse-client", "--user", "${CLICKHOUSE_USERNAME:-default}", "--password", "${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}", "--query", "SELECT 1"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
start_period: 15s
|
||||
|
||||
volumes:
|
||||
openflare_uploads:
|
||||
openflare_postgres_data:
|
||||
openflare_redis_data:
|
||||
openflare_clickhouse_data:
|
||||
```
|
||||
|
||||
Start and verify:
|
||||
|
||||
```bash
|
||||
mkdir -p ./config/clickhouse
|
||||
curl -fsSL -o ./config/clickhouse/performance.xml \
|
||||
https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/config/clickhouse/performance.xml
|
||||
curl -o .env.example https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/.env.example
|
||||
cp .env.example .env
|
||||
# edit .env and make sure APP_SESSION_SECRET password is set
|
||||
|
||||
docker compose up -d
|
||||
```
|
||||
After startup, open `http://localhost:16686` to view the Jaeger monitoring UI and system span traces.
|
||||
|
||||
---
|
||||
|
||||
## First Login
|
||||
|
||||
The Server listens on port `3000` by default; open `http://localhost:3000` in a browser after startup.
|
||||
|
||||
Default admin account:
|
||||
|
||||
| Username | Password |
|
||||
| --- | --- |
|
||||
| `admin` | `12345678` |
|
||||
|
||||
> [!WARNING]
|
||||
> For your system's security, change the default password immediately in your profile settings after the first login.
|
||||
|
||||
---
|
||||
|
||||
## Distributed Deployment
|
||||
|
||||
In large production deployments, split the Server into multiple processes by responsibility:
|
||||
|
||||
```bash
|
||||
go run main.go api # API service for admin panel and node communication only
|
||||
go run main.go worker # background task Worker service only
|
||||
go run main.go scheduler # scheduled task Scheduler service only
|
||||
```
|
||||
@@ -0,0 +1,20 @@
|
||||
# Upgrade & Maintenance
|
||||
|
||||
You will learn: how to upgrade the Server and Agent, how to clean up observability data, and which verification commands to run before and after maintenance.
|
||||
|
||||
Before upgrading, confirm the current active version, the most recent Agent apply result, and the DB backup policy. In production, don't upgrade while a config release, a large-scale Agent reconnect, or a DB migration is in progress.
|
||||
|
||||
## Server Upgrade
|
||||
|
||||
Pull the latest image and upgrade:
|
||||
|
||||
```bash
|
||||
docker compose pull
|
||||
docker compose up
|
||||
```
|
||||
|
||||
For source deployments, restart the Server and confirm the logs show no DB migration or startup errors.
|
||||
|
||||
## Agent Upgrade
|
||||
|
||||
The Agent only caches runtime config and state files locally — no business data. To upgrade, directly pull the latest image and recreate the container. For specific deployment commands and install methods, see **[Access Agent](./agent.md)**.
|
||||
@@ -0,0 +1,177 @@
|
||||
# Agent Design
|
||||
|
||||
You will learn: the Agent's design principles, core functional modules, interaction chain with the Server, and how the immutable version model and three-stage disaster recovery guarantee config-apply safety and reliability.
|
||||
|
||||
---
|
||||
|
||||
## Requirements Analysis
|
||||
|
||||
In distributed reverse-proxy and edge-security gateway scenarios, the Agent is the core bridge between the control plane (Server) and the data plane (OpenResty). Since the Agent runs on the user's actual node server, its design must satisfy these core security and HA requirements:
|
||||
|
||||
1. **Active pull (Pull model), not passive receive**: the Server doesn't hold node SSH keys and never initiates inbound connections to nodes. All control instructions and config updates are pulled upward by the Agent via heartbeat or WebSocket. This removes inbound-firewall security risks on nodes and prevents control-channel hijacking.
|
||||
2. **Minimal invasiveness**: the Agent runs as a standalone Go binary, interacting with the local OpenResty process only via file-based config rewriting and signal notifications — no interference with other system services on the node.
|
||||
3. **Strong disaster recovery and self-healing**: since network jitter, full disks, or bad configs can easily break config sync, the Agent must have zero-dependency local rollback self-healing to prevent one bad config from taking down the whole machine.
|
||||
4. **Pure data and state landing**: the Agent only carries Server-rendered files and control intent to landing; it contains no complex business validation or multi-tenant auth — control-plane duties stay on the Server, keeping the node efficient and light.
|
||||
|
||||
---
|
||||
|
||||
## Core Features
|
||||
|
||||
The Agent mainly consists of these submodules cooperating for its full lifecycle:
|
||||
|
||||
| Module | Directory | Responsibility |
|
||||
| :--- | :--- | :--- |
|
||||
| **Config sync** | `sync/` | pull full config packages, write files, trigger reloads, record and report sync state. |
|
||||
| **Heartbeat** | `heartbeat/` | periodically report node health, resource metrics, and fetch the latest active version summary. |
|
||||
| **WebSocket** | `wsclient/` | keep a long connection to the Server for second-level real-time config push and control-plane instructions. |
|
||||
| **OpenResty control** | `nginx/` | run Nginx config validation (`openresty -t`), rewriting, smooth reload, and process auto-start. |
|
||||
| **Local state** | `state/` | persist local applied version, error logs, and buffered observability metrics not yet reported. |
|
||||
| **Self-update** | `updater/` | listen for Server self-update instructions, safely fetch new binaries, and hot-upgrade in place. |
|
||||
| **Observability** | `observability/` | collect host resource readings, OpenResty health/connections, and tail access-log details for reporting; **no** business pre-aggregation like UV/TopN/throughput. See [Edge Observability & Business Traffic Stats](./observability-design.md). |
|
||||
| **GeoIP maintenance** | `geoipdata/` `geoipupdate/` | maintain and periodically update the local GeoIP DB for WAF geo filtering. |
|
||||
|
||||
---
|
||||
|
||||
## Interaction Chain with the Server
|
||||
|
||||
The Agent communicates with the control plane via **Token-based auto-registration** and **heartbeat/WebSocket dual channels** over its lifecycle.
|
||||
|
||||
### 1. Auto-Registration Flow
|
||||
If `access_token` in the local `agent.json` is empty at startup but `discovery_token` is configured, auto-registration triggers:
|
||||
1. The Agent sends a registration request to `/api/v1/agent/nodes/register` with local hardware summary, IP, and hostname.
|
||||
2. The Server validates the `discovery_token`, generates a unique `NodeID` and dedicated `AccessToken` (i.e. `agent_token`), and returns them.
|
||||
3. The Agent writes the dedicated Token into the local config file, erases the one-time `discovery_token`, and all future communication authenticates with the dedicated `AccessToken`.
|
||||
|
||||
### 2. Dual-Channel Heartbeat and Sync
|
||||
* **HTTP polling channel (fallback & probe)**: the Agent POSTs heartbeats at the configured `heartbeat_interval` by default, reporting metrics while fetching the current active version summary (Version & Checksum).
|
||||
* **WebSocket channel (real-time)**: after a successful HTTP heartbeat, the Agent auto-upgrades to WebSocket (`/api/v1/agent/ws`).
|
||||
* Once established, heartbeat and metric reporting fully move to the WS pipe, reducing network overhead.
|
||||
* When the Server releases/activates a new version, it broadcasts to Agents via WS. The Agent triggers sync **immediately** on the change event for second-level config effect.
|
||||
* If the WS link drops due to network issues, the Agent degrades to HTTP polling and retries WS with exponential backoff.
|
||||
|
||||
### 3. Interaction Sequence Diagram
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant Agent as OpenFlare Agent
|
||||
participant OR as Local OpenResty
|
||||
participant Server as OpenFlare Server
|
||||
|
||||
Note over Agent: first startup (no AccessToken)
|
||||
Agent->>Server: 1. auto-registration request (with discovery_token)
|
||||
Server-->>Agent: 2. issue NodeID and dedicated AccessToken (agent_token)
|
||||
Note over Agent: store Token in local config file
|
||||
|
||||
rect rgb(240, 248, 255)
|
||||
Note over Agent, Server: HTTP fallback and WebSocket upgrade
|
||||
Agent->>Server: 3. send HTTP Heartbeat (report system state and health)
|
||||
Server-->>Agent: 4. return ActiveConfig summary and AgentSettings
|
||||
Agent->>Server: 5. request WebSocket upgrade (/api/v1/agent/ws)
|
||||
Server-->>Agent: 6. upgrade success (bidirectional persistent real-time channel)
|
||||
end
|
||||
|
||||
rect rgb(245, 245, 245)
|
||||
Note over Agent, Server: real-time config release/apply chain
|
||||
Note over Server: admin clicks publish config in the UI
|
||||
Server->>Agent: 7. broadcast new config summary via WS (WSMessageTypeActiveConfig)
|
||||
Agent->>Server: 8. request full config details (with target Version/Checksum)
|
||||
Server-->>Agent: 9. return full config snapshot (Nginx config, certs, WAF rules, etc.)
|
||||
Note over Agent: back up old files, write new config to local temp path
|
||||
Agent->>OR: 10. run config syntax validation (openresty -t)
|
||||
OR-->>Agent: 11. return validation result (OK)
|
||||
Agent->>OR: 12. smooth reload signal (openresty -s reload)
|
||||
Agent->>Server: 13. report apply success (Apply Log & ActiveVersion)
|
||||
end
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## OpenResty Control
|
||||
|
||||
The Agent's control over the data-plane OpenResty forms an end-to-end loop: config landing, syntax validation, smooth reload, and abnormal-state capture.
|
||||
|
||||
### 1. Config File Landing Organization
|
||||
After a successful sync, the Agent writes config under `data_dir` (default relative `etc/nginx/`, `etc/openflare/`, `var/lib/openflare/`; exact paths follow `main_config_path`, `route_config_path`, `cert_dir`, `lua_dir`, `runtime_config_dir`, `pages_dir` in `agent.json`):
|
||||
* `nginx.conf`: the main config (replaces relevant placeholders, configures performance params, Shared Dictionaries, and the global Server).
|
||||
* `conf.d/openflare_routes.conf`: the route config (generated by the Agent; contains all proxied sites' Server blocks, cert paths, cache, and rate-limit directives).
|
||||
* `certs/`: certificate dir (files named `{cert_id}.crt` and `{cert_id}.key`).
|
||||
* `lua/waf/` and `lua/pow/`: dedicated Lua runtime scripts for WAF and anti-CC challenges.
|
||||
* `etc/openflare/waf_config.json` and `waf_ip_groups.json`: structured rule configs for the WAF filtering engine.
|
||||
* `pages_dir`: the Pages static site deployment dir, default `data_dir/var/lib/openflare/pages`. When the active config references a Pages **project**, the Agent requests the control plane's「latest active package」(hash + package) by `project_id`, streams to a temp file with real response-size limits and SHA-256 validation, then safely extracts to `projects/{project_id}/releases/{hash}`. After extraction it rechecks file count and total bytes; absolute hard caps are 2 GiB package, 1,000 files, 8 GiB single-file/total. It then atomically switches `current` and **immediately deletes other historical releases of the same project** (only latest kept). Switching the active deployment within a project doesn't require republishing the main config; multi-project reconciliation isolates single-project failures.
|
||||
|
||||
### 2. Fine-Grained Reload Actions
|
||||
1. **Back up current config**: before writing new files, copy existing config to a `.backup` temp dir, keeping a full scene snapshot.
|
||||
2. **Write and replace placeholders**: write the latest template, replacing absolute-path placeholders (e.g. `__OPENFLARE_LUA_DIR__`, `__OPENFLARE_PAGES_DIR__`) with local actual runtime paths.
|
||||
3. **Syntax validation**: run `openresty -t -c <temp_nginx.conf>` for strict syntax testing.
|
||||
4. **Smooth reload**: on validation pass, move the new config to the formal path and run `openresty -s reload`. If OpenResty isn't started, start the process with the current config.
|
||||
5. **Capture exceptions**: on validation/reload failure, the Agent captures command stdout/stderr as failure details for reporting.
|
||||
|
||||
---
|
||||
|
||||
## Release and Config Apply Model
|
||||
|
||||
OpenFlare uses an **immutable config version release model**, not online dynamic patching of node configs.
|
||||
|
||||
```text
|
||||
modify rules -> preview / view diff -> release -> generate full config version -> activate version -> Agent pulls -> local apply -> report result
|
||||
```
|
||||
|
||||
### 1. Core Design Principles
|
||||
* **Full release**: each release compiles all enabled routes, certs, Pages deployment references, and global/local WAF rules on the control plane in one pass, generating a full version with a unique `checksum`.
|
||||
* **Version format**: `YYYYMMDD-NNN` incrementing format for intuitive, monotonically increasing version history.
|
||||
* **Globally single active version**: only one globally active config version exists at a time. Rollback doesn't reverse-patch; just set a historical healthy version to `active`, and Agents re-pull and apply it.
|
||||
|
||||
### 2. Three-Stage Disaster Recovery Rollback
|
||||
When the Agent detects a config apply (or smooth reload) failure, it auto-activates this three-stage anti-outage chain:
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
A[config apply failed] --> B[stage 1: try local backup restore]
|
||||
B -- backup file exists --> C[write local backup files]
|
||||
C --> D[run openresty -t validation]
|
||||
D -- validation ok --> E[reload to restore old version]
|
||||
D -- validation failed --> F[enter stage 2]
|
||||
B -- no backup --> F[stage 2: write built-in safe fallback config]
|
||||
F --> G[write fallback nginx.conf: listen on 80 only]
|
||||
G --> H[enable stub_status health check]
|
||||
G --> I[other routes return 503 uniformly and block bad configs]
|
||||
G --> J[try starting OpenResty to keep basic liveness]
|
||||
J --> K[enter stage 3]
|
||||
E --> L[report Apply Warning]
|
||||
K --> M[locally block re-applying the bad version]
|
||||
M --> N[report Apply Error with detailed error]
|
||||
```
|
||||
|
||||
1. **Stage 1: local backup fallback**
|
||||
* The Agent tries restoring the main config, routes, and certs from the previously saved `.backup` dir.
|
||||
* After writing backup files, re-run `openresty -t`. On success, reload back and report `Warning` to the Server (new version apply failed; auto-rolled back to the last healthy version).
|
||||
2. **Stage 2: built-in safe fallback runtime**
|
||||
* If no local backup config exists (e.g. first deployment with a bad config), or the restored backup still fails validation, the Agent activates the final self-healing mechanism — writing the **built-in safe fallback config**.
|
||||
* **Safe fallback spec**:
|
||||
* listens only on port `80`, containing no user real reverse proxy routes.
|
||||
* Everything except the `/openflare/stub_status` health-check route (which returns normally) returns `503 Service Unavailable` with a fixed body `OpenFlare: No Valid Configuration`.
|
||||
* Tries starting OpenResty with this minimal config. This keeps the Nginx process itself alive, preserves the underlying health check/probe channel, prevents container/Pod restart loops from failed health checks, and protects sensitive routes.
|
||||
3. **Stage 3: local config blocking**
|
||||
* The Agent records the crash-causing config `version + checksum` in a blocklist in the local state store.
|
||||
* Until the control plane activates a new config (`checksum` changes), the Agent heartbeat blocks re-pulling that bad version — preventing the "heartbeat → pull crash config → crash rollback" infinite loop.
|
||||
|
||||
### 3. WAF IP Group Runtime Async Sync
|
||||
To avoid high-frequency malicious-IP blocklist changes constantly triggering full main-config releases and reloads (smooth reload still has slight CPU and connection overhead on Nginx), IP group members use an **async differential sync** decoupled from release versions:
|
||||
|
||||
* **Static release snapshot**: the released `waf_config.json` only contains rule groups' references to IP groups (`ip_whitelist_group_ids` / `ip_blacklist_group_ids`), not the concrete IP member lists.
|
||||
* **Heartbeat differential comparison**: the Agent reports the MD5 Checksum map of locally cached IP groups in heartbeat packets.
|
||||
* **Differential dispatch**: the Server compares the hashes of IP groups referenced by the current active version and only dispatches missing or changed members, written to the local `waf_ip_groups.json` for fast differential sync.
|
||||
* **WebSocket real-time notification**: when the Server manually updates an IP group, a subscription source sync succeeds, or security rules auto-trigger temporary bans, the Server immediately broadcasts the affected IP group update via WebSocket; the Agent lands it instantly — **no Nginx reload** throughout.
|
||||
|
||||
---
|
||||
|
||||
## Design Constraints
|
||||
|
||||
To guarantee the security boundary of data and control channels, Agent code and secondary development must strictly follow:
|
||||
|
||||
1. **Zero privileged command channel**: the Server is absolutely forbidden from passing arbitrary shell commands or remote script execution (exec/eval, etc.) to the Agent. All system control primitives (start, stop, reload, update) must be hardcoded inside the Agent binary.
|
||||
2. **Strict Token filtering and prefix validation**: when the Agent requests resources from the Server, endpoints are fixed under the `/api/v1/agent/` prefix and must carry `X-Agent-Token` for signature/token verification.
|
||||
3. **Node autonomy**: the Agent must have complete offline capability. While disconnected from the Server, the local OpenResty must keep reverse-proxying normally based on locally landed config.
|
||||
4. **Observability reports facts only**: access logs are reported as details; host metrics report counters/instant readings. Computing conclusion metrics like business UV, Top domains, or 24h data provided inside the Agent is forbidden (the Server aggregates). See [Edge Observability & Business Traffic Stats](./observability-design.md).
|
||||
5. **Pages consumes only control-plane artifacts**: Remote URLs, GitHub Releases, the auto scanner, and future repo checkout/build executors are all Server responsibilities. The Agent receives no external URLs, access tokens, repo credentials, or clone/install/build commands — it only pulls already-activated deployment packages with integrity metadata.
|
||||
@@ -0,0 +1,208 @@
|
||||
# System Architecture
|
||||
|
||||
You will learn: OpenFlare's overall architecture, the responsibility split of each core component (Server, Agent, OpenResty, Relay, Client), and the macro flow of the main data and request streams.
|
||||
|
||||
OpenFlare is a self-hosted OpenResty control plane. Physically it consists of the Server (control plane), the Agent (config landing), node-local OpenResty (data plane), intranet penetration components (Relay and OpenFlared, data-plane extensions), and the admin frontend.
|
||||
|
||||
---
|
||||
|
||||
## Traffic Path Overview
|
||||
|
||||
Depending on the website upstream type, OpenFlare supports three data-plane traffic paths:
|
||||
|
||||
### 1. Standard Reverse Proxy Path
|
||||
```text
|
||||
Browser
|
||||
|
|
||||
| HTTPS/HTTP request
|
||||
v
|
||||
OpenResty (WAF, TLS, Rate Limit, optional origin error page)
|
||||
|
|
||||
| reverse proxy (proxy_pass)
|
||||
v
|
||||
Origin Server (direct public/LAN upstream)
|
||||
```
|
||||
|
||||
When the origin or gateway returns an error status in the configured list, a global custom/default HTML can be returned while keeping the real HTTP status; see [Origin Error Page Design](./origin-error-page.md).
|
||||
|
||||
### 2. Intranet Penetration Path
|
||||
For origin services on firewall-restricted intranet servers:
|
||||
```text
|
||||
Browser
|
||||
|
|
||||
| HTTPS/HTTP request
|
||||
v
|
||||
OpenResty (Agent host, TLS/WAF)
|
||||
|
|
||||
| proxy_pass http://localhost:vhost_port (Host header preserved)
|
||||
v
|
||||
OpenFlareRelay (frps) <-- same host as the Agent, provides relaying
|
||||
|
|
||||
| frp tunnel protocol (Host header routing)
|
||||
v
|
||||
OpenFlared (frpc) <-- firewall-restricted intranet server
|
||||
|
|
||||
| HTTP/HTTPS forward
|
||||
v
|
||||
Internal Service (192.168.x.x)
|
||||
```
|
||||
|
||||
### 3. Pages Static Hosting Path
|
||||
For pre-built SPAs or static site hosting:
|
||||
```text
|
||||
Browser
|
||||
|
|
||||
| HTTPS/HTTP request
|
||||
v
|
||||
OpenResty (Agent, TLS/WAF)
|
||||
|
|
||||
+---> [static serving] root/try_files ---> Agent local Pages deployment dir
|
||||
|
|
||||
+---> [API proxy] proxy_pass ---> backend API service (if API proxying enabled)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Component Responsibilities
|
||||
|
||||
| Component | Responsibility | Detailed Design Reference |
|
||||
| --- | --- | --- |
|
||||
| **Server** | admin UI/API, control-plane state persistence, config compilation/rendering, release versioning, Pages deployment package storage, Cloudflare A-record pointing, access-log storage and business traffic aggregation, Uptime Kuma monitoring sync, login CAPTCHA protection | [Agent & Publish Model](./agent-design.md) / [Cloudflare DNS Pointing Design](./cloudflare-pointing.md) / [Edge Observability & Business Traffic Stats](./observability-design.md) / [Uptime Kuma Sync Design](./kuma-design.md) / [Login CAPTCHA Design](./login-captcha.md) |
|
||||
| **Agent** | periodic heartbeat & WS sync, static package pull/extraction, OpenResty config write/validate/reload and self-healing; observability reports only access details and host/health readings, no business pre-aggregation | [Agent & Publish Model](./agent-design.md) / [Edge Observability & Business Traffic Stats](./observability-design.md) |
|
||||
| **OpenResty** | receives real traffic; executes WAF filtering, PoW protection, Basic Auth, static/reverse-proxy serving, and optional origin error pages | [WAF Design](./waf-design.md) / [Pages Design](./pages-design.md) / [Origin Error Page Design](./origin-error-page.md) |
|
||||
| **Relay** | deployed on edge nodes; manages the `frps` daemon lifecycle and accepts heartbeat-dispatched penetration relay configs | [Tunnel Design](./tunnel-design.md) |
|
||||
| **OpenFlared** | deployed in the intranet; manages the `frpc` process group, establishes reverse tunnels to multiple Relays, reports connection state | [Tunnel Design](./tunnel-design.md) |
|
||||
|
||||
---
|
||||
|
||||
## Component Architecture and Division
|
||||
|
||||
### 1. Server (control plane)
|
||||
The Go backend at the repo root (module `github.com/Rain-kl/Wavelet`) is the OpenFlare control plane, built on the Wavelet full-stack scaffold:
|
||||
* Provides admin REST APIs (`/api/v1/d/*`) authenticated via **Session Cookie**, with optional `X-Access-Token`.
|
||||
* Edge node protocols go through `/api/v1/agent|relay|tunnel/*`, authenticated with `X-Agent-Token` / `X-Tunnel-Token` respectively.
|
||||
* Contains the config Compiler, uniformly compiling DB rules, certs, and global params into immutable config snapshots and OpenResty physical config file text.
|
||||
* Uniformly receives Pages local uploads, Remote URLs, and public GitHub Release pre-built artifacts, completing source checks, restricted downloads, archive validation, and immutable deployments; manual uploads create candidates awaiting explicit activation, persistent-source sync creates-or-loads and atomically activates. The Server offers controlled latest-download endpoints to Agents; the internal scanner handles limited GitHub latest checks, lease recovery, optional auto-publish, and orphan upload compensation; the generic task management entry can't modify this schedule. Future repo source builds are extended by a standalone Server build executor; the Agent never executes third-party fetch or build commands.
|
||||
* Provides the optional Cloudflare DNS pointing control plane: maintains group desired state with ZoneDomains as members, idempotently syncing a single A record to the current active node IPv4 via Asynq; node IP changes only best-effort enqueue; no auto-failover in phase 1.
|
||||
* Backend integration with the Uptime Kuma monitoring sync service auto-maintains HTTP probe tasks for available sites.
|
||||
* Startup entry: root `main.go` + `internal/cmd/` (`api` / `worker` / `scheduler` / `all`); OpenFlare business in `internal/apps/openflare/`, edge protocol handling in `internal/apps/openflare/{agent,relay,flared}/`.
|
||||
* *See: [Agent & Publish Model](./agent-design.md) and [Uptime Kuma Sync Design](./kuma-design.md)*
|
||||
|
||||
### 2. Agent (config landing)
|
||||
`openflare-agent` is the daemon running on the node:
|
||||
* Maintains periodic heartbeats with the control plane after startup, receiving real-time config release broadcasts via the optional WebSocket.
|
||||
* Pulls the latest active version's config files and certs, writes them locally, and performs safe validation via `openresty -t` before a smooth reload.
|
||||
* Handles Pages deployment package download, SHA-256 validation, and extraction switching locally.
|
||||
* *See: [Agent & Publish Model](./agent-design.md)*
|
||||
|
||||
### 3. OpenResty (data plane)
|
||||
Receives visitor traffic and performs final business landing:
|
||||
* Traffic entry, supporting HTTP/2, HTTP/3 (QUIC), and dynamic TLS certificate binding.
|
||||
* Embeds Lua logic filtering WAF rules and verifying PoW challenges efficiently in the `access_by_lua` phase, followed by connection/rate limits and basic caching (policy in [Edge Cache Strategy Design](./edge-cache-design.md)).
|
||||
* *See: [WAF Design](./waf-design.md) and [Pages Static Hosting Design](./pages-design.md)*
|
||||
|
||||
### 4. Relay and OpenFlared (tunnel components)
|
||||
Extend data-plane reverse penetration:
|
||||
* `openflare-relay` guards the local `frps`, accepts Server config dispatch, and auto-updates the relay port.
|
||||
* `openflared` guards a group of `frpc` client processes in the intranet for nearest multi-relay connections and HA disaster recovery.
|
||||
* *See: [Tunnel Design](./tunnel-design.md)*
|
||||
|
||||
---
|
||||
|
||||
## Data and Request Flow Overview
|
||||
|
||||
### 1. Config Release and Sync Flow
|
||||
```text
|
||||
admin modifies config -> release new version -> generate globally unique Checksum active version
|
||||
|
|
||||
+------------------+------------------+
|
||||
| (WebSocket broadcast or periodic Heartbeat) |
|
||||
v v
|
||||
[edge node Agent] [intranet OpenFlared]
|
||||
pull latest OpenResty config/certs pull latest Tunnel mapping config
|
||||
incrementally pull/extract Pages packages generate/rewrite frpc.toml
|
||||
validate config and smooth reload smooth reload or spawn frpc
|
||||
report apply state (Success / Error) report tunnel connection state and metrics
|
||||
```
|
||||
* *Fine-grained sync/self-healing timing and the rollback model: [Agent & Publish Model](./agent-design.md)*
|
||||
|
||||
### 2. Static Hosting and API Proxy Flow
|
||||
* Static assets are extracted to `projects/{project_id}/current` on the Agent node (pulled per project latest, only the newest package kept); OpenResty serves static resources at the edge via `root`/`index`/`try_files`.
|
||||
* With API proxying enabled, OpenResty rewrites and forwards (`proxy_pass`) API requests to the backend dynamic API based on the site's `api_proxy_path` (e.g. `/api`).
|
||||
* Admin operations and the internal scanner only generate constrained artifact candidates, reusing the unified inspect, `upload.Ingest`, and deployment pipeline. Manual uploads create a new inactive candidate; persistent-source sync/scanner creates-or-loads and atomically activates. A future repository build executor can only emit into the same artifact pipeline; the Agent is always just an active-deployment consumer.
|
||||
* *Package validation, extraction escape defense, and Nginx rule rendering: [Pages Static Hosting Design](./pages-design.md)*
|
||||
|
||||
### 3. WAF Security Filtering Flow
|
||||
* The WAF engine is embedded in the OpenResty request lifecycle.
|
||||
* WAF rules are orchestrated as a visual DAG on the control plane and compiled into a runtime graph at release; after an OpenResty reload each Worker loads it once, and subsequent requests only traverse the in-memory object.
|
||||
* Global rules always run first; route-bound rules execute in explicit order; reaching "pass" in the current rule continues to the next, reaching "block" immediately returns that node's configured block response.
|
||||
* IP group members hot-update independently: a coordinating worker checks the checksum every 5 seconds, loading the full snapshot only on change; each Worker's request path always reads the local in-memory object.
|
||||
* *IP group sources and sync: [WAF Design](./waf-design.md); graph model, execution semantics, release constraints: [WAF Orchestration Rule Design](./waf-orchestration-design.md)*
|
||||
|
||||
### 4. Edge Observability and Business Traffic Stats Flow
|
||||
```text
|
||||
OpenResty access.log (business facts)
|
||||
|
|
||||
| Agent tails incremental details (no sum/count/uniq)
|
||||
v
|
||||
Server stores via logstore (current log primary DB: PostgreSQL / SQLite / ClickHouse)
|
||||
|
|
||||
+---> global aggregation --> dashboard "data provided / requests / UV"
|
||||
+---> host∈Zone --> Zone "data provided" etc. (same semantics)
|
||||
+---> node_id filter --> node business volume
|
||||
|
||||
host /proc NIC, CPU etc. --> Agent reading snapshots --> host resource trends (displayed separately from business delivery)
|
||||
OpenResty health and connections --> edge health (instant, not 24h business totals)
|
||||
```
|
||||
* **Principle**: the Agent reports only facts; the Server interprets facts; access logs are the single truth for business traffic. `openresty_tx` and "data provided" must not run on dual tracks.
|
||||
* *Transport model, examples, and collection frequency: [Observability Transport Model](./observability-transport-model.md); field convergence and migration: [Edge Observability & Business Traffic Stats](./observability-design.md)*
|
||||
|
||||
### 5. Cloudflare DNS Pointing Flow
|
||||
|
||||
```text
|
||||
admin configures connection/group/member -> Server persists desired state -> Asynq sync tasks
|
||||
|
|
||||
v
|
||||
Cloudflare Zone / DNS API
|
||||
|
|
||||
v
|
||||
single A record -> active_node IPv4
|
||||
|
||||
node IP manually updated or Agent heartbeat change --------------------> best-effort enqueue per node
|
||||
```
|
||||
|
||||
* The Cloudflare module only manages cached or taken-over uniquely-named A records; it doesn't extend the Zone core into an authoritative DNS control plane. On multiple same-name A records it stops syncing and asks the admin to clean up in Cloudflare.
|
||||
* Group backup/active nodes are reserved for later failover; phase 1 fixes the primary node and doesn't auto-switch on heartbeat offline.
|
||||
* *Connection, model, idempotent sync, and phasing: [Cloudflare DNS Pointing Design](./cloudflare-pointing.md)*
|
||||
|
||||
---
|
||||
|
||||
## Core Objects
|
||||
|
||||
Current core system entities include:
|
||||
|
||||
* **Reverse proxy & config**: `zones` (root-domain management boundary), `zone_domains` (explicit domains with cert/route association), `proxy_routes` (route policy), `origins`, `config_versions`, `tls_certificates`. See [Zone & Domain Resource Design](./zone-design.md).
|
||||
* **Cloudflare DNS pointing**: `of_cf_connections` (global connection), `of_cf_pointing_groups` (primary/backup/active nodes and default orange-cloud), `of_cf_pointing_members` (ZoneDomain members, record cache, sync state). See [Cloudflare DNS Pointing Design](./cloudflare-pointing.md).
|
||||
* **Pages static hosting**: `of_pages_projects`, `of_pages_project_sources` / `of_pages_project_source_runtime` (mutable source config and runtime), `of_pages_deployments` (immutable deployments), `of_pages_deployment_files` (deployment file manifests).
|
||||
* **Nodes & tunnels**: `nodes`, `tunnels` (tunnel clients), `node_system_profiles`, `apply_logs`.
|
||||
* **WAF & security**: `waf_rule_groups`, `waf_ip_groups`, `waf_rule_group_bindings` (site WAF bindings).
|
||||
* **System & accounts**: `acme_accounts`, `dns_accounts`, `geoip_update_configs`.
|
||||
|
||||
---
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
| Decision | Reason |
|
||||
| --- | --- |
|
||||
| Full config versions instead of online patching | stable boundaries for preview, activation, history, and rollback; consistent node state |
|
||||
| Agent active pull | the Server needs no SSH access, lowering security risk; supports HTTP/WebSocket dual-protocol switching |
|
||||
| Globally single active version | lowers control-plane complexity, keeps all nodes consistent by default; stable one-click second-level rollback |
|
||||
| Zone domains separated from route policy | Zones provide the root-domain entry and domain boundaries; routes still reuse the same site-level policy and bind certs per domain |
|
||||
| Cloudflare pointing independent of the Zone core | ZoneDomains only provide explicit FQDNs; the Cloudflare module drives single A records from DB desired state without widening Zones into a general DNS control plane |
|
||||
| Intranet penetration integrated on frp | reuses a mature tunnel protocol, avoiding self-built tunnel stability risks; its Vhost mechanism natively fits reverse-proxy routes |
|
||||
| Runtime config decoupled from the control store | WAF rules compile at release and load with the OpenResty reload; dynamic IP groups refresh independently via checksum-driven memory snapshots |
|
||||
| Access logs as the single truth for business traffic | the Agent forbids business pre-aggregation; dashboard and Zone share Server-side aggregation, avoiding openresty_tx vs bytes_sent dual tracks |
|
||||
| Business delivery / edge health / host capacity layered | data provided ≠ host NIC outbound ≠ OpenResty connections; UI and API name and section them separately |
|
||||
| Pages artifacts separated from repo builds | current sources only import pre-built artifacts; future checkout/build happens in a Server-isolated executor reusing the artifact pipeline; the Agent never runs third-party builds |
|
||||
|
||||
---
|
||||
@@ -0,0 +1,223 @@
|
||||
# Cloudflare DNS Pointing Design
|
||||
|
||||
## Goals
|
||||
|
||||
Point **ZoneDomains (explicit FQDNs)** in OpenFlare to edge node IPs quickly via the Cloudflare API, replacing manual A-record edits in the CF console. Users organize domains into **pointing groups**: each group configures a primary node, a backup node, and a default orange-cloud (proxied) policy; members can override orange-cloud individually. The system treats the database tables as the desired state and idempotently syncs remote DNS.
|
||||
|
||||
This module is an **optional integration capability**; it does not turn Zones into an authoritative DNS control plane. Zones still only handle root-domain boundaries, domains, certificates, and reverse-proxy associations; A record create/update/delete is driven by this module through Cloudflare.
|
||||
|
||||
## Scope and Phasing
|
||||
|
||||
### Phase 1 (this design's scope)
|
||||
|
||||
* Sidebar **Cloudflare** entry with a Token-ready gate
|
||||
* Connection config: import from an existing DNS account **or** standalone entry within the module (mixed sources), stored encrypted
|
||||
* Pointing group CRUD: primary node, backup node (reserved), group default orange-cloud
|
||||
* Member management: add/remove at the granularity of `zone_domain_id`; member-level orange-cloud
|
||||
* Sync: write each member as a **single A record** on Cloudflare → current active node IPv4
|
||||
* Triggers: manual sync, member add, node/orange-cloud change, node IP change enqueue
|
||||
* Async task batch sync; member sync status with readable errors
|
||||
|
||||
### Phase 2
|
||||
|
||||
* Agent heartbeat offline detection of primary node failure → `active_node` switches to backup → whole-group auto sync
|
||||
* Optional auto failback, failure notification push
|
||||
|
||||
### Explicitly Out of Scope (later or permanent)
|
||||
|
||||
* Multiple parallel Cloudflare accounts (one global connection config)
|
||||
* AAAA / multi-A load balancing / CNAME to node hostnames
|
||||
* Managing MX/TXT/Page Rules and other non-module A records
|
||||
* Non-Cloudflare DNS providers
|
||||
* Merging DNS record management into the Zone core model
|
||||
|
||||
## Relationship with Existing Capabilities
|
||||
|
||||
| Existing Capability | Relationship |
|
||||
| --- | --- |
|
||||
| `of_zones` / `of_zone_domains` | Provide the pointable FQDN list; this module only references `zone_domain_id` |
|
||||
| `of_nodes.ip` | Source of A record `content`; recommended to restrict to edge nodes with valid IPv4 |
|
||||
| `of_dns_accounts` + `sealSensitive` | ACME DNS-01 already supports Cloudflare Token; this module can **import** the same account or store a Token standalone |
|
||||
| lego Cloudflare provider | **Only** TXT/DNS-01; this module builds its own CF HTTP client for Zone/DNS Record APIs |
|
||||
|
||||
## Core Model
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
CF_CONNECTIONS ||--o| DNS_ACCOUNTS : optional_import
|
||||
CF_POINTING_GROUPS ||--o{ CF_POINTING_MEMBERS : contains
|
||||
ZONE_DOMAINS ||--o| CF_POINTING_MEMBERS : pointed_as
|
||||
NODES ||--o{ CF_POINTING_GROUPS : primary
|
||||
NODES ||--o{ CF_POINTING_GROUPS : backup
|
||||
NODES ||--o{ CF_POINTING_GROUPS : active
|
||||
|
||||
CF_CONNECTIONS {
|
||||
uint id PK
|
||||
string source
|
||||
uint dns_account_id
|
||||
string authorization
|
||||
string status
|
||||
time verified_at
|
||||
}
|
||||
CF_POINTING_GROUPS {
|
||||
uint id PK
|
||||
string name
|
||||
uint primary_node_id
|
||||
uint backup_node_id
|
||||
uint active_node_id
|
||||
bool default_proxied
|
||||
bool enabled
|
||||
}
|
||||
CF_POINTING_MEMBERS {
|
||||
uint id PK
|
||||
uint group_id
|
||||
uint zone_domain_id UK
|
||||
bool proxied
|
||||
string cf_zone_id
|
||||
string cf_record_id
|
||||
string desired_ip
|
||||
string sync_status
|
||||
string last_error
|
||||
time synced_at
|
||||
}
|
||||
```
|
||||
|
||||
### `of_cf_connections` (one valid connection globally)
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `source` | `dns_account` \| `standalone` |
|
||||
| `dns_account_id` | references `of_dns_accounts` (type=cloudflare) when `source=dns_account` |
|
||||
| `authorization` | encrypted storage when `source=standalone`, payload shape `{"api_token":"..."}`, consistent with DNS accounts; the API **never returns it** |
|
||||
| `status` / `verified_at` | connectivity check result and time |
|
||||
|
||||
**Token resolution:** `dns_account` → decrypt the associated account; `standalone` → decrypt this row. Associated account deleted or validation failed → module not ready, sync forbidden.
|
||||
|
||||
**Recommended permissions:** Cloudflare API Token with `Zone:Read`, `DNS:Edit`.
|
||||
|
||||
### `of_cf_pointing_groups`
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `name` | display name |
|
||||
| `primary_node_id` | primary node |
|
||||
| `backup_node_id` | backup (nullable; phase 1 stores only) |
|
||||
| `active_node_id` | currently effective node; equals primary in phase 1; rewritten by phase 2 failover |
|
||||
| `default_proxied` | group default orange-cloud; **only affects newly added members** |
|
||||
| `enabled` | whether to participate in sync |
|
||||
|
||||
Constraints: primary and backup must not be the same node; the node chosen as the active target must have a valid IPv4.
|
||||
|
||||
### `of_cf_pointing_members`
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `group_id` | owning group |
|
||||
| `zone_domain_id` | globally unique: one domain belongs to at most one group |
|
||||
| `proxied` | member orange-cloud (the only runtime basis) |
|
||||
| `cf_zone_id` / `cf_record_id` | Cloudflare cache for idempotent updates |
|
||||
| `desired_ip` / `sync_status` / `last_error` / `synced_at` | desired and sync state |
|
||||
|
||||
`sync_status`: `pending` \| `syncing` \| `ok` \| `error`.
|
||||
|
||||
No physical foreign keys; `zone_domain_id` unique index; query indexes on `group_id` etc.
|
||||
|
||||
## Orange-Cloud Priority
|
||||
|
||||
1. **Member `proxied`**: the only basis written to CF during sync.
|
||||
2. **Group `default_proxied`**: copied to `proxied` when a member is **added**.
|
||||
3. Later changes to the group default **do not rewrite** existing members.
|
||||
|
||||
## Sync Semantics
|
||||
|
||||
### Desired State
|
||||
|
||||
The OpenFlare DB tables are the Source of Truth. Each member expects:
|
||||
|
||||
| Item | Value |
|
||||
| --- | --- |
|
||||
| type | `A` |
|
||||
| name | the ZoneDomain's FQDN |
|
||||
| content | the group `active_node`'s IPv4 |
|
||||
| proxied | member `proxied` |
|
||||
| ttl | forced Auto by CF when orange-cloud is on; unified default (e.g. 300) when off |
|
||||
|
||||
Phase 1 does not write AAAA. Node IP not a valid IPv4 → that member is `error`.
|
||||
|
||||
### Triggers
|
||||
|
||||
| Trigger | Behavior |
|
||||
| --- | --- |
|
||||
| Manual sync (all / group / member) | reconcile |
|
||||
| Member added | initialize `proxied`, then enqueue sync |
|
||||
| Member removed / group deleted | delete the remote A managed by this module by default (configurable keep) |
|
||||
| Primary node / active / member `proxied` changed | re-sync the corresponding scope |
|
||||
| Node IP changed (heartbeat or manual) | enqueue members whose `active_node_id` points to that node |
|
||||
| Token not ready | refuse sync |
|
||||
|
||||
Phase 1 does not do scheduled full reconciliation.
|
||||
|
||||
### Reconcile (single member, idempotent)
|
||||
|
||||
1. Resolve the CF Zone by the FQDN's registrable root domain, cache `cf_zone_id`.
|
||||
2. With a `cf_record_id`, prefer Update; if stale, list by `name+type=A`.
|
||||
3. **0 records** → Create; **exactly 1** → take over and Update; **multiple** → fail and tell the user to clean up in CF.
|
||||
4. Write back `cf_record_id`, `desired_ip`, `sync_status`, `synced_at` / `last_error`.
|
||||
5. On rate limiting, retry with bounded backoff.
|
||||
|
||||
**Ownership:** only manage records cached by this module or taken over as "the only same-name A"; do not clear the Zone or touch other record types. After a user edits in the CF console, the next sync overwrites with the OpenFlare desired state.
|
||||
|
||||
### Execution Carrier
|
||||
|
||||
* Single record: can sync on the request path.
|
||||
* Whole group / per-node batch: Asynq tasks (`cloudflare:sync_member` / `sync_group` / `sync_by_node`), registered in `bootstrap`.
|
||||
* Per-member mutex to prevent concurrent double-writes.
|
||||
* Node IP change path delivers tasks **best-effort**, not blocking the heartbeat.
|
||||
|
||||
## API (Admin Panel)
|
||||
|
||||
Prefix: `/api/v1/d/cloudflare`, Session admin auth. Package: `internal/apps/openflare/cloudflare/`; routes: `internal/router/v1/openflare/register_cloudflare.go`.
|
||||
|
||||
| Resource | Method & Path |
|
||||
| --- | --- |
|
||||
| Connection | `GET/PUT /connection`, `POST /connection/verify`, `POST /connection/clear` |
|
||||
| Overview | `GET /overview` |
|
||||
| Groups | `GET/POST /groups`, `GET /groups/:id`, `POST /groups/:id/update|delete|sync` |
|
||||
| Members | `GET/POST /groups/:id/members`, `POST .../members/:memberId/update|remove|sync` |
|
||||
| Available domains | `GET /domains/available` |
|
||||
|
||||
* Success `response.OK`; failure `response.Abort*`; the Token is **never** returned in JSON.
|
||||
* Handlers separated from `logics.go`; the CF client is abstracted behind an interface for replaceability.
|
||||
|
||||
## Frontend
|
||||
|
||||
* Navigation: `frontend/lib/navigation/openflare-nav.ts` adds **Cloudflare** → `/cloudflare` (near Website Management / DNS Accounts).
|
||||
* Routes:
|
||||
* `/cloudflare`: overview; guide to configure when not ready
|
||||
* `/cloudflare/settings`: mixed Token config and test connection
|
||||
* `/cloudflare/groups`, `/cloudflare/groups/[id]`: list and detail (members, orange-cloud, sync)
|
||||
* Services: independent service under `frontend/lib/services/openflare/`, extending `BaseService`.
|
||||
* Pages follow the existing title-bar and component-split conventions; destructive actions need double confirmation.
|
||||
* Copy that must be visible: sync overwrites module-managed A records; multiple same-name A records need manual cleanup; removal deletes remote records by default; phase 1 has no automatic failover.
|
||||
|
||||
## Errors and Security
|
||||
|
||||
* User-visible copy is module-internal constants; internal errors log via `pkg/logger`.
|
||||
* Typical: token not configured, invalid token, node without IP, no CF Zone, multiple same-name A records, rate limiting.
|
||||
* Token is only decrypted server-side for use; responses and logs must never contain plaintext tokens.
|
||||
|
||||
## Data Migration
|
||||
|
||||
* goose both dialects (PG/SQLite) create the three tables; defaults match Go zero values.
|
||||
|
||||
## Key Decision Summary
|
||||
|
||||
| Decision | Conclusion |
|
||||
| --- | --- |
|
||||
| Module shape | standalone Cloudflare pointing module, not embedded Zone fields |
|
||||
| Token | mixed: imported from DNS account or encrypted standalone |
|
||||
| Domain granularity | ZoneDomain (FQDN) |
|
||||
| Record shape | single A → active node IPv4 |
|
||||
| Failover | phase 2; heartbeat offline; phase 1 only stores backup/active |
|
||||
| Orange-cloud | member-level effective; group default only initializes |
|
||||
| SoT | DB tables as desired state drive CF |
|
||||
@@ -0,0 +1,264 @@
|
||||
# Edge Cache Strategy Design
|
||||
|
||||
You will learn: how OpenFlare's edge `proxy_cache` aligns with the Cloudflare default loop between "should cache" and "should not cache": request eligibility (extension/policy) × response shareability (origin `Cache-Control` / `Expires` / `Set-Cookie`), and the differences from the previous over-strict request bypass.
|
||||
|
||||
This design is the productized chapter on "basic caching" in [System Architecture](./architecture.md); cache results in access logs are in [Observability Data Model §3.5.1](./observability-data-model.md).
|
||||
|
||||
---
|
||||
|
||||
## 1. Goals and Non-Goals
|
||||
|
||||
### 1.1 Goals
|
||||
|
||||
* **Close to CF default out of the box**: after enabling cache on a route, **only static extensions are cached by default** — HTML is not cached by default; **request session cookies / Authorization / client Cache-Control no longer cause a blanket BYPASS**.
|
||||
* **Cacheable content hits**: a logged-in user visiting `/_app/**/*.js` and other static assets can show `MISS` → `HIT`.
|
||||
* **Non-cacheable stays blocked**: policy not eligible (equivalent to CF `DYNAMIC`); origin `private` / `no-store`; responses with **`Set-Cookie` not stored** (aligned with CF OCC default); `all` is an advanced option with documented warnings.
|
||||
* **Default Edge TTL when no origin freshness**: aligned with CF's per-status default TTL (see §3.5).
|
||||
* **Consistent observability**: keep relying on `$upstream_cache_status` → three-state `cache_status` detail.
|
||||
* **Backward compatible**: legacy route `cache_policy=url` maps to `all`; policy enum and migration rules stay in [§5](#5-兼容与迁移).
|
||||
|
||||
### 1.2 Non-Goals (later iterations)
|
||||
|
||||
* Cache Rules expression engine
|
||||
* Forced Edge TTL ignoring origin `Cache-Control` (CF Cache Rules "Ignore cache-control")
|
||||
* Purge (by URL/prefix/site-wide)
|
||||
* Browser TTL rewriting, client `CF-Cache-Status` response header
|
||||
* Full RFC conditions: `Authorization` cached only when the response has `public`/`s-maxage`/`must-revalidate` (needs Lua; this iteration deletes the request-side bypass entirely, relying on policy + origin headers)
|
||||
* HEAD → GET conversion then cache
|
||||
* Hit-rate dashboard
|
||||
|
||||
---
|
||||
|
||||
## 2. Cloudflare Decision Loop (Alignment Baseline)
|
||||
|
||||
CF default is a **two-stage decision**, **not** "request has Cookie → don't cache".
|
||||
|
||||
### 2.1 Stage A — Eligible at Request Time
|
||||
|
||||
| Condition | CF Result |
|
||||
| --- | --- |
|
||||
| Non-GET | not cached by default |
|
||||
| Extension not in default cacheable table, no Rules forcing eligible | **`DYNAMIC`** (no cache lookup) |
|
||||
| Extension in default table, or Rules eligible | continue to Stage B |
|
||||
| **Request Cookie** | **no effect by default** |
|
||||
| Cache Rules Bypass | `DYNAMIC` |
|
||||
|
||||
CF's default cacheable extensions are keyed by **extension** rather than MIME; **HTML / JSON are not cached by default**.
|
||||
|
||||
### 2.2 Stage B — Response Storeable (OCC on, Free/Pro/Biz default)
|
||||
|
||||
| Condition | Result |
|
||||
| --- | --- |
|
||||
| `Cache-Control: no-store` / `private` | not stored |
|
||||
| `public` + `max-age>0`, or future `Expires` | cacheable |
|
||||
| No Cache-Control / Expires | still cacheable with per-status **default Edge TTL** (e.g. 200 → 120m) |
|
||||
| Response **`Set-Cookie`** (default cache level + OCC) | **not stored**, status tends toward **BYPASS** |
|
||||
| Request `Authorization` | cacheable only when the response also has `public` / `s-maxage` / `must-revalidate` (full condition simplified with Nginx this iteration, see §3.4) |
|
||||
|
||||
### 2.3 Status Semantics (vs. Observability)
|
||||
|
||||
| CF | Meaning | OpenFlare `cache_status` |
|
||||
| --- | --- | --- |
|
||||
| HIT / STALE / UPDATING / REVALIDATED | hit class | same-name or equivalent |
|
||||
| MISS / EXPIRED | fetch from origin | same-name |
|
||||
| BYPASS | eligible at request time, response not cacheable | `BYPASS` → UI "not cached" |
|
||||
| DYNAMIC | not eligible at request time | policy skip mostly `BYPASS` or empty → UI "not cached" |
|
||||
|
||||
---
|
||||
|
||||
## 3. Product Semantics
|
||||
|
||||
### 3.1 Two-Level Switch (unchanged)
|
||||
|
||||
* **Global** `openresty_cache_enabled`: generates `proxy_cache_path` etc.; when off, route-level cache directives are inert.
|
||||
* **Route** `cache_enabled`: whether to enable `proxy_cache` in that site's `location`.
|
||||
|
||||
Cache logic only runs when both are on.
|
||||
|
||||
### 3.2 Policy Enum
|
||||
|
||||
| `cache_policy` | Meaning | New Default | Legacy Compatibility |
|
||||
| --- | --- | --- | --- |
|
||||
| **`static`** | only eligible when URI matches **standard static extensions** | **yes** | — |
|
||||
| **`all`** | after method bypass, no path/extension restriction (advanced; risk similar to CF Cache Everything) | no | legacy `url` → `all` |
|
||||
| **`suffix`** | custom extension list (`cache_rules`) | no | kept |
|
||||
| **`path_prefix`** | custom path prefix | no | kept |
|
||||
| **`path_exact`** | custom exact path | no | kept |
|
||||
|
||||
Render layer: historical `url` is treated as `all`; API/UI only expose the enum above.
|
||||
|
||||
### 3.3 Standard Static Extensions (built-in)
|
||||
|
||||
Aligned with CF default "no HTML/JSON caching"; keeps modern frontend-friendly enhancements:
|
||||
|
||||
```text
|
||||
css js mjs map
|
||||
ico cur gif jpg jpeg png webp avif svg svgz
|
||||
ttf otf woff woff2 eot
|
||||
mp3 mp4 webm ogg flac
|
||||
wasm pdf
|
||||
zip 7z gz tar
|
||||
```
|
||||
|
||||
* **Excludes** `html` / `htm` / **`json`** (aligned with CF not caching JSON by default).
|
||||
* **Includes** `map` / `mjs` / `wasm` (deliberate enhancement for sourcemap / ES module / WASM hits).
|
||||
* Matching: `$uri` extension, case-insensitive:
|
||||
`if ($uri !~* \.(?:css|js|…)$) { set $openflare_skip_cache 1; }`
|
||||
|
||||
### 3.4 Request-Side Bypass (after CF alignment)
|
||||
|
||||
Only kept:
|
||||
|
||||
1. `$request_method != GET` (HEAD included, consistent with current network; no CF HEAD→GET)
|
||||
|
||||
**Removed** (previously over-strict, causing low hit rates):
|
||||
|
||||
* Session-cookie regex
|
||||
* `$http_authorization != ""`
|
||||
* request `$http_cache_control` matching `no-cache|no-store|private`
|
||||
|
||||
**How security still holds:**
|
||||
|
||||
| Threat | Gate |
|
||||
| --- | --- |
|
||||
| Accidentally caching HTML/API | default `static` extensions (no html/json) |
|
||||
| Personalized content | origin `private` / `no-store` (respected by Nginx) |
|
||||
| Response writes session | **`Set-Cookie` → not stored** (§3.6) |
|
||||
| `all` too broad | UI/doc warning: needs correct origin Cache-Control |
|
||||
| API with Bearer | rely on policy (don't use `all` for APIs) + origin headers; full Auth conditional caching is later |
|
||||
|
||||
### 3.5 Default Edge TTL (no origin freshness)
|
||||
|
||||
Aligned with CF's per-status default TTL without `Cache-Control`/`Expires`, emitted in cache-enabled locations:
|
||||
|
||||
| Status | TTL |
|
||||
| --- | --- |
|
||||
| 200, 206, 301 | 120m |
|
||||
| 302, 303 | 20m |
|
||||
| 404, 410 | 3m |
|
||||
|
||||
```nginx
|
||||
proxy_cache_valid 200 206 301 120m;
|
||||
proxy_cache_valid 302 303 20m;
|
||||
proxy_cache_valid 404 410 3m;
|
||||
```
|
||||
|
||||
* When the origin provides valid `Cache-Control` / `Expires`, the origin freshness wins (no `proxy_ignore_headers`).
|
||||
* **No** forced Edge TTL override ignoring origin headers.
|
||||
|
||||
### 3.6 Response Side: Set-Cookie Not Stored
|
||||
|
||||
Aligned with CF OCC default: an eligible request whose origin returns **`Set-Cookie`** is **not written** into `proxy_cache` (read path may still have MISS/BYPASS semantics).
|
||||
|
||||
```nginx
|
||||
proxy_no_cache $openflare_skip_cache $upstream_http_set_cookie;
|
||||
```
|
||||
|
||||
(`proxy_no_cache` multi-arg: any non-empty and non-`"0"` arg means no write.)
|
||||
|
||||
`proxy_cache_bypass` still only binds `$openflare_skip_cache` (request-side skip); the response side only affects **writes**, consistent with CF "eligible but response not cacheable".
|
||||
|
||||
### 3.7 Relationship with Origin Headers
|
||||
|
||||
* **Eligibility**: policy + method bypass.
|
||||
* **Store / duration**: origin `Cache-Control` / `Expires` + default `proxy_cache_valid` + Set-Cookie gate + global `inactive`.
|
||||
|
||||
---
|
||||
|
||||
## 4. Rendering and Data Flow
|
||||
|
||||
```text
|
||||
Global cache_enabled?
|
||||
│ no → no proxy_cache_* generated
|
||||
▼ yes
|
||||
Route cache_enabled?
|
||||
│ no → location without proxy_cache
|
||||
▼ yes
|
||||
set $openflare_skip_cache 0
|
||||
→ non-GET → set 1
|
||||
→ policy if (static/all/suffix/…) → may set 1
|
||||
proxy_cache openflare_cache
|
||||
proxy_cache_methods GET
|
||||
proxy_cache_bypass $openflare_skip_cache
|
||||
proxy_no_cache $openflare_skip_cache $upstream_http_set_cookie
|
||||
proxy_cache_valid …
|
||||
→
|
||||
access.log cache_status=$upstream_cache_status
|
||||
```
|
||||
|
||||
### 4.1 Policy → Nginx Conditions
|
||||
|
||||
| Policy | Extra Condition |
|
||||
| --- | --- |
|
||||
| `static` | `$uri` not matching built-in extension table → skip |
|
||||
| `all` | no extra path condition |
|
||||
| `suffix` | not matching `cache_rules` extensions → skip |
|
||||
| `path_prefix` / `path_exact` | same as current implementation |
|
||||
|
||||
### 4.2 Code Areas Involved
|
||||
|
||||
| Area | Path |
|
||||
| --- | --- |
|
||||
| Rendering | `pkg/render/openresty/render.go` (bypass, Set-Cookie, `proxy_cache_valid`, extension constants) |
|
||||
| Validation | `internal/apps/openflare/proxy_route/helpers.go` |
|
||||
| Model/defaults | creating a route defaults `cache_policy=static`; `url`→`all` on read/write |
|
||||
| Snapshot | `config_version` snapshot normalization |
|
||||
| UI | `proxy-routes/detail/components/cache-section.tsx` |
|
||||
|
||||
---
|
||||
|
||||
## 5. Compatibility and Migration
|
||||
|
||||
| Data | Handling |
|
||||
| --- | --- |
|
||||
| `cache_policy=''` or `url` in DB (and cache enabled) | read / snapshot / render → **`all`** |
|
||||
| API write with enabled and empty policy | normalized to **`all`**; UI new-create with cache on **explicitly submits** `static` |
|
||||
| New routes | default **`static`** when cache enabled |
|
||||
| Bypass behavior change | **breaking vs. old implementation**: cookie/auth traffic goes from "not cached" to cacheable HIT; requires **republishing node configs** |
|
||||
| Default extensions | **remove `json`** from the table; sites relying on caching `*.json` can use custom `suffix` or `all` |
|
||||
|
||||
**Release note:** document this alignment with the CF default model; hit rate expected to rise; `all` and wrong origin headers need ops self-check.
|
||||
|
||||
---
|
||||
|
||||
## 6. UI Copy Points (Cache Tab)
|
||||
|
||||
* After enabling cache, default: **standard static assets** (summary extensions, **excluding HTML/JSON**; including map/mjs etc.).
|
||||
* Options: standard static / all cacheable GET (advanced) / custom suffix / path prefix / exact path.
|
||||
* CF-aligned notes:
|
||||
* login cookies are **not** separately skipped from caching;
|
||||
* origin `private` / `no-store` / response **`Set-Cookie`** are not written to the edge cache;
|
||||
* default Edge TTL used when no origin cache headers.
|
||||
* **Advanced `all`**: warn "similar to Cache Everything; personalized pages must declare private/no-store from the origin".
|
||||
* Global Performance cache master switch must be on.
|
||||
|
||||
---
|
||||
|
||||
## 7. Decision Matrix (Avoid Missed Judgments)
|
||||
|
||||
| Scenario | CF | OpenFlare (this design) |
|
||||
| --- | --- | --- |
|
||||
| GET static + session Cookie + origin public max-age | HIT | HIT |
|
||||
| GET HTML + static policy | DYNAMIC | policy skip → not cached |
|
||||
| GET + all + origin private | not stored | not stored |
|
||||
| GET static + response Set-Cookie | BYPASS (OCC) | not stored |
|
||||
| GET + Authorization + static public | conditional cache | cacheable (simplified; rely on origin not marking sensitive APIs public) |
|
||||
| GET + no-CC 200 static | default 120m | `proxy_cache_valid` 120m |
|
||||
| DevTools Disable cache (request no-cache) | edge may still HIT by default | edge may still HIT by default |
|
||||
| POST | not cached | non-GET skip |
|
||||
|
||||
---
|
||||
|
||||
## 8. Decision Record
|
||||
|
||||
| Decision | Choice | Reason |
|
||||
| --- | --- | --- |
|
||||
| Request Cookie bypass | **removed** | aligned with CF; restore static hit rate for logged-in users |
|
||||
| Request Authorization / Cache-Control bypass | **removed** | aligned with CF request-eligibility model; response gate as backstop |
|
||||
| Set-Cookie | **bind to proxy_no_cache** | aligned with CF OCC "response Set-Cookie not stored" |
|
||||
| Default Edge TTL | **per-status proxy_cache_valid** | aligned with CF default TTL when headerless, avoiding "never stored" |
|
||||
| Remove json from default table | **yes** | aligned with CF not caching JSON by default |
|
||||
| Keep map/mjs/wasm | **yes** | useful hits for modern frontend, deliberate enhancement |
|
||||
| Default cacheable scope | cache-on defaults to `static` | benchmarked to CF, reduces HTML/API mis-caching |
|
||||
| Legacy `url` | maps to `all` | doesn't narrow existing behavior |
|
||||
| Full Auth conditions / Purge / Rules | later | close the default loop first, then extend |
|
||||
@@ -0,0 +1,191 @@
|
||||
# Product Boundaries
|
||||
|
||||
You will learn: what OpenFlare is, its current stable capabilities, and the core product boundaries and repository structure layout you must follow when developing.
|
||||
|
||||
OpenFlare is a self-hosted OpenResty control plane for single-team or single-organization internal operations.
|
||||
|
||||
---
|
||||
|
||||
## Project Positioning
|
||||
|
||||
OpenFlare suits teams that need to centrally manage multiple OpenResty proxy nodes, with this positioning:
|
||||
* **Control/landing separation**: the Server control plane doesn't SSH into proxy nodes; Agents actively pull versions and apply them.
|
||||
* **Immutable config release**: full config versions are used for preview, release, activation, and one-click rollback.
|
||||
* **Integrated gateway hosting**: website reverse proxying, automatic TLS certificate issuance/renewal, WAF protection, intranet penetration (Tunnel), and Pages static hosting are integrated into one control plane.
|
||||
|
||||
**Not this product's positioning**: multi-tenant cloud platforms, Kubernetes Ingress Controllers, service meshes, or general log platforms.
|
||||
|
||||
---
|
||||
|
||||
## Current Capabilities
|
||||
|
||||
| Capability | Description | Detailed Design/Usage |
|
||||
| --- | --- | --- |
|
||||
| **Reverse proxy config management** | website rules (Proxy Route) as the aggregation boundary; multi-domain and multi-upstream load balancing | [Create a Reverse Proxy Config](../guide/proxy-config.md) |
|
||||
| **Origin error page** | globally configurable: matching origin/gateway status codes return OpenFlare default or custom HTML with the HTTP status kept | [Origin Error Page Design](./origin-error-page.md) |
|
||||
| **Edge cache** | single-node OpenResty `proxy_cache`; default static extensions + origin-header/Set-Cookie gates + default Edge TTL (benchmarked to the CF default model) | [Edge Cache Strategy Design](./edge-cache-design.md) |
|
||||
| **Zone & domain management** | registrable root domains as the management entry, aggregating explicit domains, domain certificates, and reverse proxy routes | [Zone & Domain Resource Design](./zone-design.md) |
|
||||
| **Cloudflare DNS pointing** | per ZoneDomain, idempotently point a single Cloudflare A record at an edge node IPv4; connection config, groups, member orange-cloud, and async sync; no auto-failover in phase 1 | [Cloudflare DNS Pointing Design](./cloudflare-pointing.md) |
|
||||
| **Config versioning** | global single active version with preview, release, immutable snapshot history, and second-level one-click rollback | [Agent & Publish Model](./agent-design.md) |
|
||||
| **WAF protection** | visual DAG rule orchestration, manual/auto/subscription IP groups, GeoIP matching, and PoW CC protection | [WAF Design](./waf-design.md) / [WAF Orchestration Rule Design](./waf-orchestration-design.md) / [WAF Usage Guide](../guide/waf-usage.md) |
|
||||
| **Intranet penetration** | reverse-penetrate and expose intranet web services via Relay nodes and the OpenFlared client | [Tunnel Design](./tunnel-design.md) / [Tunnel Usage Guide](../guide/tunnel-usage.md) |
|
||||
| **Pages static hosting** | upload or sync pre-built artifacts from Remote URLs or public GitHub Releases; GitHub latest can be periodically checked and optionally auto-published. Immutable deployments are pulled by edge nodes and served locally by OpenResty, supporting rollback, API proxying, and SPA Fallback | [Pages Static Hosting Design](./pages-design.md) / [Pages Usage Guide](../guide/pages-usage.md) |
|
||||
| **TLS certificate auto-renewal** | explicitly bind certificates to Zone domains; issue/renew via ACME against Let's Encrypt | [Zone & Domain Resource Design](./zone-design.md) |
|
||||
| **Multi-node monitoring & observability** | access logs as the single truth for business traffic; Agent reports only details and host readings, Server aggregates uniformly; reconciled with Zone/dashboard | [Observability Transport Model](./observability-transport-model.md) / [Edge Observability & Business Traffic Stats](./observability-design.md) / [Reporting Protocol & Tables](./observability-data-model.md) / [System Architecture](./architecture.md) |
|
||||
| **Log storage** | access logs and observability time series use the switchable log primary DB (follows the business primary DB or ClickHouse); still writable/queryable with ClickHouse off | [Log Store Decoupling](./logstore.md) |
|
||||
| **Console bilingual** | zh-CN / en without URL prefixes, `NEXT_LOCALE` cookie precedence, static-export compatible | [Frontend i18n design](../superpowers/specs/2026-07-24-frontend-i18n-design.md) |
|
||||
|
||||
---
|
||||
|
||||
## Core Product Boundaries and Constraints
|
||||
|
||||
When developing and contributing code, **you must strictly follow** these business boundaries and technical constraints; don't bypass them for temporary needs:
|
||||
|
||||
### 1. Website Config and Upstream Constraints
|
||||
* **Single-site domain sharing policy**: one route rule corresponds to one website; the site's multiple domains share rate limit, cache, and reverse-proxy upstream config. Differential per-domain service config within the same rule is not supported.
|
||||
* **Upstream type mutual exclusion**: the upstream must be one of direct address (`direct`), intranet tunnel (`tunnel`), or Pages static hosting (`pages`); mixing within one rule is not allowed.
|
||||
* **Direct type restrictions**: a direct upstream can be a single or multiple pure `http://` or `https://` addresses (multi-address only supports plain `scheme://host[:port]`); non-HTTP protocols (TCP/UDP) upstreams are not supported.
|
||||
|
||||
### 2. WAF Security Boundaries
|
||||
* **Allowlist priority**: the allowlist has absolute matching power. Only when an allowlist rule isn't hit do the global and custom blocklist filters trigger in order.
|
||||
* **GeoIP weak dependency**: geo access resolution fully depends on the node-local MaxMind DB. When GeoIP is abnormal or fails to resolve, the system must auto-ignore geo rules — **never** break IP-group filtering or the reverse-proxy main chain's availability.
|
||||
* **Runtime data decoupling**: OpenResty interception only reads Agent-synced local JSON, never talking to the Server DB. IP group member sync is decoupled from version release via Checksum differential pull for zero-reload smooth effect.
|
||||
|
||||
### 3. Intranet Penetration Boundaries
|
||||
* **HTTP traffic only**: the tunnel components only support HTTP/HTTPS (based on frp's vhost mechanism for single-port domain-route reuse); standalone TCP/UDP port allocation is not supported yet.
|
||||
* **Dynamic relay config control**: a Relay node, after connecting to the Server, dynamically pulls and syncs global system config via heartbeats (e.g. whether the embedded FRPS Web UI and its port are enabled), but isn't part of the control plane's immutable config version release system.
|
||||
* **Tunnel/Node system isolation**: Tunnel clients make outbound connections from the intranet and are independent entities from control-plane-hosted edge Nodes (public nodes), authenticated with the dedicated `tunnel_token`.
|
||||
|
||||
### 4. Pages Static Hosting Boundaries
|
||||
* **Pre-built artifact sources**: a project may stay manual-upload, or configure one Remote URL / public GitHub Release asset source. Remote and fixed tags only support manual ops; only GitHub latest enters scheduled checks and can opt into auto-update. Sources are switchable, but immutable deployments and the current production version don't get lost when editing or deleting a source.
|
||||
* **Archive and resource limits**: supports `zip`, `tar.gz` / `tgz`, `tar.xz` / `txz`, `tar.bz2` / `tbz2`, `tar`, `7z`. Archive cap controlled by `pages_max_package_size_mb` (default 100 MiB, range 1–2048); expanded single-file and total limits are 4× the package cap with a 100 MiB floor, at most 1,000 regular files. Both Server and Agent validate actual bytes and reject path traversal, symlinks/hard links, and special files.
|
||||
* **Build and runtime boundaries**: currently no source checkout or build execution from external git repos, and no edge Serverless, dynamic SSR, or preview subdomains. Future repo integration must use a separate `git_repository` Provider with a Server-side isolated build executor, emitting only restricted artifacts into the unified artifact pipeline; the Agent never receives repo credentials, external URLs, or clone/install/build commands.
|
||||
|
||||
### 5. System and Version Boundaries
|
||||
* **Globally single active version**: all nodes pull and consume the same globally active config. Per-node-group differentiated config release isn't performed.
|
||||
* **Single-tenant architecture**: OpenFlare is for a single team deploying on a trusted internal network. Single-tenant by design; fine-grained multi-user roles or multi-tenant resource isolation aren't supported.
|
||||
* **External infra dependency**: the Server **must depend on** external Redis (or Valkey) for distributed coordination, the Asynq queue, and system cache. The relational DB is PostgreSQL, or SQLite when `database.enabled` is off. ClickHouse **optional**: when off, access logs and observability time series are handled by the current log primary DB (follows the business primary DB); when on, the「Switch Log Database」task can migrate to ClickHouse. Running without Redis is not supported. See [Log Store Decoupling](./logstore.md).
|
||||
|
||||
---
|
||||
|
||||
## Repository Structure
|
||||
|
||||
OpenFlare has converged to a **single monorepo** (Go module `github.com/Rain-kl/Wavelet`). The control-plane Server and edge components (Agent, Relay, OpenFlared) share the repo, organized by Wavelet `internal/apps/` domain modules.
|
||||
|
||||
When contributing code, strictly follow this physical layering and directory division:
|
||||
|
||||
| Path | Responsibility |
|
||||
| --- | --- |
|
||||
| `main.go` | the Server's single entry, delegating to `internal/cmd/` |
|
||||
| `cmd/agent`, `cmd/relay`, `cmd/flared` | edge component CLI entries (**not** the Server) |
|
||||
| `internal/` | control-plane and edge runtime implementations |
|
||||
| `frontend/` | Next.js admin panel; build artifacts embedded into the Go Server |
|
||||
| `pkg/` | cross-component shared libs (protocol, rendering, GeoIP, etc.) |
|
||||
| `scripts/` | Swagger generation, install scripts, etc. |
|
||||
| `docs/` | VitePress docs site and design baseline |
|
||||
| `docker/` | per-component Dockerfiles |
|
||||
| `uploads/`, `data/` | runtime upload dir and static data (`.gitignore`d) |
|
||||
|
||||
### 1. Server Layering (`main.go` + `internal/`)
|
||||
|
||||
| Directory | Responsibility |
|
||||
| --- | --- |
|
||||
| `main.go` | Server startup entry |
|
||||
| `internal/cmd/` | Cobra subcommands: `api`, `worker`, `scheduler`, `all` (default fused mode) |
|
||||
| `internal/platform/bootstrap/` | cross-module assembly: task handlers, push domain events, process-level init |
|
||||
| `internal/router/` | HTTP route registration and global middleware |
|
||||
| `internal/router/v1/openflare/` | OpenFlare route registrars (`register_*.go`) |
|
||||
| `internal/apps/openflare/` | OpenFlare control-plane business domains (`routers.go` + `logics.go`) |
|
||||
| `internal/apps/{admin,user,oauth,upload,cap,...}/` | Wavelet platform capabilities (users, auth, tasks, push, etc.) |
|
||||
| `internal/apps/openflare/{agent,relay,flared}/` | **Server-side** edge protocol handlers (auth, heartbeat, WS) |
|
||||
| `internal/model/` | GORM entities / DTOs / no-IO domain rules (`openflare_*.go` + platform models); **no** DB access |
|
||||
| `internal/infra/persistence/migrator/goose/` | goose SQL migrations (PostgreSQL / SQLite / ClickHouse) |
|
||||
| `internal/repository/` | data access layer (platform + OpenFlare business CRUD, cache, `logstore` log IO); the **only** persistence entry |
|
||||
| `internal/infra/task/` | Asynq async tasks (Worker + Scheduler) |
|
||||
| `internal/infra/config/` | Viper config loading |
|
||||
| `internal/shared/` | unified API response wrapper (`response/`) |
|
||||
| `pkg/protocol/` | Relay / Tunnel shared HTTP/WS protocol structures |
|
||||
| `pkg/render/`, `pkg/geoip/`, `pkg/wsclient/` | OpenResty config rendering, GeoIP, WebSocket client |
|
||||
|
||||
**API route prefixes:**
|
||||
|
||||
| Prefix | Purpose | Auth |
|
||||
| --- | --- | --- |
|
||||
| `/api/v1/d/*` | OpenFlare admin console API | Session Cookie + optional `X-Access-Token` |
|
||||
| `/api/v1/agent/*` | Agent node protocol | `X-Agent-Token` |
|
||||
| `/api/v1/relay/*` | Relay protocol | `X-Agent-Token` |
|
||||
| `/api/v1/tunnel/*` | Tunnel client protocol | `X-Tunnel-Token` |
|
||||
| `/api/v1/admin/*` | Wavelet platform admin API | admin Session |
|
||||
|
||||
### 2. Agent Modules (`internal/apps/agent/` / `cmd/agent/`)
|
||||
| `internal/apps/agent/httpclient/` | Server communication |
|
||||
| `internal/apps/agent/wsclient/` | WebSocket client communication |
|
||||
| `internal/apps/agent/protocol/` | Agent API protocol types |
|
||||
| `internal/apps/agent/updater/` | Agent self-update logic |
|
||||
| `internal/apps/agent/logging/` | logging |
|
||||
| `internal/apps/agent/observability/`| observability (metrics, traces, etc.) |
|
||||
| `internal/apps/agent/geoipdata/` | GeoIP data handling |
|
||||
| `internal/apps/agent/geoipupdate/` | GeoIP data updates |
|
||||
| `internal/apps/agent/agent/` | core Agent logic and lifecycle |
|
||||
|
||||
### 3. Frontend Layering (`frontend/`)
|
||||
|
||||
Based on the Wavelet Next.js scaffold, OpenFlare business UI is organized route-co-located under `app/(main)/`.
|
||||
|
||||
| Directory | Responsibility |
|
||||
| --- | --- |
|
||||
| `app/` | Next.js App Router; `(main)` console, `(auth)` auth, `(docs)` docs pages |
|
||||
| `app/(main)/<domain>/` | business pages and in-domain components (route-co-located) |
|
||||
| `components/` | cross-domain reusable UI (`ui/`, `layout/`, `common/`, etc.) |
|
||||
| `lib/services/` | API service layer: `core/` base class + `openflare/` business APIs |
|
||||
| `lib/navigation/` | OpenFlare sidebar nav config (`openflare-nav.ts`) |
|
||||
| `lib/theme/` | theme parsing and switching |
|
||||
| `contexts/` | cross-page UI state (user, notifications, etc.) |
|
||||
| `hooks/`, `lib/hooks/` | reusable React Hooks |
|
||||
| `public/` | static assets and theme CSS |
|
||||
| `scripts/` | build helper scripts |
|
||||
| `proxy.ts` | dev/prod proxy: API rate limit and page auth |
|
||||
|
||||
**API conventions**: OpenFlare business APIs uniformly prefix `/api/v1/d/*`, wrapped via `OpenFlareBaseService`; page data fetching uses `@tanstack/react-query`.
|
||||
|
||||
### 4. Relay Modules (`internal/apps/relay/` / `cmd/relay/`)
|
||||
|
||||
| Module | Responsibility |
|
||||
| --- | --- |
|
||||
| `cmd/relay/` | Relay CLI entry and init main |
|
||||
| `internal/apps/relay/config/` | local config parsing and default init |
|
||||
| `internal/apps/relay/frps/` | manage frps process lifecycle, ports & Token, monitor runtime |
|
||||
| `internal/apps/relay/heartbeat/` | periodic HTTP heartbeat, report state, fetch update requests |
|
||||
| `internal/apps/relay/httpclient/` | generic Server API client helpers |
|
||||
| `internal/apps/relay/observability/` | collect local host and frps base runtime metrics with pre-aggregation |
|
||||
| `internal/apps/relay/relay/` | coordinate core lifecycle, init, and cleanup |
|
||||
| `internal/apps/relay/state/` | local runtime state, error records, persistent cache |
|
||||
| `internal/apps/relay/updater/` | Relay upgrade check, download/install, restart |
|
||||
| `internal/apps/relay/wsclient/` | long-lived WebSocket bidirectional channel with the Server |
|
||||
|
||||
### 5. OpenFlared (Client) Modules (`internal/apps/flared/` / `cmd/flared/`)
|
||||
|
||||
| Module | Responsibility |
|
||||
| --- | --- |
|
||||
| `cmd/flared/` | Client CLI entry and init main |
|
||||
| `internal/apps/flared/config/` | local client config loading and parsing |
|
||||
| `internal/apps/flared/flared/` | intranet penetration client core scheduling and state management |
|
||||
| `internal/apps/flared/frpc/` | hot-reload/dynamically generate per-Relay `frpc_{relayNodeID}.toml` and monitor frpc |
|
||||
| `internal/apps/flared/heartbeat/` | heartbeat communication with the control plane, incl. Token validation |
|
||||
| `internal/apps/flared/httpclient/` | generic client API communication (`/api/v1/tunnel/*`) |
|
||||
| `internal/apps/flared/sync/` | incrementally pull latest Tunnel route bindings, generate snapshots, apply |
|
||||
| `internal/apps/flared/updater/` | client self-update, new-version check, update landing |
|
||||
| `internal/apps/flared/wsclient/` | WS channel for real-time Server tunnel config change push |
|
||||
|
||||
> **Note**: OpenFlared has no standalone `state/` package; version and checksum are persisted by `frpc/manager.go` to `flared-state.json`.
|
||||
|
||||
---
|
||||
|
||||
## Doc Maintenance Principles
|
||||
|
||||
* Product scope or system boundary changes: update this doc ([Product Boundaries](./index.md)).
|
||||
* Log storage, log-table judgment, or switch-protocol changes: update [Log Store Decoupling](./logstore.md).
|
||||
* System structure or component division changes: update [System Architecture](./architecture.md).
|
||||
* Release, sync, rollback, or Agent model changes: update [Agent & Publish Model](./agent-design.md).
|
||||
* Deployment method changes: update [Deployment Guide](../deployment/deployment.md) and the README.
|
||||
* Config item changes: update [Configuration Reference](../reference/configuration.md).
|
||||
@@ -0,0 +1,109 @@
|
||||
# Uptime Kuma Sync Design
|
||||
|
||||
You will learn: the design background of the OpenFlare × Uptime Kuma monitoring integration, the control-flow design based on the Socket.IO protocol, the anti-pollution model centered on tag isolation, and the differential incremental sync state machine.
|
||||
|
||||
---
|
||||
|
||||
## Requirements Analysis
|
||||
|
||||
In a multi-node gateway architecture, monitoring system state and reverse proxy route state are usually disconnected:
|
||||
1. **High entry overhead**: every time the gateway control plane adds or decommissions a site, the admin must re-configure the corresponding probe address and alert policy in the monitoring system (e.g. Uptime Kuma).
|
||||
2. **Data inconsistency**: when a proxy route domain changes or switches to HTTPS, monitoring parameters are easily left un-updated, causing false positives or missed alerts.
|
||||
3. **Environment pollution risk**: a full "delete-recreate" sync in monitoring would wipe historical statistics and SLA curves, and would also affect other monitor tasks the user configured manually on the instance that are unrelated to the gateway.
|
||||
|
||||
To address these, OpenFlare introduces a **Uptime Kuma auto-monitoring sync mechanism** based on the client/server model, achieving strongly consistent, low-overhead, zero-pollution synchronization between gateway site route definitions and the availability monitoring system.
|
||||
|
||||
---
|
||||
|
||||
## Core Architecture
|
||||
|
||||
The Uptime Kuma sync subsystem runs entirely in the **Server control plane** background scheduler.
|
||||
|
||||
```text
|
||||
[ OpenFlare Control Plane / DB ] [ Uptime Kuma Instance ]
|
||||
│ │
|
||||
1. Scheduled Cron trigger (Job) │
|
||||
│ │
|
||||
2. Read proxy routes & options config │
|
||||
│ │
|
||||
3. Connect to Socket.IO <──── 4. Socket.IO handshake & login ────┤
|
||||
│ │
|
||||
├────── 5. Validate / create "OpenFlare" tag ──►│
|
||||
├────── 6. Compare site attrs vs Kuma monitor list ─►│
|
||||
│ │
|
||||
└────── 7. Execute differential ops (add / edit / delete) ─►│
|
||||
```
|
||||
|
||||
The sync subsystem does not pass through the data-plane Agent nodes; the Server talks directly to Uptime Kuma's exposed Socket.IO endpoint. This reduces edge node network overhead and keeps auth credentials (Kuma username/password) safely inside the control plane.
|
||||
|
||||
---
|
||||
|
||||
## Tag Isolation and Anti-Pollution Design
|
||||
|
||||
To run safely in a shared Uptime Kuma instance without disturbing manually created monitors, a **dedicated tag isolation mechanism** is used:
|
||||
|
||||
1. **`OpenFlare`-specific tag**:
|
||||
* On first connect, the sync routine calls `getTags` to fetch all tags in the instance.
|
||||
* It checks whether a tag named `OpenFlare` exists (default color indigo `#4f46e5`). If not, it creates it automatically via the `addTag` API.
|
||||
2. **Filtered scope**:
|
||||
* After fetching Uptime Kuma's monitor list (`monitorList`), the sync task only keeps monitors **tagged with `OpenFlare`**.
|
||||
* All modification comparisons (`editMonitor`) and offline cleanups (`deleteMonitor`) operate **only within this filtered subset**. Any monitor not bound with the `OpenFlare` tag is "invisible" to the sync routine — perfect anti-pollution isolation.
|
||||
|
||||
---
|
||||
|
||||
## Differential Sync State Machine
|
||||
|
||||
On each run, the sync routine computes a diff between OpenFlare's local config and Uptime Kuma's data, then executes different Socket.IO events based on the comparison:
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> 检查站点状态与监控范围
|
||||
|
||||
state "检查监控范围" as Scope {
|
||||
[*] --> 校验站点是否启用并且在 Scope 内
|
||||
校验站点是否启用并且在 Scope 内 --> 在Scope内 : 是
|
||||
校验站点是否启用并且在 Scope 内 --> 不在Scope内 : 否
|
||||
}
|
||||
|
||||
不在Scope内 --> 检查Kuma中是否存在同名且带标签的监控
|
||||
检查Kuma中是否存在同名且带标签的监控 --> 执行清理 : 存在
|
||||
检查Kuma中是否存在同名且带标签的监控 --> 忽略 : 不存在
|
||||
|
||||
在Scope内 --> 检查Kuma中是否存在同名监控
|
||||
|
||||
state "比对属性" as Compare {
|
||||
[*] --> 检查是否存在
|
||||
检查是否存在 --> 新建监控项 : 否
|
||||
检查是否存在 --> 比对元数据 : 是
|
||||
比对元数据 --> 属性一致 : 匹配
|
||||
比对元数据 --> 属性不一致 : 不匹配
|
||||
}
|
||||
|
||||
新建监控项 --> 发送add指令并绑定Tag
|
||||
属性不一致 --> 发送editMonitor指令
|
||||
属性一致 --> 忽略
|
||||
|
||||
执行清理 --> 发送deleteMonitor指令
|
||||
忽略 --> [*]
|
||||
```
|
||||
|
||||
### 1. Monitor URL Normalization
|
||||
A site route in OpenFlare can configure multiple domains; the sync routine automatically extracts the primary domain and assembles a standard `http://` or `https://` prefix based on whether HTTPS is enabled.
|
||||
|
||||
### 2. Compared Attribute Set
|
||||
If a same-named, tagged monitor already exists, the sync routine compares the following 5 key fields against the current gateway global option. Any mismatch triggers an update:
|
||||
* **URL**: `Url`
|
||||
* **Probe interval**: `Interval` (default 60s)
|
||||
* **Max retries**: `MaxRetries`
|
||||
* **Retry interval**: `RetryInterval` (default 60s)
|
||||
* **Request timeout**: `Timeout` (default 48s)
|
||||
|
||||
---
|
||||
|
||||
## Scheduler and High-Concurrency Protection
|
||||
|
||||
1. **Cron-based single-thread execution**:
|
||||
* The Server periodically (every 1 minute) probes via a background Cron Job whether the configured sync interval (`UptimeKumaSyncInterval`) is reached.
|
||||
* The task uses mutex locking internally. If a previous sync request is still running due to network latency, the next schedule is skipped automatically, preventing concurrent Socket.IO connections from DDOS-ing the Uptime Kuma instance.
|
||||
2. **WebSocket state listening**:
|
||||
* The sync routine uses Socket.IO's event listener; after the connection is established, it only proceeds to the differential algorithm once the full `monitorList` event list push is received, avoiding monitor deletion caused by incomplete data loading.
|
||||
@@ -0,0 +1,127 @@
|
||||
# Login CAPTCHA Integration (Cap)
|
||||
|
||||
This document describes the design of introducing **Cap** — an open-source CAPTCHA solution based on Proof-of-Work (PoW) and invisible browser fingerprint features — into the OpenFlare control plane, to protect the login API against brute-force attacks and credential-stuffing by crawlers.
|
||||
|
||||
---
|
||||
|
||||
## 1. Business Background and Product Scope
|
||||
|
||||
### Background and Pain Points
|
||||
The OpenFlare login endpoint `/api/v1/user/login` lacks user-dimension protection; attackers can use proxy pools to perform credential stuffing and brute-force attacks on high-privilege accounts (such as `root`). At the same time, standard visual CAPTCHAs are unfriendly to login-page UX and accessibility.
|
||||
|
||||
### Product Scope and Technology Choice
|
||||
* **Technology choice**: Cap (a Proof-of-Work-driven, invisible, image-free CAPTCHA solution).
|
||||
- **Core principle**: the client (Widget/page) obtains a proof-of-work (PoW) challenge from the server, computes the solution in the browser background, and sends the answer back. The server verifies the answer to complete human-machine verification.
|
||||
- **Advantages**: invisible, image-free, no dependency on external third-party API nodes (private), tiny package size.
|
||||
* **Integration scope**: the control-plane Server login API (`/api/v1/user/login`) and the frontend login page.
|
||||
* **Config granularity**: admins can toggle the CAPTCHA on/off anytime via the console Option table (`cap_login_enabled`).
|
||||
|
||||
---
|
||||
|
||||
## 2. System Architecture and Interaction Sequence
|
||||
|
||||
### 2.1 Module Responsibilities
|
||||
1. **Frontend**:
|
||||
* Introduces the `cap-widget` (React 19 custom element) on the login page.
|
||||
* On form submit, accompanies the submission with the `cap-token` solved by the Widget.
|
||||
2. **Server (control-plane backend)**:
|
||||
* Exposes `POST /api/cap/challenge` to distribute the PoW challenge and a signed JWT token to the client.
|
||||
* Exposes `POST /api/cap/redeem` to verify the submitted PoW solution and issue a login credential (Redeem Token) with an expiry time.
|
||||
* Stores the Redeem Token and its expiry in the in-memory/Redis cache.
|
||||
* In `POST /api/v1/user/login`, when CAPTCHA protection is enabled, first validates and consumes (single-use) the corresponding `cap-token`.
|
||||
|
||||
### 2.2 Verification Flow Sequence Diagram
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
actor User as User
|
||||
participant Browser as Browser (Frontend Web)
|
||||
participant Server as OpenFlare Server (Backend)
|
||||
participant Cache as Memory/Redis Cache
|
||||
|
||||
User->>Browser: Open login page
|
||||
Browser->>Server: POST /api/cap/challenge (get challenge)
|
||||
Server->>Browser: Return {challenge, token, expires} (JWT format)
|
||||
Note over Browser: Widget computes the PoW challenge in background (WASM/Worker)
|
||||
Browser->>Server: POST /api/cap/redeem (submit solutions + token)
|
||||
alt PoW solution valid
|
||||
Server->>Cache: Store Redeem Token (tokenKey:expires)
|
||||
Server->>Browser: Return {success: true, token} (i.e. cap-token)
|
||||
else validation failed
|
||||
Server->>Browser: Return {success: false, reason}
|
||||
end
|
||||
User->>Browser: Enter account/password, click login
|
||||
Browser->>Server: POST /api/v1/user/login (with X-Cap-Token in HTTP header)
|
||||
alt CapLoginEnabled = true
|
||||
Server->>Server: Middleware (CapAuth) validates and consumes X-Cap-Token
|
||||
alt token valid, not expired, not consumed
|
||||
Server->>Server: c.Next() -> normal login logic (Bcrypt password check)
|
||||
Server->>Browser: Return login success (Session Cookie)
|
||||
else token invalid or already consumed
|
||||
Server->>Browser: Intercept and return CAPTCHA error (401 Unauthorized)
|
||||
end
|
||||
else CapLoginEnabled = false
|
||||
Server->>Server: c.Next() -> normal login logic
|
||||
end
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Core APIs and Data Model
|
||||
|
||||
### 3.1 API Definitions
|
||||
|
||||
#### 1. Get Challenge (POST /api/cap/challenge)
|
||||
* **Method**: `POST`
|
||||
* **Auth**: public
|
||||
* **Response payload** (unified API envelope, `data` is the business payload):
|
||||
```json
|
||||
{
|
||||
"error_msg": "",
|
||||
"data": {
|
||||
"challenge": {
|
||||
"c": 1,
|
||||
"s": 32,
|
||||
"d": 4
|
||||
},
|
||||
"token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...",
|
||||
"expires": 1717660800000
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### 2. Redeem Challenge (POST /api/cap/redeem)
|
||||
* **Method**: `POST`
|
||||
* **Request payload**:
|
||||
```json
|
||||
{
|
||||
"token": "challenge_jwt_token_here",
|
||||
"solutions": [12345, 67890, 54321]
|
||||
}
|
||||
```
|
||||
* **Response payload (success)**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"token": "random_id:ver_token",
|
||||
"expires": 1717661000000
|
||||
}
|
||||
```
|
||||
|
||||
#### 3. Login API (POST /api/v1/user/login)
|
||||
* **Request payload unchanged**:
|
||||
```json
|
||||
{
|
||||
"username": "root",
|
||||
"password": "your_password"
|
||||
}
|
||||
```
|
||||
* **CAPTCHA carrier**: placed in the HTTP Request Header `X-Cap-Token`.
|
||||
|
||||
---
|
||||
|
||||
## 4. Replay Attack Protection and Security Trade-offs
|
||||
1. **JWT temporary state binding**: the challenge is signed into the JWT payload at generation time, including an expiry limit (10 minutes).
|
||||
2. **Replay interception (nonce consumption)**: when the client calls `/redeem` to submit the solution, the backend marks the JWT signature as used in the cache. Re-submitting the same solution package returns `already_redeemed`.
|
||||
3. **Redeem single-use (one-time invalidation)**: when the client logs in and submits the `cap-token`, the backend immediately deletes the key from the cache after validating it, preventing attackers from extracting historical valid `cap-token`s for login replay.
|
||||
4. **Seamless verification**: by tuning parameters like `c` (challenge count) and `d` (difficulty), you balance solve time against anti-crawler strength; users solve silently in the background without interrupting the login flow.
|
||||
@@ -0,0 +1,86 @@
|
||||
# Log Store Decoupling
|
||||
|
||||
You will learn: which tables are log-purpose, why they must not be pinned to ClickHouse, and which code path a new log table must follow.
|
||||
|
||||
Observability fields and the reporting protocol are still governed by [Observability Protocol & Tables](./observability-data-model.md); this document only defines **where data is stored and how to switch databases**.
|
||||
|
||||
---
|
||||
|
||||
## 1. Goals
|
||||
|
||||
* **ClickHouse optional**: when not enabled, PostgreSQL (or SQLite when the primary DB is off) fully takes over writes, queries, aggregation, and cleanup.
|
||||
* **Upper layers don't touch the underlying DB**: apps only face `internal/repository/logstore` (or the `repository` facade). `repository/analytics` and `db.ChConn` / `db.ChDB` are only used by logstore's ClickHouse implementation.
|
||||
* **Switchable**: 「Switch Log Database」in Task Management copies data between PostgreSQL/SQLite and ClickHouse and flips the primary; writes are frozen during migration, the switch only happens on success, and source data is not deleted.
|
||||
|
||||
---
|
||||
|
||||
## 2. What Counts as a Log Table
|
||||
|
||||
A table enters logstore only if it meets all of:
|
||||
|
||||
* Append-only writes, almost no row updates
|
||||
* Query or aggregate by time, deletable by retention days
|
||||
* Must still support writes and queries when ClickHouse is off
|
||||
* Does not participate in transactional consistency for websites / nodes / certificates, etc.
|
||||
|
||||
**Don't** make these log tables: Zones, nodes, config versions, task executions, upload metadata. These go through the business primary DB `repository`.
|
||||
|
||||
Current log domains:
|
||||
|
||||
| Domain | Interface | Tables |
|
||||
| --- | --- | --- |
|
||||
| Node access logs | `AccessLogStore` | `of_node_access_logs` |
|
||||
| Observability time series | `ObservabilityStore` | `of_node_metric_snapshots` / `of_node_edge_health` / `of_node_obs_frps` / `of_node_obs_frpc` |
|
||||
| User access audit | `UserAccessLogStore` | `w_user_access_logs` |
|
||||
|
||||
Hourly materialized views on ClickHouse (e.g. `of_access_log_hourly`) only serve CH query acceleration. PostgreSQL / SQLite **do not** build isomorphic aggregation tables; queries aggregate in real time from raw logs.
|
||||
|
||||
---
|
||||
|
||||
## 3. Layering
|
||||
|
||||
| Layer | Path | Responsibility |
|
||||
| --- | --- | --- |
|
||||
| Abstraction | `internal/repository/logstore` | Interfaces + `Active` / `BuildForMigration`; selects implementation by `log_database` |
|
||||
| CH implementation | `logstore/clickhouse_store.go` | Delegates to `repository/analytics` (native batch + existing aggregation SQL) |
|
||||
| Primary DB implementation | `logstore/postgres_store.go` | PostgreSQL (high-frequency tables partitioned monthly) and SQLite (plain tables) share GORM |
|
||||
| Model | `internal/model/analytics` | Entities and batch SQL, no IO |
|
||||
| Enqueue | `chwriter` / `risk_control` + `batchwriter` | `FlushFunc` calls `logstore.Active`; node logs / observability enqueue via hooks |
|
||||
| Constraint | `logstore/imports_test.go` | apps are forbidden from importing `repository/analytics` |
|
||||
|
||||
`log_database` has only two legal states: **follow the business primary DB** (`postgres` or `sqlite`) or **`clickhouse`**. "Primary PostgreSQL + log SQLite" does not exist. `log_database` / `log_db_migration` are protected and cannot be changed from the admin panel.
|
||||
|
||||
At startup: `log_database=clickhouse` but ClickHouse not enabled → startup is refused; you must re-enable ClickHouse, switch back to the primary DB, and only then turn it off.
|
||||
|
||||
---
|
||||
|
||||
## 4. Switch Protocol
|
||||
|
||||
Task type `of_log_db_switch` (admin name 「Switch Log Database」), parameter `target`.
|
||||
|
||||
1. Validate the target is legal and not the current DB.
|
||||
2. Write `log_db_migration=migrating`, drain in-flight batchwriter (`Drain`, not `Stop` writer). Writes return a clear error afterward (HTTP 503), not queued backlog.
|
||||
3. Clear the target log tables, then copy by id in pages; call `EnsurePartitions` on the PostgreSQL target before copying.
|
||||
4. Only on full success write `log_database=target` and clear the migration marker; on failure clear the marker and writes continue on the source DB.
|
||||
5. Source data is not deleted; re-clear the target before retry to guarantee idempotency.
|
||||
|
||||
Don't invent another switch protocol, and don't connect `analyticsrepo` directly inside tasks.
|
||||
|
||||
---
|
||||
|
||||
## 5. Adding a New Log Table
|
||||
|
||||
Column names must be identical across the three goose migrations (ClickHouse / PostgreSQL / SQLite). Key points:
|
||||
|
||||
* High-frequency tables: CH uses `MergeTree` + `toYYYYMM`; PG uses `PARTITION BY RANGE(time column)` with the partition key in the primary key; SQLite uses a plain table + indexes.
|
||||
* IDs use snowflake `uint64`, preserved as-is during migration.
|
||||
* Writes go through a dedicated `batchwriter`; flush calls `logstore.Active`, not `analyticsrepo.BatchInsert`.
|
||||
* The switch task's `copy*` must cover the new table; cleanup uses existing `log_retention_days_*` or `metric_retention_days`, don't use the wrong TTL.
|
||||
|
||||
Runtime config: [Configuration Reference · Log Storage](../reference/configuration.md#8-日志存储log-database).
|
||||
|
||||
---
|
||||
|
||||
## 6. Related Docs
|
||||
|
||||
* Observability fields and reporting protocol: [Observability Protocol & Tables](./observability-data-model.md)
|
||||
@@ -0,0 +1,769 @@
|
||||
# Agent Reporting Protocol and Observability Data Model
|
||||
|
||||
You will learn: the **data structures** of the refactored Agent heartbeat/WS reports, how the Server **parses and writes** them, and the **target table structures** in ClickHouse / relational DBs.
|
||||
**No protocol compatibility layer**: Agents upgrade via destroy-and-recreate or binary replacement; old fields are not parsed, old buffers are discarded wholesale.
|
||||
|
||||
This design is the **protocol & storage chapter** of [Edge Observability & Business Traffic Stats Refactor](./observability-design.md); implement against the fields and DDL in this document.
|
||||
|
||||
**First read the transport overview and examples:** [Observability Transport Model](./observability-transport-model.md).
|
||||
|
||||
---
|
||||
|
||||
## 1. Design Goals
|
||||
|
||||
| Goal | Description |
|
||||
| --- | --- |
|
||||
| Agent reports only facts | details + host readings + edge health instant state; no business pre-aggregation |
|
||||
| One business detail table | access logs are the only L1 write path |
|
||||
| Aggregation in DB/control plane | hourly summaries come from ClickHouse MV or queries; the Agent never writes summary tables |
|
||||
| No field overlap | `bytes_sent` = data provided; NIC `network_*` = host; no business `openresty_tx` anymore |
|
||||
| Evolvable | new fields optional; missing numeric values default to 0; removed legacy protocol fields are not parsed |
|
||||
|
||||
---
|
||||
|
||||
## 2. Layering and Write Overview
|
||||
|
||||
```text
|
||||
Agent NodePayload (v2)
|
||||
│
|
||||
┌───────────────┼───────────────┐
|
||||
▼ ▼ ▼
|
||||
access_logs host_metrics edge_health
|
||||
(L1 details) (L3 readings) (L2 instant)
|
||||
│ │ │
|
||||
▼ ▼ ▼
|
||||
of_node_access_logs of_node_metric_ of_node_edge_health
|
||||
│ snapshots │
|
||||
│ │ │
|
||||
▼ ▼ │
|
||||
of_access_log_hourly of_node_metric_ │
|
||||
(MV, Server side) capacity_hourly (MV) │
|
||||
│ │ │
|
||||
└─────── admin aggregation API ────┘
|
||||
|
||||
Relational DB (PostgreSQL/SQLite): node latest state, Profile, health events (not a detail lake)
|
||||
```
|
||||
|
||||
| Layer | Meaning | Agent Report Block | ClickHouse Fact Table |
|
||||
| --- | --- | --- | --- |
|
||||
| L1 | business delivery | `access_logs` | `of_node_access_logs` |
|
||||
| L2 | edge health | `edge_health` | `of_node_edge_health` |
|
||||
| L3 | host capacity | `host_metrics` | `of_node_metric_snapshots` |
|
||||
|
||||
---
|
||||
|
||||
## 3. Agent Report Data Structures (protocol v2)
|
||||
|
||||
### 3.1 Top-Level `NodePayload`
|
||||
|
||||
Transport: HTTP heartbeat body and WebSocket `status` messages share the same structure.
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n_xxx",
|
||||
"name": "edge-1",
|
||||
"ip": "1.2.3.4",
|
||||
"version": "3.3.0",
|
||||
"ext_version": "",
|
||||
"current_version": "cfg-checksum-or-version",
|
||||
"last_error": "",
|
||||
"profile": { },
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ],
|
||||
"buffered": [ ],
|
||||
"health_events": [ ],
|
||||
"waf_ip_group_checksums": { "1": "md5..." }
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `schema_version` | int | suggested | fixed to `2` (this design) |
|
||||
| `node_id` | string | ✅ | node ID |
|
||||
| `name` | string | ✅ | display name |
|
||||
| `ip` | string | ✅ | reporting IP |
|
||||
| `version` / `ext_version` | string | ✅ | Agent version |
|
||||
| `current_version` | string | | locally active config version summary |
|
||||
| `last_error` | string | | latest sync/runtime error, nullable |
|
||||
| `openresty_status` | string | ✅ (when OpenResty present) | **latest health-state authoritative field** → written to PG node table |
|
||||
| `openresty_message` | string | | **latest health-description authoritative field** → written to PG node table (**not into CH**) |
|
||||
| `profile` | object | | host overview, report on change (may throttle) |
|
||||
| `host_metrics` | object | suggested each beat | L3 resource snapshot |
|
||||
| `edge_health` | object | suggested each beat | L2 connection time series + status aligned with top level |
|
||||
| `access_logs` | array | | this beat's incremental access details |
|
||||
| `buffered` | array | | offline backfill fact batches (see §3.6) |
|
||||
| `health_events` | array | | edge health events |
|
||||
| `waf_ip_group_checksums` | map | | for differential sync, not an observability lake |
|
||||
|
||||
**Removed, Server no longer parses (no compatibility layer):**
|
||||
|
||||
| Old Field | Disposition |
|
||||
| --- | --- |
|
||||
| `traffic_report` | not in the protocol; not stored |
|
||||
| `openresty_observation` | not present; connections/status go through `edge_health` |
|
||||
| `snapshot` | not present; only `host_metrics` |
|
||||
| `buffered_observability` | not present; only `buffered` |
|
||||
|
||||
### 3.2 `profile` — Host Overview (low frequency)
|
||||
|
||||
Maps to relational `of_node_system_profiles` (or an existing equivalent), **not into the ClickHouse detail lake**.
|
||||
|
||||
```json
|
||||
{
|
||||
"hostname": "edge-1",
|
||||
"os_name": "linux",
|
||||
"os_version": "...",
|
||||
"kernel_version": "...",
|
||||
"architecture": "amd64",
|
||||
"cpu_model": "...",
|
||||
"cpu_cores": 8,
|
||||
"total_memory_bytes": 16106127360,
|
||||
"total_disk_bytes": 107374182400,
|
||||
"uptime_seconds": 864000,
|
||||
"reported_at_unix": 1720000000
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Semantics |
|
||||
| --- | --- |
|
||||
| hardware/OS description fields | factual readings |
|
||||
| `reported_at_unix` | Agent collection time (UTC seconds) |
|
||||
|
||||
### 3.3 `host_metrics` — Host Capacity (L3)
|
||||
|
||||
**All readings, no 24h business totals.**
|
||||
NIC/disk bytes are **kernel cumulative counter raw values** (monotonically increasing, may reset on restart); CPU is an instant percentage; memory/disk usage is current usage.
|
||||
|
||||
```json
|
||||
{
|
||||
"captured_at_unix": 1720000000,
|
||||
"cpu_usage_percent": 12.5,
|
||||
"memory_used_bytes": 4294967296,
|
||||
"memory_total_bytes": 16106127360,
|
||||
"storage_used_bytes": 50000000000,
|
||||
"storage_total_bytes": 107374182400,
|
||||
"disk_read_bytes": 9000000000,
|
||||
"disk_write_bytes": 12000000000,
|
||||
"network_rx_bytes": 500000000000,
|
||||
"network_tx_bytes": 800000000000
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Semantics | How Server Uses It |
|
||||
| --- | --- | --- | --- |
|
||||
| `captured_at_unix` | int64 | sampling time | `captured_at` |
|
||||
| `cpu_usage_percent` | float | instant CPU% | store directly; average for trends |
|
||||
| `memory_*` / `storage_*` | int64 | current used/total | store directly; compute usage rate |
|
||||
| `disk_read_bytes` / `disk_write_bytes` | int64 | **cumulative** IO bytes | store raw; adjacent deltas at query time |
|
||||
| `network_rx_bytes` / `network_tx_bytes` | int64 | **cumulative** NIC bytes | store raw; adjacent deltas at query time → "host NIC in/outbound" |
|
||||
|
||||
> The Agent is **forbidden** from replacing cumulative values with "this period's delta" before reporting (otherwise Server deltas would be wrong).
|
||||
|
||||
### 3.4 `edge_health` — OpenResty Edge Health (L2)
|
||||
|
||||
**Instant state only, no business throughput.**
|
||||
|
||||
```json
|
||||
{
|
||||
"captured_at_unix": 1720000000,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 42
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Semantics |
|
||||
| --- | --- | --- |
|
||||
| `status` | string | `healthy` / `unhealthy` / `unknown` (must match top-level `openresty_status`) |
|
||||
| `message` | string | status description (may be reported; **only backfills PG latest state, not into CH**) |
|
||||
| `connections` | int64 | stub_status Active connections |
|
||||
|
||||
#### Health-State Authoritative Sources (converged)
|
||||
|
||||
| Data | Authoritative Storage | Description |
|
||||
| --- | --- | --- |
|
||||
| **Current** OpenResty health + description | **PG node table** `openresty_status` / `openresty_message` | UI badges, lists, alerts use this |
|
||||
| **Time series** health status + connections | **CH** `of_node_edge_health` (`status`, `connections`) | connection curves / health history; **no message column** |
|
||||
| Agent report | top-level status/message + `edge_health` | Server normalizes both statuses aligned; message **only written to PG** |
|
||||
|
||||
So: "is it unhealthy now" → read PG; "connections over the past 24h" → read CH.
|
||||
|
||||
### 3.5 `access_logs[]` — Access Details (L1, single business truth)
|
||||
|
||||
Agent: tail access.log → parse JSON lines → report fields as-is (path may be truncated).
|
||||
|
||||
```json
|
||||
{
|
||||
"logged_at_unix": 1720000001,
|
||||
"remote_addr": "203.0.113.10",
|
||||
"host": "www.example.com",
|
||||
"path": "/api/v1/ping",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 1024,
|
||||
"request_length": 128,
|
||||
"request_time_ms": 15,
|
||||
"user_agent": "Mozilla/5.0 ...",
|
||||
"cache_status": "HIT"
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Required | Source (OpenResty) | Business Meaning |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| `logged_at_unix` | int64 | ✅ | parse `$time_iso8601` | request completion time |
|
||||
| `remote_addr` | string | ✅ | `$remote_addr` | client IP → UV |
|
||||
| `host` | string | ✅ | `$host` | domain → Zone ownership |
|
||||
| `path` | string | ✅ | `$request_uri`, Agent may truncate | path |
|
||||
| `status_code` | int | ✅ | `$status` | status code |
|
||||
| `bytes_sent` | int64 | ✅ | **`$body_bytes_sent`** | **data provided** (response body) |
|
||||
| `request_length` | int64 | suggested | `$request_length` | **data received** |
|
||||
| `request_time_ms` | int64 | optional | `$request_time * 1000` | latency; default 0 |
|
||||
| `user_agent` | string | suggested | `$http_user_agent` | UA; may truncate on store |
|
||||
| `cache_status` | string | suggested | **`$upstream_cache_status`** | edge cache result (see §3.5.1) |
|
||||
|
||||
**Explicitly not reported by Agent (written by Server):**
|
||||
|
||||
* `region` / country: GeoIP resolved at insert time
|
||||
* `id` / `created_at`: Server-generated
|
||||
* `node_id`: from payload / auth context
|
||||
|
||||
**Explicitly not reported:**
|
||||
|
||||
* `upstream_addr` / origin address / `origin_fetched`: no origin-endpoint tracking; "did it fetch from origin" is only derived from `cache_status` at the control plane (§3.5.1)
|
||||
|
||||
### 3.5.1 `cache_status` — Cache Hit and Origin Fetch (detail-first)
|
||||
|
||||
**Goal (phase 1):** access log details/list can show "cache hit / origin fetch / no cache used".
|
||||
**Caliber:** only store the raw OpenResty `$upstream_cache_status`; **no upstream address reported**.
|
||||
|
||||
#### Raw Values (stored)
|
||||
|
||||
| Value | Meaning (OpenResty) |
|
||||
| --- | --- |
|
||||
| `HIT` | cache hit |
|
||||
| `MISS` | miss, fetched from upstream |
|
||||
| `BYPASS` | cache skipped (e.g. method/cookie/policy caused `$openflare_skip_cache`) |
|
||||
| `EXPIRED` | expired then origin fetch |
|
||||
| `STALE` | served stale |
|
||||
| `UPDATING` | background updating, may return old cache |
|
||||
| `REVALIDATED` | revalidated, still used cache |
|
||||
| `-` or empty | didn't pass through `proxy_cache` (e.g. Pages local static, non-proxy location) |
|
||||
|
||||
#### UI Three-State Derivation (not stored)
|
||||
|
||||
Control-plane display uses derived enum `cache_outcome`, **not written to CH**:
|
||||
|
||||
| Three-State | Condition (`cache_status`) | Suggested List Label |
|
||||
| --- | --- | --- |
|
||||
| **Cache hit** | `HIT` / `STALE` / `REVALIDATED` / `UPDATING` | hit |
|
||||
| **Origin fetch** | `MISS` / `EXPIRED` | origin |
|
||||
| **No cache used** | `BYPASS` / `-` / `""` | not cached |
|
||||
|
||||
Details can show both the three-state and the raw `cache_status`.
|
||||
|
||||
#### Boundaries
|
||||
|
||||
* Pages static / locations without `proxy_cache`: mostly empty or `-` → **no cache used**, must not be labeled "hit".
|
||||
* Detail pages show cache state; hit-rate dashboards and hourly dimensions can extend from the same column.
|
||||
|
||||
**Per-heartbeat count suggestion:**
|
||||
|
||||
* Soft cap e.g. 2000 lines/beat; overflow goes into `buffered` next batch, **forbidden** to compress into a TrafficReport in the Agent.
|
||||
|
||||
### 3.6 `buffered[]` — Offline Backfill (facts only)
|
||||
|
||||
```json
|
||||
{
|
||||
"captured_at_unix": 1719999900,
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ]
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `captured_at_unix` | batch collection/buffer time, used for ack and dedup window |
|
||||
| `host_metrics` / `edge_health` / `access_logs` | same structures as the main payload; empty blocks may be omitted |
|
||||
|
||||
**Forbidden** to carry `traffic_report` or rx/tx throughput in buffered.
|
||||
|
||||
### 3.7 `health_events[]`
|
||||
|
||||
```json
|
||||
{
|
||||
"event_type": "openresty_unhealthy",
|
||||
"severity": "critical",
|
||||
"message": "...",
|
||||
"triggered_at_unix": 1720000000,
|
||||
"metadata": { }
|
||||
}
|
||||
```
|
||||
|
||||
Written to the relational health-event table (existing model suffices), not into the access log lake.
|
||||
|
||||
### 3.8 Go Protocol Structures
|
||||
|
||||
```go
|
||||
// pkg/protocol/agent.go (current implementation)
|
||||
|
||||
type NodePayload struct {
|
||||
SchemaVersion int `json:"schema_version,omitempty"`
|
||||
NodeID string `json:"node_id"`
|
||||
Name string `json:"name"`
|
||||
IP string `json:"ip"`
|
||||
Version string `json:"version"`
|
||||
ExtVersion string `json:"ext_version"`
|
||||
CurrentVersion string `json:"current_version"`
|
||||
LastError string `json:"last_error"`
|
||||
OpenrestyStatus string `json:"openresty_status"` // PG latest-state authority
|
||||
OpenrestyMessage string `json:"openresty_message"` // PG latest-state authority; not into CH
|
||||
Profile *NodeSystemProfile `json:"profile,omitempty"`
|
||||
HostMetrics *NodeHostMetrics `json:"host_metrics,omitempty"`
|
||||
EdgeHealth *NodeEdgeHealth `json:"edge_health,omitempty"`
|
||||
AccessLogs []NodeAccessLog `json:"access_logs,omitempty"`
|
||||
Buffered []BufferedFacts `json:"buffered,omitempty"`
|
||||
HealthEvents []NodeHealthEvent `json:"health_events"`
|
||||
WAFIPGroupChecksums map[string]string `json:"waf_ip_group_checksums,omitempty"`
|
||||
}
|
||||
|
||||
type NodeHostMetrics struct {
|
||||
CapturedAtUnix int64 `json:"captured_at_unix"`
|
||||
CPUUsagePercent float64 `json:"cpu_usage_percent"`
|
||||
MemoryUsedBytes int64 `json:"memory_used_bytes"`
|
||||
MemoryTotalBytes int64 `json:"memory_total_bytes"`
|
||||
StorageUsedBytes int64 `json:"storage_used_bytes"`
|
||||
StorageTotalBytes int64 `json:"storage_total_bytes"`
|
||||
DiskReadBytes int64 `json:"disk_read_bytes"`
|
||||
DiskWriteBytes int64 `json:"disk_write_bytes"`
|
||||
NetworkRxBytes int64 `json:"network_rx_bytes"`
|
||||
NetworkTxBytes int64 `json:"network_tx_bytes"`
|
||||
}
|
||||
|
||||
type NodeEdgeHealth struct {
|
||||
CapturedAtUnix int64 `json:"captured_at_unix"`
|
||||
Status string `json:"status"`
|
||||
Message string `json:"message"`
|
||||
Connections int64 `json:"connections"`
|
||||
}
|
||||
|
||||
type NodeAccessLog struct {
|
||||
LoggedAtUnix int64 `json:"logged_at_unix"`
|
||||
RemoteAddr string `json:"remote_addr"`
|
||||
Host string `json:"host"`
|
||||
Path string `json:"path"`
|
||||
UserAgent string `json:"user_agent,omitempty"`
|
||||
CacheStatus string `json:"cache_status,omitempty"` // $upstream_cache_status
|
||||
StatusCode int `json:"status_code"`
|
||||
BytesSent int64 `json:"bytes_sent"` // body_bytes_sent, data provided
|
||||
RequestLength int64 `json:"request_length"` // data received
|
||||
RequestTimeMs int64 `json:"request_time_ms"` // optional
|
||||
}
|
||||
|
||||
type BufferedFacts struct {
|
||||
CapturedAtUnix int64 `json:"captured_at_unix"`
|
||||
HostMetrics *NodeHostMetrics `json:"host_metrics,omitempty"`
|
||||
EdgeHealth *NodeEdgeHealth `json:"edge_health,omitempty"`
|
||||
AccessLogs []NodeAccessLog `json:"access_logs,omitempty"`
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Server Parsing and Storage Flow
|
||||
|
||||
### 4.1 Entry Points
|
||||
|
||||
* HTTP: `POST /api/v1/agent/...` heartbeat (existing path)
|
||||
* WebSocket: `type=status` payload = `NodePayload`
|
||||
* Auth: `X-Agent-Token` → binds `node_id` (payload.node_id must match the token node)
|
||||
|
||||
### 4.2 Processing Pipeline (single payload)
|
||||
|
||||
```text
|
||||
1. Deserialize NodePayload
|
||||
2. Normalize
|
||||
- schema_version < 2:
|
||||
host_metrics ← snapshot
|
||||
edge_health.status ← openresty_status
|
||||
edge_health.connections ← openresty_observation.connections (if present)
|
||||
traffic_report → drop
|
||||
openresty_observation.rx/tx → drop
|
||||
buffered ← buffered_observability
|
||||
- path truncated again, status clamped, negative bytes → 0
|
||||
3. Relational transaction (node latest state)
|
||||
- update node online time, IP, version, edge_health.status/message
|
||||
- upsert profile (if present)
|
||||
- insert health_events (if present)
|
||||
4. ClickHouse async batch (log on failure; doesn't block heartbeat response config delivery)
|
||||
a. access_logs + buffered[].access_logs
|
||||
→ fill region (GeoIP)
|
||||
→ assign snowflake id
|
||||
→ BatchInsert of_node_access_logs
|
||||
b. host_metrics + buffered[].host_metrics
|
||||
→ of_node_metric_snapshots
|
||||
c. edge_health + buffered[].edge_health
|
||||
→ of_node_edge_health (connections + status snapshot only, optional)
|
||||
5. Return heartbeat response (settings / active_config / waf diff)
|
||||
6. If using buffer ack: confirm by the list of buffered.captured_at_unix
|
||||
```
|
||||
|
||||
### 4.3 Normalization Rules (hard constraints)
|
||||
|
||||
| Rule | Behavior |
|
||||
| --- | --- |
|
||||
| `logged_at` ahead of now+5m | clamp to now or drop the line (pick one in implementation and unit-test it) |
|
||||
| `logged_at` older than now−TTL | still writable; relies on table TTL cleanup |
|
||||
| empty `host` | allowed, aggregated into "unassigned" |
|
||||
| `bytes_sent` / `request_length` < 0 | set to 0 |
|
||||
| single batch access_logs > N | truncate and alert-metric it (or only queue into buffer), never switch to pre-aggregation |
|
||||
| duplicate backfill | CH tolerates a few duplicate rows; queries approximate with sum (no forced exact dedup) |
|
||||
|
||||
### 4.4 Field Mapping Table (report → table)
|
||||
|
||||
| Report Path | Target Storage | Columns |
|
||||
| --- | --- | --- |
|
||||
| `access_logs[]` | CH `of_node_access_logs` | see §5.1 |
|
||||
| `host_metrics` | CH `of_node_metric_snapshots` | see §5.2 |
|
||||
| `edge_health` | CH `of_node_edge_health` + PG node latest state | see §5.3 / §5.6 |
|
||||
| `profile` | PG `of_node_system_profiles` | existing columns |
|
||||
| `health_events` | PG health-event table | existing model |
|
||||
| `waf_ip_group_checksums` | not into observability tables | sync logic |
|
||||
| `traffic_report` (legacy) | **not written** | — |
|
||||
| `openresty_rx/tx` (legacy) | **not written** | — |
|
||||
|
||||
### 4.5 Query Side (no new "business outbound" column)
|
||||
|
||||
| Product Metric | SQL Semantics (sketch) |
|
||||
| --- | --- |
|
||||
| Data provided | `sum(bytes_sent)` |
|
||||
| Data received | `sum(request_length)` |
|
||||
| Request count | `count()` |
|
||||
| UV | `uniqExact(remote_addr)` |
|
||||
| 5xx | `countIf(status_code >= 500)` |
|
||||
| By domain/status/region | `GROUP BY host / status_code / region` |
|
||||
| Host NIC outbound | non-negative delta over `network_tx_bytes` per node in time order, then sum |
|
||||
| OpenResty connections | `of_node_edge_health.connections` latest or average |
|
||||
|
||||
---
|
||||
|
||||
## 5. Table Structures (DDL)
|
||||
|
||||
> Engine and TTL tend to match production: access logs 90 days, metrics 30 days.
|
||||
> `id` uses control-plane Snowflake/unique UInt64.
|
||||
|
||||
### 5.1 L1 Fact Table: `of_node_access_logs`
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS of_node_access_logs
|
||||
(
|
||||
id UInt64,
|
||||
node_id String,
|
||||
logged_at DateTime64(3, 'UTC'),
|
||||
remote_addr String,
|
||||
region String, -- Server GeoIP writes, Agent doesn't send
|
||||
host String,
|
||||
path String,
|
||||
user_agent String DEFAULT '', -- $http_user_agent
|
||||
cache_status String DEFAULT '', -- $upstream_cache_status
|
||||
status_code Int32,
|
||||
bytes_sent UInt64, -- data provided (body)
|
||||
request_length UInt64 DEFAULT 0, -- data received
|
||||
request_time_ms UInt32 DEFAULT 0, -- optional
|
||||
created_at DateTime64(3, 'UTC')
|
||||
)
|
||||
ENGINE = MergeTree()
|
||||
PARTITION BY toYYYYMM(logged_at)
|
||||
ORDER BY (node_id, logged_at, host, status_code, remote_addr)
|
||||
TTL toDateTime(logged_at) + INTERVAL 90 DAY
|
||||
SETTINGS index_granularity = 8192;
|
||||
```
|
||||
|
||||
| Column | Type | Source |
|
||||
| --- | --- | --- |
|
||||
| `id` | UInt64 | Server |
|
||||
| `node_id` | String | auth/payload |
|
||||
| `logged_at` | DateTime64(3) | `logged_at_unix` |
|
||||
| `remote_addr` | String | report |
|
||||
| `region` | String | Server GeoIP |
|
||||
| `host` | String | report |
|
||||
| `path` | String | report |
|
||||
| `user_agent` | String | report (nullable) |
|
||||
| `cache_status` | String | report (nullable) → **cache status** |
|
||||
| `status_code` | Int32 | report |
|
||||
| `bytes_sent` | UInt64 | report → **data provided** |
|
||||
| `request_length` | UInt64 | report → **data received** |
|
||||
| `request_time_ms` | UInt32 | report optional |
|
||||
| `created_at` | DateTime64(3) | Server now |
|
||||
|
||||
**Migration:** the current table already has `bytes_sent` / `request_length` / `request_time_ms` / `user_agent`; cache status adds:
|
||||
|
||||
```sql
|
||||
ALTER TABLE of_node_access_logs
|
||||
ADD COLUMN IF NOT EXISTS cache_status String DEFAULT '';
|
||||
```
|
||||
|
||||
### 5.2 L1 Hourly Rollup (Server-side MV)
|
||||
|
||||
**Agent forbidden to write.** Serves dashboard/node 24h fast queries of request count, error count, bytes.
|
||||
|
||||
**Implemented choice: `SummingMergeTree` + no UV column.**
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS of_access_log_hourly
|
||||
(
|
||||
node_id String,
|
||||
hour DateTime('UTC'),
|
||||
host String,
|
||||
request_count UInt64,
|
||||
error_count UInt64,
|
||||
bytes_sent UInt64,
|
||||
request_length UInt64
|
||||
)
|
||||
ENGINE = SummingMergeTree()
|
||||
PARTITION BY toYYYYMM(hour)
|
||||
ORDER BY (node_id, hour, host)
|
||||
TTL hour + INTERVAL 90 DAY;
|
||||
|
||||
CREATE MATERIALIZED VIEW IF NOT EXISTS of_access_log_hourly_mv
|
||||
TO of_access_log_hourly
|
||||
AS
|
||||
SELECT
|
||||
node_id,
|
||||
toStartOfHour(logged_at) AS hour,
|
||||
host,
|
||||
toUInt64(count()) AS request_count,
|
||||
toUInt64(countIf(status_code >= 500)) AS error_count,
|
||||
sum(bytes_sent) AS bytes_sent,
|
||||
sum(request_length) AS request_length
|
||||
FROM of_node_access_logs
|
||||
GROUP BY node_id, hour, host;
|
||||
```
|
||||
|
||||
Historical hours (details stored before the MV existed) need a one-time backfill, see migration `202607180003_backfill_access_log_hourly.sql` (ANTI JOIN to prevent duplicates).
|
||||
|
||||
#### UV Policy (must follow)
|
||||
|
||||
| Scenario | Data Source | Algorithm | Notes |
|
||||
| --- | --- | --- | --- |
|
||||
| **Window total UV** (dashboard totals, node cards, Zone totals) | `of_node_access_logs` details | `uniqExact(remote_addr)` (`TrafficSummary` / node aggregation) | **single authority**; never sum hourly UV |
|
||||
| **24h trend line request/error/bytes** | `of_access_log_hourly` preferred, fall back to detail buckets | `sum(request_count)` etc. | hourly path **doesn't fill** `unique_visitor_count` (always 0) |
|
||||
| **24h trend per-hour UV** | detail bucket path only | in-bucket `uniqExact` | when using hourly, UI should show empty/0 or hide the UV series; **forbidden** to `sum(UV)` over hourly rows |
|
||||
|
||||
**Why hourly doesn't store UV:**
|
||||
|
||||
1. `SummingMergeTree` can only safely merge addable counts; `uniqExact` across parts needs `AggregatingMergeTree` + state, heavier to implement and query.
|
||||
2. Even if hourly UV were stored, **summing over multi-hour windows severely overestimates** (the same IP is counted once per hour).
|
||||
3. Product "24h unique visitors" only recognizes whole-window `uniqExact`; the trend chart's main series are requests/errors/bytes — per-hour UV is not a primary metric.
|
||||
|
||||
### 5.3 L3 Fact Table: `of_node_metric_snapshots` (kept, semantics clarified)
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS of_node_metric_snapshots
|
||||
(
|
||||
id UInt64,
|
||||
node_id String,
|
||||
captured_at DateTime64(3, 'UTC'),
|
||||
cpu_usage_percent Float64,
|
||||
memory_used_bytes Int64,
|
||||
memory_total_bytes Int64,
|
||||
storage_used_bytes Int64,
|
||||
storage_total_bytes Int64,
|
||||
disk_read_bytes Int64, -- cumulative raw
|
||||
disk_write_bytes Int64,
|
||||
network_rx_bytes Int64, -- cumulative raw → host NIC inbound
|
||||
network_tx_bytes Int64, -- cumulative raw → host NIC outbound
|
||||
created_at DateTime64(3, 'UTC')
|
||||
)
|
||||
ENGINE = MergeTree()
|
||||
PARTITION BY toYYYYMM(captured_at)
|
||||
ORDER BY (node_id, captured_at, id)
|
||||
TTL toDateTime(captured_at) + INTERVAL 30 DAY
|
||||
SETTINGS index_granularity = 8192;
|
||||
```
|
||||
|
||||
Columns match production; **docs and API must label `network_*` as host NIC cumulative values**.
|
||||
|
||||
### 5.4 L3 Hourly Rollup: `of_node_metric_capacity_hourly` (kept)
|
||||
|
||||
Existing min/max used for cumulative-counter hourly increment approximation + CPU/memory averages. Logic unchanged:
|
||||
|
||||
* `network_tx_max - network_tx_min` ≈ that hour's host outbound
|
||||
* **must not** be used for "data provided"
|
||||
|
||||
### 5.5 L2 Fact Table: `of_node_edge_health` (new, replaces throughput-style openresty table)
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS of_node_edge_health
|
||||
(
|
||||
id UInt64,
|
||||
node_id String,
|
||||
captured_at DateTime64(3, 'UTC'),
|
||||
status LowCardinality(String), -- healthy / unhealthy / unknown
|
||||
connections Int64,
|
||||
created_at DateTime64(3, 'UTC')
|
||||
)
|
||||
ENGINE = MergeTree()
|
||||
PARTITION BY toYYYYMM(captured_at)
|
||||
ORDER BY (node_id, captured_at, id)
|
||||
TTL toDateTime(captured_at) + INTERVAL 30 DAY
|
||||
SETTINGS index_granularity = 8192;
|
||||
```
|
||||
|
||||
| Column | Description |
|
||||
| --- | --- |
|
||||
| `status` | instant health (same source as PG current state; for time series, not the sole UI authority) |
|
||||
| `connections` | current connection count |
|
||||
|
||||
**No** `message` column (description only in PG latest state).
|
||||
**No** `openresty_rx_bytes` / `openresty_tx_bytes`.
|
||||
|
||||
### 5.6 Relational DB (node latest state, not an analytics lake)
|
||||
|
||||
Separate from the observability lake, keeping "latest one":
|
||||
|
||||
| Table (logical name) | Purpose | Key Columns |
|
||||
| --- | --- | --- |
|
||||
| `of_nodes` (or current node table) | online, version, IP | `last_seen_at`, `openresty_status`, `openresty_message`, `agent_version` |
|
||||
| `of_node_system_profiles` | profile upsert | hostname, cpu_cores, total_memory_bytes, ... |
|
||||
| health-event table | `health_events` | event_type, severity, message, triggered_at |
|
||||
|
||||
> Actual physical table names follow the repo's existing GORM models; this design doesn't force renames, only forces **business throughput no longer written into node tables**.
|
||||
|
||||
### 5.7 Deprecated Tables (stop writing → delete after TTL)
|
||||
|
||||
| Table | Reason | Replacement |
|
||||
| --- | --- | --- |
|
||||
| `of_node_request_reports` | Agent pre-aggregation | `of_node_access_logs` + hourly |
|
||||
| `of_node_traffic_hourly` + MV | depends on request_reports | `of_access_log_hourly` |
|
||||
| `of_node_obs_openresty` | contains business rx/tx | `of_node_edge_health` |
|
||||
| `of_node_openresty_hourly` + MV | business throughput deltas | `of_access_log_hourly` bytes_* |
|
||||
|
||||
Relay-specific `of_node_obs_frps` / `of_node_obs_frpc` **kept** (not this Agent's main path, but same CH observability).
|
||||
|
||||
---
|
||||
|
||||
## 6. Table-Protocol Cross-Reference
|
||||
|
||||
| Product Concept | Protocol Field | Table.Column | Aggregation |
|
||||
| --- | --- | --- | --- |
|
||||
| Data provided | `access_logs[].bytes_sent` | `of_node_access_logs.bytes_sent` | `sum` |
|
||||
| Data received | `access_logs[].request_length` | `...request_length` | `sum` |
|
||||
| Request count | row count | — | `count` |
|
||||
| UV (window total) | `remote_addr` | same details | `uniqExact` (**forbidden** to sum hourly UV) |
|
||||
| Top domains | `host` | same | `group by` |
|
||||
| Status distribution | `status_code` | same | `group by` |
|
||||
| Source region | — | `region` (Server) | `group by` |
|
||||
| Host NIC outbound | `host_metrics.network_tx_bytes` | `of_node_metric_snapshots.network_tx_bytes` | time-series delta |
|
||||
| Host NIC inbound | `network_rx_bytes` | same | delta |
|
||||
| Disk read/write | `disk_*_bytes` | same | delta |
|
||||
| CPU/memory | instant fields | same | avg |
|
||||
| OpenResty connections | `edge_health.connections` | `of_node_edge_health.connections` | latest/avg |
|
||||
| OpenResty health | `edge_health.status` | node table + optional CH | latest |
|
||||
|
||||
**Mappings that no longer exist:**
|
||||
|
||||
| Old Concept | Old Field | Disposition |
|
||||
| --- | --- | --- |
|
||||
| OpenResty outbound | `openresty_tx_bytes` | removed; use data provided |
|
||||
| OpenResty inbound | `openresty_rx_bytes` | removed; use data received |
|
||||
| Window request report | `traffic_report` | removed |
|
||||
|
||||
---
|
||||
|
||||
## 7. OpenResty Log Format (aligned with details)
|
||||
|
||||
Target `log_format` (ensures the `bytes_sent` key = body; includes UA and cache status):
|
||||
|
||||
```nginx
|
||||
log_format openflare_json escape=json
|
||||
'{"ts":"$time_iso8601","host":"$host","path":"$request_uri",'
|
||||
'"remote_addr":"$remote_addr","status":$status,'
|
||||
'"request_time":$request_time,'
|
||||
'"bytes_sent":$body_bytes_sent,"request_length":$request_length,'
|
||||
'"user_agent":"$http_user_agent",'
|
||||
'"cache_status":"$upstream_cache_status"}';
|
||||
```
|
||||
|
||||
Agent parsing:
|
||||
|
||||
* `ts` → `logged_at_unix`
|
||||
* `bytes_sent` → protocol `bytes_sent` (provided)
|
||||
* `request_length` → protocol `request_length`
|
||||
* `request_time` → optional `request_time_ms = round(sec * 1000)`
|
||||
* `user_agent` → protocol `user_agent`
|
||||
* `cache_status` → protocol `cache_status` (passed through as-is, no three-state compression)
|
||||
|
||||
---
|
||||
|
||||
## 8. Upgrade Strategy (no compatibility layer)
|
||||
|
||||
| Item | Strategy |
|
||||
| --- | --- |
|
||||
| Agent upgrade | **destroy-and-recreate** preferred; **binary replacement** allowed |
|
||||
| Protocol | only schema v2 fields; legacy JSON fields not parsed |
|
||||
| Local observability buffer | if still in old format (containing `snapshot` / `openresty_observation` / `traffic_report`) or corrupt → **delete the file wholesale**, rebuild at runtime |
|
||||
| Read path | business APIs **only read** access_logs (and hourly); current health reads PG; connection series reads CH edge_health |
|
||||
| Old Agents | must upgrade; the control plane provides no v1 dual-read path |
|
||||
|
||||
---
|
||||
|
||||
## 9. Example: Storage Result of One Heartbeat
|
||||
|
||||
**Agent report (excerpt):**
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n1",
|
||||
"host_metrics": {
|
||||
"captured_at_unix": 1720000000,
|
||||
"cpu_usage_percent": 10,
|
||||
"memory_used_bytes": 1,
|
||||
"memory_total_bytes": 2,
|
||||
"storage_used_bytes": 3,
|
||||
"storage_total_bytes": 4,
|
||||
"disk_read_bytes": 100,
|
||||
"disk_write_bytes": 200,
|
||||
"network_rx_bytes": 1000,
|
||||
"network_tx_bytes": 2000
|
||||
},
|
||||
"edge_health": {
|
||||
"captured_at_unix": 1720000000,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 5
|
||||
},
|
||||
"access_logs": [
|
||||
{
|
||||
"logged_at_unix": 1720000001,
|
||||
"remote_addr": "1.1.1.1",
|
||||
"host": "a.example.com",
|
||||
"path": "/",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 500,
|
||||
"request_length": 80
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Written:**
|
||||
|
||||
1. PG node latest state: `openresty_status` / `openresty_message` (if reported)
|
||||
2. `of_node_metric_snapshots` 1 row (network_tx=2000 cumulative)
|
||||
3. `of_node_edge_health` 1 row (status + connections=5; **no message**)
|
||||
4. `of_node_access_logs` 1 row (bytes_sent=500, request_length=80, region filled by Server)
|
||||
5. MV asynchronously counts into `of_access_log_hourly`
|
||||
|
||||
**Query 24h data provided:** `sum(bytes_sent)` → at least 500 (plus history)
|
||||
**Query host outbound:** delta over snapshots; **no forced equality** with 500.
|
||||
|
||||
---
|
||||
|
||||
## 10. Revision History
|
||||
|
||||
| Date | Notes |
|
||||
| --- | --- |
|
||||
| 2026-07-17 | initial draft: protocol v2, Server storage pipeline, CH/relational target table structures and deprecated table list |
|
||||
@@ -0,0 +1,552 @@
|
||||
# Edge Observability and Business Traffic Statistics Refactor
|
||||
|
||||
You will learn: the problems this refactor solves (dashboard "OpenResty outbound" vs "Zone data provided" inconsistency, field and aggregation redundancy), and how the target architecture makes **the Agent report only facts and the Server interpret facts**, with access logs as the single source of truth for business traffic.
|
||||
|
||||
---
|
||||
|
||||
## 1. Goals
|
||||
|
||||
### 1.1 Problems to Solve
|
||||
|
||||
1. **Dual sources of truth**: business throughput comes from both access-log aggregation and OpenResty observability deltas, and the numbers never match long-term.
|
||||
2. **Agent over-computes**: the edge pre-aggregates `TrafficReport`, throughput accumulation, and the control plane aggregates again — semantics are hard to evolve and reconcile.
|
||||
3. **Field semantic overlap**: "OpenResty outbound" and "data provided" are the same business problem for users, but the system uses two field sets and two pipelines.
|
||||
4. **Instant vs cumulative mixed**: 60-second window counts are treated as process cumulative values for 24h deltas, causing severe underestimation.
|
||||
5. **UI induces wrong comparisons**: the dashboard and Zone page use similar "traffic/data" wording without declaring scope and caliber differences.
|
||||
|
||||
### 1.2 Refactor Goals
|
||||
|
||||
| Goal | Description |
|
||||
| --- | --- |
|
||||
| **Single business truth** | request count, data provided, UV, status distribution, Top domains etc. **only** derived from access logs (and Server-side rollups) |
|
||||
| **Agent reports only facts** | detail logs + machine readings + health snapshots; **business UV/TopN/24h totals pre-aggregation is forbidden** |
|
||||
| **Field convergence** | one business concept maps to one authoritative field; machine NIC and business delivery strictly separated by name |
|
||||
| **Reconcilable** | global "data provided" ≈ sum of per-Zone "data provided" (difference only from unbound/unknown Hosts) |
|
||||
| **Evolvable** | changing time windows, TopN, ownership rules only changes the Server, not the Agent |
|
||||
|
||||
### 1.3 Non-Goals (outside this design)
|
||||
|
||||
* Building a general log platform, full-log long-term archive, or search product.
|
||||
* Replacing ClickHouse / removing the analytics DB dependency.
|
||||
* Reworking Relay / OpenFlared host metric collection (principles align, but not in this round's protocol main path).
|
||||
* Real-time streaming alert engine, APM tracing (the OpenTelemetry server side already exists and is orthogonal to this business traffic model).
|
||||
|
||||
---
|
||||
|
||||
## 2. Scope and Constraints
|
||||
|
||||
### 2.1 Product Constraints (inherited)
|
||||
|
||||
* Single-tenant, single globally active config; observability introduces no multi-tenant billing isolation.
|
||||
* Access logs and time-series observability use the switchable log primary DB (ClickHouse by default; switchable to PostgreSQL/SQLite), see [Log Store Decoupling](./logstore.md).
|
||||
* Agent has no inbound control, Pull model; during offline periods local OpenResty keeps serving, and observability can buffer locally and backfill.
|
||||
|
||||
### 2.2 Engineering Constraints
|
||||
|
||||
* Agent stays lightweight: parse log lines, read `/proc`, health checks; no business analysis.
|
||||
* Control-plane API errors still use the unified envelope and `response.Abort*`.
|
||||
* Access log field changes must update both the OpenResty `log_format` and the Agent parser simultaneously; Agent and control plane ship at the same version, no legacy protocol parsing.
|
||||
|
||||
---
|
||||
|
||||
## 3. Design Principles
|
||||
|
||||
### Principle P1: Agent Reports Facts, Server Interprets Facts
|
||||
|
||||
```text
|
||||
Agent = collection + reliable delivery (raw / near-raw)
|
||||
Server = storage + aggregation + ownership + trends + reconciliation
|
||||
```
|
||||
|
||||
**Allowed edge processing (collection)**
|
||||
|
||||
* Parsing JSON access.log lines into structured fields
|
||||
* path length caps, dropping invalid lines, skipping observability-port's own requests
|
||||
* Reading NIC/CPU/memory counters as **raw values**
|
||||
* Batching, compression, offline buffering and retries
|
||||
|
||||
**Forbidden edge processing (business computation)**
|
||||
|
||||
* UV / Top domains / status histograms / window request_count as authoritative metrics
|
||||
* Maintaining "business in/out cumulative" for the dashboard
|
||||
* Zone / domain ownership stats, country distribution (country can be resolved at Server insert time)
|
||||
|
||||
### Principle P2: Single Truth for Business Traffic = Access Logs
|
||||
|
||||
| Business Question | Single Answer |
|
||||
| --- | --- |
|
||||
| How much data was provided | `sum(bytes_sent)` |
|
||||
| How many requests | `count()` |
|
||||
| How many unique visitors | `uniqExact(remote_addr)` (or product-defined hashing) |
|
||||
| Status codes / Top domains | `group by` on logs |
|
||||
|
||||
### Principle P3: Three Metric Layers Never Mixed
|
||||
|
||||
| Layer | Name | Purpose | Typical Fields |
|
||||
| --- | --- | --- | --- |
|
||||
| L1 Business delivery | Business Traffic | user & Zone reconciliation, dashboard business trends | access log |
|
||||
| L2 Edge health | Edge Health | is OpenResty alive, current connections | status, connections |
|
||||
| L3 Host capacity | Host Capacity | capacity planning, is the machine saturated | CPU, memory, disk, **NIC** |
|
||||
|
||||
Never name L3 NIC or L2 instant counts as "data provided"; never draw L1 and L3 on the same summary card without labeling semantics.
|
||||
|
||||
### Principle P4: One Business Concept, One Field
|
||||
|
||||
* **Data provided** ≡ response body delivered ≡ "OpenResty outbound (business meaning)" in legacy copy → **keep only `bytes_sent` aggregation**
|
||||
* **Data received** (optional) ≡ request-side volume → log `request_length` aggregation
|
||||
* **Host outbound** ≡ `network_tx` delta, copy must include "host/NIC"
|
||||
|
||||
---
|
||||
|
||||
## 4. Pre-Refactor Problems (Baseline)
|
||||
|
||||
### 4.1 Pre-Refactor Data Flow (redundant)
|
||||
|
||||
```text
|
||||
One HTTP request
|
||||
│
|
||||
├─ access.log line
|
||||
│ → Agent tail → AccessLogs[]
|
||||
│ → CH of_node_access_logs
|
||||
│ → Zone "data provided" ✅
|
||||
│
|
||||
├─ Lua shared dict window/cumulative counts
|
||||
│ → /openflare/observability
|
||||
│ → TrafficReport + OpenrestyObservation(rx/tx)
|
||||
│ → CH request_reports / obs_openresty
|
||||
│ → dashboard "OpenResty in/outbound" ❌ easily inconsistent with Zone
|
||||
│
|
||||
├─ second access.log aggregation (fallback when observability endpoint fails)
|
||||
│ → yet another TrafficReport / throughput
|
||||
│
|
||||
└─ host network_rx/tx
|
||||
→ Snapshot → "host" curve in network trends
|
||||
```
|
||||
|
||||
### 4.2 Field Overlap
|
||||
|
||||
| User Perception | System Field A | System Field B | Problem |
|
||||
| --- | --- | --- | --- |
|
||||
| Outbound / provided | `openresty_tx_bytes` | `bytes_sent` | duplicate business semantics |
|
||||
| Inbound | `openresty_rx_bytes` | `request_length` (log) | duplicate business semantics |
|
||||
| Request count | `TrafficReport.request_count` | `count(access_logs)` | duplicate aggregation, window easily double-counted |
|
||||
| Outbound (machine) | `network_tx_bytes` | (no business equivalent) | should be named separately, never reconciled with business |
|
||||
|
||||
### 4.3 Typical Failure Modes
|
||||
|
||||
1. Window counts treated as cumulative deltas → 24h business throughput severely underestimated.
|
||||
2. Hourly rollup `max−min` broken for resetting counters.
|
||||
3. Zone uses logs, dashboard uses observability → users think the system is wrong.
|
||||
4. Changing caliber requires syncing Lua, Agent state accumulation, Server deltas, and frontend copy.
|
||||
|
||||
---
|
||||
|
||||
## 5. Target Architecture
|
||||
|
||||
### 5.1 Target Data Flow
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph edge [Edge Node]
|
||||
OR[OpenResty]
|
||||
LOG[access.log]
|
||||
PROC[host /proc and disk]
|
||||
STUB[stub_status connections]
|
||||
AG[Agent]
|
||||
OR -->|log_format writes line| LOG
|
||||
LOG -->|tail incremental details only| AG
|
||||
PROC -->|reading snapshots| AG
|
||||
STUB -->|instant connections| AG
|
||||
OR -->|health probe| AG
|
||||
end
|
||||
|
||||
subgraph server [Control-Plane Server]
|
||||
HB[Heartbeat / WS receive]
|
||||
CH[(ClickHouse)]
|
||||
AGG[Aggregation query layer]
|
||||
API[Admin API]
|
||||
HB --> CH
|
||||
CH --> AGG
|
||||
AGG --> API
|
||||
end
|
||||
|
||||
subgraph ui [Admin Panel]
|
||||
DASH[Dashboard: global business trends]
|
||||
ZONE[Zone: filter by domain]
|
||||
NODE[Node: host resources + health]
|
||||
end
|
||||
|
||||
AG -->|AccessLogs + HostSnapshot + Health| HB
|
||||
API --> DASH
|
||||
API --> ZONE
|
||||
API --> NODE
|
||||
```
|
||||
|
||||
### 5.2 Responsibility Matrix
|
||||
|
||||
| Capability | Agent | Server | Frontend |
|
||||
| --- | --- | --- | --- |
|
||||
| Write access.log | OpenResty | — | — |
|
||||
| Read and report details | ✅ | store | — |
|
||||
| sum/count/uniq/TopN | ❌ | ✅ | display |
|
||||
| Zone domain filtering | ❌ | ✅ | select Zone |
|
||||
| Host CPU/memory/NIC | read raw values and report | delta/average | node/dashboard resource area |
|
||||
| OpenResty connections | read instant and report | latest value | node health |
|
||||
| Business 24h in/outbound | ❌ | log aggregation | uniformly called "data provided/received" |
|
||||
|
||||
---
|
||||
|
||||
## 6. Metrics and Field Model
|
||||
|
||||
### 6.1 Authoritative Field Table (target)
|
||||
|
||||
#### L1 Business Delivery (from access logs)
|
||||
|
||||
| Concept | Storage Field | Aggregation | Display Name |
|
||||
| --- | --- | --- | --- |
|
||||
| Request time | `logged_at` | window filter | — |
|
||||
| Node | `node_id` | group | — |
|
||||
| Client IP | `remote_addr` | `uniq` → UV | Unique visitors |
|
||||
| Host | `host` | group / Zone mapping | Domain |
|
||||
| Path | `path` | optional | — |
|
||||
| Status code | `status_code` | group | Status distribution |
|
||||
| **Data provided** | **`bytes_sent`** | **`sum`** | **Data provided** |
|
||||
| **Data received** | **`request_length`** | **`sum`** | **Data received** (optional display) |
|
||||
| Region | `region` (resolved & written by Server) | group | Source region |
|
||||
|
||||
> Note: the JSON key in the OpenResty `log_format` may keep the name `bytes_sent`; the value must come from **`$body_bytes_sent`** (consistent with production), representing response body delivered, i.e. "data provided".
|
||||
|
||||
#### L2 Edge Health (instant; no 24h business totals)
|
||||
|
||||
| Concept | Field | Description |
|
||||
| --- | --- | --- |
|
||||
| OpenResty health | `openresty_status` / message | existing |
|
||||
| Current connections | `openresty_connections` | stub_status |
|
||||
| (optional) rough recent-window QPS | node detail "right now" only, **never** authoritative 24h totals | if implemented must be labeled "instant" |
|
||||
|
||||
#### L3 Host Capacity
|
||||
|
||||
| Concept | Field | Display Name |
|
||||
| --- | --- | --- |
|
||||
| CPU / memory / disk usage | `host_metrics` | keep |
|
||||
| NIC cumulative bytes | `network_rx_bytes` / `network_tx_bytes` | **Host NIC in/outbound** |
|
||||
| Disk IO cumulative | `disk_read_bytes` / `disk_write_bytes` | Disk read/write |
|
||||
|
||||
### 6.2 Removed Fields (no compatibility layer)
|
||||
|
||||
| Original Field | Disposition | Reason |
|
||||
| --- | --- | --- |
|
||||
| `openresty_tx_bytes` / `openresty_rx_bytes` | **removed** | business bytes follow access logs |
|
||||
| `TrafficReport` and TopN/window UV | **removed** | edge pre-aggregation |
|
||||
| Agent state business lifetime accumulators | removed | violates P1 |
|
||||
| Lua shared dict business throughput/window request counts | removed | not the delivery main path |
|
||||
|
||||
### 6.3 Naming Reference (frontend copy enforced)
|
||||
|
||||
| Forbidden Copy | Correct Copy | Data Source |
|
||||
| --- | --- | --- |
|
||||
| OpenResty outbound (business volume) | **Data provided** | `sum(bytes_sent)` |
|
||||
| OpenResty inbound (business volume) | **Data received** | `sum(request_length)` |
|
||||
| Network outbound (unspecified) | **Host NIC outbound** | `network_tx` delta |
|
||||
| Two cards: data provided vs outbound | **keep only one business card** | logs |
|
||||
|
||||
---
|
||||
|
||||
## 7. Agent Design
|
||||
|
||||
### 7.1 Heartbeat Payload (target protocol)
|
||||
|
||||
Keep and strengthen:
|
||||
|
||||
```text
|
||||
NodePayload
|
||||
identity / version / openresty_status / openresty_message # latest state → PG
|
||||
profile # host overview (low frequency)
|
||||
host_metrics # L3 resource readings (incl. NIC cumulative raw values)
|
||||
edge_health # L2: status + connections (CH time series; message not in CH)
|
||||
access_logs[] # L1 details (main path)
|
||||
health_events[]
|
||||
buffered[] # buffered facts above, not reports
|
||||
waf_ip_group_checksums
|
||||
```
|
||||
|
||||
Removed from the protocol (no compatibility layer):
|
||||
|
||||
```text
|
||||
traffic_report
|
||||
openresty_observation
|
||||
snapshot / buffered_observability aliases
|
||||
```
|
||||
|
||||
### 7.2 Access Log Reporting Requirements
|
||||
|
||||
Each detail at minimum contains:
|
||||
|
||||
| Field | Required | Note |
|
||||
| --- | --- | --- |
|
||||
| `logged_at_unix` | ✅ | request completion time |
|
||||
| `remote_addr` | ✅ | UV |
|
||||
| `host` | ✅ | Zone mapping |
|
||||
| `path` | ✅ | may be truncated |
|
||||
| `status_code` | ✅ | |
|
||||
| `bytes_sent` | ✅ | body bytes, data provided |
|
||||
| `request_length` | ✅ | data received |
|
||||
|
||||
Agent responsibilities:
|
||||
|
||||
1. Tail `access.log` by offset (reset offset on truncation/rotation, **only report new lines still present in the file**).
|
||||
2. Parse into structured form, batch into heartbeat / WS.
|
||||
3. Offline writes to local buffer, backfill by window once connected.
|
||||
4. **No sum/count/uniq on details.**
|
||||
|
||||
### 7.3 Host Snapshot
|
||||
|
||||
* Keep reporting NIC/disk **cumulative counter raw values** (not business pre-aggregation).
|
||||
* Server does non-negative deltas between adjacent samples → host trends.
|
||||
* This is unrelated to "data provided"; the UI must display it in a separate section.
|
||||
|
||||
### 7.4 OpenResty Local Observability
|
||||
|
||||
Converged state:
|
||||
|
||||
* Keep: health checks, `stub_status` current connections.
|
||||
* The main path no longer relies on `log.lua` shared dict business counts; `/openflare/observability` only returns health and connection snapshots, not business report sources.
|
||||
|
||||
### 7.5 Relationship with the Agent Design Doc
|
||||
|
||||
This design strengthens "pure data landing" in [Agent & Publish Model](./agent-design.md):
|
||||
|
||||
* Config and certificates: land and report applied state.
|
||||
* Observability: only carry facts, not business conclusions.
|
||||
|
||||
---
|
||||
|
||||
## 8. Server Design
|
||||
|
||||
### 8.1 Storage
|
||||
|
||||
| Input | Table | Description |
|
||||
| --- | --- | --- |
|
||||
| `access_logs[]` | `of_node_access_logs` | authoritative business details |
|
||||
| `host_metrics` | `of_node_metric_snapshots` | L3; NIC/disk cumulative |
|
||||
| `openresty_status` / `openresty_message` | **PG node table** | L2 **latest-state authority** (message only here) |
|
||||
| `edge_health` | `of_node_edge_health` | L2 time series: status + connections (**no message**) |
|
||||
|
||||
GeoIP: continue resolving `remote_addr` → `region` in the Server insert path, not in the Agent.
|
||||
|
||||
### 8.2 Aggregation Layer (unified)
|
||||
|
||||
All business trends and Zone stats share the same query semantics:
|
||||
|
||||
```text
|
||||
filter: logged_at ∈ [since, until]
|
||||
optional: node_id / host IN (...)
|
||||
metrics:
|
||||
request_count = count()
|
||||
unique_visitors = uniqExact(remote_addr)
|
||||
bytes_provided = sum(bytes_sent) -- data provided
|
||||
bytes_received = sum(request_length) -- data received
|
||||
series folded by hour/bucket
|
||||
distributions by status_code / host / region
|
||||
```
|
||||
|
||||
Implementation locations:
|
||||
|
||||
* Zone: `GET .../zones/:id/stats` (existing, align field naming)
|
||||
* Dashboard: overview traffic / business network trends **switch to the same aggregation** (global, no host filter or Top filter)
|
||||
* Node detail: business volume = the same aggregation filtered by that `node_id`; host NIC still uses metric deltas
|
||||
|
||||
### 8.3 Derived Rollups (optional performance path)
|
||||
|
||||
When detail queries over all nodes for 24h are too heavy, allow **Server-side** materialized views:
|
||||
|
||||
```text
|
||||
of_access_log_hourly
|
||||
(hour, node_id, host, request_count, bytes_sent, bytes_received, ...)
|
||||
```
|
||||
|
||||
Constraints:
|
||||
|
||||
* Derived only by CH from `of_node_access_logs`; **Agent is forbidden from writing this table directly**.
|
||||
* Zone / dashboard prefer reading the rollup, falling back to details (similar to the existing metric hourly policy).
|
||||
|
||||
### 8.4 Decommissioned Analytics Paths
|
||||
|
||||
| Path | After Migration |
|
||||
| --- | --- |
|
||||
| `BuildNetworkTrendPoints` delta on openresty_rx/tx | deleted, or keep only `network_*` host curves |
|
||||
| `of_node_obs_openresty` throughput fields | stop writing; drop table or shrink columns after TTL expiry |
|
||||
| `of_node_request_reports` + traffic hourly | business trends no longer depend on it; table can be deprecated wholesale |
|
||||
| Dashboard compact openresty_tx series | change to bytes_provided series |
|
||||
|
||||
---
|
||||
|
||||
## 9. API and Frontend
|
||||
|
||||
### 9.1 Semantically Unified Response Fields
|
||||
|
||||
Business stats APIs should uniformly use:
|
||||
|
||||
```json
|
||||
{
|
||||
"request_count": 0,
|
||||
"unique_visitors": 0,
|
||||
"bytes_provided": 0,
|
||||
"bytes_received": 0,
|
||||
"series": [
|
||||
{
|
||||
"bucket_started_at": "...",
|
||||
"request_count": 0,
|
||||
"unique_visitors": 0,
|
||||
"bytes_provided": 0,
|
||||
"bytes_received": 0
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
API business byte fields use `bytes_provided` / `bytes_received` (access-log aggregation); no more openresty throughput aliases.
|
||||
|
||||
### 9.2 Dashboard
|
||||
|
||||
* **Business area**: request trend, data provided, data received (optional), status codes, Top domains, source regions — all L1.
|
||||
* **Resource area**: CPU/memory, **host NIC**, disk IO — all L3.
|
||||
* **Forbidden**: showing "OpenResty in/outbound" in the business area as a metric reconciled with Zone.
|
||||
|
||||
Suggest splitting or retitling "24-hour network and disk trends":
|
||||
|
||||
* "24-hour business traffic" → `bytes_provided` / `bytes_received` / requests
|
||||
* "24-hour host network and disk" → `network_*` / `disk_*`
|
||||
|
||||
### 9.3 Zone `/websites/:id`
|
||||
|
||||
* Keep cards like "total data provided".
|
||||
* Data and the dashboard business area use **the same aggregation function**, only `hosts = zone domain list`.
|
||||
* Docs and UI may note: the global dashboard includes all Hosts; this page is only this Zone.
|
||||
|
||||
### 9.4 Node Detail
|
||||
|
||||
* Business throughput: that node's `sum(bytes_sent)` etc.
|
||||
* OpenResty: health + current connections.
|
||||
* NIC: clearly "host".
|
||||
|
||||
---
|
||||
|
||||
## 10. OpenResty and Log Format
|
||||
|
||||
### 10.1 Keep
|
||||
|
||||
Existing JSON `log_format` core fields:
|
||||
|
||||
```text
|
||||
ts, host, path, remote_addr, status, request_time,
|
||||
bytes_sent (= $body_bytes_sent), request_length
|
||||
```
|
||||
|
||||
### 10.2 Changes
|
||||
|
||||
* No longer rely on log-phase writes of business shared dict counts as control-plane input.
|
||||
* Observability-port requests continue not writing business stats (or `access_log off`).
|
||||
|
||||
### 10.3 Agent Parsing
|
||||
|
||||
* Protocol `NodeAccessLog` adds `request_length`.
|
||||
* Legacy log lines missing fields default to 0, not blocking the whole batch.
|
||||
|
||||
---
|
||||
|
||||
## 11. Upgrade and Migration (no compatibility layer)
|
||||
|
||||
### 11.1 Phase Review (shipped)
|
||||
|
||||
| Phase | Content |
|
||||
| --- | --- |
|
||||
| **M1–M5** | read path switches to access logs; protocol v2; stop pre-aggregation; edge_health + access_log_hourly; drop old tables and API compat fields |
|
||||
|
||||
### 11.2 Upgrade Strategy
|
||||
|
||||
* **Agent: destroy-and-recreate preferred**; binary replacement allowed.
|
||||
* On binary replacement: the local old observability buffer (including `snapshot` / `openresty_observation` / `traffic_report`) is **deleted wholesale**, rebuilt after running.
|
||||
* Server **does not** parse v1 fields, **does not** dual-read request_reports / openresty throughput.
|
||||
* Detail-missing periods: business charts are empty or partial; **never** impersonate data provided with NIC or removed openresty throughput.
|
||||
|
||||
### 11.3 Data Backfill
|
||||
|
||||
* Historical "data provided" follows access logs.
|
||||
* Before `of_access_log_hourly` is created, history is backfilled with goose SQL (ANTI JOIN to prevent duplicates).
|
||||
|
||||
### 11.4 Health-State Authority
|
||||
|
||||
* **Current state**: PG `openresty_status` / `openresty_message`.
|
||||
* **Time series**: log primary DB `of_node_edge_health` (status + connections; no message).
|
||||
|
||||
### 11.5 UV
|
||||
|
||||
* **Whole-window unique visitors**: `uniqExact(remote_addr)` (dashboard totals, Zone totals).
|
||||
* **Bucketed UV** (Zone curves): per-bucket uniq, **not summable across buckets**; UI must note it.
|
||||
* **Hourly trend path**: don't plot / fill per-hour UV (hourly table has no UV).
|
||||
|
||||
---
|
||||
|
||||
## 12. Storage and Capacity
|
||||
|
||||
* Business trends rely on details or hourly rollups; watch `of_node_access_logs` TTL and sampling.
|
||||
* If details are too large: prefer **Server-side rollup** rather than restoring Agent pre-aggregation.
|
||||
* For high-cardinality path scenarios, limit detail path length (existing); aggregation doesn't do global Top over full paths by default.
|
||||
|
||||
---
|
||||
|
||||
## 13. Risks and Trade-offs
|
||||
|
||||
| Risk | Mitigation |
|
||||
| --- | --- |
|
||||
| Large detail volume makes CH and heartbeat heavy | batching, compression, sampling policy evaluation; Server rollup; limit per-batch count |
|
||||
| Brief log loss lowers business volume | local buffer and rotation handling; monitor access log collection lag |
|
||||
| Users still compare "NIC outbound" with "data provided" | UI sections and copy enforce the "host" prefix |
|
||||
| Old Agents stay online long-term | **no compatibility layer**; Agents must be upgraded/rebuilt |
|
||||
|
||||
**Why not keep Agent pre-aggregation as an optimization?**
|
||||
|
||||
* Saving bandwidth re-splits the truth, drifts calibers, and repeats this problem.
|
||||
* Optimization belongs in Server derived tables and queries, not edge business computation.
|
||||
|
||||
---
|
||||
|
||||
## 14. Key Decision Summary
|
||||
|
||||
| Decision | Choice | Rejected Alternative |
|
||||
| --- | --- | --- |
|
||||
| Business traffic truth | access logs | OpenResty dict / TrafficReport |
|
||||
| Agent role | report only facts | edge UV/TopN/throughput accumulation |
|
||||
| "Outbound" vs "provided" | merged into data provided | dual fields and dual pipelines long-term |
|
||||
| NIC traffic | independent L3, separate copy | reconciled side-by-side with business outbound |
|
||||
| Performance | CH rollup | Agent pre-aggregation |
|
||||
| Migration | switch read path first, then slim Agent | drop details first, rely on pre-aggregation |
|
||||
|
||||
---
|
||||
|
||||
## 15. Doc and Code Mapping
|
||||
|
||||
| Area | Main Paths |
|
||||
| --- | --- |
|
||||
| Protocol | `pkg/protocol/agent.go` |
|
||||
| Agent collection | `internal/apps/agent/observability/`, `heartbeat/` |
|
||||
| OpenResty logging and Lua | `pkg/render/openresty/`, `internal/apps/agent/nginx/observability_assets.go` |
|
||||
| Server storage | `internal/apps/openflare/agent/observability.go` |
|
||||
| Log aggregation | `internal/repository/analytics/node_access_log*.go`, `internal/apps/openflare/zone/stats.go` |
|
||||
| Dashboard | `internal/apps/openflare/dashboard/`, `internal/apps/openflare/observability/analytics.go` |
|
||||
| Frontend | `frontend/app/(main)/page.tsx`, `components/dashboard/*`, `websites/.../zone-overview.tsx` |
|
||||
|
||||
**Recommended reading order:**
|
||||
|
||||
1. **[Observability Transport Model](./observability-transport-model.md)** (latest: what to send, where collected from, frequency, sample JSON)
|
||||
2. [Agent Reporting Protocol and Observability Data Model](./observability-data-model.md) (protocol fields and DDL)
|
||||
|
||||
---
|
||||
|
||||
## 16. Revision History
|
||||
|
||||
| Date | Notes |
|
||||
| --- | --- |
|
||||
| 2026-07-17 | initial draft: target architecture and migration phases for dual truth, Agent pre-aggregation, field redundancy |
|
||||
| 2026-07-17 | added protocol/table-structure chapter links `observability-data-model.md` |
|
||||
@@ -0,0 +1,502 @@
|
||||
# Edge Observability Transport Model (current target version)
|
||||
|
||||
> **This document is the latest authoritative description of "how Agent ↔ Server observability data is transmitted".**
|
||||
> After reading you should be able to answer: what is sent, where it is collected from, how often, how the Server stores it, and where product metrics are queried from.
|
||||
> Protocol fields and DDL details: [Observability Reporting Protocol & Data Model](./observability-data-model.md); background: [Edge Observability & Business Traffic Stats](./observability-design.md).
|
||||
|
||||
---
|
||||
|
||||
## 0. Remember the Three Layers First
|
||||
|
||||
| Layer | Question Answered | Single Data Source | Product Examples |
|
||||
| --- | --- | --- | --- |
|
||||
| **L1 Business delivery** | How much data provided? How many requests? | **access.log details** | data provided, request count, UV, status codes, Top domains |
|
||||
| **L2 Edge health** | Is OpenResty alive? Current connections? | **local `/openflare/observability`** | node health, current connections |
|
||||
| **L3 Host capacity** | How are CPU/memory/disk/NIC? | **OS readings** | capacity trends, host NIC |
|
||||
|
||||
**The three layers are never reconciled against each other.**
|
||||
"Data provided" ≠ "current connections" ≠ "host NIC outbound".
|
||||
|
||||
---
|
||||
|
||||
## 1. Overview: Who Collects, Who Reports, Who Aggregates
|
||||
|
||||
```text
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Edge Node │
|
||||
│ │
|
||||
│ Visitor request ──► OpenResty │
|
||||
│ │ │
|
||||
│ ├─ access.log (one line per request) ←── L1 collection point │
|
||||
│ │ │
|
||||
│ └─ connection state (maintained in-process) │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ GET /openflare/observability ←── L2 reads snapshot │
|
||||
│ (no log scanning, no business recomputation) │
|
||||
│ │
|
||||
│ OS /proc etc. ──────────────────────────── L3 reads snapshot │
|
||||
│ │
|
||||
│ ┌────────── Agent ──────────┐ │
|
||||
│ │ default: one NodePayload per 3s │ │
|
||||
│ │ · tail access.log incremental │ │
|
||||
│ │ · GET local observability │ │
|
||||
│ │ · read host_metrics │ │
|
||||
│ └────────────┬──────────────┘ │
|
||||
└─────────────────────────────│──────────────────────────────────┘
|
||||
│ HTTP heartbeat or WebSocket status
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Server (control plane) │
|
||||
│ · details → ClickHouse of_node_access_logs │
|
||||
│ · health → node latest state + of_node_edge_health │
|
||||
│ · host → of_node_metric_snapshots │
|
||||
│ · business trends / Zone stats = sum/count/uniq over access_logs only │
|
||||
└─────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
| Role | Does | Doesn't |
|
||||
| --- | --- | --- |
|
||||
| OpenResty | writes access.log; maintains connection counts | no direct reporting to the control plane |
|
||||
| Agent | **collects facts and reports them** | **does not compute** UV/TopN/24h data provided |
|
||||
| Server | stores + **aggregates/interpreters** | does not trust edge business pre-summaries |
|
||||
|
||||
---
|
||||
|
||||
## 2. Collection Frequency (defaults)
|
||||
|
||||
| Action | Default Frequency | Config |
|
||||
| --- | --- | --- |
|
||||
| Agent → Server report | **every 3 seconds** a full payload | `heartbeat_interval` / control-plane `agent_heartbeat_interval` (ms, default `3000`) |
|
||||
| Tail access.log when packing | **with report** (new lines since last report) | same |
|
||||
| GET `/openflare/observability` when packing | **with report** (reads **current** connection snapshot) | same |
|
||||
| Read host metrics when packing | **with report** | same |
|
||||
| OpenResty writes access.log | **1 line at each request end** | unrelated to heartbeat |
|
||||
| Connection counts update in-process | **on connection change** (kernel-maintained) | unrelated to heartbeat |
|
||||
| Offline replay window | keep ~**60 minutes** by default | `observability_replay_minutes` |
|
||||
| Node offline detection | ~**60s** without a successful heartbeat | `node_offline_threshold` (default `60000` ms) |
|
||||
|
||||
**Notes:**
|
||||
|
||||
- The Agent has **no separate "sampling clock"**; **sampling points = report points** (default 3s).
|
||||
- access.log is "per-request continuous writes"; the Agent only **moves incremental lines** periodically.
|
||||
- `/openflare/observability` is **not** "business stats start being counted when called"; for connections it **reads Nginx's existing instantaneous values**.
|
||||
|
||||
Transport channels:
|
||||
|
||||
- **HTTP heartbeat**: POST the full payload at the interval.
|
||||
- **WebSocket**: after connecting, sends `status` messages at the same interval (same content shape); HTTP heartbeat is not double-sent then.
|
||||
|
||||
---
|
||||
|
||||
## 3. Agent → Server Packet (NodePayload v2)
|
||||
|
||||
### 3.1 Structure Skeleton
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n_01hxyz",
|
||||
"name": "edge-shanghai-1",
|
||||
"ip": "203.0.113.10",
|
||||
"version": "3.4.0",
|
||||
"ext_version": "",
|
||||
"current_version": "20260718-abc",
|
||||
"last_error": "",
|
||||
"profile": { },
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ],
|
||||
"buffered": [ ],
|
||||
"health_events": [ ],
|
||||
"waf_ip_group_checksums": { }
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Layer | Meaning |
|
||||
| --- | --- | --- |
|
||||
| identity/version/last_error | control | who the node is, what version it runs |
|
||||
| `profile` | low-frequency overview | hostname, core count, etc. (report on change) |
|
||||
| `access_logs` | **L1** | access detail increments |
|
||||
| `edge_health` | **L2** | OpenResty health + current connections |
|
||||
| `host_metrics` | **L3** | CPU/memory/disk/NIC readings |
|
||||
| `buffered` | backfill | batches of facts accumulated while offline |
|
||||
| `health_events` | events | e.g. openresty_unhealthy |
|
||||
| `waf_ip_group_checksums` | sync | not an observability lake |
|
||||
|
||||
**Removed from the protocol (no compatibility layer; old Agents must upgrade):**
|
||||
|
||||
- `traffic_report`
|
||||
- `openresty_observation` (incl. rx/tx)
|
||||
- `snapshot` / `buffered_observability`
|
||||
- business-meaning openresty throughput fields
|
||||
|
||||
---
|
||||
|
||||
## 4. L1 Business: access_logs
|
||||
|
||||
### 4.1 Where Collection Comes From
|
||||
|
||||
| Step | Location | Description |
|
||||
| --- | --- | --- |
|
||||
| 1 | OpenResty `log_format openflare_json` | writes one JSON line per request to `access_log_path` |
|
||||
| 2 | Agent **tails increments** by file offset | new lines between two heartbeats |
|
||||
| 3 | parse and put into `access_logs[]` | overlong paths may be truncated; **no sum/count** |
|
||||
|
||||
Log format (OpenResty variables):
|
||||
|
||||
```text
|
||||
ts ← $time_iso8601
|
||||
host ← $host
|
||||
path ← $request_uri
|
||||
remote_addr ← $remote_addr
|
||||
status ← $status
|
||||
request_time ← $request_time
|
||||
bytes_sent ← $body_bytes_sent 【data provided = response body bytes】
|
||||
request_length← $request_length 【data received】
|
||||
user_agent ← $http_user_agent
|
||||
cache_status ← $upstream_cache_status 【cache status; UI can derive hit/origin/un-cached】
|
||||
```
|
||||
|
||||
Observability-port requests **don't write** business access.log (separate server with `access_log off`).
|
||||
|
||||
### 4.2 Report Example
|
||||
|
||||
```json
|
||||
"access_logs": [
|
||||
{
|
||||
"logged_at_unix": 1721289601,
|
||||
"remote_addr": "198.51.100.20",
|
||||
"host": "www.example.com",
|
||||
"path": "/api/v1/ping",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 1024,
|
||||
"request_length": 128,
|
||||
"request_time_ms": 15,
|
||||
"user_agent": "curl/8.0",
|
||||
"cache_status": "MISS"
|
||||
},
|
||||
{
|
||||
"logged_at_unix": 1721289602,
|
||||
"remote_addr": "198.51.100.21",
|
||||
"host": "www.example.com",
|
||||
"path": "/index.html",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 8192,
|
||||
"request_length": 300,
|
||||
"request_time_ms": 8,
|
||||
"user_agent": "Mozilla/5.0",
|
||||
"cache_status": "HIT"
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
| Field | Explanation |
|
||||
| --- | --- |
|
||||
| `bytes_sent` | **data provided** (single request); global/Zone totals = Server `sum` |
|
||||
| `request_length` | **data received** (single request) |
|
||||
| `logged_at_unix` | request completion time (business timeline) |
|
||||
| `host` | used for Zone domain filtering |
|
||||
| `cache_status` | `$upstream_cache_status` as-is; detail/list can derive three states (hit/origin/un-cached); **no upstream address reported** |
|
||||
| no `region` | **written by Server at insert time** via GeoIP |
|
||||
|
||||
### 4.3 How the Server Uses It (product metrics)
|
||||
|
||||
| Product Metric | Algorithm (L1 only) |
|
||||
| --- | --- |
|
||||
| Data provided | `sum(bytes_sent)` |
|
||||
| Data received | `sum(request_length)` |
|
||||
| Request count | `count()` |
|
||||
| UV | `uniqExact(remote_addr)` |
|
||||
| Status distribution | `group by status_code` |
|
||||
| Top domains | `group by host` |
|
||||
| Zone page | same + `host IN (that Zone's domains)` |
|
||||
| Dashboard business area | same, global or Top-filtered |
|
||||
|
||||
Stored in: `of_node_access_logs` (optional Server-side `of_access_log_hourly` acceleration, **Agent never writes it**).
|
||||
|
||||
### 4.4 Report Frequency
|
||||
|
||||
```text
|
||||
Request happens ──immediately──► write access.log
|
||||
Agent every 3s ──moves──► new lines in those 3s (possibly 0, possibly many)
|
||||
Server ──immediately/batched──► CH
|
||||
```
|
||||
|
||||
Business volume correctness does **not** depend on 3s alignment; 3s only affects "detail arrival latency at the control plane" and per-packet line count.
|
||||
|
||||
---
|
||||
|
||||
## 5. L2 Health: edge_health and `/openflare/observability`
|
||||
|
||||
### 5.1 Local Monitoring Endpoint
|
||||
|
||||
**Data collection endpoint:**
|
||||
|
||||
```http
|
||||
GET http://127.0.0.1:{openresty_observability_port}/openflare/observability
|
||||
```
|
||||
|
||||
Default port: **18081** (`openresty_observability_port`).
|
||||
|
||||
**Responsibility:** answers "how is OpenResty right now", **not** "how much business data was provided".
|
||||
|
||||
#### Response Example
|
||||
|
||||
```json
|
||||
{
|
||||
"ok": true,
|
||||
"captured_at_unix": 1721289600,
|
||||
"connections": {
|
||||
"active": 42,
|
||||
"reading": 0,
|
||||
"writing": 1,
|
||||
"waiting": 41
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Instant? | Source | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `ok` | this probe | returns 200 → true | liveness |
|
||||
| `captured_at_unix` | sampling time | `ngx.time()` | aligned with report |
|
||||
| `connections.active` | **instant** | Nginx connection state (original stub_status Active) | current active connections |
|
||||
| `reading` / `writing` / `waiting` | **instant** | same, subdivided | optional but recommended |
|
||||
|
||||
**Not returned (removed):**
|
||||
|
||||
| Old Field | Reason |
|
||||
| --- | --- |
|
||||
| `request_count` / `error_count` / UV / status_codes / top_domains | business window summaries, now from access log |
|
||||
| `openresty_rx_bytes` / `openresty_tx_bytes` | duplicates data provided/received and error-prone |
|
||||
| `source_countries` | never implemented; countries go through Server GeoIP |
|
||||
| `server.accepts/handled/requests` | process cumulative counters, easily confused with business requests; not on the main path |
|
||||
|
||||
**`/openflare/stub_status`:** kept; `/openflare/observability` internally reads that endpoint to assemble the connection-count JSON, and the Agent health check also probes it directly.
|
||||
|
||||
### 5.2 Collection Mechanism (read snapshot)
|
||||
|
||||
```text
|
||||
Nginx maintains Active connections etc. on connect/disconnect
|
||||
│
|
||||
Agent GET /openflare/observability
|
||||
│
|
||||
only reads "current values" and returns JSON
|
||||
```
|
||||
|
||||
- No access.log scanning, no 60-second business averages.
|
||||
- Returns an **instant gauge snapshot**.
|
||||
|
||||
### 5.3 Report Example (packed into NodePayload)
|
||||
|
||||
```json
|
||||
"edge_health": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 42
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Source |
|
||||
| --- | --- |
|
||||
| `status` / `message` | Agent health probe (config validation/process etc., may work with the observability endpoint's `ok`); must align with top-level `openresty_status` / `openresty_message` |
|
||||
| `connections` | observability endpoint `connections.active` |
|
||||
|
||||
**Storage split (authoritative sources):**
|
||||
|
||||
| Content | Written To |
|
||||
| --- | --- |
|
||||
| latest `status` + `message` | **PG node table** (UI / list / alerts) |
|
||||
| time-series `status` + `connections` | **CH `of_node_edge_health`** (**no message**) |
|
||||
|
||||
---
|
||||
|
||||
## 6. L3 Host: host_metrics
|
||||
|
||||
### 6.1 Where Collection Comes From
|
||||
|
||||
The Agent reads the local machine (e.g. `/proc`, disk stats), **once per packet**.
|
||||
|
||||
| Field | Semantics | Description |
|
||||
| --- | --- | --- |
|
||||
| `cpu_usage_percent` | instant | current CPU% |
|
||||
| `memory_*` / `storage_*` | instant used/total | usage rates computed at Server or display layer |
|
||||
| `disk_read_bytes` / `disk_write_bytes` | **cumulative counter** | kernel cumulative IO |
|
||||
| `network_rx_bytes` / `network_tx_bytes` | **cumulative counter** | **host NIC**, not data provided |
|
||||
|
||||
### 6.2 Report Example
|
||||
|
||||
```json
|
||||
"host_metrics": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"cpu_usage_percent": 12.5,
|
||||
"memory_used_bytes": 4294967296,
|
||||
"memory_total_bytes": 16106127360,
|
||||
"storage_used_bytes": 50000000000,
|
||||
"storage_total_bytes": 107374182400,
|
||||
"disk_read_bytes": 9000000000,
|
||||
"disk_write_bytes": 12000000000,
|
||||
"network_rx_bytes": 500000000000,
|
||||
"network_tx_bytes": 800000000000
|
||||
}
|
||||
```
|
||||
|
||||
### 6.3 How the Server Handles Cumulative Fields
|
||||
|
||||
```text
|
||||
store raw-value time series
|
||||
when displaying "NIC outbound over this period":
|
||||
delta = current - previous
|
||||
if delta < 0 → treat as restart/counter reset, record this segment's increment as 0, continue from new baseline
|
||||
if delta >= 0 → record into that period's increment
|
||||
```
|
||||
|
||||
- The Agent **reports raw values**, never computes 24h totals at the edge.
|
||||
- **Forbidden** to `sum` cumulative raw values as business volume.
|
||||
- Copy must be **"host NIC"**, never "data provided / OpenResty outbound".
|
||||
|
||||
Stored in: `of_node_metric_snapshots` (optional capacity hourly MV).
|
||||
|
||||
---
|
||||
|
||||
## 7. One Complete Report Example
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n_01hxyz",
|
||||
"name": "edge-shanghai-1",
|
||||
"ip": "203.0.113.10",
|
||||
"version": "3.4.0",
|
||||
"ext_version": "",
|
||||
"current_version": "20260718-abc",
|
||||
"last_error": "",
|
||||
"host_metrics": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"cpu_usage_percent": 12.5,
|
||||
"memory_used_bytes": 4294967296,
|
||||
"memory_total_bytes": 16106127360,
|
||||
"storage_used_bytes": 50000000000,
|
||||
"storage_total_bytes": 107374182400,
|
||||
"disk_read_bytes": 9000000000,
|
||||
"disk_write_bytes": 12000000000,
|
||||
"network_rx_bytes": 500000000000,
|
||||
"network_tx_bytes": 800000000000
|
||||
},
|
||||
"edge_health": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 42
|
||||
},
|
||||
"access_logs": [
|
||||
{
|
||||
"logged_at_unix": 1721289595,
|
||||
"remote_addr": "198.51.100.20",
|
||||
"host": "www.example.com",
|
||||
"path": "/",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 4096,
|
||||
"request_length": 200,
|
||||
"request_time_ms": 12
|
||||
}
|
||||
],
|
||||
"buffered": [],
|
||||
"health_events": [],
|
||||
"waf_ip_group_checksums": {
|
||||
"1": "d41d8cd98f00b204e9800998ecf8427e"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Server storage sketch:**
|
||||
|
||||
| payload block | written to |
|
||||
| --- | --- |
|
||||
| `access_logs[0]` | one CH row, `bytes_sent=4096`, `region` filled by GeoIP |
|
||||
| `edge_health` | node `openresty_status=healthy`, connections=42 |
|
||||
| `host_metrics` | one CH metric row with cumulative/instant fields |
|
||||
|
||||
**Product query sketch (24h):**
|
||||
|
||||
- data provided = `sum(bytes_sent)` over that node's (or global) logs
|
||||
- current connections = latest `edge_health.connections`
|
||||
- host NIC outbound = sum of non-negative `network_tx` deltas over metrics
|
||||
|
||||
The three numbers **need not be equal**.
|
||||
|
||||
---
|
||||
|
||||
## 8. Offline Backfill `buffered`
|
||||
|
||||
When reporting fails, the Agent caches **the same kind of facts** locally by window (default ~60 minutes), then packs them into `buffered[]` after recovery:
|
||||
|
||||
```json
|
||||
"buffered": [
|
||||
{
|
||||
"captured_at_unix": 1721289500,
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ]
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
- Only facts, no legacy TrafficReport.
|
||||
- Server processing logic is identical to the main fields.
|
||||
|
||||
---
|
||||
|
||||
## 9. End-to-End Timeline (default 3s)
|
||||
|
||||
```text
|
||||
t=0.0s visitor request completes → writes one access.log line; connection count may change
|
||||
t=0.1s another request → another log line
|
||||
…
|
||||
t=3s Agent heartbeat:
|
||||
· reads 2 access_logs lines
|
||||
· GET observability → connections=42
|
||||
· reads host_metrics
|
||||
· sends to Server
|
||||
t=3s+ Server stores; dashboard/Zone queries aggregate logs
|
||||
t=6s next round…
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 10. Old Model Comparison
|
||||
|
||||
| Old Approach | New Model |
|
||||
| --- | --- |
|
||||
| Lua dict 60s window request_count + Agent 10s pull + Server sum | **removed**; request count = log count |
|
||||
| openresty_tx as "outbound" | **removed**; data provided = `sum(bytes_sent)` |
|
||||
| Two endpoints observability + stub_status | data collection unified through observability; stub_status kept as liveness and internal read endpoint |
|
||||
| TrafficReport pre-aggregation | **removed**; no such path in protocol or API |
|
||||
| Business and NIC both called "traffic" | **separate copy, separate APIs, separate tables** |
|
||||
| health status/message | **PG latest-state authority**; CH only status+connection time series |
|
||||
|
||||
---
|
||||
|
||||
## 11. Config and Implementation Index
|
||||
|
||||
| Item | Location/Key |
|
||||
| --- | --- |
|
||||
| Heartbeat interval | Agent `heartbeat_interval`; control plane `agent_heartbeat_interval` (default 3000ms) |
|
||||
| Offline threshold | control plane `node_offline_threshold` (default 60000ms) |
|
||||
| Observability port | `openresty_observability_port` (default 18081) |
|
||||
| access.log path | `access_log_path` |
|
||||
| Replay minutes | `observability_replay_minutes` (default 60) |
|
||||
| Protocol types | `pkg/protocol/agent.go` (evolves to v2 when landing) |
|
||||
| Table DDL | [observability-data-model.md](./observability-data-model.md) |
|
||||
|
||||
---
|
||||
|
||||
## 12. Revision History
|
||||
|
||||
| Date | Notes |
|
||||
| --- | --- |
|
||||
| 2026-07-18 | initial draft: single-page "latest transport model" — three layers, frequency, sample JSON, collection sources, old-model comparison |
|
||||
| 2026-07-18 | default report interval 3s; offline threshold 60s; replay window 60 minutes |
|
||||
| 2026-07-18 | M5: edge_health table, access_log_hourly, deprecate request_reports/obs_openresty throughput tables |
|
||||
| 2026-07-18 | no compatibility layer: removed "may ignore during compat period" wording; health message only in PG, no message in CH |
|
||||
@@ -0,0 +1,204 @@
|
||||
# Origin Error Page Design
|
||||
|
||||
You will learn: when an origin or the gateway returns a specified error status code, how OpenFlare replaces the pass-through response with a globally configurable page; how the config enters the immutable config version; and how the edge OpenResty keeps the real HTTP status code while displaying it in the page.
|
||||
|
||||
This design is the productized complement of the reverse proxy traffic path in [System Architecture](./architecture.md); the config release model is in [Agent & Publish Model](./agent-design.md).
|
||||
|
||||
---
|
||||
|
||||
## 1. Goals and Non-Goals
|
||||
|
||||
### 1.1 Goals
|
||||
|
||||
* **Interceptable**: for a user-configured status code set, replace the previously pass-through origin/Nginx default error response with a unified HTML.
|
||||
* **Disableable**: when the global switch is off, behavior matches today (pass-through / Nginx default page).
|
||||
* **Visible by default**: enabled by default, default status code tag `500-599`, default minimal OpenFlare error page.
|
||||
* **Customizable**: admins can edit the full HTML online; empty HTML means the built-in default template.
|
||||
* **Status passthrough**: the HTTP response `status` keeps the original error code (e.g. 502, 522); the page body shows the same value via `{{status}}`.
|
||||
* **Globally unified**: a single config under sidebar「Website Management → Response Pages」shared by all reverse proxy routes.
|
||||
* **Consistent with release**: the config persists via Option, enters the config version snapshot, and is distributed with release/rollback.
|
||||
|
||||
### 1.2 Non-Goals
|
||||
|
||||
* Per-route / per-Zone error page overrides
|
||||
* Hosting error pages via file upload (online HTML only)
|
||||
* Modifying WAF / PoW / rate-limit's own response pages (unless the user adds those status codes to the list)
|
||||
* Pages static route error pages
|
||||
* Multi-language error pages, brand asset CDN
|
||||
|
||||
---
|
||||
|
||||
## 2. Product Behavior
|
||||
|
||||
### 2.1 When to Replace
|
||||
|
||||
| Condition | Behavior |
|
||||
| --- | --- |
|
||||
| Switch on and the response status falls in the expanded set | return custom/default HTML, **status unchanged** |
|
||||
| Switch on with GET-only enabled, non-GET request returns a matching status | pass through the origin's raw response, no replacement |
|
||||
| Switch off | no `error_page` directives generated, pass through |
|
||||
| Status not in the set | no replacement |
|
||||
| Pages upstream routes | this feature is not applied |
|
||||
| Origin returns 2xx/3xx/4xx successfully (not configured) | no replacement |
|
||||
|
||||
In all-methods mode, `proxy_intercept_errors on` is enabled on the reverse proxy `location`, so **origin-returned** matching 5xx etc. are also intercepted, not just gateway-local 502s; GET-only mode switches to Lua header/body filters that only replace GET response bodies.
|
||||
|
||||
### 2.2 Status Code Tag Syntax
|
||||
|
||||
Each Tags Input entry:
|
||||
|
||||
| Form | Example | Meaning |
|
||||
| --- | --- | --- |
|
||||
| Single code | `522` | only that code |
|
||||
| Closed range | `500-599` | expand including endpoints |
|
||||
|
||||
* Valid range: single codes and range endpoints must be in **400–599**; `lo ≤ hi`.
|
||||
* Default tag list: `["500-599"]`.
|
||||
* Persist the **raw tags** (JSON array string); expand, dedupe, and sort at render time.
|
||||
* If the expanded result is empty while enabled → save rejected.
|
||||
* Invalid tags → save rejected with a readable error.
|
||||
|
||||
### 2.3 Page Placeholders
|
||||
|
||||
| Placeholder | Meaning |
|
||||
| --- | --- |
|
||||
| `{{status}}` | the current response status code (consistent with the HTTP status) |
|
||||
| `{{host}}` | request Host |
|
||||
|
||||
Both custom HTML and the default template support these placeholders; replaced at the edge at runtime. Unused placeholders may be omitted from the template.
|
||||
|
||||
### 2.4 Default Page
|
||||
|
||||
Built-in minimal white-background OpenFlare default page: large pass-through status code, short English description, Host, and a brand footer. Supports `{{status}}` / `{{host}}`; the frontend can load prebuilt styles from the built-in template catalog on the edit page.
|
||||
|
||||
---
|
||||
|
||||
## 3. Config Model
|
||||
|
||||
### 3.1 Option Keys (`w_system_configs` / OpenFlare Option API)
|
||||
|
||||
| Key | Type | Default | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `origin_error_page_enabled` | bool string | `true` | master switch |
|
||||
| `origin_error_page_status_codes` | JSON string array | `["500-599"]` | raw tags |
|
||||
| `origin_error_page_html` | text | `""` | empty = built-in default; max **256 KiB** |
|
||||
| `origin_error_page_get_only` | bool string | `false` | replace error pages only for GET; other methods pass through |
|
||||
|
||||
Reuses APIs:
|
||||
|
||||
* `GET /api/v1/d/option`
|
||||
* `POST /api/v1/d/option/update-batch`
|
||||
|
||||
No new resource routes. goose migration writes the seed; constants defined in the `internal/model` config key area.
|
||||
|
||||
### 3.2 Validation (update-batch)
|
||||
|
||||
1. `enabled`: parseable as bool.
|
||||
2. `status_codes`: valid JSON array; each entry `^\d{3}$` or `^\d{3}-\d{3}$`; expanded values all in 400–599; non-empty when enabled.
|
||||
3. `html`: length ≤ 256 KiB (bytes); empty allowed.
|
||||
4. Parse/expand logic is a **pure function** shared by the API and `pkg/render/openresty` to avoid semantic forks.
|
||||
|
||||
No XSS sanitization on HTML: it's an admin global ops config consistent with public edge display; docs warn not to embed untrusted third-party scripts.
|
||||
|
||||
### 3.3 Config Version Snapshot
|
||||
|
||||
`ConfigSnapshot` adds fields:
|
||||
|
||||
```text
|
||||
OriginErrorPageEnabled bool
|
||||
OriginErrorPageStatusCodes []string // raw tags
|
||||
OriginErrorPageHTML string // empty => renderer uses built-in default
|
||||
OriginErrorPageGetOnly bool
|
||||
```
|
||||
|
||||
Read from Option when building the snapshot; the Agent only consumes the snapshot, never reading the control-plane DB directly.
|
||||
|
||||
---
|
||||
|
||||
## 4. Edge Rendering
|
||||
|
||||
### 4.1 Content Generated When Enabled
|
||||
|
||||
1. **SupportFile**: error page template (e.g. `error_pages/origin_error.html.tmpl`), content is the custom HTML or built-in default, keeping `{{status}}` / `{{host}}`.
|
||||
2. **Each reverse proxy server** (HTTP/HTTPS proxy; excluding Pages):
|
||||
|
||||
```nginx
|
||||
proxy_intercept_errors on;
|
||||
error_page <expanded codes...> @__openflare_origin_error;
|
||||
|
||||
location @__openflare_origin_error {
|
||||
default_type text/html;
|
||||
charset utf-8;
|
||||
content_by_lua_block {
|
||||
# read template, replace {{status}} / {{host}}, output body
|
||||
# ngx.status keeps the original error code
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 4.2 Runtime Replacement
|
||||
|
||||
Use a **named location with `content_by_lua_block`** to read the template and replace placeholders — the status is **not** baked into a static file (status differs per request). GET-only mode uses `header_filter_by_lua_block` + `body_filter_by_lua_block` inside the reverse proxy location to replace only GET response bodies; non-GET requests pass through.
|
||||
|
||||
Never rewrite the error page to HTTP 200.
|
||||
|
||||
### 4.3 When Disabled
|
||||
|
||||
Do not output `proxy_intercept_errors`, `error_page`, the internal location, or the corresponding SupportFile (or the file may be written but unreferenced). GET-only mode also omits the Lua filters.
|
||||
|
||||
### 4.4 Interaction with Cache / Stale
|
||||
|
||||
If global `proxy_cache_use_stale` returns stale cache for some error codes, **successful stale responses never enter `error_page`**. The error page is only shown when the client actually receives an error status in the configured list. Behavior depends on existing cache directives; this feature does not change stale policy.
|
||||
|
||||
---
|
||||
|
||||
## 5. Frontend
|
||||
|
||||
### 5.1 Entry
|
||||
|
||||
* Sidebar「Website Management → Response Pages」: Error Page tab (`/responses`), edit page `/responses/error-page/edit`, preview page `/responses/error-page/preview`.
|
||||
|
||||
### 5.2 Page Structure
|
||||
|
||||
* Header note: takes effect after releasing via「Version Release」.
|
||||
* **Switch + Tags Input** (shadcn-extension Tags Input: `@/components/ui/tags-input`): status code tags.
|
||||
* **HTML editor area** +「Load default template」「Restore default (clear)」+ placeholder docs.
|
||||
* **Client-side preview**: replace with sample `status=502`, `host=example.com` and preview in sandbox/iframe.
|
||||
* Save: `OptionService.updateBatch`; permissions same as the performance tuning page (admin).
|
||||
|
||||
### 5.3 Component Dependencies
|
||||
|
||||
Tags Input and the HTML editor reuse existing shadcn/ui components, consistent with the existing UI style.
|
||||
|
||||
---
|
||||
|
||||
## 6. Data Flow
|
||||
|
||||
```text
|
||||
Admin /responses (Error Page tab)
|
||||
→ Option update-batch (validate tags & HTML)
|
||||
→ w_system_configs
|
||||
|
||||
Release config version
|
||||
→ snapshot writes OriginErrorPage*
|
||||
→ render OpenResty conf + SupportFile
|
||||
→ Agent pulls and reloads
|
||||
|
||||
Visitor requests a proxied domain
|
||||
→ origin/gateway produces a matching status code
|
||||
→ error_page → named location
|
||||
→ replace placeholders, keep original status, return HTML
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Decision Record
|
||||
|
||||
| Decision | Choice | Reason |
|
||||
| --- | --- | --- |
|
||||
| Config scope | global | product requirement; simple implementation and ops |
|
||||
| Storage | Option + config version | consistent with performance tuning, rollbackable |
|
||||
| Status input | tags: single code and range | default whole 5xx, but can name 522 |
|
||||
| Response status | keep original | correct for monitoring/SEO/client semantics |
|
||||
| Runtime replacement | internal + lightweight template replacement | status differs per request |
|
||||
| Customization | online HTML | flexible without a file-upload chain |
|
||||
@@ -0,0 +1,252 @@
|
||||
# Pages Static Hosting Design
|
||||
|
||||
You will learn: the architecture design of OpenFlare Pages static hosting, the immutable deployment and secure extraction flow, the OpenResty static serving and API reverse proxy config rendering, and the cooperative workflow between the control plane and the Agent.
|
||||
|
||||
---
|
||||
|
||||
## Requirements Analysis
|
||||
|
||||
In modern web operations, besides reverse proxying dynamic applications, deploying and hosting static frontend sites (SPA apps built with React/Vue, or static generator output like Hugo/VitePress) is extremely common. Traditional approaches suffer from:
|
||||
1. **Release disconnected from proxy config**: after uploading frontend build artifacts to the Nginx host, you still need to modify the Nginx vhost config manually or via other scripts — error-prone and without version control.
|
||||
2. **Multi-node distribution is hard**: with multiple edge nodes managed by the control plane, syncing static files to all nodes consistently requires complex sync scripts (e.g. rsync).
|
||||
3. **Rollback lacks consistency**: once a new frontend package fails or has serious defects, you must restore both the static files and the proxy rules — atomic rollback is hard.
|
||||
|
||||
To solve these, OpenFlare introduces **Pages static hosting**, inspired by Cloudflare Pages. It brings "pre-built artifact import" and "website proxy rule config" into the same control plane, leveraging OpenFlare's pull-based cooperative architecture to achieve eventual convergence across multiple Agents via immutable deployments, single-node atomic switching, and periodic reconciliation, with fast rollback.
|
||||
|
||||
---
|
||||
|
||||
## Core Features
|
||||
|
||||
The Pages static hosting subsystem includes:
|
||||
* **Pre-built artifact deployment**: upload a static resource archive directly, or save a Remote URL or public GitHub Release asset source for a project. External sources are only accessed by the Server; on successful sync they uniformly create or reuse an immutable deployment and activate it atomically.
|
||||
* **Immutable deployment snapshots**: each local upload creates a new candidate deployment; persistent-source sync creates or reuses a deployment by source identity/revision and activates it. All deployments have a unique ID and whole-package SHA-256, support keeping the most recent N historical versions per system config, and can be rolled back anytime.
|
||||
* **Check and auto-update**: GitHub latest can be checked periodically per project; by default it only hints at available updates; only after an admin explicitly enables it does it auto-sync and publish by the exact revision found.
|
||||
* **SPA Fallback**: supports fallback routing for single-page apps; when a static file isn't found, requests redirect to the entry file.
|
||||
* **Built-in API reverse proxy**: enables API proxying within Pages rules with one click, eliminating cross-origin issues by forwarding requests to a designated backend.
|
||||
* **Secure package validation and extraction**: built-in path-traversal defense, symlink-hijack protection, file size/count limits, and configurable upload package size control to keep nodes physically safe.
|
||||
* **Configurable limits**: admins can adjust the "deployment package size limit" and "historical deployment retention" in ops settings.
|
||||
|
||||
### Deployment Sources
|
||||
|
||||
Projects currently support three source views: manual, Remote URL, and GitHub Release. No source record means manual; switching or deleting a source doesn't delete historical deployments or change the current active deployment. Remote URL only allows manual "Sync and Publish"; GitHub Release supports manual check/sync for latest/tag, and only latest can opt into scheduled checking and auto-update.
|
||||
|
||||
A source is mutable config; a deployment is an immutable fact. Source config and runtime cursor/state/lease are stored separately; deployments only save the security provenance snapshot at creation time. All artifacts reuse the "download or receive artifact → real-byte and entry validation → `upload.Ingest` → deployment" artifact pipeline: manual uploads stop at candidate, waiting for explicit admin activation; persistent-source sync creates-or-loads and atomically activates in the same business transaction.
|
||||
|
||||
The admin project detail is organized as "current production deployment → deployment source → deployment history".
|
||||
|
||||
---
|
||||
|
||||
## Pages Static Hosting Architecture
|
||||
|
||||
Pages hosting is logically split into a **Control Plane** and a **Data Plane**.
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
%% data flow
|
||||
Browser[1. Browser / Visitor] -->|HTTPS request / traffic| OpenResty[2. OpenResty / WAF]
|
||||
OpenResty -->|1. static serving try_files| StaticFiles[3. Edge node local static dir current]
|
||||
OpenResty -->|2. forward API proxy| BackEnd[4. Backend API service]
|
||||
|
||||
%% control flow & heartbeat
|
||||
Admin[Admin / CI] -->|upload or configure source| Server[OpenFlare Server control plane]
|
||||
Providers[Remote / GitHub Provider] -->|restricted artifact candidate| Server
|
||||
Scanner[internal scanner / action task] -->|check & auto sync| Server
|
||||
Server <-->|Agent API / Heartbeat| Agent[openflare-agent process]
|
||||
Server -.->|unified upload.Ingest| UploadStore[(platform upload backend)]
|
||||
|
||||
Agent -->|1. discover new version| Server
|
||||
Agent -->|2. download deployment package| Server
|
||||
Agent -->|3. validate, extract, atomic switch| StaticFiles
|
||||
```
|
||||
|
||||
* **Control Plane**: the Server receives local uploads or fetches Remote/GitHub pre-built artifacts via restricted providers; action tasks and the internal scanner handle checking, syncing, and auto-updates. All artifacts pass unified inspection and `upload.Ingest` into the platform storage backend; manual uploads create a new candidate, persistent-source sync creates-or-loads a deployment and activates it atomically. Config release only compiles the stable project anchor and static serving metadata.
|
||||
* **Data Plane**: the Agent discovers Pages projects referenced by config during heartbeat/WS reconciliation, pulls the project's currently active package via a dedicated API, and performs validation and extraction. OpenResty serves static files locally; the Agent doesn't know whether the artifact came from upload, Remote, or GitHub.
|
||||
|
||||
---
|
||||
|
||||
## Data Model and Metadata Design
|
||||
|
||||
### 1. Core DB Entities
|
||||
* **Pages project (`of_pages_projects`)**:
|
||||
* Records the business name, Slug (URL-friendly), enabled state, static serving root dir (RootDir, nullable), entry filename (EntryFile, default `index.html`), SPA Fallback settings, and API reverse proxy config (APIProxyPath, APIProxyPass, APIProxyRewrite).
|
||||
* **Source config (`of_pages_project_sources`)**:
|
||||
* At most one mutable source config per project, distinguishing Remote URL and GitHub Release by `source_type`. `config_version` fences stale tasks; the full Remote URL lives only in the config table — never in responses, logs, task payloads, or deployment provenance. V2 does not promise DB column encryption.
|
||||
* **Source runtime (`of_pages_project_source_runtime`)**:
|
||||
* 1:1 with the source, storing ETag, seen/applied revision, last check/sync, next check, errors, and lease. State is fixed to `idle | checking | update_available | syncing | failed | attention`; queued/completed state is carried by `TaskExecution`.
|
||||
* **Pages deployment (`of_pages_deployments`)**:
|
||||
* Records immutable deployment facts: project-incremental deployment number, whole-package SHA-256, `upload_id`, file count/total bytes, creator, and nullable source identity/revision, source security snapshot, and trigger. `artifact_path` is only a legacy-compat field and is no longer the storage truth for new deployments.
|
||||
* **Deployment file manifest (`of_pages_deployment_files`)**:
|
||||
* Stores the full regular-file path list and actual byte counts per deployment for console display and statistics.
|
||||
* No longer computes content hashes per file; integrity is guaranteed by the **whole-package** SHA-256 (`of_pages_deployments.checksum`), verified by the Agent when pulling.
|
||||
* Control-plane inspection reads the archive via file handles, streaming each regular file body and checking declared vs actual size — no whole-package `ReadFile` into memory, no per-file disk write for hashing.
|
||||
|
||||
### 2. Route Association and Snapshot
|
||||
`proxy_routes` rules associate with a Pages project via `upstream_type = "pages"` and `pages_project_id`. A route may join the release flow only when its type is `pages` and the project has an activated deployment.
|
||||
The version snapshot emitted at release includes `snapshotPagesDeployment`:
|
||||
```json
|
||||
{
|
||||
"project_id": 1,
|
||||
"project_slug": "my-spa-app",
|
||||
"deployment_id": 12,
|
||||
"deployment_number": 3,
|
||||
"checksum": "a7b3c2...",
|
||||
"entry_file": "index.html",
|
||||
"spa_fallback_enabled": true,
|
||||
"spa_fallback_path": "/index.html",
|
||||
"api_proxy_enabled": true,
|
||||
"api_proxy_path": "/api",
|
||||
"api_proxy_pass": "http://api.internal:8000",
|
||||
"api_proxy_rewrite": "/api/(.*) /$1",
|
||||
"local_root": "__OPENFLARE_PAGES_DIR__/projects/1/current"
|
||||
}
|
||||
```
|
||||
|
||||
### 3. Dual-Track Relationship with the Main Config Version (project anchor + latest pull)
|
||||
* The **main config version** and **Pages deployments** are two independent version systems.
|
||||
* The stable anchor of a Pages route in the main config is **`pages_project_id` (project ID)**, not a deployment ID.
|
||||
* The OpenResty `root` uses the project-level path `__OPENFLARE_PAGES_DIR__/projects/{project_id}/current`; the path stays unchanged on activation switch, so swapping packages never requires republishing the main config.
|
||||
* The Agent requests the "latest active package" per project (like `github/release/latest`):
|
||||
* `GET /api/v1/agent/pages/projects/:project_id/latest/hash`
|
||||
* `GET /api/v1/agent/pages/projects/:project_id/latest/package`
|
||||
* The control plane returns the deployment ID, hash, package size, and expanded manifest metadata for the project's **currently active deployment**. The Agent uses the deployment ID and other latest metadata to detect pointer races during download, but the stable anchor of the main config and local dir remains the project ID.
|
||||
* Therefore: switching the active deployment within a project **does not require publishing the main config**; the Agent polls the latest hash during periodic reconciliation, downloads on change, and switches `current`.
|
||||
* The `pages_deployment` field in the snapshot still records release-time metadata (entry file, SPA/API proxy, etc.) but does not lock the Agent's package version.
|
||||
|
||||
---
|
||||
|
||||
## Server (Control Plane) Responsibilities and Lifecycle
|
||||
|
||||
### 1. Deployment Package Security Validation and Analysis
|
||||
To protect the server from untrusted artifacts, the control plane applies the same strict validation to local uploads and all external sources:
|
||||
* **Format support**: `zip`, `tar.gz` / `tgz`, `tar.xz` / `txz`, `tar.bz2` / `tbz2`, `tar`, `7z`.
|
||||
* **Size limits**: archive size is controlled by system config `pages_max_package_size_mb` (default 100 MiB, range 1–2048); expanded single-file and total limits are "package size × 4" with a floor of 100 MiB. Inspection always streams regular file bodies, checking declared vs actual size and enforcing limits on actual values.
|
||||
* **Count limit**: at most 1,000 static files per package.
|
||||
* **Symlink blocking**: any symlink detected while walking the archive immediately errors and rejects the upload, defending against symlink-hijack attacks.
|
||||
* **Path traversal defense**: every archive file path is `Clean`ed and checked for `..` or leading `/`, defending against directory-traversal writes to sensitive system paths.
|
||||
* **Entry file validation**: the project's entry file (e.g. `index.html`, possibly under `project.RootDir`) must exist in the package, otherwise the upload is rejected.
|
||||
* **Common root prefix stripping**: many packaging tools add a redundant top folder as a common root prefix; the control plane auto-detects and safely strips it.
|
||||
* **Whole-package integrity**: SHA-256 is computed once over the archive bytes at upload/import and written to the deployment record; the Agent reconciles against the whole-package hash after pulling. No per-file content hashes.
|
||||
* **Actual size recheck**: `InspectOptions.VerifySizes` is kept only for compatibility; current inspection always reads regular file bodies, checks declared values, and accumulates actual sizes, but still does not compute per-file content hashes.
|
||||
* **History retention**: system config `pages_max_history_count` (default 20; 0 = unlimited) trims after successful deployment. Semantics: **each project keeps at most N deployments**; the currently active deployment is always kept, remaining slots fill from newest to oldest by deployment ID. With `history_count=1`, manual uploads temporarily keep both the active and the newest candidate; the next upload replaces the old candidate; after the candidate activates, the strict limit resumes. Exceeding non-active deployments and their file manifests are deleted; the corresponding upload record is soft-deleted idempotently via platform primitives — Pages never physically deletes blobs that may be shared by dedup. If trimming fails after a successful deployment, it only logs and doesn't roll back activation; concurrent operations may temporarily exceed N and converge on later trims. Main config version rollback does not depend on old Pages packages (see the dual-track section above).
|
||||
|
||||
### 2. Deployment Package Storage Planning
|
||||
The control plane stores local, Remote, and GitHub artifacts into the configured local/S3 backend via the unified upload framework (`upload.Ingest`), recording `upload_id` and the file manifest in the DB. **Large static packages never enter `config_versions` records or any config push channel**, keeping control-plane data sync lightweight.
|
||||
|
||||
### 3. Source Check, Auto-Update, and Upload Compensation
|
||||
|
||||
* `openflare:pages_source_action` executes admin check/sync or scanner-dispatched exact-revision sync; payloads never carry URL, Token, ETag, or lease tokens. Manual sync only accepts real user actors; auto sync only accepts the system actor with the `scheduled_auto_update` trigger.
|
||||
* `openflare:pages_source_scan` is a fixed `*/5 * * * *` internal-only TaskHandler accepting only `{}`; it never appears in generic task types or the schedule management UI. Each round runs "recover expired leases → compensate orphan uploads → scan due sources".
|
||||
* The scanner sorts stably by `next_check_at, source_id`, serially checking at most 20 GitHub latest sources per batch; ETag/304 still advances the check time; 403/429 record the status code and the actual backoff deadline; a single source failure doesn't block subsequent sources.
|
||||
* On finding an update, the seen cursor is always saved first. Only with `auto_update_enabled=true` and a normal `update_available` state does it dispatch sync with the exact revision found this check; `attention`, Remote, and fixed tags never auto-publish. Manually activating another deployment fences in-flight tasks and disables auto.
|
||||
* Orphan compensation checks at most 100 upload records per round that have been quarantined for at least 2 hours, requiring a system owner, Pages retention type, V2 marker, and no deployment references. Candidates are re-checked in the `project → source → runtime → upload` lock order and only soft-deleted via the upload framework with stat updates — never physically deleting blobs possibly shared by dedup.
|
||||
|
||||
---
|
||||
|
||||
## Agent (Data Landing) Responsibilities and Self-Healing
|
||||
|
||||
The Agent runs on each edge proxy node: on first applying config referencing a Pages project, and on subsequent periodic latest reconciliation, it "atomically" pulls the currently active static assets to the node.
|
||||
|
||||
### 1. Pull Latest Per Project
|
||||
1. The Agent parses routes with `UpstreamType == "pages"` from the active main config and collects the stable anchor **`pages_project_id`**.
|
||||
2. For each project it calls `GET /api/v1/agent/pages/projects/:project_id/latest/hash` to get the control plane's currently active package hash (a latest pointer).
|
||||
3. If the local `projects/{project_id}/releases/{hash}` isn't ready, it streams `.../latest/package` to a temp file, enforcing real response limits and SHA-256; after download it **requests the hash again** to avoid activation-switch races, retrying a bounded number of times on mismatch.
|
||||
4. The request carries the node's `X-Agent-Token`.
|
||||
|
||||
### 2. Secure Extraction, Atomic Switch, Keep Only Latest
|
||||
1. Absolute package cap is 2 GiB; the downloaded content's SHA-256 must match the latest hash from the post-download re-query; the whole package never enters `[]byte`.
|
||||
2. Extract into a random staging dir `projects/{project_id}/releases/.{hash}-<random>.tmp` (zip / tar.* / 7z supported), rejecting path traversal, links, and special files. The Agent obeys both Server metadata limits and local absolute limits: at most 1,000 files, single file and total at most 8 GiB.
|
||||
3. After extraction, walk the actual file tree and precisely recheck file count and total bytes against the Server metadata; mismatch → refuse to switch.
|
||||
4. Write `.openflare-pages.json`, then rename to `releases/{hash}`.
|
||||
5. **Atomic switch** `projects/{project_id}/current` to the new release (symlink preferred, copy on failure).
|
||||
6. **Only after the new package is ready and current has switched successfully**, delete other `releases/*` (including `.tmp`) under the project — **historical deployment packages are never kept on the edge**. Each project always keeps exactly one latest content per node.
|
||||
7. Multi-project reconciliation **isolates failures**: a single project failure logs and continues with others, finally aggregating errors.
|
||||
|
||||
---
|
||||
|
||||
## OpenResty (Static Serving and Proxy) Config Rendering
|
||||
|
||||
For Pages-hosted sites, the control plane renders the corresponding `server` block, replacing the regular proxy route's `proxy_pass`.
|
||||
|
||||
### 1. Static Serving Directive Rendering
|
||||
* **`root` and `index`**:
|
||||
The Server points `root` at the project-level placeholder path `__OPENFLARE_PAGES_DIR__/projects/{project_id}/current` (optionally appending `RootDir`). Activation switching only changes directory contents, not the path, so swapping packages never requires republishing the main config.
|
||||
```nginx
|
||||
server {
|
||||
listen 80;
|
||||
server_name myapp.example.com;
|
||||
|
||||
root "/var/lib/openflare/pages/projects/3/current";
|
||||
index "index.html";
|
||||
...
|
||||
}
|
||||
```
|
||||
|
||||
### 2. try_files and SPA Fallback
|
||||
* **SPA Fallback disabled (default)**:
|
||||
only match physically existing files, otherwise strict 404:
|
||||
```nginx
|
||||
location / {
|
||||
try_files $uri $uri/ =404;
|
||||
}
|
||||
```
|
||||
* **SPA Fallback enabled**:
|
||||
if the requested file doesn't exist, redirect to the project's configured entry fallback (usually `/index.html`):
|
||||
```nginx
|
||||
location / {
|
||||
try_files $uri $uri/ /index.html;
|
||||
}
|
||||
```
|
||||
|
||||
### 3. API Reverse Proxy and Rewrite Rendering
|
||||
When a static frontend needs backend API access without cross-origin issues, enable the API proxy. The OpenResty renderer nests a dedicated API `location` branch inside the static `server` block:
|
||||
```nginx
|
||||
server {
|
||||
listen 80;
|
||||
server_name myapp.example.com;
|
||||
...
|
||||
# API proxy path match
|
||||
location /api {
|
||||
# apply rewrite rules when configured
|
||||
rewrite ^/api/(.*)$ /v1/$1 break;
|
||||
rewrite ^/api$ / break;
|
||||
|
||||
proxy_pass http://api.internal:8000;
|
||||
proxy_http_version 1.1;
|
||||
proxy_set_header Host $http_host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
proxy_set_header X-Forwarded-Proto $scheme;
|
||||
proxy_set_header Upgrade $http_upgrade;
|
||||
proxy_set_header Connection $connection_upgrade;
|
||||
}
|
||||
|
||||
location / {
|
||||
try_files $uri $uri/ /index.html;
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Interaction Logic and Sync Flow
|
||||
|
||||
A full pre-built artifact import and activation lifecycle looks like this. Binding a project for the first time requires publishing the main config; subsequent active deployment changes converge independently via the project latest:
|
||||
|
||||
```text
|
||||
[Admin / scanner] [Server control plane] [Agent] [OpenResty]
|
||||
| | | |
|
||||
|-- manual upload ----->|-- inspect / Ingest ---->| |
|
||||
| |-- create candidate | |
|
||||
|-- explicit activate -->|-- switch active | |
|
||||
| | | |
|
||||
|-- source sync ------->|-- inspect / Ingest | |
|
||||
| |-- create/load + atomic activation | |
|
||||
| | | |
|
||||
|-- first bind & publish|-- broadcast project anchor -->|-- write/reload route -->|
|
||||
| | | |
|
||||
|-- later activate/sync/rollback -->|-- active latest changed -->| |
|
||||
| |<-- latest metadata reconciliation ----| |
|
||||
| |--- stream package -------------------->| |
|
||||
| | |-- validate, extract, recheck --|
|
||||
| | |-- atomic switch current ------>|
|
||||
```
|
||||
@@ -0,0 +1,126 @@
|
||||
# Tunnel & Intranet Penetration Design
|
||||
|
||||
You will learn: the architecture design of OpenFlare's intranet penetration tunnels, the internal principles of the dual-end control components (Relay and Client), interaction logic, and the data-plane / control-plane communication flows.
|
||||
|
||||
---
|
||||
|
||||
## Requirements Analysis
|
||||
|
||||
In typical web-hosting scenarios, many origins are deployed in intranet environments (local dev machines, LAN servers, or firewall-restricted intranet clusters). These servers usually:
|
||||
1. **Have no public IP**: cannot be directly reached by public traffic.
|
||||
2. **Compliance restrictions**: port mapping (NAT) on border routers is not freely allowed.
|
||||
3. **Dynamic IP changes**: traditional DDNS is high-latency and unstable.
|
||||
|
||||
To let intranet origins seamlessly join the OpenFlare global data gateway and enjoy value-added services like WAF geo protection and TLS certificate management, OpenFlare designs a **reverse-relay tunnel penetration** solution. In this architecture, public edge nodes act as the reverse-proxy entry and traffic relay; the intranet side only needs outbound secure connections to safely and stably reverse-penetrate public traffic to intranet origins.
|
||||
|
||||
---
|
||||
|
||||
## Core Features
|
||||
|
||||
The intranet penetration subsystem includes:
|
||||
|
||||
* **Dynamic Relay node management**: the control plane dynamically dispatches the relay service (frps), dynamically distributing service ports and auth tokens.
|
||||
* **Multi-tunnel reverse proxy mapping**: map multiple intranet web ports on a single intranet client, binding multi-domain routes to corresponding relay nodes.
|
||||
* **Independent process lifecycle management**: both relay and client are standalone Go binaries that spawn, monitor, self-heal, and hot-upgrade the underlying frp engines.
|
||||
* **Token-based auth isolation**: the relay uses `agent_token`; the intranet client uses its dedicated `tunnel_token` — permissions and route boundaries isolated.
|
||||
* **Config validation and incremental hot reload**: config files are rewritten and processes reloaded only when tunnel bindings, certificates, or Relay topology actually change, reducing runtime overhead.
|
||||
|
||||
---
|
||||
|
||||
## Tunnel Architecture
|
||||
|
||||
The subsystem integrates the mature `frp` high-performance tunnel protocol, split into a **Control Plane** and a **Data Plane**.
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
%% data flow
|
||||
Browser[1. Browser / Visitor] -->|HTTPS request| Agent[2. OpenResty / Agent]
|
||||
Agent -->|local forward proxy_pass| RelayFrps[3. OpenFlare Relay / frps]
|
||||
RelayFrps -->|encrypted tunnel protocol| FlaredFrpc[4. OpenFlared / frpc]
|
||||
FlaredFrpc -->|forward local request| LocalOrigin[5. Intranet origin 192.168.x.x]
|
||||
|
||||
%% control flow & heartbeat
|
||||
Server[OpenFlare Server control plane] <-->|Relay API / Heartbeat| RelayManager[openflare-relay process]
|
||||
Server <-->|Client API / Heartbeat| ClientManager[openflared process]
|
||||
|
||||
RelayManager -.->|manage process & config| RelayFrps
|
||||
ClientManager -.->|manage multiple Relay processes| FlaredFrpc
|
||||
```
|
||||
|
||||
* **Control Plane**: the Server maintains DB state; `openflare-relay` on relay nodes and `openflared` on intranet servers sync tunnel config via HTTP heartbeats and WebSocket long channels.
|
||||
* **Data Plane**: public traffic first enters the public-edge Agent (OpenResty), where HTTPS handshake, TLS termination, and WAF filtering happen; then `proxy_pass` forwards to the same-host `openflare-relay (frps)`. `frps` encapsulates the request and sends it through the persistent tunnel established with the intranet `openflared (frpc)`, which finally unpacks and dispatches to the actual intranet origin.
|
||||
|
||||
---
|
||||
|
||||
## Relay Design
|
||||
|
||||
`openflare-relay` is the relay manager deployed at the public edge, running on `tunnel_relay`-type nodes.
|
||||
|
||||
### 1. Core Architecture & Logic
|
||||
* **Process guard**: the Relay process holds the `frps` binary, spawns `frps -c frps.toml` via `exec.Command`, and starts a goroutine asynchronously watching its exit state. If `frps` exits abnormally, it auto-restarts with backoff.
|
||||
* **Dynamic config rendering**: syncs state to the control plane via HTTP heartbeat and fetches the current `RelayConfig`, mainly:
|
||||
* `bindPort`: frps's public control port listening for intranet frpc client connections.
|
||||
* `vhostHTTPPort`: vhost HTTP traffic port; the Agent's proxy_pass points here.
|
||||
* `authToken`: security credential for client handshake validation.
|
||||
* `webServer`: enables the frps dashboard API; the Relay collects real-time active tunnel counts and traffic metrics from this or the admin control port.
|
||||
* **State reporting**: each heartbeat reports the underlying `frps` active connections, registered client count, per-proxy real-time state, and Relay version.
|
||||
|
||||
---
|
||||
|
||||
## Openflared (Client) Design
|
||||
|
||||
`openflared` is the client manager on the user's intranet server, authenticated with its dedicated `tunnel_token`.
|
||||
|
||||
### 1. Core Mechanisms
|
||||
* **Multiple Relay support (multiplexing)**:
|
||||
for HA or nearest access, the control plane may schedule a client across multiple public Relays. `openflared` reads the Relays list in `TunnelConfig`, generates a dedicated config per Relay locally (named `frpc_<relay_node_id>.toml`), and assigns each Relay process an independent cancelable context.
|
||||
* **Independent child-process monitoring**:
|
||||
`openflared` maintains a `processes` map for per-`frpc` lifecycle management. When the control plane adds or removes a Relay, the client incrementally spawns new processes or gracefully shuts down old ones without affecting other working tunnels.
|
||||
* **Dynamic TOML generation**:
|
||||
when rendering the TOML for each Relay, the client iterates the Proxies list and writes each intranet service's `LocalAddr`, `LocalPort`, and bound `CustomDomains` into `[[proxies]]` blocks.
|
||||
|
||||
---
|
||||
|
||||
## Interaction Logic and Traffic Model
|
||||
|
||||
The subsystem implements consistent versioning and state feedback.
|
||||
|
||||
### 1. Control-Plane Release and Sync Flow
|
||||
|
||||
```text
|
||||
Admin modifies tunnel/intranet port mapping -> submit release -> generate new Tunnel version and Checksum
|
||||
|
|
||||
v (push or heartbeat pull)
|
||||
+-------------------------------------------+-------------------------------------------+
|
||||
| |
|
||||
v (relay side) v (intranet client)
|
||||
openflare-relay heartbeat detects frps port/Token changes openflared heartbeat detects tunnel_version change
|
||||
re-render local frps.toml request latest proxy mapping package
|
||||
kill and restart the frps process re-render frpc_<relay_id>.toml
|
||||
report health state healthy restart changed Relay processes with hot reload
|
||||
report apply result (Apply Success/Error)
|
||||
```
|
||||
|
||||
1. **Versioned control**: tunnel routes and mappings are versioned like the main routing system, dispatching `version` and `checksum` so clients don't rewrite or reload processes redundantly.
|
||||
2. **Apply-result loop**: after applying new config, the client reports the result in its heartbeat. If frpc can't connect (intranet port unreachable or wrong cert config), the client captures process output and reports `LastError`, letting admins see penetration failure reasons directly in the Server.
|
||||
|
||||
### 2. Data-Plane Traffic Model
|
||||
1. **Public entry (Agent)**:
|
||||
```nginx
|
||||
server {
|
||||
listen 443 ssl;
|
||||
server_name intranet.example.com;
|
||||
# ... TLS cert & WAF filtering logic ...
|
||||
location / {
|
||||
proxy_pass http://127.0.0.1:8080; # points to the local frps vhost port
|
||||
proxy_set_header Host $host; # must keep the original Host; frps routes by Host
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
}
|
||||
}
|
||||
```
|
||||
2. **Relay node (frps)**:
|
||||
`frps` receives the HTTP request on the vhost port (default `8080`), reads the `Host: intranet.example.com` header, and looks up the registered active-tunnel table for the matching encrypted TCP connection (established by the intranet frpc).
|
||||
3. **Encrypted tunnel transport (TCP)**:
|
||||
`frps` encapsulates the HTTP request into the internal TCP tunnel protocol and sends it to the intranet `frpc` client.
|
||||
4. **Intranet client dispatch (frpc)**:
|
||||
the `frpc` managed by `openflared` receives the packet, opens a local TCP connection per local config (`localIP = "127.0.0.1"`, `localPort = 8080`), forwards to the intranet web service, and returns the response along the same path to the public user.
|
||||
@@ -0,0 +1,15 @@
|
||||
# WAF Design
|
||||
|
||||
OpenFlare's current WAF rule model is a visual DAG. Node semantics, graph constraints, multi-rule ordering, release compilation, and migration boundaries are all governed by [WAF Orchestration Rule Design](./waf-orchestration-design.md).
|
||||
|
||||
## System Boundaries
|
||||
|
||||
The Server stores the edit graph with coordinates and a revision number; at release it re-validates and compiles it into a compact runtime graph; the Agent atomically writes the snapshot and reloads OpenResty; the request hot path only traverses the immutable in-memory graph in the Worker.
|
||||
|
||||
IP groups update independently of rule topology. Manual, subscription, and auto IP groups are maintained by the control plane; the Agent atomically replaces the JSON first and updates the checksum last. A coordinating worker checks the checksum every 5 seconds, reading and distributing the full snapshot only on change; on failure it keeps the previous valid data. The full runtime snapshot is capped at 20 MiB; Server release/sync and Agent disk writes use the same serialization validation; OpenResty uses a separate 64 MiB shared dict with non-evicting writes, refusing new versions on capacity shortage without breaking committed snapshots.
|
||||
|
||||
Geo nodes use Country and City MMDB. Docker images bundle the database files; bare-binary installs have the Agent download missing files at first startup and update them periodically per config; request handling always reads the DB already loaded by OpenResty. When the DB is unavailable, geo match returns `false` with a rate-limited warning; other execution errors must not be accidentally allowed through due to data corruption.
|
||||
|
||||
## Security Ordering
|
||||
|
||||
Enabled global rules always run first; route rules execute by binding sequence. A block node terminates immediately; a pass node only ends the current rule; only after all rules pass does traffic enter the origin chain. Unknown nodes, missing outlets, or step-limit overruns always block the request.
|
||||
@@ -0,0 +1,118 @@
|
||||
# WAF Orchestration Rule Design
|
||||
|
||||
This document defines the target architecture, data model, execution semantics, release model, and migration boundaries for reworking OpenFlare WAF from a fixed decision chain into a visual directed acyclic graph (DAG). IP group sources and membership computation still follow [WAF Design](./waf-design.md); this document only changes how rules are composed and executed.
|
||||
|
||||
## Goals and Boundaries
|
||||
|
||||
When a user adds a WAF rule, they only enter a name. The Server immediately creates a legal default graph `start → pass`, and the frontend enters a standalone orchestration page based on React Flow. Users build policies by adding processing units, configuring nodes, and connecting branches — no more filling in fixed-order allow/block lists and PoW forms.
|
||||
|
||||
Supported nodes:
|
||||
|
||||
| Node | Count Constraint | Inputs | Outputs | Config |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Start | exactly one per graph | none | `next` | none |
|
||||
| Pass | exactly one per graph | one or more | none | none |
|
||||
| Block | multiple allowed | one or more | none | HTTP status code, HTML response body |
|
||||
| IP match | multiple allowed | one or more | `true`, `false` | IP, CIDR, IP group ID |
|
||||
| Geo match | multiple allowed | one or more | `true`, `false` | country code, region code |
|
||||
| UA check | multiple allowed | one or more | `true`, `false` | UA required, browser/OS allowlist with and/or, block crawlers/abnormal UAs (excluding crawlers)/custom regex |
|
||||
| Security | multiple allowed | one or more | `true`, `false` | basic signature detection (path traversal/file inclusion on by default; SQL/XSS/command injection/SSRF/upload/XXE/CRLF toggleable); any enabled rule hit → false |
|
||||
| PoW | multiple allowed | one or more | `next` | algorithm, difficulty, session TTL, challenge TTL |
|
||||
|
||||
IP match, geo match, UA check, and security do not distinguish allowlist vs blocklist. `true` only means the request passed that node's judgment; `false` only means it did not; the business meaning of allow vs block is entirely determined by the wiring. UA check evaluation order: require UA → block crawlers/abnormal UAs → allowlist match. Security performs signature matching on the request Path/Query/Header/Cookie/Body (bounded). After PoW verification succeeds, execution continues along `next`; when incomplete, the challenge page takes over the request and no `false` branch is produced.
|
||||
|
||||
Loops, script nodes, arbitrary expression nodes, subgraph calls, and cross-rule jumps are not implemented.
|
||||
|
||||
## Control-Plane Architecture
|
||||
|
||||
The rule graph uses a dual model separating control-plane edit state and data-plane runtime state:
|
||||
|
||||
1. The React Flow editor submits a versioned graph JSON containing node IDs, node types, display names, coordinates, typed configs, and edges.
|
||||
2. The Server performs authoritative validation of the whole graph; on success it saves the graph in a single transaction and increments the revision number.
|
||||
3. On config release, the Server validates all enabled rules again, compiles the graph into a compact runtime DAG without UI fields (coordinates, labels, etc.), and collects the referenced IP group IDs.
|
||||
4. The Agent atomically writes the full release snapshot and reloads OpenResty. New Workers only load and parse the rule JSON once at startup.
|
||||
5. The request hot path only traverses the immutable in-memory runtime graph in the Worker — no file reads, no checksum computation, no JSON parsing.
|
||||
|
||||
The edit-state JSON uses an explicit `schema_version`. Node configs use per-node-type structures; unconstrained key-value objects bypassing Server validation are not allowed. Initial safety limits: 128 nodes, 256 edges, and 256 KiB of edit-state JSON per rule; these limits are enforced by both the API and the release compiler.
|
||||
|
||||
## Graph Structure Constraints
|
||||
|
||||
Rule save and release must satisfy all constraints:
|
||||
|
||||
* The graph is a DAG; self-loops and arbitrary cycles are forbidden.
|
||||
* Exactly one start node and one pass node exist; multiple block nodes are allowed.
|
||||
* The start node has no incoming edges and exactly one `next` outlet; pass and block nodes have no outlets.
|
||||
* IP match, geo match, UA check, and security must each connect their `true` and `false` outlets once; PoW's `next` must connect once.
|
||||
* No dangling outlets except terminal nodes; every non-start node has at least one incoming edge.
|
||||
* All nodes must be reachable from start, and every executable node must be able to reach pass or block.
|
||||
* Edge source ports must belong to the source node type; the same source port must not connect to multiple targets.
|
||||
* Node IDs are unique within the graph, edge IDs are unique within the graph, and all referenced nodes must exist.
|
||||
* Node configs must pass per-type field, range, reference-existence, and size validation.
|
||||
|
||||
The frontend provides instant validation and connection restrictions to improve UX, but the Server is the only authoritative validator. When deleting a node, the frontend synchronously removes related edges and marks the rule unsaved; saving is forbidden until the graph is legal again.
|
||||
|
||||
## Multi-Rule Execution Semantics
|
||||
|
||||
A route can bind multiple custom rules. The binding is an ordered list following this order:
|
||||
|
||||
1. Enabled global rules always execute first and do not participate in route-side ordering.
|
||||
2. Enabled rules bound to the route execute in binding order.
|
||||
3. When the current rule reaches a block node, it immediately outputs that node's configured response and terminates the request.
|
||||
4. When the current rule reaches a pass node, that only means the current rule finished; if more rules remain, execution continues.
|
||||
5. Only after all rules reach a pass node is the request truly allowed to proceed into the OpenResty/origin chain.
|
||||
|
||||
The runtime graph is fully validated before release. If the Lua executor still encounters an unknown node, unknown port, missing target, or exceeds the node-step limit, it logs a rate-limited error and blocks the request, preventing a broken security config from accidentally allowing traffic.
|
||||
|
||||
## IP Group Memory Refresh
|
||||
|
||||
Rule topology only takes effect on release + OpenResty reload; IP group membership can still be updated independently by manual, subscription, or auto tasks without release or reload.
|
||||
|
||||
IP groups use a two-level cache of coordinating worker, shared snapshot, and worker-local objects:
|
||||
|
||||
1. Requests always read the IP group object in the current worker's memory — no file access or shared-dict JSON parsing.
|
||||
2. Every 5 seconds only one worker holding a shared lock reads the lightweight checksum file.
|
||||
3. If the checksum is unchanged, it ends immediately without reading the full `waf_ip_groups.json`.
|
||||
4. On checksum change, the coordinating worker reads and validates the full JSON once, writes the raw snapshot to a dedicated 64 MiB `ngx.shared.openflare_waf_ip_groups` keyed by checksum, then updates the commit pointer.
|
||||
5. Other workers detecting a shared version change fetch the snapshot from shared memory, parse it, and atomically replace their local object — no repeated disk reads.
|
||||
6. On refresh failure, keep using the previous valid object, log a rate-limited error, and retry next cycle.
|
||||
|
||||
The Agent must atomically replace the IP group JSON first, then atomically update the checksum last, so workers never recognize a half-written file as a new version. Server release/sync and Agent disk writes jointly enforce the 20 MiB aggregate snapshot limit; shared dict uses a safe write that never force-evicts old keys, keeping the current and previous immutable snapshots on failure.
|
||||
|
||||
## API and Editor
|
||||
|
||||
The create API only accepts a rule name and returns the rule detail with the default graph. Rule metadata, graph save, and route binding use separate operations, so toggling enabled state or binding does not overwrite the canvas.
|
||||
|
||||
The graph detail includes `revision`. Save requests submit `revision + graph`; the Server only updates and increments the revision when it matches. On mismatch it returns a conflict, the frontend prompts a reload, and silent overwrites of another page's changes are forbidden. The route binding API accepts an ordered array of rule IDs.
|
||||
|
||||
The React Flow editor page uses a full-width canvas and a fixed right property panel:
|
||||
|
||||
* Top bar: back, rule name, enabled state, validation state, and save.
|
||||
* The canvas uses compact height with a smaller initial fit scale; zoom, pan, box-select, delete, auto-layout, MiniMap/Controls and other necessary navigation are supported. Node dragging is handled by React Flow's local controlled state in real time; coordinates are written back to the edit graph only after the drag ends.
|
||||
* "Add processing unit" offers IP match, geo match, UA check, security, PoW, and block; start and pass are provided by the default graph and cannot be deleted or duplicated.
|
||||
* Selecting a normal node or edge allows deletion via the canvas delete button or Delete/Backspace; deleting a node synchronously removes associated edges.
|
||||
* The right property panel is hidden by default, shown only when a node is selected; it collapses when clicking an edge or blank canvas.
|
||||
* Geo-match properties use the full country and ISO 3166-2 first-level administrative division data; country options show both localized names and codes; administrative divisions support search by country name, division name, or code to avoid rendering thousands of options at once.
|
||||
* Leaving the page with unsaved changes must prompt; save conflicts and Server validation errors should locate the relevant node or edge.
|
||||
|
||||
The WAF list shows rule name, enabled state, node count, bound route count, and update time. The new-rule dialog only has the name field and navigates to the orchestration page immediately on success.
|
||||
|
||||
## Persistence and Migration
|
||||
|
||||
Rule records add a versioned graph JSON and a revision number; binding records add execution order. The graph is saved as a single aggregate (not split into node/edge tables) to keep edit operations transactional and let new node types avoid frequent DB schema extensions.
|
||||
|
||||
When upgrading existing installs:
|
||||
|
||||
* Keep rule names, global flags, enabled state, and route bindings.
|
||||
* All rule graphs reset to `start → pass`; legacy IP/geo lists, PoW, or block-response configs are not migrated.
|
||||
* Existing bindings are written into the order field in a stable sequence; global rules remain fixed in front.
|
||||
* Once the new graph and runtime stabilize, remove the legacy rule fields, fixed-order compile logic, and old frontend forms — do not maintain dual executors long-term.
|
||||
|
||||
This migration stops the old protection config from taking effect; the upgrade notes must prominently tell admins to re-orchestrate rules before releasing the next version.
|
||||
|
||||
## Release, Failure, and Rollback
|
||||
|
||||
Rule graphs only take effect on config release. If validation or compilation fails before release, the release is refused and the current active version stays unchanged. If Agent write, OpenResty config check, or reload fails, the apply flow fails and restores the previous valid released version.
|
||||
|
||||
New workers only accept complete, parseable runtime rule configs. During an OpenResty graceful reload, old workers keep the old in-memory graph and new workers use the new graph, so requests never observe a half-updated state.
|
||||
|
||||
When the geo database is unavailable, geo match returns `false` with a rate-limited warning, preserving current behavior. On IP group refresh failure, the old in-memory snapshot is kept. An incomplete PoW is taken over by the challenge module, not treated as an execution error; PoW node config is first written to OpenResty shared memory with a short-lived key, then passed to the internal challenge handler via explicit `ngx.exec` parameters — it cannot rely on internal redirects preserving `ngx.ctx` or implicitly inherited request params. Empty rule bindings in the release snapshot must be encoded as JSON empty arrays; at runtime, legacy `null` optional arrays in old snapshots are treated as empty arrays, so `cjson`'s `ngx.null` userdata never breaks the request.
|
||||
@@ -0,0 +1,93 @@
|
||||
# Zone & Domain Resource Design
|
||||
|
||||
## Goals
|
||||
|
||||
Refactor "websites" into a Zone management experience keyed by registrable root domains. A Zone like `example.com` is a stable management boundary; users enter the Zone through a stable ID path to view and maintain its explicitly declared domains, the reverse proxy routes and certificates bound to those domains, and route-level WAF, Pages and other capabilities.
|
||||
|
||||
This design replaces the concept, tables, and APIs of `managed_domains`. The Zone core does **not** include authoritative DNS record management; to point ZoneDomain A records at edge nodes, use the optional module [Cloudflare DNS Pointing](./cloudflare-pointing.md).
|
||||
|
||||
## Scope and Constraints
|
||||
|
||||
* Zone root domains are resolved with the Public Suffix List, e.g. `api.example.co.uk` belongs to `example.co.uk`.
|
||||
* URLs use IDs: list at `/websites`, detail at `/websites/:zoneId`; domains are not used as URL parameters.
|
||||
* Zone domains must be explicit FQDNs; `*.example.com` is not allowed. TLS certificates may still contain wildcard SANs and cover explicit Zone domains.
|
||||
* A Zone domain is associated with at most one reverse proxy route; a route may associate with multiple Zone domains, thus sharing the same upstream, cache, rate limit, WAF and Pages config across Zones.
|
||||
* The Zone model itself adds no DNS records, edge functions, preview subdomains, or tenant isolation. Creating/updating external DNS A records is handled by the separate Cloudflare pointing module and does not change the Zone / ZoneDomain table responsibilities.
|
||||
|
||||
## Core Model
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
ZONES ||--o{ ZONE_DOMAINS : contains
|
||||
PROXY_ROUTES ||--o{ ZONE_DOMAINS : serves
|
||||
TLS_CERTIFICATES ||--o{ ZONE_DOMAINS : secures
|
||||
PROXY_ROUTES ||--o{ WAF_RULE_GROUP_BINDINGS : applies
|
||||
PAGES_PROJECTS ||--o{ PROXY_ROUTES : backs
|
||||
|
||||
ZONES {
|
||||
uint id PK
|
||||
string domain UK
|
||||
}
|
||||
ZONE_DOMAINS {
|
||||
uint id PK
|
||||
uint zone_id
|
||||
uint proxy_route_id
|
||||
string domain UK
|
||||
uint cert_id
|
||||
}
|
||||
```
|
||||
|
||||
### `of_zones`
|
||||
|
||||
Stores the root domain, created time, and updated time. Root domains are globally unique and cannot be modified in place after creation; to change one, create a new Zone and migrate the domains. Before deleting a Zone, all of its Zone domains must be cleared first.
|
||||
|
||||
### `of_zone_domains`
|
||||
|
||||
Stores `zone_id`, explicit `domain`, nullable `proxy_route_id`, nullable `cert_id`, and timestamps. `domain` is globally unique; all relationship fields are indexed but no physical foreign keys are created. `proxy_route_id` may be null to host historical domains that have a certificate prepared but no reverse proxy configured yet.
|
||||
|
||||
`of_proxy_routes` gradually removes the domain/certificate redundancy columns `domain`, `domains`, `cert_id`, `cert_ids`, and `domain_cert_ids`. Routes must no longer specify any TLS certificate; the route name `site_name` becomes the stable human-readable identifier, and the compiler reads `server_name` and its `cert_id` from the associated Zone domains. This gives each explicit domain a single certificate source.
|
||||
|
||||
## Business and API
|
||||
|
||||
New Zone resources in the admin panel:
|
||||
|
||||
* `GET/POST /api/v1/d/zones`
|
||||
* `GET/POST /api/v1/d/zones/:id/update`
|
||||
* `POST /api/v1/d/zones/:id/delete`
|
||||
* `POST /api/v1/d/zones/:id/domains` (list returned via overview)
|
||||
* `POST /api/v1/d/zones/:id/domains/:domainID/update`
|
||||
* `POST /api/v1/d/zones/:id/domains/:domainID/delete`
|
||||
* `GET /api/v1/d/zones/:id/overview`
|
||||
|
||||
Reverse proxy route create/update requests switch to `zone_domain_ids` and no longer submit `domains`, `cert_id`, `cert_ids`, or `domain_cert_ids`. The server validates domain ownership, global uniqueness, and certificate SAN coverage in a transaction; failures are returned uniformly via `response.Abort*`. Deleting a Zone domain bound to a route requires unbinding or deleting the route first; deleting a Zone that still has domains must be rejected.
|
||||
|
||||
WAF, Pages, upstream, and release versions remain part of `proxy_routes`. The Zone overview only aggregates the route state associated with its domains and does not copy or redefine those configs.
|
||||
|
||||
## Frontend Experience
|
||||
|
||||
`/websites` shows only Zone root domains with configured domain count, route count and status, plus search, create, and action menus. Clicking enters `/websites/:zoneId`.
|
||||
|
||||
The detail page includes:
|
||||
|
||||
* Overview: domain, route, and valid certificate statistics; domain—route—certificate summary; route-level WAF and Pages summary.
|
||||
* Domains: a list of explicit FQDNs, certificate selection, and associated routes; wildcard domains are not shown or accepted.
|
||||
* Routes: routes filtered to the current Zone, linking to existing route details.
|
||||
* Certificates: certificates actually referenced by the current Zone's domains.
|
||||
* Settings: Zone notes and a protected delete operation.
|
||||
|
||||
When adding a route, select from Zone domains; users can also register domains in the Zone first, then bind a route. The global reverse proxy route entry remains but uses the same Zone domain selector.
|
||||
|
||||
## Data Migration
|
||||
|
||||
This rework ships in two release phases to avoid SQL using a wrong "last two labels" rule for multi-level public suffixes. Operation details: [Zone Domain Migration and Release Acceptance](../guide/zone-domain-migration.md).
|
||||
|
||||
1. **Phase 1 DDL**: PostgreSQL and SQLite goose create `of_zones` / `of_zone_domains` at the same version; `of_managed_domains` and route redundancy columns are temporarily kept.
|
||||
2. **Data import (automatic)**: at Server startup `migrator.Migrate()` first applies goose SQL up to `202607120002`, then automatically imports legacy route domains / `managed_domains` (registering root domains via `publicsuffix` parsing, writing `cert_id` and `proxy_route_id`), then continues with the remaining SQL. Conflicts fail startup; fixing and restarting retries idempotently. No manual command needed.
|
||||
3. **Code switch**: control-plane APIs, config snapshots, rendering, and frontend all use Zone domains as the single source; route writes only use `zone_domain_ids`.
|
||||
4. **Phase 2 cleanup**: goose SQL `202607130001_drop_legacy_route_domain_columns` drops `of_managed_domains` and the `of_proxy_routes` redundancy columns. Down only restores an empty dev-DB structure and does not backfill historical data.
|
||||
|
||||
### Runtime Model Boundaries
|
||||
|
||||
* Persistence: domains and certificates exist only in `of_zone_domains`; `of_proxy_routes` only stores route policy (upstream, cache, rate limit, WAF binding keys, etc.).
|
||||
* Rendering: the config snapshot assembles temporary `Domains` / `DomainCertIDs` in memory for OpenResty rendering and does not write back to the database.
|
||||
* Structure migration only uses `internal/infra/persistence/migrator/goose/{postgres,sqlite}/*.sql`; legacy domains are imported automatically at startup, and after phase 2 the old columns no longer exist so it is a no-op.
|
||||
@@ -0,0 +1,68 @@
|
||||
# TLS Certificates and Auto-Renewal
|
||||
|
||||
This guide explains how to manage TLS certificates in OpenFlare. To secure traffic with HTTPS, you need to configure the corresponding certificate. OpenFlare supports **manually importing existing certificates** and **automatic issuance and managed renewal via ACME**.
|
||||
|
||||
---
|
||||
|
||||
## Method 1: Manually Import an Existing Certificate
|
||||
|
||||
If you have obtained a free or paid certificate from a third-party provider (such as Tencent Cloud, Alibaba Cloud, etc.), or generated a self-signed certificate locally:
|
||||
|
||||
1. Log in to the admin panel, go to **「Website Management」->「TLS Certificates」** in the left navigation.
|
||||
2. Click **「Import Certificate」** in the top-right corner.
|
||||
3. Fill in the configuration:
|
||||
* **Certificate Name**: Enter an easily recognizable alias (e.g. `my-domain-cert`).
|
||||
* **Certificate Content (PEM)**: Paste the PEM-format certificate public key (usually starts with `-----BEGIN CERTIFICATE-----`).
|
||||
* **Private Key (KEY)**: Paste the certificate private key (usually starts with `-----BEGIN PRIVATE KEY-----` or `-----BEGIN RSA PRIVATE KEY-----`).
|
||||
4. Click **「Save」**. After a successful import, the certificate can be directly bound when configuring domains.
|
||||
|
||||
---
|
||||
|
||||
## Method 2: Automatic Issuance and Auto-Renewal (ACME)
|
||||
|
||||
OpenFlare has a built-in ACME client integrated with the **Asynq async task queue**. With the DNS API of your cloud DNS provider, the system can automatically complete DNS-01 challenge validation, apply for wildcard/single-domain certificates from a CA (Let's Encrypt by default), and **automatically trigger renewal 7 days before expiry**.
|
||||
|
||||
### Step 1: Create a DNS API Token in Cloudflare
|
||||
|
||||
To let OpenFlare automatically add TXT records under your domain for DNS validation, you need a Cloudflare API Token with specific permissions.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> For security, **it is strongly recommended to use a permission-scoped API Token** rather than the Global API Key.
|
||||
|
||||
1. Log in to the [Cloudflare dashboard](https://dash.cloudflare.com/).
|
||||
2. Click the user avatar in the top-right corner and select **「My Profile」**.
|
||||
3. In the left menu select **「API Tokens」**, then click **「Create Token」**.
|
||||
4. Find the **「Edit Zone DNS」** template and click **「Use template」**.
|
||||
5. Configure the token permissions and scope (keep defaults or restrict as needed):
|
||||
* **Permissions**:
|
||||
* `Zone` - `DNS` - `Edit` (required, for ACME to write TXT records)
|
||||
* `Zone` - `Zone` - `Read` (required, to list and retrieve zone IDs)
|
||||
* **Zone Resources**:
|
||||
* Select **「Include」** -> **「All zones」**, or select **「Specific zone」** and point to the specific domain you manage.
|
||||
6. Click **「Continue to summary」**, confirm, then click **「Create Token」**.
|
||||
7. Copy the generated **API Token** string. It is only shown once, so save it carefully.
|
||||
|
||||
### Step 2: Add a DNS Account in the Control Plane
|
||||
|
||||
1. Log in to the OpenFlare admin panel, go to **「Website Management」->「DNS Accounts」**.
|
||||
2. Click **「Add Account」**.
|
||||
3. Fill in the configuration:
|
||||
* **Account Name**: e.g. `cloudflare-main`.
|
||||
* **DNS Provider**: Select `Cloudflare`.
|
||||
* **API Token**: Paste the API token copied from Cloudflare (stored encrypted automatically).
|
||||
4. Click **「Save」**.
|
||||
|
||||
### Step 3: Submit a Certificate Application Task
|
||||
|
||||
1. Go to **「Website Management」->「TLS Certificates」**, click **「Apply for Certificate」** in the top-right corner.
|
||||
2. Fill in the application form:
|
||||
* **Certificate Name**: Custom name (e.g. `wildcard-example-cert`).
|
||||
* **Primary Domain**: The domain to apply for (wildcards supported, e.g. `example.com` or `*.example.com`).
|
||||
* **Associated Domains**: Append more domains if any (wildcards supported, comma-separated).
|
||||
* **DNS Account**: Select the DNS account just added from the dropdown (e.g. `cloudflare-main`).
|
||||
3. Click **「Save and Apply」**.
|
||||
|
||||
### Step 4: Track Application Progress and Renewal Status
|
||||
|
||||
- **Real-time progress**: After saving, the system delivers a certificate renewal/application task (`of_ssl_single_renew`) to the Asynq queue. You can view detailed step-by-step logs (adding TXT records, DNS record global propagation probing, ACME validation, certificate issuance, etc.) in the admin task or node log pages.
|
||||
- **Automatic renewal**: All certificates issued via ACME are automatically managed by the system. The background Scheduler scans certificate validity daily and automatically triggers renewal via async tasks 7 days before expiry — no manual maintenance needed.
|
||||
@@ -0,0 +1,29 @@
|
||||
# Credits
|
||||
|
||||
OpenFlare draws on the excellent ideas, architectures, and technical implementations of many open-source projects during design and development. Below are the key open-source projects referenced in OpenFlare's core underlying engines, security mechanisms, and frontend/backend frameworks. We thank these projects and their communities.
|
||||
|
||||
---
|
||||
|
||||
### 1. OpenResty
|
||||
* **Positioning**: a high-performance web platform based on Nginx and Lua.
|
||||
* **Role in OpenFlare**: the edge gateway of the global data plane. All public web traffic is first received by OpenResty, where high-concurrency HTTPS handshakes, WAF security rule matching, anti-CC human verification, and finally reverse proxy forwarding are performed.
|
||||
* **Link**: [OpenResty official site](https://openresty.org/)
|
||||
|
||||
### 2. FRP (Fast Reverse Proxy)
|
||||
* **Positioning**: a high-performance reverse proxy application focused on intranet penetration.
|
||||
* **Role in OpenFlare**: the underlying tunnel engine of the intranet penetration subsystem. The relay-side manager `openflare-relay` guards and schedules the `frps` engine, while the intranet client `openflared` auto-generates TOML config locally and guards multiplexed `frpc` child processes.
|
||||
* **Link**: [fatedier/frp (GitHub)](https://github.com/fatedier/frp)
|
||||
|
||||
---
|
||||
|
||||
### 3. Anubis (PoW solution)
|
||||
* **Positioning**: a lightweight human-verification protection solution based on Proof of Work.
|
||||
* **Role in OpenFlare**: provides the core **invisible anti-CC human challenge** capability for the gateway WAF.
|
||||
|
||||
---
|
||||
|
||||
### 4. gin-template
|
||||
* **Positioning**: a modern full-stack development scaffold template based on Go Gin and frontend builds.
|
||||
* **Role in OpenFlare**: provides a canonical, unified frontend/backend system architecture prototype for the OpenFlare control plane (Server).
|
||||
|
||||
---
|
||||
@@ -0,0 +1,72 @@
|
||||
# Publish First Configuration
|
||||
|
||||
You will learn: how to create the first reverse proxy rule in the simplest way, publish a config version, and confirm the Agent has pulled and applied the config.
|
||||
|
||||
OpenFlare's release chain centers on "immutable config versions". After you modify rules in the admin panel, you must publish and activate a new version for online Agents to auto-sync and apply.
|
||||
|
||||
---
|
||||
|
||||
## Pre-Release Checks
|
||||
|
||||
Before starting, make sure the following conditions are met:
|
||||
|
||||
| Check | Required State |
|
||||
| --- | --- |
|
||||
| **Server** | control panel started normally and you can log in to the admin panel |
|
||||
| **Agent** | at least one Agent node online (confirmable in「Node Management」) |
|
||||
| **Origin** | your backend origin service is reachable from the Agent host |
|
||||
| **Domain/testing** | the domain's DNS resolves, or you're ready to test with local hosts / curl Host header on the client |
|
||||
|
||||
---
|
||||
|
||||
## Step 1: Create the First Website Config
|
||||
|
||||
For a quick verification, deploy a basic HTTP reverse proxy site first:
|
||||
|
||||
1. Log in to the control panel, go to **「Website Management」->「Domain List」** in the left navigation, click **「Add Zone」**.
|
||||
2. Fill in the domain config:
|
||||
* **Domain**: enter the test domain (e.g. `first.example.com`).
|
||||
* **Bind Certificate**: choose not to bind a certificate (for HTTP quick verification).
|
||||
* Click save to complete domain registration.
|
||||
3. Go to **「Rule Management」**, click **「New Rule」**:
|
||||
* **Rule Name**: enter a simple identifier (e.g. `first-app-route`).
|
||||
* **Domain Match**: fill in your test domain (e.g. `first.example.com`).
|
||||
* In the **「Reverse Proxy」** tab below, set the origin mode to「Direct Upstream」.
|
||||
* **Upstream Address**: fill in the backend service address (e.g. the test-only `http://httpbin.org`).
|
||||
* Click save to create the rule.
|
||||
|
||||
> [!TIP]
|
||||
> **About HTTPS and certificate preparation**
|
||||
> This section only guides the quick deployment of a basic HTTP rule. To import an existing SSL certificate or auto-issue one from Let's Encrypt via ACME and enable HTTPS proxying on port 443, go to [Create a Reverse Proxy Config](./proxy-config.md) for detailed steps.
|
||||
|
||||
---
|
||||
|
||||
## Step 2: Preview and Publish a Config Version
|
||||
|
||||
The new website config is still a draft in the Server database and needs a released version to be distributed to the data plane:
|
||||
|
||||
1. Click the **「Preview and Publish」** button in the top-right of the control panel; the system shows the physical config file diff for the newly added route.
|
||||
2. After confirming the rendered config is correct, click **「Confirm Publish」**.
|
||||
3. The control plane generates a unique config version number (format `YYYYMMDD-NNN`).
|
||||
|
||||
---
|
||||
|
||||
## Step 3: Verify the Agent Applied It
|
||||
|
||||
After publishing, the control plane immediately notifies online Agents via WebSocket (if the WebSocket is offline, the Agent detects it as a diff in its heartbeat):
|
||||
|
||||
1. **Admin-side verification**: go to「Node Management」-> click the node to open details; check that the**current version number** has changed to the just-published latest active version and the「Apply Records」show success.
|
||||
2. **Edge node verification**: check application via logs on the Agent host:
|
||||
```bash
|
||||
# If the Agent is Docker-deployed
|
||||
docker logs openflare-agent
|
||||
|
||||
# If the Agent is deployed with local systemd
|
||||
journalctl -u openflare-agent -n 50 --no-pager
|
||||
```
|
||||
3. **Connectivity test**:
|
||||
On the client machine, use `curl` with a test Host header against the Agent node's IP for final verification:
|
||||
```bash
|
||||
curl -I -H "Host: first.example.com" http://AGENT_NODE_IP
|
||||
```
|
||||
If the returned status code matches the backend origin's response, your first reverse proxy rule has successfully landed on the edge node!
|
||||
@@ -0,0 +1,50 @@
|
||||
# Guide
|
||||
|
||||
You will learn: how the OpenFlare docs are organized, which pages to read on first run, and where to start for deployment, usage, and troubleshooting.
|
||||
|
||||
OpenFlare is a self-hosted OpenResty control plane. It brings reverse proxy website configs, config version release, Agent node sync, TLS certificates, and basic observability into one admin panel — suitable for a single team or organization managing multiple proxy nodes.
|
||||
|
||||
## Recommended Reading Path
|
||||
|
||||
If you're new to OpenFlare, read in this order:
|
||||
|
||||
1. [Quick Start](./quick-start.md): start the Server with Docker Compose, log in to the admin panel, and connect your first Agent.
|
||||
2. [Publish First Configuration](./first-site.md): quickly create a basic HTTP reverse proxy site rule and verify the node applied it.
|
||||
3. [Create a Reverse Proxy Config](./proxy-config.md): step by step, from certificate import and application to HTTPS, upstream origins, and edge cache.
|
||||
4. [Zone Domain Migration](./zone-domain-migration.md): upgrade from legacy managed domains / inline route domains to the Zone model (automatic goose import), with backup, acceptance, and rollback notes.
|
||||
5. [Pages Static Hosting Usage](./pages-usage.md): static project ZIP upload limits, SPA Fallback, and built-in API reverse proxy config.
|
||||
6. [Tunnel & Intranet Penetration](./tunnel-usage.md): deploy Relay and Client for secure reverse penetration without a public IP.
|
||||
7. [WAF Security Protection](./waf-usage.md): configure WAF rule groups; master IP allow/block lists, auto/subscription IP groups, geo restrictions, and PoW CC protection.
|
||||
8. [WAF Auto IP Group Expressions](./waf-ip-group-expr.md): write auto IP group Expr rules; understand keyword meanings and preset rules.
|
||||
9. [Uptime Kuma Monitoring Sync](./uptime-kuma.md): configure Uptime Kuma auto differential sync and monitor scope control.
|
||||
10. [SSO Login Configuration](./sso.md): configure OIDC for third-party single sign-on (SSO).
|
||||
11. [Troubleshooting](./troubleshooting.md): troubleshoot login, database, node sync, OpenResty, and edge cache hit issues by symptom.
|
||||
12. [Credits](./credits.md): the excellent open-source projects and community acknowledgments this system depends on.
|
||||
|
||||
## Find by Role
|
||||
|
||||
| What you want to do | Recommended Entry |
|
||||
| --- | --- |
|
||||
| Get the admin panel running in 5 minutes | [Quick Start](./quick-start.md) |
|
||||
| Publish your first reverse proxy config | [Publish First Configuration](./first-site.md) |
|
||||
| Configure domain certs, reverse proxy, and edge cache | [Create a Reverse Proxy Config](./proxy-config.md) (incl. cache notes) |
|
||||
| Static assets not hitting cache | [Troubleshooting · Edge Cache](./troubleshooting.md#edge-cache-hit-rate-anomalies) |
|
||||
| Host an SPA or static website | [Pages Static Hosting Usage](./pages-usage.md) |
|
||||
| Configure intranet penetration mapping | [Tunnel & Intranet Penetration](./tunnel-usage.md) |
|
||||
| Configure anti-CC and IP group blocking | [WAF Security Protection](./waf-usage.md) |
|
||||
| Write auto IP group rules | [WAF Auto IP Group Expressions](./waf-ip-group-expr.md) |
|
||||
| Auto-sync monitored site status | [Uptime Kuma Monitoring Sync](./uptime-kuma.md) |
|
||||
| Connect or reinstall a node Agent | [Access Agent](../deployment/agent.md) |
|
||||
| Start the Server from source | [Start the Server](../deployment/server.md) |
|
||||
| Configure OIDC login | [SSO Login Configuration](./sso.md) |
|
||||
| Upgrade Server or Agent | [Upgrade & Maintenance](../deployment/upgrade.md) |
|
||||
| Understand the architecture and release model | [System Architecture](../design/architecture.md) and [Agent & Publish Model](../design/agent-design.md) |
|
||||
| See open-source references and acknowledgments | [Credits](./credits.md) |
|
||||
|
||||
## Doc Sections
|
||||
|
||||
`guide/` targets users and deployers with executable steps from install to daily operations.
|
||||
|
||||
`reference/` consolidates stable facts: config fields, commands, API response conventions, and repository structure.
|
||||
|
||||
`design/` targets maintainers and contributors, describing product boundaries, system architecture, the Agent & publish model, and engineering constraints. Before adding capabilities or changing boundaries, update the corresponding design doc first.
|
||||
@@ -0,0 +1,116 @@
|
||||
# Pages Static Hosting Usage
|
||||
|
||||
You will learn: how to deploy pre-built static sites via local upload, Remote URL, or public GitHub Release assets; configure SPA Fallback and API reverse proxy; and safely check for updates, auto-publish, and roll back.
|
||||
|
||||
---
|
||||
|
||||
## Core Mechanics and Page Structure
|
||||
|
||||
OpenFlare Pages is inspired by Cloudflare Pages' Direct Upload and deployment history interaction, but currently handles **pre-built artifacts** rather than building from repository source. The project detail is organized as "current production deployment → deployment source → deployment history": source configuration can change, while created deployments stay immutable.
|
||||
|
||||
```text
|
||||
Local upload ─> unified validation / upload.Ingest ─> new candidate ─> admin explicit activation ─┐
|
||||
Remote URL ── Server restricted download ─────────────┐ │
|
||||
GitHub Release asset ─ Server resolves ───────────────┴─> create/load deployment ────────────────┤
|
||||
└─> source sync atomic activation ────────┘
|
||||
|
|
||||
v
|
||||
Agent pulls per-project latest
|
||||
|
|
||||
v
|
||||
OpenResty local static serving
|
||||
```
|
||||
|
||||
External URLs, GitHub metadata, and auto-checks are handled only by the Server. The Agent only pulls the currently active deployment package from the control plane; it does not receive external source credentials, nor does it run `git clone`, dependency installation, or build commands.
|
||||
|
||||
## Step 1: Create a Project
|
||||
|
||||
1. Log in to the admin panel, go to **「Pages」**, click **「Create Project」**.
|
||||
2. Fill in the project name and a unique Slug.
|
||||
3. Configure the content entry:
|
||||
* **Entry file name**: default `index.html`.
|
||||
* **Static asset root path (RootDir)**: fill in the relative path when artifacts are in a subdirectory like `dist/`; leave empty when artifacts are at the archive root.
|
||||
4. Set SPA Fallback and API proxy as needed. RootDir and entry file are project-level configs applied uniformly to all sources.
|
||||
|
||||
## Step 2: Choose a Deployment Source
|
||||
|
||||
### 1. Manual Upload
|
||||
|
||||
Without a persistent source configured, the project stays in manual mode. Click **「Upload Deployment Package」** to select a pre-built archive; a successful upload creates a candidate deployment, which you then explicitly activate from the deployment history. Re-uploading does not modify existing deployments.
|
||||
|
||||
Supported formats: `zip`, `tar.gz` / `tgz`, `tar.xz` / `txz`, `tar.bz2` / `tbz2`, `tar`, and `7z`.
|
||||
|
||||
### 2. Remote URL
|
||||
|
||||
In the deployment source card select **Remote URL**, fill in the HTTP(S) address and choose a network policy:
|
||||
|
||||
* **public**: default policy; rejects loopback, private network, link-local addresses, DNS rebinding, self-signed TLS, and redirects to non-public targets.
|
||||
* **trusted_internal**: only for explicitly trusted intranet or self-signed services; requires a second risk confirmation before saving.
|
||||
|
||||
After saving, the address is only displayed masked. You don't need to re-enter it when editing other configs; only submit a new URL when choosing to change the address. Remote sources only offer **「Sync and Publish」**: the Server downloads, validates, and atomically activates each time — no "check for updates", scheduled checks, or auto-updates.
|
||||
|
||||
### 3. GitHub Release
|
||||
|
||||
GitHub sources only support public `github.com` repositories. Fill in:
|
||||
|
||||
* A repository address in `https://github.com/{owner}/{repo}` format;
|
||||
* **Latest Release** or a **fixed Tag**;
|
||||
* An exact, case-sensitive Release Asset filename, default `dist.zip`.
|
||||
|
||||
Both options support manual **「Check for Updates」** and **「Sync and Publish」**. Differences:
|
||||
|
||||
* **latest**: supports a check interval of 5–1440 minutes, default 1440 minutes (24 hours); auto-update is off by default. When enabled, the scanner asynchronously syncs and publishes only when a new revision is found.
|
||||
* **tag**: only supports manual admin checks and sync; does not participate in the scheduled scanner.
|
||||
|
||||
"Check for updates" only resolves the Release/asset and advances the version cursor without downloading the deployment package; "Sync and publish" downloads, validates, creates or reuses a deployment, and activates it. If the asset under the same Release is replaced, the source enters **「Needs Confirmation」** — you must confirm the exact revision shown before publishing, to avoid silent overwrites.
|
||||
|
||||
GitHub Release sources only import pre-built artifacts; they do not build from repository source.
|
||||
|
||||
### 4. Switch or Delete a Source
|
||||
|
||||
You can switch between Manual, Remote, and GitHub Release. Modifying or deleting a source does not delete the current production deployment or historical deployments; switching back to manual mode lets you continue uploading and explicitly activating.
|
||||
|
||||
## Deployment Package Security Limits
|
||||
|
||||
Deployment packages must satisfy these constraints:
|
||||
|
||||
* Archive size is controlled by the system config `pages_max_package_size_mb`, default 100 MiB, configurable 1–2048 MiB.
|
||||
* Expanded single-file and total size limits are "package size limit × 4", with a floor of 100 MiB; at most 1,000 regular files.
|
||||
* The control plane streams regular file bodies, checking declared size against actual bytes, and validates the project entry file.
|
||||
* Absolute paths, `..` path traversal, symlinks, hard links, and special files in archives are all rejected.
|
||||
|
||||
The Agent also verifies SHA-256, real response byte limits, and post-extraction file count and total size on download; failures do not switch the existing `current`.
|
||||
|
||||
## Step 3: Configure Advanced Routing Rules
|
||||
|
||||
### 1. SPA Fallback
|
||||
|
||||
When using front-end routing like React Router or Vue Router, enable **「SPA Fallback」** and set the entry path (usually `/index.html`). When a visitor accesses a physical path that doesn't exist, OpenResty falls back to the entry file for the front-end router to handle.
|
||||
|
||||
### 2. API Reverse Proxy
|
||||
|
||||
Pages can forward a specified prefix to a backend API under the same domain:
|
||||
|
||||
* **APIProxyPath**: match prefix, e.g. `/api`.
|
||||
* **APIProxyPass**: backend address, e.g. `http://10.0.0.5:8080`.
|
||||
* **APIProxyRewrite**: optional path rewrite rule.
|
||||
|
||||
Requests matching the API prefix go through the reverse proxy; other requests continue to be served by the static site.
|
||||
|
||||
## Step 4: Bind a Route and First Publish
|
||||
|
||||
1. Create or edit a proxy rule.
|
||||
2. Set the origin type to **Pages** and select the Pages **project**.
|
||||
3. Preview the config, then publish and activate.
|
||||
|
||||
The route binds to a stable project ID, not a specific deployment. The first publish gives the Agent the project anchor; afterwards, local uploads, source syncs, auto-updates, or manual rollbacks only change the project's active deployment — the Agent converges via the latest hash reconciliation without needing to republish the main config.
|
||||
|
||||
## Operations, Status, and Rollback
|
||||
|
||||
* The source card shows the last check/sync time, found vs. applied revision, next check time, and security errors. While a check or sync task runs, the page polls the task status; when latest is idle, it refreshes at low frequency only near the check time.
|
||||
* A failed auto-update does not replace the old active deployment; a single source failure does not block the scanner from processing other projects.
|
||||
* Activating another deployment in the history is a manual rollback. The system fences in-flight source tasks and disables that source's auto-update to avoid the next latest round overwriting your manual choice; re-activating the current version is a no-op.
|
||||
* The Agent downloads to a temp file, verifies SHA-256, extracts safely, then atomically switches `current`. Any failure keeps the old content; with multi-project reconciliation, a single project failure does not affect others.
|
||||
|
||||
> [!TIP]
|
||||
> For the source state machine, auto scanner, upload compensation, immutable deployments, and Agent atomic switching, see [Pages Static Hosting Design](../design/pages-design.md).
|
||||
@@ -0,0 +1,117 @@
|
||||
# Create a Reverse Proxy Config
|
||||
|
||||
You will learn: how to create and publish a reverse proxy website configuration from scratch, step by step, in OpenFlare. This guide walks you through certificate import and application, origin definition, route rule configuration, version release, and connectivity verification.
|
||||
|
||||
---
|
||||
|
||||
## Recommended Workflow
|
||||
|
||||
In the gateway control plane, follow these steps to add a new reverse proxy rule:
|
||||
|
||||
```text
|
||||
[ Step 1. Certificate Management ] ──► [ Step 2. Origin Definition (optional) ] ──► [ Step 3. Add Website Config ]
|
||||
│
|
||||
[ Step 5. Verify Access ] ◄── [ Step 4. Publish & Activate Version ] ◄───────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Step 1: Prepare Certificates
|
||||
|
||||
Before using HTTPS-secured traffic, you need to prepare the corresponding TLS certificate (supports manually importing an existing certificate, or automatically applying from a CA via DNS validation with managed renewal).
|
||||
|
||||
To keep this guide concise, the certificate details (including creating a dedicated DNS API Token in Cloudflare) have been split into a dedicated guide. First go to **[TLS Certificates & Auto-Renewal](./certificates.md)** to prepare the certificate, then come back to continue.
|
||||
|
||||
---
|
||||
|
||||
## Step 2: Prepare the Upstream Origin (optional)
|
||||
|
||||
An Origin represents the backend real service address being proxied. Although you can fill in an IP directly when creating a website, it is recommended to register origins in the origin library first for reuse and maintenance:
|
||||
|
||||
1. Go to **「Website Management」->「Origin Addresses」** in the left navigation, click **「Add Origin」**.
|
||||
2. Fill in the origin name (e.g. `production-api`).
|
||||
3. Enter a valid upstream address (e.g. `http://10.0.0.10:8080`) and save.
|
||||
|
||||
---
|
||||
|
||||
## Step 3: Create the Website Config
|
||||
|
||||
Once the certificate and origin are ready, create the core website proxy route:
|
||||
|
||||
1. Go to **「Website Management」->「Domain List」**, click **「Add Zone」**:
|
||||
* **Domain**: Enter the domain bound to this site.
|
||||
* **Bind Certificate**: Select the certificate prepared or applied for in Step 1.
|
||||
2. Configure the request route rule: go to the **「Rule Management」** page, click **「New Rule」** or edit an existing rule:
|
||||
* **Rule Name**: Enter a unique simple identifier (e.g. `app-portal-route`).
|
||||
* **Domain Match**: Enter the corresponding domain (wildcards or exact domains supported; must match the registered domain above).
|
||||
* In the **「Reverse Proxy」** tab below, select the origin mode as「Direct Upstream」.
|
||||
* **Origin Selection**: Select the origin created in Step 2 from the dropdown; or choose manual input and fill in `http://10.0.0.20:9000`.
|
||||
3. Click save to create the config.
|
||||
|
||||
---
|
||||
|
||||
## Step 4: Publish and Apply the Config
|
||||
|
||||
Website configs added in the admin panel are only saved in the Server database — **they do not take effect immediately**. You must generate a config version snapshot and distribute it to the Agent edge nodes:
|
||||
|
||||
1. Click the **「Preview and Publish」** button in the top-right corner of the control panel.
|
||||
2. Review the config file diff, confirming the newly added `server` block and certificate binding rules are correct.
|
||||
3. Click **「Confirm Publish」**.
|
||||
4. **Agent application mechanism**:
|
||||
* The Agent node on the data plane detects the active version Checksum change in its heartbeat, and automatically pulls the full OpenResty config files and certificate bundle locally.
|
||||
* It automatically runs a local config validation (similar to `openresty -t`); after confirming no syntax errors, it performs a smooth reload.
|
||||
* If reload or validation fails, the Agent safely blocks and rolls back to the previous stable version to keep the node highly available.
|
||||
|
||||
---
|
||||
|
||||
## Step 5: Connectivity and Rollback Verification
|
||||
|
||||
### 1. Verify Access
|
||||
You can verify the new config takes effect as follows:
|
||||
* **Browser access**: Open `https://your-domain.com` directly in a browser and check whether it proxies the backend successfully.
|
||||
* **CLI verification** (recommended): probe with `curl`:
|
||||
```bash
|
||||
curl -I https://your-domain.com
|
||||
```
|
||||
* **Bypass DNS validation**: if your domain is not yet resolvable, temporarily send a `Host` header request to the Agent node's physical IP:
|
||||
```bash
|
||||
curl -I -H "Host: your-domain.com" https://AGENT_NODE_IP --insecure
|
||||
```
|
||||
|
||||
### 2. One-Click Second-Level Rollback
|
||||
If the released config causes an online business issue:
|
||||
1. Navigate to the **「Version Release」** menu on the left.
|
||||
2. Find the previous stable version before the release in the history list.
|
||||
3. Click **「Activate」**.
|
||||
4. All online Agent nodes will automatically reload the historical config within seconds for second-level risk avoidance.
|
||||
|
||||
---
|
||||
|
||||
## Edge Cache (optional)
|
||||
|
||||
The **「Cache」** page in the site details can enable edge `proxy_cache` (requires **Performance Settings → Global OpenResty Cache** to be enabled at the same time). Behavior mirrors the Cloudflare default model; see [Edge Cache Strategy Design](../design/edge-cache-design.md).
|
||||
|
||||
### Recommended Settings
|
||||
|
||||
| Item | Suggestion |
|
||||
| --- | --- |
|
||||
| Strategy | **Standard static assets** (recommended default): only css/js/map/images/fonts, **not HTML/JSON** |
|
||||
| Login Cookie | Is **not** separately skipped from caching; users with sessions can still hit static assets |
|
||||
| Origin | Static assets: `Cache-Control: public, max-age=…`; dynamic/personalized must be `private` or `no-store` |
|
||||
| Response Set-Cookie | Is not written to the edge cache |
|
||||
| No origin cache headers | Uses default Edge TTL by status code (e.g. ~120 min for 200) |
|
||||
|
||||
### Advanced Strategy「All Cacheable GET」
|
||||
|
||||
Similar to Cloudflare Cache Everything: the path is no longer limited by extension. If the origin does not declare `private`/`no-store` for HTML, **personalized pages may be cached and served across users**. Use only when origin cache headers are correct or content is globally consistent.
|
||||
|
||||
### How It Takes Effect
|
||||
|
||||
The cache switch and strategy are written into the config snapshot. After saving the site, you must **publish and activate the config version** for the Agent to apply it. Changing the UI only without publishing leaves nodes on the old rules.
|
||||
|
||||
### Quick Self-Check
|
||||
|
||||
1. Global cache on, site cache on, strategy「Standard static assets」.
|
||||
2. Publish the config and confirm nodes applied successfully.
|
||||
3. Request the same `/assets/app.js` (or a hashed immutable path) twice with a login cookie; the `cache_status` in access logs should be **HIT** on the second request.
|
||||
4. If still「not cached」: check whether the strategy matches the path extension, whether it is a non-GET request, whether the origin returns `Set-Cookie` / `private`, and whether the node applied the new version. More in [Troubleshooting · Edge Cache](./troubleshooting.md#edge-cache-hit-rate-anomalies).
|
||||
@@ -0,0 +1,241 @@
|
||||
# Quick Start
|
||||
|
||||
You will learn: how to start the OpenFlare Server with Docker Compose, complete the first login, connect your first Agent, and verify that a config has been published to a node.
|
||||
|
||||
OpenFlare's minimal runtime consists of:
|
||||
|
||||
| Component | Responsibility |
|
||||
| --- | --- |
|
||||
| Server | admin UI, admin API, Agent API, config rendering, version release, and state storage |
|
||||
| Agent | runs on proxy nodes; pulls config, writes OpenResty, executes validation and reload |
|
||||
| OpenResty | actually receives traffic and reverse proxies to origins |
|
||||
|
||||
The Agent uniformly controls the runtime via the OpenResty binary. Local deployment requires an `openresty` executable on the node; Docker deployment can directly run the Agent image with a built-in OpenResty.
|
||||
|
||||
## Environment Requirements
|
||||
|
||||
| Item | Requirement |
|
||||
| --- | --- |
|
||||
| Docker / Docker Compose | starts the Server and its PostgreSQL, Valkey dependencies; also runs the Agent if using the Docker Agent |
|
||||
| OpenResty | local Agent installs need an executable `openresty`, or specify the path in the install script |
|
||||
| Reachable port | Server listens on `3000` by default; Agent nodes must be able to reach the Server address |
|
||||
|
||||
---
|
||||
|
||||
## 1. Start the Server
|
||||
|
||||
Quick start recommends the standard **PostgreSQL + Valkey** deployment.
|
||||
|
||||
Create a `docker-compose.yaml` in an empty directory:
|
||||
|
||||
```yaml
|
||||
version: '3.8'
|
||||
|
||||
services:
|
||||
openflare:
|
||||
image: ghcr.io/rain-kl/openflare:latest
|
||||
container_name: openflare-server
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "3000:3000"
|
||||
volumes:
|
||||
- openflare_uploads:/app/uploads
|
||||
environment:
|
||||
TZ: Asia/Shanghai
|
||||
APP_SESSION_SECRET: 'replace-with-a-long-random-string' # replace with a long random string in production
|
||||
DB_ENABLED: "true"
|
||||
DB_HOST: "postgres"
|
||||
DB_PORT: "5432"
|
||||
DB_USERNAME: "${DB_USERNAME:-openflare}"
|
||||
DB_PASSWORD: "${DB_PASSWORD:-replace-with-strong-password}"
|
||||
DB_NAME: "${DB_NAME:-openflare}"
|
||||
REDIS_ENABLED: "true"
|
||||
REDIS_ADDR: "redis:6379"
|
||||
depends_on:
|
||||
postgres:
|
||||
condition: service_healthy
|
||||
redis:
|
||||
condition: service_healthy
|
||||
|
||||
postgres:
|
||||
image: postgres:17-alpine
|
||||
restart: unless-stopped
|
||||
environment:
|
||||
POSTGRES_DB: ${DB_NAME:-openflare}
|
||||
POSTGRES_USER: ${DB_USERNAME:-openflare}
|
||||
POSTGRES_PASSWORD: ${DB_PASSWORD:-replace-with-strong-password}
|
||||
volumes:
|
||||
- openflare_postgres_data:/var/lib/postgresql/data
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "pg_isready -U ${DB_USERNAME:-openflare} -d ${DB_NAME:-openflare}"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
|
||||
redis:
|
||||
image: valkey/valkey:8.0-alpine
|
||||
restart: unless-stopped
|
||||
command: ["valkey-server", "--appendonly", "yes"]
|
||||
volumes:
|
||||
- openflare_redis_data:/data
|
||||
healthcheck:
|
||||
test: ["CMD", "valkey-cli", "ping"]
|
||||
interval: 10s
|
||||
timeout: 5s
|
||||
retries: 5
|
||||
|
||||
|
||||
volumes:
|
||||
openflare_uploads:
|
||||
openflare_postgres_data:
|
||||
openflare_redis_data:
|
||||
```
|
||||
|
||||
Start the services:
|
||||
|
||||
```bash
|
||||
docker compose up -d
|
||||
```
|
||||
|
||||
Confirm the containers are running:
|
||||
|
||||
```bash
|
||||
docker compose ps
|
||||
docker compose logs -f openflare
|
||||
```
|
||||
|
||||
After seeing `server listening` and the `openflare-server` container status running, open in a browser:
|
||||
|
||||
```text
|
||||
http://localhost:3000
|
||||
```
|
||||
|
||||
Default account:
|
||||
|
||||
| Username | Password |
|
||||
| --- | --- |
|
||||
| `admin` | `12345678` |
|
||||
|
||||
> [!WARNING]
|
||||
> For your system's security, change the default password immediately after the first login.
|
||||
|
||||
If you forget the password and no password-recovery channel is configured, reset it with:
|
||||
|
||||
```bash
|
||||
go run main.go reset-passwd --user admin
|
||||
```
|
||||
|
||||
Without `--password`, the command auto-generates a random password and prints it to the terminal; you can also explicitly specify a new password with `--password`.
|
||||
|
||||
---
|
||||
|
||||
## 2. Prepare an Agent Token
|
||||
|
||||
Agents can connect with two credential types:
|
||||
|
||||
| Credential | Use Case |
|
||||
| --- | --- |
|
||||
| `discovery_token` | first-time auto registration; the Server exchanges it for a node-specific Token |
|
||||
| `agent_token` | node already created or assigned in the admin panel; use the node-specific Token directly |
|
||||
|
||||
Prepare one of these credentials in the admin panel, then continue.
|
||||
|
||||
- **`discovery_token`** menu path:「System Settings」->「OpenFlare」tab ->「Discovery Token & Deployment」→ Discovery Token
|
||||
- **`agent_token`** menu path: after creating a node in「Node Management」, click into the node detail page to see its dedicated Token.
|
||||
|
||||
---
|
||||
|
||||
## 3. Install / Run the Agent
|
||||
|
||||
Docker image deployment is recommended; you can also deploy to the local host via the install script.
|
||||
|
||||
### Option A: Run the Agent with Docker (recommended)
|
||||
|
||||
Run the Agent image directly on the proxy node:
|
||||
|
||||
```bash
|
||||
docker pull ghcr.io/rain-kl/openflare-agent:latest
|
||||
docker rm -f openflare-agent 2>/dev/null || true
|
||||
docker run -d --name openflare-agent --restart unless-stopped \
|
||||
-p 80:80 -p 443:443/tcp -p 443:443/udp \
|
||||
-v openflare-agent-pages:/data/var/lib/openflare/pages \
|
||||
-e OPENFLARE_SERVER_URL=http://your-server:3000 \
|
||||
-e OPENFLARE_AGENT_TOKEN=YOUR_AGENT_TOKEN \
|
||||
ghcr.io/rain-kl/openflare-agent:latest
|
||||
```
|
||||
|
||||
### Option B: Run the install script (local deployment)
|
||||
|
||||
Run the install script on the proxy node.
|
||||
|
||||
With `discovery_token`:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--discovery-token YOUR_DISCOVERY_TOKEN
|
||||
```
|
||||
|
||||
With the node-specific `agent_token`:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--agent-token YOUR_AGENT_TOKEN
|
||||
```
|
||||
|
||||
The script defaults:
|
||||
|
||||
| Item | Default |
|
||||
| --- | --- |
|
||||
| Install directory | `/opt/openflare-agent` |
|
||||
| Config file | `/opt/openflare-agent/agent.json` |
|
||||
| systemd service | `openflare-agent.service` |
|
||||
| OpenResty path | auto-finds `openresty` when unspecified |
|
||||
|
||||
Confirm the Agent service state:
|
||||
|
||||
```bash
|
||||
systemctl status openflare-agent
|
||||
journalctl -u openflare-agent -f
|
||||
```
|
||||
|
||||
Without systemd, the script prints a manual start command.
|
||||
|
||||
---
|
||||
|
||||
## 4. Next Steps
|
||||
|
||||
After starting the control panel and connecting an Agent node, you've successfully built the base runtime environment of the OpenFlare gateway. Continue with these two guides to deploy your first reverse proxy site:
|
||||
|
||||
1. **Publish your first website**:
|
||||
* See [Publish First Configuration](./first-site.md). It guides you to publish your first proxy rule in the simplest way (plain HTTP) and verify the node applied it.
|
||||
2. **Full reverse proxy config (HTTPS & origin management)**:
|
||||
* See [Create a Reverse Proxy Config](./proxy-config.md). It guides you from certificate import/application to domain HTTPS certificate binding, origin management, and preview release.
|
||||
|
||||
---
|
||||
|
||||
## When You Hit Problems
|
||||
|
||||
Handle in this order:
|
||||
|
||||
1. Upgrade Server and Agent to the latest version; confirm whether the problem persists.
|
||||
2. Re-publish and activate a config version, wait for the node to apply.
|
||||
3. Run「Force Sync」on the target node in the node detail page to push an immediate config pull.
|
||||
4. Rebuild or reinstall the Agent (re-run the install script).
|
||||
5. If none of the above works, file a [GitHub Issue](https://github.com/Rain-kl/OpenFlare/issues) with the Server logs and node apply records.
|
||||
|
||||
More troubleshooting: [Troubleshooting](./troubleshooting.md).
|
||||
|
||||
---
|
||||
|
||||
## Advanced Deployment Guides
|
||||
|
||||
After completing the quick start and getting familiar with OpenFlare, read these advanced deployment docs to put components into production:
|
||||
|
||||
* **Server production deployment**: read [Start the Server](../deployment/server.md) for building the frontend from source, system env vars, and Docker Compose.
|
||||
* **Agent production access**: read [Access Agent](../deployment/agent.md) for systemd service management, detailed local config file fields, and troubleshooting.
|
||||
* **Intranet relay deployment**: read [Deploy Relay](../deployment/relay.md) for configuring public relay nodes (frps) for tunnels.
|
||||
* **Intranet client deployment**: read [Deploy OpenFlared](../deployment/openflared.md) for running the tunnel daemon client (frpc) on the intranet server.
|
||||
* **Production topology reference**: read [Deployment Guide](../deployment/deployment.md) for production HA topology and overall network planning.
|
||||
* **Upgrades and maintenance**: read [Upgrade & Maintenance](../deployment/upgrade.md) for smooth upgrades of the Server and Agent nodes.
|
||||
@@ -0,0 +1,85 @@
|
||||
# SSO Login Configuration
|
||||
|
||||
You will learn: how to configure an OIDC third-party login entry for OpenFlare, fill in the callback URL, and how third-party accounts bind to local users.
|
||||
|
||||
OpenFlare connects third-party login through OIDC auth sources. Any service providing standard OIDC Discovery (Google, Keycloak, authentik, Logto, Casdoor, etc.) can be integrated.
|
||||
|
||||
After an auth source is configured and enabled, it appears in the third-party account login area on the login page. Users can log in with a third-party account, or bind a third-party account to the current local account while logged in.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
| Item | Description |
|
||||
| --- | --- |
|
||||
| Server access URL | configured in admin「System Settings」->「System Settings」tab ->「General Settings」; must match the address users' browsers actually visit (protocol, domain, port) |
|
||||
| Auth source name | unique identifier inside OpenFlare, e.g. `company-oidc` |
|
||||
| Client ID | provided after creating the app on the third-party platform |
|
||||
| Client Secret | provided after creating the app on the third-party platform |
|
||||
| OIDC Discovery URL | e.g. `https://idp.example.com/.well-known/openid-configuration` |
|
||||
|
||||
The auth source name may only contain letters, digits, hyphens, or underscores, and must start with a letter or digit.
|
||||
|
||||
## Callback URL
|
||||
|
||||
The Redirect URI / Callback URL on the third-party platform is fixed to:
|
||||
|
||||
```text
|
||||
<server access URL>/login
|
||||
```
|
||||
|
||||
For example, with a server access URL of `https://openflare.example.com`:
|
||||
|
||||
```text
|
||||
https://openflare.example.com/login
|
||||
```
|
||||
|
||||
The callback URL only relates to the「server access URL」and does not include the auth source name. After the third-party platform completes authorization, it redirects here, and the OpenFlare login page uses the authorization code to complete login or binding.
|
||||
|
||||
## Configure OIDC Login
|
||||
|
||||
1. Create an app or client on the OIDC Provider; choose Web / Confidential Client as the app type.
|
||||
2. Set the Redirect URI / Callback URL to `<server access URL>/login`.
|
||||
3. Copy the Client ID and Client Secret.
|
||||
4. Get the Provider's Discovery URL, usually ending in `/.well-known/openid-configuration`.
|
||||
5. Log in to the OpenFlare admin panel, go to **「System Settings」**, select the **「Security Settings」** tab, and add an auth source in the **「Auth Source Management」** section.
|
||||
6. Choose type `OIDC`; fill in the auth source name, display name, Client ID, Client Secret, and OIDC Discovery URL.
|
||||
7. Scope defaults to `openid profile email`. If the Provider restricts scopes, adjust according to the Provider's allowed values.
|
||||
8. Save and enable the auth source.
|
||||
|
||||
Once enabled, the login page shows the corresponding third-party login button.
|
||||
|
||||
## Login and Binding Behavior
|
||||
|
||||
When a third-party account returns to OpenFlare, it is handled as follows:
|
||||
|
||||
| Scenario | Behavior |
|
||||
| --- | --- |
|
||||
| Third-party account already bound to a local user | log in directly |
|
||||
| User already logged in and initiates third-party authorization | bind to the current local user |
|
||||
| Third-party account not bound, and registration allowed | auto-create a normal user and bind |
|
||||
| Third-party account not bound, and registration disabled | require entering an existing local account password to bind |
|
||||
|
||||
To allow only existing users to use SSO, turn off user registration. Unbound third-party accounts then enter the bind-existing-account flow.
|
||||
|
||||
## Modifying an Auth Source
|
||||
|
||||
When editing an auth source, leaving the Client Secret input empty keeps the existing secret; entering a new value overwrites and saves it.
|
||||
|
||||
Changing the auth source name does not affect the callback URL, so the third-party platform config doesn't need to change.
|
||||
|
||||
## FAQ
|
||||
|
||||
### Returns `invalid_scope`
|
||||
|
||||
The third-party platform doesn't allow the currently configured scope. The OIDC default scope is `openid profile email`. Adjust the scope on the auth source edit page, or allow the scope on the third-party platform.
|
||||
|
||||
### Callback URL mismatch
|
||||
|
||||
Check that the Redirect URI / Callback URL on the third-party platform exactly matches `<server access URL>/login`. Protocol, domain, port, and path must all match.
|
||||
|
||||
### Login page doesn't show the third-party login button
|
||||
|
||||
Check that the auth source is enabled and that the Client ID and Client Secret are saved. OpenFlare validates these fields before enabling the auth source.
|
||||
|
||||
### Client Secret saved, but the list doesn't show the plaintext
|
||||
|
||||
This is expected. OpenFlare never echoes the Client Secret through the API; it only shows whether a secret is configured.
|
||||
@@ -0,0 +1,221 @@
|
||||
# Troubleshooting
|
||||
|
||||
You will learn: how to troubleshoot OpenFlare Server, database, login, Agent, OpenResty, and config release issues by symptom.
|
||||
|
||||
First determine which layer the problem is in: browser, Server, database, Agent, OpenResty, origin, or DNS. OpenFlare configs are not written to all nodes online directly — only after the active version changes do Agents detect and apply it in their heartbeat.
|
||||
|
||||
## Quick Locate
|
||||
|
||||
| Symptom | Look Here First |
|
||||
| --- | --- |
|
||||
| Admin panel won't open | Server container/process logs, port listening |
|
||||
| Login abnormal | default account, Session Cookie, Server logs |
|
||||
| Data won't save | DB connection, SQLite file permissions, PostgreSQL health |
|
||||
| Agent offline | Agent logs, Token, Server address, network connectivity |
|
||||
| Node not updated after release | active version, node heartbeat, apply records |
|
||||
| OpenResty apply failure | apply records, Agent logs, certificates, upstream addresses, port usage |
|
||||
| Access analytics empty | OpenResty container state, observability port, Agent backfill logs |
|
||||
| Static assets never hit cache | global/site cache switch, policy extensions, whether config released, access log `cache_status`, origin Set-Cookie / Cache-Control |
|
||||
|
||||
## Server Won't Start
|
||||
|
||||
1. Check the logs:
|
||||
|
||||
```bash
|
||||
docker compose logs -n 200 openflare
|
||||
```
|
||||
|
||||
For source runs, check terminal output.
|
||||
|
||||
2. Check port usage:
|
||||
|
||||
```bash
|
||||
lsof -i :3000
|
||||
```
|
||||
|
||||
3. If using PostgreSQL, confirm DB health:
|
||||
|
||||
```bash
|
||||
docker compose ps postgres
|
||||
docker compose logs -n 100 postgres
|
||||
```
|
||||
|
||||
4. If using SQLite, confirm the DB file directory is writable:
|
||||
|
||||
```bash
|
||||
ls -ld "$(dirname /path/to/openflare.db)"
|
||||
```
|
||||
|
||||
Common causes:
|
||||
|
||||
| Log or Symptom | Handling |
|
||||
| --- | --- |
|
||||
| DB connection failed | check `DB_HOST`, `DB_PORT`, `DB_USERNAME`, `DB_PASSWORD`, `DB_NAME`, `DB_SSL_MODE` consistency |
|
||||
| SQLite can't create file | check the `SQLITE_PATH` directory exists and is writable |
|
||||
| Port occupied | change `PORT` or `--port`, or stop the process holding the port |
|
||||
|
||||
## Admin Panel Won't Open or Is Blank
|
||||
|
||||
1. Confirm the Server is listening:
|
||||
|
||||
```bash
|
||||
curl -I http://127.0.0.1:3000
|
||||
```
|
||||
|
||||
2. Check that the browser access address matches the reverse proxy config.
|
||||
|
||||
## Default Account Can't Log In
|
||||
|
||||
The default account is `admin` / `12345678`. If the password was changed after first login, use the changed one.
|
||||
|
||||
Steps:
|
||||
|
||||
1. Confirm you're connected to the intended database — avoid `SQLITE_PATH` or `DB_HOST` / `DB_NAME` pointing at another environment.
|
||||
2. Check whether the Server log uses `sqlite` or `postgres`.
|
||||
3. In browser dev tools, confirm admin API requests carry the Session Cookie correctly.
|
||||
4. Clear browser cache and cookies, then log in again.
|
||||
|
||||
### Emergency Admin Password Reset
|
||||
|
||||
If you forget the `admin` password, reset it with the `reset-passwd` command (supports SQLite and PostgreSQL):
|
||||
|
||||
```bash
|
||||
go run main.go reset-passwd --user admin --password your-new-password
|
||||
```
|
||||
|
||||
With SQLite, stop the Server process first to avoid DB file lock conflicts. Without `--password`, the command generates a random password and prints it. After resetting, log in and change the password immediately.
|
||||
|
||||
## Agent Can't Register or Stays Offline
|
||||
|
||||
On the Agent node:
|
||||
|
||||
```bash
|
||||
curl -I http://your-server:3000
|
||||
```
|
||||
|
||||
Check Agent logs:
|
||||
|
||||
```bash
|
||||
journalctl -u openflare-agent -n 200 --no-pager
|
||||
```
|
||||
|
||||
Check the config file:
|
||||
|
||||
```bash
|
||||
sed -n '1,160p' /opt/openflare-agent/agent.json
|
||||
```
|
||||
|
||||
Confirm:
|
||||
|
||||
| Config | Description |
|
||||
| --- | --- |
|
||||
| `server_url` | must be a Server address reachable by the Agent node |
|
||||
| `agent_token` / `discovery_token` | at least one filled in |
|
||||
| `heartbeat_interval` | supports millisecond integer or Go duration string |
|
||||
| `request_timeout` | increase for slow networks |
|
||||
|
||||
If logs say the Token is invalid, prepare a new Token in the admin panel, update `agent.json`, then restart:
|
||||
|
||||
```bash
|
||||
systemctl restart openflare-agent
|
||||
```
|
||||
|
||||
## Node Didn't Apply the New Version After Release
|
||||
|
||||
Check in order:
|
||||
|
||||
1. Is the target version activated in the version page?
|
||||
2. Is the node online, and did the last heartbeat time update?
|
||||
3. Do the apply records show success, warning, or failure for the target version?
|
||||
4. Is the website config enabled? Disabled sites don't participate in release rendering.
|
||||
5. Do Agent logs show pull, validation, reload, or rollback messages?
|
||||
|
||||
View Agent logs:
|
||||
|
||||
```bash
|
||||
journalctl -u openflare-agent -f
|
||||
```
|
||||
|
||||
Note: once a target `version + checksum` fails to apply and rolls back, the Agent blocks retrying that target in local state. After fixing the config, republish to generate a new checksum, or activate an older version to roll back.
|
||||
|
||||
If this is the Agent's first config apply with no historical `nginx.conf` to roll back to, the failed target is still blocked, but the Agent enters a safe fallback runtime. The apply records and Agent logs will contain `fallback runtime started`; OpenResty only listens on port `80` and returns `503` with `OpenFlare: No Valid Configuration` for everything, while keeping the local `stub_status` health endpoint. After fixing the config and republishing a new version, the Agent overwrites the fallback config and resumes normal proxying.
|
||||
|
||||
## OpenResty Apply Failure
|
||||
|
||||
Common causes:
|
||||
|
||||
| Cause | Troubleshooting |
|
||||
| --- | --- |
|
||||
| Domain or server block conflict | check whether the same domain is used by multiple site configs |
|
||||
| Invalid upstream address | confirm all upstreams are `http://` or `https://` |
|
||||
| Multi-upstream format violates constraints | multi-upstream must be plain `scheme://host[:port]` |
|
||||
| Certificate missing or wrong path | check whether the domain is bound to a cert and the Agent cert dir is writable |
|
||||
| Port occupied | check local `80`, `443` ports |
|
||||
|
||||
OpenResty config validation:
|
||||
|
||||
```bash
|
||||
openresty -t -c /path/to/openflare/data/etc/nginx/nginx.conf
|
||||
```
|
||||
|
||||
OpenResty running state:
|
||||
|
||||
```bash
|
||||
ps aux | grep openresty
|
||||
```
|
||||
|
||||
The Agent's periodic health check probes the local `http://127.0.0.1:<openresty_observability_port>/openflare/stub_status` to judge OpenResty liveness — it does not repeatedly run `openresty -t`. If a node is marked unhealthy, first confirm that local observability port is listening; if `host not found in upstream` appears only during config apply, the failure comes from config validation or reload, not the periodic health probe.
|
||||
|
||||
Actual binary and main config paths follow `openresty_path` and `main_config_path` in `agent.json`.
|
||||
|
||||
## HTTPS Not Taking Effect
|
||||
|
||||
1. Confirm the certificate is uploaded or managed.
|
||||
2. Confirm the site config's domain is bound to a certificate.
|
||||
3. Confirm a new version was published and activated.
|
||||
4. Check whether the apply records succeeded.
|
||||
5. Inspect the certificate and status code with `curl`:
|
||||
|
||||
```bash
|
||||
curl -Iv https://your-domain
|
||||
```
|
||||
|
||||
Domains without a bound certificate are not auto-added to the HTTPS config — that's expected.
|
||||
|
||||
## Access Analytics Empty
|
||||
|
||||
1. Confirm the node successfully applied config including the observability Lua assets.
|
||||
2. Confirm OpenResty is running.
|
||||
3. Check Agent logs for observability collection or backfill failures.
|
||||
4. Check whether `openresty_observability_port` is occupied (default `18081`).
|
||||
5. Confirm the Server's DB cleanup policy hasn't deleted the relevant time window.
|
||||
|
||||
## Edge Cache Hit-Rate Anomalies
|
||||
|
||||
Access log cache three states: **hit** (HIT/STALE/REVALIDATED/UPDATING), **origin** (MISS/EXPIRED), **not cached** (BYPASS or empty — request didn't enter a cacheable path or response wasn't stored). Design: [Edge Cache Strategy Design](../design/edge-cache-design.md).
|
||||
|
||||
### Checklist
|
||||
|
||||
1. Global OpenResty cache enabled in **Performance Settings**.
|
||||
2. Site **Cache** enabled and policy matches the path (「Standard static assets」covers only built-in extensions, **not HTML/JSON**; `.js.map`'s extension is `map`, in the default table).
|
||||
3. Config version **published and activated**, node apply records succeeded (changing cache rules without publishing leaves nodes on old bypass logic).
|
||||
4. Request method is **GET** (non-GET is never cached).
|
||||
5. Origin doesn't return **`Set-Cookie`** for the target URL (if so, not written to the edge).
|
||||
6. Origin doesn't declare **`Cache-Control: private` / `no-store`** (shared caches won't store).
|
||||
7. Browser DevTools "Disable cache" only affects the browser; whether the edge HITs is judged by access log `cache_status`, not the Network panel.
|
||||
|
||||
### Common Misconceptions
|
||||
|
||||
| Symptom | Explanation |
|
||||
| --- | --- |
|
||||
| Everything「not cached」after login, never republished | old config bypassed session cookies; after upgrade you must republish node configs |
|
||||
| `/api/foo` or `/index.html` not cached under `static` | expected (extension not in the default cacheable table) |
|
||||
| HTML cross-user leakage after switching to `all` | origin didn't forbid shared caching; switch back to `static` or add `private`/`no-store` to dynamic responses |
|
||||
| URLs with `?v=` have low hit rates | default cache key includes the full `$request_uri`; different query = different object |
|
||||
| First MISS, second still MISS | check whether the origin sets `Set-Cookie`/`private` every time, or node disk/cache `inactive` is too short |
|
||||
|
||||
### Expected Behavior (aligned with Cloudflare defaults)
|
||||
|
||||
* A logged-in user accessing `/_app/**/*.js` static assets: **can HIT**.
|
||||
* Response with `Set-Cookie` or `private`: **not stored**.
|
||||
* Cacheable status codes without origin cache headers: use the default Edge TTL (e.g. ~120 min for 200).
|
||||
@@ -0,0 +1,172 @@
|
||||
# Tunnel & Intranet Penetration
|
||||
|
||||
You will learn: the design principles of OpenFlare's intranet penetration tunnels, core concepts (relay nodes and tunnel clients), and how to publish an intranet dev environment or private cloud service to a public domain step by step, securely and stably.
|
||||
|
||||
In many real development and ops scenarios, origin services run inside a LAN, on a local dev machine, or in a private VPC — with no public IP and no way to configure port mapping on the border firewall or router.
|
||||
|
||||
OpenFlare provides a complete **reverse-relay tunnel penetration** solution. You only initiate an outbound secure connection from the intranet to a public relay node — no inbound ports need to be configured — and public web traffic is routed into the intranet origin, while enjoying the gateway's automatic TLS certificate management and WAF protection.
|
||||
|
||||
---
|
||||
|
||||
## Core Concepts
|
||||
|
||||
Before using intranet penetration, get familiar with these components:
|
||||
|
||||
| Component | Description | Corresponding Entity |
|
||||
| --- | --- | --- |
|
||||
| **Relay node** | a traffic relay service deployed at the public edge; listens for the intranet client's long connections and bridges gateway Agent (OpenResty) and intranet traffic | `tunnel_relay` node guarded by `openflare-relay` |
|
||||
| **Tunnel** | a logical penetration client instance with a globally unique ID and an auth token, identifying one concrete intranet environment | `tunnel_client` node created in「Node Management」, assigned a dedicated Tunnel Token |
|
||||
| **Tunnel client** | a lightweight controller running in the intranet; auto-manages the underlying frpc tunnel child processes based on Server-dispatched config | `openflared` container or standalone binary deployed in the intranet |
|
||||
| **Tunnel upstream** | a special reverse proxy type in route rules. With this type, the gateway forwards public traffic to the local relay's Vhost port, eventually reaching the intranet origin | reverse proxy type configured in the「Rule Management」detail page, origin mode「Intranet Tunnel」with a bound Tunnel node |
|
||||
|
||||
---
|
||||
|
||||
## Recommended Order
|
||||
|
||||
To publish an intranet service to the public, follow this order:
|
||||
|
||||
1. Register and deploy at least one public **Relay node** and keep it online.
|
||||
2. Go to **「Node Management」**, create a node of type **Tunnel node (tunnel_client)**, and get the dedicated Token.
|
||||
3. Deploy and start the **tunnel client (OpenFlared)** on the intranet server.
|
||||
4. Confirm the Tunnel node's status shows「Online」in the admin panel.
|
||||
5. Add or edit a rule in **「Rule Management」**; in the「Reverse Proxy」tab choose origin mode **「Intranet Tunnel」**, bind the Tunnel node, and enter the intranet service port (e.g. `127.0.0.1:8080`).
|
||||
6. Publish and activate the new version.
|
||||
7. Access via the public domain to verify the tunnel link.
|
||||
|
||||
---
|
||||
|
||||
## Detailed Steps
|
||||
|
||||
### Step 1: Prepare a Relay Node
|
||||
|
||||
Intranet traffic needs a public relay node to transit. Before starting, make sure you have a usable relay server on the public network.
|
||||
|
||||
1. Log in to the admin panel, go to **「Node Management」**.
|
||||
2. Add a new node and set **Node Type** to **Relay node (tunnel_relay)**.
|
||||
3. After saving, copy the node's dedicated `agent_token`.
|
||||
4. Start `openflare-relay` on your public server. Docker quick run:
|
||||
|
||||
```bash
|
||||
docker run -d --name openflare-relay --restart unless-stopped \
|
||||
-p 7000:7000 \
|
||||
-e OPENFLARE_SERVER_URL=http://<your-Server-public-IP>:3000 \
|
||||
-e OPENFLARE_AGENT_TOKEN=<the-AgentToken-you-copied> \
|
||||
-v openflare-relay-data:/var/lib/openflare-relay \
|
||||
ghcr.io/rain-kl/openflare-relay:latest
|
||||
```
|
||||
|
||||
> [!IMPORTANT]
|
||||
> Open port `7000` (the frpc client connection control port, default `relay_bind_port`) in the cloud server's security group. If your Server and relay node are on the same machine, `OPENFLARE_SERVER_URL` here should point to the Server's public or intranet communication IP.
|
||||
|
||||
### Step 2: Create a Tunnel Node in the Admin Panel
|
||||
|
||||
1. Navigate to **「Node Management」** in the admin sidebar.
|
||||
2. Click **「Add Node」**; in the dialog choose node type **「Tunnel node (tunnel_client)」**.
|
||||
3. Fill in the node name and description, click save.
|
||||
4. Click into the Tunnel node's detail page; find the dedicated **Tunnel Token** and the one-click client deployment command.
|
||||
|
||||
### Step 3: Deploy the Intranet Client (OpenFlared)
|
||||
|
||||
Back on your intranet server, run the client with the copied deployment command.
|
||||
|
||||
#### Option A: Deploy with Docker (recommended)
|
||||
|
||||
The official `openflared` image bundles the supervisor daemon and the `frpc` runtime — out of the box, no extra dependencies:
|
||||
|
||||
```bash
|
||||
docker run -d --name openflared --restart unless-stopped \
|
||||
-e OPENFLARE_SERVER_URL=http://<your-Server-public-IP>:3000 \
|
||||
-e OPENFLARE_TUNNEL_TOKEN=<the-TunnelToken-you-copied> \
|
||||
-v openflared-data:/app/data \
|
||||
ghcr.io/rain-kl/openflared:latest
|
||||
```
|
||||
|
||||
#### Option B: Run the host binary manually
|
||||
|
||||
If Docker isn't convenient, download or build the `flared` binary yourself:
|
||||
|
||||
1. Create a `flared.json` config file next to the program on the intranet machine:
|
||||
```json
|
||||
{
|
||||
"server_url": "http://<your-Server-public-IP>:3000",
|
||||
"tunnel_token": "<the-TunnelToken-you-copied>",
|
||||
"frpc_path": "/usr/local/bin/frpc",
|
||||
"data_dir": "./data"
|
||||
}
|
||||
```
|
||||
2. Start it:
|
||||
```bash
|
||||
./flared -config ./flared.json
|
||||
```
|
||||
|
||||
#### Status Confirmation
|
||||
|
||||
After starting, the intranet client sends heartbeat syncs to the control plane over outbound connections:
|
||||
1. Refresh the **「Node Management」** list; the Tunnel node's status light should turn green **「Online」**.
|
||||
2. Click into the node detail page to see which public Relay nodes the intranet client is connected to.
|
||||
|
||||
### Step 4: Configure the Route and Bind the Tunnel Upstream
|
||||
|
||||
Now configure public reverse proxying and domain access for your intranet service:
|
||||
|
||||
1. First go to **「Website Management」->「Domain List」** and register the domain you want to expose.
|
||||
2. Go to **「Rule Management」**, click **「New Rule」** or edit an existing rule.
|
||||
3. In the **「Reverse Proxy」** tab, switch the **origin mode** to **「Intranet Tunnel」**.
|
||||
4. Select the online **Tunnel node** from the dropdown.
|
||||
5. Fill in the **intranet target address** (a local address/port reachable by the intranet client, e.g. `127.0.0.1:8080`) and **intranet protocol** (usually `http`).
|
||||
6. Configure other regular site options and save.
|
||||
|
||||
### Step 5: Publish and Apply
|
||||
|
||||
To let the gateway's OpenResty match and route domain traffic, publish a new config version:
|
||||
|
||||
1. Click **「Preview and Publish」** in the top-right nav; confirm the generated site config is correct.
|
||||
2. In the dialog, click **「Confirm Publish」**.
|
||||
3. The public-edge Agent now pulls the latest route: it forwards requests to the same-host `openflare-relay (frps)` vhost port.
|
||||
4. The intranet client `openflared (frpc)` receives the relayed packets and safely forwards them to the intranet `127.0.0.1:8080` service, returning the response along the same path.
|
||||
5. Visit the domain in a public browser to confirm the intranet service displays.
|
||||
|
||||
---
|
||||
|
||||
## Advanced Scenarios
|
||||
|
||||
### 1. One Tunnel, Multiple Services (multi-port mapping)
|
||||
|
||||
You don't need a separate `openflared` container for every intranet service.
|
||||
|
||||
To map multiple services in one intranet environment (e.g. `127.0.0.1:80` blog, `127.0.0.1:8080` API, `192.168.1.120:9000` intranet drive):
|
||||
1. Keep this one `openflared` client online.
|
||||
2. Create three separate website configs in the admin panel (each bound to its own public domain).
|
||||
3. Select **the same tunnel** as the origin mode for all three.
|
||||
4. Fill in the corresponding different ports or LAN IPs in each intranet target address (e.g. `127.0.0.1:80`, `127.0.0.1:8080`, `192.168.1.120:9000`).
|
||||
5. Publish and activate — one tunnel, many uses.
|
||||
|
||||
### 2. Gateway Security Features Stack Seamlessly
|
||||
|
||||
Because all public traffic first enters the public Agent node — HTTPS/TLS handshake and WAF engine interception happen there — then travels through the secure tunnel to the intranet:
|
||||
|
||||
Your intranet service needs **zero modification** to enjoy:
|
||||
* **One-click HTTPS**: select or apply an SSL certificate for the domain directly in the admin panel; data is encrypted end-to-end.
|
||||
* **Global/custom WAF protection**: enable SQL injection blocking, XSS injection defense, and malicious geo-IP blocking.
|
||||
* **Human challenge (CC PoW)**: one-click defense against malicious CC requests to intranet APIs.
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### 1. Tunnel shows「Offline」in the admin panel
|
||||
|
||||
* **Check the Token**: verify the `tunnel_token` in `flared` logs or env vars matches the one generated in the admin panel.
|
||||
* **Check network connectivity**: the intranet server must be able to reach the Server address over outbound connections.
|
||||
* **Relay firewall not open**: check that the relay node's public `7000` port (or custom `relay_bind_port`) is opened to the public in the security group.
|
||||
|
||||
### 2. Public domain returns 502 Bad Gateway / 504 Gateway Timeout
|
||||
|
||||
* **Intranet service not running**: confirm the service at the intranet target address is started and listening on the intranet server.
|
||||
* **Target address unreachable**: if the intranet address is `127.0.0.1:8080`, ensure the service runs on the same host as `openflared`; if it's a LAN IP `192.168.x.x`, test LAN reachability from inside the `openflared` container.
|
||||
* **Check node state and logs**: view the Tunnel node detail and「Apply Records」in the admin panel; frpc process errors are logged in detail in the `flared` logs on the intranet host.
|
||||
|
||||
### 3. Multi-relay network flapping or retry failures
|
||||
|
||||
* When the control plane is associated with multiple Relay nodes, `openflared` spawns a separate frpc supervisor per Relay and periodically pulls topology state from the control plane within the `sync_interval` configured in `flared.json` (default 30s).
|
||||
* If a relay node frequently drops due to network jitter, the system auto-triggers exponential backoff retries (initial 1s, cap 60s). You may see `frpc process missing, starting` in the host logs — that's normal process self-healing; it reconnects automatically after the network recovers.
|
||||
@@ -0,0 +1,49 @@
|
||||
# Uptime Kuma Monitoring Sync
|
||||
|
||||
You will learn: how to enable and configure the Uptime Kuma auto-sync integration, control the sync scope and heartbeat probe parameters for monitored sites, and the underlying principles of differential synchronization between OpenFlare and Uptime Kuma.
|
||||
|
||||
---
|
||||
|
||||
## Feature Overview
|
||||
|
||||
In edge multi-node operations, knowing the availability of each proxied site in time is critical. To avoid manually re-entering site information into a monitoring system, OpenFlare provides deep integration with the open-source monitoring service **Uptime Kuma**.
|
||||
|
||||
Once enabled, OpenFlare starts a background sync scheduler that automatically syncs the proxy sites configured in the admin panel as HTTP monitor tasks in Uptime Kuma. It supports scope filtering, differential attribute updates, and automatic cleanup of decommissioned sites.
|
||||
|
||||
---
|
||||
|
||||
## Step 1: Configure the Integration in System Settings
|
||||
|
||||
1. Log in to the admin panel, go to **「System Settings」** in the left navigation, select the **「OpenFlare」** tab, and configure the **「Uptime Kuma Integration」** section.
|
||||
2. Configure the following core connection parameters:
|
||||
* **Enabled**: Turn on the integration switch.
|
||||
* **Instance URL**: Your Uptime Kuma service address, e.g. `http://192.168.1.100:3001` or `https://kuma.example.com` (the protocol prefix `http://` or `https://` is required).
|
||||
* **Username** and **Password**: Credentials for a Uptime Kuma account with admin privileges, used for API authentication.
|
||||
|
||||
---
|
||||
|
||||
## Step 2: Control Monitor Scope and Heartbeat Parameters
|
||||
|
||||
In the integration panel you can finely control the monitor scope and probe behavior:
|
||||
|
||||
### 1. Monitor Scope
|
||||
* **All sites**: Default option. OpenFlare automatically syncs all **enabled** proxy route sites. When a new site is created and enabled, or an old site is disabled, the monitor list is updated automatically.
|
||||
* **Selected sites**: Only monitor specified sites. After selecting this mode, click the **「Select Monitored Sites」** dialog. Inside the dialog you can filter sites by search and check the ones you want. Sites that are unchecked or not checked will not be synced (and will be automatically cleaned up if they already exist).
|
||||
|
||||
### 2. Probe Frequency and Heartbeat Settings
|
||||
You can specify uniform probe parameters for auto-generated monitors:
|
||||
* **Sync Interval**: Frequency (minutes) of automatic differential sync, default `5` minutes. The control plane compares state with Uptime Kuma every 5 minutes.
|
||||
* **Heartbeat Interval**: Frequency (seconds) at which Uptime Kuma probes sites, default `60` seconds.
|
||||
* **Retry**: Maximum number of retries before a failed probe is judged Down, default `0`.
|
||||
* **Retry Interval**: Seconds to wait between retries, default `60` seconds.
|
||||
* **Request Timeout**: Seconds after which a probe request is judged timed out, default `48` seconds.
|
||||
|
||||
---
|
||||
|
||||
## Sync and Cleanup Mechanism
|
||||
|
||||
* **Dedicated tag isolation**: All auto-created monitors are bound with the `OpenFlare`-specific tag (purple-blue). The sync and cleanup routines only operate on monitors with this tag, and will not interfere with or damage other monitors you created manually in Uptime Kuma.
|
||||
* **Differential incremental sync**: The sync routine periodically compares monitor metadata. When a domain or heartbeat configuration change is detected, only a differential update is performed to avoid interrupting historical statistics; when a site is disabled or moved out of scope, it is automatically taken offline and cleaned up.
|
||||
|
||||
> [!TIP]
|
||||
> For details on the Socket.IO control flow, anti-pollution tag model, and differential comparison algorithm of Uptime Kuma monitoring sync, see [Uptime Kuma Sync Design](../design/kuma-design.md).
|
||||
@@ -0,0 +1,188 @@
|
||||
# WAF Auto IP Group Rule Syntax
|
||||
|
||||
Auto IP groups aggregate metrics per client IP from request logs, then use Expr expressions to decide whether to add an IP to the group list. Auto IP groups can be referenced by a WAF rule group's IP blocklist or allowlist; on config release, the Server only writes the IP group reference IDs into `waf_config.json` — IP group members are synced independently by the Agent to the local runtime file.
|
||||
|
||||
## Config Structure
|
||||
|
||||
An auto IP group config is a JSON object:
|
||||
|
||||
```json
|
||||
{
|
||||
"lookback": "1h",
|
||||
"rules": [
|
||||
{
|
||||
"name": "Single-IP high-frequency 404 scanning",
|
||||
"expr": "request_count > 100 && StatusRatio(404) >= 0.8"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Field reference:
|
||||
|
||||
| Field | Type | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `lookback` | string | lookback window duration in Go Duration syntax, e.g. `30m`, `1h`, `90m`. Defaults to `1h`, max 30 days. Compatible with the legacy `lookback_minutes` (integer minutes). |
|
||||
| `rules` | array | auto-rule list. If any rule matches, the IP enters the auto IP group list. |
|
||||
| `rules[].name` | string | rule name, only for UI display and error messages. |
|
||||
| `rules[].expr` | string | Expr expression, must return a boolean. |
|
||||
|
||||
## Execution Semantics
|
||||
|
||||
Auto IP groups first aggregate metrics per client IP, then run rule expressions against each IP:
|
||||
|
||||
1. The Server reads request logs from the last `lookback` window.
|
||||
2. Groups by normalized `remote_addr` IP.
|
||||
3. Computes per-IP metrics: request count, 404 count, direct-IP Host count, etc.
|
||||
4. Runs `rules[].expr` per IP.
|
||||
5. If an IP matches any rule, it is written into the auto IP group's `IP / IP segment` list.
|
||||
|
||||
Whether a Host is "accessed via IP" is judged by the `Host` field in request logs: if the Host is an IPv4 or IPv6 literal, e.g. `203.0.113.10`, `[2001:db8::10]`, `203.0.113.10:443`, it counts toward `ip_host_count`.
|
||||
|
||||
## Available Keywords
|
||||
|
||||
The expression can use these fields directly:
|
||||
|
||||
| Keyword | Type | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `ip` | string | the client IP currently being judged. |
|
||||
| `request_count` | number | the current IP's total requests in the lookback window. |
|
||||
| `status_404_count` | number | the current IP's requests returning 404 in the window. |
|
||||
| `status_404_ratio` | number | 404 ratio, computed as `status_404_count / request_count`. |
|
||||
| `ip_host_count` | number | requests where the current IP accessed via an IP-literal Host. |
|
||||
| `ip_host_ratio` | number | ratio of IP-address access, computed as `ip_host_count / request_count`. |
|
||||
| `client_error_count` | number | the current IP's requests returning 4xx. |
|
||||
| `server_error_count` | number | the current IP's requests returning 5xx. |
|
||||
| `last_seen_unix` | number | the current IP's last request Unix timestamp (seconds) in the window. |
|
||||
|
||||
Ratio fields are decimals between `0` and `1`. 80% is written `0.8`, 50% is `0.5`.
|
||||
|
||||
### Custom Status Code Matching
|
||||
|
||||
If the built-in `status_404_count` / `status_404_ratio` don't fit, use these built-in methods to match arbitrary status codes:
|
||||
|
||||
* **`StatusCount(code)`**: request count of the current IP returning the given status code (or class) in the window.
|
||||
* exact code: `StatusCount(403) > 10`
|
||||
* status class (`1xx`–`5xx`, case-insensitive): `StatusCount("4xx") > 50`
|
||||
* **`StatusRatio(code)`**: the ratio of the above count to the IP's total requests.
|
||||
* exact code: `StatusRatio(502) >= 0.5`
|
||||
* status class: `StatusRatio("4xx") >= 0.8`, `StatusRatio("5xx") >= 0.3`
|
||||
|
||||
A status class aggregates all codes in that hundred range, e.g. `"4xx"` covers 400–499, `"2xx"` covers 200–299.
|
||||
|
||||
## Common Expr Patterns
|
||||
|
||||
Auto IP groups use Expr syntax; the expression must return a boolean.
|
||||
|
||||
Common operators:
|
||||
|
||||
| Pattern | Purpose | Example |
|
||||
| --- | --- | --- |
|
||||
| `>`、`>=`、`<`、`<=` | numeric comparison | `request_count > 100` |
|
||||
| `==`、`!=` | equal / not equal | `ip != "127.0.0.1"` |
|
||||
| `&&` | and | `request_count > 100 && StatusRatio(404) >= 0.8` |
|
||||
| `||` | or | `StatusRatio(404) >= 0.8 || server_error_count > 20` |
|
||||
| `!` | negation | `!(ip == "127.0.0.1")` |
|
||||
| `in` | value in list | `ip in ["203.0.113.10", "198.51.100.20"]` |
|
||||
| `not in` | value not in list | `ip not in ["127.0.0.1"]` |
|
||||
| `()` | grouping precedence | `(request_count > 100 && StatusRatio(404) >= 0.8) || server_error_count > 50` |
|
||||
|
||||
## Built-in Presets
|
||||
|
||||
The admin panel ships two preset rules that can be added and then adjusted:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "Single-IP high-frequency 404 scanning",
|
||||
"expr": "request_count > 100 && StatusRatio(404) >= 0.8"
|
||||
}
|
||||
```
|
||||
|
||||
Meaning: a single IP has over 100 requests in the window and a 404 ratio of at least 80%.
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "Single-IP direct-access anomaly",
|
||||
"expr": "ip_host_count > 50 && ip_host_ratio > 0.5"
|
||||
}
|
||||
```
|
||||
|
||||
Meaning: a single IP accessed via IP-literal Host over 50 times, and that access ratio exceeds 50%.
|
||||
|
||||
## Examples
|
||||
|
||||
High-frequency 404 scanning:
|
||||
|
||||
```json
|
||||
{
|
||||
"lookback": "1h",
|
||||
"rules": [
|
||||
{
|
||||
"name": "High-frequency 404 scanning",
|
||||
"expr": "request_count > 100 && StatusRatio(404) >= 0.8"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
IP direct-access anomaly:
|
||||
|
||||
```json
|
||||
{
|
||||
"lookback": "30m",
|
||||
"rules": [
|
||||
{
|
||||
"name": "IP direct-access anomaly",
|
||||
"expr": "ip_host_count > 50 && ip_host_ratio > 0.5"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Catching both high 4xx and high 5xx:
|
||||
|
||||
```json
|
||||
{
|
||||
"lookback": "2h",
|
||||
"rules": [
|
||||
{
|
||||
"name": "Abnormal error rate",
|
||||
"expr": "(client_error_count > 80 && request_count > 100) || server_error_count > 30"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Using status-class syntax (equivalent thinking to `client_error_count` / `server_error_count`):
|
||||
|
||||
```json
|
||||
{
|
||||
"lookback": "2h",
|
||||
"rules": [
|
||||
{
|
||||
"name": "High 4xx or 5xx ratio",
|
||||
"expr": "request_count > 100 && (StatusRatio(\"4xx\") >= 0.8 || StatusRatio(\"5xx\") >= 0.3)"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Excluding trusted IPs:
|
||||
|
||||
```json
|
||||
{
|
||||
"lookback": "1h",
|
||||
"rules": [
|
||||
{
|
||||
"name": "404 scanning excluding trusted IPs",
|
||||
"expr": "ip not in [\"203.0.113.10\", \"198.51.100.20\"] && request_count > 100 && StatusRatio(404) >= 0.8"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Usage Tips
|
||||
|
||||
Start with a shorter lookback and higher thresholds to observe hits, then tune gradually. The admin IP group page supports clicking **Test Rule** before saving to directly view the IPs hit in the current window; when the auto IP group actually runs, it overwrites the group's IP list. To keep certain addresses long-term, put them in a manual IP group and reference both the manual and auto groups in the WAF rule group.
|
||||
|
||||
Auto IP group updates don't require republishing a config version. Online Agents receive the changed IP groups via WebSocket and update the local `waf_ip_groups.json`; when the WebSocket is unavailable, the Agent reports the local IP group checksum in the next heartbeat and the Server only returns the groups whose checksums differ.
|
||||
@@ -0,0 +1,31 @@
|
||||
# WAF Security Protection
|
||||
|
||||
OpenFlare WAF orchestrates rules with a visual directed acyclic graph. When creating a rule you only enter a name; the system creates the default "start → pass" graph and enters the editor.
|
||||
|
||||
## Nodes and Connections
|
||||
|
||||
- **Start**: unique per rule; enters the graph along `next`.
|
||||
- **Pass**: ends the current rule; if the route has more rules, execution continues.
|
||||
- **Block**: immediately terminates the request with the configured status code and HTML response.
|
||||
- **IP match**: configure an IP, CIDR, or IP group; connect `true` / `false` respectively.
|
||||
- **Geo match**: branch by country or ISO 3166-2 first-level division code; the country list shows both localized names and codes; divisions are searchable by country name, division name, or code. Country and City MMDB are provided by disk files (Docker images COPY them to the default path; bare binaries download from the configured URL on first startup) and update on the configured cycle. When City MMDB is unavailable, it is treated as no match.
|
||||
- **PoW**: takes over the request when the challenge is incomplete; after verification passes, continues along `next`.
|
||||
|
||||
The server rejects cycles, dangling outlets, unreachable nodes, duplicate port connections, and invalid configs. Save carries the `revision` obtained at page load; on a 409 conflict, reload to avoid overwriting others' changes.
|
||||
|
||||
Select a normal node or edge, then click the delete button at the canvas top-right, or press Delete/Backspace. Deleting a node also deletes associated edges; the single "start" and "pass" nodes cannot be deleted. Dragging a node only records the final coordinates on release; the canvas state is not rebuilt repeatedly during movement.
|
||||
|
||||
The right node property panel is hidden by default and shows after clicking a node; it auto-collapses when clicking an edge or blank canvas.
|
||||
|
||||
The orchestration area defaults to a compact height and smaller initial zoom; you can still zoom freely with the wheel or canvas Controls.
|
||||
|
||||
## Binding and Effect
|
||||
|
||||
Enabled global rules always execute first; route-bound custom rules execute strictly in list order. After adjusting the order, you must publish a config version for the rule topology to take effect with the OpenResty reload.
|
||||
|
||||
IP group members are dynamic resources. The Agent checks the checksum every 5 seconds and updates in-memory snapshots across workers on change — no rule republish or reload needed. Manual, subscription, and auto IP groups can all be referenced by IP-match nodes. A single full IP group runtime snapshot is capped at 20 MiB; exceeding the cap makes publish or sync return an error and keeps using the previous valid snapshot.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> When upgrading from the legacy fixed allow/block-list, geo, or PoW forms, rule graphs reset to "start → pass" and old policy fields are not migrated. Re-orchestrate and verify each rule before publishing a new version.
|
||||
|
||||
Architecture, graph validation, and failure rollback details: [WAF Orchestration Rule Design](../design/waf-orchestration-design.md).
|
||||
@@ -0,0 +1,52 @@
|
||||
# Zone Domain Migration and Release Acceptance
|
||||
|
||||
When migrating from the legacy `managed_domains` / inline domain columns of reverse proxy routes to the Zone + Zone Domain model, data import and table structure upgrades are both completed by the **automatic goose migration at Server startup** — no separate import command is needed.
|
||||
|
||||
## What Happens During Upgrade
|
||||
|
||||
When starting (or rolling-upgrading) a Server version that includes the Zone rework, **no manual command is required**; `migrator.Migrate()` automatically:
|
||||
|
||||
1. Applies goose SQL: creates `of_zones` / `of_zone_domains` (if they do not yet exist).
|
||||
2. **Automatically imports** the legacy route domain columns (and `of_managed_domains` when routes have no domains) as Zone / Zone Domains, binding `proxy_route_id` / `cert_id` (registering root domains via public suffix list parsing).
|
||||
3. Continues goose SQL: drops the redundant domain/certificate columns from `of_managed_domains` and `of_proxy_routes`.
|
||||
|
||||
The import is idempotent: existing domains are skipped or have their route binding back-filled.
|
||||
|
||||
**If historical data cannot be parsed (conflicting domains, invalid root domains, missing certificates, etc.), startup fails.** Fix the data or restore a backup and start again to retry.
|
||||
|
||||
## Recommended Actions
|
||||
|
||||
### 1. Back Up Before Upgrading
|
||||
|
||||
```bash
|
||||
# PostgreSQL example
|
||||
pg_dump "$DATABASE_URL" > openflare-pre-zone-$(date +%Y%m%d).sql
|
||||
|
||||
# Or copy the backup volume / snapshot; for SQLite, copy the database file in the data directory
|
||||
```
|
||||
|
||||
Optional: note down the current **active config version number** and checksum in the admin panel for config rollback comparison.
|
||||
|
||||
### 2. Upgrade and Start the Server
|
||||
|
||||
Deploy the new version and start it. Watch the goose success messages in the startup log; if "Zone migration failed (N conflicts)" appears, fix the source data according to the conflicts listed in the log and restart.
|
||||
|
||||
### 3. Post-Upgrade Checks
|
||||
|
||||
1. Admin panel **Websites** `/websites`: check whether Zone root domains and domain counts are reasonable.
|
||||
2. Zone details: domains, certificates, associated route IDs.
|
||||
3. **Reverse proxy routes**: domain bindings come from Zone Domains, not legacy hand-written fields.
|
||||
|
||||
### 4. Config Preview and Release
|
||||
|
||||
1. Review the config diff / preview in the admin panel.
|
||||
2. Verify **per route**: `server_name` set, certificate paths, WAF Route ID, Pages references.
|
||||
3. **Allow** the redundant `domain` / `domains` / `cert_ids` on routes in old snapshot JSON to disappear.
|
||||
4. **Do not allow** data-plane semantic changes.
|
||||
5. After the preview passes, release it; if needed, activate the pre-upgrade version in config versions for config rollback. For database rollback, use the pre-upgrade backup (down migrations do not backfill business domain data).
|
||||
|
||||
## Related Docs
|
||||
|
||||
* [Zone & Domain Resource Design](../design/zone-design.md)
|
||||
* [Create a Reverse Proxy Config](./proxy-config.md)
|
||||
* [Publish First Configuration](./first-site.md)
|
||||
@@ -0,0 +1,38 @@
|
||||
---
|
||||
layout: home
|
||||
|
||||
hero:
|
||||
name: OpenFlare
|
||||
text: Open-source CDN Orchestration & Edge Security Platform
|
||||
tagline: Supports reverse proxy, centralized configuration synchronization, Pages static hosting, secure intranet penetration (Tunnels), dynamic WAF protection, and anti-CC challenges.
|
||||
actions:
|
||||
- theme: brand
|
||||
text: Quick Start
|
||||
link: /en/guide/quick-start
|
||||
- theme: alt
|
||||
text: Design Boundaries
|
||||
link: /en/design/
|
||||
- theme: alt
|
||||
text: GitHub
|
||||
link: https://github.com/Rain-kl/OpenFlare
|
||||
|
||||
features:
|
||||
- icon: 🛰️
|
||||
title: Centralized Config Sync
|
||||
details: Sync configurations across all nodes in real time via WebSockets and heartbeats with sub-second hot reload. Instantly retrieve alerts and statuses.
|
||||
- icon: 🌐
|
||||
title: Distributed CDN Orchestration
|
||||
details: Orchestrate scattered and independent OpenResty nodes into a highly collaborative CDN fleet with multi-upstream load balancing.
|
||||
- icon: 📄
|
||||
title: Pages Static Hosting
|
||||
details: Upload pre-built frontend zip assets directly; edge nodes pull, extract, and serve them locally at high performance with API proxying.
|
||||
- icon: 🚇
|
||||
title: Secure Intranet Penetration (Tunnels)
|
||||
details: An open-source alternative to Cloudflare Tunnels. Expose local intranet services securely to the public network without a public IP or open inbound ports.
|
||||
- icon: 🛡️
|
||||
title: Edge WAF Protection
|
||||
details: Dynamic WAF rules with differential syncing of IP groups to Lua shared memory without Nginx reloads, plus country-level regional access control.
|
||||
- icon: 🧩
|
||||
title: Anti-CC & Bot Defense (PoW)
|
||||
details: Built-in high-performance client-side cryptographic Proof of Work challenges (similar to Turnstile) to intercept botnets and scrapers at the edge.
|
||||
---
|
||||
@@ -0,0 +1,167 @@
|
||||
# Commands & Scripts
|
||||
|
||||
You will learn: common start, build, test, install, and uninstall commands for the OpenFlare Server, admin frontend, Agent, Relay, OpenFlared, Swagger, and the docs site.
|
||||
|
||||
> All commands run at the **repo root** unless noted otherwise.
|
||||
|
||||
## Server
|
||||
|
||||
Source startup:
|
||||
|
||||
```bash
|
||||
cp config.example.yaml config.yaml
|
||||
go run main.go all
|
||||
```
|
||||
|
||||
Split processes:
|
||||
|
||||
```bash
|
||||
go run main.go api # HTTP API only
|
||||
go run main.go worker # Asynq Worker only
|
||||
go run main.go scheduler # scheduled tasks only
|
||||
```
|
||||
|
||||
Build the binary:
|
||||
|
||||
```bash
|
||||
make build-backend
|
||||
# output: bin/openflare-server
|
||||
```
|
||||
|
||||
Tests:
|
||||
|
||||
```bash
|
||||
GOCACHE=/tmp/openflare-go-cache go test ./...
|
||||
```
|
||||
|
||||
Quality gate:
|
||||
|
||||
```bash
|
||||
make code-check
|
||||
```
|
||||
|
||||
Auto-format backend Go source (organize imports) and frontend source:
|
||||
|
||||
```bash
|
||||
make format
|
||||
```
|
||||
|
||||
This command uses `goimports` to organize backend Go imports and the repo-pinned Prettier version to format `frontend/` source; build artifacts, dependencies, public static assets, and lock files are ignored.
|
||||
|
||||
## Frontend
|
||||
|
||||
Dev:
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
pnpm install
|
||||
pnpm dev
|
||||
```
|
||||
|
||||
Build the embedded artifact (hosted by the Go Server):
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
pnpm build:embed
|
||||
# or at repo root: make build-embedded
|
||||
```
|
||||
|
||||
Checks:
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
pnpm lint
|
||||
pnpm tsc --noEmit --jsx preserve
|
||||
pnpm check:i18n
|
||||
```
|
||||
|
||||
## Agent
|
||||
|
||||
Source run:
|
||||
|
||||
```bash
|
||||
go run ./cmd/agent -config /path/to/agent.json
|
||||
```
|
||||
|
||||
Build:
|
||||
|
||||
```bash
|
||||
make build-agent
|
||||
# or: go build -o bin/openflare-agent ./cmd/agent
|
||||
```
|
||||
|
||||
Tests:
|
||||
|
||||
```bash
|
||||
GOCACHE=/tmp/openflare-go-cache go test ./internal/apps/agent/...
|
||||
```
|
||||
|
||||
## Relay
|
||||
|
||||
Source run:
|
||||
|
||||
```bash
|
||||
go run ./cmd/relay -config /path/to/relay.json
|
||||
```
|
||||
|
||||
Build:
|
||||
|
||||
```bash
|
||||
make build-relay
|
||||
# or: go build -o bin/openflare-relay ./cmd/relay
|
||||
```
|
||||
|
||||
## OpenFlared (Tunnel client)
|
||||
|
||||
Source run:
|
||||
|
||||
```bash
|
||||
go run ./cmd/flared -config /path/to/flared.json
|
||||
```
|
||||
|
||||
Build:
|
||||
|
||||
```bash
|
||||
make build-flared
|
||||
# or: go build -o bin/flared ./cmd/flared
|
||||
```
|
||||
|
||||
## Install Agent
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
|
||||
--server-url http://your-server:3000 \
|
||||
--agent-token YOUR_AGENT_TOKEN
|
||||
```
|
||||
|
||||
## Uninstall Agent
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/uninstall-agent.sh | bash
|
||||
```
|
||||
|
||||
## Swagger
|
||||
|
||||
Regenerate the Swagger docs:
|
||||
|
||||
```bash
|
||||
make swagger
|
||||
```
|
||||
|
||||
Access: `http://localhost:3000/api/swagger/index.html` (default `api_prefix` `/api`; only mounted in non-production)
|
||||
|
||||
## Docs
|
||||
|
||||
Local preview:
|
||||
|
||||
```bash
|
||||
cd docs
|
||||
pnpm dev
|
||||
```
|
||||
|
||||
Build:
|
||||
|
||||
```bash
|
||||
cd docs
|
||||
pnpm build
|
||||
```
|
||||
@@ -0,0 +1,407 @@
|
||||
# Configuration
|
||||
|
||||
You will learn: which config sources OpenFlare Server, frontend build, Agent, Relay, and OpenFlared support, what config fields and env vars exist, and their defaults and behavior.
|
||||
|
||||
This document summarizes all configuration items supported by the current OpenFlare version.
|
||||
|
||||
---
|
||||
|
||||
## Config Sources
|
||||
|
||||
### 1. Server Config Sources
|
||||
- **Config file**: reads `config.yaml` in the same directory at startup by default (overridable via the `CONFIG_PATH` env var).
|
||||
- **Env vars**: every field in the config file can be overridden by an `UPPER_SNAKE_CASE` env var (env vars take precedence over `config.yaml`).
|
||||
- **System runtime config**: stored in the `w_system_configs` table in the relational DB. These can be hot-updated and take effect dynamically via the admin UI or system API.
|
||||
|
||||
### 2. Agent / Relay / OpenFlared Config Sources
|
||||
- **CLI args**: `-config` specifies the config file (JSON format).
|
||||
- **Config file**: e.g. `agent.json`, `relay.json`, `flared.json`.
|
||||
- **Override env vars**: specific env vars can override connection addresses and Token credentials in the config file.
|
||||
|
||||
---
|
||||
|
||||
## Config File Locations
|
||||
|
||||
| Component | Default Location | Notes |
|
||||
| --- | --- | --- |
|
||||
| Server config file | `./config.yaml` | overridable via `CONFIG_PATH` |
|
||||
| Server SQLite DB | `openflare.db` | overridable via `database.sqlite_path` / `SQLITE_PATH` |
|
||||
| Agent config file | `./agent.json` | overridable via `-config` |
|
||||
| One-click install Agent config | `/opt/openflare-agent/agent.json` | default path generated by the install script |
|
||||
| Agent data dir | `data` next to the config file | overridable via `data_dir` |
|
||||
| Relay config file | `./relay.json` | overridable via `-config` |
|
||||
| One-click install Relay config | `/opt/openflare-relay/relay.json` | default path generated by the install script |
|
||||
| Client config file | `./flared.json` | overridable via `-config` |
|
||||
| One-click install Client config | `/opt/openflared/flared.json` | default path generated by the install script |
|
||||
|
||||
---
|
||||
|
||||
## Server CLI Args
|
||||
|
||||
```bash
|
||||
# specify a config file when starting the Server
|
||||
CONFIG_PATH=/path/to/custom-config.yaml ./openflare-server all
|
||||
```
|
||||
|
||||
Supported sub-service commands (fused/single-process mode):
|
||||
- `all`: starts all services in one process (API + Worker + Scheduler, default).
|
||||
- `api`: starts only the API service for the admin panel and node communication.
|
||||
- `worker`: starts only the background-task Worker service.
|
||||
- `scheduler`: starts only the scheduled-task Scheduler service.
|
||||
|
||||
---
|
||||
|
||||
## Server Env Vars vs Config File
|
||||
|
||||
All Server core base config is defined in `config.yaml`, and every field supports env-var overrides (env vars take precedence over the YAML file).
|
||||
|
||||
### 1. App Basic Config (`app:`)
|
||||
| YAML path | Override env var | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `app.app_name` | `APP_NAME` | application identifier name | `openflare` |
|
||||
| `app.env` | `APP_ENV` | runtime env (`development` / `testing` / `production`) | `production` |
|
||||
| `app.addr` | `APP_ADDR` | service listen address and port | `:3000` |
|
||||
| `app.node_id` | `APP_NODE_ID` | Snowflake node ID (0-1023); must be unique in multi-instance deploys | `1` |
|
||||
| `app.api_prefix` | `APP_API_PREFIX` | route prefix for admin panel and API | `/api` |
|
||||
| `app.graceful_shutdown_timeout` | `APP_GRACEFUL_SHUTDOWN_TIMEOUT` | graceful shutdown wait timeout (seconds) | `30` |
|
||||
| `app.session_cookie_name` | `APP_SESSION_COOKIE_NAME` | session cookie name | `openflare_session_id` |
|
||||
| `app.session_secret` | `APP_SESSION_SECRET` | session signature secret; **must be a random long string in production** | none (random) |
|
||||
| `app.session_domain` | `APP_SESSION_DOMAIN` | shared-session cookie scope domain | empty |
|
||||
| `app.session_age` | `APP_SESSION_AGE` | browser session lifetime (seconds) | `86400` (24h) |
|
||||
| `app.session_http_only` | `APP_SESSION_HTTP_ONLY` | enable the cookie's HttpOnly attribute | `false` |
|
||||
| `app.session_secure` | `APP_SESSION_SECURE` | enable the cookie's Secure attribute (HTTPS) | `false` |
|
||||
|
||||
### 2. Relational DB Config (`database:`)
|
||||
| YAML path | Override env var | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `database.enabled` | `DB_ENABLED` | enable PostgreSQL; `false` falls back to SQLite | `true` |
|
||||
| `database.sqlite_path` | `SQLITE_PATH` | SQLite DB file path when PostgreSQL disabled | `openflare.db` |
|
||||
| `database.host` | `DB_HOST` | PostgreSQL host | `127.0.0.1` |
|
||||
| `database.port` | `DB_PORT` | PostgreSQL port | `5432` |
|
||||
| `database.username` | `DB_USERNAME` | PostgreSQL username | `openflare` |
|
||||
| `database.password` | `DB_PASSWORD` | PostgreSQL password | `replace-with-strong-password` |
|
||||
| `database.database` | `DB_NAME` | PostgreSQL database name | `openflare` |
|
||||
| `database.ssl_mode` | `DB_SSL_MODE` | PostgreSQL SSL mode | `disable` |
|
||||
| `database.time_zone` | `DB_TIMEZONE` | DB session timezone | `UTC` |
|
||||
| `database.log_level` | `DB_LOG_LEVEL` | GORM SQL log level (`info` / `warn` / `error` / `silent`) | `info` |
|
||||
| `database.max_idle_conn` | `DB_MAX_IDLE_CONN` | connection pool max idle | `16` |
|
||||
| `database.max_open_conn` | `DB_MAX_OPEN_CONN` | connection pool max open | `128` |
|
||||
|
||||
### 3. Redis Config (`redis:`)
|
||||
| YAML path | Override env var | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `redis.enabled` | `REDIS_ENABLED` | enable Redis. **The async queue and sync depend on it; must be on** | `true` |
|
||||
| `redis.addrs` | `REDIS_ADDR` | Redis single/cluster address array (env var sets a single address) | `["127.0.0.1:6379"]` |
|
||||
| `redis.username` | `REDIS_USERNAME` | Redis username (if any) | empty |
|
||||
| `redis.password` | `REDIS_PASSWORD` | Redis password | empty |
|
||||
| `redis.db` | `REDIS_DB` | Redis logical DB number | `0` |
|
||||
| `redis.key_prefix` | `REDIS_KEY_PREFIX` | system key prefix in Redis | `openflare:` |
|
||||
| `redis.pool_size` | `REDIS_POOL_SIZE` | Redis connection pool size | `100` |
|
||||
| `redis.maint_notifications` | `REDIS_MAINT_NOTIFICATIONS` | enable Redis maintenance-notifications negotiation at startup; keep off when compatibility is unclear; restart needed after change | `false` |
|
||||
|
||||
### 4. ClickHouse Config (`clickhouse:`)
|
||||
|
||||
> **Note**: these are OpenFlare **client** connection params. For ClickHouse **server** small-host tuning: curl `performance.xml` to `./config/clickhouse/`, then mount it as a single file at the container's `config.d/performance.xml` — see [Start the Server](../deployment/server.md).
|
||||
|
||||
| YAML path | Override env var | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `clickhouse.enabled` | `CLICKHOUSE_ENABLED` | enable ClickHouse. **Node metrics and access logs are written here at scale**. Off by default: missing config or `false` disables it and the primary DB handles logs/metrics; explicit `true` or setting `CLICKHOUSE_HOST` enables it | `false` |
|
||||
| `clickhouse.hosts` | `CLICKHOUSE_HOST` | ClickHouse cluster address array (env var sets a single address) | `["127.0.0.1:9000"]` |
|
||||
| `clickhouse.username` | `CLICKHOUSE_USERNAME` | ClickHouse username | `default` |
|
||||
| `clickhouse.password` | `CLICKHOUSE_PASSWORD` | ClickHouse password | `replace-with-clickhouse-password` |
|
||||
| `clickhouse.database` | `CLICKHOUSE_NAME` | ClickHouse database name | `openflare` |
|
||||
| `clickhouse.max_idle_conn` | - | client idle connections (low by default for small hosts) | `8` |
|
||||
| `clickhouse.max_open_conn` | - | client max open connections | `16` |
|
||||
| `clickhouse.conn_max_lifetime` | - | connection max lifetime (seconds) | `3600` |
|
||||
| `clickhouse.dial_timeout` | - | dial timeout (seconds) | `5` |
|
||||
| `clickhouse.block_buffer_size` | - | native-protocol block buffer rows | `32` |
|
||||
|
||||
### 5. System Log Config (`log:`)
|
||||
| YAML path | Override env var | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `log.level` | `LOG_LEVEL` | global log level (`debug` / `info` / `warn` / `error` / `fatal`) | `info` |
|
||||
| `log.format` | `LOG_FORMAT` | log format (`console` readable / `json` structured) | `console` |
|
||||
| `log.output` | `LOG_OUTPUT` | log output (`stdout` / `file`) | `stdout` |
|
||||
| `log.file_path` | - | log file path when output is file | `./logs/app.log` |
|
||||
| `log.max_size` | - | max single log file size (MB); auto-rotates beyond | `100` |
|
||||
| `log.max_age` | - | max days to keep rotated log files | `30` |
|
||||
|
||||
### 6. Async Task Worker Queue Config (`worker:`)
|
||||
| YAML path | Override env var | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `worker.concurrency` | `WORKER_CONCURRENCY` | max concurrent tasks consumed by the background Worker | `20` |
|
||||
| `worker.strict_priority`| `WORKER_STRICT_PRIORITY` | strictly assign consumer threads by queue priority (else weighted round-robin) | `false` |
|
||||
|
||||
### 7. Tracing OpenTelemetry Config (`otel:`)
|
||||
| YAML path | Override env var | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `otel.sampling_rate` | `OTEL_SAMPLING_RATE` | global OTel sampling rate. `0.0` no sampling, `1.0` full tracing | `0.0` |
|
||||
| `otel.tracer_name` | `OTEL_TRACER_NAME` | global OTel tracer instance name | `github.com/Rain-kl/OpenFlare` |
|
||||
|
||||
---
|
||||
|
||||
## Runtime System Config (SystemConfig)
|
||||
|
||||
These items are stored in the `w_system_configs` table. Changes actively notify Redis cache invalidation for dynamic hot-update; admins manage them via the admin UI.
|
||||
|
||||
### 1. Base & Business Runtime Config
|
||||
| Key | Type | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `site_name` | `string` | admin platform display name | `OpenFlare` |
|
||||
| `server_address` | `string` | public access address of the admin console, used to assemble OAuth callbacks and download links | empty |
|
||||
| `password_login_enabled` | `bool` | allow admin login with normal username/password | `true` |
|
||||
| `registration_enabled` | `bool` | allow self-service new-user registration (off by default; root invites or distributes) | `false` |
|
||||
| `password_register_enabled` | `bool` | allow direct email/password registration on the frontend | `false` |
|
||||
| `oidc_login_enabled` | `bool` | enable OIDC (SSO) third-party passwordless login | `true` |
|
||||
| `max_api_keys_per_user` | `int` | max API keys (API Tokens) per admin user | `5` |
|
||||
| `login_session_ttl_hours` | `int` | user session lifetime in the browser cookie (hours). 0 = clear on browser close | `0` |
|
||||
| `upload_allowed_extensions` | `string` | allowed upload file extensions (comma-separated; empty = unlimited) | `jpg,png,webp` |
|
||||
| `file_access_whitelist` | `json` | file business types allowed for public download/access without login (JSON array) | `["avatar"]` |
|
||||
| `disk_cache_max_size_mb` | `int` | platform local disk cache max storage (MB) | `100` |
|
||||
| `disk_cache_ttl_minutes` | `int` | local disk cache object default TTL (minutes) | `60` |
|
||||
| `disk_cache_lru_enabled` | `bool` | use LRU eviction when local disk cache space is low | `true` |
|
||||
| `update_upstream_repository` | `string` | GitHub repo for self-update detection | `Rain-kl/OpenFlare` |
|
||||
| `storage_config` | `json` | object-storage structured config (JSON): local disk and AWS S3-compatible storage | local-storage mode |
|
||||
| `relay_frps_web_ui_enabled` | `bool` | enable the embedded frps traffic-monitoring Web UI on relay nodes | `false` |
|
||||
| `relay_frps_web_ui_port` | `int` | host port the relay frps monitoring panel listens on | `17500` |
|
||||
| `search_engine_indexing_enabled` | `bool` | allow search engines to crawl/index the site | `false` |
|
||||
| `menu_display_config` | `string` | menu display structured config (JSON string, format `{url: enabled}`) | `{}` |
|
||||
| `pages_max_package_size_mb` | `int` | Pages deployment package upload size cap (MiB, range 1–2048) | `100` |
|
||||
| `pages_max_history_count` | `int` | max historical deployments kept per Pages project (0 = unlimited) | `20` |
|
||||
|
||||
### 2. Human Verification (PoW Captcha)
|
||||
| Key | Type | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `cap_login_enabled` | `bool` | require local PoW anti-brute-force human verification on the login page | `false` |
|
||||
| `cap_auto_solve` | `bool` | auto-start background PoW computation on page load (no manual click) | `true` |
|
||||
| `cap_challenge_count` | `int` | number of PoW challenges required. More = longer compute (recommended 1–5) | `1` |
|
||||
| `cap_challenge_difficulty`| `int`| PoW hash prefix-match difficulty per challenge. Recommended 3-5 | `4` |
|
||||
| `cap_challenge_size` | `int` | challenge salt length | `32` |
|
||||
| `cap_challenge_ttl_seconds`| `int`| max valid time to submit the computed challenge (seconds); auto-invalidates on timeout | `600` |
|
||||
| `cap_token_ttl_seconds` | `int` | validity of the login credential after solving (seconds); must log in within the window | `1200` |
|
||||
|
||||
### 3. SMTP Email Config
|
||||
| Key | Type | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `smtp_host` | `string` | SMTP server address | empty |
|
||||
| `smtp_port` | `int` | SMTP port (usually 465 SSL or 587 STARTTLS) | `465` |
|
||||
| `smtp_username` | `string` | SMTP account email | empty |
|
||||
| `smtp_password` | `string` | SMTP account auth password/cert key (encrypted on save, never echoed) | empty |
|
||||
| `email_login_verification_enabled` | `bool` | send a one-time 6-digit code for second-factor auth on email login | `false` |
|
||||
| `email_register_verification_enabled` | `bool` | force email verification with a registration code for self-service registration | `false` |
|
||||
|
||||
### 4. Node & Agent Ops Runtime
|
||||
| Key | Type | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `agent_discovery_token` | `string` | global discovery Token for one-click first-time node registration | none (auto-generated on first visit) |
|
||||
| `agent_heartbeat_interval`| `int` | standard heartbeat interval dispatched to all Agents (ms) | `3000` (3s) |
|
||||
| `agent_websocket_upgrade_enabled` | `bool` | allow Agents to upgrade to a persistent WebSocket connection after HTTP heartbeat handshake | `true` |
|
||||
| `node_offline_threshold` | `int` | no-response threshold (ms) after which a node is marked offline in the admin panel | `60000` (60s) |
|
||||
| `agent_update_repo` | `string` | Release repo source for Agent self-binary updates | `Rain-kl/OpenFlare` |
|
||||
| `geoip_provider` | `string` | GeoIP provider, e.g. `maxmind`, for WAF geo analysis | `ipinfo` |
|
||||
|
||||
### 5. Uptime Kuma Monitoring Sync
|
||||
| Key | Type | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `uptime_kuma_enabled` | `bool` | auto-sync generated Uptime Kuma HTTP monitors after config release/activation | `false` |
|
||||
| `uptime_kuma_url` | `string` | Uptime Kuma instance access URL (with port and path) | empty |
|
||||
| `uptime_kuma_username` | `string` | Uptime Kuma admin username for sync API auth | empty |
|
||||
| `uptime_kuma_password` | `string` | Uptime Kuma login password (encrypted on save, never echoed) | empty |
|
||||
| `uptime_kuma_monitor_scope`| `string` | route scope for auto-generated monitors (`all` sites or `selected`) | `all` |
|
||||
| `uptime_kuma_selected_sites`| `string` | selected proxied-site Site Name list to monitor (comma-separated) | empty |
|
||||
| `uptime_kuma_sync_interval`| `int` | differential scan/calibration sync frequency to the Uptime Kuma instance (minutes) | `5` |
|
||||
| `uptime_kuma_interval` | `int` | HTTP GET probe period of generated monitors (seconds) | `60` |
|
||||
| `uptime_kuma_retry` | `int` | max reconnect retries after probe connection failures | `0` |
|
||||
| `uptime_kuma_retry_interval`| `int` | pause between failed reconnect retries (seconds) | `60` |
|
||||
| `uptime_kuma_timeout` | `int` | timeout for an HTTP GET monitor request (seconds) | `48` |
|
||||
|
||||
### 6. OpenResty Core Main Config & Rendering Options
|
||||
| Key | Type | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `openresty_default_server_return_status` | `int` | status code returned for requests hitting no matching route by default | `421` |
|
||||
| `openresty_worker_processes` | `string` | nginx `worker_processes`; fixed integer or `auto` | `auto` |
|
||||
| `openresty_worker_connections` | `int` | nginx `worker_connections` per-process max connections | `4096` |
|
||||
| `openresty_worker_rlimit_nofile` | `int` | nginx `worker_rlimit_nofile` max open file descriptors | `65535` |
|
||||
| `openresty_events_use` | `string` | event polling engine (e.g. `epoll` preferred on Linux) | `epoll` |
|
||||
| `openresty_events_multi_accept_enabled` | `bool` | accept all pending connection handshakes in one batch | `true` |
|
||||
| `openresty_keepalive_timeout` | `int` | nginx `keepalive_timeout` (seconds) | `20` |
|
||||
| `openresty_keepalive_requests` | `int` | max requests per reused TCP connection | `1000` |
|
||||
| `openresty_client_header_timeout` | `int` | read timeout for the client Request Header (seconds) | `15` |
|
||||
| `openresty_client_body_timeout` | `int` | read timeout for the client Request Body (seconds) | `15` |
|
||||
| `openresty_client_max_body_size` | `string` | max client request Body size, with a unit like `10m`/`50m` | `64m` |
|
||||
| `openresty_large_client_header_buffers` | `string` | buffers for oversized request headers (e.g. `4 16k`) | `4 16k` |
|
||||
| `openresty_send_timeout` | `int` | max interval for sending Response data to the client (seconds) | `30` |
|
||||
| `openresty_resolvers` | `string` | DNS resolver addresses/params for dynamic name resolution on nodes | empty |
|
||||
| `openresty_proxy_connect_timeout` | `int` | TCP handshake timeout to the origin (seconds) | `3` |
|
||||
| `openresty_proxy_send_timeout` | `int` | max interval for writing request data to the origin (seconds) | `60` |
|
||||
| `openresty_proxy_read_timeout` | `int` | max wait for origin response data (seconds) | `60` |
|
||||
| `openresty_websocket_enabled` | `bool` | auto-load WebSocket-supporting globals and headers in the HTTP section | `true` |
|
||||
| `openresty_http3_enabled` | `bool` | render HTTP/3 QUIC dual-stack listen capability in generated nginx listens | `true` |
|
||||
| `openresty_proxy_request_buffering_enabled`| `bool` | fully buffer the client Request Body before forwarding to the origin | `false` |
|
||||
| `openresty_proxy_buffering_enabled` | `bool` | buffer large origin Response data before forwarding to the user | `true` |
|
||||
| `openresty_proxy_buffers` | `string` | nginx proxy response buffer count/size (e.g. `16 16k`) | `16 16k` |
|
||||
| `openresty_proxy_buffer_size` | `string` | buffer for origin Response Header | `8k` |
|
||||
| `openresty_proxy_busy_buffers_size` | `string` | Busy-state buffer cap when response stream is oversized | `64k` |
|
||||
| `openresty_gzip_enabled` | `bool` | enable gzip real-time compression on eligible content | `true` |
|
||||
| `openresty_gzip_min_length` | `int` | file size threshold for gzip; below it, skip to save CPU | `1024` (1KB) |
|
||||
| `openresty_gzip_comp_level` | `int` | gzip level 1-9; higher = more compression, more CPU | `5` |
|
||||
| `openresty_cache_enabled` | `bool` | initialize the proxy cache region (Proxy Cache Path) in global config | `false` |
|
||||
| `openresty_cache_path` | `string` | proxy cache temp physical dir on the node | `__OPENFLARE_PROXY_CACHE_PATH__` |
|
||||
| `openresty_cache_levels` | `string` | proxy cache directory tree level layout | `1:2` |
|
||||
| `openresty_cache_inactive` | `string` | time after which an unaccessed cache file is invalidated from disk | `30m` |
|
||||
| `openresty_cache_max_size` | `string` | max disk quota for the proxy cache region on a node | `1g` |
|
||||
| `openresty_cache_key_template` | `string` | default proxy cache key template | `$scheme$host$request_uri` |
|
||||
| `openresty_cache_lock_enabled` | `bool` | queue/lock origin connections on high-concurrency cache misses for the same expired resource | `true` |
|
||||
| `openresty_cache_lock_timeout` | `string` | max queue wait for the proxy cache lock | `5s` |
|
||||
| `openresty_cache_use_stale` | `string` | serve stale cache on specific origin errors (500/502/504 etc.) | `error timeout updating http_500 http_502 http_503 http_504` |
|
||||
| `openresty_default_limit_conn_per_server` | `int` | default concurrent-connection cap per server when a site has no config; `0` = off | `0` |
|
||||
| `openresty_default_limit_conn_per_ip` | `int` | default per-IP concurrent cap when a site has no config; `0` = off | `0` |
|
||||
| `openresty_default_limit_rate` | `string` | default per-request bandwidth when a site has no config (e.g. `512k`); empty = off | empty |
|
||||
| `openresty_default_limit_req_per_ip` | `string` | default per-IP request rate limit when a site has no config (e.g. `10r/s`, `100r/m`); empty = off | empty |
|
||||
| `openresty_main_config_template` | `string` | fully rewrite the OpenResty nginx.conf skeleton template | empty (built-in default skeleton) |
|
||||
|
||||
### 7. Origin Error Page
|
||||
|
||||
Global origin error page config; written into the config version snapshot and distributed to edge Agents on release/rollback. Only affects **reverse proxy** routes; Pages static routes are unaffected. Admin entry:「Website Management → Error Page」. Design: [Origin Error Page Design](../design/origin-error-page.md).
|
||||
|
||||
| Key | Type | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `origin_error_page_enabled` | `bool` | enable the global origin error page. When on, matching status codes from origin/gateway are replaced by custom/default HTML with the **HTTP status kept**; when off, no directives are generated and pass-through resumes. Requires a config release to take effect | `true` |
|
||||
| `origin_error_page_get_only` | `bool` | only apply to **GET** requests. When on, only matching GET error statuses return the custom error page; POST/PUT and other methods **pass through the origin response** (original status and body unchanged) | `false` |
|
||||
| `origin_error_page_status_codes` | `json` | status code tag JSON array triggering the error page. Supports single codes (e.g. `522`) and closed ranges (e.g. `500-599`); single codes and range endpoints must be in **400–599**, with `lo ≤ hi`. The expanded result must not be empty when enabled | `["500-599"]` |
|
||||
| `origin_error_page_html` | `string` | custom error page HTML. Empty uses the built-in OpenFlare default template (minimal white); supports `{{status}}` (matches the HTTP status) and `{{host}}` (request Host). Max **256 KiB** (bytes). Don't embed untrusted third-party scripts | empty |
|
||||
|
||||
---
|
||||
|
||||
### 8. Log Database
|
||||
|
||||
Runtime config after log-store decoupling: the log primary DB is managed by the「Switch Log Database」task (internal/protected keys, forbidden for manual admin edits); access-log retention days are set per storage DB in business config; performance metrics (CPU/memory/disk/network) decay fast and use a shared short retention independent of the access-log retention config.
|
||||
|
||||
| Key | Type | Description | Default |
|
||||
| --- | --- | --- | --- |
|
||||
| `log_database` | `string` | current log primary DB (`postgres` / `sqlite` / `clickhouse`). **Internal protected key**: only written by the「Switch Log Database」migration task; admins can't create/modify manually | follows the primary DB (`postgres` when PostgreSQL enabled, else `sqlite`; `clickhouse` preferred when ClickHouse enabled) |
|
||||
| `log_db_migration` | `string` | log migration freeze marker (`migrating` or empty). **Internal protected key**: only the migration task writes it; while set, log writes return 503「log DB migrating, not writable」 | empty |
|
||||
| `log_retention_days_postgres` | `int` | access-log retention days in the PostgreSQL log DB (expired logs deleted by the daily garbage-collection task) | `30` |
|
||||
| `log_retention_days_sqlite` | `int` | access-log retention days in the SQLite log DB | `30` |
|
||||
| `log_retention_days_clickhouse` | `int` | access-log retention days in the ClickHouse log DB | `30` |
|
||||
| `metric_retention_days` | `int` | performance metric (CPU/memory/disk/network) retention days; shared short retention across the three DBs (independent of access-log retention) | `3` |
|
||||
|
||||
---
|
||||
|
||||
## Frontend Build Env Vars
|
||||
|
||||
| Env var | Purpose | Default |
|
||||
| --- | --- | --- |
|
||||
| `WAVELET_BACKEND_URL` | backend address for server-side rendering and dev proxy | `http://localhost:3000` |
|
||||
| `NEXT_PUBLIC_WAVELET_BACKEND_URL` | backend address for browser API requests; empty = same-origin | empty |
|
||||
| `NEXT_PUBLIC_APP_VERSION` | displayed frontend version | `dev` |
|
||||
|
||||
---
|
||||
|
||||
## Agent Env Vars
|
||||
|
||||
| Env var | Purpose | Default |
|
||||
| --- | --- | --- |
|
||||
| `LOG_LEVEL` | Agent log level | `info` |
|
||||
| `OPENFLARE_SERVER_URL` | control-plane address, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_AGENT_TOKEN` | node-specific auth Token, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_DISCOVERY_TOKEN` | first-time auto-registration Token, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_NODE_NAME` | node name, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_NODE_IP` | node IP, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_DATA_DIR` | Agent data dir, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_OPENRESTY_PATH` | OpenResty binary path, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_PAGES_DIR` | Pages static deployment dir, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_HEARTBEAT_INTERVAL` | heartbeat interval, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_REQUEST_TIMEOUT` | request timeout, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_OPENRESTY_OBSERVABILITY_PORT` | local observability port, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_MMDB_PATH` | WAF GeoIP mmdb path, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_MMDB_UPDATE_INTERVAL` | WAF GeoIP mmdb update interval, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_MMDB_DOWNLOAD_URL` | WAF GeoIP mmdb download URL, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_CITY_MMDB_PATH` | WAF City MMDB path, overrides `agent.json` | empty |
|
||||
| `OPENFLARE_CITY_MMDB_DOWNLOAD_URL` | WAF City MMDB download URL, overrides `agent.json` | empty |
|
||||
|
||||
---
|
||||
|
||||
## Agent CLI Args and Config Fields
|
||||
|
||||
### CLI Args
|
||||
- `-config`: Agent config file path, default `./agent.json`.
|
||||
|
||||
### Config File Fields (agent.json)
|
||||
| Field | Purpose | Required | Default/Behavior |
|
||||
| --- | --- | --- | --- |
|
||||
| `server_url` | control-plane address | yes | none |
|
||||
| `agent_token` | node-specific auth Token | one of two with `discovery_token` | empty |
|
||||
| `discovery_token` | global Token for first-time auto-registration | one of two with `agent_token` | empty |
|
||||
| `node_name` | node name | no | hostname automatically |
|
||||
| `node_ip` | node IP | no | auto-detected, prefers the public egress IP; falls back to local NIC detection on failure |
|
||||
| `openresty_path` | OpenResty binary path | no | `openresty` |
|
||||
| `openresty_observability_port` | local observability & OpenResty health-check port | no | `18081` |
|
||||
| `data_dir` | Agent data dir | no | `data` next to the config file |
|
||||
| `main_config_path` | OpenResty main config write path | no | `data_dir/etc/nginx/nginx.conf` |
|
||||
| `route_config_path` | route config write path | no | `data_dir/etc/nginx/conf.d/openflare_routes.conf` |
|
||||
| `access_log_path` | OpenResty access log path | no | `data_dir/var/log/openflare/access.log` |
|
||||
| `cert_dir` | certificate write dir | no | `data_dir/etc/nginx/certs` |
|
||||
| `openresty_cert_dir` | certificate dir read by the OpenResty config | no | same as `cert_dir` |
|
||||
| `lua_dir` | Lua scripts & static assets write dir | no | `data_dir/etc/nginx/lua` |
|
||||
| `openresty_lua_dir` | Lua dir read by the OpenResty config | no | same as `lua_dir` |
|
||||
| `runtime_config_dir` | Agent runtime config write dir, e.g. `pow_config.json` | no | `data_dir/etc/openflare` |
|
||||
| `pages_dir` | Pages deployment package extraction & current deploy dir | no | `data_dir/var/lib/openflare/pages` |
|
||||
| `mmdb_path` | WAF GeoIP mmdb file path | no | `data_dir/etc/openflare/GeoLite2-Country.mmdb` |
|
||||
| `city_mmdb_path` | WAF City MMDB file path | no | `data_dir/etc/openflare/GeoLite2-City.mmdb` |
|
||||
| `mmdb_update_interval` | WAF GeoIP mmdb update interval | no | `86400000` ms (24h) |
|
||||
| `mmdb_download_url` | WAF GeoIP mmdb periodic update URL | no | GeoLite2 Country update URL; first download when the disk file is missing (Docker images COPY default-path files) |
|
||||
| `city_mmdb_download_url` | WAF City MMDB periodic update URL | no | GeoLite2 City update URL; first download when the disk file is missing (Docker images COPY default-path files) |
|
||||
| `observability_buffer_path` | observability backfill buffer file path | no | `data_dir/var/lib/openflare/observability-buffer.json` |
|
||||
| `observability_replay_minutes` | auto-backfill recent observability window (minutes) | no | `60` |
|
||||
| `state_path` | Agent local state file path | no | `data_dir/var/lib/openflare/agent-state.json` |
|
||||
| `heartbeat_interval` | heartbeat interval | no | `3000` ms |
|
||||
| `request_timeout` | HTTP request timeout | no | `10000` ms |
|
||||
|
||||
---
|
||||
|
||||
## Relay Env Vars and Config Fields
|
||||
|
||||
### Env Vars
|
||||
- `LOG_LEVEL`: Relay log level, default `info`.
|
||||
- Supports `OPENFLARE_SERVER_URL`, `OPENFLARE_AGENT_TOKEN`, `OPENFLARE_DISCOVERY_TOKEN`, `OPENFLARE_NODE_NAME`, `OPENFLARE_NODE_IP`, `OPENFLARE_DATA_DIR`, `OPENFLARE_FRPS_PATH` overrides.
|
||||
|
||||
### CLI Args
|
||||
- `-config`: Relay config file path, default `./relay.json`.
|
||||
|
||||
### Config File Fields (relay.json)
|
||||
| Field | Purpose | Required | Default/Behavior |
|
||||
| --- | --- | --- | --- |
|
||||
| `server_url` | control-plane address | yes | none |
|
||||
| `agent_token` | relay node-specific auth Token | one of two with `discovery_token` | empty |
|
||||
| `discovery_token` | global Token for first-time auto-registration | one of two with `agent_token` | empty |
|
||||
| `node_name` | node name | no | hostname automatically |
|
||||
| `node_ip` | relay node IP for receiving tunnel traffic | no | auto-detected, prefers the public egress IP; falls back to NIC detection on failure |
|
||||
| `frps_path` | frps binary path | no | `frps` (found on system PATH) |
|
||||
| `data_dir` | Relay runtime data dir | no | `data` next to the config file |
|
||||
| `state_path` | Relay local state file path | no | `data_dir/relay-state.json` |
|
||||
| `heartbeat_interval` | heartbeat interval | no | `10000` ms |
|
||||
| `request_timeout` | HTTP request timeout | no | `10000` ms |
|
||||
|
||||
---
|
||||
|
||||
## OpenFlared (Client) Env Vars and Config Fields
|
||||
|
||||
### Env Vars
|
||||
- `LOG_LEVEL`: Client log level, default `info`.
|
||||
- Supports `OPENFLARE_SERVER_URL`, `OPENFLARE_TUNNEL_TOKEN`, `OPENFLARE_DATA_DIR`, `OPENFLARE_FRPC_PATH` overrides.
|
||||
|
||||
### CLI Args
|
||||
- `-config`: Client config file path, default `./flared.json`.
|
||||
|
||||
### Config File Fields (flared.json)
|
||||
| Field | Purpose | Required | Default/Behavior |
|
||||
| --- | --- | --- | --- |
|
||||
| `server_url` | control-plane address | yes | none |
|
||||
| `tunnel_token` | tunnel-specific auth Token | yes | none |
|
||||
| `frpc_path` | frpc binary path | no | `frpc` (found on system PATH) |
|
||||
| `data_dir` | Client runtime data dir | no | `data` next to the config file |
|
||||
| `state_path` | Client local state file path | no | `data_dir/flared-state.json` |
|
||||
| `heartbeat_interval` | heartbeat interval | no | `10000` ms |
|
||||
| `sync_interval` | config pull/sync interval | no | `30000` ms |
|
||||
| `request_timeout` | HTTP request timeout | no | `10000` ms |
|
||||
@@ -0,0 +1,11 @@
|
||||
# Reference
|
||||
|
||||
You will learn: which information counts as stable reference material, and where to look up config, commands, API, and repository structure.
|
||||
|
||||
This section consolidates stable runtime, interface, and repository-level information for quick reference during deployment, integration, and troubleshooting.
|
||||
|
||||
| Page | Content |
|
||||
| --- | --- |
|
||||
| [Configuration](./configuration.md) | Server env vars, CLI args, runtime Options, and Agent config fields |
|
||||
| [Commands & Scripts](./cli.md) | common start, build, test, install, and uninstall commands |
|
||||
| [Repository Structure](../design/index.md#repository-structure) | monorepo directory responsibilities and layering (`main.go`, `cmd/`, `internal/apps/`, `frontend/`, etc.) |
|
||||
+17
-17
@@ -12,35 +12,35 @@
|
||||
2. 点击右上角的 **「导入证书」**。
|
||||
3. 填写配置信息:
|
||||
* **证书名称**:输入一个易于识别的别名(如 `my-domain-cert`)。
|
||||
* **证书内容 (PEM)**:复制并粘贴 PEM 格式 of 证书公钥内容(通常以 `-----BEGIN CERTIFICATE-----` 开头)。
|
||||
* **证书私钥 (KEY)**:复制并粘贴证书的私钥内容(通常以 `-----BEGIN PRIVATE KEY-----` 或 `-----BEGIN RSA PRIVATE KEY-----` 开头)。
|
||||
* **证书内容(PEM)**:复制并粘贴 PEM 格式的证书公钥内容(通常以 `-----BEGIN CERTIFICATE-----` 开头)。
|
||||
* **证书私钥(KEY)**:复制并粘贴证书的私钥内容(通常以 `-----BEGIN PRIVATE KEY-----` 或 `-----BEGIN RSA PRIVATE KEY-----` 开头)。
|
||||
4. 点击 **「保存」**。导入成功后,该证书即可在配置域名时直接绑定使用。
|
||||
|
||||
---
|
||||
|
||||
## 方式二:自动申请与到期自动续签 (ACME)
|
||||
## 方式二:自动申请与到期自动续签(ACME)
|
||||
|
||||
OpenFlare 内置了 ACME 客户端并对接了 **Asynq 异步任务队列**。通过配合云解析服务商的 DNS API,系统能自动完成 DNS-01 挑战(Challenge)校验,并向 CA(默认 Let's Encrypt)申请通配符/单域名证书,并在**到期前 30 天自动触发后台秒级续签**。
|
||||
OpenFlare 内置了 ACME 客户端并对接了 **Asynq 异步任务队列**。通过配合云解析服务商的 DNS API,系统能自动完成 DNS-01 挑战(Challenge)校验,并向 CA(默认 Let's Encrypt)申请通配符/单域名证书,并在**到期前 7 天自动触发续签**。
|
||||
|
||||
### 第一步:在 Cloudflare 申请 DNS API Token
|
||||
|
||||
为了使 OpenFlare 能够自动在你的域名下添加 TXT 记录以完成 DNS 校验,你需要准备一个具有特定权限的 Cloudflare API Token。
|
||||
|
||||
> [!IMPORTANT]
|
||||
> 安全起见,**强烈建议使用限定权限的 API Token**,而非全局 API Key (Global API Key)。
|
||||
> 安全起见,**强烈建议使用限定权限的 API Token**,而非全局 API Key(Global API Key)。
|
||||
|
||||
1. 登录 [Cloudflare 控制台](https://dash.cloudflare.com/)。
|
||||
2. 点击右上角的用户头像,选择 **「我的个人资料 (My Profile)」**。
|
||||
3. 在左侧菜单中选择 **「API 令牌 (API Tokens)」**,然后点击 **「创建令牌 (Create Token)」**。
|
||||
4. 找到 **「编辑区域 DNS (Edit Zone DNS)」** 模板,点击 **「使用模板 (Use template)」**。
|
||||
2. 点击右上角的用户头像,选择 **「我的个人资料(My Profile)」**。
|
||||
3. 在左侧菜单中选择 **「API 令牌(API Tokens)」**,然后点击 **「创建令牌(Create Token)」**。
|
||||
4. 找到 **「编辑区域 DNS(Edit Zone DNS)」** 模板,点击 **「使用模板(Use template)」**。
|
||||
5. 配置令牌权限与范围(保持默认或根据实际情况限定):
|
||||
* **权限 (Permissions)**:
|
||||
* `区域 (Zone)` - `DNS` - `编辑 (Edit)` (必须,ACME 写入 TXT 记录用)
|
||||
* `区域 (Zone)` - `区域 (Zone)` - `读取 (Read)` (必须,用于列出和检索区域 ID)
|
||||
* **区域资源 (Zone Resources)**:
|
||||
* 选择 **「包括 (Include)」** -> **「所有区域 (All zones)」**,或者选择 **「特定区域 (Specific zone)」** 并指向你托管的特定域名。
|
||||
6. 点击 **「继续以转到摘要 (Continue to summary)」**,确认无误后点击 **「创建令牌 (Create Token)」**。
|
||||
7. 复制生成的 **API 令牌 (Token)** 字符串。*注意:该令牌仅展示一次,请妥善保存*。
|
||||
* **权限(Permissions)**:
|
||||
* `区域(Zone)` - `DNS` - `编辑(Edit)`(必须,ACME 写入 TXT 记录用)
|
||||
* `区域(Zone)` - `区域(Zone)` - `读取(Read)`(必须,用于列出和检索区域 ID)
|
||||
* **区域资源(Zone Resources)**:
|
||||
* 选择 **「包括(Include)」** -> **「所有区域(All zones)」**,或者选择 **「特定区域(Specific zone)」** 并指向你托管的特定域名。
|
||||
6. 点击 **「继续以转到摘要(Continue to summary)」**,确认无误后点击 **「创建令牌(Create Token)」**。
|
||||
7. 复制生成的 **API 令牌(Token)** 字符串。该令牌仅展示一次,请妥善保存。
|
||||
|
||||
### 第二步:在控制端添加 DNS 账号
|
||||
|
||||
@@ -49,7 +49,7 @@ OpenFlare 内置了 ACME 客户端并对接了 **Asynq 异步任务队列**。
|
||||
3. 填写配置信息:
|
||||
* **账号名称**:如 `cloudflare-main`。
|
||||
* **DNS 服务商**:选择 `Cloudflare`。
|
||||
* **API Token**:填入刚刚在 Cloudflare 复制的 API 令牌(该值在入库时会自动加密存储,保障安全)。
|
||||
* **API Token**:填入刚刚在 Cloudflare 复制的 API 令牌(该值在入库时会自动加密存储)。
|
||||
4. 点击 **「保存」**。
|
||||
|
||||
### 第三步:提交证书申请任务
|
||||
@@ -65,4 +65,4 @@ OpenFlare 内置了 ACME 客户端并对接了 **Asynq 异步任务队列**。
|
||||
### 第四步:查看申请进度与续期状态
|
||||
|
||||
- **查看实时进度**:保存后,系统会向 Asynq 队列投递单证书续期/申请任务(`of_ssl_single_renew`)。你可以进入管理后台的任务或节点日志页面,实时查看每一步(添加 TXT 记录、DNS 记录全球生效探测、ACME 验证、证书颁发落地等)的详细日志。
|
||||
- **自动续期**:所有通过 ACME 申请的证书都会被系统自动托管。后台的 Scheduler 每日会自动扫描证书有效期,在到期前 30 天自动通过异步任务触发续签,无需任何手动维护。
|
||||
- **自动续期**:所有通过 ACME 申请的证书都会被系统自动托管。后台的 Scheduler 每日会自动扫描证书有效期,在到期前 7 天自动通过异步任务触发续签,无需任何手动维护。
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# 引用与致谢
|
||||
|
||||
OpenFlare 本质上是一个方案整合项目, 在设计与实现过程中借鉴了众多开源项目的优秀理念、架构设计和技术实现。以下是 OpenFlare 在核心底层引擎、安全防护机制以及前后端系统框架等方面所引用的关键开源项目,以及对这些项目及其社区的感谢。
|
||||
OpenFlare 在设计与实现过程中借鉴了众多开源项目的优秀理念、架构设计和技术实现。以下是 OpenFlare 在核心底层引擎、安全防护机制以及前后端系统框架等方面所引用的关键开源项目,在此对这些项目及其社区表示感谢。
|
||||
|
||||
---
|
||||
|
||||
@@ -9,14 +9,14 @@ OpenFlare 本质上是一个方案整合项目, 在设计与实现过程中借
|
||||
* **在 OpenFlare 中的作用**:作为全局数据面(Data Plane)的边缘网关。所有的公网 Web 流量均首先由 OpenResty 接收,在此处进行高并发的 HTTPS 握手、WAF 安全规则比对、防 CC 人机验证,并最终执行反向代理转发。
|
||||
* **项目链接**:[OpenResty 官网](https://openresty.org/)
|
||||
|
||||
### 2. FRP (Fast Reverse Proxy)
|
||||
### 2. FRP(Fast Reverse Proxy)
|
||||
* **项目定位**:高性能的反向代理应用,专注于内网穿透。
|
||||
* **在 OpenFlare 中的作用**:作为内网穿透子系统的底层隧道引擎。中继端管理器 `openflare-relay` 负责守护和调度 `frps` 引擎,而内网客户端 `openflared` 则负责在本地自动生成 TOML 配置并守护多路复用 `frpc` 子进程。
|
||||
* **项目链接**:[fatedier/frp (GitHub)](https://github.com/fatedier/frp)
|
||||
|
||||
---
|
||||
|
||||
### 3. Anubis (PoW 方案)
|
||||
### 3. Anubis(PoW 方案)
|
||||
* **项目定位**:基于工作量证明(Proof of Work)的轻量级人机验证防护方案。
|
||||
* **在 OpenFlare 中的作用**:为网关 WAF 提供了核心的**无感防 CC 人机挑战**能力。
|
||||
|
||||
|
||||
@@ -23,15 +23,15 @@ OpenFlare 的发布链路以“不可变配置版本”为核心。你在管理
|
||||
|
||||
为了快速验证,我们首先部署一个最基础的 HTTP 反代站点:
|
||||
|
||||
1. 登录控制面板,进入左侧导航 **「网站管理」->「域名列表」**,点击 **「新增网站」**。
|
||||
1. 登录控制面板,进入左侧导航 **「网站管理」->「域名列表」**,点击 **「新增 Zone」**。
|
||||
2. 填写域名配置:
|
||||
* **域名**:输入用于测试的域名(如 `first.example.com`)。
|
||||
* **绑定证书**:选择不绑定证书(作为 HTTP 快速验证)。
|
||||
* 点击保存,完成网站登记。
|
||||
3. 进入左侧导航 **「规则管理」**,点击 **「新增规则」**:
|
||||
* 点击保存,完成域名登记。
|
||||
3. 进入左侧导航 **「规则管理」**,点击 **「新建规则」**:
|
||||
* **规则名称**:输入简易标识(如 `first-app-route`)。
|
||||
* **域名匹配**:填入你的测试域名(如 `first.example.com`)。
|
||||
* 在下方 **「反向代理」** 选项卡中,配置 **源站类型** 为「标准反代」 (Direct)。
|
||||
* 在下方 **「反向代理」** 选项卡中,配置 **回源方式** 为「直连上游」。
|
||||
* **上游地址**:填写后端服务地址(如测试专用的 `http://httpbin.org`)。
|
||||
* 点击保存创建规则。
|
||||
|
||||
@@ -45,8 +45,8 @@ OpenFlare 的发布链路以“不可变配置版本”为核心。你在管理
|
||||
|
||||
新增的网站配置仍保存在 Server 的数据库中,处于草稿状态,需要通过发布版本分发到数据面:
|
||||
|
||||
1. 点击控制面板右上角的 **「配置预览」** 按钮,系统会展示本次新增路由的物理配置文件 Diff 差异。
|
||||
2. 确认渲染出的配置内容正确无误后,点击 **「发布并激活」**。
|
||||
1. 点击控制面板右上角的 **「预览并发布」** 按钮,系统会展示本次新增路由的物理配置文件 Diff 差异。
|
||||
2. 确认渲染出的配置内容正确无误后,点击 **「确认发布」**。
|
||||
3. 控制面将生成一个唯一的配置版本号(格式为 `YYYYMMDD-NNN`)。
|
||||
|
||||
---
|
||||
|
||||
+4
-4
@@ -1,6 +1,6 @@
|
||||
# 指南
|
||||
|
||||
你会学到:OpenFlare 文档如何组织、首次运行应该读哪些页面,以及部署、使用、排查和开发分别从哪里开始。
|
||||
你会学到:OpenFlare 文档如何组织、首次运行应该读哪些页面,以及部署、使用、排查分别从哪里开始。
|
||||
|
||||
OpenFlare 是一套自托管的 OpenResty 控制面。它把反向代理网站配置、配置版本发布、Agent 节点同步、TLS 证书和基础观测放到一个管理端中,适合单团队或单组织管理多台代理节点。
|
||||
|
||||
@@ -17,8 +17,8 @@ OpenFlare 是一套自托管的 OpenResty 控制面。它把反向代理网站
|
||||
7. [WAF 安全防护使用](./waf-usage.md):配置 WAF 规则组,掌握 IP 黑白名单、自动/订阅 IP 组、地域限制与 PoW CC 防护。
|
||||
8. [WAF 自动 IP 组语法](./waf-ip-group-expr.md):编写自动 IP 组 Expr 规则,了解关键字含义和预设规则。
|
||||
9. [Uptime Kuma 监控同步](./uptime-kuma.md):配置并使用 Uptime Kuma 自动差分同步和监控范围控制。
|
||||
10. [SSO 登录配置](./sso.md):配置 GitHub 或 OIDC 实现第三方单点登录 (SSO) 接入。
|
||||
11. [故障排查](./troubleshooting.md):按症状排查登录、数据库、节点同步、OpenResty、边缘缓存命中与前端构建问题。
|
||||
10. [SSO 登录配置](./sso.md):配置 OIDC 实现第三方单点登录(SSO)接入。
|
||||
11. [故障排查](./troubleshooting.md):按症状排查登录、数据库、节点同步、OpenResty 与边缘缓存命中问题。
|
||||
12. [引用与致谢](./credits.md):查看系统依赖的优秀开源项目与社区致谢清单。
|
||||
|
||||
## 按角色查找
|
||||
@@ -36,7 +36,7 @@ OpenFlare 是一套自托管的 OpenResty 控制面。它把反向代理网站
|
||||
| 自动同步监测站点状态 | [Uptime Kuma 监控同步](./uptime-kuma.md) |
|
||||
| 接入或重装节点 Agent | [接入 Agent](../deployment/agent.md) |
|
||||
| 从源码启动 Server | [启动 Server](../deployment/server.md) |
|
||||
| 配置 GitHub 或 OIDC 登录 | [SSO 登录配置](./sso.md) |
|
||||
| 配置 OIDC 登录 | [SSO 登录配置](./sso.md) |
|
||||
| 升级 Server 或 Agent | [升级与维护](../deployment/upgrade.md) |
|
||||
| 理解架构和发布模型 | [系统架构](../design/architecture.md) 与 [Agent 与发布模型](../design/agent-design.md) |
|
||||
| 查看开源引用与致谢 | [引用与致谢](./credits.md) |
|
||||
|
||||
@@ -59,12 +59,12 @@ GitHub 来源仅支持公开 `github.com` 仓库。填写:
|
||||
|
||||
两种选择都可手动 **「检查更新」** 和 **「同步并发布」**。区别如下:
|
||||
|
||||
* **latest**:可设置 5~1440 分钟检查间隔,默认 60 分钟;自动更新默认关闭。开启后,scanner 发现新 revision 才会异步同步并发布。
|
||||
* **latest**:可设置 5~1440 分钟检查间隔,默认 1440 分钟(24 小时);自动更新默认关闭。开启后,scanner 发现新 revision 才会异步同步并发布。
|
||||
* **tag**:只支持管理员手动检查和同步,不参与定时 scanner。
|
||||
|
||||
“检查更新”只解析 Release/asset 并更新版本游标,不下载部署包;“同步并发布”才会下载、校验、创建或复用 deployment 并激活。如果同一个 Release 下的 asset 被替换,来源会进入 **「需要确认」**,必须确认页面显示的精确 revision 后才能发布,避免静默覆盖。
|
||||
|
||||
GitHub Release 在这里是预构建产物源,不等同于连接代码仓库自动构建。未来仓库集成会使用独立的 `git_repository` 来源和 Server build executor,再把构建产物送入同一部署管线。
|
||||
GitHub Release 来源只导入预构建产物,不执行仓库源码构建。
|
||||
|
||||
### 4. 切换或删除来源
|
||||
|
||||
|
||||
+10
-10
@@ -9,7 +9,7 @@
|
||||
在网关控制面中,建议遵循以下步骤新增反代规则:
|
||||
|
||||
```text
|
||||
[ 步骤 1. 证书管理 ] ──► [ 步骤 2. 源站定义 (可选) ] ──► [ 步骤 3. 新增网站配置 ]
|
||||
[ 步骤 1. 证书管理 ] ──► [ 步骤 2. 源站定义(可选) ] ──► [ 步骤 3. 新增网站配置 ]
|
||||
│
|
||||
[ 步骤 5. 验证访问 ] ◄── [ 步骤 4. 发布与激活版本 ] ◄───────────────┘
|
||||
```
|
||||
@@ -28,7 +28,7 @@
|
||||
|
||||
源站(Origin)代表被代理的后端真实服务地址。虽然在新建网站时可以直接填写 IP,但推荐先在源站库中进行注册,以便后续复用与维护:
|
||||
|
||||
1. 进入左侧导航 **「网站管理」->「源站地址」**,点击 **「创建源站」**。
|
||||
1. 进入左侧导航 **「网站管理」->「源站地址」**,点击 **「新增源站」**。
|
||||
2. 填写源站名称(如 `production-api`)。
|
||||
3. 填入合法的上游地址(如 `http://10.0.0.10:8080`),点击保存。
|
||||
|
||||
@@ -38,13 +38,13 @@
|
||||
|
||||
证书和源站就绪后,即可创建核心网站代理路由:
|
||||
|
||||
1. 进入左侧导航 **「网站管理」->「域名列表」**,点击 **「新增网站」**:
|
||||
1. 进入左侧导航 **「网站管理」->「域名列表」**,点击 **「新增 Zone」**:
|
||||
* **域名**:输入该站点绑定的域名。
|
||||
* **绑定证书**:选择第一步准备或申请好的证书。
|
||||
2. 配置请求路由规则:进入 **「规则管理」** 页面,点击 **「新增规则」** 或编辑已有规则:
|
||||
2. 配置请求路由规则:进入 **「规则管理」** 页面,点击 **「新建规则」** 或编辑已有规则:
|
||||
* **规则名称**:输入规则的唯一简易标识(如 `app-portal-route`)。
|
||||
* **域名匹配**:填入对应的域名(支持通配符或精确域名,需与上面登记的域名一致)。
|
||||
* 在下方 **「反向代理」** 选项卡下,选择源站类型为 **「标准反代」**。
|
||||
* 在下方 **「反向代理」** 选项卡下,选择 **回源方式** 为「直连上游」。
|
||||
* **源站选择**:从下拉框中选择第二步创建的源站;或者选择手动输入并填入 `http://10.0.0.20:9000`。
|
||||
3. 点击保存创建配置。
|
||||
|
||||
@@ -54,13 +54,13 @@
|
||||
|
||||
你在管理端新增的网站配置仅保存在 Server 数据库中,**不会立即生效**。必须生成配置版本快照并分发到 Agent 边缘节点:
|
||||
|
||||
1. 点击控制面板右上角的 **「配置预览」** 按钮。
|
||||
1. 点击控制面板右上角的 **「预览并发布」** 按钮。
|
||||
2. 检查配置文件的 Diff 差异,确认你刚刚新增的 `server` 块以及证书绑定规则无误。
|
||||
3. 点击 **「发布并激活」** 按钮。
|
||||
3. 点击 **「确认发布」** 按钮。
|
||||
4. **Agent 落地机制**:
|
||||
* 数据面的 Agent 节点在心跳中发现激活的版本 Checksum 变更,会自动拉取完整的 OpenResty 配置文件和证书包到本地。
|
||||
* 自动在本地执行配置校验(类似于 `openresty -t`),确认无语法错误后,执行平滑重载(`reload`)。
|
||||
* *如果重载或校验失败,Agent 会安全阻断并回滚至上一稳定版本,保证节点高可用。*
|
||||
* 如果重载或校验失败,Agent 会安全阻断并回滚至上一稳定版本,保证节点高可用。
|
||||
|
||||
---
|
||||
|
||||
@@ -80,9 +80,9 @@
|
||||
|
||||
### 2. 一键秒级回滚
|
||||
如果发布的新配置导致了线上业务异常:
|
||||
1. 导航至左侧 **「配置版本」** 菜单。
|
||||
1. 导航至左侧 **「版本发布」** 菜单。
|
||||
2. 在历史列表中找到发布前的上一个稳定版本。
|
||||
3. 点击 **「激活此版本」**。
|
||||
3. 点击 **「激活」**。
|
||||
4. 所有在线 Agent 节点将在秒级自动重载回历史配置,实现秒级避险。
|
||||
|
||||
---
|
||||
|
||||
+16
-18
@@ -19,16 +19,12 @@ Agent 统一通过 OpenResty 二进制控制运行时。本地部署需要节点
|
||||
| Docker / Docker Compose | 用于启动 Server 及其依赖的 PostgreSQL、Valkey;如采用 Docker Agent,也用于运行 Agent |
|
||||
| OpenResty | 本地安装 Agent 时需要可执行 `openresty`,或在安装脚本中指定路径 |
|
||||
| 可访问端口 | Server 默认监听 `3000`,Agent 节点需要能访问 Server 地址 |
|
||||
| 浏览器 | 用于访问管理端 |
|
||||
|
||||
- **Docker**:`20.10.0+`
|
||||
- **Docker Compose**:`2.0.0+`
|
||||
|
||||
---
|
||||
|
||||
## 1. 启动 Server
|
||||
|
||||
快速开始推荐采用 **PostgreSQL + Redis ** 标准部署方案。
|
||||
快速开始推荐采用 **PostgreSQL + Valkey** 标准部署方案。
|
||||
|
||||
在空目录中创建 `docker-compose.yaml`:
|
||||
|
||||
@@ -123,12 +119,14 @@ http://localhost:3000
|
||||
> [!WARNING]
|
||||
> 为了你的系统安全,首次登录后请立即修改默认密码。
|
||||
|
||||
如果忘记密码并且没有配置找回密码渠道, 可以使用命令进行重置
|
||||
如果忘记密码并且没有配置找回密码渠道,可以使用命令重置:
|
||||
|
||||
```bash
|
||||
go run main.go reset-paswd # 重置管理员密码
|
||||
go run main.go reset-passwd --user admin
|
||||
```
|
||||
|
||||
未指定 `--password` 时命令会自动生成随机密码并输出到终端;也可以使用 `--password` 显式指定新密码。
|
||||
|
||||
---
|
||||
|
||||
## 2. 准备 Agent Token
|
||||
@@ -142,14 +140,14 @@ Agent 可以用两类凭证接入:
|
||||
|
||||
在管理端准备其中一种凭证后,进入下一步。
|
||||
|
||||
- **`discovery_token`** 获取菜单路径:「系统设置」 (Settings) -> 「OpenFlare」选项卡 -> 「自动注册」凭证
|
||||
- **`discovery_token`** 获取菜单路径:「系统设置」->「OpenFlare」选项卡 ->「Discovery Token 与部署」中的 Discovery Token
|
||||
- **`agent_token`** 获取菜单路径:在「节点管理」中创建节点后,点击进入节点详情页即可查看到对应的专属 Token。
|
||||
|
||||
---
|
||||
|
||||
## 3. 安装/运行 Agent
|
||||
|
||||
Agent 部署方式推荐使用 Docker 部署(即直接运行内置 OpenResty 的 Agent 镜像);亦支持通过安装脚本将 Agent 部署在本地宿主机上。
|
||||
推荐使用 Docker 镜像部署 Agent;也可以通过安装脚本部署到本地宿主机。
|
||||
|
||||
### 方式 A:Docker 运行 Agent(推荐)
|
||||
|
||||
@@ -217,17 +215,17 @@ journalctl -u openflare-agent -f
|
||||
|
||||
---
|
||||
|
||||
## 常见失败原因
|
||||
## 遇到问题时
|
||||
|
||||
| 现象 | 排查方向 |
|
||||
| --- | --- |
|
||||
| 浏览器打不开管理端 | 确认 `docker compose ps` 中 Server 正在运行,宿主机 `3000` 端口没有被占用 |
|
||||
| 登录后数据无法保存/提示报错 | 检查 PostgreSQL 容器健康状态,以及 `DB_PASSWORD` / 密码等连接参数是否一致 |
|
||||
| Agent 无法注册 | 确认 Agent 节点能访问 `--server-url`,并检查 Token 是否填错或已失效 |
|
||||
| Agent 在线但没有应用配置 | 确认网站配置已启用,并且已经发布并激活版本 |
|
||||
| OpenResty 应用失败 | 查看节点应用记录和 `journalctl -u openflare-agent`,重点检查域名、证书、上游地址和端口占用 |
|
||||
按以下顺序处理:
|
||||
|
||||
更多排查路径见 [故障排查](./troubleshooting.md)。
|
||||
1. 将 Server 与 Agent 升级到最新版本,确认问题是否仍然存在。
|
||||
2. 重新发布并激活配置版本,等待节点应用。
|
||||
3. 在节点详情页对目标节点执行「强制同步」,推动节点立即拉取最新配置。
|
||||
4. 重建或重装 Agent(重新执行安装脚本)。
|
||||
5. 上述步骤均无效时,携带 Server 日志与节点应用记录提交 [GitHub Issue](https://github.com/Rain-kl/OpenFlare/issues)。
|
||||
|
||||
更多排查思路见 [故障排查](./troubleshooting.md)。
|
||||
|
||||
---
|
||||
|
||||
|
||||
+23
-44
@@ -1,72 +1,51 @@
|
||||
# SSO 登录配置
|
||||
|
||||
你会学到:如何为 OpenFlare 配置 GitHub OAuth 或标准 OIDC 登录入口,如何填写回调地址,以及第三方账号如何绑定本地用户。
|
||||
你会学到:如何为 OpenFlare 配置 OIDC 第三方登录入口、填写回调地址,以及第三方账号如何绑定本地用户。
|
||||
|
||||
OpenFlare 支持通过认证源配置第三方登录入口。当前支持 GitHub OAuth 与标准 OIDC Provider,例如 Logto、authentik、Keycloak、Casdoor 等。
|
||||
OpenFlare 通过 OIDC 认证源接入第三方登录。任意提供标准 OIDC Discovery 的服务(如 Google、Keycloak、authentik、Logto、Casdoor 等)都可以接入。
|
||||
|
||||
认证源配置完成并启用后,会显示在登录页的第三方账号登录区域。用户可以通过第三方账号登录,也可以在已登录状态下把第三方账号绑定到当前本地账号。
|
||||
|
||||
## 使用前准备
|
||||
|
||||
你需要先准备:
|
||||
|
||||
| 项目 | 说明 |
|
||||
| --- | --- |
|
||||
| OpenFlare 访问地址 | 用户浏览器实际访问的地址,例如 `https://openflare.example.com` |
|
||||
| 认证源名称 | OpenFlare 内部唯一标识,例如 `github`、`company-oidc` |
|
||||
| 服务器访问地址 | 在管理端「系统设置」->「系统设置」选项卡 ->「通用设置」中配置,须与用户浏览器实际访问的地址一致(协议、域名、端口) |
|
||||
| 认证源名称 | OpenFlare 内部唯一标识,例如 `company-oidc` |
|
||||
| Client ID | 第三方平台创建应用后提供 |
|
||||
| Client Secret | 第三方平台创建应用后提供 |
|
||||
| OIDC Discovery URL | 仅 OIDC 需要,例如 `https://idp.example.com/.well-known/openid-configuration` |
|
||||
| OIDC Discovery URL | 例如 `https://idp.example.com/.well-known/openid-configuration` |
|
||||
|
||||
**确认系统设置->通用设置->服务器地址能正确和域名匹配**
|
||||
|
||||
认证源名称只能包含字母、数字、短横线或下划线,并且必须以字母或数字开头。认证源名称会出现在回调地址中,保存后如需修改名称,也必须同步修改第三方平台中的回调地址。
|
||||
认证源名称只能包含字母、数字、短横线或下划线,并且必须以字母或数字开头。
|
||||
|
||||
## 回调地址
|
||||
|
||||
第三方平台中的 Redirect URI / Callback URL 填写格式为:
|
||||
第三方平台中的 Redirect URI / Callback URL 固定填写:
|
||||
|
||||
```text
|
||||
<OpenFlare 访问地址>/oauth/<认证源名称>
|
||||
<服务器访问地址>/login
|
||||
```
|
||||
|
||||
示例:
|
||||
例如服务器访问地址为 `https://openflare.example.com` 时:
|
||||
|
||||
```text
|
||||
https://openflare.example.com/oauth/github
|
||||
https://openflare.example.com/oauth/company-oidc
|
||||
https://openflare.example.com/login
|
||||
```
|
||||
|
||||
在管理端新增或修改认证源时,表单会根据当前浏览器访问地址和你输入的认证源名称自动显示应填写的回调地址。
|
||||
|
||||
## 配置 GitHub 登录
|
||||
|
||||
1. 在 GitHub 创建 OAuth App。
|
||||
2. `Homepage URL` 填写 OpenFlare 访问地址。
|
||||
3. `Authorization callback URL` 填写 OpenFlare 显示的回调地址,例如 `https://openflare.example.com/oauth/github`。
|
||||
4. 复制 GitHub 提供的 Client ID 和 Client Secret。
|
||||
5. 登录 OpenFlare 管理端,进入左侧导航 **「系统设置」** (Settings),选择 **「安全设置」** 选项卡,在 **「认证源管理」** 栏目中进行配置。
|
||||
6. 新增认证源,类型选择 `GitHub`。
|
||||
7. 填写认证源名称、展示名称、Client ID、Client Secret。
|
||||
8. Scope 默认使用 `user:email`,通常无需修改。
|
||||
9. 保存并启用认证源。
|
||||
|
||||
启用后,登录页会显示对应的 GitHub 登录按钮。
|
||||
回调地址只与「服务器访问地址」相关,不包含认证源名称。第三方平台授权完成后会跳转到该地址,OpenFlare 登录页携带授权码完成登录或绑定。
|
||||
|
||||
## 配置 OIDC 登录
|
||||
|
||||
1. 在 OIDC Provider 中创建应用或客户端。
|
||||
2. 应用类型选择 Web / Confidential Client。
|
||||
3. Redirect URI / Callback URL 填写 OpenFlare 显示的回调地址,例如 `https://openflare.example.com/oauth/company-oidc`。
|
||||
4. 复制 Client ID 和 Client Secret。
|
||||
5. 获取 Provider 的 Discovery URL,通常以 `/.well-known/openid-configuration` 结尾。
|
||||
6. 登录 OpenFlare 管理端,进入左侧导航 **「系统设置」** (Settings),选择 **「安全设置」** 选项卡,在 **「认证源管理」** 栏目中进行配置。
|
||||
7. 新增认证源,类型选择 `OIDC`。
|
||||
8. 填写认证源名称、展示名称、Client ID、Client Secret、OIDC Discovery URL。
|
||||
9. Scope 默认使用 `openid profile email`。如果 Provider 限制了 scope,请按 Provider 允许的值调整。
|
||||
10. 保存并启用认证源。
|
||||
1. 在 OIDC Provider 中创建应用或客户端,应用类型选择 Web / Confidential Client。
|
||||
2. Redirect URI / Callback URL 填写 `<服务器访问地址>/login`。
|
||||
3. 复制 Client ID 和 Client Secret。
|
||||
4. 获取 Provider 的 Discovery URL,通常以 `/.well-known/openid-configuration` 结尾。
|
||||
5. 登录 OpenFlare 管理端,进入左侧导航 **「系统设置」**,选择 **「安全设置」** 选项卡,在 **「认证源管理」** 栏目中新增认证源。
|
||||
6. 类型选择 `OIDC`,填写认证源名称、展示名称、Client ID、Client Secret、OIDC Discovery URL。
|
||||
7. Scope 默认使用 `openid profile email`。如果 Provider 限制了 scope,请按 Provider 允许的值调整。
|
||||
8. 保存并启用认证源。
|
||||
|
||||
启用后,登录页会显示对应的 OIDC 登录按钮。
|
||||
启用后,登录页会显示对应的第三方登录按钮。
|
||||
|
||||
## 登录与绑定行为
|
||||
|
||||
@@ -85,17 +64,17 @@ https://openflare.example.com/oauth/company-oidc
|
||||
|
||||
修改认证源时,Client Secret 输入框留空表示保留已有密钥;填写新值则会覆盖保存。
|
||||
|
||||
如果修改了认证源名称,回调地址也会随之变化。你必须到第三方平台同步修改 Redirect URI / Callback URL,否则第三方平台会拒绝回调或返回错误。
|
||||
修改认证源名称不会影响回调地址,无需同步修改第三方平台配置。
|
||||
|
||||
## 常见问题
|
||||
|
||||
### 返回 `invalid_scope`
|
||||
|
||||
说明第三方平台不允许当前配置的 Scope。OIDC 默认 Scope 是 `openid profile email`,GitHub 默认 Scope 是 `user:email`。请到认证源编辑页调整 Scope,或在第三方平台放行对应 Scope。
|
||||
说明第三方平台不允许当前配置的 Scope。OIDC 默认 Scope 是 `openid profile email`。请到认证源编辑页调整 Scope,或在第三方平台放行对应 Scope。
|
||||
|
||||
### 提示回调地址不匹配
|
||||
|
||||
检查第三方平台中配置的 Redirect URI / Callback URL 是否与 OpenFlare 表单提示完全一致。协议、域名、端口和路径都必须一致。
|
||||
检查第三方平台中配置的 Redirect URI / Callback URL 是否与 `<服务器访问地址>/login` 完全一致。协议、域名、端口和路径都必须一致。
|
||||
|
||||
### 登录页没有显示第三方登录按钮
|
||||
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# 故障排查
|
||||
|
||||
你会学到:如何按症状排查 OpenFlare Server、数据库、登录、Agent、OpenResty、配置发布和前端构建问题。
|
||||
你会学到:如何按症状排查 OpenFlare Server、数据库、登录、Agent、OpenResty 和配置发布问题。
|
||||
|
||||
排查时先确认问题发生在哪一层:浏览器、Server、数据库、Agent、OpenResty、源站或 DNS。OpenFlare 的配置不会直接在线写入所有节点,只有激活版本变化后,Agent 才会在 heartbeat 中发现并应用。
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
| 现象 | 先看哪里 |
|
||||
| --- | --- |
|
||||
| 管理端打不开 | Server 容器或进程日志、端口监听 |
|
||||
| 登录异常 | 默认账号、OPENFLARE_TOKEN、浏览器请求、Server 日志 |
|
||||
| 登录异常 | 默认账号、Session Cookie、Server 日志 |
|
||||
| 数据无法保存 | 数据库连接、SQLite 文件权限、PostgreSQL 健康状态 |
|
||||
| Agent 离线 | Agent 日志、Token、Server 地址、网络连通性 |
|
||||
| 发布后节点未更新 | 激活版本、节点 heartbeat、应用记录 |
|
||||
@@ -50,7 +50,7 @@ ls -ld "$(dirname /path/to/openflare.db)"
|
||||
|
||||
| 日志或现象 | 处理 |
|
||||
| --- | --- |
|
||||
| 数据库连接失败 | 检查 `DSN` 中用户名、密码、主机、端口、库名和 `sslmode` |
|
||||
| 数据库连接失败 | 检查 `DB_HOST`、`DB_PORT`、`DB_USERNAME`、`DB_PASSWORD`、`DB_NAME`、`DB_SSL_MODE` 是否一致 |
|
||||
| SQLite 无法创建文件 | 检查 `SQLITE_PATH` 所在目录是否存在且可写 |
|
||||
| 端口被占用 | 修改 `PORT` 或 `--port`,或停止占用端口的进程 |
|
||||
|
||||
@@ -62,21 +62,7 @@ ls -ld "$(dirname /path/to/openflare.db)"
|
||||
curl -I http://127.0.0.1:3000
|
||||
```
|
||||
|
||||
2. 如果是源码运行,确认已经构建前端静态产物:
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
pnpm build
|
||||
```
|
||||
|
||||
3. 检查浏览器访问地址是否与反向代理配置一致。
|
||||
|
||||
4. 如果通过前端开发服务器访问,确认后端代理地址:
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
NEXT_DEV_BACKEND_URL=http://127.0.0.1:3000 pnpm dev
|
||||
```
|
||||
2. 检查浏览器访问地址是否与反向代理配置一致。
|
||||
|
||||
## 默认账号无法登录
|
||||
|
||||
@@ -84,32 +70,20 @@ NEXT_DEV_BACKEND_URL=http://127.0.0.1:3000 pnpm dev
|
||||
|
||||
排查步骤:
|
||||
|
||||
1. 确认连接的是预期数据库,避免 `SQLITE_PATH` 或 `DSN` 指向了另一个环境。
|
||||
1. 确认连接的是预期数据库,避免 `SQLITE_PATH` 或 `DB_HOST` / `DB_NAME` 指向了另一个环境。
|
||||
2. 查看 Server 日志中使用的是 `sqlite` 还是 `postgres`。
|
||||
3. 在浏览器开发者工具中确认管理端 API 请求已正确携带 Session Cookie。
|
||||
4. 清理浏览器缓存及 Cookie 后重新登录。
|
||||
|
||||
### 应急重置管理员密码
|
||||
|
||||
如果忘记了 `admin` 账户的密码,可以通过直接更新数据库中的密码哈希值将其重置为 `12345678`(登录后请务必立即修改):
|
||||
忘记 `admin` 账户密码时,使用 `reset-passwd` 命令重置(支持 SQLite 与 PostgreSQL):
|
||||
|
||||
#### 1. 若使用 SQLite 数据库
|
||||
停止 Server 运行,使用 sqlite3 客户端打开数据库文件:
|
||||
```bash
|
||||
sqlite3 /path/to/openflare.db
|
||||
go run main.go reset-passwd --user admin --password your-new-password
|
||||
```
|
||||
执行以下 SQL 语句:
|
||||
```sql
|
||||
UPDATE users SET password = '$2a$10$eXpE9i/6S3gPT94/G0mu0.B8ser66ARETFz5NWYSYcrQ4JmtSrMXu' WHERE username = 'admin';
|
||||
```
|
||||
输入 `.exit` 退出并重新启动 Server。
|
||||
|
||||
#### 2. 若使用 PostgreSQL 数据库
|
||||
通过您的数据库连接工具(如 psql、pgAdmin 或 DBeaver)连接到 PostgreSQL 实例,选择对应的 `openflare` 数据库,执行以下 SQL 语句:
|
||||
```sql
|
||||
UPDATE users SET password = '$2a$10$eXpE9i/6S3gPT94/G0mu0.B8ser66ARETFz5NWYSYcrQ4JmtSrMXu' WHERE username = 'admin';
|
||||
```
|
||||
执行成功后即可使用默认密码 `12345678` 重新登录管理后台。
|
||||
若使用 SQLite,建议先停止 Server 进程再执行,避免数据库文件锁冲突。未指定 `--password` 时命令会生成随机密码并输出到终端。重置成功后请立即登录并修改密码。
|
||||
|
||||
## Agent 无法注册或一直离线
|
||||
|
||||
@@ -216,29 +190,6 @@ curl -Iv https://your-domain
|
||||
4. 检查 `openresty_observability_port` 是否被占用,默认是 `18081`。
|
||||
5. 确认 Server 侧没有因数据库清理策略删除对应时间窗口数据。
|
||||
|
||||
## 前端构建失败
|
||||
|
||||
执行:
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
corepack enable
|
||||
pnpm install
|
||||
pnpm lint
|
||||
pnpm typecheck
|
||||
pnpm test
|
||||
pnpm build
|
||||
```
|
||||
|
||||
常见原因:
|
||||
|
||||
| 现象 | 处理 |
|
||||
| --- | --- |
|
||||
| pnpm 版本不一致 | 使用 `corepack enable` 后重新安装 |
|
||||
| 类型错误 | 先运行 `pnpm typecheck` 定位具体文件 |
|
||||
| API 类型不一致 | 检查 `lib/api/` 和 `types/` 中的响应结构 |
|
||||
| E2E 失败 | 确认 Server 和前端开发服务器都已启动 |
|
||||
|
||||
## 边缘缓存命中率异常
|
||||
|
||||
访问日志中缓存三态:**命中**(HIT/STALE/REVALIDATED/UPDATING)、**回源**(MISS/EXPIRED)、**未缓存**(BYPASS 或空,请求时未进入可缓存路径或响应未入库)。设计说明见 [边缘缓存策略设计](../design/edge-cache-design.md)。
|
||||
@@ -269,12 +220,3 @@ pnpm build
|
||||
* 响应带 `Set-Cookie` 或 `private`:**不入库**。
|
||||
* 无源站缓存头的可缓存状态码:使用默认 Edge TTL(如 200 约 120 分钟)。
|
||||
|
||||
## 文档站构建失败
|
||||
|
||||
```bash
|
||||
cd docs
|
||||
pnpm install
|
||||
pnpm build
|
||||
```
|
||||
|
||||
如果是链接错误,检查新增页面是否已经加入 `docs/config.ts` 侧边栏,或者相对链接是否指向存在的 Markdown 文件。
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user