Compare commits

...

24 Commits

Author SHA1 Message Date
ryan 01ed2c5e36 chore(release): v3.5.2
修复几个遗漏bug
2026-08-09 14:08:04 +08:00
ryan 80696c12fa fix: lint 2026-08-09 13:47:40 +08:00
ryan 3d4d99081e fix(log): PG 日志库批量写入为零 ID 行生成雪花 ID
PostgreSQL 日志表 id 为 NOT NULL 且无默认值,而 GORM 将零值 uint64
主键视为自增并省略 id 列,导致 node access log / 可观测指标等批量
落库持续报 "null value in column id violates not-null constraint"。
在 BatchInsert* 落库前为零 ID 行生成雪花 ID(与 ClickHouse 写入路径
一致),并新增单元回归与 PG 集成回归测试覆盖六张日志表。
2026-08-09 13:47:13 +08:00
ryan 0639855653 fix(openresty): 修复源站错误页「仅针对 GET 请求」未生效
error_page 的 URI 内部重定向会把请求方法改写成 GET,导致内部
Lua 中 ngx.req.get_method() 恒为 GET,get_only 判断永不命中,
POST/PUT 等请求仍返回自定义错误页。

改为命名 location(@__openflare_origin_error)承载错误页:
命名 location 保留原始请求方法与原始错误状态码,非 GET 请求
直接以原状态码退出、不再注入自定义 HTML。附带回归断言,禁止
回退到 URI 内部重定向形式。
2026-08-09 13:42:38 +08:00
ryan 0524ae1da4 chore(release): v3.5.1
### ✨ 新功能
- 日志存储解耦:ClickHouse 变为可选项,不启用时由 PostgreSQL/SQLite 承担全部日志功能;新增「切换日志数据库」任务支持 PostgreSQL/SQLite 与 ClickHouse 间数据迁移(迁移期间冻结日志写入,成功后自动切换主库并保留源数据);`log_database` / `log_db_migration` 设为受保护配置;ClickHouse 改为默认关闭。
- 新增 PostgreSQL/SQLite 日志存储实现:节点访问日志按月分区,统计查询合并为单次扫描、IP 汇总归属地取查询窗口内最新记录、WAF 按 IP 聚合减少扫描次数,并新增 `logged_at` 前导索引与主机名小写表达式索引;过期清理直接删除完全过期的整月分区,启动时兜底预建当月及未来 2 个月分区。
- 性能指标与访问日志的保留时长解耦:新增三库共用的 `metric_retention_days` 配置(默认 3 天),每日垃圾清理按独立短留存清理指标快照。

### 🛠 修复
- 修复 UptimeKuma 同步调试日志泄露凭据:Socket.IO 事件日志不再打印 payload 内容(仅记录长度),避免凭据进入日志。
- 修复日志保留天数配置继承旧键导致的误删风险:`log_retention_days_*` 不再继承 `database_auto_cleanup_retention_days`,统一默认 30 天。

### ⚡️ 优化与改进
- 系统定期垃圾清理由每 2 小时改为每日执行一次(凌晨 3 点,Asia/Shanghai),降低非必要高频扫描。

### 💄 其他/体验
- 服务工作者(SW)注入挑战页改为前台无感知:不再显示「加载中…」文案,页面空白,仅通过浏览器控制台输出 `[sw-challenge]` 调试信息,注入过程不打扰访客。
- 用户访问日志(`w_user_access_logs`)记录禁用:不再采集与写入新的用户访问日志,存量数据与管理端访问日志统计页面保留。
2026-08-09 11:39:28 +08:00
ryan adee4f7b27 docs: update 2026-08-09 11:22:35 +08:00
ryan 1d0f2d6342 fix(log): hard-set log retention days default to 30, drop legacy inheritance
log_retention_days_* 迁移不再继承旧键 database_auto_cleanup_retention_days
的值,统一默认 30 天。此前若旧键残留异常小值(如 2 天)会被静默带入,
导致升级后首次垃圾清理把大部分日志直接删掉。PG/SQLite 双方言同步修改,
文档默认值 90 -> 30。
2026-08-09 10:48:45 +08:00
ryan e3f603f72a fix(security): stop logging UptimeKuma socket payload content 2026-08-09 10:42:21 +08:00
ryan 3b010bb15e feat(log): disable user access log recording
- 移除全局用户访问日志采集中间件与批写入 writer(risk_control 包整包删除),
  不再写入 w_user_access_logs;存量数据与管理端访问日志统计页面保留
- 日志库迁移任务不再排空用户访问日志队列,状态接口不再展示其缓冲队列统计
- 迁移测试的系统配置 seed 计数断言更新为当前实际值(86 → 95),
  注释改为提示新增配置 seed 时同步更新
2026-08-09 10:35:39 +08:00
ryan f530cd4025 perf(log): optimize PG log store queries and expired partition cleanup
- Count/节点访问日志统计改为单次扫描聚合,WAF 按 IP 聚合由三次扫描合并为两次
- IPSummaries 归属地改为取过滤窗口内最新记录(对齐 ClickHouse argMax 口径),
  子查询带窗口条件,可分区裁剪并命中索引
- 新增 goose 迁移:of_node_access_logs (logged_at DESC, id DESC) 前导索引与
  lower(trim(host)) 表达式索引,加速列表排序与主机过滤
- 过期日志清理先按数据校验直接 DROP 完全过期整月分区,再对边界月逐行删除;
  启动时兜底预建当月及未来 2 个月分区,跨月停机重启后首次写入不再报
  "no partition of relation found"
2026-08-09 10:20:58 +08:00
ryan 08d28c2c8e feat(log): drop empty old-month PG partitions during cleanup
系统垃圾清理任务删除过期日志后,顺带清理旧月份空分区表:
PostgreSQL 按月分区的访问日志表(节点/用户)在数据删除后若该月
分区已无数据,则自动删除对应分区表,避免历史分区表无限累积。

- 仅删除「当前月之前」且为空的月份分区,当月/未来月及仍有数据的分区保留
- ClickHouse/SQLite 为 no-op(CH 分区随数据删除自动消失)
- 修复既有集成测试 pg_inherits 查询(inhrelid → inhparent)
- 新增单元测试与 PG 集成测试
2026-08-09 09:33:59 +08:00
ryan 34a0896ff8 chore(task): run system garbage cleanup once daily
系统定期垃圾清理 cron 由每 2 小时(0 */2 * * *)改为每日凌晨 3 点
(0 3 * * *,Asia/Shanghai),降低非必要高频扫描。新增 PG/SQLite
双方言 goose 迁移(含 Down 回滚)与迁移测试。
2026-08-09 09:25:30 +08:00
ryan 0c22e76f4b fix(frontend): optimization 2026-08-09 09:14:54 +08:00
ryan bd2183c8bb feat(log): add independent short retention for performance metrics
性能指标(CPU/内存/磁盘/网络)不再跟随 log_retention_days_*,新增三库共用
的 metric_retention_days 配置(默认 3 天),系统垃圾清理按独立短留存清理
指标快照;访问日志保留时长不变。新增 PG/SQLite 双方言 goose 迁移 seed。
2026-08-09 09:04:33 +08:00
ryan 9df2437e47 fix(config): set ClickHouse to disabled by default and update related documentation 2026-08-09 08:54:26 +08:00
ryan e8c414aa12 fix: ch migrate 2026-08-09 08:47:56 +08:00
ryan 7e8aa5fa0f Merge branch 'codex/log-database-decoupling'
# Conflicts:
#	docs/changelog/index.md
#	frontend/app/(main)/error-pages/page.tsx
#	internal/infra/persistence/migrator/migrator_test.go
2026-08-09 08:37:37 +08:00
Ryan 93ec3096f3 Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-08-08 20:48:54 +08:00
Ryan 0e34301c92 Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-08-08 20:48:37 +08:00
Ryan 12b5271f92 Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-08-08 20:47:33 +08:00
Ryan 8ee966434d Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-08-08 20:46:54 +08:00
ryan 7d93d3d2a1 fix(log): address remaining CodeRabbit suggestions for log database switch and migrations 2026-08-08 20:36:04 +08:00
ryan a6fc2b7737 fix(log): address CodeRabbit review findings for log database decoupling 2026-08-08 20:15:36 +08:00
ryan 7d71f1e4e1 feat(log): decouple log storage from ClickHouse with switchable logstore
- New internal/repository/logstore abstraction: exported domain interfaces
  (AccessLogStore/ObservabilityStore/UserAccessLogStore/StatusStore),
  config-driven provider (Active/Build/Migrating/SetConfigReader), GORM
  implementation for PostgreSQL/SQLite (incl. hourly rollups computed in
  real time, migration listers, PG partition maintenance), and a ClickHouse
  wrapper preserving the native batch path; repository facade delegates to
  logstore; import-lint test enforces apps never import analyticsrepo.
- ClickHouse is now optional: the log DB is either the main DB (postgres
  when database.enabled, else sqlite) or clickhouse; boot validation +
  first-run seed; log_database / log_db_migration are protected keys.
- New user task 切换日志数据库 (of_log_db_switch): freeze log writes,
  drain batch writers, copy all 6 raw log tables by id (preserving IDs)
  with target-partition pre-creation for PG, flip log_database on success,
  clear the freeze flag on failure.
- Per-store retention (log_retention_days_*) with expiry cleanup folded
  into the daily system_cleanup task; legacy database_auto_cleanup_* and
  of_database_auto_cleanup decommissioned.
- goose migrations: 6 log tables in PG (2 monthly-partitioned) + SQLite,
  retention config seeds, schedule cleanup; GET
  /api/v1/admin/status/log-database endpoint; frontend retention settings,
  switch-task UI and status badge; changelog and docs updated.

docs(plan): log database decoupling implementation plan

docs(design): log database decoupling design (ClickHouse optional)
2026-08-08 19:43:01 +08:00
135 changed files with 12188 additions and 4901 deletions
+12
View File
@@ -38,6 +38,7 @@ description: "Wavelet 项目专用:根据自上一个正式版本 Tag 以来
固定使用以下分类:
```text
### 新增
### 🛠 修复
### ⚡️ 优化与改进
### 💄 其他/体验
@@ -45,15 +46,26 @@ description: "Wavelet 项目专用:根据自上一个正式版本 Tag 以来
分类规则:
- 新功能、新能力、新配置、新任务:放入 ### 新增
- Bug、异常行为、错误逻辑:放入 ### 🛠 修复
- 性能、稳定性、接口、架构、兼容性:放入 ### ⚡️ 优化与改进
- 日志、文案、UI、文档、开发体验:放入 ### 💄 其他/体验
「修复/优化」与「新增」的判定(关键):
- **判定标准是“该功能在上一正式版本中是否已存在”**:
- 已存在 → 本次对其 bug 的修正可计入「🛠 修复」,对其行为/性能的改进可计入「⚡️ 优化与改进」;
- 不存在(本版本新增)→ 该功能的一切内容——包括开发过程中修的 bug、做的性能优化、补的索引——都只属于新功能开发的一部分,不应该在发布说明中提及。
- 禁止把新功能的开发期修复/优化写进「修复」或「优化」:新功能此前版本没有,谈不上“修复/优化了旧行为”。
示例:
```
chore(release): v3.3.0
### ✨ 新功能
- 新增笔记库快照备份功能,支持定时备份与手动一键恢复(仅说明新增的功能, 禁止提及新功能开发时期的优化修复等内容)。
### 🛠 修复
- 修复了通过 MCP 接口操作时笔记库范围限制未正确生效的问题。
- 修复了 MCP 接口返回数据格式不一致的问题。
-19
View File
@@ -1,19 +0,0 @@
root = true
[*]
indent_style = space
indent_size = 4
charset = utf-8
end_of_line = lf
trim_trailing_whitespace = true
insert_final_newline = true
[*.{json,yml,yaml}]
indent_size = 2
[*.md]
insert_final_newline = false
trim_trailing_whitespace = false
[*.{js,ts,css,html,jsx,tsx,vue}]
indent_size = 2
+1 -1
View File
@@ -56,7 +56,7 @@ REDIS_MAINT_NOTIFICATIONS=false
# ─── ClickHouse(必需)────────────────────────────────────────────────────────
# CLICKHOUSE_HOST 设置后会自动启用;测试环境可显式 CLICKHOUSE_ENABLED=true 做 live 联调
CLICKHOUSE_ENABLED=true
CLICKHOUSE_ENABLED=false
# compose 内:clickhouse:9000;本机连映射端口:127.0.0.1:9000
CLICKHOUSE_HOST=clickhouse:9000
CLICKHOUSE_USERNAME=default
+61 -106
View File
@@ -2,9 +2,9 @@
# OpenFlare
**[English](./README.en.md) | [📖 中文](./README.md)**
**[📖 中文](./README.md) | [English](./README.en.md)**
OpenFlare is an open-source CDN orchestration and edge security platform. It supports reverse proxies, centralized configuration synchronization, secure intranet penetration (Tunnels), dynamic WAF protection, and anti-CC challenges.
OpenFlare is an open-source CDN orchestration and edge security platform. It supports reverse proxy, centralized configuration synchronization, in-network tunneling (Tunnels), dynamic WAF protection, and CC defense challenges.
</div>
@@ -21,37 +21,67 @@ OpenFlare is an open-source CDN orchestration and edge security platform. It sup
</p>
> [!WARNING]
> After logging in for the first time with the `root` user, make sure to change the default password `123456`.
>
> The BETA version is a temporary product for the development and testing phase. It may contain unknown issues and should not be used in production environments.
> After the first login with the `admin` user, you must change the default password `12345678`.
>
> The BETA version is a temporary product in the development and testing stage and may have unknown issues. It should not be used in production environments.
## Documentation
**https://open-flare.pages.dev**
Quick links:
Common entry points:
* [Quick Start](https://open-flare.pages.dev/en/guide/quick-start)
* [Deployment Guide](https://open-flare.pages.dev/en/deployment/deployment)
* [Quick Start](https://open-flare.pages.dev/guide/quick-start)
* [Deployment Guide](https://open-flare.pages.dev/deployment/deployment)
* [Configuration Reference](https://open-flare.pages.dev/reference/configuration)
* [System Design](https://open-flare.pages.dev/design/)
## Core Features
## Core Capabilities
* **Reverse Proxy Management**: Website rules as the aggregation boundary, supporting multi-domain binding and multi-upstream load balancing with unified management of all OpenResty node configurations.
* **Immutable Config Version Control**: Full-snapshot publish model based on version numbers (`YYYYMMDD-NNN`), with pre-publish diff preview, a single globally active version, and one-click sub-second rollback.
* **Secure Intranet Penetration (Tunnels)**: An open-source alternative to Cloudflare Tunnels. Securely expose local intranet Web services to the public network via Relay and OpenFlared clients — no public IP or open inbound ports required.
* **Edge WAF Safety Protection**: Provides global and custom rule groups, supporting manual/automatic/subscription IP groups, MaxMind GeoIP country-level access control, Checksum-based differential IP group sync (no Nginx reload), and custom block responses.
* **Anti-CC & Human-Machine Challenge (PoW)**: Built-in high-performance client-side cryptographic Proof of Work challenges (similar to Turnstile) to block and intercept botnets and scrapers at the gateway edge in seconds.
* **Pages Static Hosting**: Upload pre-built ZIP packages directly; edge Agents pull and serve them via local OpenResty, with SPA Fallback and built-in API reverse proxy configuration.
* **Automated TLS Certificate Management**: Supports dynamic certificate upload, automatic multi-domain certificate matching and binding, and ACME-based automatic issuance and renewal via Let's Encrypt.
* **Uptime Kuma Monitoring Sync**: Integrates with Uptime Kuma to automatically sync the monitoring site list using differential updates, providing real-time awareness of node availability and service health.
* **SSO Single Sign-On**: Supports GitHub OAuth and standard OIDC protocol for seamless integration with enterprise identity providers.
* **Unified Observability**: Aggregates node request metrics, real-time access log details, host/Nginx resource snapshots, health events, and a re-upload buffer for network fluctuations.
* **Reverse Proxy Configuration Management**: Uses website rules as the aggregation boundary, supports multi-domain binding and multi-upstream load balancing, and centrally manages reverse proxy configurations for all OpenResty nodes.
* **Secure In-Network Tunneling (Tunnels)**: Open-source version of Cloudflare Tunnels. No public IP or exposed inbound ports are required. Securely reverse-proxy internal web services to the public internet through Relay relay nodes and OpenFlared clients.
* **Edge WAF Security Protection**: Provides global and custom rule groups, supports manual/auto/subscription-type IP groups, MaxMind GeoIP national-level geographic access control, IP group member Checksum differential synchronization (no Nginx reload required), and custom blocking responses.
* **CC Defense and Human-Computer Challenge (PoW)**: Built-in high-performance client-side cryptography Proof of Work challenge (similar to Turnstile). Secures high-speed interception and blocking of zombie networks and crawlers at the gateway edge.
* **Pages Static Hosting**: Supports uploading or synchronizing pre-built artifacts from restricted Remote URLs or public GitHub Release assets. GitHub latest can be checked periodically and optionally auto-published. All sources are unified to generate immutable deployments, pulled by the edge Agent and served locally by OpenResty, supporting rollbacks, SPA Fallback, and API reverse proxy.
* **TLS Certificate Automation**: Supports dynamic certificate uploads, automatic multi-domain certificate matching and binding, and automatic issuance and renewal of certificates from Let's Encrypt via the ACME protocol.
* **Uptime Kuma Monitoring Synchronization**: Integrated with Uptime Kuma to automatically perform differential synchronization of monitoring site lists, real-time awareness of node availability and service status.
* **SSO Single Sign-On**: Supports GitHub OAuth and standard OIDC protocol for seamless integration with enterprise identity providers to achieve unified login.
* **Unified Observability**: Aggregates node request metrics, real-time access log details, host and Nginx resource snapshots, health events, and network fluctuation replenishment buffers.
## Interface Preview
### Dashboard Overview
![OpenFlare dashboard overview](./docs/assets/readme/dashboard-overview.png)
### Access Logs
![OpenFlare version release](./docs/assets/readme/domain_overview.png)
### WAF Protection
![OpenFlare version release](./docs/assets/readme/waf.png)
## Quick Start
### 1. Launch Server
### Hardware Configuration Recommendations
| Component | Minimum Hardware Requirements | Recommended Hardware Requirements | Notes |
|------------------------|-----------------------------------|-----------------------------------|-------|
| **Server Control Plane** | 1 CPU core / 2 GB RAM / 20 GB disk | 2 CPU cores / 4 GB RAM / 50 GB+ disk | Disk usage should be expanded reasonably based on access log retention duration and concurrent traffic |
| **Agent Data Plane** | 1 CPU core / 512 MB RAM / 2 GB disk | 2 CPU cores / 2 GB RAM / 10 GB+ disk | Expanded based on OpenResty concurrent proxy connections and WAF interception processing |
| **Relay Relay Node** | 1 CPU core / 1 GB RAM / 5 GB disk | 2 CPU cores / 2 GB RAM / 20 GB disk | frps transmission relay throughput is mainly limited by bandwidth and CPU throughput |
| **OpenFlared Client** | 1 CPU core / 256 MB RAM / 1 GB disk | 1 CPU core / 512 MB RAM / 5 GB disk | Runs independently on the internal network with extremely low resource consumption; only network throughput needs to be guaranteed |
### 1. Start the Server
Use `docker-compose`:
```bash
# Download environment variable template and create .env file
curl -o .env.example https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/.env.example
cp .env.example .env
```
```yaml
services:
@@ -70,8 +100,6 @@ services:
condition: service_healthy
redis:
condition: service_healthy
clickhouse:
condition: service_healthy
postgres:
image: postgres:17-alpine
@@ -101,118 +129,45 @@ services:
retries: 5
start_period: 5s
clickhouse:
image: clickhouse/clickhouse-server:25.3-alpine
restart: unless-stopped
environment:
CLICKHOUSE_DB: ${CLICKHOUSE_NAME:-openflare}
CLICKHOUSE_USER: ${CLICKHOUSE_USERNAME:-default}
CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1
TZ: ${TZ:-Asia/Shanghai}
volumes:
- openflare_clickhouse_data:/var/lib/clickhouse
healthcheck:
test: ["CMD", "clickhouse-client", "--user", "${CLICKHOUSE_USERNAME:-default}", "--password", "${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}", "--query", "SELECT 1"]
interval: 10s
timeout: 5s
retries: 5
start_period: 15s
volumes:
openflare_uploads:
openflare_postgres_data:
openflare_redis_data:
openflare_clickhouse_data:
```
```bash
docker compose up -d
```
See the [deployment documentation](https://open-flare.pages.dev/deployment/deployment) for details.
Access at: `http://localhost:3000`
Access address: `http://localhost:3000`
Default credentials:
Default account:
* Username: `root`
* Password: `123456`
* Username: `admin`
* Password: `12345678`
### 2. Install Agent
Before installing an Agent, please install OpenResty on the target node first, or use the Agent Docker image with OpenResty built-in.
Before installing the Agent, first install OpenResty on the node or use the built-in OpenResty Agent Docker image.
You can copy the installation command from **Node Management -> Details -> Node Info -> Node Token & Deployment** in the control panel, or directly use the scripts below:
You can copy the installation command from the control panel's **Nodes Management -> Details -> Node Information -> Node ID and Deployment**, or use the script below:
#### Docker Deployment
For Docker deployment, you can directly run the Agent image:
Docker deployment can directly run the Agent image:
```bash
docker pull ghcr.io/rain-kl/openflare-agent:latest
docker rm -f openflare-agent 2>/dev/null || true
docker run -d --name openflare-agent --restart unless-stopped \
-p 80:80 -p 443:443/tcp -p 443:443/udp \
-v openflare-agent-pages:/data/var/lib/openflare/pages \
-e OPENFLARE_SERVER_URL=http://your-server:3000 \
-e OPENFLARE_AGENT_TOKEN=YOUR_AGENT_TOKEN \
ghcr.io/rain-kl/openflare-agent:latest
```
#### Local Installation
## Open Source License
Using `discovery_token` to register:
```bash
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
--server-url http://your-server:3000 \
--discovery-token YOUR_DISCOVERY_TOKEN
```
Using node-specific `agent_token`:
```bash
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/install-agent.sh | bash -s -- \
--server-url http://your-server:3000 \
--agent-token YOUR_AGENT_TOKEN
```
The installation script defaults to `/opt/openflare-agent`, creates a `openflare-agent.service`, automatically searches for `openresty`, and can be executed repeatedly to reinstall or upgrade the Agent.
### 3. Uninstall Agent
To completely uninstall the Agent and clear local data, run:
```bash
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/uninstall-agent.sh | bash
```
The uninstallation script will stop and remove the `openflare-agent.service`, and delete the entire `/opt/openflare-agent` directory. It will not delete the local OpenResty installation.
### 4. Publish Your First Configuration
1. Log in to the management panel and add a reverse proxy rule.
2. View the preview or change summary before publishing.
3. Activate the new version.
4. Agents will receive the configuration and apply it via WebSocket notification or subsequent heartbeats.
The version number format is fixed as `YYYYMMDD-NNN`. Historical versions are immutable, and rollback is achieved by reactivating an older version.
## UI Preview
### Dashboard Overview
![OpenFlare dashboard overview](./docs/assets/readme/dashboard-overview.png)
### Node Details
![OpenFlare node detail](./docs/assets/readme/node-detail.png)
### Proxy Configuration
![OpenFlare version release](./docs/assets/readme/proxy-route-detail.png)
## License
This project is licensed under [Apache License 2.0](./LICENSE).
This project is licensed under the [Apache License 2.0](./LICENSE).
## Star History
+13 -35
View File
@@ -54,16 +54,25 @@ OpenFlare 是开源 CDN 编排与边缘安全平台。它支持反向代理、
![OpenFlare dashboard overview](./docs/assets/readme/dashboard-overview.png)
### 节点详情
### 访问日志
![OpenFlare node detail](./docs/assets/readme/node-detail.png)
![OpenFlare version release](./docs/assets/readme/domain_overview.png)
### 配置新增
### WAF 防护
![OpenFlare version release](./docs/assets/readme/proxy-route-detail.png)
![OpenFlare version release](./docs/assets/readme/waf.png)
## 快速开始
### 硬件配置推荐
| 组件 | 最低硬件配额 | 推荐硬件配额 | 说明 |
| --- |-------------------------------| --- | --- |
| **Server 控制面** | 1 核 CPU / 2 GB 内存 / 20 GB 磁盘 | 2 核 CPU / 4 GB 内存 / 50 GB+ 磁盘 | 磁盘用量需根据访问日志留存时长与并发流量合理扩容 |
| **Agent 数据面** | 1 核 CPU / 512 MB 内存 / 2 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 10 GB+ 磁盘 | 根据 OpenResty 的并发代理连接量与 WAF 拦截处理扩容 |
| **Relay 中继节点**| 1 核 CPU / 1 GB 内存 / 5 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 20 GB 磁盘 | frps 传输中继吞吐量主要受带宽与 CPU 吞吐能力限制 |
| **OpenFlared 客户端**| 1 核 CPU / 256 MB 内存 / 1 GB 磁盘 | 1 核 CPU / 512 MB 内存 / 5 GB 磁盘 | 独立运行于内网,自身资源占用极小,保障网络吞吐即可 |
### 1. 启动 Server
使用 docker-compose
@@ -72,11 +81,6 @@ OpenFlare 是开源 CDN 编排与边缘安全平台。它支持反向代理、
# 下载环境变量模板并创建 .env 文件
curl -o .env.example https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/.env.example
cp .env.example .env
# ClickHouse 服务端:curl performance.xml 到 ./config/clickhouse,并以单文件方式挂载到 config.d
mkdir -p ./config/clickhouse
curl -fsSL -o ./config/clickhouse/performance.xml \
https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/config/clickhouse/performance.xml
```
```yaml
@@ -96,8 +100,6 @@ services:
condition: service_healthy
redis:
condition: service_healthy
clickhouse:
condition: service_healthy
postgres:
image: postgres:17-alpine
@@ -127,34 +129,10 @@ services:
retries: 5
start_period: 5s
clickhouse:
image: clickhouse/clickhouse-server:25.3-alpine
restart: unless-stopped
environment:
CLICKHOUSE_DB: ${CLICKHOUSE_NAME:-openflare}
CLICKHOUSE_USER: ${CLICKHOUSE_USERNAME:-default}
CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1
TZ: ${TZ:-Asia/Shanghai}
ulimits:
nofile:
soft: 262144
hard: 262144
volumes:
- openflare_clickhouse_data:/var/lib/clickhouse
- ./config/clickhouse/performance.xml:/etc/clickhouse-server/config.d/performance.xml:ro
healthcheck:
test: ["CMD", "clickhouse-client", "--user", "${CLICKHOUSE_USERNAME:-default}", "--password", "${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}", "--query", "SELECT 1"]
interval: 10s
timeout: 5s
retries: 5
start_period: 15s
volumes:
openflare_uploads:
openflare_postgres_data:
openflare_redis_data:
openflare_clickhouse_data:
```
详细部署说明见 [部署文档](https://open-flare.pages.dev/deployment/deployment)。
+4 -2
View File
@@ -99,10 +99,12 @@ otel:
tracer_name: "github.com/Rain-kl/OpenFlare" # Global tracer instrumentation name
# ─── ClickHouse (required) ──────────────────────────────────────────────────────
# ─── ClickHouse (optional) ─────────────────────────────────────────────────────
# Analytics / observability OLAP store. Telemetry writes are best-effort (async batch).
# 默认关闭:缺失本配置块或 enabled: false 时不启用 ClickHouse,日志/指标由主库承担;
# 设置 CLICKHOUSE_HOST 或 CLICKHOUSE_ENABLED=true 可经环境变量启用。
clickhouse:
enabled: true
enabled: false
hosts:
- "127.0.0.1:9000" # compose 内应用可用 clickhouse:9000(经 CLICKHOUSE_HOST)
username: "default"
Binary file not shown.

After

Width:  |  Height:  |  Size: 58 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 67 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 77 KiB

+33 -16
View File
@@ -11,32 +11,49 @@ sidebar: false
## 重大变更
> [!IMPORTANT]
>
> 3.1.2 版本更新了 CLickHouse 部署配置。
>
> 3.0.0 版本为 Wavelet 平台迁移与架构重构版本,涉及数据库表结构、环境变量以及前后端底层架构的重大变更。请务必在升级前备份数据库,并且更新到 V2.3.4。
> 目前已知的兼容性问题:
>
> - Pages 无法迁移, 升级前请先手动下载并备份 Pages 静态站点的 ZIP 包,升级后重新创建。
> - 性能调优参数重置, 升级后请重新配置
>
>3.5.1 版本解耦了日志存储,ClickHouse 变为可选项,如果想切换数据库, 点击 「任务管理」 -> 「切换日志数据库」任务,按提示迁移数据并切换主库。
>
## [Unreleased]
## [v3.5.0] - 2026-08-08
### 🛠 修复
- 修复源站错误页「仅针对 GET 请求」未生效:`error_page` 内部重定向会把请求方法改写成 GET,导致内部 Lua 无法识别 POST/PUT 等原始方法、仍返回自定义错误页;现改为命名 location(`@__openflare_origin_error`)承载错误页,保留原始请求方法与错误状态码,非 GET 请求不再返回自定义错误页。
- 修复 PostgreSQL 作为日志库时节点访问日志/可观测指标/用户访问日志批量写入失败:GORM 对零值 `uint64` 主键会省略 `id` 列,而 PG 日志表 `id` 无默认值,导致持续报「null value in column id violates not-null constraint」;现于落库前为零 ID 行生成雪花 ID(与 ClickHouse 写入路径一致),并新增回归测试覆盖六张日志表。
## [v3.5.1] - 2026-08-09
### 新增
- 日志存储解耦:ClickHouse 变为可选项,不启用时由 PostgreSQL/SQLite 承担全部日志功能;新增「切换日志数据库」任务支持 PostgreSQL/SQLite 与 ClickHouse 间数据迁移(迁移期间冻结日志写入,成功后自动切换主库并保留源数据);`log_database` / `log_db_migration` 设为受保护配置;ClickHouse 改为默认关闭。
- 新增 PostgreSQL/SQLite 日志存储实现:节点访问日志按月分区,统计查询合并为单次扫描、IP 汇总归属地取查询窗口内最新记录、WAF 按 IP 聚合减少扫描次数,并新增 `logged_at` 前导索引与主机名小写表达式索引;过期清理直接删除完全过期的整月分区,启动时兜底预建当月及未来 2 个月分区。
- 性能指标与访问日志的保留时长解耦:新增三库共用的 `metric_retention_days` 配置(默认 3 天),每日垃圾清理按独立短留存清理指标快照。
- 支持 Service Worker 离线兜底:为启用 HTTPS 的网站下发 Service Worker 并缓存离线页,域名无法访问时浏览器展示离线兜底页面,减少用户流失。可指定生效域名范围(仅对选中的 HTTPS 域名生效),配置位于「响应页面」-「离线页」,可在版本发布中批量生效。
### 🛠 修复
- 修复 UptimeKuma 同步调试日志泄露凭据:Socket.IO 事件日志不再打印 payload 内容(仅记录长度),避免凭据进入日志。
- 修复日志保留天数配置继承旧键导致的误删风险:`log_retention_days_*` 不再继承 `database_auto_cleanup_retention_days`,统一默认 30 天。
### 修复
### ⚡️ 优化与改进
- 系统定期垃圾清理由每 2 小时改为每日执行一次(凌晨 3 点,Asia/Shanghai),降低非必要高频扫描。
- 修复 PoW 挑战页潜在 XSS:错误提示与状态文案改用纯文本渲染,挑战通过后的 `redir` 跳转参数仅允许 http/https 协议,防止异常文本被当作 HTML 执行或跳转到危险协议。
- 修复邮件发送的邮件头注入风险:标题、发件人、收件人在写入邮件头前清除 CR/LF 换行符,防止注入额外邮件头(CWE-93)。
- 修复 UptimeKuma 同步调试日志泄露凭据:输出日志前对 password/token/secret 等敏感字段打码,避免凭据进入日志。
### 💄 其他/体验
- 服务工作者(SW)注入挑战页改为前台无感知:不再显示「加载中…」文案,页面空白,仅通过浏览器控制台输出 `[sw-challenge]` 调试信息,注入过程不打扰访客。
- 用户访问日志(`w_user_access_logs`)记录禁用:不再采集与写入新的用户访问日志,存量数据与管理端访问日志统计页面保留。
### 改进
- 离线页预制模板支持:新增离线页内置预制模板套件(「极简白底」、「线框拓扑」、「包豪斯」),与源站错误页模板风格保持一致,可在编辑界面一键加载与预览。
## [v3.5.0] - 2026-08-08
### 🛠 修复
- 修复 PoW 挑战页潜在 XSS 风险,状态与错误文案改用纯文本渲染,并限制跳转 URL 仅允许 http/https 协议。
- 修复邮件发送的邮件头注入风险,写入邮件头前自动清除 CR/LF 换行符(CWE-93)。
- 修复 UptimeKuma 同步调试日志泄露凭据问题,输出日志前对密码和 Token 等敏感字段打码。
### ⚡️ 优化与改进
- 新增 Service Worker 离线兜底功能,为启用 HTTPS 的网站自动下发 Service Worker 并缓存离线页,域名不可达时展示离线兜底页面。
- 重构响应页面设置,将源站错误页与 Service Worker 离线页整合至统一的「响应页面」(/responses)标签页,并增加 URL 查询参数 tab 状态同步。
### 💄 其他/体验
- 新增离线页内置预制模板套件(「极简白底」、「线框拓扑」、「包豪斯」),与源站错误页模板风格保持一致,支持编辑界面一键加载与预览。
## [v3.4.5] - 2026-08-08
+7 -80
View File
@@ -136,55 +136,6 @@ docker run -d --name openflare-agent --restart unless-stopped \
> **Pages 持久化**
> 默认将 Pages 部署目录挂载到 Docker 命名卷 `openflare-agent-pages`(容器内路径 `/data/var/lib/openflare/pages`)。重建或升级 Agent 容器时无需重新拉取静态站点包。
> [!NOTE]
> **非 Root 安全加固运行**
> Agent 容器内部已完成安全加固,在启动后会统一以低权限非 root 用户 `openflare` 运行。
> 容器已内置了 `cap_net_bind_service` 内核能力,使得低权限进程依然能够正常监听宿主机的 `80` 和 `443` 特权端口。
> 同时,OpenResty 运行时所需的各种临时路径(包括 PID 路径、各类临时缓存目录如 `client_body_temp_path`、`proxy_temp_path` 等)都由 Agent 控制器动态渲染并自动重定向至容器内的 `/data` 目录,彻底避免在非 root 权限运行时写入默认系统路径而导致的权限拒绝错误(Permission Denied)。
> 具体物理缓存写入路径为:
> * 临时缓存目录:`/data/var/cache/nginx`
> * 代理缓存目录:`/data/var/cache/openflare_proxy`
## 启动与验证
systemd 环境:
```bash
systemctl status openflare-agent
journalctl -u openflare-agent -f
```
手动启动:
```bash
/opt/openflare-agent/openflare-agent -config /opt/openflare-agent/agent.json
```
源码运行:
```bash
export LOG_LEVEL='info'
go run ./cmd/agent -config /path/to/agent.json
```
编译后二进制运行:
```bash
go build -o openflare-agent ./cmd/agent
export LOG_LEVEL='info'
./openflare-agent -config /path/to/agent.json
```
在管理端确认:
| 位置 | 期望结果 |
| --- | --- |
| 节点列表 | 节点在线 |
| 节点详情 | 能看到心跳时间、当前版本和基础资源信息 |
| 应用记录 | 发布配置后出现应用结果 |
## 卸载
### 交互式卸载 (推荐)
@@ -195,38 +146,14 @@ export LOG_LEVEL='info'
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/uninstall-agent.sh | bash
```
### 自动化 (非交互式) 卸载
### 卸载
使用命令行传参进行无人值守卸载。
本地卸载(默认):
```bash
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/uninstall-agent.sh | bash -s -- --install-dir /opt/openflare-agent
```
Docker 容器卸载:
```bash
curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/uninstall-agent.sh | bash -s -- --docker
```
支持参数:
| 参数 | 说明 |
| --- | --- |
| `--install-dir` | 安装目录,默认 `/opt/openflare-agent`(仅本地卸载生效) |
| `--service-name` | systemd 服务名,默认 `openflare-agent`(仅本地卸载生效) |
| `--docker` | 使用 Docker 容器方式卸载 |
| `--method` | 卸载方式,可选 `local` 或 `docker`(默认 `local`) |
本地卸载只会移除 Agent 服务、进程和安装目录,不会删除本机 OpenResty。Docker 卸载会停止并删除 `openflare-agent` 容器,交互模式下还可以选择是否清理对应的 Docker 镜像。
停止并删除 `openflare-agent` 容器即可
## 常见问题
| 现象 | 处理步骤 |
| --- | --- |
| `agent_token 和 discovery_token 不能同时为空` | 检查 `agent.json` 至少配置了一个 Token |
| 节点一直离线 | 在 Agent 节点执行 `curl -I http://your-server:3000`,确认 Server 地址可达 |
| OpenResty 没有启动 | 查看 `journalctl -u openflare-agent`,确认 `openresty_path` 可执行,80/443 端口未被占用,且运行用户(如 `openflare`)对数据目录具有读写权限 |
| 发布后重复失败 | Agent 会阻断同一 `version + checksum` 的重复应用;需要修正配置后重新发布,或激活旧版本回滚 |
| 现象 | 处理步骤 |
| --- |---------------------------------------------------------------------------------------------------------|
| `agent_token 和 discovery_token 不能同时为空` | 检查 `agent.json` 至少配置了一个 Token |
| 节点一直离线 | 在 Agent 节点执行 `curl -I http://your-server:3000`,确认 Server 地址可达 |
| 发布后重复失败 | Agent 会阻断同一 `version + checksum` 的重复应用;在节点尝试强制同步,或者重新发布版本 |
+4 -62
View File
@@ -47,33 +47,13 @@ Internal Service (192.168.x.x)
## 前置条件
Server:
| 项目 | 要求 |
| --- | --- |
| Go | `1.25+`,仅源码运行需要 |
| Node.js | `18+`,仅源码构建管理端需要 |
| 数据库 | 可写 SQLite 文件目录,或可访问的 PostgreSQL 实例 |
| 端口 | 默认监听 `3000` |
Agent:
| 项目 | 要求 |
| --- | --- |
| 系统 | 安装脚本支持 Linux 和 macOS;systemd 服务仅在 Linux + systemd 环境创建 |
| 架构 | `amd64` 或 `arm64` |
| OpenResty | 本地部署需要可执行 `openresty`,或通过 `--openresty-path` 指定路径 |
| Docker | 仅 Docker 部署 Agent 镜像时需要 |
| 网络 | Agent 节点必须能访问 Server 地址 |
| GeoIP | WAF 地域规则使用 OpenResty 读取本地 MaxMind mmdb;镜像内置文件或首次下载,Agent 负责周期更新 |
### 硬件配置推荐
| 组件 | 最低硬件配额 | 推荐硬件配额 | 说明 |
| --- | --- | --- | --- |
| **Server 控制面** | 1 核 CPU / 1 GB 内存 / 10 GB 磁盘 | 2 核 CPU / 4 GB 内存 / 50 GB+ 磁盘 | 磁盘用量需根据访问日志留存时长与并发流量合理扩容 |
| 组件 | 最低硬件配额 | 推荐硬件配额 | 说明 |
| --- |-------------------------------| --- | --- |
| **Server 控制面** | 1 核 CPU / 2 GB 内存 / 20 GB 磁盘 | 2 核 CPU / 4 GB 内存 / 50 GB+ 磁盘 | 磁盘用量需根据访问日志留存时长与并发流量合理扩容 |
| **Agent 数据面** | 1 核 CPU / 512 MB 内存 / 2 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 10 GB+ 磁盘 | 根据 OpenResty 的并发代理连接量与 WAF 拦截处理扩容 |
| **Relay 中继节点**| 1 核 CPU / 1 GB 内存 / 5 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 20 GB 磁盘 | frps 传输中继吞吐量主要受带宽与 CPU 吞吐能力限制 |
| **Relay 中继节点**| 1 核 CPU / 1 GB 内存 / 5 GB 磁盘 | 2 核 CPU / 2 GB 内存 / 20 GB 磁盘 | frps 传输中继吞吐量主要受带宽与 CPU 吞吐能力限制 |
| **OpenFlared 客户端**| 1 核 CPU / 256 MB 内存 / 1 GB 磁盘 | 1 核 CPU / 512 MB 内存 / 5 GB 磁盘 | 独立运行于内网,自身资源占用极小,保障网络吞吐即可 |
## Docker Compose 部署 Server
@@ -169,41 +149,3 @@ curl -fsSL https://raw.githubusercontent.com/Rain-kl/OpenFlare/main/scripts/inst
systemctl status openflare-agent
journalctl -u openflare-agent -f
```
## 手动运行 Agent
源码运行:
```bash
export LOG_LEVEL='info'
go run ./cmd/agent -config /path/to/agent.json
```
编译后二进制运行:
```bash
go build -o openflare-agent ./cmd/agent
export LOG_LEVEL='info'
./openflare-agent -config /path/to/agent.json
```
最小 `agent.json` 示例:
```json
{
"server_url": "http://127.0.0.1:3000",
"agent_token": "replace-with-node-auth-token",
"data_dir": "./data",
"openresty_path": "openresty",
"heartbeat_interval": 3000,
"request_timeout": 10000
}
```
未配置 `openresty_path` 时,Agent 默认调用 `openresty`。
默认情况下,Agent 在 HTTP 心跳成功后会尝试升级为 WebSocket。升级成功时,Server 发布或激活配置会立即通知 Agent;如果 WebSocket 无法建立或意外断开,Agent 会自动退回 HTTP 心跳同步。
WAF 地域规则依赖 Agent 本地 `GeoLite2-Country.mmdb` / `GeoLite2-City.mmdb`(OpenResty `resty.maxminddb` 读磁盘路径)。Docker 镜像会将 MMDB COPY 到 `data_dir/etc/openflare/`;裸二进制安装时若文件缺失则首次启动按配置 URL 下载。Agent 按配置周期尝试更新;更新失败只记录警告,不影响配置同步与 OpenResty reload。MMDB **不**再嵌入 agent 二进制。Server 控制面可选 MaxMind 提供方仍**仅内嵌 Country**(约 9MB,不含 City)用于离线 seed。
+1 -35
View File
@@ -34,7 +34,7 @@
---
## Docker 运行(推荐)
## Docker 运行
Docker 部署是内网运行最简单也最安全的方式。官方的 `openflared` 镜像已经内置了客户端控制器以及 `frpc v0.69.0` 二进制运行时,无需额外搭建环境。
@@ -51,40 +51,6 @@ docker run -d --name openflared --restart unless-stopped \
---
## 宿主机手动运行
如果您需要直接在内网的 Linux/macOS/Windows 宿主机上独立运行:
### 1. 编译二进制
```bash
go build -o bin/flared ./cmd/flared
```
### 2. 准备 `flared.json`
在程序同级目录下创建 `flared.json` 配置文件:
```json
{
"server_url": "http://your-server-ip:3000",
"tunnel_token": "your-tunnel-auth-token",
"frpc_path": "/usr/local/bin/frpc",
"data_dir": "./data",
"heartbeat_interval": "10s",
"sync_interval": "30s"
}
```
### 3. 运行服务
```bash
export LOG_LEVEL='info'
./flared -config ./flared.json
```
---
## 启动与验证
### 1. 自动同步逻辑
+3 -41
View File
@@ -15,7 +15,7 @@
- 必须确保 `bindPort`(frpc 连接端口,默认 `7000`)可被公网/内网客户端访问。
- 必须确保 `vhostHTTPPort`(HTTP Vhost 端口,默认 `8080`)处于空闲状态,Agent 将在此端口上与 frps 进行流量传递。
3. **软件依赖**(仅限宿主机直接部署):
- 本地需有可执行的 `frps` 二进制文件(建议版本为 `v0.61.0+` 或最新稳定版 `v0.69.0`),或通过参数显式指定路径。
- 本地需有可执行的 `frps` 二进制文件,或通过参数显式指定路径。
---
@@ -40,9 +40,9 @@
---
## Docker 运行(推荐)
## Docker 运行)
Docker 运行是 TunnelRelay 节点最便捷的部署方案。官方镜像内置了 `openflare-relay` 控制器与 `frps v0.69.0` 运行时,开箱即用。
Docker 运行是 TunnelRelay 节点最便捷的部署方案。官方镜像内置了 `openflare-relay` 控制器与 `frps` 运行时,开箱即用。
```bash
docker pull ghcr.io/rain-kl/openflare-relay:latest
@@ -67,39 +67,6 @@ docker run -d --name openflare-relay --restart unless-stopped \
---
## 宿主机手动运行
如果您倾向于在物理机或虚拟机上直接运行:
### 1. 编译二进制
```bash
go build -o bin/openflare-relay ./cmd/relay
```
### 2. 准备 `relay.json`
在程序同级目录下创建 `relay.json` 配置文件:
```json
{
"server_url": "http://127.0.0.1:3000",
"agent_token": "your-relay-node-agent-token",
"frps_path": "/usr/local/bin/frps",
"data_dir": "./data",
"heartbeat_interval": "10s",
"request_timeout": "10s"
}
```
### 3. 运行服务
```bash
export LOG_LEVEL='info'
./openflare-relay -config ./relay.json
```
---
## 启动与验证
@@ -110,11 +77,6 @@ export LOG_LEVEL='info'
docker logs -f openflare-relay
```
如果是在 Linux 上通过 Systemd 托管的,可执行:
```bash
journalctl -u openflare-relay -f
```
### 2. 验证运行状态
启动成功后,Relay 将进行以下工作:
+13 -151
View File
@@ -6,7 +6,8 @@ OpenFlare Server 是 Gin + GORM 单体控制面,负责管理端 UI、管理 AP
> [!IMPORTANT]
> **关于外部依赖**:
> OpenFlare 系统内建了对后台异步任务(Asynq 框架)及海量节点日志分析与度量指标(观测面板)的支持。因此,**无论采用何种部署模式,系统都必须依赖 Redis(或 Valkey)与 ClickHouse 的运行**。各个部署方案的主要差异在于主关系型数据库的选择(SQLite vs PostgreSQL)以及是否启用链路追踪服务(Jaeger)。
> OpenFlare 系统内建了对后台异步任务(Asynq 框架)的支持。因此,**无论采用何种部署模式,系统都必须依赖 Redis(或 Valkey)**。各个部署方案的主要差异在于主关系型数据库的选择(SQLite vs PostgreSQL)以及是否启用链路追踪服务(Jaeger)。
> 若业务流量过大, 建议使用 ClickHouse 存储日志。
> [!TIP]
> **ClickHouse 服务端性能配置(推荐挂载)**
@@ -37,11 +38,11 @@ volumes:
使用 Docker 部署可以免去本地配置 Go 与 Node.js 前端构建环境的麻烦。根据你的服务器硬件配置及业务需求,你可以选择以下三种方案之一:
### 1. 快速启动 (SQLite + Redis + ClickHouse)
### 1. 快速启动 (SQLite + Redis)
> **适用场景**:测试体验、轻量化单机部署。
>
> **特点**:主关系型数据库使用内建的 SQLite 文件
> **特点**:主关系型数据库使用 SQLite
创建 `docker-compose.yaml` 文件:
@@ -70,8 +71,6 @@ services:
depends_on:
redis:
condition: service_healthy
clickhouse:
condition: service_healthy
redis:
image: valkey/valkey:8.0-alpine
@@ -84,47 +83,13 @@ services:
interval: 10s
timeout: 5s
retries: 5
clickhouse:
image: clickhouse/clickhouse-server:25.3-alpine
restart: unless-stopped
environment:
CLICKHOUSE_DB: openflare
CLICKHOUSE_USER: default
CLICKHOUSE_PASSWORD: 123456
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1
TZ: Asia/Shanghai
ulimits:
nofile:
soft: 262144
hard: 262144
volumes:
- ./data/clickhouse_data:/var/lib/clickhouse
- ./config/clickhouse/performance.xml:/etc/clickhouse-server/config.d/performance.xml:ro
healthcheck:
test: ["CMD", "clickhouse-client", "--user", "default", "--password", "123456", "--query", "SELECT 1"]
interval: 10s
timeout: 5s
retries: 5
start_period: 15s
```
运行启动命令:
```bash
mkdir -p ./config/clickhouse
curl -fsSL -o ./config/clickhouse/performance.xml \
https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/config/clickhouse/performance.xml
docker compose up -d
```
---
### 2. 生产推荐 (PostgreSQL + Redis + ClickHouse)
### 2. 小流量业务场景 (PostgreSQL + Redis)
> **适用场景**:生产环境、多节点集群管理、高并发高可用要求。
>
> **特点**:完全分层架构。启用专用的 PostgreSQL 服务作为主关系数据库,Redis 负责高并发分布式锁、会话缓存与异步队列,ClickHouse 承载海量日志异步 Flush 与观测指标。
> **适用场景**:生产环境、业务流量中小, PostgreSQL 不会成为日志记录的瓶颈。
创建 `docker-compose.yaml` 文件:
@@ -145,8 +110,6 @@ services:
condition: service_healthy
redis:
condition: service_healthy
clickhouse:
condition: service_healthy
postgres:
image: postgres:17-alpine
@@ -176,45 +139,18 @@ services:
retries: 5
start_period: 5s
clickhouse:
image: clickhouse/clickhouse-server:25.3-alpine
restart: unless-stopped
environment:
CLICKHOUSE_DB: ${CLICKHOUSE_NAME:-openflare}
CLICKHOUSE_USER: ${CLICKHOUSE_USERNAME:-default}
CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1
TZ: ${TZ:-Asia/Shanghai}
ulimits:
nofile:
soft: 262144
hard: 262144
volumes:
- openflare_clickhouse_data:/var/lib/clickhouse
- ./config/clickhouse/performance.xml:/etc/clickhouse-server/config.d/performance.xml:ro
healthcheck:
test: ["CMD", "clickhouse-client", "--user", "${CLICKHOUSE_USERNAME:-default}", "--password", "${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}", "--query", "SELECT 1"]
interval: 10s
timeout: 5s
retries: 5
start_period: 15s
volumes:
openflare_uploads:
openflare_postgres_data:
openflare_redis_data:
openflare_clickhouse_data:
```
创建对应的 `.env` 文件来配置系统环境变量(可复制并修改根目录下的 `.env.example`):
```bash
mkdir -p ./config/clickhouse
curl -fsSL -o ./config/clickhouse/performance.xml \
https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/config/clickhouse/performance.xml
curl -o .env.example https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/.env.example
cp .env.example .env
# 编辑 .env 文件,填入对应的数据库、Redis、ClickHouse 连接地址、密码与 APP_SESSION_SECRET
# 编辑 .env 文件,填入对应的数据库、Redis、密码与 APP_SESSION_SECRET
docker compose up -d
```
@@ -223,9 +159,9 @@ docker compose up -d
### 3. 进阶版 (含 Jaeger 链路追踪的完整编排)
> **适用场景**:开发者调试、系统深度性能诊断、高级可观测性追溯。
> **适用场景**:大流量场景, 需要进行链路性能指标追踪。
>
> **特点**:在“生产推荐”全家桶的基础上,联动拉起 Jaeger 作为 OpenTelemetry (OTel) 链路追踪的后端,收集 Server 运行时各个 API 请求的 Span Trace 信息。
> **特点**:在“生产推荐”全家桶的基础上,使用 ClickHouse 存储日志, 联动 Jaeger 作为 OpenTelemetry (OTel) 链路追踪的后端。
创建 `docker-compose.yaml` 文件:
@@ -340,64 +276,6 @@ docker compose up -d
---
## 方式二:本地部署 (源码/二进制启动)
如果你不希望使用 Docker,也可以直接在本地或虚拟机上从源码构建和运行 Server。由于后台异步任务和可观测指标分析为系统核心防线,**本地部署时依然需要连接外部 Redis 与 ClickHouse 实例**。
### 前置条件
| 项目 | 要求 |
| --- | --- |
| Go | `1.25+` |
| Node.js | `18+` |
| pnpm | 推荐通过 `corepack enable` 使用项目声明的 pnpm |
| 外部服务 | 必须在本地或远端运行 Redis (Valkey) 和 ClickHouse 实例;ClickHouse 建议挂载仓库提供的 `performance.xml`(见上文「ClickHouse 服务端性能配置」) |
### 1. 构建管理端前端
Go Server 运行时需要嵌入前端静态资源。编译 Go 二进制前需要先构建前端静态产物并输出到 Go 服务目录:
```bash
cd frontend
corepack enable
pnpm install
pnpm build:embed
cd ..
```
> **常用前端代码检查命令**:
> * `pnpm lint`
> * `pnpm typecheck`
### 2. 使用 SQLite 启动
关系数据库存储在本地 SQLite 文件,但依然需要提供 Redis 和 ClickHouse 连接配置:
```bash
cp config.example.yaml config.yaml
# 编辑 config.yaml:
# 1. 设置 app.session_secret 为一个随机的长字符串
# 2. 将 database.enabled 设为 false 以启用内置 SQLite
# 3. 将 redis.addrs 与 clickhouse.hosts 修改为你的本地/局域网服务连接信息
# 启动 Server(默认融合模式)
go run main.go all
```
### 3. 使用 PostgreSQL 启动
```bash
cp config.example.yaml config.yaml
# 编辑 config.yaml:
# 1. 设置 app.session_secret
# 2. 将 database.enabled 设为 true,并完整设置 database.*、redis.*、clickhouse.* 字段连接参数
# 启动 Server(默认融合模式)
go run main.go all
```
---
## 首次登录
Server 默认监听 `3000` 端口,启动成功后可以使用浏览器访问:`http://localhost:3000`。
@@ -413,28 +291,12 @@ Server 默认监听 `3000` 端口,启动成功后可以使用浏览器访问
---
## 常用运维指南
### 1. 命令行子服务分进程启动
## 分布式部署
在大型生产部署中,你可以选择将 Server 按职责拆分为多个进程运行:
```bash
go run main.go api # 仅启动管理端与节点通信的 API 服务
go run main.go worker # 仅启动后台任务的 Worker 服务
go run main.go scheduler # 仅启动定时任务的 Scheduler 服务
go run main.go all # 融合模式(在一进程内运行上述所有服务,默认)
```
### 2. 状态验证
```bash
# 验证编译是否通过
go build ./...
# 运行内部单元测试
go test ./internal/apps/openflare/... -count=1
# 检查服务健康状态
curl http://127.0.0.1:3000/api/v1/d/status
go run main.go api # 仅启动管理端与节点通信的 API 服务
go run main.go worker # 仅启动后台任务的 Worker 服务
go run main.go scheduler # 仅启动定时任务的 Scheduler 服务
```
+89 -167
View File
@@ -1024,7 +1024,7 @@ const docTemplate = `{
"SessionCookie": []
}
],
"description": "分页并按照用户、接口路径、时间范围等维度检索 ClickHouse 用户访问日志列表(需要管理员权限,ClickHouse 未启用时报错)",
"description": "分页并按照用户、接口路径、时间范围等维度检索用户访问日志列表(需要管理员权限,日志存储未启用时报错)",
"produces": [
"application/json"
],
@@ -1092,7 +1092,7 @@ const docTemplate = `{
}
},
"400": {
"description": "ClickHouse 未启用或参数错误",
"description": "日志存储未启用或参数错误",
"schema": {
"$ref": "#/definitions/response.Any"
}
@@ -1119,7 +1119,7 @@ const docTemplate = `{
"SessionCookie": []
}
],
"description": "聚合统计最近 7 天的每日访问趋势、浏览器分布以及前 10 名最活跃用户排行(需要管理员权限,ClickHouse 未启用时报错)",
"description": "聚合统计最近 7 天的每日访问趋势、浏览器分布以及前 10 名最活跃用户排行(需要管理员权限,日志存储未启用时报错)",
"produces": [
"application/json"
],
@@ -1147,7 +1147,7 @@ const docTemplate = `{
}
},
"400": {
"description": "ClickHouse 未启用",
"description": "日志存储未启用",
"schema": {
"$ref": "#/definitions/response.Any"
}
@@ -1877,21 +1877,21 @@ const docTemplate = `{
}
}
},
"/api/v1/admin/status/clickhouse": {
"/api/v1/admin/status/log-database": {
"get": {
"security": [
{
"SessionCookie": []
}
],
"description": "返回 ClickHouse parts、mutation、async_insert 队列及进程内 batch writer 指标,需要管理员权限",
"description": "返回当前日志主库、迁移状态、各库保留天数与合法迁移目标,需要管理员权限",
"produces": [
"application/json"
],
"tags": [
"admin"
],
"summary": "获取 ClickHouse 运行指标",
"summary": "获取日志数据库状态",
"responses": {
"200": {
"description": "获取成功",
@@ -1904,19 +1904,13 @@ const docTemplate = `{
"type": "object",
"properties": {
"data": {
"$ref": "#/definitions/analytics.ClickHouseOperationalStats"
"$ref": "#/definitions/status.LogDatabaseStatus"
}
}
}
]
}
},
"400": {
"description": "ClickHouse 未启用",
"schema": {
"$ref": "#/definitions/response.Any"
}
},
"401": {
"description": "未登录",
"schema": {
@@ -2393,6 +2387,18 @@ const docTemplate = `{
"name": "task_type",
"in": "query"
},
{
"type": "string",
"description": "任务类型前缀筛选(与 task_type / task_types 互斥,精确类型优先)",
"name": "task_type_prefix",
"in": "query"
},
{
"type": "string",
"description": "逗号分隔的精确任务类型列表(IN 筛选,优先于前缀)",
"name": "task_types",
"in": "query"
},
{
"type": "integer",
"default": 1,
@@ -8305,80 +8311,6 @@ const docTemplate = `{
}
}
},
"/api/v1/d/option/database/cleanup": {
"post": {
"security": [
{
"SessionCookie": []
}
],
"description": "按目标与保留天数清理可观测性相关数据表,需要管理员权限",
"consumes": [
"application/json"
],
"produces": [
"application/json"
],
"tags": [
"openflare-option"
],
"summary": "清理可观测性数据库",
"parameters": [
{
"description": "清理参数",
"name": "request",
"in": "body",
"schema": {
"$ref": "#/definitions/option.databaseCleanupInput"
}
}
],
"responses": {
"200": {
"description": "清理结果",
"schema": {
"allOf": [
{
"$ref": "#/definitions/response.Any"
},
{
"type": "object",
"properties": {
"data": {
"$ref": "#/definitions/option.databaseCleanupResult"
}
}
}
]
}
},
"400": {
"description": "参数错误",
"schema": {
"$ref": "#/definitions/response.Any"
}
},
"401": {
"description": "未登录",
"schema": {
"$ref": "#/definitions/response.Any"
}
},
"404": {
"description": "无权限或不存在",
"schema": {
"$ref": "#/definitions/response.Any"
}
},
"500": {
"description": "内部错误",
"schema": {
"$ref": "#/definitions/response.Any"
}
}
}
}
},
"/api/v1/d/option/geoip/lookup": {
"post": {
"security": [
@@ -14906,33 +14838,26 @@ const docTemplate = `{
}
}
},
"analytics.ClickHouseOperationalStats": {
"analytics.BatchWriterStats": {
"type": "object",
"properties": {
"active_parts": {
"cap": {
"type": "integer"
},
"async_insert_bytes": {
"depth": {
"type": "integer"
},
"async_insert_queue": {
"drops": {
"type": "integer"
},
"batch_writers": {
"description": "BatchWriters reports in-process queue depth/drops/flush errors for CH writers.",
"type": "array",
"items": {
"$ref": "#/definitions/batchwriter.Stats"
}
"flush_errors": {
"type": "integer"
},
"database": {
"name": {
"type": "string"
},
"pending_mutations": {
"type": "integer"
},
"total_rows": {
"type": "integer"
"running": {
"type": "boolean"
}
}
},
@@ -15024,29 +14949,6 @@ const docTemplate = `{
}
}
},
"batchwriter.Stats": {
"type": "object",
"properties": {
"cap": {
"type": "integer"
},
"depth": {
"type": "integer"
},
"drops": {
"type": "integer"
},
"flush_errors": {
"type": "integer"
},
"name": {
"type": "string"
},
"running": {
"type": "boolean"
}
}
},
"cache.updateCacheConfigRequest": {
"type": "object",
"required": [
@@ -15105,6 +15007,9 @@ const docTemplate = `{
"id": {
"type": "integer"
},
"zone_domain": {
"type": "string"
},
"zone_id": {
"type": "integer"
}
@@ -15826,6 +15731,36 @@ const docTemplate = `{
}
}
},
"github_com_Rain-kl_Wavelet_internal_model_analytics.ClickHouseOperationalStats": {
"type": "object",
"properties": {
"active_parts": {
"type": "integer"
},
"async_insert_bytes": {
"type": "integer"
},
"async_insert_queue": {
"type": "integer"
},
"batch_writers": {
"description": "BatchWriters reports in-process queue depth/drops/flush errors for CH writers.",
"type": "array",
"items": {
"$ref": "#/definitions/analytics.BatchWriterStats"
}
},
"database": {
"type": "string"
},
"pending_mutations": {
"type": "integer"
},
"total_rows": {
"type": "integer"
}
}
},
"github_com_Rain-kl_Wavelet_pkg_protocol.ActiveConfigMeta": {
"type": "object",
"properties": {
@@ -18481,46 +18416,6 @@ const docTemplate = `{
}
}
},
"option.databaseCleanupInput": {
"type": "object",
"properties": {
"retention_days": {
"type": "integer"
},
"target": {
"type": "string"
}
}
},
"option.databaseCleanupResult": {
"type": "object",
"properties": {
"cleanup_mode": {
"type": "string"
},
"delete_all": {
"type": "boolean"
},
"deleted_count": {
"type": "integer"
},
"eligible_count": {
"type": "integer"
},
"retention_days": {
"type": "integer"
},
"table_ttl_days": {
"type": "integer"
},
"target": {
"type": "string"
},
"target_label": {
"type": "string"
}
}
},
"option.geoIPLookupRequest": {
"type": "object",
"properties": {
@@ -19859,6 +19754,33 @@ const docTemplate = `{
}
}
},
"status.LogDatabaseStatus": {
"type": "object",
"properties": {
"active_database": {
"type": "string"
},
"available_targets": {
"type": "array",
"items": {
"type": "string"
}
},
"clickhouse": {
"$ref": "#/definitions/github_com_Rain-kl_Wavelet_internal_model_analytics.ClickHouseOperationalStats"
},
"migration": {
"description": "idle | migrating",
"type": "string"
},
"retention_days": {
"type": "object",
"additionalProperties": {
"type": "integer"
}
}
}
},
"status.SystemStatusResponse": {
"type": "object",
"properties": {
+13 -42
View File
@@ -14,12 +14,12 @@ Agent 统一通过 OpenResty 二进制控制运行时。本地部署需要节点
## 环境要求
| 项目 | 要求 |
| --- | --- |
| Docker / Docker Compose | 用于启动 Server 及其依赖的 PostgreSQL、Redis 和 ClickHouse 容器;如采用 Docker Agent,也用于运行 Agent |
| OpenResty | 本地安装 Agent 时需要可执行 `openresty`,或在安装脚本中指定路径 |
| 可访问端口 | Server 默认监听 `3000`,Agent 节点需要能访问 Server 地址 |
| 浏览器 | 用于访问管理端 |
| 项目 | 要求 |
| --- |------------------------------------------------------------------|
| Docker / Docker Compose | 用于启动 Server 及其依赖的 PostgreSQL、Valkey;如采用 Docker Agent,也用于运行 Agent |
| OpenResty | 本地安装 Agent 时需要可执行 `openresty`,或在安装脚本中指定路径 |
| 可访问端口 | Server 默认监听 `3000`,Agent 节点需要能访问 Server 地址 |
| 浏览器 | 用于访问管理端 |
- **Docker**:`20.10.0+`
- **Docker Compose**:`2.0.0+`
@@ -28,15 +28,7 @@ Agent 统一通过 OpenResty 二进制控制运行时。本地部署需要节点
## 1. 启动 Server
为了保证异步任务队列(Asynq 框架)及可观测流量看板功能完整运行,快速开始推荐采用 **PostgreSQL + Redis + ClickHouse** 经典单机版编排。
先拉取 ClickHouse 服务端性能配置到 `./config/clickhouse`,并以单文件方式挂载:
```bash
mkdir -p ./config/clickhouse
curl -fsSL -o ./config/clickhouse/performance.xml \
https://raw.githubusercontent.com/Rain-kl/OpenFlare/refs/heads/main/config/clickhouse/performance.xml
```
快速开始推荐采用 **PostgreSQL + Redis ** 标准部署方案。
在空目录中创建 `docker-compose.yaml`:
@@ -63,15 +55,11 @@ services:
DB_NAME: "${DB_NAME:-openflare}"
REDIS_ENABLED: "true"
REDIS_ADDR: "redis:6379"
CLICKHOUSE_ENABLED: "true"
CLICKHOUSE_HOST: "clickhouse:9000"
depends_on:
postgres:
condition: service_healthy
redis:
condition: service_healthy
clickhouse:
condition: service_healthy
postgres:
image: postgres:17-alpine
@@ -100,34 +88,11 @@ services:
timeout: 5s
retries: 5
clickhouse:
image: clickhouse/clickhouse-server:25.3-alpine
restart: unless-stopped
environment:
CLICKHOUSE_DB: openflare
CLICKHOUSE_USER: default
CLICKHOUSE_PASSWORD: ${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1
TZ: Asia/Shanghai
ulimits:
nofile:
soft: 262144
hard: 262144
volumes:
- openflare_clickhouse_data:/var/lib/clickhouse
- ./config/clickhouse/performance.xml:/etc/clickhouse-server/config.d/performance.xml:ro
healthcheck:
test: ["CMD", "clickhouse-client", "--user", "default", "--password", "${CLICKHOUSE_PASSWORD:-replace-with-clickhouse-password}", "--query", "SELECT 1"]
interval: 10s
timeout: 5s
retries: 5
start_period: 15s
volumes:
openflare_uploads:
openflare_postgres_data:
openflare_redis_data:
openflare_clickhouse_data:
```
启动服务:
@@ -158,6 +123,12 @@ http://localhost:3000
> [!WARNING]
> 为了你的系统安全,首次登录后请立即修改默认密码。
如果忘记密码并且没有配置找回密码渠道, 可以使用命令进行重置
```bash
go run main.go reset-paswd # 重置管理员密码
```
---
## 2. 准备 Agent Token
+16 -3
View File
@@ -105,7 +105,7 @@ Server 的所有核心基础配置定义在 `config.yaml` 中,且均支持环
| 配置文件 YAML 路径 | 对应覆盖环境变量 | 作用说明 | 默认值 |
| --- | --- | --- | --- |
| `clickhouse.enabled` | `CLICKHOUSE_ENABLED` | 是否启用 ClickHouse。**系统节点指标与访问日志在此进行海量写入** | `true` |
| `clickhouse.enabled` | `CLICKHOUSE_ENABLED` | 是否启用 ClickHouse。**系统节点指标与访问日志在此进行海量写入**。默认关闭:缺失本配置项或为 `false` 时不启用,日志/指标由主库承担;显式 `true` 或设置 `CLICKHOUSE_HOST` 时启用 | `false` |
| `clickhouse.hosts` | `CLICKHOUSE_HOST` | ClickHouse 集群连接地址数组(环境变量仅设置单地址) | `["127.0.0.1:9000"]` |
| `clickhouse.username` | `CLICKHOUSE_USERNAME` | ClickHouse 账号用户名 | `default` |
| `clickhouse.password` | `CLICKHOUSE_PASSWORD` | ClickHouse 密码 | `replace-with-clickhouse-password` |
@@ -197,8 +197,6 @@ Server 的所有核心基础配置定义在 `config.yaml` 中,且均支持环
| `node_offline_threshold` | `int` | 在管理后台中判定节点失去心跳并标注为离线状态的无响应阈值(毫秒) | `60000` (60s) |
| `agent_update_repo` | `string` | Agent 节点更新下载自身二进制的 Release 仓库源 | `Rain-kl/OpenFlare` |
| `geoip_provider` | `string` | GeoIP 提供商,支持 `maxmind` 等,用于 WAF 防护时地域分析 | `ipinfo` |
| `database_auto_cleanup_enabled` | `bool` | 是否在每天凌晨 3:00 自动清理过期观测历史日志(降低数据库空间) | `true` |
| `database_auto_cleanup_retention_days` | `int` | 自动清理观测数据(访问日志、度量曲线、审计等)的默认保留天数 | `30` |
### 5. Uptime Kuma 监控联动同步
| 配置键 (Key) | 数据类型 | 作用说明 | 默认值 |
@@ -272,6 +270,21 @@ Server 的所有核心基础配置定义在 `config.yaml` 中,且均支持环
---
### 8. 日志存储(Log Database)
日志存储解耦后的运行时配置:日志主库由「切换日志数据库」任务管理(内部/受保护 key,禁止管理员手动修改),访问日志保留天数按存储库分别在业务配置中设置;性能指标(CPU/内存/磁盘/网络)价值衰减快,按三库共用的独立短留存清理。
| 配置键 (Key) | 数据类型 | 作用说明 | 默认值 |
| --- | --- | --- | --- |
| `log_database` | `string` | 当前日志主库(`postgres` / `sqlite` / `clickhouse`)。**内部受保护 key**:仅「切换日志数据库」迁移任务写入,管理员不可手动创建/修改 | 随主库(PostgreSQL 启用时为 `postgres`,否则 `sqlite`;ClickHouse 启用时优先 `clickhouse`) |
| `log_db_migration` | `string` | 日志迁移冻结标记(`migrating` 或空)。**内部受保护 key**:仅迁移任务写入,置位期间日志写入返回 503「日志数据库迁移中,暂不可写」 | 空 |
| `log_retention_days_postgres` | `int` | PostgreSQL 日志库的访问日志过期清理保留天数(过期日志由系统垃圾清理每日任务删除) | `30` |
| `log_retention_days_sqlite` | `int` | SQLite 日志库的访问日志过期清理保留天数 | `30` |
| `log_retention_days_clickhouse` | `int` | ClickHouse 日志库的访问日志过期清理保留天数 | `30` |
| `metric_retention_days` | `int` | 性能指标(CPU/内存/磁盘/网络)保留天数,三库共用独立短留存(不随访问日志保留配置) | `3` |
---
## 前端构建环境变量
| 环境变量 | 作用 | 默认值 |
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,177 @@
# 日志数据库解耦设计(ClickHouse 可选化)
> 状态:已与用户逐段确认,待用户复核。
> 日期:2026-08-08
## 1. 背景与目标
当前系统日志/分析(访问日志、可观测时序)完全绑定 ClickHouse:`internal/repository/analytics` 直接操作 `db.ChConn`/`db.ChDB`,apps 层(`chwriter`、`risk_control`、`admin/logs`、`admin/status`)依赖 `config.ClickHouse.Enabled` 判断可用性。业务流量小、主机性能低时 ClickHouse 负担大。
目标:
1. **解耦**:ClickHouse 变为可选项;不启用时,主库(PostgreSQL;禁用时 SQLite)完整承接全部日志功能(写入、查询、聚合、清理)。
2. **代码级约束**:上层应用写日志不能直接调用底层库(`analyticsrepo` / `db.ChConn`),用接口 + import-lint 测试保证,而非 AGENTS.md 口头约束。
3. **可迁移**:提供用户触发的「切换日志数据库」任务,支持 PostgreSQL/SQLite ↔ ClickHouse 数据迁移。
4. **表结构**:CH 日志表迁入 PG/SQLite;CH 保持只有日志表的 SQL 脚本;PG/SQLite 包含全部表。
## 2. 现状要点
- 连接:`internal/infra/persistence/clickhouse.go`(`ChConn` 原生批量写 + `ChDB` GORM 查询),`init()` 依据 `clickhouse.enabled`。
- 分析域:`internal/repository/analytics/` 直接读写 CH;apps 通过 `batchwriter` 异步 flush(`chwriter`、`risk_control`)。
- 已有抽象雏形:`internal/repository/openflare_access_log_store.go` / `openflare_observability_store.go` 中的未导出 `accessLogStore` / `observabilityStore` 接口,默认 `clickhouseAccessLogStore{}`,测试可换 memory 实现——默认写死 CH、不可配置切换、接口未导出。
- 迁移:主库 goose(`goose/postgres` + `goose/sqlite` 双方言)与 CH 单方言(`goose/clickhouse`)分离。
- 历史:PG/SQLite 曾有过 `of_node_metric_snapshots`、`of_node_access_logs` 等观测表(`202606190010_create_of_observability_tables.sql`),后由 `202606200005_drop_of_node_observability_timeseries.sql` 删除(迁去 CH)。**旧 DDL 可复活改造**。
- 任务:Asynq + `task.RegisterHandler`/`RegisterTaskMeta`;`system_cleanup`(系统垃圾清理)每日任务已存在;`of_database_auto_cleanup`(可观测清理,schedule id=102)存在。
- 系统配置:`system_configs` 表(key/type/visibility),现有 `database_auto_cleanup_enabled` / `database_auto_cleanup_retention_days`(business)。
## 3. 已确认的核心决策
| # | 决策 |
|---|---|
| 1 | 范围:CH 不启用时,PG(或 SQLite)承担**全部**日志功能;聚合在 PG/SQLite 查询时实时计算,不物理建 MV 同构表。 |
| 2 | 实现:接口定义在 repository 层;PG 用 GORM 全新实现;CH 保留现有原生批量优化(`PrepareBatch`)包进同一接口。 |
| 3 | SQLite 是一等公民:`log_database` ∈ {`postgres`, `sqlite`, `clickhouse`};迁移方向 PG→CH、SQLite→CH、CH→PG、CH→SQLite。 |
| 4 | 日志库只有两种合法状态:**随主库**(`database.enabled` → postgres,否则 sqlite)或 **clickhouse**;不存在主库 PG + 日志 SQLite 的组合。 |
| 5 | 迁移任务「切换日志数据库」:纯复制、**源数据不删除**、可重试;迁移期间**冻结日志写入**(拒绝,不排队积压);全部成功才翻转主库标记。 |
| 6 | 清理统一到 `system_cleanup`(每日一次,日志过期无需实时);保留时间按**存储库**配置(`type=business`)。 |
## 4. 包结构与接口(方案一)
新增 `internal/repository/logstore/`,职责唯一:日志存储抽象。
```
internal/repository/logstore/
├── logstore.go # 导出接口:AccessLogStore / ObservabilityStore / UserAccessLogStore / CleanupStore / StatusStore
├── provider.go # Open(ctx) 按当前日志主库返回实现;ActiveDatabase() 供状态/UI;测试可注入
├── postgres_store.go # GORM 实现(PG 与 SQLite 共用一套,方言差异只在 goose DDL + dialect_* 小文件)
├── dialect_postgres.go # PG 方言 SQL 片段(date_trunc / FILTER / 分区清理)
├── dialect_sqlite.go # SQLite 方言 SQL 片段(strftime / unixepoch)
└── clickhouse_store.go # 把现有 analyticsrepo 原生批量 + GORM 查询包进接口(零性能损耗)
```
- **接口划分**(避免 40+ 方法巨型接口,合成 `logstore.Store` 结构体持有):
- `AccessLogStore`:节点访问日志的 InsertBatch / List / Count / RegionCounts / BucketAggregates / CountBuckets / BucketDimensions / IPAggregates / IPSummaries / CountIPSummaries / WAFIPAggregates / IPTrend / TrafficSummary / ValueCounts / NodeAggregates / DeleteAll / DeleteBefore / DeleteByNodeBefore。
- `ObservabilityStore`:4 表(metric snapshots / edge health / frps / frpc)的 Insert / List / Delete。
- `UserAccessLogStore`:`w_user_access_logs` 的 BatchInsert / Count / List / 统计(DailyTrend / BrowserDistribution / TopActiveUsers 等)。
- `CleanupStore`:按保留天数清理过期数据(PG=分区 DROP + 分批 DELETE;SQLite=分批 DELETE;CH=MODIFY TTL + materialize)。
- `StatusStore`:当前库状态、CH 运行指标(激活时)、GORM 写入器状态。
- **消费面**:`internal/repository` 现有公开函数(`ListOpenFlareAccessLogs`、`InsertOpenFlareAccessLogsBatch`、`InsertOpenFlareMetricSnapshot` 等)**保留签名、改为一行委托 `logstore`**,apps 调用面几乎不动;apps 里现有 `analyticsrepo` 直连(`risk_control`、`chwriter`、`tasks/database_cleanup.go`、`observability/access_log_logics.go`、`admin/logs`、`admin/status`)全部改走 repository/logstore。
- **import-lint 测试**:新增 `go test`,扫描 `internal/apps/**` 的 import,发现 `internal/repository/analytics` 或 `internal/infra/persistence`(`batchwriter` 白名单除外)即失败。这是「代码层面规避」的验收。
- `analyticsrepo` 保留,仅被 `logstore/clickhouse_store.go` 引用(CH 实现细节)。
### 主库标记与启动校验
- `system_configs` 新增内部 key:
- `log_database`(`postgres`/`sqlite`/`clickhouse`):当前日志主库,仅迁移任务写入。
- `log_db_migration`(`"migrating"`/空):迁移冻结标记,仅迁移任务写入。
- **首次 seed**(bootstrap Go 侧,因依赖运行时主库选择):key 缺失时,`clickhouse.enabled` → `clickhouse`(保持现状、不丢现有 CH 数据);否则 → 当前主库(`database.enabled` → `postgres`,否则 `sqlite`)。
- **启动校验**(bootstrap):
- `log_database=clickhouse` 但 `clickhouse.enabled=false` → 启动报错:「当前日志主库为 ClickHouse 但 ClickHouse 未启用。请先重新启用 ClickHouse 配置并启动,在任务管理运行『切换日志数据库』迁移到 PostgreSQL/SQLite 后再禁用 ClickHouse」。
- `log_database=postgres` 但 `database.enabled=false`,或 `log_database=sqlite` 但 `database.enabled=true` → 启动报错(违反「随主库或随 CH」规则)。
- **key 保护**:`log_database`、`log_db_migration` 在配置更新接口(admin system-configs / option 校验)拒绝修改;仅迁移任务可写;启动校验兜底被篡改组合。
- **热切换**:`logstore` 通过系统配置缓存(Redis,更新即失效)读取 `log_database`;翻转后 API 进程自动切到新实现,无需自定义跨进程协议。
## 5. PG/SQLite 表结构与优化
**新建原始日志表(PG + SQLite 双方言 goose,同版本号)**——只建原始表,**不建** CH 物化视图/聚合表(`of_access_log_hourly`、`of_node_metric_capacity_hourly` 等),PG/SQLite 查询时实时聚合:
| 表 | 说明 |
|---|---|
| `w_user_access_logs` | 用户访问日志 |
| `of_node_access_logs` | 节点访问日志(含 user_agent/cache_status/bytes_sent/request_length/request_time_ms 现行列) |
| `of_node_metric_snapshots` | 资源指标 |
| `of_node_edge_health` | 边缘健康 |
| `of_node_obs_frps` | FRPS 观测 |
| `of_node_obs_frpc` | FRPC 观测 |
- **ID**:沿用 snowflake uint64(DDL 用 BIGINT,与 CH UInt64 对齐);不换自增,保证迁移 ID 原样保留、无冲突。
- **时间**:PG `TIMESTAMPTZ`;SQLite `DATETIME`。
- **复合主键**:分区表主键 `(id, 时间列)`(满足 PG 分区键进唯一索引要求)。
### PG 优化
1. **分区**:仅 `of_node_access_logs`、`w_user_access_logs` 两个高频表用 PG 原生 `PARTITION BY RANGE` **按月分区**;可观测 4 表数据量小,普通表 + 索引。SQLite 无原生分区 → 普通表 + 组合索引(方言差异只留在 goose DDL,运行时 GORM 代码共用)。
2. **批量写入**:PG/SQLite 统一 GORM `CreateInBatches`(批次 500–1000);CH 维持原生 `PrepareBatch`。
3. **索引**:
- `of_node_access_logs`:`(logged_at DESC)`、`(node_id, logged_at DESC)`、`(host, logged_at DESC)`;
- `w_user_access_logs`:`(created_at DESC)`、`(user_id, created_at DESC)`;
- 可观测表:`(node_id, captured_at DESC)`。
4. **聚合查询重写**:PG 用 `date_trunc` / `count(DISTINCT)` / `FILTER (WHERE ...)` 等价替换 CH 的 `toStartOfHour` / `uniqExact` / `countIf`;SQLite 用 `strftime` / `unixepoch`。时间分桶等少量方言 SQL 拆到 `dialect_postgres.go` / `dialect_sqlite.go`,store 主体方言中立。
### goose 迁移
- PG/SQLite 各新增一组建表迁移(复活并改造 `202606190010` 旧 DDL,按 database-migration 技能双方言、同版本号规则)。
- CH 目录不动(本来就只有日志表脚本,满足「CH 保持只有日志表 SQL」)。
## 6. 清理(并入 system_cleanup)
- 日志过期清理并入 `system_cleanup`(系统垃圾清理)每日任务;`of_database_auto_cleanup` 专用 schedule(id=102)与任务下线。
- 新增 `type=business` 配置(替换旧 `database_auto_cleanup_enabled` / `database_auto_cleanup_retention_days`):
- `log_retention_days_postgres`(默认 90)
- `log_retention_days_sqlite`(默认 90)
- `log_retention_days_clickhouse`(默认 90)
- `CleanupStore` 按当前生效库读取对应值执行:
- PG:分区 DROP(整月)+ 分批 DELETE(不满月);
- SQLite:分批 DELETE;
- CH:`ALTER TABLE ... MODIFY TTL toDateTime(...) + INTERVAL N DAY` + materialize(保留期由配置驱动,不再依赖 DDL 写死)。
- 旧 key `database_auto_cleanup_*` 由 goose 迁移删除,前端同步清理。
## 7. 迁移任务「切换日志数据库」
**元数据**:Asynq `openflare:log_db_switch`,管理类型 `of_log_db_switch`,名称「切换日志数据库」,参数 `target`(`postgres`/`sqlite`/`clickhouse`),`Retryable: true`。UI 按当前日志主库只展示合法目标(当前=CH → 「主库」;当前=主库 → 「ClickHouse」)。
**执行流程(worker 进程)**:
1. **校验**:`target == 当前主库` → 拒绝;`target=clickhouse` 但 CH 未启用 / `target=postgres` 但 `database.enabled=false` / `target=sqlite` 但 `database.enabled=true` → 拒绝。
2. **写冻结**:写 `log_db_migration = "migrating"`;先让 batchwriter 把在途批次 flush 完;此后 API 进程所有日志写入路径(`risk_control`、`chwriter` 队列、agent 上报落库)检查该 key → 返回明确错误(HTTP 503「日志数据库迁移中,暂不可写」),不排队积压。
3. **复制**:6 张原始日志表逐表、按 id 分批(每批 ~1000)读源 → 写目标(CH→主库用 GORM `CreateInBatches`;主库→CH 用原生 `PrepareBatch`);ID 原样保留;每表/每批 `task.AppendLog` 进度。
- **幂等前提**:开始复制前**清空目标库日志表**(任务参数「覆盖目标库已有日志」默认开启;目标库通常为空,仅「切回去」场景有旧数据)——保证失败重试可重跑不重复。
4. **翻转**:全部成功 → 更新 `log_database = target`、清除迁移标记 → `logstore` 缓存失效自动切到新实现 → 写入恢复(走新库)。
5. **失败**:返回错误触发 Asynq 重试;**失败时清除迁移标记**,写入继续走源库(不丢功能);重试时重新清空目标 + 复制。
**双进程一致性**:迁移标记与主库标记落在 `system_configs`(Redis 缓存,worker 更新后 API 进程自动失效重读)。
## 8. API 与前端
**后端**:
- `GET /api/v1/admin/status/log-database`(改造现有 `/clickhouse` 状态端点):返回当前日志主库、迁移状态(`idle`/`migrating`)、各库保留天数、当前合法迁移目标;CH 为主时附带现有 CH 运行指标,主库为主时附带 GORM 写入器状态。
- 任务「切换日志数据库」走现有任务管理通用派发 API(`RegisterTaskMeta` + Params),无需新派发接口;执行记录/进度复用任务框架。
- 系统配置:新增 3 个 `log_retention_days_*`(business)图形化 + 参数表可见;新增内部 `log_database`、`log_db_migration`(system、visibility=0、受保护);下线 `database_auto_cleanup_*`。
**前端**:
- 任务管理页:出现「切换日志数据库」,参数下拉只显示合法目标;页面展示当前日志主库与迁移状态。
- `/admin/settings` 业务配置:新增「日志保留时间」分组(PG/SQLite/CH 三个数字输入)。
- 状态/仪表盘:日志库状态卡片(当前库 + 迁移中提示)。
## 9. 测试与验证
- **import-lint 测试**:`internal/apps/**` 不得 import `internal/repository/analytics`、`internal/infra/persistence`(`batchwriter` 白名单除外),违规即失败。
- **logstore 单测**:GORM 实现用 SQLite 全量跑;PG 专属(分区 DROP 等)走既有集成测试路径;CH 实现复用现有 analyticsrepo 测试。
- **迁移任务测试**:目标/组合校验、批处理与 ID 保留、清空目标、翻转标记、失败清标记回退、冻结期写入拒绝——用 memory/sqlite 双端模拟,不依赖真实 CH。
- **清理测试**:`system_cleanup` 日志清理步骤(PG 分区 DROP / SQLite 分批 DELETE / CH TTL 修改)与保留配置读取。
- **迁移验证**:goose 空库 Up 全量(PG/SQLite/CH 三套)、`go test ./...`、`make swagger`(API 变更)、`make code-check`、`make format`。
## 10. 非目标(YAGNI)
- 不在 PG/SQLite 物理建聚合/物化视图表(查询实时聚合)。
- 不做 PG ↔ SQLite 日志互迁(非法组合,启动校验拒绝)。
- 迁移成功不自动删除源库数据(保留,后续提供手动清理入口)。
- 不引入 PG COPY 协议(GORM `CreateInBatches` 对低流量足够)。
- 不引入自定义跨进程迁移协议(`system_configs` + Redis 缓存即可)。
## 11. 里程碑建议(供实现计划分解)
1. **M1 抽象与改造**:`logstore` 接口 + PG/SQLite 实现 + `clickhouse_store` 包装 + import-lint 测试 + repository 委托改造 + apps 直连改造 + `log_database`/`log_db_migration` key 与启动校验。
2. **M2 表与清理**:goose 双方言建表迁移 + 保留配置 key + `system_cleanup` 日志清理步骤 + 下线 `of_database_auto_cleanup` 与旧配置。
3. **M3 迁移任务与展示**:迁移任务 Handler + 状态端点 + 任务管理页/业务配置前端 + 日志库状态卡片。
4. **M4 收尾**:全量验证(goose 三套、单测、`make code-check`/`swagger`/`format`)、文档同步(中文)、changelog `[Unreleased]`。
## 12. 实现归档说明(Task 18,2026-08-08)
- 设计稿第 4 节 provider 入口写作 `Open(ctx)`,实现命名为 `Active(ctx)`(按 `log_database` 解析并缓存,配置翻转后重建),另导出 `Build(ctx, database)` / `BuildForMigration(ctx, database)` 供迁移任务构造目标库 store;`ActiveDatabase(ctx)` 供状态端点。
- 设计稿第 4 节列出的 `CleanupStore` 接口未单独落地:清理实现为包级 `CleanupExpired(ctx)`(按当前激活库保留天数删除过期日志并预建 PG 分区),由 `system_cleanup` 每日任务调用。
- 设计稿第 4 节列举的 `tasks/database_cleanup.go` 已随 M2 下线(`of_database_auto_cleanup` 配置与前端 UI 一并移除),日志清理职责并入 `system_cleanup`。
- 迁移复制按 id 升序分页,`copyObservability` 以每批最后一条 id 作为下一批游标(修正计划中 `lastID += n` 的近似写法);失败回退由 `defer setMigrationFlag("")` 保证源库恢复可写,重试前先清空目标库保证幂等。
- 其余实现决策(`SetConfigReader` 注入、`ensureWritable` 统一冻结、解析 helper 迁至 `model/analytics` 等)见计划「自检记录」,与本文档一致。
+89 -167
View File
@@ -1017,7 +1017,7 @@
"SessionCookie": []
}
],
"description": "分页并按照用户、接口路径、时间范围等维度检索 ClickHouse 用户访问日志列表(需要管理员权限,ClickHouse 未启用时报错)",
"description": "分页并按照用户、接口路径、时间范围等维度检索用户访问日志列表(需要管理员权限,日志存储未启用时报错)",
"produces": [
"application/json"
],
@@ -1085,7 +1085,7 @@
}
},
"400": {
"description": "ClickHouse 未启用或参数错误",
"description": "日志存储未启用或参数错误",
"schema": {
"$ref": "#/definitions/response.Any"
}
@@ -1112,7 +1112,7 @@
"SessionCookie": []
}
],
"description": "聚合统计最近 7 天的每日访问趋势、浏览器分布以及前 10 名最活跃用户排行(需要管理员权限,ClickHouse 未启用时报错)",
"description": "聚合统计最近 7 天的每日访问趋势、浏览器分布以及前 10 名最活跃用户排行(需要管理员权限,日志存储未启用时报错)",
"produces": [
"application/json"
],
@@ -1140,7 +1140,7 @@
}
},
"400": {
"description": "ClickHouse 未启用",
"description": "日志存储未启用",
"schema": {
"$ref": "#/definitions/response.Any"
}
@@ -1870,21 +1870,21 @@
}
}
},
"/api/v1/admin/status/clickhouse": {
"/api/v1/admin/status/log-database": {
"get": {
"security": [
{
"SessionCookie": []
}
],
"description": "返回 ClickHouse parts、mutation、async_insert 队列及进程内 batch writer 指标,需要管理员权限",
"description": "返回当前日志主库、迁移状态、各库保留天数与合法迁移目标,需要管理员权限",
"produces": [
"application/json"
],
"tags": [
"admin"
],
"summary": "获取 ClickHouse 运行指标",
"summary": "获取日志数据库状态",
"responses": {
"200": {
"description": "获取成功",
@@ -1897,19 +1897,13 @@
"type": "object",
"properties": {
"data": {
"$ref": "#/definitions/analytics.ClickHouseOperationalStats"
"$ref": "#/definitions/status.LogDatabaseStatus"
}
}
}
]
}
},
"400": {
"description": "ClickHouse 未启用",
"schema": {
"$ref": "#/definitions/response.Any"
}
},
"401": {
"description": "未登录",
"schema": {
@@ -2386,6 +2380,18 @@
"name": "task_type",
"in": "query"
},
{
"type": "string",
"description": "任务类型前缀筛选(与 task_type / task_types 互斥,精确类型优先)",
"name": "task_type_prefix",
"in": "query"
},
{
"type": "string",
"description": "逗号分隔的精确任务类型列表(IN 筛选,优先于前缀)",
"name": "task_types",
"in": "query"
},
{
"type": "integer",
"default": 1,
@@ -8298,80 +8304,6 @@
}
}
},
"/api/v1/d/option/database/cleanup": {
"post": {
"security": [
{
"SessionCookie": []
}
],
"description": "按目标与保留天数清理可观测性相关数据表,需要管理员权限",
"consumes": [
"application/json"
],
"produces": [
"application/json"
],
"tags": [
"openflare-option"
],
"summary": "清理可观测性数据库",
"parameters": [
{
"description": "清理参数",
"name": "request",
"in": "body",
"schema": {
"$ref": "#/definitions/option.databaseCleanupInput"
}
}
],
"responses": {
"200": {
"description": "清理结果",
"schema": {
"allOf": [
{
"$ref": "#/definitions/response.Any"
},
{
"type": "object",
"properties": {
"data": {
"$ref": "#/definitions/option.databaseCleanupResult"
}
}
}
]
}
},
"400": {
"description": "参数错误",
"schema": {
"$ref": "#/definitions/response.Any"
}
},
"401": {
"description": "未登录",
"schema": {
"$ref": "#/definitions/response.Any"
}
},
"404": {
"description": "无权限或不存在",
"schema": {
"$ref": "#/definitions/response.Any"
}
},
"500": {
"description": "内部错误",
"schema": {
"$ref": "#/definitions/response.Any"
}
}
}
}
},
"/api/v1/d/option/geoip/lookup": {
"post": {
"security": [
@@ -14899,33 +14831,26 @@
}
}
},
"analytics.ClickHouseOperationalStats": {
"analytics.BatchWriterStats": {
"type": "object",
"properties": {
"active_parts": {
"cap": {
"type": "integer"
},
"async_insert_bytes": {
"depth": {
"type": "integer"
},
"async_insert_queue": {
"drops": {
"type": "integer"
},
"batch_writers": {
"description": "BatchWriters reports in-process queue depth/drops/flush errors for CH writers.",
"type": "array",
"items": {
"$ref": "#/definitions/batchwriter.Stats"
}
"flush_errors": {
"type": "integer"
},
"database": {
"name": {
"type": "string"
},
"pending_mutations": {
"type": "integer"
},
"total_rows": {
"type": "integer"
"running": {
"type": "boolean"
}
}
},
@@ -15017,29 +14942,6 @@
}
}
},
"batchwriter.Stats": {
"type": "object",
"properties": {
"cap": {
"type": "integer"
},
"depth": {
"type": "integer"
},
"drops": {
"type": "integer"
},
"flush_errors": {
"type": "integer"
},
"name": {
"type": "string"
},
"running": {
"type": "boolean"
}
}
},
"cache.updateCacheConfigRequest": {
"type": "object",
"required": [
@@ -15098,6 +15000,9 @@
"id": {
"type": "integer"
},
"zone_domain": {
"type": "string"
},
"zone_id": {
"type": "integer"
}
@@ -15819,6 +15724,36 @@
}
}
},
"github_com_Rain-kl_Wavelet_internal_model_analytics.ClickHouseOperationalStats": {
"type": "object",
"properties": {
"active_parts": {
"type": "integer"
},
"async_insert_bytes": {
"type": "integer"
},
"async_insert_queue": {
"type": "integer"
},
"batch_writers": {
"description": "BatchWriters reports in-process queue depth/drops/flush errors for CH writers.",
"type": "array",
"items": {
"$ref": "#/definitions/analytics.BatchWriterStats"
}
},
"database": {
"type": "string"
},
"pending_mutations": {
"type": "integer"
},
"total_rows": {
"type": "integer"
}
}
},
"github_com_Rain-kl_Wavelet_pkg_protocol.ActiveConfigMeta": {
"type": "object",
"properties": {
@@ -18474,46 +18409,6 @@
}
}
},
"option.databaseCleanupInput": {
"type": "object",
"properties": {
"retention_days": {
"type": "integer"
},
"target": {
"type": "string"
}
}
},
"option.databaseCleanupResult": {
"type": "object",
"properties": {
"cleanup_mode": {
"type": "string"
},
"delete_all": {
"type": "boolean"
},
"deleted_count": {
"type": "integer"
},
"eligible_count": {
"type": "integer"
},
"retention_days": {
"type": "integer"
},
"table_ttl_days": {
"type": "integer"
},
"target": {
"type": "string"
},
"target_label": {
"type": "string"
}
}
},
"option.geoIPLookupRequest": {
"type": "object",
"properties": {
@@ -19852,6 +19747,33 @@
}
}
},
"status.LogDatabaseStatus": {
"type": "object",
"properties": {
"active_database": {
"type": "string"
},
"available_targets": {
"type": "array",
"items": {
"type": "string"
}
},
"clickhouse": {
"$ref": "#/definitions/github_com_Rain-kl_Wavelet_internal_model_analytics.ClickHouseOperationalStats"
},
"migration": {
"description": "idle | migrating",
"type": "string"
},
"retention_days": {
"type": "object",
"additionalProperties": {
"type": "integer"
}
}
}
},
"status.SystemStatusResponse": {
"type": "object",
"properties": {
+66 -112
View File
@@ -169,26 +169,20 @@ definitions:
$ref: '#/definitions/github_com_Rain-kl_Wavelet_pkg_protocol.WAFIPGroup'
type: array
type: object
analytics.ClickHouseOperationalStats:
analytics.BatchWriterStats:
properties:
active_parts:
cap:
type: integer
async_insert_bytes:
depth:
type: integer
async_insert_queue:
drops:
type: integer
batch_writers:
description: BatchWriters reports in-process queue depth/drops/flush errors
for CH writers.
items:
$ref: '#/definitions/batchwriter.Stats'
type: array
database:
flush_errors:
type: integer
name:
type: string
pending_mutations:
type: integer
total_rows:
type: integer
running:
type: boolean
type: object
apply_log.CleanupInput:
properties:
@@ -247,21 +241,6 @@ definitions:
is_active:
type: boolean
type: object
batchwriter.Stats:
properties:
cap:
type: integer
depth:
type: integer
drops:
type: integer
flush_errors:
type: integer
name:
type: string
running:
type: boolean
type: object
cache.updateCacheConfigRequest:
properties:
lru_enabled:
@@ -301,6 +280,8 @@ definitions:
type: string
id:
type: integer
zone_domain:
type: string
zone_id:
type: integer
type: object
@@ -774,6 +755,27 @@ definitions:
token:
type: string
type: object
github_com_Rain-kl_Wavelet_internal_model_analytics.ClickHouseOperationalStats:
properties:
active_parts:
type: integer
async_insert_bytes:
type: integer
async_insert_queue:
type: integer
batch_writers:
description: BatchWriters reports in-process queue depth/drops/flush errors
for CH writers.
items:
$ref: '#/definitions/analytics.BatchWriterStats'
type: array
database:
type: string
pending_mutations:
type: integer
total_rows:
type: integer
type: object
github_com_Rain-kl_Wavelet_pkg_protocol.ActiveConfigMeta:
properties:
checksum:
@@ -2533,32 +2535,6 @@ definitions:
window_started_at:
type: string
type: object
option.databaseCleanupInput:
properties:
retention_days:
type: integer
target:
type: string
type: object
option.databaseCleanupResult:
properties:
cleanup_mode:
type: string
delete_all:
type: boolean
deleted_count:
type: integer
eligible_count:
type: integer
retention_days:
type: integer
table_ttl_days:
type: integer
target:
type: string
target_label:
type: string
type: object
option.geoIPLookupRequest:
properties:
ip:
@@ -3442,6 +3418,24 @@ definitions:
version:
type: string
type: object
status.LogDatabaseStatus:
properties:
active_database:
type: string
available_targets:
items:
type: string
type: array
clickhouse:
$ref: '#/definitions/github_com_Rain-kl_Wavelet_internal_model_analytics.ClickHouseOperationalStats'
migration:
description: idle | migrating
type: string
retention_days:
additionalProperties:
type: integer
type: object
type: object
status.SystemStatusResponse:
properties:
alloc:
@@ -4935,7 +4929,7 @@ paths:
- admin
/api/v1/admin/logs/access:
get:
description: 分页并按照用户、接口路径、时间范围等维度检索 ClickHouse 用户访问日志列表(需要管理员权限,ClickHouse 未启用时报错)
description: 分页并按照用户、接口路径、时间范围等维度检索用户访问日志列表(需要管理员权限,日志存储未启用时报错)
parameters:
- default: 1
description: 页码
@@ -4976,7 +4970,7 @@ paths:
$ref: '#/definitions/logs.accessLogsResponse'
type: object
"400":
description: ClickHouse 未启用或参数错误
description: 日志存储未启用或参数错误
schema:
$ref: '#/definitions/response.Any'
"401":
@@ -4994,7 +4988,7 @@ paths:
- admin
/api/v1/admin/logs/analytics:
get:
description: 聚合统计最近 7 天的每日访问趋势、浏览器分布以及前 10 名最活跃用户排行(需要管理员权限,ClickHouse 未启用时报错)
description: 聚合统计最近 7 天的每日访问趋势、浏览器分布以及前 10 名最活跃用户排行(需要管理员权限,日志存储未启用时报错)
produces:
- application/json
responses:
@@ -5008,7 +5002,7 @@ paths:
$ref: '#/definitions/logs.logsAnalyticsResponse'
type: object
"400":
description: ClickHouse 未启用
description: 日志存储未启用
schema:
$ref: '#/definitions/response.Any'
"401":
@@ -5434,9 +5428,9 @@ paths:
summary: 获取系统状态信息
tags:
- admin
/api/v1/admin/status/clickhouse:
/api/v1/admin/status/log-database:
get:
description: 返回 ClickHouse parts、mutation、async_insert 队列及进程内 batch writer 指标,需要管理员权限
description: 返回当前日志主库、迁移状态、各库保留天数与合法迁移目标,需要管理员权限
produces:
- application/json
responses:
@@ -5447,12 +5441,8 @@ paths:
- $ref: '#/definitions/response.Any'
- properties:
data:
$ref: '#/definitions/analytics.ClickHouseOperationalStats'
$ref: '#/definitions/status.LogDatabaseStatus'
type: object
"400":
description: ClickHouse 未启用
schema:
$ref: '#/definitions/response.Any'
"401":
description: 未登录
schema:
@@ -5467,7 +5457,7 @@ paths:
$ref: '#/definitions/response.Any'
security:
- SessionCookie: []
summary: 获取 ClickHouse 运行指标
summary: 获取日志数据库状态
tags:
- admin
/api/v1/admin/system-configs:
@@ -5738,6 +5728,14 @@ paths:
in: query
name: task_type
type: string
- description: 任务类型前缀筛选(与 task_type / task_types 互斥,精确类型优先)
in: query
name: task_type_prefix
type: string
- description: 逗号分隔的精确任务类型列表(IN 筛选,优先于前缀)
in: query
name: task_types
type: string
- default: 1
description: 页码
in: query
@@ -9285,50 +9283,6 @@ paths:
summary: 列出 OpenFlare 配置项
tags:
- openflare-option
/api/v1/d/option/database/cleanup:
post:
consumes:
- application/json
description: 按目标与保留天数清理可观测性相关数据表,需要管理员权限
parameters:
- description: 清理参数
in: body
name: request
schema:
$ref: '#/definitions/option.databaseCleanupInput'
produces:
- application/json
responses:
"200":
description: 清理结果
schema:
allOf:
- $ref: '#/definitions/response.Any'
- properties:
data:
$ref: '#/definitions/option.databaseCleanupResult'
type: object
"400":
description: 参数错误
schema:
$ref: '#/definitions/response.Any'
"401":
description: 未登录
schema:
$ref: '#/definitions/response.Any'
"404":
description: 无权限或不存在
schema:
$ref: '#/definitions/response.Any'
"500":
description: 内部错误
schema:
$ref: '#/definitions/response.Any'
security:
- SessionCookie: []
summary: 清理可观测性数据库
tags:
- openflare-option
/api/v1/d/option/geoip/lookup:
post:
consumes:
@@ -18,8 +18,6 @@ export type OpenFlareOpsFields = {
uptime_kuma_retry: string;
uptime_kuma_retry_interval: string;
uptime_kuma_timeout: string;
database_auto_cleanup_enabled: boolean;
database_auto_cleanup_retention_days: string;
pages_max_package_size_mb: string;
pages_max_history_count: string;
};
@@ -42,8 +40,6 @@ export const defaultOpenFlareOpsFields: OpenFlareOpsFields = {
uptime_kuma_retry: '0',
uptime_kuma_retry_interval: '60',
uptime_kuma_timeout: '48',
database_auto_cleanup_enabled: false,
database_auto_cleanup_retention_days: '30',
pages_max_package_size_mb: '100',
pages_max_history_count: '20',
};
@@ -88,12 +84,6 @@ export function mapOptionsToOpsFields(
uptime_kuma_retry: optionMap.uptime_kuma_retry ?? '0',
uptime_kuma_retry_interval: optionMap.uptime_kuma_retry_interval ?? '60',
uptime_kuma_timeout: optionMap.uptime_kuma_timeout ?? '48',
database_auto_cleanup_enabled: toBoolean(
optionMap.database_auto_cleanup_enabled,
false,
),
database_auto_cleanup_retention_days:
optionMap.database_auto_cleanup_retention_days ?? '30',
pages_max_package_size_mb: optionMap.pages_max_package_size_mb ?? '100',
pages_max_history_count: optionMap.pages_max_history_count ?? '20',
};
@@ -165,16 +155,6 @@ export function validateUptimeKumaFields(fields: OpenFlareOpsFields) {
throw new Error('请求超时必须为正整数。');
}
export function validateDatabaseAutoCleanup(fields: OpenFlareOpsFields) {
const retentionDays = Number.parseInt(
fields.database_auto_cleanup_retention_days,
10,
);
if (Number.isNaN(retentionDays) || retentionDays < 1) {
throw new Error('自动清理保留天数至少为 1 天。');
}
}
export function agentOptionEntries(fields: OpenFlareOpsFields): OptionItem[] {
validateAgentFields(fields);
return [
@@ -220,22 +200,6 @@ export function uptimeKumaOptionEntries(
];
}
export function databaseAutoCleanupEntries(
fields: OpenFlareOpsFields,
): OptionItem[] {
validateDatabaseAutoCleanup(fields);
return [
{
key: 'database_auto_cleanup_enabled',
value: String(fields.database_auto_cleanup_enabled),
},
{
key: 'database_auto_cleanup_retention_days',
value: fields.database_auto_cleanup_retention_days,
},
];
}
export function validatePagesFields(fields: OpenFlareOpsFields) {
const packageSize = Number.parseInt(fields.pages_max_package_size_mb, 10);
const historyCount = Number.parseInt(fields.pages_max_history_count, 10);
@@ -11,20 +11,8 @@ import {
RotateCw,
Save,
Server,
Trash2,
} from 'lucide-react';
import { toast } from 'sonner';
import {
AlertDialog,
AlertDialogAction,
AlertDialogCancel,
AlertDialogContent,
AlertDialogDescription,
AlertDialogFooter,
AlertDialogHeader,
AlertDialogTitle,
} from '@/components/ui/alert-dialog';
import { Button } from '@/components/ui/button';
import {
Card,
@@ -46,7 +34,6 @@ import { Switch } from '@/components/ui/switch';
import { Textarea } from '@/components/ui/textarea';
import { ErrorInline } from '@/components/layout/error';
import { LoadingStateWithBorder } from '@/components/layout/loading';
import type { DatabaseCleanupTarget } from '@/lib/services/openflare';
import {
NodeService,
OptionService,
@@ -57,7 +44,6 @@ import {
import {
agentOptionEntries,
buildDiscoveryCommand,
databaseAutoCleanupEntries,
defaultOpenFlareOpsFields,
formatDurationLabel,
getBrowserOrigin,
@@ -72,31 +58,6 @@ import { UptimeKumaSiteSelectModal } from './uptimekuma-site-modal';
const optionsQueryKey = ['openflare', 'options'] as const;
const openflarePublicStatusQueryKey = ['openflare', 'public-status'] as const;
const cleanupTargets: Array<{
target: DatabaseCleanupTarget;
label: string;
description: string;
}> = [
{
target: 'node_access_logs',
label: '访问日志',
description:
'清理 node_access_logs,影响访问明细与 IP 汇总;表 TTL 为 90 天。',
},
{
target: 'node_metric_snapshots',
label: '性能快照',
description:
'清理 node_metric_snapshots,影响节点资源趋势;表 TTL 为 30 天。',
},
{
target: 'node_edge_health',
label: 'OpenResty 健康',
description:
'清理 node_edge_health(OpenResty 连接/健康快照);表 TTL 为 30 天。业务流量请清理访问日志。',
},
];
async function copyText(value: string) {
await navigator.clipboard.writeText(value);
}
@@ -109,11 +70,6 @@ export function OpenFlareOpsSettings() {
const [savingSection, setSavingSection] = useState<string | null>(null);
const [geoIPTestIP, setGeoIPTestIP] = useState('8.8.8.8');
const [uptimeKumaModalOpen, setUptimeKumaModalOpen] = useState(false);
const [cleanupTarget, setCleanupTarget] = useState<{
target: DatabaseCleanupTarget;
label: string;
} | null>(null);
const [cleanupRetentionDays, setCleanupRetentionDays] = useState('');
const optionsQuery = useQuery({
queryKey: optionsQueryKey,
@@ -194,25 +150,6 @@ export function OpenFlareOpsSettings() {
toast.error(error instanceof Error ? error.message : '同步失败'),
});
const cleanupMutation = useMutation({
mutationFn: (payload: {
target: DatabaseCleanupTarget;
retention_days?: number;
}) => OptionService.cleanupDatabase(payload),
onSuccess: (result) => {
setCleanupTarget(null);
setCleanupRetentionDays('');
toast.success(
result.delete_all
? `已清空${result.target_label},共删除 ${result.deleted_count} 条。`
: `已清理${result.target_label},共删除 ${result.deleted_count} 条。`,
);
},
onError: (error) => {
toast.error(error instanceof Error ? error.message : '清理失败');
},
});
const discoveryToken = bootstrapQuery.data?.discovery_token ?? '';
const discoveryCommand = useMemo(() => {
if (!fields.server_address || !discoveryToken) return '';
@@ -248,17 +185,6 @@ export function OpenFlareOpsSettings() {
}
};
const saveDatabaseAutoCleanup = () => {
try {
saveMutation.mutate({
section: 'database-auto',
entries: databaseAutoCleanupEntries(fields),
});
} catch (error) {
toast.error(error instanceof Error ? error.message : '参数校验失败');
}
};
const savePagesSettings = () => {
try {
saveMutation.mutate({
@@ -690,86 +616,6 @@ export function OpenFlareOpsSettings() {
</CardContent>
</Card>
<div className='grid gap-6 xl:grid-cols-2'>
<Card className='border-dashed shadow-none'>
<CardHeader className='flex flex-row items-center justify-between gap-4'>
<div>
<CardTitle className='text-base'>数据库自动清理</CardTitle>
<CardDescription>
每天凌晨 3 点物化 ClickHouse 表 TTL;访问日志至少保留 90
天,其它观测数据至少保留 30 天。
</CardDescription>
</div>
<Button
size='sm'
disabled={savingSection === 'database-auto'}
onClick={saveDatabaseAutoCleanup}
>
保存
</Button>
</CardHeader>
<CardContent className='space-y-4'>
<ToggleRow
label='启用每日自动清理'
checked={fields.database_auto_cleanup_enabled}
onChange={(value) =>
updateField('database_auto_cleanup_enabled', value)
}
/>
<div className='space-y-1.5'>
<FieldInput
label='期望保留天数'
value={fields.database_auto_cleanup_retention_days}
type='number'
onChange={(value) =>
updateField('database_auto_cleanup_retention_days', value)
}
/>
<p className='text-xs text-muted-foreground'>
小于表 TTL 时自动按下限执行:访问日志 90 天,其它观测数据 30
天。
</p>
</div>
</CardContent>
</Card>
<Card className='border-dashed shadow-none'>
<CardHeader>
<CardTitle className='text-base'>手动数据清理</CardTitle>
<CardDescription>
按表 TTL 清理;输入天数不能小于对应表
TTL,留空时将删除该类数据的全部历史记录。
</CardDescription>
</CardHeader>
<CardContent className='space-y-3'>
{cleanupTargets.map((item) => (
<div
key={item.target}
className='flex items-start justify-between gap-3 rounded-lg border border-dashed p-3'
>
<div>
<p className='text-sm font-medium'>{item.label}</p>
<p className='mt-1 text-xs text-muted-foreground'>
{item.description}
</p>
</div>
<Button
type='button'
variant='destructive'
size='sm'
onClick={() =>
setCleanupTarget({ target: item.target, label: item.label })
}
>
<Trash2 className='size-3.5 mr-1' />
清理
</Button>
</div>
))}
</CardContent>
</Card>
</div>
<UptimeKumaSiteSelectModal
open={uptimeKumaModalOpen}
selectedSites={
@@ -782,56 +628,6 @@ export function OpenFlareOpsSettings() {
updateField('uptime_kuma_selected_sites', sites.join(','))
}
/>
<AlertDialog
open={cleanupTarget !== null}
onOpenChange={(open) => !open && setCleanupTarget(null)}
>
<AlertDialogContent>
<AlertDialogHeader>
<AlertDialogTitle>清理{cleanupTarget?.label}</AlertDialogTitle>
<AlertDialogDescription>
输入保留天数后会按该表 TTL 物化过期数据;小于表 TTL
的天数会被拒绝。留空则删除全部历史记录,操作不可恢复。
</AlertDialogDescription>
</AlertDialogHeader>
<FieldInput
label='保留天数'
value={cleanupRetentionDays}
type='number'
onChange={setCleanupRetentionDays}
placeholder='留空则全部删除'
/>
<AlertDialogFooter>
<AlertDialogCancel disabled={cleanupMutation.isPending}>
取消
</AlertDialogCancel>
<AlertDialogAction
disabled={cleanupMutation.isPending}
onClick={(event) => {
event.preventDefault();
if (!cleanupTarget) return;
const trimmed = cleanupRetentionDays.trim();
if (trimmed !== '') {
const retentionDays = Number.parseInt(trimmed, 10);
if (Number.isNaN(retentionDays) || retentionDays < 1) {
toast.error('保留天数至少为 1 天');
return;
}
cleanupMutation.mutate({
target: cleanupTarget.target,
retention_days: retentionDays,
});
return;
}
cleanupMutation.mutate({ target: cleanupTarget.target });
}}
>
{cleanupMutation.isPending ? '清理中...' : '确认清理'}
</AlertDialogAction>
</AlertDialogFooter>
</AlertDialogContent>
</AlertDialog>
</div>
);
}
@@ -1,6 +1,6 @@
'use client';
import { useCallback, useEffect, useState } from 'react';
import { useCallback, useEffect, useMemo, useState } from 'react';
import { toast } from 'sonner';
import { Button } from '@/components/ui/button';
import { Input } from '@/components/ui/input';
@@ -8,6 +8,13 @@ import { Label } from '@/components/ui/label';
import { Textarea } from '@/components/ui/textarea';
import { Switch } from '@/components/ui/switch';
import { Spinner } from '@/components/ui/spinner';
import {
Select,
SelectContent,
SelectItem,
SelectTrigger,
SelectValue,
} from '@/components/ui/select';
import {
Dialog,
DialogContent,
@@ -19,12 +26,17 @@ import {
import {
Calendar as CalendarIcon,
Clock,
Database,
Info,
Layers,
Play,
} from 'lucide-react';
import type { DispatchTaskRequest, TaskMeta } from '@/lib/services/admin';
import type {
DispatchTaskRequest,
LogDatabaseStatus,
TaskMeta,
} from '@/lib/services/admin';
import services from '@/lib/services';
import { buildTaskPayload } from '@/lib/task-param-utils';
import { ErrorInline } from '@/components/layout/error';
@@ -68,6 +80,18 @@ const TASK_CONFIGS: Record<
gradient:
'from-rose-500/10 via-rose-500/5 to-transparent border-rose-200/50 dark:border-rose-800/50 hover:border-rose-400 dark:hover:border-rose-500',
},
of_log_db_switch: {
icon: Database,
color: 'text-teal-600 dark:text-teal-400',
gradient:
'from-teal-500/10 via-teal-500/5 to-transparent border-teal-200/50 dark:border-teal-800/50 hover:border-teal-400 dark:hover:border-teal-500',
},
};
const LOG_DATABASE_LABELS: Record<string, string> = {
postgres: 'PostgreSQL(主库)',
sqlite: 'SQLite(主库)',
clickhouse: 'ClickHouse',
};
const DEFAULT_TASK_CONFIG = {
@@ -192,9 +216,38 @@ export function TaskManager() {
}
}, []);
const [logDbStatus, setLogDbStatus] = useState<LogDatabaseStatus | null>(
null,
);
// 日志库状态用于「切换日志数据库」卡片与 target 下拉;获取失败不阻塞任务列表。
const fetchLogDbStatus = useCallback(async () => {
try {
const data = await services.adminStatus.getLogDatabaseStatus();
setLogDbStatus(data);
} catch {
setLogDbStatus(null);
}
}, []);
useEffect(() => {
fetchTaskTypes();
}, [fetchTaskTypes]);
fetchLogDbStatus();
}, [fetchTaskTypes, fetchLogDbStatus]);
const availableLogDbTargets = useMemo(
() => logDbStatus?.available_targets ?? [],
[logDbStatus],
);
const retentionSummary = useMemo(() => {
const days = logDbStatus?.retention_days ?? {};
const parts: string[] = [];
if (days.postgres != null) parts.push(`PG ${days.postgres}`);
if (days.sqlite != null) parts.push(`SQLite ${days.sqlite}`);
if (days.clickhouse != null) parts.push(`CH ${days.clickhouse}`);
return parts.join(' / ');
}, [logDbStatus]);
useEffect(() => {
if (selectedTaskType) {
@@ -344,6 +397,45 @@ export function TaskManager() {
</div>
</div>
{task.type === 'of_log_db_switch' && logDbStatus && (
<div className='pt-3 mt-3 border-t border-border/50 space-y-1.5'>
<div className='flex items-center justify-between gap-2'>
<span className='text-[10px] text-muted-foreground shrink-0'>
日志主库
</span>
<span className='text-[10px] font-mono text-foreground truncate'>
{LOG_DATABASE_LABELS[logDbStatus.active_database] ||
logDbStatus.active_database}
</span>
</div>
<div className='flex items-center justify-between gap-2'>
<span className='text-[10px] text-muted-foreground shrink-0'>
保留天数
</span>
<span className='text-[10px] font-mono text-muted-foreground truncate'>
{retentionSummary || '-'}
</span>
</div>
<div className='flex items-center justify-between gap-2'>
<span className='text-[10px] text-muted-foreground shrink-0'>
迁移状态
</span>
<Badge
variant={
logDbStatus.migration === 'migrating'
? 'default'
: 'outline'
}
className='text-[10px] h-5 px-1.5'
>
{logDbStatus.migration === 'migrating'
? '迁移中'
: '空闲'}
</Badge>
</div>
</div>
)}
<div className='pt-4 mt-1'>
<Button
className='w-full h-7 text-xs'
@@ -436,70 +528,104 @@ export function TaskManager() {
}
return (
<div className='space-y-4'>
{targetTask.params.map((param) => (
<div key={param.name} className='grid gap-2'>
<Label
htmlFor={`param-${param.name}`}
className='flex items-center gap-1'
>
{param.label}
{param.required && (
<span className='text-destructive font-bold'>*</span>
)}
</Label>
{param.type === 'text' ? (
<Textarea
id={`param-${param.name}`}
placeholder={param.placeholder}
className='text-xs min-h-[80px]'
value={paramValues[param.name] || ''}
onChange={(e) =>
setParamValues((prev) => ({
...prev,
[param.name]: e.target.value,
}))
}
/>
) : param.type === 'boolean' ? (
<div className='flex items-center gap-2 pt-1 h-9'>
<Switch
{targetTask.params.map((param) => {
const isSwitchTarget =
param.name === 'target' &&
getSelectedTaskMeta()?.type === 'of_log_db_switch';
return (
<div key={param.name} className='grid gap-2'>
<Label
htmlFor={`param-${param.name}`}
className='flex items-center gap-1'
>
{param.label}
{param.required && (
<span className='text-destructive font-bold'>
*
</span>
)}
</Label>
{param.type === 'text' ? (
<Textarea
id={`param-${param.name}`}
checked={paramValues[param.name] === 'true'}
onCheckedChange={(checked) =>
placeholder={param.placeholder}
className='text-xs min-h-[80px]'
value={paramValues[param.name] || ''}
onChange={(e) =>
setParamValues((prev) => ({
...prev,
[param.name]: checked ? 'true' : 'false',
[param.name]: e.target.value,
}))
}
/>
<span className='text-xs text-muted-foreground'>
{paramValues[param.name] === 'true'
? '开启'
: '关闭'}
</span>
</div>
) : (
<Input
id={`param-${param.name}`}
type={param.type === 'number' ? 'number' : 'text'}
placeholder={param.placeholder}
className='text-xs'
value={paramValues[param.name] || ''}
onChange={(e) =>
setParamValues((prev) => ({
...prev,
[param.name]: e.target.value,
}))
}
/>
)}
{param.description && (
<p className='text-[10px] text-muted-foreground'>
{param.description}
</p>
)}
</div>
))}
) : param.type === 'boolean' ? (
<div className='flex items-center gap-2 pt-1 h-9'>
<Switch
id={`param-${param.name}`}
checked={paramValues[param.name] === 'true'}
onCheckedChange={(checked) =>
setParamValues((prev) => ({
...prev,
[param.name]: checked ? 'true' : 'false',
}))
}
/>
<span className='text-xs text-muted-foreground'>
{paramValues[param.name] === 'true'
? '开启'
: '关闭'}
</span>
</div>
) : isSwitchTarget &&
availableLogDbTargets.length > 0 ? (
<Select
value={paramValues[param.name] || ''}
onValueChange={(value) =>
setParamValues((prev) => ({
...prev,
[param.name]: value,
}))
}
disabled={dispatching}
>
<SelectTrigger
id={`param-${param.name}`}
className='w-full text-xs'
size='sm'
>
<SelectValue placeholder='选择目标日志库...' />
</SelectTrigger>
<SelectContent>
{availableLogDbTargets.map((target) => (
<SelectItem key={target} value={target}>
{LOG_DATABASE_LABELS[target] || target}
</SelectItem>
))}
</SelectContent>
</Select>
) : (
<Input
id={`param-${param.name}`}
type={param.type === 'number' ? 'number' : 'text'}
placeholder={param.placeholder}
className='text-xs'
value={paramValues[param.name] || ''}
onChange={(e) =>
setParamValues((prev) => ({
...prev,
[param.name]: e.target.value,
}))
}
/>
)}
{param.description && (
<p className='text-[10px] text-muted-foreground'>
{param.description}
</p>
)}
</div>
);
})}
</div>
);
})()}
@@ -94,7 +94,7 @@ export function ErrorPageTab({
<div className='space-y-1.5'>
<CardTitle className='text-base'>触发策略</CardTitle>
<CardDescription>
配置源站错误页的启用条件与触发状态码。修改后需点击保存。
配置源站错误页的启用条件与触发状态码。
</CardDescription>
</div>
<Button
@@ -133,8 +133,7 @@ export function ErrorPageTab({
<div className='space-y-1'>
<Label className='text-sm font-medium'>仅针对 GET 请求</Label>
<p className='text-sm text-muted-foreground'>
开启后仅对 GET 请求的匹配错误状态码返回自定义错误页;POST/PUT
等其它方法直接透传源站响应。
开启后仅对 GET 请求的匹配错误状态码返回自定义错误页。
</p>
</div>
<Switch
@@ -157,8 +156,7 @@ export function ErrorPageTab({
触发状态码
</Label>
<p className='text-sm text-muted-foreground'>
支持单码(如 502)或闭区间(如 500-599),范围 400–599。默认
500-599。
支持单码(如 502)或闭区间(如 500-599)。
</p>
</div>
<TagsInput
@@ -189,7 +187,7 @@ export function ErrorPageTab({
<Button variant='outline' size='sm' asChild>
<Link href='/responses/error-page/preview'>
<Expand className='size-3.5' />
真实预览
预览
</Link>
</Button>
<Button size='sm' asChild>
@@ -91,10 +91,10 @@ export function OfflinePageTab({
<Card className='border-dashed shadow-none'>
<CardHeader className='flex flex-row items-start justify-between gap-4 space-y-0'>
<div className='space-y-1.5'>
<CardTitle className='text-base'>离线兜底</CardTitle>
<CardTitle className='text-base'>离线页</CardTitle>
<CardDescription>
启用后给启用 HTTPS 的网站下发 Service
Worker,域名被墙时浏览器从缓存展示此离线页。
Worker,域名无法访问时浏览器从缓存展示此离线页。
</CardDescription>
</div>
<Button
@@ -115,10 +115,10 @@ export function OfflinePageTab({
<div className='flex items-start justify-between gap-6'>
<div className='space-y-1'>
<Label className='text-sm font-medium'>
启用 Service Worker 离线兜底
启用 Service Worker 离线页
</Label>
<p className='text-sm text-muted-foreground'>
仅对 HTTPS 网站生效;未启用的站点不受影响。
仅对 HTTPS 网站生效。
</p>
</div>
<Switch
@@ -126,7 +126,7 @@ export function OfflinePageTab({
onCheckedChange={(enabled) =>
setFields((prev) => ({ ...prev, enabled }))
}
aria-label='启用离线兜底'
aria-label='启用离线页'
className='mt-0.5 shrink-0'
/>
</div>
@@ -194,7 +194,7 @@ export function OfflinePageTab({
<Button variant='outline' size='sm' asChild>
<Link href='/responses/offline/preview'>
<Expand className='size-3.5' />
真实预览
预览
</Link>
</Button>
<Button size='sm' asChild>
@@ -87,7 +87,7 @@ export default function ErrorPagePreviewPage() {
</Link>
</Button>
<div className='min-w-0'>
<p className='text-sm font-semibold truncate'>错误页真实预览</p>
<p className='text-sm font-semibold truncate'>错误页预览</p>
<p className='text-[11px] text-muted-foreground font-mono truncate'>
{'{{status}}'}→502 · {'{{host}}'}→example.com · 全屏展示
</p>
@@ -108,7 +108,7 @@ export default function ErrorPagePreviewPage() {
</div>
</div>
<iframe
title='源站错误页真实预览'
title='源站错误页预览'
sandbox=''
srcDoc={previewSrcDoc}
className='flex-1 w-full border-0 bg-background min-h-0'
@@ -88,7 +88,7 @@ export default function OfflinePagePreviewPage() {
</Link>
</Button>
<div className='min-w-0'>
<p className='text-sm font-semibold truncate'>离线页真实预览</p>
<p className='text-sm font-semibold truncate'>离线页预览</p>
<p className='text-[11px] text-muted-foreground font-mono truncate'>
全屏展示
</p>
@@ -109,7 +109,7 @@ export default function OfflinePagePreviewPage() {
</div>
</div>
<iframe
title='离线页真实预览'
title='离线页预览'
sandbox=''
srcDoc={html}
className='flex-1 w-full border-0 bg-background min-h-0'
+1 -1
View File
@@ -115,7 +115,7 @@ function ResponsesPageContent() {
<div>
<h1 className='text-2xl font-semibold tracking-tight'>响应页面</h1>
<p className='text-sm text-muted-foreground'>
配置源站错误页与离线页。保存后需发布配置版本后生效。
配置源站错误页与离线页。
</p>
</div>
</div>
@@ -139,7 +139,7 @@ export function HtmlEditorWorkspace({
>
<Link href='/error-pages/preview'>
<Expand className='size-3' />
真实预览
预览
</Link>
</Button>
) : null}
@@ -1,13 +1,13 @@
'use client';
import { useMemo } from 'react';
import { useEffect, useMemo, useState } from 'react';
import {
useMutation,
useQuery,
useQueryClient,
type UseQueryResult,
} from '@tanstack/react-query';
import { KeyRound, ShieldAlert, X } from 'lucide-react';
import { Database, KeyRound, Save, ShieldAlert, X } from 'lucide-react';
import {
Card,
CardContent,
@@ -16,6 +16,10 @@ import {
CardTitle,
} from '@/components/ui/card';
import { Badge } from '@/components/ui/badge';
import { Button } from '@/components/ui/button';
import { Input } from '@/components/ui/input';
import { Spinner } from '@/components/ui/spinner';
import { Label } from '@/components/ui/label';
import {
Select,
SelectContent,
@@ -28,6 +32,24 @@ import type { SystemConfig } from '@/lib/services/admin';
import { TemplatesManager } from './templates';
import { toast } from 'sonner';
const LOG_RETENTION_FIELDS = [
{
key: 'log_retention_days_postgres',
label: 'PostgreSQL',
description: '访问日志与可观测指标统一保留天数',
},
{
key: 'log_retention_days_sqlite',
label: 'SQLite',
description: 'SQLite 日志保留天数',
},
{
key: 'log_retention_days_clickhouse',
label: 'ClickHouse',
description: 'ClickHouse 日志保留天数',
},
] as const;
interface OperationTabProps {
configs: Record<string, SystemConfig>;
systemConfigsQuery: UseQueryResult<SystemConfig[], Error>;
@@ -44,6 +66,65 @@ export function OperationTab({
queryFn: () => services.adminSystemConfig.listUploadTypes(),
});
const businessConfigsQuery = useQuery({
queryKey: ['admin', 'system-configs', 'business'],
queryFn: () => services.adminSystemConfig.listSystemConfigs('business'),
});
const businessConfigs = useMemo(() => {
return (businessConfigsQuery.data ?? []).reduce<
Record<string, SystemConfig>
>((acc, config) => {
acc[config.key] = config;
return acc;
}, {});
}, [businessConfigsQuery.data]);
const [retentionValues, setRetentionValues] = useState<
Record<string, string>
>({});
useEffect(() => {
if (!businessConfigsQuery.data) return;
setRetentionValues((prev) => {
const next: Record<string, string> = {};
LOG_RETENTION_FIELDS.forEach((field) => {
const config = businessConfigs[field.key];
next[field.key] = config?.value || prev[field.key] || '90';
});
return next;
});
}, [businessConfigsQuery.data, businessConfigs]);
const updateRetentionMutation = useMutation({
mutationFn: async (values: Record<string, string>) => {
for (const field of LOG_RETENTION_FIELDS) {
const raw = (values[field.key] ?? '').trim();
const num = Number(raw);
if (!raw || !Number.isInteger(num) || num < 1) {
throw new Error(`${field.label}必须为大于等于 1 的整数`);
}
const config = businessConfigs[field.key];
if (!config) {
throw new Error(`缺少配置项: ${field.key}`);
}
await services.adminSystemConfig.updateSystemConfig(field.key, {
value: String(num),
description: config.description,
});
}
},
onSuccess: async () => {
await queryClient.invalidateQueries({
queryKey: ['admin', 'system-configs'],
});
toast.success('日志保留时间已更新');
},
onError: (error: Error) => {
toast.error(error.message || '更新日志保留时间失败');
},
});
const updateWhitelistMutation = useMutation({
mutationFn: async (newValue: string) => {
const config = configs['file_access_whitelist'];
@@ -208,6 +289,77 @@ export function OperationTab({
</CardContent>
</Card>
{/* 日志保留时间设置 */}
<Card className='border border-dashed shadow-sm'>
<CardHeader className='border-b border-dashed pb-4'>
<div className='flex items-center gap-2'>
<div className='p-1.5 rounded-lg bg-primary/10 text-primary'>
<Database className='size-4' />
</div>
<div>
<CardTitle className='text-base font-semibold'>
日志保留时间
</CardTitle>
<CardDescription className='text-xs'>
配置各日志数据库的日志保留天数,切换日志数据库后自动按对应配置清理过期日志。
</CardDescription>
</div>
</div>
</CardHeader>
<CardContent className='pt-6'>
<div className='grid gap-4 sm:grid-cols-3'>
{LOG_RETENTION_FIELDS.map((field) => (
<div key={field.key} className='grid gap-2'>
<Label htmlFor={field.key}>{field.label}</Label>
<div className='flex items-center gap-2'>
<Input
id={field.key}
type='number'
min={1}
className='text-xs'
value={retentionValues[field.key] ?? ''}
disabled={
updateRetentionMutation.isPending ||
businessConfigsQuery.isPending
}
onChange={(e) =>
setRetentionValues((prev) => ({
...prev,
[field.key]: e.target.value,
}))
}
/>
<span className='text-xs text-muted-foreground whitespace-nowrap'>
天
</span>
</div>
<p className='text-[10px] text-muted-foreground'>
{field.description}
</p>
</div>
))}
</div>
<div className='mt-4 flex justify-end'>
<Button
type='button'
size='sm'
onClick={() => updateRetentionMutation.mutate(retentionValues)}
disabled={
updateRetentionMutation.isPending ||
businessConfigsQuery.isPending
}
>
{updateRetentionMutation.isPending ? (
<Spinner className='size-3' />
) : (
<Save className='size-3' />
)}
保存
</Button>
</div>
</CardContent>
</Card>
{/* 通知模板管理 */}
<TemplatesManager />
</div>
+1 -2
View File
@@ -55,7 +55,6 @@ import {
LogOut,
Settings,
ShieldCheck,
Terminal,
UserRound,
} from 'lucide-react';
@@ -72,7 +71,7 @@ const data = {
{ title: '存储管理', url: '/admin/files', icon: FolderOpen },
{ title: '数据管理', url: '/admin/database', icon: Database },
{ title: '通知推送', url: '/admin/push', icon: Bell },
{ title: '系统日志', url: '/admin/logs', icon: Terminal },
// { title: '系统日志', url: '/admin/logs', icon: Terminal },
{ title: '系统配置', url: '/admin/system', icon: ShieldCheck },
{ title: '系统设置', url: '/admin/settings', icon: Settings },
],
+1
View File
@@ -29,6 +29,7 @@ export type {
UpdateUserStatusRequest,
UpdateUserRequest,
SystemStatus,
LogDatabaseStatus,
AppUpdateStatus,
Schedule,
CreateScheduleRequest,
@@ -1,5 +1,5 @@
import { BaseService } from '@/lib/services/core';
import type { AppUpdateStatus, SystemStatus } from './types';
import type { AppUpdateStatus, LogDatabaseStatus, SystemStatus } from './types';
export class AdminStatusService extends BaseService {
protected static readonly basePath = '/api/v1/admin';
@@ -8,6 +8,10 @@ export class AdminStatusService extends BaseService {
return this.get<SystemStatus>('/status');
}
static async getLogDatabaseStatus(): Promise<LogDatabaseStatus> {
return this.get<LogDatabaseStatus>('/status/log-database');
}
static async getUpdateStatus(): Promise<AppUpdateStatus> {
return this.get<AppUpdateStatus>('/update');
}
+16
View File
@@ -405,6 +405,22 @@ export interface ToggleAuthSourceRequest {
is_active: boolean;
}
/**
* 日志数据库状态
*/
export interface LogDatabaseStatus {
/** 当前日志主库:postgres | sqlite | clickhouse */
active_database: string;
/** 迁移状态:idle | migrating */
migration: string;
/** 各日志库保留天数 */
retention_days: Record<string, number>;
/** 当前主库的合法迁移目标 */
available_targets: string[];
/** ClickHouse 运行指标(仅主库为 ClickHouse 时) */
clickhouse?: unknown;
}
/**
* 系统状态信息
*/
+1 -1
View File
@@ -158,6 +158,7 @@ export type {
CreateUserRequest,
UpdateUserRequest,
SystemStatus,
LogDatabaseStatus,
AppUpdateStatus,
Schedule,
CreateScheduleRequest,
@@ -260,6 +261,5 @@ export type {
AccessLogOverview,
OptionItem,
GeoIPLookupResult,
DatabaseCleanupResult,
OpenFlarePublicStatus,
} from './openflare';
-3
View File
@@ -104,9 +104,6 @@ export type {
FoldedAccessLogList,
OptionItem,
GeoIPLookupResult,
DatabaseCleanupPayload,
DatabaseCleanupResult,
DatabaseCleanupTarget,
OpenFlarePublicStatus,
OriginDetail,
OriginItem,
@@ -1,7 +1,5 @@
import { OpenFlareBaseService } from './base.service';
import type {
DatabaseCleanupPayload,
DatabaseCleanupResult,
GeoIPLookupResult,
OptionBatchPayload,
OptionItem,
@@ -26,10 +24,4 @@ export class OptionService extends OpenFlareBaseService {
static lookupGeoIP(provider: string, ip: string): Promise<GeoIPLookupResult> {
return this.post<GeoIPLookupResult>('/geoip/lookup', { provider, ip });
}
static cleanupDatabase(
payload: DatabaseCleanupPayload,
): Promise<DatabaseCleanupResult> {
return this.post<DatabaseCleanupResult>('/database/cleanup', payload);
}
}
-21
View File
@@ -866,27 +866,6 @@ export interface GeoIPLookupResult {
longitude?: number | null;
}
export type DatabaseCleanupTarget =
| 'node_access_logs'
| 'node_metric_snapshots'
| 'node_edge_health'
| 'node_obs_frps'
| 'node_obs_frpc';
export interface DatabaseCleanupPayload {
target: DatabaseCleanupTarget;
retention_days?: number;
}
export interface DatabaseCleanupResult {
target: DatabaseCleanupTarget;
target_label: string;
deleted_count: number;
delete_all: boolean;
retention_days?: number;
cutoff?: string;
}
export interface OpenFlarePublicStatus {
version: string;
start_time: number;
+20 -17
View File
@@ -14,10 +14,9 @@ import (
"time"
"github.com/Rain-kl/Wavelet/internal/apps/admin"
"github.com/Rain-kl/Wavelet/internal/infra/config"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/internal/repository"
analyticsrepo "github.com/Rain-kl/Wavelet/internal/repository/analytics"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
"github.com/Rain-kl/Wavelet/pkg/logger"
"github.com/gin-gonic/gin"
@@ -158,8 +157,8 @@ type accessLogsResponse struct {
List []accessLogItem `json:"list"`
}
func buildAccessLogFilter(ctx context.Context, c *gin.Context) (analyticsrepo.AccessLogFilter, error) {
filter := analyticsrepo.AccessLogFilter{}
func buildAccessLogFilter(ctx context.Context, c *gin.Context) (analyticsmodel.AccessLogFilter, error) {
filter := analyticsmodel.AccessLogFilter{}
username := c.Query("username")
if username != "" {
@@ -227,7 +226,7 @@ func enrichAccessLogsWithUsers(ctx context.Context, list []accessLogItem) {
// GetAccessLogs 获取 ClickHouse 异步采集的访问日志
// @Summary 获取用户访问日志
// @Description 分页并按照用户、接口路径、时间范围等维度检索 ClickHouse 用户访问日志列表(需要管理员权限,ClickHouse 未启用时报错)
// @Description 分页并按照用户、接口路径、时间范围等维度检索用户访问日志列表(需要管理员权限,日志存储未启用时报错)
// @Tags admin
// @Produce json
// @Security SessionCookie
@@ -238,14 +237,16 @@ func enrichAccessLogsWithUsers(ctx context.Context, list []accessLogItem) {
// @Param start_time query string false "起始时间(RFC3339 或 YYYY-MM-DD HH:MM:SS)"
// @Param end_time query string false "结束时间(RFC3339 或 YYYY-MM-DD HH:MM:SS)"
// @Success 200 {object} response.Any{data=logs.accessLogsResponse} "访问日志列表"
// @Failure 400 {object} response.Any "ClickHouse 未启用或参数错误"
// @Failure 400 {object} response.Any "日志存储未启用或参数错误"
// @Failure 401 {object} response.Any "未登录"
// @Failure 403 {object} response.Any "无管理员权限"
// @Router /api/v1/admin/logs/access [get]
func GetAccessLogs(c *gin.Context) {
ctx := c.Request.Context()
if !config.Config.ClickHouse.Enabled || !db.ChConnReady() {
response.AbortWithError(c, http.StatusBadRequest, "ClickHouse 存储服务未启用,无法检索访问日志")
store, err := logstore.Active(ctx)
if err != nil {
logger.ErrorF(ctx, "获取日志存储实例失败: %v", err)
response.AbortWithError(c, http.StatusBadRequest, "日志存储未启用,无法检索访问日志")
return
}
@@ -271,7 +272,7 @@ func GetAccessLogs(c *gin.Context) {
return
}
logs, total, err := analyticsrepo.ListAccessLogs(ctx, filter, page, pageSize)
logs, total, err := store.UserAccessLogs.List(ctx, filter, page, pageSize)
if err != nil {
response.AbortWithError(c, http.StatusInternalServerError, err.Error())
return
@@ -333,25 +334,27 @@ type logsAnalyticsResponse struct {
// GetLogsAnalytics 获取 ClickHouse 访问日志图表聚合指标
// @Summary 获取访问日志分析数据
// @Description 聚合统计最近 7 天的每日访问趋势、浏览器分布以及前 10 名最活跃用户排行(需要管理员权限,ClickHouse 未启用时报错)
// @Description 聚合统计最近 7 天的每日访问趋势、浏览器分布以及前 10 名最活跃用户排行(需要管理员权限,日志存储未启用时报错)
// @Tags admin
// @Produce json
// @Security SessionCookie
// @Success 200 {object} response.Any{data=logs.logsAnalyticsResponse} "分析统计数据"
// @Failure 400 {object} response.Any "ClickHouse 未启用"
// @Failure 400 {object} response.Any "日志存储未启用"
// @Failure 401 {object} response.Any "未登录"
// @Failure 403 {object} response.Any "无管理员权限"
// @Router /api/v1/admin/logs/analytics [get]
func GetLogsAnalytics(c *gin.Context) {
ctx := c.Request.Context()
if !config.Config.ClickHouse.Enabled || !db.ChConnReady() {
response.AbortWithError(c, http.StatusBadRequest, "ClickHouse 存储服务未启用,无法获取分析数据")
store, err := logstore.Active(ctx)
if err != nil {
logger.ErrorF(ctx, "获取日志存储实例失败: %v", err)
response.AbortWithError(c, http.StatusBadRequest, "日志存储未启用,无法获取分析数据")
return
}
startTime := time.Now().AddDate(0, 0, -(analyticsDays - 1)).Truncate(hoursInDay * time.Hour)
trendPoints, err := analyticsrepo.GetDailyTrend(ctx, analyticsDays)
trendPoints, err := store.UserAccessLogs.GetDailyTrend(ctx, analyticsDays)
if err != nil {
response.AbortWithError(c, http.StatusInternalServerError, "查询访问趋势失败: "+err.Error())
return
@@ -364,7 +367,7 @@ func GetLogsAnalytics(c *gin.Context) {
}
}
browserPoints, err := analyticsrepo.GetBrowserDistribution(ctx, startTime)
browserPoints, err := store.UserAccessLogs.GetBrowserDistribution(ctx, startTime)
if err != nil {
response.AbortWithError(c, http.StatusInternalServerError, "查询浏览器分布失败: "+err.Error())
return
@@ -377,7 +380,7 @@ func GetLogsAnalytics(c *gin.Context) {
}
}
topUserPoints, err := analyticsrepo.GetTopActiveUsers(ctx, startTime, topActiveLimit)
topUserPoints, err := store.UserAccessLogs.GetTopActiveUsers(ctx, startTime, topActiveLimit)
if err != nil {
response.AbortWithError(c, http.StatusInternalServerError, "查询活跃用户失败: "+err.Error())
return
+99 -17
View File
@@ -4,43 +4,126 @@
package status
import (
"context"
"errors"
"net/http"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/chwriter"
"github.com/Rain-kl/Wavelet/internal/apps/risk_control"
"github.com/Rain-kl/Wavelet/internal/infra/config"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/infra/persistence/batchwriter"
analyticsrepo "github.com/Rain-kl/Wavelet/internal/repository/analytics"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
"github.com/Rain-kl/Wavelet/internal/shared/response"
"github.com/Rain-kl/Wavelet/pkg/logger"
"github.com/gin-gonic/gin"
"gorm.io/gorm"
)
// GetClickHouseStatus returns ClickHouse operational metrics for administrators.
// @Summary 获取 ClickHouse 运行指标
// @Description 返回 ClickHouse parts、mutation、async_insert 队列及进程内 batch writer 指标,需要管理员权限
// 日志库名取值(与 model 配置值、logstore provider 分支保持一致)。
const (
logDBNamePostgres = "postgres"
logDBNameSQLite = "sqlite"
logDBNameClickHouse = "clickhouse"
)
// defaultLogRetentionDays 日志保留天数配置缺失时的兜底值(与 seed 默认一致)。
const defaultLogRetentionDays = 90
// LogDatabaseStatus 日志库状态。
type LogDatabaseStatus struct {
ActiveDatabase string `json:"active_database"`
Migration string `json:"migration"` // idle | migrating
RetentionDays map[string]int `json:"retention_days"`
AvailableTargets []string `json:"available_targets"`
ClickHouse *analyticsmodel.ClickHouseOperationalStats `json:"clickhouse,omitempty"`
}
// GetLogDatabaseStatus 返回当前日志库状态。
// @Summary 获取日志数据库状态
// @Description 返回当前日志主库、迁移状态、各库保留天数与合法迁移目标,需要管理员权限
// @Tags admin
// @Produce json
// @Security SessionCookie
// @Success 200 {object} response.Any{data=analyticsrepo.ClickHouseOperationalStats} "获取成功"
// @Failure 400 {object} response.Any "ClickHouse 未启用"
// @Success 200 {object} response.Any{data=status.LogDatabaseStatus} "获取成功"
// @Failure 401 {object} response.Any "未登录"
// @Failure 403 {object} response.Any "无管理员权限"
// @Failure 500 {object} response.Any "内部错误"
// @Router /api/v1/admin/status/clickhouse [get]
func GetClickHouseStatus(c *gin.Context) {
if !config.Config.ClickHouse.Enabled || !db.ChConnReady() {
response.AbortWithError(c, http.StatusBadRequest, "ClickHouse 存储服务未启用")
// @Router /api/v1/admin/status/log-database [get]
func GetLogDatabaseStatus(c *gin.Context) {
ctx := c.Request.Context()
store, err := logstore.Active(ctx)
if err != nil {
logger.ErrorF(ctx, "获取日志存储实例失败: %v", err)
response.AbortInternal(c, "日志存储初始化失败")
return
}
stats, err := analyticsrepo.GetClickHouseOperationalStats(c.Request.Context())
// 分支判定复用同一 store 实例的 ActiveDatabase,避免与 Active 解析之间出现 TOCTOU。
activeDB, err := store.Status.ActiveDatabase(ctx)
if err != nil {
response.AbortInternal(c, "获取 ClickHouse 运行指标失败")
logger.ErrorF(ctx, "获取日志库状态失败: %v", err)
response.AbortInternal(c, "获取日志库状态失败")
return
}
stats.BatchWriters = collectBatchWriterStats()
c.JSON(http.StatusOK, response.OK(stats))
migration := "idle"
if logstore.Migrating(ctx) {
migration = "migrating"
}
out := LogDatabaseStatus{
ActiveDatabase: activeDB,
Migration: migration,
RetentionDays: map[string]int{
logDBNamePostgres: retentionOr(ctx, model.ConfigKeyLogRetentionDaysPostgres),
logDBNameSQLite: retentionOr(ctx, model.ConfigKeyLogRetentionDaysSQLite),
logDBNameClickHouse: retentionOr(ctx, model.ConfigKeyLogRetentionDaysClickHouse),
},
AvailableTargets: availableTargets(activeDB),
}
if activeDB == logDBNameClickHouse {
stats, err := store.Status.ClickHouseOperationalStats(ctx)
if err != nil {
logger.ErrorF(ctx, "获取 ClickHouse 运行指标失败: %v", err)
} else {
stats.BatchWriters = collectBatchWriterStats()
out.ClickHouse = stats
}
}
c.JSON(http.StatusOK, response.OK(out))
}
// retentionOr 读取保留天数配置,缺失或非法时返回默认值。
func retentionOr(ctx context.Context, key string) int {
v, err := repository.GetIntByKey(ctx, key)
if err != nil {
if !errors.Is(err, gorm.ErrRecordNotFound) {
logger.ErrorF(ctx, "读取日志保留天数配置失败 key=%s: %v", key, err)
}
return defaultLogRetentionDays
}
if v < 1 {
return defaultLogRetentionDays
}
return v
}
// availableTargets 返回当前日志主库的合法迁移目标(复用调用方已解析的 active):
// 当前为 clickhouse 时目标为主库(postgres/sqlite 按启动配置);当前为主库时目标为 clickhouse(仅 CH 启用时)。
func availableTargets(active string) []string {
if active == logDBNameClickHouse {
if config.Config.Database.Enabled {
return []string{logDBNamePostgres}
}
return []string{logDBNameSQLite}
}
if config.Config.ClickHouse.Enabled {
return []string{logDBNameClickHouse}
}
return []string{}
}
func collectBatchWriterStats() []batchwriter.Stats {
@@ -48,6 +131,5 @@ func collectBatchWriterStats() []batchwriter.Stats {
if out == nil {
out = make([]batchwriter.Stats, 0, 1)
}
out = append(out, risk_control.LogWriterStats())
return out
}
@@ -0,0 +1,117 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package status
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"reflect"
"testing"
"github.com/Rain-kl/Wavelet/internal/infra/config"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
"github.com/gin-gonic/gin"
)
// restoreConfig 恢复测试中临时修改的全局配置。
func restoreConfig(t *testing.T) {
t.Helper()
dbEnabled := config.Config.Database.Enabled
chEnabled := config.Config.ClickHouse.Enabled
t.Cleanup(func() {
config.Config.Database.Enabled = dbEnabled
config.Config.ClickHouse.Enabled = chEnabled
})
}
func TestAvailableTargets(t *testing.T) {
restoreConfig(t)
logstore.ResetForTest()
t.Cleanup(logstore.ResetForTest)
// 当前 clickhouse → 主库(postgres/sqlite 按启动配置)。
logstore.SetConfigReader(func(_ context.Context, key string) (string, error) {
if key == "log_database" {
return logDBNameClickHouse, nil
}
return "", nil
})
config.Config.Database.Enabled = true
config.Config.ClickHouse.Enabled = true
if got := availableTargets(logDBNameClickHouse); !reflect.DeepEqual(got, []string{logDBNamePostgres}) {
t.Fatalf("clickhouse active + postgres main: got %v, want [postgres]", got)
}
config.Config.Database.Enabled = false
if got := availableTargets(logDBNameClickHouse); !reflect.DeepEqual(got, []string{logDBNameSQLite}) {
t.Fatalf("clickhouse active + sqlite main: got %v, want [sqlite]", got)
}
// 当前主库 → clickhouse(CH 启用时)。
logstore.SetConfigReader(func(_ context.Context, key string) (string, error) {
if key == "log_database" {
return logDBNamePostgres, nil
}
return "", nil
})
config.Config.ClickHouse.Enabled = true
if got := availableTargets(logDBNamePostgres); !reflect.DeepEqual(got, []string{logDBNameClickHouse}) {
t.Fatalf("postgres active: got %v, want [clickhouse]", got)
}
// CH 禁用时排除 clickhouse。
config.Config.ClickHouse.Enabled = false
if got := availableTargets(logDBNamePostgres); len(got) != 0 {
t.Fatalf("CH disabled: got %v, want empty", got)
}
}
// TestGetLogDatabaseStatusSmoke 覆盖 handler 的 CH 激活分支(无需 DB/CH 连接)。
func TestGetLogDatabaseStatusSmoke(t *testing.T) {
restoreConfig(t)
config.Config.Database.Enabled = false
config.Config.ClickHouse.Enabled = true
logstore.ResetForTest()
t.Cleanup(logstore.ResetForTest)
logstore.SetConfigReader(func(_ context.Context, key string) (string, error) {
if key == "log_database" {
return logDBNameClickHouse, nil
}
return "", nil
})
gin.SetMode(gin.TestMode)
w := httptest.NewRecorder()
c, _ := gin.CreateTestContext(w)
c.Request = httptest.NewRequest(http.MethodGet, "/api/v1/admin/status/log-database", nil)
GetLogDatabaseStatus(c)
if w.Code != http.StatusOK {
t.Fatalf("status code = %d, want 200; body=%s", w.Code, w.Body.String())
}
var resp struct {
Data LogDatabaseStatus `json:"data"`
}
if err := json.Unmarshal(w.Body.Bytes(), &resp); err != nil {
t.Fatalf("unmarshal body: %v", err)
}
if resp.Data.ActiveDatabase != logDBNameClickHouse {
t.Fatalf("active_database = %q, want clickhouse", resp.Data.ActiveDatabase)
}
if resp.Data.Migration != "idle" {
t.Fatalf("migration = %q, want idle", resp.Data.Migration)
}
for _, key := range []string{logDBNamePostgres, logDBNameSQLite, logDBNameClickHouse} {
if got := resp.Data.RetentionDays[key]; got != defaultLogRetentionDays {
t.Fatalf("retention_days[%s] = %d, want default %d", key, got, defaultLogRetentionDays)
}
}
if got := resp.Data.AvailableTargets; !reflect.DeepEqual(got, []string{logDBNameSQLite}) {
t.Fatalf("available_targets = %v, want [sqlite] (test main DB disabled)", got)
}
}
@@ -17,6 +17,11 @@ import (
)
func createSystemConfig(ctx context.Context, req CreateSystemConfigRequest) error {
// 防御:受保护 key(log_database / log_db_migration)仅允许内部写入,Handler 已拦截。
if isProtectedConfigKey(req.Key) {
return errors.New(protectedConfigKeyMessage)
}
exists, err := repository.SystemConfigExists(ctx, req.Key)
if err != nil {
return err
@@ -27,6 +27,17 @@ import (
const maskedConfigValue = "******"
// protectedConfigKeyMessage 命中受保护 key 时返回给管理员的业务错误文案。
const protectedConfigKeyMessage = "该配置项由系统任务管理,禁止手动修改"
// protectedConfigKeys 仅允许内部(迁移任务/bootstrap)写入的 key。
var protectedConfigKeys = map[string]bool{
model.ConfigKeyLogDatabase: true,
model.ConfigKeyLogDBMigration: true,
}
func isProtectedConfigKey(key string) bool { return protectedConfigKeys[key] }
// CreateSystemConfigRequest 创建系统配置请求
type CreateSystemConfigRequest struct {
Key string `json:"key" binding:"required,max=64"`
@@ -64,11 +75,21 @@ func CreateSystemConfig(c *gin.Context) {
return
}
// 与 PUT 路径一致:log_database / log_db_migration 仅允许内部(迁移任务/bootstrap)写入。
if isProtectedConfigKey(req.Key) {
response.AbortBadRequest(c, protectedConfigKeyMessage)
return
}
if err := createSystemConfig(c.Request.Context(), req); err != nil {
if err.Error() == ConfigKeyExists {
response.AbortBadRequest(c, ConfigKeyExists)
return
}
if err.Error() == protectedConfigKeyMessage {
response.AbortBadRequest(c, protectedConfigKeyMessage)
return
}
response.AbortInternal(c, err.Error())
return
}
@@ -155,6 +176,10 @@ func UpdateSystemConfig(c *gin.Context) {
}
key := c.Param("key")
if isProtectedConfigKey(key) {
response.AbortBadRequest(c, protectedConfigKeyMessage)
return
}
if err := updateSystemConfig(c.Request.Context(), key, req); err != nil {
if errors.Is(err, gorm.ErrRecordNotFound) {
response.AbortNotFound(c, SystemConfigNotFound)
@@ -27,7 +27,7 @@ import (
"github.com/Rain-kl/Wavelet/internal/shared/response"
)
const expectedDefaultConfigsCount = 34
const expectedDefaultConfigsCount = 37
func setupTestRouter(authUser *model.User) *gin.Engine {
r := testhelper.NewTestGinEngine()
@@ -168,8 +168,8 @@ func TestListSystemConfigs(t *testing.T) {
var configs []model.SystemConfig
_ = json.Unmarshal(dataBytes, &configs)
if len(configs) != 5 {
t.Errorf("expected 5 business configs, got %d: %v", len(configs), configs)
if len(configs) != 8 {
t.Errorf("expected 8 business configs, got %d: %v", len(configs), configs)
}
})
}
+8 -3
View File
@@ -78,22 +78,27 @@ ngx.say([[<!DOCTYPE html>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="robots" content="noindex,nofollow">
<title>加载中...</title>
<title></title>
<script>
console.debug("[sw-challenge] challenge page loaded, redirect target: ]] .. redir .. [[");
if ("serviceWorker" in navigator) {
console.debug("[sw-challenge] registering service worker /sw.js");
navigator.serviceWorker.register("/sw.js").then(function () {
console.debug("[sw-challenge] service worker registered");
document.cookie = "__openflare_sw=1; Path=/; Max-Age=31536000; Secure; SameSite=Lax";
location.replace("]] .. redir .. [[");
}).catch(function () {
}).catch(function (err) {
console.debug("[sw-challenge] service worker registration failed, redirecting anyway: ", err);
location.replace("]] .. redir .. [[");
});
} else {
console.debug("[sw-challenge] service worker unsupported, redirecting");
document.cookie = "__openflare_sw=1; Path=/; Max-Age=31536000; Secure; SameSite=Lax";
location.replace("]] .. redir .. [[");
}
</script>
</head>
<body>正在加载...</body>
<body></body>
</html>]])
`
+21 -62
View File
@@ -21,11 +21,6 @@ const (
// TaskTypeSSLRenew is the admin task type for SSL renewal.
TaskTypeSSLRenew = "of_ssl_renew"
// DatabaseAutoCleanupTask prunes observability tables by retention policy.
DatabaseAutoCleanupTask = "openflare:database_auto_cleanup"
// TaskTypeDatabaseAutoCleanup is the admin task type for observability cleanup.
TaskTypeDatabaseAutoCleanup = "of_database_auto_cleanup"
// WAFIPGroupSyncTask syncs due automatic/subscription WAF IP groups.
WAFIPGroupSyncTask = "openflare:waf_ip_group_sync"
// TaskTypeWAFIPGroupSync is the admin task type for WAF IP group sync.
@@ -35,6 +30,11 @@ const (
UptimeKumaSyncTask = "openflare:uptime_kuma_sync"
// TaskTypeUptimeKumaSync is the admin task type for Uptime Kuma sync.
TaskTypeUptimeKumaSync = "of_uptime_kuma_sync"
// LogDBSwitchTask 切换日志数据库任务标识。
LogDBSwitchTask = "openflare:log_db_switch"
// TaskTypeLogDBSwitch is the admin task type for log database switch.
TaskTypeLogDBSwitch = "of_log_db_switch"
)
var (
@@ -54,18 +54,6 @@ var SSLRenewMeta = task.TaskMeta{
Retryable: true,
}
// DatabaseAutoCleanupMeta describes the observability auto-cleanup task.
var DatabaseAutoCleanupMeta = task.TaskMeta{
Type: TaskTypeDatabaseAutoCleanup,
AsynqTask: DatabaseAutoCleanupTask,
Name: "OpenFlare 可观测数据自动清理",
Description: "按保留天数清理访问日志、性能快照与请求聚合数据",
SupportsTime: false,
MaxRetry: task.DefaultMaxRetry,
Queue: task.QueueDefault,
Retryable: true,
}
// WAFIPGroupSyncMeta describes the WAF IP group sync task.
var WAFIPGroupSyncMeta = task.TaskMeta{
Type: TaskTypeWAFIPGroupSync,
@@ -90,6 +78,22 @@ var UptimeKumaSyncMeta = task.TaskMeta{
Retryable: true,
}
// LogDBSwitchMeta 描述切换日志数据库任务。
var LogDBSwitchMeta = task.TaskMeta{
Type: TaskTypeLogDBSwitch,
AsynqTask: LogDBSwitchTask,
Name: "切换日志数据库",
Description: "复制迁移日志数据并在成功后切换日志主库(期间禁止日志写入)",
SupportsTime: false,
MaxRetry: task.DefaultMaxRetry,
Queue: task.QueueDefault,
Retryable: true,
Params: []task.TaskParam{
{Name: "target", Label: "目标日志库", Type: "string", Required: true,
Placeholder: "postgres|sqlite|clickhouse", Description: "迁移目标:postgres(主库为 PG 时)、sqlite(主库为 SQLite 时)或 clickhouse"},
},
}
// SSLRenewHandler renews due TLS certificates.
type SSLRenewHandler struct{}
@@ -105,51 +109,6 @@ func (h *SSLRenewHandler) Execute(ctx context.Context, _ []byte) (*task.TaskResu
return &task.TaskResult{Message: msg}, nil
}
// DatabaseAutoCleanupHandler prunes observability data when auto-cleanup is enabled.
type DatabaseAutoCleanupHandler struct{}
// Execute runs retention-based cleanup for all observability targets.
func (h *DatabaseAutoCleanupHandler) Execute(ctx context.Context, _ []byte) (*task.TaskResult, error) {
// 从 SystemConfig 读取自动清理配置
enabled, _ := repository.GetBoolByKey(ctx, model.ConfigKeyDatabaseAutoCleanupEnabled)
if !enabled {
msg := "自动清理未启用,跳过执行"
task.AppendLog(ctx, "%s", msg)
return &task.TaskResult{Message: msg}, nil
}
retentionDays, _ := repository.GetIntByKey(ctx, model.ConfigKeyDatabaseAutoCleanupRetentionDays)
if retentionDays <= 0 {
retentionDays = 30 // 默认保留 30 天
}
task.AppendLog(ctx, "开始执行可观测数据自动清理,保留天数=%d", retentionDays)
summary, err := tasks.RunDatabaseAutoCleanupOnce(ctx, time.Now())
if err != nil {
task.AppendLog(ctx, "可观测数据自动清理失败: %v", err)
return nil, err
}
if summary == nil {
msg := "自动清理未启用,跳过执行"
task.AppendLog(ctx, "%s", msg)
return &task.TaskResult{Message: msg}, nil
}
var totalDeleted int64
for _, item := range summary.Results {
totalDeleted += item.DeletedCount
task.AppendLog(ctx, "清理 %s:删除 %d 条", item.TargetLabel, item.DeletedCount)
}
msg := fmt.Sprintf(
"可观测数据自动清理完成,保留 %d 天,共删除 %d 条",
summary.RetentionDays,
totalDeleted,
)
task.AppendLog(ctx, "%s", msg)
return &task.TaskResult{Message: msg}, nil
}
// WAFIPGroupSyncHandler syncs due WAF IP groups to agents.
type WAFIPGroupSyncHandler struct{}
@@ -6,7 +6,6 @@ package openflare
import (
"context"
"testing"
"time"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
@@ -17,63 +16,6 @@ import (
"gorm.io/gorm"
)
func TestDatabaseAutoCleanupHandlerSkipsWhenDisabled(t *testing.T) {
sqliteDB, err := gorm.Open(sqlite.Open(":memory:"), &gorm.Config{
DisableForeignKeyConstraintWhenMigrating: true,
})
require.NoError(t, err)
require.NoError(t, sqliteDB.AutoMigrate(&model.SystemConfig{}))
db.SetDB(sqliteDB)
t.Cleanup(func() { db.SetDB(nil) })
ctx := context.Background()
require.NoError(t, repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyDatabaseAutoCleanupEnabled, "false"))
result, err := (&DatabaseAutoCleanupHandler{}).Execute(ctx, nil)
require.NoError(t, err)
require.NotNil(t, result)
assert.Contains(t, result.Message, "未启用")
}
func TestDatabaseAutoCleanupHandlerDeletesRowsWhenEnabled(t *testing.T) {
sqliteDB, err := gorm.Open(sqlite.Open(":memory:"), &gorm.Config{
DisableForeignKeyConstraintWhenMigrating: true,
})
require.NoError(t, err)
require.NoError(t, sqliteDB.AutoMigrate(&model.SystemConfig{}))
db.SetDB(sqliteDB)
resetAccessLogStore := repository.SetAccessLogStoreForTest(repository.NewMemoryAccessLogStore())
resetObservabilityStore := repository.SetObservabilityStoreForTest(repository.NewMemoryObservabilityStore())
t.Cleanup(func() {
resetObservabilityStore()
resetAccessLogStore()
db.SetDB(nil)
})
ctx := context.Background()
now := time.Now().UTC()
require.NoError(t, repository.InsertOpenFlareAccessLogsBatch(ctx, []*model.OpenFlareAccessLog{{
NodeID: "node-a",
LoggedAt: now.Add(-95 * 24 * time.Hour),
RemoteAddr: "203.0.113.10",
Host: "example.com",
Path: "/access",
StatusCode: 200,
}}))
require.NoError(t, repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyDatabaseAutoCleanupEnabled, "true"))
require.NoError(t, repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyDatabaseAutoCleanupRetentionDays, "1"))
result, err := (&DatabaseAutoCleanupHandler{}).Execute(ctx, nil)
require.NoError(t, err)
require.NotNil(t, result)
assert.Contains(t, result.Message, "共删除")
rows, err := repository.ListOpenFlareAccessLogs(ctx, model.OpenFlareAccessLogQuery{Page: 0, PageSize: 10})
require.NoError(t, err)
assert.Empty(t, rows)
}
func TestUptimeKumaSyncHandlerSkipsWhenDisabled(t *testing.T) {
sqliteDB, err := gorm.Open(sqlite.Open(":memory:"), &gorm.Config{
DisableForeignKeyConstraintWhenMigrating: true,
+102 -20
View File
@@ -11,13 +11,10 @@ import (
"sync"
"time"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/infra/config"
"github.com/Rain-kl/Wavelet/internal/infra/persistence/batchwriter"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/internal/platform/lifecycle"
analyticsrepo "github.com/Rain-kl/Wavelet/internal/repository/analytics"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
"github.com/Rain-kl/Wavelet/pkg/logger"
)
@@ -56,12 +53,10 @@ var (
frpcDedup *dedupSet
)
// Init starts OpenFlare ClickHouse batch writers. Safe to call multiple times.
// Init starts OpenFlare log batch writers. Safe to call multiple times.
// Writers always initialize regardless of ClickHouse.enabled; the active log
// store is resolved via logstore at flush time (PG/SQLite when CH is not active).
func Init(ctx context.Context) {
if !config.Config.ClickHouse.Enabled {
return
}
initOnce.Do(func() {
metricSnapshotDedup = newDedupSet()
edgeHealthDedup = newDedupSet()
@@ -70,25 +65,25 @@ func Init(ctx context.Context) {
metricSnapshotWriter = mustNewObservabilityWriter(
"metric_snapshots",
withFlushRetries(analyticsrepo.BatchInsertNodeMetricSnapshots),
withFlushRetries(flushNodeMetricSnapshots),
metricSnapshotDedup,
metricSnapshotKey,
)
edgeHealthWriter = mustNewObservabilityWriter(
"edge_health",
withFlushRetries(analyticsrepo.BatchInsertNodeEdgeHealth),
withFlushRetries(flushNodeEdgeHealth),
edgeHealthDedup,
edgeHealthKey,
)
frpsWriter = mustNewObservabilityWriter(
"frps_obs",
withFlushRetries(analyticsrepo.BatchInsertNodeObsFrps),
withFlushRetries(flushNodeObsFrps),
frpsDedup,
frpsKey,
)
frpcWriter = mustNewObservabilityWriter(
"frpc_obs",
withFlushRetries(analyticsrepo.BatchInsertNodeObsFrpc),
withFlushRetries(flushNodeObsFrpc),
frpcDedup,
frpcKey,
)
@@ -129,6 +124,51 @@ func Stop(ctx context.Context) error {
return firstErr
}
// Drain 等待所有 OpenFlare 日志 writer 的在途批次落库:轮询队列 Depth 归零后
// 再保持一个最大 flush 周期(observabilityFlushEvery)持续为空才返回;
// 不停止 writer(迁移冻结后由 ensureWritable 拒绝新写入)。未初始化时直接返回 nil。
func Drain(ctx context.Context) error {
return drainWriters(ctx, WriterStats, observabilityFlushEvery)
}
// drainWriters 轮询 stats 直至所有队列 Depth=0 并持续 quietPeriod 无新积压。
func drainWriters(ctx context.Context, stats func() []batchwriter.Stats, quietPeriod time.Duration) error {
if !running() {
return nil
}
ticker := time.NewTicker(drainPollInterval)
defer ticker.Stop()
var quietSince time.Time
for {
if allDepthZero(stats()) {
if quietSince.IsZero() {
quietSince = time.Now()
} else if time.Since(quietSince) >= quietPeriod {
return nil
}
} else {
quietSince = time.Time{}
}
select {
case <-ctx.Done():
return ctx.Err()
case <-ticker.C:
}
}
}
// drainPollInterval 队列轮询间隔。
const drainPollInterval = 50 * time.Millisecond
func allDepthZero(stats []batchwriter.Stats) bool {
for _, s := range stats {
if s.Depth > 0 {
return false
}
}
return true
}
// WriterStats returns queue depth and failure counters for all OpenFlare writers.
func WriterStats() []batchwriter.Stats {
writers := []statsProvider{
@@ -237,7 +277,7 @@ func mustNewNodeAccessLogWriter() *batchwriter.Writer[analyticsmodel.NodeAccessL
}
writer, err := batchwriter.New[analyticsmodel.NodeAccessLog](
cfg,
withFlushRetries(analyticsrepo.BatchInsertNodeAccessLogs),
withFlushRetries(flushNodeAccessLogs),
batchwriter.WithDropHandler[analyticsmodel.NodeAccessLog](func(item analyticsmodel.NodeAccessLog) {
logger.WarnF(context.Background(), "[OpenFlare] node access log queue full, dropping log for node %s path %s", item.NodeID, item.Path)
}),
@@ -280,17 +320,59 @@ func withFlushRetries[T any](flush batchwriter.FlushFunc[T]) batchwriter.FlushFu
}
func wireModelInsertHooks() {
repository.SetObservabilityInsertHooks(repository.ObservabilityInsertHooks{
QueueMetricSnapshot: QueueMetricSnapshot,
QueueEdgeHealth: QueueEdgeHealth,
QueueFrpsObservation: QueueFrpsObservation,
QueueFrpcObservation: QueueFrpcObservation,
logstore.SetObservabilityHooks(logstore.ObservabilityHooks{
QueueMetricSnapshot: QueueMetricSnapshot,
QueueEdgeHealth: QueueEdgeHealth,
QueueNodeObsFrps: QueueFrpsObservation,
QueueNodeObsFrpc: QueueFrpcObservation,
})
repository.SetAccessLogInsertHooks(repository.AccessLogInsertHooks{
logstore.SetAccessLogHooks(logstore.AccessLogHooks{
QueueNodeAccessLogs: QueueNodeAccessLogs,
})
}
// 以下 flush 函数作为 batchwriter 的落库目标:激活库由 logstore 在 flush 时决定。
func flushNodeMetricSnapshots(ctx context.Context, rows []analyticsmodel.NodeMetricSnapshot) error {
s, err := logstore.Active(ctx)
if err != nil {
return err
}
return s.Observability.BatchInsertNodeMetricSnapshots(ctx, rows)
}
func flushNodeEdgeHealth(ctx context.Context, rows []analyticsmodel.NodeEdgeHealth) error {
s, err := logstore.Active(ctx)
if err != nil {
return err
}
return s.Observability.BatchInsertNodeEdgeHealth(ctx, rows)
}
func flushNodeObsFrps(ctx context.Context, rows []analyticsmodel.NodeObsFrps) error {
s, err := logstore.Active(ctx)
if err != nil {
return err
}
return s.Observability.BatchInsertNodeObsFrps(ctx, rows)
}
func flushNodeObsFrpc(ctx context.Context, rows []analyticsmodel.NodeObsFrpc) error {
s, err := logstore.Active(ctx)
if err != nil {
return err
}
return s.Observability.BatchInsertNodeObsFrpc(ctx, rows)
}
func flushNodeAccessLogs(ctx context.Context, rows []analyticsmodel.NodeAccessLog) error {
s, err := logstore.Active(ctx)
if err != nil {
return err
}
return s.AccessLogs.BatchInsertNodeAccessLogs(ctx, rows)
}
func metricSnapshotKey(snapshot analyticsmodel.NodeMetricSnapshot) string {
return fmt.Sprintf("%s|%d", snapshot.NodeID, snapshot.CapturedAt.UTC().UnixNano())
}
@@ -9,6 +9,7 @@ import (
"time"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/testhelper"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
@@ -26,11 +27,8 @@ func setupDashboardTestDB(t *testing.T) func() {
require.NoError(t, sqliteDB.AutoMigrate(&model.OpenFlareNode{}))
db.SetDB(sqliteDB)
resetAccessLogStore := repository.SetAccessLogStoreForTest(repository.NewMemoryAccessLogStore())
resetObservabilityStore := repository.SetObservabilityStoreForTest(repository.NewMemoryObservabilityStore())
testhelper.SetupLogStoresForTest(t)
return func() {
resetObservabilityStore()
resetAccessLogStore()
db.SetDB(nil)
}
}
@@ -40,15 +40,12 @@ func setupProtocolTestEnv(t *testing.T) (*gin.Engine, func()) {
db.SetDB(sqliteDB)
agent.ResetAuthCacheForTest()
resetAccessLogStore := repository.SetAccessLogStoreForTest(repository.NewMemoryAccessLogStore())
resetObservabilityStore := repository.SetObservabilityStoreForTest(repository.NewMemoryObservabilityStore())
testhelper.SetupLogStoresForTest(t)
engine := testhelper.NewTestGinEngine()
mountOpenFlareTestRoutes(engine)
cleanup := func() {
resetObservabilityStore()
resetAccessLogStore()
db.SetDB(nil)
agent.ResetAuthCacheForTest()
}
+2 -4
View File
@@ -15,6 +15,7 @@ import (
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/testhelper"
"github.com/glebarez/sqlite"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
@@ -41,12 +42,9 @@ func setupNodeTestDB(t *testing.T) func() {
))
db.SetDB(sqliteDB)
resetAccessLogStore := repository.SetAccessLogStoreForTest(repository.NewMemoryAccessLogStore())
resetObservabilityStore := repository.SetObservabilityStoreForTest(repository.NewMemoryObservabilityStore())
testhelper.SetupLogStoresForTest(t)
return func() {
resetObservabilityStore()
resetAccessLogStore()
db.SetDB(nil)
}
}
@@ -11,7 +11,7 @@ import (
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsrepo "github.com/Rain-kl/Wavelet/internal/repository/analytics"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/pkg/logger"
)
@@ -448,9 +448,9 @@ func buildAccessLogUADistributions(
for _, row := range uaRows {
ua := row.Key
count := row.Value
deviceAcc[analyticsrepo.ParseDeviceType(ua)] += count
browserAcc[analyticsrepo.ParseBrowserName(ua)] += count
osAcc[analyticsrepo.ParseOSName(ua)] += count
deviceAcc[analyticsmodel.ParseDeviceType(ua)] += count
browserAcc[analyticsmodel.ParseBrowserName(ua)] += count
osAcc[analyticsmodel.ParseOSName(ua)] += count
}
topUserAgents = make([]DistributionItem, 0, accessLogOverviewTopLimit)
-43
View File
@@ -10,7 +10,6 @@ import (
"strings"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/geoip"
oftasks "github.com/Rain-kl/Wavelet/internal/apps/openflare/tasks"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/uptimekuma"
"github.com/Rain-kl/Wavelet/internal/buildinfo"
"github.com/Rain-kl/Wavelet/internal/model"
@@ -50,22 +49,6 @@ type geoIPLookupView struct {
Longitude *float64 `json:"longitude,omitempty"`
}
type databaseCleanupInput struct {
Target string `json:"target"`
RetentionDays *int `json:"retention_days"`
}
type databaseCleanupResult struct {
Target string `json:"target"`
TargetLabel string `json:"target_label"`
DeletedCount int64 `json:"deleted_count"`
EligibleCount int64 `json:"eligible_count,omitempty"`
CleanupMode string `json:"cleanup_mode,omitempty"`
TableTTLDays int `json:"table_ttl_days,omitempty"`
DeleteAll bool `json:"delete_all"`
RetentionDays *int `json:"retention_days,omitempty"`
}
type optionBatchPayload struct {
Options []model.OpenFlareOption `json:"options"`
}
@@ -183,32 +166,6 @@ func lookupGeoIP(_ context.Context, provider, rawIP string) (*geoIPLookupView, e
}, nil
}
func cleanupDatabaseObservability(ctx context.Context, input databaseCleanupInput) (*databaseCleanupResult, error) {
target := strings.TrimSpace(input.Target)
if target == "" {
return nil, errors.New(errInvalidParams)
}
result, err := oftasks.CleanupDatabaseObservability(ctx, oftasks.DatabaseCleanupInput{
Target: target,
RetentionDays: input.RetentionDays,
})
if err != nil {
return nil, err
}
return &databaseCleanupResult{
Target: result.Target,
TargetLabel: result.TargetLabel,
DeletedCount: result.DeletedCount,
EligibleCount: result.EligibleCount,
CleanupMode: result.CleanupMode,
TableTTLDays: result.TableTTLDays,
DeleteAll: result.DeleteAll,
RetentionDays: result.RetentionDays,
}, nil
}
func syncUptimeKuma(ctx context.Context) error {
return uptimekuma.SyncToUptimeKuma(ctx)
}
@@ -6,7 +6,6 @@ package option
import (
"context"
"testing"
"time"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
@@ -117,55 +116,3 @@ func TestLookupGeoIPDisabledProvider(t *testing.T) {
assert.Equal(t, "disabled", view.Provider)
assert.Equal(t, "8.8.8.8", view.IP)
}
func TestCleanupDatabaseObservabilityDeletesRows(t *testing.T) {
cleanup := setupOptionTestDB(t)
defer cleanup()
ctx := context.Background()
resetAccessLogStore := repository.SetAccessLogStoreForTest(repository.NewMemoryAccessLogStore())
defer resetAccessLogStore()
now := time.Now().UTC()
require.NoError(t, repository.InsertOpenFlareAccessLogsBatch(ctx, []*model.OpenFlareAccessLog{
{
NodeID: "node-a",
LoggedAt: now.Add(-10 * 24 * time.Hour),
RemoteAddr: "203.0.113.1",
Host: "example.com",
Path: "/old",
StatusCode: 200,
},
{
NodeID: "node-a",
LoggedAt: now.Add(-2 * time.Hour),
RemoteAddr: "203.0.113.2",
Host: "example.com",
Path: "/recent",
StatusCode: 200,
},
}))
// Retention shorter than table TTL (90d for access logs) must be rejected.
shortRetention := 7
_, err := cleanupDatabaseObservability(ctx, databaseCleanupInput{
Target: "node_access_logs",
RetentionDays: &shortRetention,
})
require.Error(t, err)
// Full truncate still hard-deletes all rows.
result, err := cleanupDatabaseObservability(ctx, databaseCleanupInput{
Target: "node_access_logs",
})
require.NoError(t, err)
assert.Equal(t, "node_access_logs", result.Target)
assert.Equal(t, "访问日志", result.TargetLabel)
assert.Equal(t, int64(2), result.DeletedCount)
assert.True(t, result.DeleteAll)
assert.Equal(t, "truncate", result.CleanupMode)
rows, err := repository.ListOpenFlareAccessLogs(ctx, model.OpenFlareAccessLogQuery{Page: 0, PageSize: 10})
require.NoError(t, err)
assert.Empty(t, rows)
}
-37
View File
@@ -4,9 +4,6 @@
package option
import (
"encoding/json"
"errors"
"io"
"net/http"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/apiutil"
@@ -128,33 +125,6 @@ func LookupGeoIPHandler(c *gin.Context) {
c.JSON(http.StatusOK, response.OK(view))
}
// CleanupDatabaseHandler 清理可观测性数据库数据。
// @Summary 清理可观测性数据库
// @Description 按目标与保留天数清理可观测性相关数据表,需要管理员权限
// @Tags openflare-option
// @Accept json
// @Produce json
// @Security SessionCookie
// @Param request body option.databaseCleanupInput false "清理参数"
// @Success 200 {object} response.Any{data=option.databaseCleanupResult} "清理结果"
// @Failure 400 {object} response.Any "参数错误"
// @Failure 401 {object} response.Any "未登录"
// @Failure 404 {object} response.Any "无权限或不存在"
// @Failure 500 {object} response.Any "内部错误"
// @Router /api/v1/d/option/database/cleanup [post]
func CleanupDatabaseHandler(c *gin.Context) {
var input databaseCleanupInput
if err := bindOptionalJSON(c.Request.Body, &input); err != nil {
response.AbortBadRequest(c, errInvalidParams)
return
}
result, err := cleanupDatabaseObservability(c.Request.Context(), input)
if apiutil.AbortBadRequestOnError(c, err) {
return
}
c.JSON(http.StatusOK, response.OK(result))
}
// SyncUptimeKumaHandler 同步 Uptime Kuma 监控。
// @Summary 同步 Uptime Kuma
// @Description 将 OpenFlare 节点同步到 Uptime Kuma,需要管理员权限
@@ -174,10 +144,3 @@ func SyncUptimeKumaHandler(c *gin.Context) {
}
c.JSON(http.StatusOK, response.OK("同步成功"))
}
func bindOptionalJSON(body io.Reader, target any) error {
if err := json.NewDecoder(body).Decode(target); err != nil && !errors.Is(err, io.EOF) {
return err
}
return nil
}
+17 -5
View File
@@ -27,6 +27,17 @@ var (
const optionValueTrue = "true"
// protectedConfigKeyMessage 命中受保护 key 时返回给管理员的业务错误文案。
const protectedConfigKeyMessage = "该配置项由系统任务管理,禁止手动修改"
// protectedConfigKeys 仅允许内部(迁移任务/bootstrap)写入的 key。
var protectedConfigKeys = map[string]bool{
model.ConfigKeyLogDatabase: true,
model.ConfigKeyLogDBMigration: true,
}
func isProtectedConfigKey(key string) bool { return protectedConfigKeys[key] }
func buildOptionValidationState(ctx context.Context, options []model.OpenFlareOption) map[string]string {
// 从 SystemConfig 读取所有业务配置构建状态
configs, err := repository.ListAdminSystemConfigs(ctx, "business")
@@ -53,7 +64,7 @@ func validateOptionWithState(ctx context.Context, option model.OpenFlareOption,
if err := validateGeoIPOption(option.Key, option.Value); err != nil {
return err
}
if err := validateDatabaseCleanupOption(option.Key, option.Value); err != nil {
if err := validateLogRetentionOption(option.Key, option.Value); err != nil {
return err
}
if err := validateAgentOption(option.Key, option.Value); err != nil {
@@ -100,11 +111,9 @@ func validateGeoIPOption(key, value string) error {
return fmt.Errorf("%s 仅支持 disabled、mmdb、ip-api、geojs、ipinfo", key)
}
func validateDatabaseCleanupOption(key, value string) error {
func validateLogRetentionOption(key, value string) error {
switch key {
case model.ConfigKeyDatabaseAutoCleanupEnabled:
return validateBooleanOption(key, value)
case model.ConfigKeyDatabaseAutoCleanupRetentionDays:
case model.ConfigKeyLogRetentionDaysPostgres, model.ConfigKeyLogRetentionDaysSQLite, model.ConfigKeyLogRetentionDaysClickHouse:
intValue, err := strconv.Atoi(value)
if err != nil || intValue < 1 {
return fmt.Errorf("%s 必须为大于等于 1 的整数天", key)
@@ -221,6 +230,9 @@ func validateOptions(ctx context.Context, options []model.OpenFlareOption) error
if strings.TrimSpace(option.Key) == "" {
return errors.New(errInvalidParams)
}
if isProtectedConfigKey(option.Key) {
return errors.New(protectedConfigKeyMessage)
}
if err := validateOptionWithState(ctx, option, state); err != nil {
return err
}
+2 -2
View File
@@ -10,6 +10,7 @@ import (
"time"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/testhelper"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/agent"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
@@ -38,10 +39,9 @@ func setupRelayTestDB(t *testing.T) func() {
db.SetDB(sqliteDB)
agent.ResetAuthCacheForTest()
resetObservabilityStore := repository.SetObservabilityStoreForTest(repository.NewMemoryObservabilityStore())
testhelper.SetupLogStoresForTest(t)
return func() {
resetObservabilityStore()
db.SetDB(nil)
agent.ResetAuthCacheForTest()
}
@@ -1,245 +0,0 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package tasks
import (
"context"
"errors"
"fmt"
"strings"
"time"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/Rain-kl/Wavelet/internal/repository"
analyticsrepo "github.com/Rain-kl/Wavelet/internal/repository/analytics"
)
const (
// DatabaseCleanupTargetAccessLogs is the API cleanup target for access logs.
DatabaseCleanupTargetAccessLogs = "node_access_logs"
// DatabaseCleanupTargetMetricSnapshots is the API cleanup target for metric snapshots.
DatabaseCleanupTargetMetricSnapshots = "node_metric_snapshots"
// DatabaseCleanupTargetEdgeHealth is the API cleanup target for OpenResty edge health (connections).
DatabaseCleanupTargetEdgeHealth = "node_edge_health"
// DatabaseCleanupTargetObsFrps is the API cleanup target for FRPS observations.
DatabaseCleanupTargetObsFrps = "node_obs_frps"
// DatabaseCleanupTargetObsFrpc is the API cleanup target for FRPC observations.
DatabaseCleanupTargetObsFrpc = "node_obs_frpc"
)
var databaseCleanupTargets = map[string]string{
DatabaseCleanupTargetAccessLogs: "访问日志",
DatabaseCleanupTargetMetricSnapshots: "性能快照",
DatabaseCleanupTargetEdgeHealth: "OpenResty 健康(连接)",
DatabaseCleanupTargetObsFrps: "FRPS 观测",
DatabaseCleanupTargetObsFrpc: "FRPC 观测",
}
// databaseCleanupTableTTLDays maps API targets to ClickHouse DDL TTL days.
var databaseCleanupTableTTLDays = map[string]int{
DatabaseCleanupTargetAccessLogs: analyticsrepo.TableTTLDaysNodeAccessLogs,
DatabaseCleanupTargetMetricSnapshots: analyticsrepo.TableTTLDaysNodeMetricSnapshots,
DatabaseCleanupTargetEdgeHealth: analyticsrepo.TableTTLDaysNodeObs,
DatabaseCleanupTargetObsFrps: analyticsrepo.TableTTLDaysNodeObs,
DatabaseCleanupTargetObsFrpc: analyticsrepo.TableTTLDaysNodeObs,
}
// DatabaseCleanupInput describes a manual observability cleanup request.
type DatabaseCleanupInput struct {
Target string `json:"target"`
RetentionDays *int `json:"retention_days"`
}
// DatabaseCleanupResult summarizes a manual observability cleanup run.
//
// Semantics:
// - delete_all / cleanup_mode=truncate: DeletedCount is hard-deleted rows (TRUNCATE).
// - retention path / cleanup_mode=ttl_materialize: DeletedCount is always 0;
// EligibleCount estimates rows past the table DDL TTL (not an arbitrary younger cutoff).
type DatabaseCleanupResult struct {
Target string `json:"target"`
TargetLabel string `json:"target_label"`
DeletedCount int64 `json:"deleted_count"`
EligibleCount int64 `json:"eligible_count,omitempty"`
CleanupMode string `json:"cleanup_mode,omitempty"`
TableTTLDays int `json:"table_ttl_days,omitempty"`
DeleteAll bool `json:"delete_all"`
RetentionDays *int `json:"retention_days,omitempty"`
Cutoff *time.Time `json:"cutoff,omitempty"`
}
// DatabaseAutoCleanupSummary summarizes a scheduled auto-cleanup run.
type DatabaseAutoCleanupSummary struct {
RetentionDays int `json:"retention_days"`
ExecutedAt time.Time `json:"executed_at"`
Results []DatabaseCleanupResult `json:"results"`
}
// TableTTLDaysForCleanupTarget returns the DDL TTL days for a cleanup target.
func TableTTLDaysForCleanupTarget(target string) (int, bool) {
days, ok := databaseCleanupTableTTLDays[strings.TrimSpace(target)]
return days, ok
}
// CleanupDatabaseObservability deletes observability rows for the given target.
//
// When RetentionDays is nil, rows are hard-deleted via TRUNCATE.
// When RetentionDays is set, ClickHouse only force-materializes the table TTL policy:
// retention_days shorter than the table TTL is rejected (do not fake success).
func CleanupDatabaseObservability(ctx context.Context, input DatabaseCleanupInput) (*DatabaseCleanupResult, error) {
target := strings.TrimSpace(input.Target)
targetLabel, ok := databaseCleanupTargets[target]
if !ok {
return nil, errors.New("unsupported cleanup target")
}
if input.RetentionDays != nil && *input.RetentionDays <= 0 {
return nil, errors.New("retention_days 必须为大于 0 的整数")
}
tableTTLDays := databaseCleanupTableTTLDays[target]
result := &DatabaseCleanupResult{
Target: target,
TargetLabel: targetLabel,
DeleteAll: input.RetentionDays == nil,
TableTTLDays: tableTTLDays,
}
if input.RetentionDays == nil {
deleted, mode, err := deleteAllObservabilityRows(ctx, target)
if err != nil {
return nil, err
}
result.DeletedCount = deleted
result.EligibleCount = deleted
result.CleanupMode = mode
return result, nil
}
retentionDays := *input.RetentionDays
if retentionDays < tableTTLDays {
return nil, fmt.Errorf(
"retention_days 不能小于表 TTL(%d 天);ClickHouse 仅支持按表 TTL 物化过期,更短保留请使用清空全部或调整 DDL",
tableTTLDays,
)
}
// MATERIALIZE TTL only enforces DDL policy; cutoff reported is the table TTL boundary.
tableCutoff := time.Now().UTC().Add(-time.Duration(tableTTLDays) * 24 * time.Hour)
eligible, mode, err := materializeObservabilityTableTTL(ctx, target)
if err != nil {
return nil, err
}
result.DeletedCount = 0
result.EligibleCount = eligible
result.CleanupMode = mode
result.RetentionDays = &retentionDays
result.Cutoff = &tableCutoff
return result, nil
}
// RunDatabaseAutoCleanupOnce runs retention-based cleanup for all observability targets.
//
// Configured retention shorter than a target's table TTL is clamped up to the table TTL
// so the scheduled job can force-materialize each table policy without failing.
func RunDatabaseAutoCleanupOnce(ctx context.Context, now time.Time) (*DatabaseAutoCleanupSummary, error) {
enabled, err := repository.GetBoolByKey(ctx, model.ConfigKeyDatabaseAutoCleanupEnabled)
if err != nil {
return nil, fmt.Errorf("failed to read database_auto_cleanup_enabled: %w", err)
}
if !enabled {
return nil, nil
}
retentionDays, err := repository.GetIntByKey(ctx, model.ConfigKeyDatabaseAutoCleanupRetentionDays)
if err != nil || retentionDays <= 0 {
// Use default value 30 if config read fails or value is invalid
retentionDays = 30
}
results := make([]DatabaseCleanupResult, 0, len(databaseCleanupTargets))
for _, target := range []string{
DatabaseCleanupTargetAccessLogs,
DatabaseCleanupTargetMetricSnapshots,
DatabaseCleanupTargetEdgeHealth,
DatabaseCleanupTargetObsFrps,
DatabaseCleanupTargetObsFrpc,
} {
effectiveDays := retentionDays
if ttl, ok := databaseCleanupTableTTLDays[target]; ok && effectiveDays < ttl {
effectiveDays = ttl
}
result, err := CleanupDatabaseObservability(ctx, DatabaseCleanupInput{
Target: target,
RetentionDays: &effectiveDays,
})
if err != nil {
return nil, err
}
results = append(results, *result)
}
return &DatabaseAutoCleanupSummary{
RetentionDays: retentionDays,
ExecutedAt: now.UTC(),
Results: results,
}, nil
}
func deleteAllObservabilityRows(ctx context.Context, target string) (int64, string, error) {
var (
deleted int64
err error
)
switch target {
case DatabaseCleanupTargetAccessLogs:
deleted, err = repository.DeleteAllOpenFlareAccessLogs(ctx)
case DatabaseCleanupTargetMetricSnapshots:
deleted, err = repository.DeleteAllOpenFlareMetricSnapshots(ctx)
case DatabaseCleanupTargetEdgeHealth:
deleted, err = repository.DeleteAllOpenFlareEdgeHealth(ctx)
case DatabaseCleanupTargetObsFrps:
deleted, err = repository.DeleteAllOpenFlareNodeObservationFrps(ctx)
case DatabaseCleanupTargetObsFrpc:
deleted, err = repository.DeleteAllOpenFlareNodeObservationFrpc(ctx)
default:
return 0, "", errors.New("unsupported cleanup target")
}
if err != nil {
return 0, "", err
}
return deleted, analyticsrepo.CleanupModeTruncate, nil
}
// materializeObservabilityTableTTL triggers table-TTL materialize (or memory-store delete-before
// with the table TTL cutoff for tests) and returns the eligible/estimate row count.
func materializeObservabilityTableTTL(ctx context.Context, target string) (int64, string, error) {
ttlDays, ok := databaseCleanupTableTTLDays[target]
if !ok {
return 0, "", errors.New("unsupported cleanup target")
}
cutoff := time.Now().UTC().Add(-time.Duration(ttlDays) * 24 * time.Hour)
var (
eligible int64
err error
)
switch target {
case DatabaseCleanupTargetAccessLogs:
eligible, err = repository.DeleteOpenFlareAccessLogsBefore(ctx, cutoff)
case DatabaseCleanupTargetMetricSnapshots:
eligible, err = repository.DeleteOpenFlareMetricSnapshotsBefore(ctx, cutoff)
case DatabaseCleanupTargetEdgeHealth:
eligible, err = repository.DeleteOpenFlareEdgeHealthBefore(ctx, cutoff)
case DatabaseCleanupTargetObsFrps:
eligible, err = repository.DeleteOpenFlareNodeObservationFrpsBefore(ctx, cutoff)
case DatabaseCleanupTargetObsFrpc:
eligible, err = repository.DeleteOpenFlareNodeObservationFrpcBefore(ctx, cutoff)
default:
return 0, "", errors.New("unsupported cleanup target")
}
if err != nil {
return 0, "", err
}
return eligible, analyticsrepo.CleanupModeTTLMaterialize, nil
}
@@ -1,208 +0,0 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package tasks
import (
"context"
"testing"
"time"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/Rain-kl/Wavelet/internal/repository"
analyticsrepo "github.com/Rain-kl/Wavelet/internal/repository/analytics"
"github.com/glebarez/sqlite"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"gorm.io/gorm"
)
func setupDatabaseCleanupTestDB(t *testing.T) context.Context {
t.Helper()
sqliteDB, err := gorm.Open(sqlite.Open(":memory:"), &gorm.Config{
DisableForeignKeyConstraintWhenMigrating: true,
})
require.NoError(t, err)
require.NoError(t, sqliteDB.AutoMigrate(&model.SystemConfig{}))
db.SetDB(sqliteDB)
resetAccessLogStore := repository.SetAccessLogStoreForTest(repository.NewMemoryAccessLogStore())
resetObservabilityStore := repository.SetObservabilityStoreForTest(repository.NewMemoryObservabilityStore())
t.Cleanup(func() {
resetObservabilityStore()
resetAccessLogStore()
db.SetDB(nil)
})
return context.Background()
}
func TestCleanupDatabaseObservabilityRejectsRetentionShorterThanTableTTL(t *testing.T) {
ctx := setupDatabaseCleanupTestDB(t)
retentionDays := 7 // metric snapshots DDL TTL is 30 days
result, err := CleanupDatabaseObservability(ctx, DatabaseCleanupInput{
Target: DatabaseCleanupTargetMetricSnapshots,
RetentionDays: &retentionDays,
})
require.Error(t, err)
assert.Nil(t, result)
assert.Contains(t, err.Error(), "不能小于表 TTL")
assert.Contains(t, err.Error(), "30")
}
func TestCleanupDatabaseObservabilityRejectsAccessLogRetentionShorterThanTableTTL(t *testing.T) {
ctx := setupDatabaseCleanupTestDB(t)
retentionDays := 30 // access logs DDL TTL is 90 days
result, err := CleanupDatabaseObservability(ctx, DatabaseCleanupInput{
Target: DatabaseCleanupTargetAccessLogs,
RetentionDays: &retentionDays,
})
require.Error(t, err)
assert.Nil(t, result)
assert.Contains(t, err.Error(), "90")
}
func TestCleanupDatabaseObservabilityMaterializeDoesNotClaimHardDelete(t *testing.T) {
ctx := setupDatabaseCleanupTestDB(t)
now := time.Now().UTC()
// One row past metric table TTL (30d), one still inside the window.
require.NoError(t, repository.InsertOpenFlareMetricSnapshot(ctx, &model.OpenFlareMetricSnapshot{
NodeID: "node-a",
CapturedAt: now.Add(-40 * 24 * time.Hour),
CPUUsagePercent: 10,
}))
require.NoError(t, repository.InsertOpenFlareMetricSnapshot(ctx, &model.OpenFlareMetricSnapshot{
NodeID: "node-a",
CapturedAt: now.Add(-12 * time.Hour),
CPUUsagePercent: 20,
}))
retentionDays := analyticsrepo.TableTTLDaysNodeMetricSnapshots
result, err := CleanupDatabaseObservability(ctx, DatabaseCleanupInput{
Target: DatabaseCleanupTargetMetricSnapshots,
RetentionDays: &retentionDays,
})
require.NoError(t, err)
assert.False(t, result.DeleteAll)
assert.Equal(t, analyticsrepo.CleanupModeTTLMaterialize, result.CleanupMode)
assert.Equal(t, analyticsrepo.TableTTLDaysNodeMetricSnapshots, result.TableTTLDays)
// MATERIALIZE is not a counted hard delete.
assert.Equal(t, int64(0), result.DeletedCount)
assert.Equal(t, int64(1), result.EligibleCount)
require.NotNil(t, result.Cutoff)
assert.True(t, result.Cutoff.Before(now.Add(-29*24*time.Hour)))
// Memory store applies the table-TTL cutoff for tests; only the recent row remains.
rows, err := repository.ListOpenFlareMetricSnapshotsSince(ctx, "", time.Time{}, 0)
require.NoError(t, err)
require.Len(t, rows, 1)
assert.Equal(t, float64(20), rows[0].CPUUsagePercent)
}
func TestCleanupDatabaseObservabilityDeletesAllRowsWhenRetentionMissing(t *testing.T) {
ctx := setupDatabaseCleanupTestDB(t)
now := time.Now().UTC()
require.NoError(t, repository.InsertOpenFlareAccessLogsBatch(ctx, []*model.OpenFlareAccessLog{
{
NodeID: "node-a",
LoggedAt: now.Add(-3 * time.Hour),
RemoteAddr: "203.0.113.1",
Host: "example.com",
Path: "/one",
StatusCode: 200,
},
{
NodeID: "node-a",
LoggedAt: now.Add(-2 * time.Hour),
RemoteAddr: "203.0.113.2",
Host: "example.com",
Path: "/two",
StatusCode: 502,
},
}))
result, err := CleanupDatabaseObservability(ctx, DatabaseCleanupInput{
Target: DatabaseCleanupTargetAccessLogs,
})
require.NoError(t, err)
assert.True(t, result.DeleteAll)
assert.Equal(t, analyticsrepo.CleanupModeTruncate, result.CleanupMode)
assert.Equal(t, int64(2), result.DeletedCount)
assert.Equal(t, int64(2), result.EligibleCount)
rows, err := repository.ListOpenFlareAccessLogs(ctx, model.OpenFlareAccessLogQuery{Page: 0, PageSize: 10})
require.NoError(t, err)
assert.Empty(t, rows)
}
func TestRunDatabaseAutoCleanupOnceClampsRetentionToTableTTL(t *testing.T) {
ctx := setupDatabaseCleanupTestDB(t)
now := time.Now().UTC()
// Access logs TTL=90d, metrics TTL=30d. Config retention=1 must clamp, not reject.
require.NoError(t, repository.InsertOpenFlareAccessLogsBatch(ctx, []*model.OpenFlareAccessLog{{
NodeID: "node-a",
LoggedAt: now.Add(-100 * 24 * time.Hour),
RemoteAddr: "203.0.113.10",
Host: "example.com",
Path: "/access",
StatusCode: 200,
}}))
require.NoError(t, repository.InsertOpenFlareMetricSnapshot(ctx, &model.OpenFlareMetricSnapshot{
NodeID: "node-a",
CapturedAt: now.Add(-40 * 24 * time.Hour),
CPUUsagePercent: 10,
}))
require.NoError(t, repository.InsertOpenFlareEdgeHealth(ctx, &model.OpenFlareEdgeHealth{
NodeID: "node-a",
CapturedAt: now.Add(-40 * 24 * time.Hour),
Status: "healthy",
Connections: 2,
}))
require.NoError(t, repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyDatabaseAutoCleanupEnabled, "true"))
require.NoError(t, repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyDatabaseAutoCleanupRetentionDays, "1"))
summary, err := RunDatabaseAutoCleanupOnce(ctx, now)
require.NoError(t, err)
require.NotNil(t, summary)
require.Len(t, summary.Results, 5)
assert.Equal(t, 1, summary.RetentionDays)
for _, result := range summary.Results {
assert.Equal(t, analyticsrepo.CleanupModeTTLMaterialize, result.CleanupMode)
assert.Equal(t, int64(0), result.DeletedCount, "target %s must not claim hard delete", result.Target)
assert.GreaterOrEqual(t, result.TableTTLDays, 30)
require.NotNil(t, result.RetentionDays)
assert.GreaterOrEqual(t, *result.RetentionDays, result.TableTTLDays)
}
accessLogs, err := repository.ListOpenFlareAccessLogs(ctx, model.OpenFlareAccessLogQuery{Page: 0, PageSize: 10})
require.NoError(t, err)
assert.Empty(t, accessLogs)
metricSnapshots, err := repository.ListOpenFlareMetricSnapshotsSince(ctx, "", time.Time{}, 0)
require.NoError(t, err)
assert.Empty(t, metricSnapshots)
edgeHealth, err := repository.ListOpenFlareEdgeHealth(ctx, "", time.Time{}, 0)
require.NoError(t, err)
assert.Empty(t, edgeHealth)
}
func TestTableTTLDaysForCleanupTarget(t *testing.T) {
days, ok := TableTTLDaysForCleanupTarget(DatabaseCleanupTargetAccessLogs)
require.True(t, ok)
assert.Equal(t, 90, days)
days, ok = TableTTLDaysForCleanupTarget(DatabaseCleanupTargetMetricSnapshots)
require.True(t, ok)
assert.Equal(t, 30, days)
_, ok = TableTTLDaysForCleanupTarget("unknown")
assert.False(t, ok)
}
@@ -0,0 +1,398 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package tasks
import (
"context"
"encoding/json"
"errors"
"fmt"
"time"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/chwriter"
"github.com/Rain-kl/Wavelet/internal/infra/config"
"github.com/Rain-kl/Wavelet/internal/infra/task"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
"github.com/Rain-kl/Wavelet/pkg/logger"
)
const copyBatchSize = 1000
// 迁移目标库名常量(normalizeTarget 归一化后的取值)。
const (
targetPostgres = "postgres"
targetSQLite = "sqlite"
targetClickHouse = "clickhouse"
)
type logDBSwitchPayload struct {
Target string `json:"target"`
}
// LogDBSwitchHandler 切换日志数据库任务处理器。
type LogDBSwitchHandler struct{}
// ValidatePayload 校验并规范化参数。
func (h *LogDBSwitchHandler) ValidatePayload(payload []byte) ([]byte, error) {
var p logDBSwitchPayload
if err := json.Unmarshal(payload, &p); err != nil {
return nil, fmt.Errorf("参数解析失败: %w", err)
}
p.Target = normalizeTarget(p.Target)
if !validTarget(p.Target) {
return nil, fmt.Errorf("目标日志库不合法: %s", p.Target)
}
out, err := json.Marshal(p)
if err != nil {
return nil, err
}
return out, nil
}
func normalizeTarget(v string) string {
switch v {
case targetPostgres, "postgresql":
return targetPostgres
case targetSQLite, "sqlite3":
return targetSQLite
case targetClickHouse, "ch":
return targetClickHouse
}
return v
}
func validTarget(v string) bool {
return v == targetPostgres || v == targetSQLite || v == targetClickHouse
}
// Execute 执行迁移。
func (h *LogDBSwitchHandler) Execute(ctx context.Context, payload []byte) (*task.TaskResult, error) {
var p logDBSwitchPayload
if err := json.Unmarshal(payload, &p); err != nil {
return nil, fmt.Errorf("参数解析失败: %w", err)
}
p.Target = normalizeTarget(p.Target)
if err := validateSwitch(ctx, p.Target); err != nil {
return nil, err
}
source, err := currentLogDatabase(ctx)
if err != nil {
task.AppendLog(ctx, "读取日志主库失败: %v", err)
return nil, err
}
task.AppendLog(ctx, "开始切换日志数据库:%s -> %s", source, p.Target)
// 设置迁移冻结标记(置位后由 ensureWritable 拒绝新写入)。
if err := setMigrationFlag(ctx, "migrating"); err != nil {
return nil, err
}
// 失败也清除(SaveOrUpdateSystemConfig 会失效 RAM 缓存并广播),保持源库可写。
defer func() {
if err := setMigrationFlag(ctx, ""); err != nil {
logger.ErrorF(ctx, "清除日志迁移冻结标记失败: %v", err)
}
}()
// 冻结标记置位后再排空在途批次(chwriter + 用户访问日志 writer),
// 保证排空完成后不再有新批次进入源库。
if err := drainLogWriters(ctx); err != nil {
return nil, fmt.Errorf("排空日志写入队列失败: %w", err)
}
src, err := logstore.Active(ctx)
if err != nil {
return nil, err
}
dst, err := buildTargetStore(ctx, p.Target)
if err != nil {
return nil, err
}
// 清空目标库日志表(幂等重试前提)。
if err := clearTargetLogTables(ctx, dst); err != nil {
return nil, err
}
// PG 目标:按源库时间范围预建分区,避免历史数据复制报 "no partition of relation found"。
if err := ensureTargetPartitions(ctx, src, dst, p.Target); err != nil {
return nil, err
}
// 逐表复制(6 张日志表)。
if err := copyAccessLogs(ctx, src, dst); err != nil {
return nil, err
}
if err := copyUserAccessLogs(ctx, src, dst); err != nil {
return nil, err
}
if err := copyObservability(ctx, src, dst); err != nil {
return nil, err
}
// 翻转主库标记。
if err := flipLogDatabase(ctx, p.Target); err != nil {
return nil, err
}
task.AppendLog(ctx, "日志数据库已切换为 %s,写入恢复", p.Target)
return &task.TaskResult{Message: fmt.Sprintf("日志数据库已从 %s 切换为 %s", source, p.Target)}, nil
}
func validateSwitch(ctx context.Context, target string) error {
source, err := currentLogDatabase(ctx)
if err != nil {
return err
}
if source == target {
return errors.New("目标日志库与当前日志库相同,无需迁移")
}
switch target {
case "clickhouse":
if !config.Config.ClickHouse.Enabled {
return errors.New("ClickHouse 未启用,无法迁移到 ClickHouse")
}
case "postgres":
if !config.Config.Database.Enabled {
return errors.New("PostgreSQL 未启用(当前主库为 SQLite),无法迁移到 PostgreSQL")
}
case "sqlite":
if config.Config.Database.Enabled {
return errors.New("当前主库为 PostgreSQL,日志库不能设置为 SQLite")
}
}
return nil
}
func currentLogDatabase(ctx context.Context) (string, error) {
cfg, err := repository.GetSystemConfigByKey(ctx, model.ConfigKeyLogDatabase)
if err != nil {
return "", fmt.Errorf("读取日志主库失败: %w", err)
}
if cfg.Value == "" {
return "", errors.New("日志主库配置为空")
}
return cfg.Value, nil
}
// drainLogWriters 等待 chwriter(节点访问日志 + 可观测 4 表)的在途批次全部落库。
// 见设计 §7.2:先排空再冻结。用户访问日志(w_user_access_logs)记录已禁用,无在途批次。
func drainLogWriters(ctx context.Context) error {
return chwriter.Drain(ctx)
}
// setMigrationFlag 写入迁移冻结标记。用 SaveOrUpdateSystemConfig:行缺失时 upsert,
// 并失效 RAM 缓存 + 广播其他节点,保证 logstore.Migrating/resolveDatabase 立即生效。
func setMigrationFlag(ctx context.Context, v string) error {
return repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyLogDBMigration, v)
}
// flipLogDatabase 翻转日志主库。同上用 SaveOrUpdateSystemConfig,确保各进程缓存失效后指向新库。
func flipLogDatabase(ctx context.Context, target string) error {
return repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyLogDatabase, target)
}
// buildTargetStore 构造目标库 Store(不经过 Active 缓存,直接 Build)。
// 迁移期间冻结标记已置位,目标库的清空/复制写入必须放行,故使用 BuildForMigration。
func buildTargetStore(ctx context.Context, database string) (*logstore.Store, error) {
return logstore.BuildForMigration(ctx, database)
}
func clearTargetLogTables(ctx context.Context, dst *logstore.Store) error {
// 依次清空 6 张表:AccessLogs.DeleteAll、UserAccessLogs.DeleteAll、Observability.DeleteAll*
// (SQLite/PG 用 DeleteAll;CH 用 TRUNCATE 语义)。
if _, err := dst.AccessLogs.DeleteAll(ctx); err != nil {
return fmt.Errorf("清空目标访问日志失败: %w", err)
}
if _, err := dst.UserAccessLogs.DeleteAll(ctx); err != nil {
return fmt.Errorf("清空目标用户访问日志失败: %w", err)
}
for _, fn := range []func(context.Context) (int64, error){
dst.Observability.DeleteAllMetricSnapshots,
dst.Observability.DeleteAllEdgeHealth,
dst.Observability.DeleteAllNodeObservationFrps,
dst.Observability.DeleteAllNodeObservationFrpc,
} {
if _, err := fn(ctx); err != nil {
return err
}
}
return nil
}
// ensureTargetPartitions 目标为 PG 时,按源库时间范围(两表合并)预建分区,
// 否则复制历史数据会报 "no partition of relation found";目标非 PG 为 no-op。
func ensureTargetPartitions(ctx context.Context, src, dst *logstore.Store, target string) error {
if target != targetPostgres {
return nil
}
from, to, err := migrationRange(ctx, src)
if err != nil {
return err
}
if from.IsZero() || to.IsZero() {
task.AppendLog(ctx, "源库无日志数据,跳过分区预建")
return nil
}
if err := dst.AccessLogs.EnsurePartitions(ctx, from, to.AddDate(0, 1, 0)); err != nil {
return fmt.Errorf("预建目标 PG 分区失败: %w", err)
}
task.AppendLog(ctx, "已为目标 PG 预建分区 %s ~ %s", from.Format("2006-01"), to.Format("2006-01"))
return nil
}
// migrationRange 合并源库节点访问日志(logged_at)与用户访问日志(created_at)
// 的最小/最大时间;任一表为空时忽略该表。
func migrationRange(ctx context.Context, src *logstore.Store) (time.Time, time.Time, error) {
fromAccess, toAccess, err := src.AccessLogs.MigrationRange(ctx)
if err != nil {
return time.Time{}, time.Time{}, fmt.Errorf("读取源访问日志时间范围失败: %w", err)
}
fromUser, toUser, err := src.UserAccessLogs.MigrationRange(ctx)
if err != nil {
return time.Time{}, time.Time{}, fmt.Errorf("读取源用户访问日志时间范围失败: %w", err)
}
return minTime(fromAccess, fromUser), maxTime(toAccess, toUser), nil
}
func minTime(a, b time.Time) time.Time {
switch {
case a.IsZero():
return b
case b.IsZero():
return a
case a.Before(b):
return a
default:
return b
}
}
func maxTime(a, b time.Time) time.Time {
switch {
case a.IsZero():
return b
case b.IsZero():
return a
case a.After(b):
return a
default:
return b
}
}
// copyAccessLogs 从 src 复制节点访问日志到 dst。
func copyAccessLogs(ctx context.Context, src, dst *logstore.Store) error {
// 注意:迁移期间 src 已冻结,但复制读取不受冻结影响;每批按 id 升序扫描。
var lastID uint64
for {
rows, err := listNodeAccessLogsByID(ctx, src, lastID, copyBatchSize)
if err != nil {
return err
}
if len(rows) == 0 {
break
}
if err := dst.AccessLogs.BatchInsertNodeAccessLogs(ctx, rows); err != nil {
return fmt.Errorf("写入目标访问日志失败(批 %d): %w", lastID, err)
}
task.AppendLog(ctx, "已复制访问日志 %d 条(截至 id=%d)", len(rows), rows[len(rows)-1].ID)
lastID = rows[len(rows)-1].ID
if len(rows) < copyBatchSize {
break
}
}
return nil
}
func listNodeAccessLogsByID(ctx context.Context, src *logstore.Store, afterID uint64, limit int) ([]analyticsmodel.NodeAccessLog, error) {
return src.AccessLogs.ListForMigration(ctx, afterID, limit)
}
// copyUserAccessLogs 从 src 复制用户访问日志到 dst(按 id 升序分批)。
func copyUserAccessLogs(ctx context.Context, src, dst *logstore.Store) error {
var lastID uint64
for {
rows, err := src.UserAccessLogs.ListForMigration(ctx, lastID, copyBatchSize)
if err != nil {
return err
}
if len(rows) == 0 {
return nil
}
if err := dst.UserAccessLogs.BatchInsert(ctx, rows); err != nil {
return fmt.Errorf("写入目标用户访问日志失败(批 %d): %w", lastID, err)
}
lastID = rows[len(rows)-1].ID
task.AppendLog(ctx, "已复制用户访问日志 %d 条(截至 id=%d)", len(rows), lastID)
if len(rows) < copyBatchSize {
return nil
}
}
}
// copyObservability 复制 4 张可观测表,每张表按 id 升序分批复制,
// 以每批最后一条 id 作为下一批游标(不使用 len 近似)。
func copyObservability(ctx context.Context, src, dst *logstore.Store) error {
if err := copyObsTable(ctx, "metric_snapshots",
src.Observability.ListMetricSnapshotsForMigration,
dst.Observability.BatchInsertNodeMetricSnapshots,
lastMetricSnapshotID); err != nil {
return err
}
if err := copyObsTable(ctx, "edge_health",
src.Observability.ListEdgeHealthForMigration,
dst.Observability.BatchInsertNodeEdgeHealth,
lastEdgeHealthID); err != nil {
return err
}
if err := copyObsTable(ctx, "obs_frps",
src.Observability.ListNodeObsFrpsForMigration,
dst.Observability.BatchInsertNodeObsFrps,
lastObsFrpsID); err != nil {
return err
}
if err := copyObsTable(ctx, "obs_frpc",
src.Observability.ListNodeObsFrpcForMigration,
dst.Observability.BatchInsertNodeObsFrpc,
lastObsFrpcID); err != nil {
return err
}
return nil
}
// copyObsTable 按 id 升序分批复制单张可观测表;idOf 返回批内最后一条 id。
func copyObsTable[T any](ctx context.Context, name string,
list func(context.Context, uint64, int) ([]T, error),
insert func(context.Context, []T) error,
idOf func([]T) uint64,
) error {
var lastID uint64
for {
rows, err := list(ctx, lastID, copyBatchSize)
if err != nil {
return fmt.Errorf("复制 %s 失败: %w", name, err)
}
if len(rows) == 0 {
return nil
}
if err := insert(ctx, rows); err != nil {
return fmt.Errorf("复制 %s 失败: %w", name, err)
}
lastID = idOf(rows)
task.AppendLog(ctx, "已复制 %s %d 条(截至 id=%d)", name, len(rows), lastID)
if len(rows) < copyBatchSize {
return nil
}
}
}
func lastMetricSnapshotID(rows []analyticsmodel.NodeMetricSnapshot) uint64 {
return rows[len(rows)-1].ID
}
func lastEdgeHealthID(rows []analyticsmodel.NodeEdgeHealth) uint64 { return rows[len(rows)-1].ID }
func lastObsFrpsID(rows []analyticsmodel.NodeObsFrps) uint64 { return rows[len(rows)-1].ID }
func lastObsFrpcID(rows []analyticsmodel.NodeObsFrpc) uint64 { return rows[len(rows)-1].ID }
@@ -0,0 +1,397 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package tasks
import (
"context"
"encoding/json"
"errors"
"fmt"
"sync/atomic"
"testing"
"time"
"github.com/glebarez/sqlite"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"gorm.io/gorm"
"gorm.io/gorm/logger"
"github.com/Rain-kl/Wavelet/internal/infra/config"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
)
var logDBSwitchDBSeq int64
// newLogDBSwitchDB 构造内存 sqlite 库(含日志 5 表 + 系统配置表)。
func newLogDBSwitchDB(t *testing.T) *gorm.DB {
t.Helper()
dsn := fmt.Sprintf("file:log-db-switch-%d?mode=memory&cache=shared", atomic.AddInt64(&logDBSwitchDBSeq, 1))
gdb, err := gorm.Open(sqlite.Open(dsn), &gorm.Config{Logger: logger.Default.LogMode(logger.Silent)})
require.NoError(t, err)
require.NoError(t, gdb.AutoMigrate(
&model.SystemConfig{},
&analyticsmodel.NodeAccessLog{},
&analyticsmodel.NodeMetricSnapshot{},
&analyticsmodel.NodeEdgeHealth{},
&analyticsmodel.NodeObsFrps{},
&analyticsmodel.NodeObsFrpc{},
&analyticsmodel.UserAccessLog{},
))
return gdb
}
// TestCopyAccessLogsPreservesIDs sqlite→sqlite 模拟:源 store 3 条,目标空库,
// copyAccessLogs 后 ID 保留、数量一致。
func TestCopyAccessLogsPreservesIDs(t *testing.T) {
oldDB, oldCH := config.Config.Database.Enabled, config.Config.ClickHouse.Enabled
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = false, false
t.Cleanup(func() {
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = oldDB, oldCH
})
logstore.ResetForTest()
defer logstore.ResetForTest()
ctx := context.Background()
srcDB := newLogDBSwitchDB(t)
dstDB := newLogDBSwitchDB(t)
db.SetDB(srcDB)
src, err := logstore.Active(ctx) // 无 reader 时按 seed 规则解析为 sqlite
require.NoError(t, err)
db.SetDB(dstDB)
dst, err := logstore.BuildForMigration(ctx, "sqlite")
require.NoError(t, err)
t.Cleanup(func() { db.SetDB(nil) })
now := time.Now().UTC()
rows := []analyticsmodel.NodeAccessLog{
{ID: 101, NodeID: "n1", LoggedAt: now, RemoteAddr: "1.1.1.1", Host: "a.example.com", Path: "/"},
{ID: 202, NodeID: "n2", LoggedAt: now, RemoteAddr: "2.2.2.2", Host: "b.example.com", Path: "/x"},
{ID: 303, NodeID: "n1", LoggedAt: now, RemoteAddr: "3.3.3.3", Host: "c.example.com", Path: "/y"},
}
require.NoError(t, src.AccessLogs.BatchInsertNodeAccessLogs(ctx, rows))
require.NoError(t, copyAccessLogs(ctx, src, dst))
var got []analyticsmodel.NodeAccessLog
require.NoError(t, dstDB.Order("id ASC").Find(&got).Error)
require.Len(t, got, 3)
for i, wantID := range []uint64{101, 202, 303} {
assert.Equal(t, wantID, got[i].ID, "row %d id preserved", i)
}
assert.Equal(t, "n1", got[0].NodeID)
assert.Equal(t, "n2", got[1].NodeID)
assert.Equal(t, "n1", got[2].NodeID)
assert.Equal(t, "1.1.1.1", got[0].RemoteAddr)
// 源库保持不变。
var srcCount int64
require.NoError(t, srcDB.Model(&analyticsmodel.NodeAccessLog{}).Count(&srcCount).Error)
assert.Equal(t, int64(3), srcCount)
}
// TestCopyUserAccessLogsPreservesIDs sqlite→sqlite 模拟:源库用户访问日志按 id 升序
// 复制到目标库,ID 保留、数量一致,且源库保持不变。
func TestCopyUserAccessLogsPreservesIDs(t *testing.T) {
oldDB, oldCH := config.Config.Database.Enabled, config.Config.ClickHouse.Enabled
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = false, false
t.Cleanup(func() {
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = oldDB, oldCH
})
logstore.ResetForTest()
defer logstore.ResetForTest()
ctx := context.Background()
srcDB := newLogDBSwitchDB(t)
dstDB := newLogDBSwitchDB(t)
db.SetDB(srcDB)
src, err := logstore.Active(ctx)
require.NoError(t, err)
db.SetDB(dstDB)
dst, err := logstore.BuildForMigration(ctx, "sqlite")
require.NoError(t, err)
t.Cleanup(func() { db.SetDB(nil) })
now := time.Now().UTC()
rows := []analyticsmodel.UserAccessLog{
{ID: 11, UserID: 1, Path: "/a", CreatedAt: now},
{ID: 22, UserID: 2, Path: "/b", CreatedAt: now.Add(time.Second)},
{ID: 33, UserID: 1, Path: "/c", CreatedAt: now.Add(2 * time.Second)},
}
require.NoError(t, src.UserAccessLogs.BatchInsert(ctx, rows))
require.NoError(t, copyUserAccessLogs(ctx, src, dst))
var got []analyticsmodel.UserAccessLog
require.NoError(t, dstDB.Order("id ASC").Find(&got).Error)
require.Len(t, got, 3)
for i, wantID := range []uint64{11, 22, 33} {
assert.Equal(t, wantID, got[i].ID, "row %d id preserved", i)
}
var srcCount int64
require.NoError(t, srcDB.Model(&analyticsmodel.UserAccessLog{}).Count(&srcCount).Error)
assert.Equal(t, int64(3), srcCount)
}
// TestClearTargetLogTablesClearsUserAccessLogs 验证清空目标包含用户访问日志表
// (6 张日志表之一),迁移「覆盖目标库已有日志」幂等前提成立。
func TestClearTargetLogTablesClearsUserAccessLogs(t *testing.T) {
oldDB, oldCH := config.Config.Database.Enabled, config.Config.ClickHouse.Enabled
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = false, false
t.Cleanup(func() {
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = oldDB, oldCH
})
logstore.ResetForTest()
defer logstore.ResetForTest()
ctx := context.Background()
dstDB := newLogDBSwitchDB(t)
db.SetDB(dstDB)
t.Cleanup(func() { db.SetDB(nil) })
dst, err := logstore.BuildForMigration(ctx, "sqlite")
require.NoError(t, err)
now := time.Now().UTC()
require.NoError(t, dst.UserAccessLogs.BatchInsert(ctx, []analyticsmodel.UserAccessLog{
{ID: 1, UserID: 1, Path: "/a", CreatedAt: now},
{ID: 2, UserID: 2, Path: "/b", CreatedAt: now},
}))
require.NoError(t, clearTargetLogTables(ctx, dst))
var count int64
require.NoError(t, dstDB.Model(&analyticsmodel.UserAccessLog{}).Count(&count).Error)
assert.Zero(t, count, "用户访问日志应被清空")
}
// TestClearTargetLogTablesDuringMigration 回归:冻结标记置位后,BuildForMigration 构造的
// 目标 store 必须放行用户访问日志清空/写入。skipFreeze 未传播到 UserAccessLogs store 时
// DeleteAll 会误报 ErrMigrating,导致真实切换任务在清空目标库阶段失败。
func TestClearTargetLogTablesDuringMigration(t *testing.T) {
logstore.ResetForTest()
defer logstore.ResetForTest()
gdb := newLogDBSwitchDB(t)
db.SetDB(gdb)
t.Cleanup(func() { db.SetDB(nil) })
ctx := context.Background()
logstore.SetConfigReader(func(ctx context.Context, key string) (string, error) {
cfg, err := repository.GetSystemConfigByKey(ctx, key)
if err != nil {
return "", err
}
return cfg.Value, nil
})
// 预置目标库已有日志(迁移「覆盖目标库已有日志」幂等前提)。
now := time.Now().UTC()
require.NoError(t, gdb.Create(&analyticsmodel.UserAccessLog{ID: 1, UserID: 1, Path: "/a", CreatedAt: now}).Error)
// 冻结标记置位(与真实任务 Execute 流程一致)。
require.NoError(t, setMigrationFlag(ctx, "migrating"))
t.Cleanup(func() { _ = setMigrationFlag(ctx, "") })
require.True(t, logstore.Migrating(ctx))
dst, err := logstore.BuildForMigration(ctx, "sqlite")
require.NoError(t, err)
require.NoError(t, clearTargetLogTables(ctx, dst), "迁移冻结期间目标库清空必须放行")
var count int64
require.NoError(t, gdb.Model(&analyticsmodel.UserAccessLog{}).Count(&count).Error)
assert.Zero(t, count, "用户访问日志应被清空")
}
// TestValidateSwitch 各非法组合报错。
func TestValidateSwitch(t *testing.T) {
oldDB, oldCH := config.Config.Database.Enabled, config.Config.ClickHouse.Enabled
t.Cleanup(func() {
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = oldDB, oldCH
})
gdb := newLogDBSwitchDB(t)
db.SetDB(gdb)
t.Cleanup(func() { db.SetDB(nil) })
ctx := context.Background()
setLogDB := func(v string) {
require.NoError(t, repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyLogDatabase, v))
}
t.Run("same target rejected", func(t *testing.T) {
setLogDB("sqlite")
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = false, false
err := validateSwitch(ctx, "sqlite")
require.Error(t, err)
assert.Contains(t, err.Error(), "相同")
})
t.Run("clickhouse disabled rejected", func(t *testing.T) {
setLogDB("sqlite")
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = false, false
err := validateSwitch(ctx, "clickhouse")
require.Error(t, err)
assert.Contains(t, err.Error(), "ClickHouse 未启用")
})
t.Run("postgres requires main db enabled", func(t *testing.T) {
setLogDB("sqlite")
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = false, false
err := validateSwitch(ctx, "postgres")
require.Error(t, err)
assert.Contains(t, err.Error(), "PostgreSQL 未启用")
})
t.Run("sqlite rejected when main db is postgres", func(t *testing.T) {
setLogDB("postgres")
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = true, false
err := validateSwitch(ctx, "sqlite")
require.Error(t, err)
assert.Contains(t, err.Error(), "SQLite")
})
t.Run("valid postgres migration", func(t *testing.T) {
setLogDB("sqlite")
config.Config.Database.Enabled, config.Config.ClickHouse.Enabled = true, false
require.NoError(t, validateSwitch(ctx, "postgres"))
})
}
// TestLogDBSwitchValidatePayload 参数归一化与非法值拒绝。
func TestLogDBSwitchValidatePayload(t *testing.T) {
h := &LogDBSwitchHandler{}
cases := []struct {
name string
in string
want string
ok bool
}{
{name: "postgresql normalized", in: `{"target":"postgresql"}`, want: "postgres", ok: true},
{name: "sqlite3 normalized", in: `{"target":"sqlite3"}`, want: "sqlite", ok: true},
{name: "ch normalized", in: `{"target":"ch"}`, want: "clickhouse", ok: true},
{name: "postgres passthrough", in: `{"target":"postgres"}`, want: "postgres", ok: true},
{name: "invalid target", in: `{"target":"mysql"}`, ok: false},
{name: "malformed json", in: `not-json`, ok: false},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
out, err := h.ValidatePayload([]byte(c.in))
if !c.ok {
require.Error(t, err)
return
}
require.NoError(t, err)
var p logDBSwitchPayload
require.NoError(t, json.Unmarshal(out, &p))
assert.Equal(t, c.want, p.Target)
})
}
}
// TestExecuteFailureClearsMigrationFlag 迁移失败后 log_db_migration 冻结标记被清除。
// 在 FRESH DB(不预置 log_db_migration 行)上验证:setMigrationFlag 必须 upsert 建行,
// 且失败后经缓存路径(GetSystemConfigByKey)可观察为空。
func TestExecuteFailureClearsMigrationFlag(t *testing.T) {
oldDB := config.Config.Database.Enabled
config.Config.Database.Enabled = true
t.Cleanup(func() { config.Config.Database.Enabled = oldDB })
logstore.ResetForTest()
defer logstore.ResetForTest()
gdb := newLogDBSwitchDB(t)
db.SetDB(gdb)
t.Cleanup(func() { db.SetDB(nil) })
ctx := context.Background()
// FRESH DB:log_db_migration 行不存在(不预置),log_database 预置为 sqlite。
require.NoError(t, repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyLogDatabase, "sqlite"))
_, err := repository.GetSystemConfigByKey(ctx, model.ConfigKeyLogDBMigration)
require.ErrorIs(t, err, gorm.ErrRecordNotFound)
// configReader 对 log_database 报错,使 logstore.Active 在冻结标记置位后失败。
logstore.SetConfigReader(func(_ context.Context, key string) (string, error) {
if key == model.ConfigKeyLogDatabase {
return "", errors.New("reader error")
}
return "", nil
})
_, err = (&LogDBSwitchHandler{}).Execute(ctx, []byte(`{"target":"postgres"}`))
require.Error(t, err)
assert.Contains(t, err.Error(), "reader error")
// 冻结标记必须被 upsert 持久化(行存在)并经缓存路径可观察为空,源库恢复可写。
cfg, err := repository.GetSystemConfigByKey(ctx, model.ConfigKeyLogDBMigration)
require.NoError(t, err, "setMigrationFlag 应 upsert 创建 log_db_migration 行")
assert.Empty(t, cfg.Value, "失败后冻结标记必须清除,源库保持可写")
assert.False(t, logstore.Migrating(ctx))
}
// TestSetMigrationFlagObservableThroughCache 在 FRESH DB 上验证 setMigrationFlag 写入
// 经缓存路径(logstore.Migrating → repository 读取)实时反映:置位 true、清除 false。
func TestSetMigrationFlagObservableThroughCache(t *testing.T) {
logstore.ResetForTest()
defer logstore.ResetForTest()
gdb := newLogDBSwitchDB(t)
db.SetDB(gdb)
t.Cleanup(func() { db.SetDB(nil) })
ctx := context.Background()
// 按 bootstrap 同款注入 repository 读取,走 RAM 缓存路径。
logstore.SetConfigReader(func(ctx context.Context, key string) (string, error) {
cfg, err := repository.GetSystemConfigByKey(ctx, key)
if err != nil {
return "", err
}
return cfg.Value, nil
})
// FRESH DB:行缺失 → fail-open false。
assert.False(t, logstore.Migrating(ctx))
require.NoError(t, setMigrationFlag(ctx, "migrating"))
assert.True(t, logstore.Migrating(ctx), "置位后缓存路径必须立即观察到 migrating")
require.NoError(t, setMigrationFlag(ctx, ""))
assert.False(t, logstore.Migrating(ctx), "清除后缓存路径必须立即观察到非 migrating")
}
// TestFlipLogDatabaseRefreshesCachedConfig 验证翻转日志主库后缓存路径立即反映新库
// (logstore.ActiveDatabase / GetSystemConfigByKey),防止各进程继续写旧库(split-brain)。
func TestFlipLogDatabaseRefreshesCachedConfig(t *testing.T) {
logstore.ResetForTest()
defer logstore.ResetForTest()
gdb := newLogDBSwitchDB(t)
db.SetDB(gdb)
t.Cleanup(func() { db.SetDB(nil) })
ctx := context.Background()
logstore.SetConfigReader(func(ctx context.Context, key string) (string, error) {
cfg, err := repository.GetSystemConfigByKey(ctx, key)
if err != nil {
return "", err
}
return cfg.Value, nil
})
require.NoError(t, repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyLogDatabase, "sqlite"))
active, err := logstore.ActiveDatabase(ctx)
require.NoError(t, err)
assert.Equal(t, "sqlite", active) // 预热缓存
require.NoError(t, flipLogDatabase(ctx, "postgres"))
active, err = logstore.ActiveDatabase(ctx)
require.NoError(t, err)
assert.Equal(t, "postgres", active, "翻转后缓存路径必须立即反映新库")
cfg, err := repository.GetSystemConfigByKey(ctx, model.ConfigKeyLogDatabase)
require.NoError(t, err)
assert.Equal(t, "postgres", cfg.Value)
}
+1 -37
View File
@@ -17,35 +17,6 @@ import (
"time"
)
// redactSensitiveJSON masks values of sensitive keys (password, token, secret,
// api key) so credentials never appear in logs.
func redactSensitiveJSON(v any) any {
switch t := v.(type) {
case map[string]any:
for k, val := range t {
if isSensitiveLogKey(k) {
t[k] = "***"
} else {
t[k] = redactSensitiveJSON(val)
}
}
case []any:
for i, val := range t {
t[i] = redactSensitiveJSON(val)
}
}
return v
}
func isSensitiveLogKey(k string) bool {
switch strings.ToLower(k) {
case "password", "passwd", "secret", "token", "access_token", "api_key", "apikey":
return true
default:
return false
}
}
const emitAckTimeout = 10 * time.Second
// Monitor represents a monitor entry from Uptime Kuma.
@@ -321,14 +292,7 @@ func (c *SocketIOClient) Emit(event string, args ...any) (string, error) {
}
body := fmt.Sprintf("42%d%s", id, string(bs))
logPayload := string(bs)
var decoded any
if err := json.Unmarshal(bs, &decoded); err == nil {
if redacted, err := json.Marshal(redactSensitiveJSON(decoded)); err == nil {
logPayload = string(redacted)
}
}
slog.Debug("Emitting Socket.IO event", "event", event, "ackID", id, "payload", logPayload)
slog.Debug("Emitting Socket.IO event", "event", event, "ackID", id, "payload_len", len(bs))
u := fmt.Sprintf("%s/socket.io/?EIO=4&transport=polling&sid=%s", c.baseURL, c.sid)
req, err := http.NewRequestWithContext(c.ctx, http.MethodPost, u, strings.NewReader(body))
@@ -1,28 +0,0 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package uptimekuma
import (
"encoding/json"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
func TestRedactSensitiveJSON(t *testing.T) {
payload := []byte(`["login",{"username":"admin","password":"s3cret","nested":{"token":"abc","label":"keep"}},"plain"]`)
var decoded any
require.NoError(t, json.Unmarshal(payload, &decoded))
out, err := json.Marshal(redactSensitiveJSON(decoded))
require.NoError(t, err)
s := string(out)
assert.NotContains(t, s, "s3cret")
assert.NotContains(t, s, "abc")
assert.Contains(t, s, `"password":"***"`)
assert.Contains(t, s, `"token":"***"`)
assert.Contains(t, s, `"label":"keep"`)
assert.Contains(t, s, `"username":"admin"`)
}
@@ -12,6 +12,7 @@ import (
"time"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/testhelper"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
@@ -34,9 +35,8 @@ func setupIPGroupSyncTestDB(t *testing.T) func() {
))
db.SetDB(sqliteDB)
resetAccessLogStore := repository.SetAccessLogStoreForTest(repository.NewMemoryAccessLogStore())
testhelper.SetupLogStoresForTest(t)
return func() {
resetAccessLogStore()
db.SetDB(nil)
}
}
+2 -2
View File
@@ -9,6 +9,7 @@ import (
"time"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/testhelper"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
@@ -61,8 +62,7 @@ func TestLegacyImportUsesEffectiveTLDPlusOne(t *testing.T) {
func TestGetStatsAggregatesZoneHosts(t *testing.T) {
ctx := setupZoneDB(t)
reset := repository.SetAccessLogStoreForTest(repository.NewMemoryAccessLogStore())
t.Cleanup(reset)
testhelper.SetupLogStoresForTest(t)
zone, err := Create(ctx, Input{Domain: "example.com"})
require.NoError(t, err)
-129
View File
@@ -1,129 +0,0 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package risk_control
import (
"context"
"sync"
"time"
"github.com/Rain-kl/Wavelet/internal/infra/config"
"github.com/Rain-kl/Wavelet/internal/infra/persistence/batchwriter"
"github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/internal/platform/lifecycle"
analyticsrepo "github.com/Rain-kl/Wavelet/internal/repository/analytics"
"github.com/Rain-kl/Wavelet/pkg/logger"
)
const (
// Bound visibility lag for sparse access-log traffic when MinBatchSize is not met.
accessLogMaxFlushWait = 3 * time.Second
)
var (
logWriterMu sync.RWMutex
logWriter *batchwriter.Writer[*analytics.UserAccessLog]
)
// InitLogWriter initializes the ClickHouse access-log batch writer.
func InitLogWriter(ctx context.Context) {
if !config.Config.ClickHouse.Enabled {
return
}
logWriterMu.Lock()
defer logWriterMu.Unlock()
if logWriter != nil {
return
}
cfg := batchwriter.DefaultConfig()
cfg.Name = "user_access_logs"
cfg.MaxFlushWait = accessLogMaxFlushWait
writer, err := batchwriter.New[*analytics.UserAccessLog](cfg, func(ctx context.Context, items []*analytics.UserAccessLog) error {
rows := make([]analytics.UserAccessLog, 0, len(items))
for _, item := range items {
if item == nil {
continue
}
rows = append(rows, *item)
}
return analyticsrepo.BatchInsert(ctx, rows)
},
batchwriter.WithDropHandler[*analytics.UserAccessLog](func(item *analytics.UserAccessLog) {
path := ""
if item != nil {
path = item.Path
}
logger.WarnF(context.Background(), "[RiskControl] Log queue full, dropping log item for path: %s", path)
}),
batchwriter.WithFlushErrorHandler[*analytics.UserAccessLog](func(ctx context.Context, items []*analytics.UserAccessLog, err error) {
logger.ErrorF(ctx, "[RiskControl] Send ClickHouse batch failed (batch=%d): %v", len(items), err)
}),
)
if err != nil {
logger.ErrorF(ctx, "[RiskControl] init log writer failed: %v", err)
return
}
writer.Start(ctx)
logWriter = writer
lifecycle.OnShutdown("risk_control_log_writer", StopLogWriter)
}
// StopLogWriter stops the ClickHouse access-log batch writer and drains pending logs.
func StopLogWriter(ctx context.Context) error {
writer := currentLogWriter()
if writer == nil {
return nil
}
return writer.Stop(ctx)
}
// IsBufferFull reports whether the access-log queue has no remaining capacity.
func IsBufferFull() bool {
writer := currentLogWriter()
if writer == nil {
return false
}
return writer.IsFull()
}
// LogWriterStats returns queue depth and failure counters for the access-log writer.
// When the writer is not initialized, it returns a zero-value Stats with the expected name.
func LogWriterStats() batchwriter.Stats {
writer := currentLogWriter()
if writer == nil {
return batchwriter.Stats{Name: "user_access_logs"}
}
return writer.Stats()
}
// QueueAccessLog enqueues an access log without blocking.
func QueueAccessLog(logItem *analytics.UserAccessLog) {
writer := currentLogWriter()
if writer == nil || logItem == nil {
return
}
writer.TryEnqueue(logItem)
}
// SetLogWriterForTest swaps the access-log writer for unit tests.
func SetLogWriterForTest(writer *batchwriter.Writer[*analytics.UserAccessLog]) func() {
logWriterMu.Lock()
previous := logWriter
logWriter = writer
logWriterMu.Unlock()
return func() {
logWriterMu.Lock()
logWriter = previous
logWriterMu.Unlock()
}
}
func currentLogWriter() *batchwriter.Writer[*analytics.UserAccessLog] {
logWriterMu.RLock()
defer logWriterMu.RUnlock()
return logWriter
}
-30
View File
@@ -1,30 +0,0 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package risk_control
import (
"testing"
"time"
)
func TestAccessLogMaxFlushWaitInRange(t *testing.T) {
t.Parallel()
if accessLogMaxFlushWait < 2*time.Second || accessLogMaxFlushWait > 5*time.Second {
t.Fatalf("accessLogMaxFlushWait = %v, want in [2s, 5s]", accessLogMaxFlushWait)
}
}
func TestLogWriterStatsWhenNil(t *testing.T) {
t.Parallel()
reset := SetLogWriterForTest(nil)
t.Cleanup(reset)
stats := LogWriterStats()
if stats.Name != "user_access_logs" {
t.Fatalf("LogWriterStats().Name = %q, want user_access_logs", stats.Name)
}
if stats.Running {
t.Fatal("LogWriterStats().Running = true for nil writer, want false")
}
}
-134
View File
@@ -1,134 +0,0 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
// Package risk_control 提供风险控制中间件
package risk_control
import (
"crypto/sha256"
"encoding/hex"
"encoding/json"
"net/http"
"strings"
"time"
"github.com/Rain-kl/Wavelet/internal/apps/oauth"
"github.com/Rain-kl/Wavelet/internal/infra/config"
"github.com/Rain-kl/Wavelet/internal/infra/persistence/idgen"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/internal/shared/response"
"github.com/gin-gonic/gin"
)
const maxAuditLogHeadersBytes = 2 * 1024
var auditLogHeaderAllowlist = map[string]struct{}{
"Authorization": {},
"Cookie": {},
"X-Forwarded-For": {},
"X-Real-Ip": {},
"User-Agent": {},
"Content-Type": {},
}
func marshalAuditLogHeaders(headers http.Header) string {
if headers == nil {
return ""
}
filtered := make(http.Header)
for key, values := range headers {
if _, ok := auditLogHeaderAllowlist[key]; !ok {
continue
}
filtered[key] = redactAuditLogHeaderValues(key, values)
}
headersBytes, err := json.Marshal(filtered)
if err != nil {
return ""
}
if len(headersBytes) <= maxAuditLogHeadersBytes {
return string(headersBytes)
}
return string(headersBytes[:maxAuditLogHeadersBytes])
}
func redactAuditLogHeaderValues(key string, values []string) []string {
switch key {
case "Authorization", "Cookie":
redacted := make([]string, len(values))
for i, value := range values {
redacted[i] = hashAuditLogSensitiveValue(value)
}
return redacted
default:
return values
}
}
func hashAuditLogSensitiveValue(value string) string {
if strings.TrimSpace(value) == "" {
return ""
}
sum := sha256.Sum256([]byte(value))
return "sha256:" + hex.EncodeToString(sum[:8])
}
// RiskControlMiddleware 全局日志采集中间件
func RiskControlMiddleware() gin.HandlerFunc {
return func(c *gin.Context) {
// 如果未启用 ClickHouse,直接放行
if !config.Config.ClickHouse.Enabled {
c.Next()
return
}
// 1. 限流背压检测(检测本地缓冲队列是否已满)
if IsBufferFull() {
response.AbortTooManyRequests(c, "系统繁忙,请稍后再试")
return
}
start := time.Now()
// 2. 执行后续请求(穿过业务处理和认证中间件)
c.Next()
// 3. 后置身份检查:仅记录通过认证的请求
userObj, exists := oauth.GetFromContext[*model.User](c, oauth.UserObjKey)
if !exists || userObj == nil {
return
}
// 4. 计算耗时并异步推送到缓冲队列
latency := time.Since(start).Milliseconds()
headersStr := marshalAuditLogHeaders(c.Request.Header)
const maxHTTPStatus = 999
status := c.Writer.Status()
if status < 0 {
status = 0
} else if status > maxHTTPStatus {
status = maxHTTPStatus
}
logItem := &analytics.UserAccessLog{
ID: idgen.NextUint64ID(),
UserID: userObj.ID, // 直接从 Context 获取已登录用户ID,避免数据库查询
Path: c.Request.URL.Path,
Method: c.Request.Method,
IP: c.ClientIP(),
UserAgent: c.Request.UserAgent(),
Headers: headersStr,
Status: int32(status),
Latency: latency,
CreatedAt: time.Now(),
}
// 非阻塞地推入缓存队列
QueueAccessLog(logItem)
}
}
@@ -1,178 +0,0 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package risk_control
import (
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"testing"
"time"
"github.com/Rain-kl/Wavelet/internal/apps/oauth"
"github.com/Rain-kl/Wavelet/internal/infra/config"
"github.com/Rain-kl/Wavelet/internal/infra/persistence/batchwriter"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/internal/testhelper"
"github.com/gin-gonic/gin"
"github.com/stretchr/testify/assert"
)
const testLogQueueSize = 10_000
func newTestLogWriter(t *testing.T, queueSize int) (*batchwriter.Writer[*analytics.UserAccessLog], chan *analytics.UserAccessLog) {
t.Helper()
received := make(chan *analytics.UserAccessLog, queueSize)
cfg := batchwriter.Config{
QueueSize: queueSize,
MaxBatchSize: 1,
FlushInterval: 5 * time.Millisecond,
}
writer, err := batchwriter.New[*analytics.UserAccessLog](cfg, func(_ context.Context, items []*analytics.UserAccessLog) error {
for _, item := range items {
if item != nil {
received <- item
}
}
return nil
})
assert.NoError(t, err)
writer.Start(context.Background())
t.Cleanup(func() {
stopCtx, cancel := context.WithTimeout(context.Background(), time.Second)
defer cancel()
_ = writer.Stop(stopCtx)
})
return writer, received
}
func TestRiskControlMiddleware(t *testing.T) {
gin.SetMode(gin.TestMode)
t.Run("ClickHouse disabled", func(t *testing.T) {
config.Config.ClickHouse.Enabled = false
defer func() { config.Config.ClickHouse.Enabled = false }()
r := testhelper.NewTestGinEngine(RiskControlMiddleware())
r.GET("/test", func(c *gin.Context) {
c.String(http.StatusOK, "ok")
})
w := httptest.NewRecorder()
req, _ := http.NewRequest(http.MethodGet, "/test", nil)
r.ServeHTTP(w, req)
assert.Equal(t, http.StatusOK, w.Code)
assert.Equal(t, "ok", w.Body.String())
})
t.Run("ClickHouse enabled - Normal Authenticated Request", func(t *testing.T) {
config.Config.ClickHouse.Enabled = true
writer, received := newTestLogWriter(t, testLogQueueSize)
resetWriter := SetLogWriterForTest(writer)
defer func() {
config.Config.ClickHouse.Enabled = false
resetWriter()
}()
r := gin.New()
r.Use(func(c *gin.Context) {
user := &model.User{ID: 12345}
oauth.SetToContext(c, oauth.UserObjKey, user)
c.Next()
})
r.Use(RiskControlMiddleware())
r.GET("/test", func(c *gin.Context) {
c.String(http.StatusOK, "ok")
})
w := httptest.NewRecorder()
req, _ := http.NewRequest(http.MethodGet, "/test", nil)
req.Header.Set("X-Test-Header", "hello")
req.Header.Set("Cookie", "session_id=abcdef123456")
req.Header.Set("Authorization", "Bearer secret-token")
req.Header.Set("Content-Type", "application/json")
r.ServeHTTP(w, req)
assert.Equal(t, http.StatusOK, w.Code)
assert.Equal(t, "ok", w.Body.String())
select {
case logItem := <-received:
assert.Equal(t, uint64(12345), logItem.UserID)
assert.Equal(t, "/test", logItem.Path)
assert.Equal(t, http.MethodGet, logItem.Method)
assert.Equal(t, int32(http.StatusOK), logItem.Status)
assert.NotEmpty(t, logItem.Headers)
assert.NotContains(t, logItem.Headers, "X-Test-Header")
assert.Contains(t, logItem.Headers, "Content-Type")
assert.Contains(t, logItem.Headers, "sha256:")
assert.NotContains(t, logItem.Headers, "secret-token")
assert.NotContains(t, logItem.Headers, "session_id=abcdef123456")
case <-time.After(200 * time.Millisecond):
t.Fatal("expected flushed log item, but got none")
}
})
t.Run("ClickHouse enabled - Unauthenticated Request", func(t *testing.T) {
config.Config.ClickHouse.Enabled = true
writer, received := newTestLogWriter(t, testLogQueueSize)
resetWriter := SetLogWriterForTest(writer)
defer func() {
config.Config.ClickHouse.Enabled = false
resetWriter()
}()
r := testhelper.NewTestGinEngine(RiskControlMiddleware())
r.GET("/test", func(c *gin.Context) {
c.String(http.StatusOK, "ok")
})
w := httptest.NewRecorder()
req, _ := http.NewRequest(http.MethodGet, "/test", nil)
r.ServeHTTP(w, req)
assert.Equal(t, http.StatusOK, w.Code)
assert.Equal(t, "ok", w.Body.String())
select {
case <-received:
t.Fatal("expected no log item for unauthenticated request")
case <-time.After(50 * time.Millisecond):
}
})
t.Run("ClickHouse enabled - Buffer Full Rate Limiting", func(t *testing.T) {
config.Config.ClickHouse.Enabled = true
writer, _ := newTestLogWriter(t, 2)
resetWriter := SetLogWriterForTest(writer)
defer func() {
config.Config.ClickHouse.Enabled = false
resetWriter()
}()
for range 2 {
assert.True(t, writer.TryEnqueue(&analytics.UserAccessLog{}))
}
r := testhelper.NewTestGinEngine(RiskControlMiddleware())
r.GET("/test", func(c *gin.Context) {
c.String(http.StatusOK, "ok")
})
w := httptest.NewRecorder()
req, _ := http.NewRequest(http.MethodGet, "/test", nil)
r.ServeHTTP(w, req)
assert.Equal(t, http.StatusTooManyRequests, w.Code)
var resp map[string]interface{}
err := json.Unmarshal(w.Body.Bytes(), &resp)
assert.NoError(t, err)
assert.Contains(t, resp["error_msg"], "系统繁忙")
})
}
+13
View File
@@ -18,6 +18,7 @@ import (
"github.com/Rain-kl/Wavelet/internal/infra/task"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
"github.com/Rain-kl/Wavelet/pkg/logger"
)
@@ -129,6 +130,18 @@ func (h *SystemCleanupHandler) Execute(ctx context.Context, _ []byte) (*task.Tas
)
}
task.AppendLog(ctx, "开始清理过期日志(访问日志按当前日志库保留天数,性能指标按独立短留存)...")
summary, err := logstore.CleanupExpired(ctx)
switch {
case err != nil:
logger.ErrorF(ctx, "清理过期日志失败: %v", err)
task.AppendLog(ctx, "清理过期日志失败: %v", err)
case summary.Deleted == 0:
task.AppendLog(ctx, "没有需要清理的过期日志 (访问日志保留 %d 天,性能指标保留 %d 天)", summary.RetentionDays, summary.MetricRetentionDays)
default:
task.AppendLog(ctx, "日志清理完成:访问日志保留 %d 天,性能指标保留 %d 天,删除 %d 条", summary.RetentionDays, summary.MetricRetentionDays, summary.Deleted)
}
msg := fmt.Sprintf("系统清理完成。成功清理未使用的上传文件 %d/%d 个;清理历史推送审计日志 %d 条;清理任务执行日志 %d 条。",
totalDeleted,
totalProcessed,
+2 -3
View File
@@ -130,9 +130,8 @@ func applyClickHouseDefaults(c *configModel) {
}
// Keep Enabled=true from env and continue applying host/pool defaults.
}
if !c.ClickHouse.Enabled {
c.ClickHouse.Enabled = true
}
// 未显式启用(缺省或 enabled: false)时保持关闭,不再强制打开;
// 显式启用后补齐连接默认参数。
if c.ClickHouse.Database == "" {
c.ClickHouse.Database = "openflare"
}
+32
View File
@@ -12,3 +12,35 @@ func TestApplyEnvOverridesRedisMaintNotifications(t *testing.T) {
t.Fatal("REDIS_MAINT_NOTIFICATIONS=true was not applied")
}
}
// TestApplyClickHouseDefaultsRespectsDisabledConfig 回归:配置未显式启用(缺省或
// enabled: false)时 ClickHouse 必须保持关闭,不得被 applyClickHouseDefaults 强制打开。
// CLICKHOUSE_ENABLED=true 仅用于绕过测试分支(isTest 默认强制关闭),以便测真实默认逻辑。
func TestApplyClickHouseDefaultsRespectsDisabledConfig(t *testing.T) {
t.Setenv("CLICKHOUSE_ENABLED", "true")
cfg := &configModel{ClickHouse: clickHouseConfig{Enabled: false}}
applyClickHouseDefaults(cfg)
if cfg.ClickHouse.Enabled {
t.Fatal("ClickHouse must stay disabled when config does not enable it")
}
}
// TestApplyClickHouseDefaultsEnablesWhenConfigured 显式 enabled: true 时保持启用并补齐默认连接参数。
func TestApplyClickHouseDefaultsEnablesWhenConfigured(t *testing.T) {
t.Setenv("CLICKHOUSE_ENABLED", "true")
cfg := &configModel{ClickHouse: clickHouseConfig{Enabled: true}}
applyClickHouseDefaults(cfg)
if !cfg.ClickHouse.Enabled {
t.Fatal("ClickHouse must stay enabled when explicitly configured")
}
if cfg.ClickHouse.Database != "openflare" {
t.Fatalf("database default not applied: %q", cfg.ClickHouse.Database)
}
if len(cfg.ClickHouse.Hosts) != 1 || cfg.ClickHouse.Hosts[0] != "127.0.0.1:9000" {
t.Fatalf("hosts default not applied: %v", cfg.ClickHouse.Hosts)
}
}
@@ -11,6 +11,8 @@ import (
"sync"
"sync/atomic"
"time"
"github.com/Rain-kl/Wavelet/internal/model/analytics"
)
// FlushFunc persists a batch of queued items. It is invoked from the worker goroutine.
@@ -22,14 +24,8 @@ type FlushFunc[T any] func(ctx context.Context, items []T) error
type FlushErrorHandler[T any] func(ctx context.Context, items []T, err error)
// Stats is a point-in-time snapshot of Writer queue and failure counters.
type Stats struct {
Name string `json:"name"`
Depth int `json:"depth"`
Cap int `json:"cap"`
Drops int64 `json:"drops"`
FlushErrors int64 `json:"flush_errors"`
Running bool `json:"running"`
}
// It is an alias of analyticsmodel.BatchWriterStats (moved to keep model pure data).
type Stats = analytics.BatchWriterStats
// Writer buffers items and flushes them by size or interval.
type Writer[T any] struct {
@@ -0,0 +1,115 @@
-- +goose Up
-- 节点访问日志:按月 RANGE 分区,复合主键 (id, logged_at) 满足分区键进唯一索引要求。
CREATE TABLE IF NOT EXISTS of_node_access_logs (
id BIGINT NOT NULL,
node_id VARCHAR(64) NOT NULL DEFAULT '',
logged_at TIMESTAMPTZ NOT NULL,
remote_addr VARCHAR(128) NOT NULL DEFAULT '',
region VARCHAR(128) NOT NULL DEFAULT '',
host VARCHAR(255) NOT NULL DEFAULT '',
path VARCHAR(2048) NOT NULL DEFAULT '',
user_agent TEXT NOT NULL DEFAULT '',
cache_status VARCHAR(64) NOT NULL DEFAULT '',
status_code INTEGER NOT NULL DEFAULT 0,
bytes_sent BIGINT NOT NULL DEFAULT 0,
request_length BIGINT NOT NULL DEFAULT 0,
request_time_ms INTEGER NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (id, logged_at)
) PARTITION BY RANGE (logged_at);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_node_id ON of_node_access_logs (node_id, logged_at DESC);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_host ON of_node_access_logs (host, logged_at DESC);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_remote_addr ON of_node_access_logs (remote_addr, logged_at DESC);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_status_code ON of_node_access_logs (status_code, logged_at DESC);
-- 用户访问日志:按月分区。
CREATE TABLE IF NOT EXISTS w_user_access_logs (
id BIGINT NOT NULL,
user_id BIGINT NOT NULL DEFAULT 0,
path VARCHAR(2048) NOT NULL DEFAULT '',
method VARCHAR(16) NOT NULL DEFAULT '',
ip VARCHAR(128) NOT NULL DEFAULT '',
user_agent TEXT NOT NULL DEFAULT '',
headers TEXT NOT NULL DEFAULT '',
status INTEGER NOT NULL DEFAULT 0,
latency BIGINT NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (id, created_at)
) PARTITION BY RANGE (created_at);
CREATE INDEX IF NOT EXISTS idx_w_user_access_logs_user_id ON w_user_access_logs (user_id, created_at DESC);
-- 可观测 4 表:普通表 + (node_id, captured_at DESC) 索引。
CREATE TABLE IF NOT EXISTS of_node_metric_snapshots (
id BIGINT NOT NULL PRIMARY KEY,
node_id VARCHAR(64) NOT NULL DEFAULT '',
captured_at TIMESTAMPTZ NOT NULL,
cpu_usage_percent DOUBLE PRECISION NOT NULL DEFAULT 0,
memory_used_bytes BIGINT NOT NULL DEFAULT 0,
memory_total_bytes BIGINT NOT NULL DEFAULT 0,
storage_used_bytes BIGINT NOT NULL DEFAULT 0,
storage_total_bytes BIGINT NOT NULL DEFAULT 0,
disk_read_bytes BIGINT NOT NULL DEFAULT 0,
disk_write_bytes BIGINT NOT NULL DEFAULT 0,
network_rx_bytes BIGINT NOT NULL DEFAULT 0,
network_tx_bytes BIGINT NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_metric_snapshots_node ON of_node_metric_snapshots (node_id, captured_at DESC);
CREATE TABLE IF NOT EXISTS of_node_edge_health (
id BIGINT NOT NULL PRIMARY KEY,
node_id VARCHAR(64) NOT NULL DEFAULT '',
captured_at TIMESTAMPTZ NOT NULL,
status VARCHAR(64) NOT NULL DEFAULT '',
connections BIGINT NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_edge_health_node ON of_node_edge_health (node_id, captured_at DESC);
CREATE TABLE IF NOT EXISTS of_node_obs_frps (
id BIGINT NOT NULL PRIMARY KEY,
node_id VARCHAR(64) NOT NULL DEFAULT '',
captured_at TIMESTAMPTZ NOT NULL,
frps_connections INTEGER NOT NULL DEFAULT 0,
frps_proxy_count INTEGER NOT NULL DEFAULT 0,
frps_client_count INTEGER NOT NULL DEFAULT 0,
frps_proxies TEXT NOT NULL DEFAULT '',
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_obs_frps_node ON of_node_obs_frps (node_id, captured_at DESC);
CREATE TABLE IF NOT EXISTS of_node_obs_frpc (
id BIGINT NOT NULL PRIMARY KEY,
node_id VARCHAR(64) NOT NULL DEFAULT '',
captured_at TIMESTAMPTZ NOT NULL,
tunnel_status VARCHAR(16) NOT NULL DEFAULT '',
connected_relays_count INTEGER NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_obs_frpc_node ON of_node_obs_frpc (node_id, captured_at DESC);
-- 分区预建:创建当月及未来 2 个月分区(共 3 个月)。
-- +goose StatementBegin
DO $$
DECLARE
d date;
BEGIN
FOR d IN SELECT generate_series(date_trunc('month', now())::date, (date_trunc('month', now()) + interval '2 months')::date, interval '1 month')::date
LOOP
EXECUTE format('CREATE TABLE IF NOT EXISTS of_node_access_logs_%s PARTITION OF of_node_access_logs FOR VALUES FROM (%L) TO (%L)',
to_char(d, 'YYYYMM'), d, d + interval '1 month');
EXECUTE format('CREATE TABLE IF NOT EXISTS w_user_access_logs_%s PARTITION OF w_user_access_logs FOR VALUES FROM (%L) TO (%L)',
to_char(d, 'YYYYMM'), d, d + interval '1 month');
END LOOP;
END $$;
-- +goose StatementEnd
-- +goose Down
DROP TABLE IF EXISTS w_user_access_logs;
DROP TABLE IF EXISTS of_node_access_logs;
DROP TABLE IF EXISTS of_node_metric_snapshots;
DROP TABLE IF EXISTS of_node_edge_health;
DROP TABLE IF EXISTS of_node_obs_frps;
DROP TABLE IF EXISTS of_node_obs_frpc;
@@ -0,0 +1,23 @@
-- +goose Up
-- 日志保留天数配置(business),替换旧的 database_auto_cleanup_* 键。
-- 默认统一 30 天:不继承旧键值,避免旧配置被静默带入导致日志被过度清理。
INSERT INTO w_system_configs (key, value, type, visibility, description, created_at, updated_at)
SELECT k, '30', 'business', 0, descr, CURRENT_TIMESTAMP, CURRENT_TIMESTAMP
FROM (VALUES
('log_retention_days_postgres', 'PostgreSQL 日志保留天数(访问日志与可观测统一)'),
('log_retention_days_sqlite', 'SQLite 日志保留天数'),
('log_retention_days_clickhouse', 'ClickHouse 日志保留天数')
) AS v(k, descr)
ON CONFLICT (key) DO NOTHING;
DELETE FROM w_system_configs WHERE key IN ('database_auto_cleanup_enabled', 'database_auto_cleanup_retention_days');
-- +goose Down
INSERT INTO w_system_configs (key, value, type, visibility, description, created_at, updated_at)
VALUES
('database_auto_cleanup_enabled', 'true', 'business', 0, '数据库自动清理开关', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP),
('database_auto_cleanup_retention_days', COALESCE((SELECT value FROM w_system_configs WHERE key = 'log_retention_days_postgres'), '30'), 'business', 0, '数据库保留天数', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP)
ON CONFLICT (key) DO NOTHING;
DELETE FROM w_system_configs WHERE key IN ('log_retention_days_postgres', 'log_retention_days_sqlite', 'log_retention_days_clickhouse');
@@ -0,0 +1,19 @@
-- +goose Up
CREATE TABLE IF NOT EXISTS w_schedules_backup_of_database_auto_cleanup AS
SELECT * FROM w_schedules WHERE task_type = 'of_database_auto_cleanup';
DELETE FROM w_schedules WHERE task_type = 'of_database_auto_cleanup';
-- +goose Down
INSERT INTO w_schedules (id, name, task_type, cron, payload, is_active, created_at, updated_at)
SELECT id, name, task_type, cron, payload, is_active, created_at, updated_at
FROM w_schedules_backup_of_database_auto_cleanup
ON CONFLICT (id) DO NOTHING;
INSERT INTO w_schedules (id, name, task_type, cron, payload, is_active, created_at, updated_at)
SELECT 102, 'OpenFlare 可观测数据自动清理', 'of_database_auto_cleanup', '0 3 * * *', '{}', TRUE, CURRENT_TIMESTAMP, CURRENT_TIMESTAMP
WHERE NOT EXISTS (SELECT 1 FROM w_schedules WHERE task_type = 'of_database_auto_cleanup')
ON CONFLICT (id) DO NOTHING;
DROP TABLE IF EXISTS w_schedules_backup_of_database_auto_cleanup;
@@ -0,0 +1,9 @@
-- +goose Up
-- 性能指标(CPU/内存/磁盘/网络)保留天数:三库共用、独立短留存(默认 3 天),
-- 不随 log_retention_days_*(访问日志保留配置)变化。
INSERT INTO w_system_configs (key, value, type, visibility, description, created_at, updated_at)
VALUES ('metric_retention_days', '3', 'business', 0, '性能指标(CPU/内存/磁盘/网络)保留天数(三库共用)', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP)
ON CONFLICT (key) DO NOTHING;
-- +goose Down
DELETE FROM w_system_configs WHERE key = 'metric_retention_days';
@@ -0,0 +1,13 @@
-- +goose Up
-- 系统定期垃圾清理改为每日执行一次(凌晨 3 点,Asia/Shanghai)。
-- 此前为每 2 小时(0 */2 * * *)高频扫描,日常清理收益有限,降频减少非必要扫描。
UPDATE w_schedules
SET cron = '0 3 * * *',
updated_at = CURRENT_TIMESTAMP
WHERE task_type = 'system_cleanup';
-- +goose Down
UPDATE w_schedules
SET cron = '0 */2 * * *',
updated_at = CURRENT_TIMESTAMP
WHERE task_type = 'system_cleanup';
@@ -0,0 +1,9 @@
-- +goose Up
-- 日志列表默认排序与 retention 清理的时间范围条件:补 logged_at 前导索引。
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_logged_at ON of_node_access_logs (logged_at DESC, id DESC);
-- hosts 过滤为 lower(trim(host)) IN:补表达式索引,普通 (host,...) 索引无法命中函数表达式。
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_host_lower ON of_node_access_logs (lower(trim(host)));
-- +goose Down
DROP INDEX IF EXISTS idx_of_node_access_logs_host_lower;
DROP INDEX IF EXISTS idx_of_node_access_logs_logged_at;
@@ -0,0 +1,93 @@
-- +goose Up
-- 节点访问日志:普通表(同 PG 语义,索引名保持一致)。
CREATE TABLE IF NOT EXISTS of_node_access_logs (
id INTEGER PRIMARY KEY AUTOINCREMENT,
node_id TEXT NOT NULL DEFAULT '',
logged_at DATETIME NOT NULL,
remote_addr TEXT NOT NULL DEFAULT '',
region TEXT NOT NULL DEFAULT '',
host TEXT NOT NULL DEFAULT '',
path TEXT NOT NULL DEFAULT '',
user_agent TEXT NOT NULL DEFAULT '',
cache_status TEXT NOT NULL DEFAULT '',
status_code INTEGER NOT NULL DEFAULT 0,
bytes_sent INTEGER NOT NULL DEFAULT 0,
request_length INTEGER NOT NULL DEFAULT 0,
request_time_ms INTEGER NOT NULL DEFAULT 0,
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_node_id ON of_node_access_logs (node_id, logged_at DESC);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_host ON of_node_access_logs (host, logged_at DESC);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_remote_addr ON of_node_access_logs (remote_addr, logged_at DESC);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_status_code ON of_node_access_logs (status_code, logged_at DESC);
CREATE TABLE IF NOT EXISTS w_user_access_logs (
id INTEGER PRIMARY KEY AUTOINCREMENT,
user_id INTEGER NOT NULL DEFAULT 0,
path TEXT NOT NULL DEFAULT '',
method TEXT NOT NULL DEFAULT '',
ip TEXT NOT NULL DEFAULT '',
user_agent TEXT NOT NULL DEFAULT '',
headers TEXT NOT NULL DEFAULT '',
status INTEGER NOT NULL DEFAULT 0,
latency INTEGER NOT NULL DEFAULT 0,
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_w_user_access_logs_user_id ON w_user_access_logs (user_id, created_at DESC);
CREATE TABLE IF NOT EXISTS of_node_metric_snapshots (
id INTEGER PRIMARY KEY AUTOINCREMENT,
node_id TEXT NOT NULL DEFAULT '',
captured_at DATETIME NOT NULL,
cpu_usage_percent REAL NOT NULL DEFAULT 0,
memory_used_bytes INTEGER NOT NULL DEFAULT 0,
memory_total_bytes INTEGER NOT NULL DEFAULT 0,
storage_used_bytes INTEGER NOT NULL DEFAULT 0,
storage_total_bytes INTEGER NOT NULL DEFAULT 0,
disk_read_bytes INTEGER NOT NULL DEFAULT 0,
disk_write_bytes INTEGER NOT NULL DEFAULT 0,
network_rx_bytes INTEGER NOT NULL DEFAULT 0,
network_tx_bytes INTEGER NOT NULL DEFAULT 0,
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_metric_snapshots_node ON of_node_metric_snapshots (node_id, captured_at DESC);
CREATE TABLE IF NOT EXISTS of_node_edge_health (
id INTEGER PRIMARY KEY AUTOINCREMENT,
node_id TEXT NOT NULL DEFAULT '',
captured_at DATETIME NOT NULL,
status TEXT NOT NULL DEFAULT '',
connections INTEGER NOT NULL DEFAULT 0,
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_edge_health_node ON of_node_edge_health (node_id, captured_at DESC);
CREATE TABLE IF NOT EXISTS of_node_obs_frps (
id INTEGER PRIMARY KEY AUTOINCREMENT,
node_id TEXT NOT NULL DEFAULT '',
captured_at DATETIME NOT NULL,
frps_connections INTEGER NOT NULL DEFAULT 0,
frps_proxy_count INTEGER NOT NULL DEFAULT 0,
frps_client_count INTEGER NOT NULL DEFAULT 0,
frps_proxies TEXT NOT NULL DEFAULT '',
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_obs_frps_node ON of_node_obs_frps (node_id, captured_at DESC);
CREATE TABLE IF NOT EXISTS of_node_obs_frpc (
id INTEGER PRIMARY KEY AUTOINCREMENT,
node_id TEXT NOT NULL DEFAULT '',
captured_at DATETIME NOT NULL,
tunnel_status TEXT NOT NULL DEFAULT '',
connected_relays_count INTEGER NOT NULL DEFAULT 0,
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_obs_frpc_node ON of_node_obs_frpc (node_id, captured_at DESC);
-- +goose Down
DROP TABLE IF EXISTS of_node_obs_frpc;
DROP TABLE IF EXISTS of_node_obs_frps;
DROP TABLE IF EXISTS of_node_edge_health;
DROP TABLE IF EXISTS of_node_metric_snapshots;
DROP TABLE IF EXISTS w_user_access_logs;
DROP TABLE IF EXISTS of_node_access_logs;
@@ -0,0 +1,20 @@
-- +goose Up
-- 日志保留天数配置(business),替换旧的 database_auto_cleanup_* 键。
-- 默认统一 30 天:不继承旧键值,避免旧配置被静默带入导致日志被过度清理。
INSERT OR IGNORE INTO w_system_configs (key, value, type, visibility, description, created_at, updated_at)
SELECT 'log_retention_days_postgres', '30', 'business', 0, 'PostgreSQL 日志保留天数(访问日志与可观测统一)', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP
UNION ALL
SELECT 'log_retention_days_sqlite', '30', 'business', 0, 'SQLite 日志保留天数', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP
UNION ALL
SELECT 'log_retention_days_clickhouse', '30', 'business', 0, 'ClickHouse 日志保留天数', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP;
DELETE FROM w_system_configs WHERE key IN ('database_auto_cleanup_enabled', 'database_auto_cleanup_retention_days');
-- +goose Down
INSERT OR IGNORE INTO w_system_configs (key, value, type, visibility, description, created_at, updated_at)
SELECT 'database_auto_cleanup_enabled', 'true', 'business', 0, '数据库自动清理开关', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP
UNION ALL
SELECT 'database_auto_cleanup_retention_days', COALESCE((SELECT value FROM w_system_configs WHERE key = 'log_retention_days_sqlite'), '30'), 'business', 0, '数据库保留天数', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP;
DELETE FROM w_system_configs WHERE key IN ('log_retention_days_postgres', 'log_retention_days_sqlite', 'log_retention_days_clickhouse');
@@ -0,0 +1,17 @@
-- +goose Up
CREATE TABLE IF NOT EXISTS w_schedules_backup_of_database_auto_cleanup AS
SELECT * FROM w_schedules WHERE task_type = 'of_database_auto_cleanup';
DELETE FROM w_schedules WHERE task_type = 'of_database_auto_cleanup';
-- +goose Down
INSERT OR IGNORE INTO w_schedules (id, name, task_type, cron, payload, is_active, created_at, updated_at)
SELECT id, name, task_type, cron, payload, is_active, created_at, updated_at
FROM w_schedules_backup_of_database_auto_cleanup;
INSERT OR IGNORE INTO w_schedules (id, name, task_type, cron, payload, is_active, created_at, updated_at)
SELECT 102, 'OpenFlare 可观测数据自动清理', 'of_database_auto_cleanup', '0 3 * * *', '{}', 1, CURRENT_TIMESTAMP, CURRENT_TIMESTAMP
WHERE NOT EXISTS (SELECT 1 FROM w_schedules WHERE task_type = 'of_database_auto_cleanup');
DROP TABLE IF EXISTS w_schedules_backup_of_database_auto_cleanup;
@@ -0,0 +1,8 @@
-- +goose Up
-- 性能指标(CPU/内存/磁盘/网络)保留天数:三库共用、独立短留存(默认 3 天),
-- 不随 log_retention_days_*(访问日志保留配置)变化。
INSERT OR IGNORE INTO w_system_configs (key, value, type, visibility, description, created_at, updated_at)
VALUES ('metric_retention_days', '3', 'business', 0, '性能指标(CPU/内存/磁盘/网络)保留天数(三库共用)', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP);
-- +goose Down
DELETE FROM w_system_configs WHERE key = 'metric_retention_days';
@@ -0,0 +1,13 @@
-- +goose Up
-- 系统定期垃圾清理改为每日执行一次(凌晨 3 点,Asia/Shanghai)。
-- 此前为每 2 小时(0 */2 * * *)高频扫描,日常清理收益有限,降频减少非必要扫描。
UPDATE w_schedules
SET cron = '0 3 * * *',
updated_at = CURRENT_TIMESTAMP
WHERE task_type = 'system_cleanup';
-- +goose Down
UPDATE w_schedules
SET cron = '0 */2 * * *',
updated_at = CURRENT_TIMESTAMP
WHERE task_type = 'system_cleanup';
@@ -0,0 +1,9 @@
-- +goose Up
-- 日志列表默认排序与 retention 清理的时间范围条件:补 logged_at 前导索引。
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_logged_at ON of_node_access_logs (logged_at DESC, id DESC);
-- hosts 过滤为 lower(trim(host)) IN:补表达式索引,普通 (host,...) 索引无法命中函数表达式。
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_host_lower ON of_node_access_logs (lower(trim(host)));
-- +goose Down
DROP INDEX IF EXISTS idx_of_node_access_logs_host_lower;
DROP INDEX IF EXISTS idx_of_node_access_logs_logged_at;
@@ -19,12 +19,11 @@ import (
"gorm.io/gorm"
)
// expectedMigratedSystemConfigCount 包含初始 32 项系统配置、202606220004
// 从 of_options 迁移过来的 48 项业务配置、Pages 的 2 项业务配置、
// OpenResty 默认限流的 3 项业务配置、单 IP 请求频率限制 1 项业务配置、
// 源站错误页的 4 项业务配置,以及 Service Worker 离线兜底的 2 项业务配置、
// SW 离线兜底生效域名的 1 项业务配置。
const expectedMigratedSystemConfigCount = 93
// expectedMigratedSystemConfigCount 为全新库执行全部迁移后 w_system_configs 的行数
// (初始系统配置 + 各期配置迁移/新增 seed:of_options 迁移、文件白名单、磁盘缓存、
// 登录会话 TTL、升级源、存储、FRPS Web UI、Pages、OpenResty 限流、单 IP 限频、
// 错误页、SW 离线、日志保留期、指标保留期等);新增配置 seed 迁移时需同步更新本常量。
const expectedMigratedSystemConfigCount = 95
func TestMigrateInitializesSQLiteDatabase(t *testing.T) {
sqliteDB, err := gorm.Open(sqlite.Open(":memory:"), &gorm.Config{
@@ -0,0 +1,95 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package migrator
import (
"database/sql"
"fmt"
"os"
"strings"
"testing"
"time"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/glebarez/sqlite"
"github.com/pressly/goose/v3"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"gorm.io/driver/postgres"
"gorm.io/gorm"
)
const (
systemCleanupPreviousMigration = int64(202608090001)
systemCleanupMigration = int64(202608090002)
systemCleanupTaskType = "system_cleanup"
)
func TestSystemCleanupScheduleMigrationSQLite(t *testing.T) {
dbPath := t.TempDir() + "/system-cleanup-migration.db"
gormDB, err := gorm.Open(sqlite.Open(dbPath), &gorm.Config{
DisableForeignKeyConstraintWhenMigrating: true,
})
require.NoError(t, err)
sqlDB, err := gormDB.DB()
require.NoError(t, err)
sqlDB.SetMaxOpenConns(1)
t.Cleanup(func() { require.NoError(t, sqlDB.Close()) })
runSystemCleanupScheduleMigration(t, gormDB, sqlDB, dialectSqlite, "goose/sqlite")
}
func TestSystemCleanupScheduleMigrationPostgres(t *testing.T) {
dsn := strings.TrimSpace(os.Getenv("OPENFLARE_TEST_POSTGRES_DSN"))
if dsn == "" {
t.Skip("OPENFLARE_TEST_POSTGRES_DSN is not set")
}
gormDB, err := gorm.Open(postgres.Open(dsn), &gorm.Config{
DisableForeignKeyConstraintWhenMigrating: true,
})
require.NoError(t, err)
sqlDB, err := gormDB.DB()
require.NoError(t, err)
sqlDB.SetMaxOpenConns(1)
schema := fmt.Sprintf("system_cleanup_migration_%d", time.Now().UnixNano())
require.Regexp(t, `^[a-z0-9_]+$`, schema)
require.NoError(t, gormDB.Exec(`CREATE SCHEMA "`+schema+`"`).Error)
require.NoError(t, gormDB.Exec(`SET search_path TO "`+schema+`"`).Error)
t.Cleanup(func() {
assert.NoError(t, gormDB.Exec("SET search_path TO public").Error)
assert.NoError(t, gormDB.Exec(`DROP SCHEMA IF EXISTS "`+schema+`" CASCADE`).Error)
assert.NoError(t, sqlDB.Close())
})
runSystemCleanupScheduleMigration(t, gormDB, sqlDB, dialectPostgres, "goose/postgres")
}
func runSystemCleanupScheduleMigration(
t *testing.T,
gormDB *gorm.DB,
sqlDB *sql.DB,
dialect string,
dir string,
) {
t.Helper()
goose.SetBaseFS(migrationFS)
require.NoError(t, goose.SetDialect(dialect))
require.NoError(t, goose.UpTo(sqlDB, dir, systemCleanupPreviousMigration))
assertSystemCleanupCron(t, gormDB, "0 */2 * * *", "迁移前应为每 2 小时")
require.NoError(t, goose.UpTo(sqlDB, dir, systemCleanupMigration))
assertSystemCleanupCron(t, gormDB, "0 3 * * *", "迁移后应为每日凌晨 3 点")
require.NoError(t, goose.DownTo(sqlDB, dir, systemCleanupPreviousMigration))
assertSystemCleanupCron(t, gormDB, "0 */2 * * *", "回滚后恢复每 2 小时")
}
func assertSystemCleanupCron(t *testing.T, gormDB *gorm.DB, wantCron, msg string) {
t.Helper()
var schedule model.Schedule
require.NoError(t, gormDB.Where("task_type = ?", systemCleanupTaskType).First(&schedule).Error)
assert.Equal(t, wantCron, schedule.Cron, msg)
}
+4 -3
View File
@@ -10,6 +10,7 @@ import (
"github.com/Rain-kl/Wavelet/internal/apps/openflare"
cf "github.com/Rain-kl/Wavelet/internal/apps/openflare/cloudflare"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/pages"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/tasks"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/tls"
"github.com/Rain-kl/Wavelet/internal/apps/upload"
"github.com/Rain-kl/Wavelet/internal/apps/user"
@@ -44,15 +45,15 @@ func Register() {
task.RegisterHandler(openflare.SSLRenewTask, &openflare.SSLRenewHandler{})
task.RegisterTaskMeta(openflare.SSLRenewMeta)
task.RegisterHandler(openflare.DatabaseAutoCleanupTask, &openflare.DatabaseAutoCleanupHandler{})
task.RegisterTaskMeta(openflare.DatabaseAutoCleanupMeta)
task.RegisterHandler(openflare.WAFIPGroupSyncTask, &openflare.WAFIPGroupSyncHandler{})
task.RegisterTaskMeta(openflare.WAFIPGroupSyncMeta)
task.RegisterHandler(openflare.UptimeKumaSyncTask, &openflare.UptimeKumaSyncHandler{})
task.RegisterTaskMeta(openflare.UptimeKumaSyncMeta)
task.RegisterHandler(openflare.LogDBSwitchTask, &tasks.LogDBSwitchHandler{})
task.RegisterTaskMeta(openflare.LogDBSwitchMeta)
task.RegisterHandler(cf.SyncMemberTask, &cf.SyncMemberTaskHandler{})
task.RegisterTaskMeta(cf.SyncMemberMeta)
task.RegisterHandler(cf.SyncGroupTask, &cf.SyncGroupTaskHandler{})
+114
View File
@@ -0,0 +1,114 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
// Package analytics defines ClickHouse analytics domain models and query DTOs
// (pure data, no IO).
package analytics
import "time"
// AccessLogFilter scopes user access log queries.
// 单一权威字段集(CH 原字段,Task 1 迁入):禁止追加仅某实现使用的字段(避免双字段集分叉)。
type AccessLogFilter struct {
// UserIDs filters by user IDs. nil means no user filter; an empty slice means no matches.
UserIDs []uint64
Path string
// StartTime filters created_at >= StartTime when non-nil.
StartTime *time.Time
// EndTime filters created_at <= EndTime when non-nil(闭区间,与 CH/GORM 实现一致)。
EndTime *time.Time
}
// NodeAccessLogFilter scopes ClickHouse node access log queries.
type NodeAccessLogFilter struct {
NodeID string
RemoteAddr string
Host string
// Hosts exact-matches any host (case-insensitive). Prefer over Host for multi-domain scopes.
Hosts []string
Path string
Since time.Time
Until time.Time
Page int
PageSize int
SortBy string
SortOrder string
}
// NodeObservabilityFilter scopes ClickHouse node observability queries.
type NodeObservabilityFilter struct {
NodeID string
Since time.Time
Limit int
}
// DailyTrend is a single day's access count.
type DailyTrend struct {
Date string
Count uint64
}
// BrowserShare is a browser group's share of access logs.
type BrowserShare struct {
Browser string
Count uint64
}
// TopUser is an active user ranked by access count.
type TopUser struct {
UserID uint64
Count uint64
}
// NodeAccessLogRegionCount aggregates access log regions.
type NodeAccessLogRegionCount struct {
Region string
Count int64
}
// NodeAccessLogTrafficSummary is a window-level access log traffic summary.
type NodeAccessLogTrafficSummary struct {
RequestCount int64
ErrorCount int64
UniqueIPCount int64
BytesSent int64
RequestLength int64
NodeCount int64
}
// NodeAccessLogValueCount is a grouped value count (status_code, host, ...).
type NodeAccessLogValueCount struct {
Value string
Count int64
}
// NodeAccessLogNodeAggregate is per-node traffic over a window.
type NodeAccessLogNodeAggregate struct {
NodeID string
RequestCount int64
ErrorCount int64
UniqueIPCount int64
}
// BatchWriterStats is a point-in-time snapshot of a batch writer queue and failure counters.
type BatchWriterStats struct {
Name string `json:"name"`
Depth int `json:"depth"`
Cap int `json:"cap"`
Drops int64 `json:"drops"`
FlushErrors int64 `json:"flush_errors"`
Running bool `json:"running"`
}
// ClickHouseOperationalStats summarizes ClickHouse merge/mutation pressure
// and in-process batch writer queue health.
type ClickHouseOperationalStats struct {
Database string `json:"database"`
ActiveParts int64 `json:"active_parts"`
TotalRows int64 `json:"total_rows"`
PendingMutations int64 `json:"pending_mutations"`
AsyncInsertQueue int64 `json:"async_insert_queue"`
AsyncInsertBytes int64 `json:"async_insert_bytes"`
// BatchWriters reports in-process queue depth/drops/flush errors for CH writers.
BatchWriters []BatchWriterStats `json:"batch_writers,omitempty"`
}
@@ -90,6 +90,34 @@ type AccessLogHourly struct {
RequestLength int64 `gorm:"column:request_length"`
}
// NodeTrafficHourly is an hourly traffic rollup row.
//
// UniqueVisitorCount is always 0 when sourced from of_access_log_hourly
// (true UV requires raw uniqExact on access logs).
type NodeTrafficHourly struct {
NodeID string
Hour time.Time
RequestCount int64
ErrorCount int64
UniqueVisitorCount int64
}
// NodeMetricHourly is an hourly metric snapshot aggregation row.
//
// Disk and host network counters are cumulative. Prefer pre-aggregated min/max
// deltas from of_node_metric_capacity_hourly; raw fallback uses consecutive
// lagInFrame samples per node (negative deltas after counter reset are dropped).
type NodeMetricHourly struct {
Hour time.Time
AverageCPUUsagePercent float64
AverageMemoryUsagePercent float64
NetworkRxBytes int64
NetworkTxBytes int64
DiskReadBytes int64
DiskWriteBytes int64
ReportedNodes int
}
// NodeObsFrps stores FRPS observability snapshots in ClickHouse.
type NodeObsFrps struct {
ID uint64 `gorm:"column:id"`
@@ -1,7 +1,6 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
// Package analytics defines ClickHouse analytics domain models.
package analytics
import (
+124
View File
@@ -0,0 +1,124 @@
// Copyright 2026 Arctel.net
// SPDX-License-Identifier: Apache-2.0
package analytics
import "strings"
// User-Agent 浏览器/OS/设备分类(纯函数,无 IO)。
// 与 internal/repository/analytics/browser.go 的判定逻辑保持一致(Task 4 复制,
// 因为 model 不得 import analyticsrepo);后续若移除旧 CH 实现,可让 analyticsrepo 改以别名复用本包。
const (
uaLabelUnknown = "Unknown"
uaLabelBot = "Bot"
uaLabelOther = "Other"
uaTokenBot = "bot"
uaTokenAndroid = "android"
uaTokenSpider = "spider"
uaTokenCrawler = "crawler"
)
type uaMatchRule struct {
label string
contains []string
allOf []string
noneOf []string
}
func matchUARules(uaLower string, rules []uaMatchRule, fallback string) string {
if uaLower == "" {
return uaLabelUnknown
}
for _, rule := range rules {
matched := false
for _, token := range rule.contains {
if strings.Contains(uaLower, token) {
matched = true
break
}
}
if !matched && len(rule.allOf) > 0 {
matched = true
for _, token := range rule.allOf {
if !strings.Contains(uaLower, token) {
matched = false
break
}
}
}
if !matched {
continue
}
excluded := false
for _, token := range rule.noneOf {
if strings.Contains(uaLower, token) {
excluded = true
break
}
}
if excluded {
continue
}
return rule.label
}
return fallback
}
var browserRules = []uaMatchRule{
{label: "WeChat", contains: []string{"micromessenger"}},
{label: "Postman", contains: []string{"postman"}},
{label: "CLI", contains: []string{"curl/", "wget/"}},
{label: "Edge", contains: []string{"edg/", "edgios/", "edga/"}},
{label: "Opera", contains: []string{"opr/", "opera"}},
{label: "Firefox", contains: []string{"firefox", "fxios"}},
{label: "Chrome", contains: []string{"crios", "chrome"}, noneOf: []string{"chromium"}},
{label: "Chromium", contains: []string{"chromium"}},
{label: "Safari", contains: []string{"safari"}},
{label: uaLabelBot, contains: []string{uaTokenBot, uaTokenSpider, uaTokenCrawler, "slurp"}},
}
var osRules = []uaMatchRule{
{label: "Android", contains: []string{uaTokenAndroid}},
{label: "iOS", contains: []string{"iphone", "ipad", "ipod", "ios"}},
{label: "Windows", contains: []string{"windows"}},
{label: "macOS", contains: []string{"mac os x", "macintosh", "macos"}},
{label: "Chrome OS", contains: []string{"cros"}},
{label: "Linux", contains: []string{"linux"}},
{label: uaLabelBot, contains: []string{uaTokenBot, uaTokenSpider, uaTokenCrawler}},
}
var deviceRules = []uaMatchRule{
{
label: uaLabelBot,
contains: []string{uaTokenBot, uaTokenSpider, uaTokenCrawler, "slurp", "curl/", "wget/", "python-requests", "go-http-client", "postman"},
},
{
label: "Tablet",
contains: []string{"ipad", "tablet"},
},
{
label: "Tablet",
allOf: []string{uaTokenAndroid},
noneOf: []string{"mobile"},
},
{
label: "Mobile",
contains: []string{"mobi", "iphone", "ipod", uaTokenAndroid},
},
}
// ParseBrowserName performs lightweight User-Agent browser identification.
func ParseBrowserName(ua string) string {
return matchUARules(strings.ToLower(ua), browserRules, uaLabelOther)
}
// ParseOSName performs lightweight User-Agent OS identification.
func ParseOSName(ua string) string {
return matchUARules(strings.ToLower(ua), osRules, uaLabelOther)
}
// ParseDeviceType performs lightweight User-Agent device type identification.
func ParseDeviceType(ua string) string {
return matchUARules(strings.ToLower(ua), deviceRules, "Desktop")
}
+18 -8
View File
@@ -41,14 +41,12 @@ const (
ConfigKeyRelayFRPSWebUIPort = "relay_frps_web_ui_port" // FRPS 内置 Web 界面端口
// OpenFlare 业务配置(从 of_options 迁移)
ConfigKeyAgentDiscoveryToken = "agent_discovery_token" //nolint:gosec // false positive: config key name. Agent 发现令牌
ConfigKeyAgentHeartbeatInterval = "agent_heartbeat_interval" // Agent 心跳间隔(毫秒)
ConfigKeyAgentWebsocketUpgradeEnabled = "agent_websocket_upgrade_enabled" // Agent WebSocket 升级开关
ConfigKeyNodeOfflineThreshold = "node_offline_threshold" // 节点离线阈值(毫秒)
ConfigKeyAgentUpdateRepo = "agent_update_repo" // Agent 更新仓库
ConfigKeyGeoIPProvider = "geoip_provider" // GeoIP 服务商
ConfigKeyDatabaseAutoCleanupEnabled = "database_auto_cleanup_enabled" // 数据库自动清理开关
ConfigKeyDatabaseAutoCleanupRetentionDays = "database_auto_cleanup_retention_days" // 数据库保留天数
ConfigKeyAgentDiscoveryToken = "agent_discovery_token" //nolint:gosec // false positive: config key name. Agent 发现令牌
ConfigKeyAgentHeartbeatInterval = "agent_heartbeat_interval" // Agent 心跳间隔(毫秒)
ConfigKeyAgentWebsocketUpgradeEnabled = "agent_websocket_upgrade_enabled" // Agent WebSocket 升级开关
ConfigKeyNodeOfflineThreshold = "node_offline_threshold" // 节点离线阈值(毫秒)
ConfigKeyAgentUpdateRepo = "agent_update_repo" // Agent 更新仓库
ConfigKeyGeoIPProvider = "geoip_provider" // GeoIP 服务商
// Pages 静态托管配置
ConfigKeyPagesMaxPackageSizeMB = "pages_max_package_size_mb" // Pages 部署包上传大小上限(MiB)
@@ -122,6 +120,18 @@ const (
ConfigKeySWOfflineDomains = "sw_offline_domains" // 离线兜底生效域名列表(JSON 数组,空则仅总开关无效)
)
// 日志数据库解耦
const (
ConfigKeyLogDatabase = "log_database" // 当前日志主库:postgres|sqlite|clickhouse(仅迁移任务写入)
ConfigKeyLogDBMigration = "log_db_migration" // 迁移冻结标记:"migrating" 或空
ConfigKeyLogRetentionDaysPostgres = "log_retention_days_postgres" // PostgreSQL 日志保留天数
ConfigKeyLogRetentionDaysSQLite = "log_retention_days_sqlite" // SQLite 日志保留天数
ConfigKeyLogRetentionDaysClickHouse = "log_retention_days_clickhouse" // ClickHouse 日志保留天数
// ConfigKeyMetricRetentionDays 性能指标(CPU/内存/磁盘/网络)保留天数,三库共用;
// 性能数据价值衰减快,默认短留存(3 天),不随访问日志保留配置。
ConfigKeyMetricRetentionDays = "metric_retention_days"
)
const (
// ConfigVisibilityHidden 表示配置不通过公共配置接口暴露
ConfigVisibilityHidden = 0
+63 -2
View File
@@ -7,18 +7,24 @@ package bootstrap
import (
"context"
"errors"
"fmt"
"log"
"sync"
admin_push "github.com/Rain-kl/Wavelet/internal/apps/admin/push"
"github.com/Rain-kl/Wavelet/internal/apps/admin/push/custom_events"
"github.com/Rain-kl/Wavelet/internal/apps/openflare/chwriter"
ofgeoip "github.com/Rain-kl/Wavelet/internal/apps/openflare/geoip"
"github.com/Rain-kl/Wavelet/internal/apps/risk_control"
"github.com/Rain-kl/Wavelet/internal/infra/config"
taskhandlers "github.com/Rain-kl/Wavelet/internal/infra/task/handlers"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/Rain-kl/Wavelet/internal/platform/lifecycle"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
"github.com/Rain-kl/Wavelet/pkg/cache/ram"
"github.com/Rain-kl/Wavelet/pkg/logger"
"gorm.io/gorm"
)
// Options selects role-specific runtime bootstrap steps for the current process.
@@ -129,6 +135,21 @@ func RegisterAll() {
// Call from cmd entry points after wiring registration and database migration, not from router.
func Init(ctx context.Context, opts Options) {
initRuntimeOnce.Do(func() {
if err := validateAndSeedLogDatabase(ctx); err != nil {
logger.ErrorF(ctx, "[Bootstrap] 日志主库配置校验失败: %v", err)
log.Fatalf("[Bootstrap] 日志主库配置校验失败: %v", err)
}
// 注入 logstore 配置读取(避免 logstore ↔ repository 循环依赖),并预热激活 store。
logstore.SetConfigReader(func(ctx context.Context, key string) (string, error) {
cfg, err := repository.GetSystemConfigByKey(ctx, key)
if err != nil {
return "", err
}
return cfg.Value, nil
})
logstore.Init(ctx)
// Register config cache loader
RegisterCache(repository.ConfigCacheType, CacheRegistry{
Loader: repository.ConfigLoader{},
@@ -146,12 +167,52 @@ func Init(ctx context.Context, opts Options) {
logger.ErrorF(ctx, "[Bootstrap] sync push events failed: %v", err)
}
if opts.API {
risk_control.InitLogWriter(ctx)
chwriter.Init(ctx)
}
})
}
// validateAndSeedLogDatabase 校验日志主库标记与运行配置的一致性,首次启动 seed。
func validateAndSeedLogDatabase(ctx context.Context) error {
cfg, err := repository.GetSystemConfigByKey(ctx, model.ConfigKeyLogDatabase)
if err != nil && !errors.Is(err, gorm.ErrRecordNotFound) {
return fmt.Errorf("读取日志主库配置失败: %w", err)
}
current := cfg.Value
if current == "" {
// 首次启动 seed:CH 启用 → clickhouse;否则随主库。
current = "sqlite"
if config.Config.Database.Enabled {
current = "postgres"
}
if config.Config.ClickHouse.Enabled {
current = "clickhouse"
}
// 行缺失时 UpdateSystemConfigFields 仅为 UPDATE 无法插入,改用可创建可更新的 SaveOrUpdateSystemConfig。
if err := repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyLogDatabase, current); err != nil {
return fmt.Errorf("初始化日志主库配置失败: %w", err)
}
return nil
}
switch current {
case "clickhouse":
if !config.Config.ClickHouse.Enabled {
return errors.New("当前日志主库为 ClickHouse 但 ClickHouse 未启用。请先重新启用 ClickHouse 配置并启动,在任务管理运行『切换日志数据库』迁移到 PostgreSQL/SQLite 后再禁用 ClickHouse")
}
case "postgres":
if !config.Config.Database.Enabled {
return errors.New("当前日志主库为 PostgreSQL 但 PostgreSQL 未启用(当前为 SQLite 主库)。请运行『切换日志数据库』迁回 SQLite 或启用 PostgreSQL")
}
case "sqlite":
if config.Config.Database.Enabled {
return errors.New("当前日志主库为 SQLite 但当前主库为 PostgreSQL。请运行『切换日志数据库』迁移到 PostgreSQL")
}
default:
return fmt.Errorf("未知的日志主库配置: %s", current)
}
return nil
}
// Stop stops all batch writers and background resources.
func Stop(ctx context.Context) {
lifecycle.Stop(ctx)
@@ -5,10 +5,15 @@ package bootstrap
import (
"context"
"strings"
"testing"
admin_push "github.com/Rain-kl/Wavelet/internal/apps/admin/push"
"github.com/Rain-kl/Wavelet/internal/infra/config"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
"github.com/Rain-kl/Wavelet/internal/testhelper"
)
@@ -50,3 +55,170 @@ func TestInitSyncsPushEventsOnce(t *testing.T) {
t.Fatalf("admin_login name = %q, want %q", adminLogin.Name, "管理员登录")
}
}
func TestValidateAndSeedLogDatabaseSeedsDefault(t *testing.T) {
_, _, cleanup := testhelper.SetupTestEnvironment(t)
defer cleanup()
prevDB := config.Config.Database.Enabled
prevCH := config.Config.ClickHouse.Enabled
t.Cleanup(func() {
config.Config.Database.Enabled = prevDB
config.Config.ClickHouse.Enabled = prevCH
})
tests := []struct {
name string
dbEnabled bool
chEnabled bool
want string
}{
{name: "sqlite default", dbEnabled: false, chEnabled: false, want: "sqlite"},
{name: "postgres default", dbEnabled: true, chEnabled: false, want: "postgres"},
{name: "clickhouse default", dbEnabled: true, chEnabled: true, want: "clickhouse"},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
config.Config.Database.Enabled = tt.dbEnabled
config.Config.ClickHouse.Enabled = tt.chEnabled
ctx := context.Background()
// 清掉标记行,模拟首次启动。
if err := db.DB(ctx).Where("key = ?", model.ConfigKeyLogDatabase).Delete(&model.SystemConfig{}).Error; err != nil {
t.Fatalf("delete log_database marker failed: %v", err)
}
repository.ResetSystemConfigRAMCacheForTest()
if err := validateAndSeedLogDatabase(ctx); err != nil {
t.Fatalf("validateAndSeedLogDatabase() error = %v", err)
}
repository.ResetSystemConfigRAMCacheForTest()
cfg, err := repository.GetSystemConfigByKey(ctx, model.ConfigKeyLogDatabase)
if err != nil {
t.Fatalf("GetSystemConfigByKey(%s) error = %v", model.ConfigKeyLogDatabase, err)
}
if cfg.Value != tt.want {
t.Fatalf("seeded log_database = %q, want %q", cfg.Value, tt.want)
}
})
}
}
func TestValidateAndSeedLogDatabaseUpdatesEmptyMarker(t *testing.T) {
_, _, cleanup := testhelper.SetupTestEnvironment(t)
defer cleanup()
dbPrev := config.Config.Database.Enabled
chPrev := config.Config.ClickHouse.Enabled
config.Config.Database.Enabled = true
config.Config.ClickHouse.Enabled = false
t.Cleanup(func() {
config.Config.Database.Enabled = dbPrev
config.Config.ClickHouse.Enabled = chPrev
})
ctx := context.Background()
// 标记行已存在但值为空,等同首次启动,应写入默认值(走更新路径)。
if err := repository.CreateSystemConfig(ctx, &model.SystemConfig{Key: model.ConfigKeyLogDatabase, Value: "", Type: "system"}); err != nil {
t.Fatalf("create empty marker failed: %v", err)
}
repository.ResetSystemConfigRAMCacheForTest()
if err := validateAndSeedLogDatabase(ctx); err != nil {
t.Fatalf("validateAndSeedLogDatabase() error = %v", err)
}
repository.ResetSystemConfigRAMCacheForTest()
cfg, err := repository.GetSystemConfigByKey(ctx, model.ConfigKeyLogDatabase)
if err != nil {
t.Fatalf("GetSystemConfigByKey(%s) error = %v", model.ConfigKeyLogDatabase, err)
}
want := "postgres"
if cfg.Value != want {
t.Fatalf("log_database = %q, want %q", cfg.Value, want)
}
}
func TestValidateAndSeedLogDatabaseRejectsInconsistentConfig(t *testing.T) {
_, _, cleanup := testhelper.SetupTestEnvironment(t)
defer cleanup()
prevDB := config.Config.Database.Enabled
prevCH := config.Config.ClickHouse.Enabled
t.Cleanup(func() {
config.Config.Database.Enabled = prevDB
config.Config.ClickHouse.Enabled = prevCH
})
seedMarker := func(t *testing.T, value string) {
t.Helper()
ctx := context.Background()
if err := db.DB(ctx).Where("key = ?", model.ConfigKeyLogDatabase).Delete(&model.SystemConfig{}).Error; err != nil {
t.Fatalf("delete log_database marker failed: %v", err)
}
if err := repository.CreateSystemConfig(ctx, &model.SystemConfig{Key: model.ConfigKeyLogDatabase, Value: value, Type: "system"}); err != nil {
t.Fatalf("create log_database marker failed: %v", err)
}
repository.ResetSystemConfigRAMCacheForTest()
}
tests := []struct {
name string
marker string
dbEnabled bool
chEnabled bool
wantErr string
}{
{name: "clickhouse marker but disabled", marker: "clickhouse", dbEnabled: true, chEnabled: false, wantErr: "ClickHouse 未启用"},
{name: "postgres marker but disabled", marker: "postgres", dbEnabled: false, chEnabled: false, wantErr: "PostgreSQL 未启用"},
{name: "sqlite marker but postgres primary", marker: "sqlite", dbEnabled: true, chEnabled: false, wantErr: "SQLite"},
{name: "unknown marker", marker: "mysql", dbEnabled: false, chEnabled: false, wantErr: "未知的日志主库配置"},
{name: "consistent sqlite", marker: "sqlite", dbEnabled: false, chEnabled: false, wantErr: ""},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
config.Config.Database.Enabled = tt.dbEnabled
config.Config.ClickHouse.Enabled = tt.chEnabled
seedMarker(t, tt.marker)
err := validateAndSeedLogDatabase(context.Background())
if tt.wantErr == "" {
if err != nil {
t.Fatalf("validateAndSeedLogDatabase() error = %v, want nil", err)
}
return
}
if err == nil {
t.Fatalf("validateAndSeedLogDatabase() = nil, want error containing %q", tt.wantErr)
}
if !strings.Contains(err.Error(), tt.wantErr) {
t.Fatalf("validateAndSeedLogDatabase() error = %q, want contains %q", err.Error(), tt.wantErr)
}
})
}
}
func TestInitWiresLogstoreConfigReader(t *testing.T) {
ResetInitRuntimeOnceForTest()
t.Cleanup(ResetInitRuntimeOnceForTest)
logstore.ResetForTest()
t.Cleanup(logstore.ResetForTest)
_, _, cleanup := testhelper.SetupTestEnvironment(t)
defer cleanup()
ctx := context.Background()
// 插入迁移标记:bootstrap 注入的 reader 应能经 repository 读到该值(区分未装配时的兜底行为)。
if err := repository.CreateSystemConfig(ctx, &model.SystemConfig{Key: model.ConfigKeyLogDBMigration, Value: "migrating", Type: "system"}); err != nil {
t.Fatalf("create log_db_migration marker failed: %v", err)
}
repository.ResetSystemConfigRAMCacheForTest()
Init(ctx, Options{})
repository.ResetSystemConfigRAMCacheForTest()
if !logstore.Migrating(ctx) {
t.Fatal("logstore config reader not wired after bootstrap.Init: Migrating() = false, want true")
}
}
@@ -38,6 +38,18 @@ func CountAccessLogs(ctx context.Context, filter AccessLogFilter) (uint64, error
return count, nil
}
// DeleteAllUserAccessLogs hard-deletes all user access logs via TRUNCATE.
func DeleteAllUserAccessLogs(ctx context.Context) (int64, error) {
if err := userAccessLogConn(); err != nil {
return 0, err
}
outcome, err := truncateClickHouseTable(ctx, db.ChConn, analyticsmodel.UserAccessLog{}.TableName())
if err != nil {
return 0, err
}
return outcome.DeletedCount, nil
}
// ListAccessLogs returns paginated access logs and the total match count.
func ListAccessLogs(ctx context.Context, filter AccessLogFilter, page, pageSize int) ([]analyticsmodel.UserAccessLog, uint64, error) {
clause, args, ok := buildUserAccessLogFilterClause(filter)
@@ -6,21 +6,14 @@ package analytics
import (
"fmt"
"strings"
"time"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
)
const userAccessLogFilterClauseCapacity = 4
// AccessLogFilter scopes ClickHouse user access log queries.
type AccessLogFilter struct {
// UserIDs filters by user IDs. nil means no user filter; an empty slice means no matches.
UserIDs []uint64
Path string
// StartTime filters created_at >= StartTime when non-nil.
StartTime *time.Time
// EndTime filters created_at <= EndTime when non-nil.
EndTime *time.Time
}
type AccessLogFilter = analyticsmodel.AccessLogFilter
func buildUserAccessLogFilterClause(filter AccessLogFilter) (string, []any, bool) {
if filter.UserIDs != nil && len(filter.UserIDs) == 0 {
@@ -16,22 +16,13 @@ import (
const hoursInDay = 24
// DailyTrend is a single day's access count.
type DailyTrend struct {
Date string
Count uint64
}
type DailyTrend = analyticsmodel.DailyTrend
// BrowserShare is a browser group's share of access logs.
type BrowserShare struct {
Browser string
Count uint64
}
type BrowserShare = analyticsmodel.BrowserShare
// TopUser is an active user ranked by access count.
type TopUser struct {
UserID uint64
Count uint64
}
type TopUser = analyticsmodel.TopUser
// GetDailyTrend returns per-day access counts for the last days days (inclusive of today).
func GetDailyTrend(ctx context.Context, days int) ([]DailyTrend, error) {

Some files were not shown because too many files have changed in this diff Show More