113 KiB
日志数据库解耦(ClickHouse 可选化)实现计划
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: 让日志/分析存储从 ClickHouse 解耦——新增 internal/repository/logstore 抽象(PG/SQLite 用 GORM、CH 用现有原生优化),ClickHouse 变为可选;提供「切换日志数据库」迁移任务与按库保留时间配置。
Architecture: repository 层导出接口 + 配置驱动 provider(log_database 系统配置决定激活实现);apps 只面向 logstore/repository 公开函数,import-lint 测试强制约束;CH 实现包住现有 analyticsrepo(零性能损耗);PG/SQLite 共用一套 GORM 实现(方言 SQL 拆 dialect_* 小文件)。
Tech Stack: Go 1.25+、GORM、PostgreSQL/SQLite(主库 goose 双方言)、ClickHouse(原生 driver + 单方言 goose)、Asynq 任务框架、Next.js/TypeScript/shadcn。
Global Constraints
- 模块:
github.com/Rain-kl/Wavelet;Go 1.25.7。 - 分层:
apps → repository → model;model禁止 importrepository/db;pkg/util/禁止 Gin/GORM/sessions。 - 路由仅注册于
internal/router/router.go;Serve()禁止进程级初始化。 - 迁移:PG/SQLite 双方言同版本号 goose SQL(
internal/infra/persistence/migrator/goose/{postgres,sqlite});CH 单方言(goose/clickhouse);禁止 GORM AutoMigrate(生产;单测可用 sqlite AutoMigrate 建测试表)。 - 任务/推送注册:
bootstrap.RegisterTasks()等显式装配,禁止init()注册跨模块集成。 - API 错误:
response.Abort*+ErrorHandlerMiddleware;禁止 Handler 直接c.JSON(..., response.Err(...))。 - 系统配置:key 常量在
internal/model/system_configs.go;值存字符串;type∈ {system,business};visibility0/1;goose 双方言 seed。 - 前端:shadcn
variant+ CSS 变量;页面根w-full;标题h1 text-2xl font-semibold tracking-tight;service 继承BaseService,回调用箭头函数。 - 日志库合法状态:
log_database∈ {postgres,sqlite,clickhouse},且postgres仅当database.enabled、sqlite仅当!database.enabled、clickhouse仅当clickhouse.enabled。 - 完成标准:
go test ./...、make swagger(API 变更时)、make code-check、make format;goose 三套空库 Up 全量通过。
里程碑与文件总览
| 文件 | 职责 |
|---|---|
internal/model/analytics/filter.go(新) |
从 analyticsrepo 迁入的过滤/结果 DTO(纯数据) |
internal/model/system_configs.go |
新增 ConfigKeyLogDatabase、ConfigKeyLogDBMigration、ConfigKeyLogRetentionDaysPostgres/SQLite/ClickHouse |
internal/repository/logstore/logstore.go(新) |
导出接口 + Store 结构体 + ErrMigrating |
internal/repository/logstore/provider.go(新) |
Init(ctx)/Active(ctx)/Migrating(ctx)/Reload/测试注入 |
internal/repository/logstore/postgres_store.go(新) |
GORM 实现(PG/SQLite 共用) |
internal/repository/logstore/dialect_postgres.go、dialect_sqlite.go(新) |
方言 SQL 片段 |
internal/repository/logstore/clickhouse_store.go(新) |
CH 实现(委托 analyticsrepo) |
internal/repository/logstore/hooks.go(新) |
AccessLogInsertHooks/ObservabilityInsertHooks 注册表(从 repository 迁入) |
internal/repository/logstore/imports_test.go(新) |
import-lint 测试 |
internal/repository/openflare_access_log_store.go、openflare_observability_store.go |
删除(被 logstore 吸收) |
internal/repository/openflare_access_log.go、openflare_observability.go |
改为一行委托 logstore |
internal/apps/risk_control/logics.go、internal/apps/openflare/chwriter/writer.go |
flush func 与入口改为 logstore;冻结检查 |
internal/apps/openflare/tasks/database_cleanup.go |
清理逻辑迁入 system_cleanup;任务下线 |
internal/apps/admin/logs/routers.go、internal/apps/admin/status/clickhouse.go |
改走 logstore;状态端点改造 |
internal/apps/upload/task/cleanup.go |
新增日志清理步骤 |
internal/apps/openflare/async_tasks.go、internal/infra/task/handlers/register.go |
注册「切换日志数据库」任务;下线清理任务 |
internal/apps/openflare/tasks/log_db_switch.go(新) |
迁移任务 Handler |
internal/platform/bootstrap/bootstrap.go |
启动校验 + logstore 初始化 |
internal/infra/config/model.go |
(无新启动配置;校验仅用现有字段) |
goose:postgres/20260808NNNN_create_log_tables.sql、sqlite/20260808NNNN_create_log_tables.sql |
6 张原始日志表(PG 分区) |
goose:postgres/20260808NNNN_log_retention_configs.sql、sqlite/... |
保留配置 + 旧 key 下线 |
goose:postgres/20260808NNNN_drop_database_cleanup_schedule.sql、sqlite/... |
下线 of_database_auto_cleanup schedule |
internal/apps/admin/system_config/routers.go、internal/apps/openflare/option/validate.go |
log_database/log_db_migration key 保护 |
frontend/... |
任务管理页日志库状态、业务配置「日志保留时间」分组 |
docs/changelog/index.md |
[Unreleased] 中文条目 |
M1:抽象层与主库日志读写
Task 1: DTO 类型迁入 model/analytics
Files:
- Create:
internal/model/analytics/filter.go - Modify:
internal/repository/analytics/access_log.go、node_access_log.go、node_observability.go、access_log_stats.go、node_access_log_stats.go、node_observability_delete.go等(删除本地类型定义,改 import model/analytics) - Test:
internal/model/analytics/filter_test.go
Interfaces:
-
Consumes: 现有 analyticsrepo 包内类型定义位置。
-
Produces:
analyticsmodel.AccessLogFilter、analyticsmodel.NodeAccessLogFilter、analyticsmodel.NodeObservabilityFilter、analyticsmodel.DailyTrend、analyticsmodel.BrowserShare、analyticsmodel.TopUser、analyticsmodel.NodeAccessLogRegionCount、analyticsmodel.NodeAccessLogTrafficSummary、analyticsmodel.NodeAccessLogValueCount、analyticsmodel.NodeAccessLogNodeAggregate(字段逐一从 analyticsrepo 原定义复制)。 -
Step 1: 在
internal/model/analytics/filter.go定义迁移类型
// Package analytics 定义分析域模型与查询 DTO(纯数据,无 IO)。
package analytics
import "time"
// AccessLogFilter 用户访问日志查询条件。
type AccessLogFilter struct {
UserID uint64
Path string
Method string
IP string
Status int32
Since time.Time
Until time.Time
Page int
PageSize int
}
// NodeAccessLogFilter 节点访问日志查询条件。
type NodeAccessLogFilter struct {
NodeID string
RemoteAddr string
Host string
Hosts []string
Path string
Since time.Time
Until time.Time
Page int
PageSize int
SortBy string
SortOrder string
}
// NodeObservabilityFilter 可观测查询条件。
type NodeObservabilityFilter struct {
NodeID string
Since time.Time
Limit int
}
// DailyTrend 每日访问趋势。
type DailyTrend struct {
Date string
Cnt uint64
}
// BrowserShare 浏览器占比。
type BrowserShare struct {
Browser string
Cnt uint64
}
// TopUser 活跃用户排行。
type TopUser struct {
UserID uint64
Cnt uint64
}
// NodeAccessLogRegionCount 地区访问计数。
type NodeAccessLogRegionCount struct {
Region string
Count uint64
}
// NodeAccessLogTrafficSummary 流量汇总。
type NodeAccessLogTrafficSummary struct {
RequestCount uint64
ErrorCount uint64
UniqueIPCount uint64
BytesSent uint64
RequestLength uint64
NodeCount uint64
}
// NodeAccessLogValueCount 维度值计数。
type NodeAccessLogValueCount struct {
Value string
Count uint64
}
// NodeAccessLogNodeAggregate 按节点聚合。
type NodeAccessLogNodeAggregate struct {
NodeID string
RequestCount uint64
ErrorCount uint64
UniqueIPCount uint64
}
注意:以上字段必须与
internal/repository/analytics/中同名类型逐字段一致(比对access_log.go、node_access_log.go、node_access_log_stats.go、access_log_stats.go)。若原类型字段与这里不同,以原类型为准修改本文件,保持语义不变。
- Step 2: 让 analyticsrepo 使用新类型——在每个原类型定义处删除定义,替换为类型别名,保证包内调用点零改动:
// internal/repository/analytics/access_log.go 顶部
import analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
type AccessLogFilter = analyticsmodel.AccessLogFilter
对 NodeAccessLogFilter、NodeObservabilityFilter、DailyTrend、BrowserShare、TopUser、NodeAccessLogRegionCount、NodeAccessLogTrafficSummary、NodeAccessLogValueCount、NodeAccessLogNodeAggregate、ClickHouseOperationalStats(及 ClickHouseOperationalStats 的字段结构体,含 BatchWriters []batchwriter.Stats)同样处理(原类型定义删除,替换为别名)。ClickHouseOperationalStats 迁入 model/analytics 后,logstore 状态接口可直接引用,CH 实现仍由 analyticsrepo 填充。
- Step 3: 编译验证 运行
go build ./internal/...,确认无重定义/未使用错误。 - Step 4: 提交
git add internal/model/analytics/filter.go internal/repository/analytics/ && git commit -m "refactor(analytics): move filter/result DTOs to model/analytics"
Task 2: logstore 接口与 provider 骨架
Files:
- Create:
internal/repository/logstore/logstore.go - Create:
internal/repository/logstore/provider.go - Create:
internal/repository/logstore/provider_test.go
Interfaces:
-
Consumes:
analyticsmodel.*DTO(Task 1)、model.ConfigKeyLogDatabase/ConfigKeyLogDBMigration(Task 8 定义,本任务先用字符串常量占位并加注释)、db.DB(ctx)(internal/infra/persistence的 GORM 句柄)、repository.GetSystemConfigByKey。 -
Produces: 接口
AccessLogStore/ObservabilityStore/UserAccessLogStore、结构体Store、ErrMigrating、Init(ctx)/Active(ctx)/Migrating(ctx)/ResetForTest。 -
Step 1: 写接口与
Store结构体(logstore.go)
// Package logstore 提供日志/分析存储抽象:上层只面向本包接口,
// 禁止直接 import internal/repository/analytics 或触碰 db.ChConn/db.ChDB。
package logstore
import (
"context"
"errors"
"time"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
)
// ErrMigrating 表示日志数据库正在迁移,当前禁止写入。
var ErrMigrating = errors.New("log database is migrating, writes are disabled")
// AccessLogStore 节点访问日志(of_node_access_logs)。
type AccessLogStore interface {
// InsertBatch 为写入入口:冻结检查 + 经 hook 入队(异步),不直接落库。
InsertBatch(ctx context.Context, records []*model.OpenFlareAccessLog) error
// BatchInsertNodeAccessLogs 为 batchwriter flush 目标:直接批量写入当前存储。
BatchInsertNodeAccessLogs(ctx context.Context, rows []analyticsmodel.NodeAccessLog) error
List(ctx context.Context, query model.OpenFlareAccessLogQuery) ([]*model.OpenFlareAccessLog, error)
Count(ctx context.Context, query model.OpenFlareAccessLogQuery) (int64, int64, int64, error)
RegionCounts(ctx context.Context, nodeID string, since time.Time, limit int) ([]*model.OpenFlareAccessLogRegionCount, error)
BucketAggregates(ctx context.Context, filter model.OpenFlareAccessLogQuery, bucketSeconds int64) ([]analyticsmodel.NodeAccessLogBucketAggregate, error)
CountBuckets(ctx context.Context, filter model.OpenFlareAccessLogQuery, bucketSeconds int64) (int64, error)
BucketDimensions(ctx context.Context, filter model.OpenFlareAccessLogQuery, column string, bucketSeconds int64) ([]analyticsmodel.NodeAccessLogBucketDimension, error)
IPAggregates(ctx context.Context, filter model.OpenFlareAccessLogQuery, exactRemoteAddr bool) ([]analyticsmodel.NodeAccessLogIPAggregate, error)
IPSummaries(ctx context.Context, filter model.OpenFlareAccessLogQuery, recentSince time.Time) ([]analyticsmodel.NodeAccessLogIPSummary, error)
CountIPSummaries(ctx context.Context, filter model.OpenFlareAccessLogQuery) (int64, error)
WAFIPAggregates(ctx context.Context, filter model.OpenFlareAccessLogQuery) ([]analyticsmodel.NodeAccessLogWAFIPAggregate, error)
IPTrend(ctx context.Context, filter model.OpenFlareAccessLogQuery, bucketSeconds int64) ([]analyticsmodel.NodeAccessLogIPTrend, error)
TrafficSummary(ctx context.Context, filter model.OpenFlareAccessLogQuery) (model.OpenFlareAccessLogTrafficSummary, error)
ValueCounts(ctx context.Context, filter model.OpenFlareAccessLogQuery, column string, limit int) ([]model.OpenFlareAccessLogValueCount, error)
NodeAggregates(ctx context.Context, filter model.OpenFlareAccessLogQuery) ([]model.OpenFlareAccessLogNodeAggregate, error)
DeleteAll(ctx context.Context) (int64, error)
DeleteBefore(ctx context.Context, cutoff time.Time) (int64, error)
DeleteByNodeBefore(ctx context.Context, nodeID string, before time.Time) (int64, error)
// ListForMigration 按 id 升序分页读取(迁移复制用)。
ListForMigration(ctx context.Context, afterID uint64, limit int) ([]analyticsmodel.NodeAccessLog, error)
}
// ObservabilityStore 可观测 4 表(metric snapshots / edge health / frps / frpc)。
type ObservabilityStore interface {
InsertMetricSnapshot(ctx context.Context, record *model.OpenFlareMetricSnapshot) error
ListMetricSnapshots(ctx context.Context, nodeID string, since time.Time, limit int) ([]*model.OpenFlareMetricSnapshot, error)
DeleteAllMetricSnapshots(ctx context.Context) (int64, error)
DeleteMetricSnapshotsBefore(ctx context.Context, cutoff time.Time) (int64, error)
BatchInsertNodeMetricSnapshots(ctx context.Context, rows []analyticsmodel.NodeMetricSnapshot) error
InsertEdgeHealth(ctx context.Context, record *model.OpenFlareEdgeHealth) error
ListEdgeHealth(ctx context.Context, nodeID string, since time.Time, limit int) ([]*model.OpenFlareEdgeHealth, error)
DeleteAllEdgeHealth(ctx context.Context) (int64, error)
DeleteEdgeHealthBefore(ctx context.Context, cutoff time.Time) (int64, error)
BatchInsertNodeEdgeHealth(ctx context.Context, rows []analyticsmodel.NodeEdgeHealth) error
InsertNodeObservationFrps(ctx context.Context, record *model.OpenFlareNodeObservationFrps) error
ListNodeObservationFrps(ctx context.Context, nodeID string, since time.Time, limit int) ([]*model.OpenFlareNodeObservationFrps, error)
DeleteAllNodeObservationFrps(ctx context.Context) (int64, error)
DeleteNodeObservationFrpsBefore(ctx context.Context, cutoff time.Time) (int64, error)
BatchInsertNodeObsFrps(ctx context.Context, rows []analyticsmodel.NodeObsFrps) error
InsertNodeObservationFrpc(ctx context.Context, record *model.OpenFlareNodeObservationFrpc) error
ListNodeObservationFrpc(ctx context.Context, nodeID string, since time.Time, limit int) ([]*model.OpenFlareNodeObservationFrpc, error)
DeleteAllNodeObservationFrpc(ctx context.Context) (int64, error)
DeleteNodeObservationFrpcBefore(ctx context.Context, cutoff time.Time) (int64, error)
BatchInsertNodeObsFrpc(ctx context.Context, rows []analyticsmodel.NodeObsFrpc) error
// 迁移复制用:按 id 升序分页读取。
ListMetricSnapshotsForMigration(ctx context.Context, afterID uint64, limit int) ([]analyticsmodel.NodeMetricSnapshot, error)
ListEdgeHealthForMigration(ctx context.Context, afterID uint64, limit int) ([]analyticsmodel.NodeEdgeHealth, error)
ListNodeObsFrpsForMigration(ctx context.Context, afterID uint64, limit int) ([]analyticsmodel.NodeObsFrps, error)
ListNodeObsFrpcForMigration(ctx context.Context, afterID uint64, limit int) ([]analyticsmodel.NodeObsFrpc, error)
}
// UserAccessLogStore 用户访问日志(w_user_access_logs)。
type UserAccessLogStore interface {
BatchInsert(ctx context.Context, logs []analyticsmodel.UserAccessLog) error
Count(ctx context.Context, filter analyticsmodel.AccessLogFilter) (uint64, error)
List(ctx context.Context, filter analyticsmodel.AccessLogFilter, page, pageSize int) ([]analyticsmodel.UserAccessLog, uint64, error)
GetDailyTrend(ctx context.Context, days int) ([]analyticsmodel.DailyTrend, error)
GetBrowserDistribution(ctx context.Context, startTime time.Time) ([]analyticsmodel.BrowserShare, error)
GetTopActiveUsers(ctx context.Context, startTime time.Time, limit int) ([]analyticsmodel.TopUser, error)
}
// StatusStore 日志库状态(供管理端状态端点)。
type StatusStore interface {
ActiveDatabase(ctx context.Context) (string, error)
ClickHouseOperationalStats(ctx context.Context) (*analyticsmodel.ClickHouseOperationalStats, error) // 仅 CH 激活时非 nil
}
// Store 聚合当前生效日志库的全部域存储。
type Store struct {
AccessLogs AccessLogStore
Observability ObservabilityStore
UserAccessLogs UserAccessLogStore
Status StatusStore
}
- Step 2: 写 provider(provider.go)
package logstore
import (
"context"
"errors"
"fmt"
"sync"
"github.com/Rain-kl/Wavelet/internal/infra/config"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
)
// logDatabaseKey / logMigrationKey 暂用字符串,Task 8 换为 model.ConfigKey*。
const (
logDatabaseKey = "log_database"
logMigrationKey = "log_db_migration"
)
// ConfigReader 读取系统配置字符串值,由 bootstrap 注入(避免 logstore ↔ repository 循环依赖)。
type ConfigReader func(ctx context.Context, key string) (string, error)
var (
configReader ConfigReader
storeMu sync.RWMutex
active *Store
activeDB string
)
// SetConfigReader 注入系统配置读取函数(bootstrap 调用,测试可注入内存实现)。
func SetConfigReader(fn ConfigReader) { configReader = fn }
func getConfig(ctx context.Context, key string) (string, error) {
if configReader == nil {
return "", errors.New("logstore: config reader not wired")
}
return configReader(ctx, key)
}
// Active 返回当前生效的日志库 Store。按 log_database 系统配置惰性解析并缓存,
// 配置更新(含迁移任务翻转)后自动重建。
func Active(ctx context.Context) (*Store, error) {
current, err := resolveDatabase(ctx)
if err != nil {
return nil, err
}
storeMu.RLock()
if active != nil && activeDB == current {
s := active
storeMu.RUnlock()
return s, nil
}
storeMu.RUnlock()
storeMu.Lock()
defer storeMu.Unlock()
if active != nil && activeDB == current {
return active, nil
}
s, err := buildStore(ctx, current)
if err != nil {
return nil, err
}
active = s
activeDB = current
return s, nil
}
// Migrating 返回日志库是否处于迁移冻结状态。
func Migrating(ctx context.Context) bool {
v, err := getConfig(ctx, logMigrationKey)
if err != nil {
return false
}
return v == "migrating"
}
// Init 在 bootstrap 阶段预热一次激活 store(幂等,失败不致命——首次使用时再解析)。
func Init(ctx context.Context) {
_, _ = Active(ctx)
}
// ResetForTest 清空缓存的激活 store 与 reader,便于测试注入。
func ResetForTest() {
storeMu.Lock()
active = nil
activeDB = ""
storeMu.Unlock()
}
// Build 直接按目标构造 store(迁移任务复制到目标库时使用,不经 Active 缓存)。
func Build(ctx context.Context, database string) (*Store, error) {
return buildStore(ctx, database)
}
// ActiveDatabase 返回当前日志主库名(postgres|sqlite|clickhouse)。
func ActiveDatabase(ctx context.Context) (string, error) {
return resolveDatabase(ctx)
}
// resolveDatabase 读取 log_database,缺失时按启动规则 seed 并返回。
func resolveDatabase(ctx context.Context) (string, error) {
v, err := getConfig(ctx, logDatabaseKey)
if err == nil && v != "" {
return v, nil
}
// 首次启动 seed:CH 启用 → clickhouse;否则随主库。
defaultDB := "sqlite"
if config.Config.Database.Enabled {
defaultDB = "postgres"
}
if config.Config.ClickHouse.Enabled {
defaultDB = "clickhouse"
}
return defaultDB, nil
}
// buildStore 按目标构造实现(Task 3-5 提供构造函数)。
func buildStore(ctx context.Context, database string) (*Store, error) {
switch database {
case "clickhouse":
ch := newClickHouseStore()
return &Store{AccessLogs: ch, Observability: ch, UserAccessLogs: ch, Status: ch}, nil
case "postgres", "sqlite":
g := newGormStore(db.DB(ctx))
return &Store{AccessLogs: g, Observability: g, UserAccessLogs: g, Status: g}, nil
default:
return nil, fmt.Errorf("unsupported log database: %s", database)
}
}
(db.DB(ctx) 返回 *gorm.DB,见 internal/infra/persistence/postgres.go;newGormStore/newClickHouseStore 在 Task 3-5 实现。)
- Step 3: 写 provider 单测(provider_test.go)——用
SetStoreForTest注入 fake 验证Active缓存与切换:
package logstore
import (
"context"
"testing"
)
func TestMigratingReadsConfig(t *testing.T) {
ResetForTest()
SetConfigReader(func(_ context.Context, key string) (string, error) {
if key == logMigrationKey {
return "migrating", nil
}
return "", nil
})
if !Migrating(context.Background()) {
t.Fatal("Migrating() = false, want true when key=migrating")
}
SetConfigReader(func(_ context.Context, key string) (string, error) {
return "", nil
})
if Migrating(context.Background()) {
t.Fatal("Migrating() = true, want false when key empty")
}
}
func TestResolveDatabaseDefaults(t *testing.T) {
ResetForTest()
// 配置缺失时按主库规则 seed(config.Config 默认值由既有测试基建决定)。
got, err := resolveDatabase(context.Background())
if err != nil {
t.Fatalf("resolveDatabase: %v", err)
}
if got != "postgres" && got != "sqlite" && got != "clickhouse" {
t.Fatalf("unexpected default log database: %s", got)
}
}
- Step 4: 运行测试
go test ./internal/repository/logstore/期望 PASS。 - Step 5: 提交
git add internal/repository/logstore/ && git commit -m "feat(logstore): add log store interfaces and provider skeleton"
Task 3: GORM 实现——节点访问日志(AccessLogStore)
Files:
- Create:
internal/repository/logstore/postgres_store.go - Create:
internal/repository/logstore/dialect_postgres.go - Create:
internal/repository/logstore/dialect_sqlite.go - Create:
internal/repository/logstore/postgres_store_test.go
Interfaces:
-
Consumes:
db.DB(ctx)、analyticsmodel.*、model.OpenFlareAccessLog*、hooks注册表(Task 5 提供QueueNodeAccessLogs)。 -
Produces:
newGormStore(db *gorm.DB) *gormLogStore(实现AccessLogStore/ObservabilityStore/UserAccessLogStore)。 -
Step 1: 写 dialect 小文件
dialect_postgres.go:
package logstore
import "gorm.io/gorm"
// timeBucketSQL 返回 PG 时间分桶表达式(epoch 秒 -> 分桶起点)。
func timeBucketSQL(column string, bucketSeconds int64) string {
return "to_timestamp(floor(extract(epoch from " + column + ")/" + itoa(bucketSeconds) + ")*" + itoa(bucketSeconds) + ")"
}
// gormDBForWrite 返回写句柄(PG/SQLite 相同)。
func gormDBForWrite(db *gorm.DB) *gorm.DB { return db }
dialect_sqlite.go:
package logstore
import (
"strconv"
"gorm.io/gorm"
)
func timeBucketSQL(column string, bucketSeconds int64) string {
return "(floor(unixepoch(" + column + ")/" + strconv.FormatInt(bucketSeconds, 10) + ")*" + strconv.FormatInt(bucketSeconds, 10) + ")"
}
func gormDBForWrite(db *gorm.DB) *gorm.DB { return db }
若需要精确到毫秒的分桶(现有 CH 用秒级分桶即可),以现有
node_access_log_stats.go的 bucket 语义为准,两种方言输出同一语义。
- Step 2: 写
postgres_store.go(节点访问日志部分)
package logstore
import (
"context"
"errors"
"fmt"
"time"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
"gorm.io/gorm"
)
// gormLogStore 是 PG/SQLite 共用的 GORM 日志存储实现。
type gormLogStore struct {
db *gorm.DB
}
func newGormStore(db *gorm.DB) *gormLogStore { return &gormLogStore{db: db} }
// ensureWritable 冻结期拒绝写入。
func (s *gormLogStore) ensureWritable(ctx context.Context) error {
if Migrating(ctx) {
return ErrMigrating
}
return nil
}
// InsertBatch 节点访问日志写入入口:冻结检查后经 hook 入队(异步),与现状一致。
func (s *gormLogStore) InsertBatch(ctx context.Context, records []*model.OpenFlareAccessLog) error {
if err := s.ensureWritable(ctx); err != nil {
return err
}
rows := make([]analyticsmodel.NodeAccessLog, 0, len(records))
for _, r := range records {
if r == nil {
continue
}
rows = append(rows, toAnalyticsNodeAccessLog(r))
}
if h := currentAccessLogHooks().QueueNodeAccessLogs; h != nil {
h(rows)
}
return nil
}
// BatchInsertNodeAccessLogs 是 batchwriter flush 目标:GORM 分批落库。
func (s *gormLogStore) BatchInsertNodeAccessLogs(ctx context.Context, rows []analyticsmodel.NodeAccessLog) error {
if len(rows) == 0 {
return nil
}
if err := s.ensureWritable(ctx); err != nil {
return err
}
return s.db.WithContext(ctx).CreateInBatches(rows, 500).Error
}
func (s *gormLogStore) List(ctx context.Context, query model.OpenFlareAccessLogQuery) ([]*model.OpenFlareAccessLog, error) {
f := toNodeAccessLogFilter(query)
var rows []analyticsmodel.NodeAccessLog
q := s.db.WithContext(ctx).Model(&analyticsmodel.NodeAccessLog{})
if f.Since.IsZero() == false {
q = q.Where("logged_at >= ?", f.Since)
}
if f.Until.IsZero() == false {
q = q.Where("logged_at <= ?", f.Until)
}
if f.NodeID != "" {
q = q.Where("node_id = ?", f.NodeID)
}
if f.RemoteAddr != "" {
q = q.Where("remote_addr = ?", f.RemoteAddr)
}
if len(f.Hosts) > 0 {
q = q.Where("host IN ?", f.Hosts)
}
if f.Host != "" {
q = q.Where("host = ?", f.Host)
}
if f.Path != "" {
q = q.Where("path = ?", f.Path)
}
order := "logged_at DESC, id DESC"
if f.SortOrder == "asc" {
order = "logged_at ASC, id ASC"
}
if err := q.Order(order).Limit(limitOr(f.PageSize, 100)).Offset(offsetOf(f.Page, f.PageSize)).Find(&rows).Error; err != nil {
return nil, err
}
return fromAnalyticsNodeAccessLogs(rows), nil
}
func (s *gormLogStore) Count(ctx context.Context, query model.OpenFlareAccessLogQuery) (int64, int64, int64, error) {
f := toNodeAccessLogFilter(query)
var total, uniqIP, bytesSent int64
q := s.db.WithContext(ctx).Model(&analyticsmodel.NodeAccessLog{})
if !f.Since.IsZero() {
q = q.Where("logged_at >= ?", f.Since)
}
if !f.Until.IsZero() {
q = q.Where("logged_at <= ?", f.Until)
}
if f.NodeID != "" {
q = q.Where("node_id = ?", f.NodeID)
}
if f.RemoteAddr != "" {
q = q.Where("remote_addr = ?", f.RemoteAddr)
}
if len(f.Hosts) > 0 {
q = q.Where("host IN ?", f.Hosts)
}
if f.Host != "" {
q = q.Where("host = ?", f.Host)
}
if f.Path != "" {
q = q.Where("path = ?", f.Path)
}
if err := q.Count(&total).Error; err != nil {
return 0, 0, 0, err
}
if err := q.Distinct("remote_addr").Count(&uniqIP).Error; err != nil {
return 0, 0, 0, err
}
if err := q.Select("COALESCE(SUM(bytes_sent),0)").Scan(&bytesSent).Error; err != nil {
return 0, 0, 0, err
}
return total, uniqIP, bytesSent, nil
}
func (s *gormLogStore) TrafficSummary(ctx context.Context, query model.OpenFlareAccessLogQuery) (model.OpenFlareAccessLogTrafficSummary, error) {
f := toNodeAccessLogFilter(query)
var out struct {
RequestCount int64
ErrorCount int64
UniqueIPCount int64
BytesSent int64
RequestLength int64
NodeCount int64
}
q := s.db.WithContext(ctx).Model(&analyticsmodel.NodeAccessLog{})
if !f.Since.IsZero() {
q = q.Where("logged_at >= ?", f.Since)
}
if !f.Until.IsZero() {
q = q.Where("logged_at <= ?", f.Until)
}
if f.NodeID != "" {
q = q.Where("node_id = ?", f.NodeID)
}
if f.Host != "" {
q = q.Where("host = ?", f.Host)
}
err := q.Select(`
COUNT(*) AS request_count,
COUNT(*) FILTER (WHERE status_code >= 500) AS error_count,
COUNT(DISTINCT remote_addr) AS unique_ip_count,
COALESCE(SUM(bytes_sent),0) AS bytes_sent,
COALESCE(SUM(request_length),0) AS request_length,
COUNT(DISTINCT node_id) AS node_count`).Scan(&out).Error
if err != nil {
return model.OpenFlareAccessLogTrafficSummary{}, err
}
return model.OpenFlareAccessLogTrafficSummary{
RequestCount: out.RequestCount,
ErrorCount: out.ErrorCount,
UniqueIPCount: out.UniqueIPCount,
BytesSent: out.BytesSent,
RequestLength: out.RequestLength,
NodeCount: out.NodeCount,
}, nil
}
func (s *gormLogStore) ValueCounts(ctx context.Context, query model.OpenFlareAccessLogQuery, column string, limit int) ([]model.OpenFlareAccessLogValueCount, error) {
col, ok := nodeAccessLogValueColumn(column)
if !ok {
return nil, fmt.Errorf("unsupported value count column: %s", column)
}
f := toNodeAccessLogFilter(query)
q := s.db.WithContext(ctx).Model(&analyticsmodel.NodeAccessLog{}).
Select(col+" AS value, COUNT(*) AS count")
if !f.Since.IsZero() {
q = q.Where("logged_at >= ?", f.Since)
}
if !f.Until.IsZero() {
q = q.Where("logged_at <= ?", f.Until)
}
if f.NodeID != "" {
q = q.Where("node_id = ?", f.NodeID)
}
if f.Host != "" {
q = q.Where("host = ?", f.Host)
}
type row struct {
Value string
Count int64
}
var rows []row
if err := q.Group(col).Order("count DESC").Limit(limitOr(limit, 10)).Scan(&rows).Error; err != nil {
return nil, err
}
out := make([]model.OpenFlareAccessLogValueCount, len(rows))
for i, r := range rows {
out[i] = model.OpenFlareAccessLogValueCount{Value: r.Value, Count: r.Count}
}
return out, nil
}
func (s *gormLogStore) NodeAggregates(ctx context.Context, query model.OpenFlareAccessLogQuery) ([]model.OpenFlareAccessLogNodeAggregate, error) {
f := toNodeAccessLogFilter(query)
type row struct {
NodeID string
RequestCount int64
ErrorCount int64
UniqueIPCount int64
}
var rows []row
q := s.db.WithContext(ctx).Model(&analyticsmodel.NodeAccessLog{}).
Select("node_id, COUNT(*) AS request_count, COUNT(*) FILTER (WHERE status_code >= 500) AS error_count, COUNT(DISTINCT remote_addr) AS unique_ip_count")
if !f.Since.IsZero() {
q = q.Where("logged_at >= ?", f.Since)
}
if !f.Until.IsZero() {
q = q.Where("logged_at <= ?", f.Until)
}
if f.NodeID != "" {
q = q.Where("node_id = ?", f.NodeID)
}
if f.Host != "" {
q = q.Where("host = ?", f.Host)
}
if err := q.Group("node_id").Order("request_count DESC").Scan(&rows).Error; err != nil {
return nil, err
}
out := make([]model.OpenFlareAccessLogNodeAggregate, len(rows))
for i, r := range rows {
out[i] = model.OpenFlareAccessLogNodeAggregate{NodeID: r.NodeID, RequestCount: r.RequestCount, ErrorCount: r.ErrorCount, UniqueIPCount: r.UniqueIPCount}
}
return out, nil
}
func (s *gormLogStore) RegionCounts(ctx context.Context, nodeID string, since time.Time, limit int) ([]*model.OpenFlareAccessLogRegionCount, error) {
type row struct {
Region string
Count int64
}
var rows []row
q := s.db.WithContext(ctx).Model(&analyticsmodel.NodeAccessLog{}).
Select("region, COUNT(*) AS count").
Where("node_id = ? AND region <> '' AND logged_at >= ?", nodeID, since)
if err := q.Group("region").Order("count DESC").Limit(limitOr(limit, 10)).Scan(&rows).Error; err != nil {
return nil, err
}
out := make([]*model.OpenFlareAccessLogRegionCount, len(rows))
for i, r := range rows {
out[i] = &model.OpenFlareAccessLogRegionCount{Region: r.Region, Count: r.Count}
}
return out, nil
}
func (s *gormLogStore) BucketAggregates(ctx context.Context, query model.OpenFlareAccessLogQuery, bucketSeconds int64) ([]analyticsmodel.NodeAccessLogBucketAggregate, error) {
f := toNodeAccessLogFilter(query)
expr := timeBucketSQL("logged_at", bucketSeconds)
type row struct {
Bucket int64
RequestCount int64
ErrorCount int64
}
var rows []row
q := s.db.WithContext(ctx).Model(&analyticsmodel.NodeAccessLog{}).
Select(expr+" AS bucket, COUNT(*) AS request_count, COUNT(*) FILTER (WHERE status_code >= 500) AS error_count")
if !f.Since.IsZero() {
q = q.Where("logged_at >= ?", f.Since)
}
if !f.Until.IsZero() {
q = q.Where("logged_at <= ?", f.Until)
}
if f.NodeID != "" {
q = q.Where("node_id = ?", f.NodeID)
}
if f.Host != "" {
q = q.Where("host = ?", f.Host)
}
if err := q.Group(expr).Order("bucket ASC").Scan(&rows).Error; err != nil {
return nil, err
}
out := make([]analyticsmodel.NodeAccessLogBucketAggregate, len(rows))
for i, r := range rows {
out[i] = analyticsmodel.NodeAccessLogBucketAggregate{Bucket: r.Bucket, RequestCount: r.RequestCount, ErrorCount: r.ErrorCount}
}
return out, nil
}
func (s *gormLogStore) DeleteAll(ctx context.Context) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
res := s.db.WithContext(ctx).Where("1 = 1").Delete(&analyticsmodel.NodeAccessLog{})
return res.RowsAffected, res.Error
}
func (s *gormLogStore) DeleteBefore(ctx context.Context, cutoff time.Time) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
res := s.db.WithContext(ctx).Where("logged_at < ?", cutoff).Delete(&analyticsmodel.NodeAccessLog{})
return res.RowsAffected, res.Error
}
func (s *gormLogStore) DeleteByNodeBefore(ctx context.Context, nodeID string, before time.Time) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
res := s.db.WithContext(ctx).Where("node_id = ? AND logged_at < ?", nodeID, before).Delete(&analyticsmodel.NodeAccessLog{})
return res.RowsAffected, res.Error
}
- Step 2b: 补齐 AccessLogStore 剩余聚合方法(必须全部实现 + 编译期断言)
gormLogStore 必须实现 AccessLogStore 的全部 20 个方法(当前 Step 1 只含 13 个)。补齐:CountBuckets、BucketDimensions、IPAggregates、IPSummaries、CountIPSummaries、WAFIPAggregates、IPTrend。语义以 internal/repository/analytics/node_access_log_stats.go(及 node_access_log.go 中对应函数)为准,用 GORM/方言 SQL 等价实现:
-
时间分桶统一返回 epoch 秒整型:PG
(floor(extract(epoch from <col>)/<N>)*<N>)::bigint;SQLite(floor(unixepoch(<col>)/<N>)*<N>)(修正timeBucketSQL,保证 PG/SQLite 输出同为 int64 epoch,与BucketEpoch扫描类型一致)。 -
CountBuckets:SELECT COUNT(*) FROM (SELECT 1 FROM t WHERE ... GROUP BY bucket) x。 -
BucketDimensions:GROUP BY bucket, <column>返回维度计数。 -
IPAggregates:按 remote_addr(或精确 remote_addr)聚合 request_count / error_count / unique host 等,字段对照NodeAccessLogIPAggregate。 -
IPSummaries/CountIPSummaries:按 IP 汇总近窗口(含最近活跃时间),字段对照NodeAccessLogIPSummary。 -
WAFIPAggregates:按 IP 聚合状态码分布,字段对照NodeAccessLogWAFIPAggregate。 -
IPTrend:按 IP × 时间桶聚合,字段对照NodeAccessLogIPTrend。 -
过滤语义对齐 CH(
node_access_log_filter.go):remote_addr/host/path 用LIKE trim(value)+'%'前缀匹配;hosts 用lower(trim(host)) IN (...);until 用开区间<;node_id 先 trim。 -
文件底部加编译期断言:
var _ AccessLogStore = (*gormLogStore)(nil)。 -
测试:
postgres_store_test.go至少覆盖CountBuckets/IPTrend(sqlite 内存库写入若干行后断言分桶数量与趋势),其余方法以编译期断言 + 既有语义测试兜底。 -
Step 3: 写 helper(postgres_store.go 同文件底部)
func limitOr(v, def int) int {
if v <= 0 {
return def
}
return v
}
func offsetOf(page, pageSize int) int {
if page < 1 {
page = 1
}
if pageSize < 1 {
pageSize = 20
}
return (page - 1) * pageSize
}
func nodeAccessLogValueColumn(column string) (string, bool) {
switch column {
case "remote_addr":
return "remote_addr", true
case "host":
return "host", true
case "path":
return "path", true
case "region":
return "region", true
case "status_code":
return "status_code", true
case "user_agent":
return "user_agent", true
case "cache_status":
return "cache_status", true
}
return "", false
}
toAnalyticsNodeAccessLog/fromAnalyticsNodeAccessLogs/toNodeAccessLogFilter从internal/repository/openflare_access_log_store.go复制(含 math 边界保护逻辑);Task 6 删除旧文件后这些 helper 不再冲突。
- Step 4: 写单测(postgres_store_test.go,sqlite 内存库 + AutoMigrate)
package logstore
import (
"context"
"testing"
"time"
"github.com/glebarez/sqlite"
"gorm.io/gorm"
"gorm.io/gorm/logger"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
)
func newTestGormStore(t *testing.T) *gormLogStore {
t.Helper()
db, err := gorm.Open(sqlite.Open("file::memory:?cache=shared"), &gorm.Config{Logger: logger.Default.LogMode(logger.Silent)})
if err != nil {
t.Fatalf("open sqlite: %v", err)
}
if err := db.AutoMigrate(&analyticsmodel.NodeAccessLog{}); err != nil {
t.Fatalf("automigrate: %v", err)
}
return newGormStore(db)
}
func TestGormBatchInsertAndCount(t *testing.T) {
ResetForTest()
SetConfigReader(func(_ context.Context, _ string) (string, error) { return "", nil })
s := newTestGormStore(t)
now := time.Now()
rows := []analyticsmodel.NodeAccessLog{
{ID: 1, NodeID: "n1", LoggedAt: now, RemoteAddr: "1.1.1.1", StatusCode: 200, BytesSent: 100},
{ID: 2, NodeID: "n1", LoggedAt: now, RemoteAddr: "2.2.2.2", StatusCode: 500, BytesSent: 200},
}
if err := s.BatchInsertNodeAccessLogs(context.Background(), rows); err != nil {
t.Fatalf("insert: %v", err)
}
total, uniqIP, bytesSent, err := s.Count(context.Background(), model.OpenFlareAccessLogQuery{NodeID: "n1"})
if err != nil {
t.Fatalf("count: %v", err)
}
if total != 2 || uniqIP != 2 || bytesSent != 300 {
t.Fatalf("count got total=%d uniq=%d bytes=%d", total, uniqIP, bytesSent)
}
}
(nodeQuery 返回 model.OpenFlareAccessLogQuery{NodeID: "n1"};InsertBatch 冻结与 hook 测试放 Task 6。)
- Step 5: 运行测试
go test ./internal/repository/logstore/期望 PASS。 - Step 6: 提交
git add internal/repository/logstore/ && git commit -m "feat(logstore): GORM node access log store"
Task 4: GORM 实现——可观测 4 表 + 用户访问日志
Files:
- Modify:
internal/repository/logstore/postgres_store.go(追加方法) - Modify:
internal/repository/logstore/postgres_store_test.go
Interfaces:
-
Consumes:
model.OpenFlareMetricSnapshot/OpenFlareEdgeHealth/OpenFlareNodeObservationFrps/OpenFlareNodeObservationFrpc、analyticsmodel.NodeMetricSnapshot等、currentObservabilityHooks()(Task 5)。 -
Produces:
gormLogStore完整实现ObservabilityStore与UserAccessLogStore。 -
Step 1: 可观测写入入口 + flush + 查询(追加到 postgres_store.go)
// ---- ObservabilityStore ----
func (s *gormLogStore) InsertMetricSnapshot(ctx context.Context, record *model.OpenFlareMetricSnapshot) error {
if record == nil {
return nil
}
if err := s.ensureWritable(ctx); err != nil {
return err
}
if h := currentObservabilityHooks().QueueMetricSnapshot; h != nil {
h(toAnalyticsNodeMetricSnapshot(record))
}
return nil
}
func (s *gormLogStore) ListMetricSnapshots(ctx context.Context, nodeID string, since time.Time, limit int) ([]*model.OpenFlareMetricSnapshot, error) {
var rows []analyticsmodel.NodeMetricSnapshot
q := s.db.WithContext(ctx).Where("node_id = ? AND captured_at >= ?", nodeID, since).Order("captured_at DESC, id DESC")
if err := q.Limit(limitOr(limit, 100)).Find(&rows).Error; err != nil {
return nil, err
}
return fromAnalyticsNodeMetricSnapshots(rows), nil
}
func (s *gormLogStore) DeleteAllMetricSnapshots(ctx context.Context) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
res := s.db.WithContext(ctx).Where("1 = 1").Delete(&analyticsmodel.NodeMetricSnapshot{})
return res.RowsAffected, res.Error
}
func (s *gormLogStore) DeleteMetricSnapshotsBefore(ctx context.Context, cutoff time.Time) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
res := s.db.WithContext(ctx).Where("captured_at < ?", cutoff).Delete(&analyticsmodel.NodeMetricSnapshot{})
return res.RowsAffected, res.Error
}
func (s *gormLogStore) BatchInsertNodeMetricSnapshots(ctx context.Context, rows []analyticsmodel.NodeMetricSnapshot) error {
if len(rows) == 0 {
return nil
}
if err := s.ensureWritable(ctx); err != nil {
return err
}
return s.db.WithContext(ctx).CreateInBatches(rows, 500).Error
}
// InsertEdgeHealth 等 8 个 entry/list/delete + 3 个 flush 全部与 metric snapshots 同构。
// 完整模板(以 edge health 为例):
func (s *gormLogStore) InsertEdgeHealth(ctx context.Context, record *model.OpenFlareEdgeHealth) error {
if record == nil {
return nil
}
if err := s.ensureWritable(ctx); err != nil {
return err
}
if h := currentObservabilityHooks().QueueEdgeHealth; h != nil {
h(toAnalyticsNodeEdgeHealth(record))
}
return nil
}
func (s *gormLogStore) ListEdgeHealth(ctx context.Context, nodeID string, since time.Time, limit int) ([]*model.OpenFlareEdgeHealth, error) {
var rows []analyticsmodel.NodeEdgeHealth
if err := s.db.WithContext(ctx).Where("node_id = ? AND captured_at >= ?", nodeID, since).
Order("captured_at DESC, id DESC").Limit(limitOr(limit, 100)).Find(&rows).Error; err != nil {
return nil, err
}
return fromAnalyticsNodeEdgeHealths(rows), nil
}
func (s *gormLogStore) DeleteAllEdgeHealth(ctx context.Context) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
res := s.db.WithContext(ctx).Where("1 = 1").Delete(&analyticsmodel.NodeEdgeHealth{})
return res.RowsAffected, res.Error
}
func (s *gormLogStore) DeleteEdgeHealthBefore(ctx context.Context, cutoff time.Time) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
res := s.db.WithContext(ctx).Where("captured_at < ?", cutoff).Delete(&analyticsmodel.NodeEdgeHealth{})
return res.RowsAffected, res.Error
}
func (s *gormLogStore) BatchInsertNodeEdgeHealth(ctx context.Context, rows []analyticsmodel.NodeEdgeHealth) error {
if len(rows) == 0 {
return nil
}
if err := s.ensureWritable(ctx); err != nil {
return err
}
return s.db.WithContext(ctx).CreateInBatches(rows, 500).Error
}
// FRPS/FRPC 两组按同一模板,替换映射如下:
// FRPS: model.OpenFlareNodeObservationFrps ↔ analyticsmodel.NodeObsFrps;hook=QueueNodeObsFrps;转换 toAnalyticsNodeObsFrps
// FRPC: model.OpenFlareNodeObservationFrpc ↔ analyticsmodel.NodeObsFrpc;hook=QueueNodeObsFrpc;转换 toAnalyticsNodeObsFrpc
// list 列名统一 captured_at;delete 统一 captured_at < cutoff。
// 转换函数(toAnalyticsNodeEdgeHealth/fromAnalyticsNodeEdgeHealths/toAnalyticsNodeObsFrps/toAnalyticsNodeObsFrpc)
// 从旧 openflare_observability_store.go 复制。
逐方法补齐(8 个 entry/list/delete + 3 个 flush),表名/模型:
analyticsmodel.NodeEdgeHealth、analyticsmodel.NodeObsFrps、analyticsmodel.NodeObsFrpc;model 侧OpenFlareEdgeHealth、OpenFlareNodeObservationFrps、OpenFlareNodeObservationFrpc。toAnalyticsNodeEdgeHealth等转换函数从旧openflare_observability_store.go复制。
- Step 2: 用户访问日志(追加)
// ---- UserAccessLogStore ----
func (s *gormLogStore) BatchInsert(ctx context.Context, logs []analyticsmodel.UserAccessLog) error {
if len(logs) == 0 {
return nil
}
if err := s.ensureWritable(ctx); err != nil {
return err
}
return s.db.WithContext(ctx).CreateInBatches(logs, 500).Error
}
func (s *gormLogStore) Count(ctx context.Context, filter analyticsmodel.AccessLogFilter) (uint64, error) {
var total int64
q := s.db.WithContext(ctx).Model(&analyticsmodel.UserAccessLog{})
if filter.UserID != 0 {
q = q.Where("user_id = ?", filter.UserID)
}
if filter.Path != "" {
q = q.Where("path = ?", filter.Path)
}
if filter.Method != "" {
q = q.Where("method = ?", filter.Method)
}
if filter.IP != "" {
q = q.Where("ip = ?", filter.IP)
}
if filter.Status != 0 {
q = q.Where("status = ?", filter.Status)
}
if !filter.Since.IsZero() {
q = q.Where("created_at >= ?", filter.Since)
}
if !filter.Until.IsZero() {
q = q.Where("created_at <= ?", filter.Until)
}
if err := q.Count(&total).Error; err != nil {
return 0, err
}
return uint64(total), nil
}
func (s *gormLogStore) List(ctx context.Context, filter analyticsmodel.AccessLogFilter, page, pageSize int) ([]analyticsmodel.UserAccessLog, uint64, error) {
total, err := s.Count(ctx, filter)
if err != nil {
return nil, 0, err
}
if total == 0 {
return []analyticsmodel.UserAccessLog{}, 0, nil
}
var rows []analyticsmodel.UserAccessLog
q := s.db.WithContext(ctx).Where(buildUserAccessLogWhere(filter)).Order("created_at DESC, id DESC")
if err := q.Limit(pageSize).Offset(offsetOf(page, pageSize)).Find(&rows).Error; err != nil {
return nil, 0, err
}
return rows, total, nil
}
func (s *gormLogStore) GetDailyTrend(ctx context.Context, days int) ([]analyticsmodel.DailyTrend, error) {
if days <= 0 {
days = 7
}
// 镜像 CH access_log_stats.go:起点 = (days-1) 天前当日零点;必须返回恰好 days 个日历日并补零。
start := time.Now().AddDate(0, 0, -(days - 1)).Truncate(24 * time.Hour)
type row struct {
Date string
Cnt uint64
}
var rows []row
err := s.db.WithContext(ctx).Model(&analyticsmodel.UserAccessLog{}).
Select(dailyTrendDateSQL()+" AS date, COUNT(*) AS cnt").
Where("created_at >= ?", start).
Group("date").Order("date ASC").Scan(&rows).Error
if err != nil {
return nil, err
}
counts := make(map[string]uint64, len(rows))
for _, r := range rows {
counts[r.Date] = r.Cnt
}
out := make([]analyticsmodel.DailyTrend, 0, days)
for i := 0; i < days; i++ {
d := start.AddDate(0, 0, i).Format("2006-01-02")
out = append(out, analyticsmodel.DailyTrend{Date: d, Cnt: counts[d]})
}
return out, nil
}
func (s *gormLogStore) GetBrowserDistribution(ctx context.Context, startTime time.Time) ([]analyticsmodel.BrowserShare, error) {
return s.userAgentGroupCount(ctx, startTime, "browser")
}
func (s *gormLogStore) GetTopActiveUsers(ctx context.Context, startTime time.Time, limit int) ([]analyticsmodel.TopUser, error) {
type row struct {
UserID uint64
Cnt uint64
}
var rows []row
err := s.db.WithContext(ctx).Model(&analyticsmodel.UserAccessLog{}).
Select("user_id, COUNT(*) AS cnt").
Where("user_id <> 0 AND created_at >= ?", startTime).
Group("user_id").Order("cnt DESC").Limit(limitOr(limit, 10)).Scan(&rows).Error
if err != nil {
return nil, err
}
out := make([]analyticsmodel.TopUser, len(rows))
for i, r := range rows {
out[i] = analyticsmodel.TopUser{UserID: r.UserID, Cnt: r.Cnt}
}
return out, nil
}
buildUserAccessLogWhere与Count内联条件一致。AccessLogFilter 使用单一权威字段集(Task 1 迁入的 CH 原字段):UserIDs []uint64、Path、StartTime/EndTime *time.Time。GORM 的 Count/List 必须用该字段集并镜像 CH 过滤语义(user_id IN ?、path LIKE '%..%'、StartTime >=、EndTime <)——禁止在 AccessLogFilter 上追加仅 GORM 使用的字段(会造成双字段集静默分叉)。GetDailyTrend的日期格式化拆到 dialect 文件:dailyTrendDateSQL()返回 PGto_char(created_at,'YYYY-MM-DD')/ SQLitestrftime('%Y-%m-%d', created_at)。userAgentGroupCount用现有analyticsrepo.ParseBrowserName语义改为 SQL 侧CASE或复用 helper——实现时对照access_log_stats.go的浏览器判定逻辑,保持统计口径一致。
- Step 3: 单测追加——
TestGormUserAccessLogCountList、TestGormObservabilityInsertList(sqlite AutoMigrate 对应模型,断言写入/查询/删除)。 - Step 4: 运行
go test ./internal/repository/logstore/PASS。 - Step 5: 提交
git add internal/repository/logstore/ && git commit -m "feat(logstore): GORM observability and user access log store"
Task 5: CH 包装实现 + hooks 注册表迁入 logstore
Files:
- Create:
internal/repository/logstore/clickhouse_store.go - Create:
internal/repository/logstore/hooks.go - Modify:
internal/repository/openflare_access_log_store.go、internal/repository/openflare_observability_store.go(删除,被吸收)
Interfaces:
-
Consumes:
analyticsrepo.*全部现成函数、db.ChConn/db.ChDB。 -
Produces:
newClickHouseStore() *clickhouseLogStore;SetAccessLogHooks/SetObservabilityHooks/currentAccessLogHooks/currentObservabilityHooks。 -
Step 1: hooks.go
package logstore
import (
"sync"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
)
// AccessLogHooks 节点访问日志异步入队回调(由 chwriter 装配)。
type AccessLogHooks struct {
QueueNodeAccessLogs func(logs []analyticsmodel.NodeAccessLog)
}
// ObservabilityHooks 可观测异步入队回调(由 chwriter 装配)。
type ObservabilityHooks struct {
QueueMetricSnapshot func(record analyticsmodel.NodeMetricSnapshot)
QueueEdgeHealth func(record analyticsmodel.NodeEdgeHealth)
QueueNodeObsFrps func(record analyticsmodel.NodeObsFrps)
QueueNodeObsFrpc func(record analyticsmodel.NodeObsFrpc)
}
var (
hooksMu sync.RWMutex
accessLogHooks AccessLogHooks
observabilityHooks ObservabilityHooks
)
func SetAccessLogHooks(h AccessLogHooks) {
hooksMu.Lock()
accessLogHooks = h
hooksMu.Unlock()
}
func SetObservabilityHooks(h ObservabilityHooks) {
hooksMu.Lock()
observabilityHooks = h
hooksMu.Unlock()
}
func currentAccessLogHooks() AccessLogHooks {
hooksMu.RLock()
defer hooksMu.RUnlock()
return accessLogHooks
}
func currentObservabilityHooks() ObservabilityHooks {
hooksMu.RLock()
defer hooksMu.RUnlock()
return observabilityHooks
}
旧
AccessLogInsertHooks/ObservabilityInsertHooks及SetAccessLogInsertHooks等在 repository 包删除,chwriter 改为调用logstore.SetAccessLogHooks(Task 9)。
- Step 2: clickhouse_store.go——逐方法委托 analyticsrepo(仅列代表,全部方法照此)
package logstore
import (
"context"
"errors"
"time"
db "github.com/Rain-kl/Wavelet/internal/infra/persistence"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
analyticsrepo "github.com/Rain-kl/Wavelet/internal/repository/analytics"
)
type clickhouseLogStore struct{}
func newClickHouseStore() *clickhouseLogStore { return &clickhouseLogStore{} }
func chConnErr() error {
if !db.ChConnReady() {
return errors.New("clickhouse connection is not initialized")
}
return nil
}
// ---- AccessLogStore ----
func (s *clickhouseLogStore) InsertBatch(ctx context.Context, records []*model.OpenFlareAccessLog) error {
if err := s.ensureWritable(ctx); err != nil {
return err
}
rows := make([]analyticsmodel.NodeAccessLog, 0, len(records))
for _, r := range records {
if r == nil {
continue
}
rows = append(rows, toAnalyticsNodeAccessLog(r))
}
if h := currentAccessLogHooks().QueueNodeAccessLogs; h != nil {
h(rows)
}
return nil
}
func (s *clickhouseLogStore) ensureWritable(ctx context.Context) error {
if Migrating(ctx) {
return ErrMigrating
}
return nil
}
func (s *clickhouseLogStore) BatchInsertNodeAccessLogs(ctx context.Context, rows []analyticsmodel.NodeAccessLog) error {
if err := s.ensureWritable(ctx); err != nil {
return err
}
return analyticsrepo.BatchInsertNodeAccessLogs(ctx, rows)
}
func (s *clickhouseLogStore) List(ctx context.Context, query model.OpenFlareAccessLogQuery) ([]*model.OpenFlareAccessLog, error) {
rows, err := analyticsrepo.ListNodeAccessLogs(ctx, toNodeAccessLogFilter(query))
if err != nil {
return nil, err
}
return fromAnalyticsNodeAccessLogs(rows), nil
}
func (s *clickhouseLogStore) Count(ctx context.Context, query model.OpenFlareAccessLogQuery) (int64, int64, int64, error) {
return analyticsrepo.CountNodeAccessLogs(ctx, toNodeAccessLogFilter(query))
}
func (s *clickhouseLogStore) RegionCounts(ctx context.Context, nodeID string, since time.Time, limit int) ([]*model.OpenFlareAccessLogRegionCount, error) {
rows, err := analyticsrepo.RegionCountsNodeAccessLogs(ctx, nodeID, since, limit)
if err != nil {
return nil, err
}
out := make([]*model.OpenFlareAccessLogRegionCount, len(rows))
for i, r := range rows {
out[i] = &model.OpenFlareAccessLogRegionCount{Region: r.Region, Count: r.Count}
}
return out, nil
}
func (s *clickhouseLogStore) BucketAggregates(ctx context.Context, query model.OpenFlareAccessLogQuery, bucketSeconds int64) ([]analyticsmodel.NodeAccessLogBucketAggregate, error) {
return analyticsrepo.BucketAggregatesNodeAccessLogs(ctx, toNodeAccessLogFilter(query), bucketSeconds)
}
func (s *clickhouseLogStore) CountBuckets(ctx context.Context, query model.OpenFlareAccessLogQuery, bucketSeconds int64) (int64, error) {
return analyticsrepo.CountBucketAggregatesNodeAccessLogs(ctx, toNodeAccessLogFilter(query), bucketSeconds)
}
func (s *clickhouseLogStore) BucketDimensions(ctx context.Context, query model.OpenFlareAccessLogQuery, column string, bucketSeconds int64) ([]analyticsmodel.NodeAccessLogBucketDimension, error) {
return analyticsrepo.BucketDimensionsNodeAccessLogs(ctx, toNodeAccessLogFilter(query), column, bucketSeconds)
}
func (s *clickhouseLogStore) IPAggregates(ctx context.Context, query model.OpenFlareAccessLogQuery, exactRemoteAddr bool) ([]analyticsmodel.NodeAccessLogIPAggregate, error) {
return analyticsrepo.IPAggregatesNodeAccessLogs(ctx, toNodeAccessLogFilter(query), exactRemoteAddr)
}
func (s *clickhouseLogStore) IPSummaries(ctx context.Context, query model.OpenFlareAccessLogQuery, recentSince time.Time) ([]analyticsmodel.NodeAccessLogIPSummary, error) {
return analyticsrepo.IPSummariesNodeAccessLogs(ctx, toNodeAccessLogFilter(query), recentSince)
}
func (s *clickhouseLogStore) CountIPSummaries(ctx context.Context, query model.OpenFlareAccessLogQuery) (int64, error) {
return analyticsrepo.CountIPSummaryNodeAccessLogs(ctx, toNodeAccessLogFilter(query))
}
func (s *clickhouseLogStore) WAFIPAggregates(ctx context.Context, query model.OpenFlareAccessLogQuery) ([]analyticsmodel.NodeAccessLogWAFIPAggregate, error) {
return analyticsrepo.IPAggregatesForWAFNodeAccessLogs(ctx, toNodeAccessLogFilter(query))
}
func (s *clickhouseLogStore) IPTrend(ctx context.Context, query model.OpenFlareAccessLogQuery, bucketSeconds int64) ([]analyticsmodel.NodeAccessLogIPTrend, error) {
return analyticsrepo.IPTrendNodeAccessLogs(ctx, toNodeAccessLogFilter(query), bucketSeconds)
}
func (s *clickhouseLogStore) TrafficSummary(ctx context.Context, query model.OpenFlareAccessLogQuery) (model.OpenFlareAccessLogTrafficSummary, error) {
row, err := analyticsrepo.TrafficSummaryNodeAccessLogs(ctx, toNodeAccessLogFilter(query))
if err != nil {
return model.OpenFlareAccessLogTrafficSummary{}, err
}
return model.OpenFlareAccessLogTrafficSummary{
RequestCount: int64(row.RequestCount),
ErrorCount: int64(row.ErrorCount),
UniqueIPCount: int64(row.UniqueIPCount),
BytesSent: int64(row.BytesSent),
RequestLength: int64(row.RequestLength),
NodeCount: int64(row.NodeCount),
}, nil
}
func (s *clickhouseLogStore) ValueCounts(ctx context.Context, query model.OpenFlareAccessLogQuery, column string, limit int) ([]model.OpenFlareAccessLogValueCount, error) {
rows, err := analyticsrepo.ValueCountsNodeAccessLogs(ctx, toNodeAccessLogFilter(query), column, limit)
if err != nil {
return nil, err
}
out := make([]model.OpenFlareAccessLogValueCount, len(rows))
for i, r := range rows {
out[i] = model.OpenFlareAccessLogValueCount{Value: r.Value, Count: int64(r.Count)}
}
return out, nil
}
func (s *clickhouseLogStore) NodeAggregates(ctx context.Context, query model.OpenFlareAccessLogQuery) ([]model.OpenFlareAccessLogNodeAggregate, error) {
rows, err := analyticsrepo.NodeAggregatesNodeAccessLogs(ctx, toNodeAccessLogFilter(query))
if err != nil {
return nil, err
}
out := make([]model.OpenFlareAccessLogNodeAggregate, len(rows))
for i, r := range rows {
out[i] = model.OpenFlareAccessLogNodeAggregate{NodeID: r.NodeID, RequestCount: int64(r.RequestCount), ErrorCount: int64(r.ErrorCount), UniqueIPCount: int64(r.UniqueIPCount)}
}
return out, nil
}
func (s *clickhouseLogStore) DeleteAll(ctx context.Context) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
return analyticsrepo.DeleteAllNodeAccessLogs(ctx)
}
func (s *clickhouseLogStore) DeleteBefore(ctx context.Context, cutoff time.Time) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
return analyticsrepo.DeleteNodeAccessLogsBefore(ctx, cutoff)
}
func (s *clickhouseLogStore) DeleteByNodeBefore(ctx context.Context, nodeID string, before time.Time) (int64, error) {
if err := s.ensureWritable(ctx); err != nil {
return 0, err
}
return analyticsrepo.DeleteNodeAccessLogsByNodeBefore(ctx, nodeID, before)
}
// ---- ObservabilityStore(entry=ensureWritable+hook;flush/query/delete 委托 analyticsrepo)
// InsertMetricSnapshot / ListMetricSnapshots / DeleteAllMetricSnapshots / DeleteMetricSnapshotsBefore / BatchInsertNodeMetricSnapshots
// ...(同构,参照旧 clickhouseObservabilityStore 委托)
// ---- UserAccessLogStore
// BatchInsert -> analyticsrepo.BatchInsert
// Count/List -> analyticsrepo.CountAccessLogs / ListAccessLogs
// GetDailyTrend / GetBrowserDistribution / GetTopActiveUsers -> analyticsrepo.GetDailyTrend / GetBrowserDistribution / GetTopActiveUsers
转换函数
toAnalyticsNodeAccessLog/fromAnalyticsNodeAccessLogs/toNodeAccessLogFilter/toAnalyticsNodeMetricSnapshot等集中放postgres_store.go或本文件共享区域(两个实现共用)。
- Step 3: 删除旧 store 文件——删
internal/repository/openflare_access_log_store.go、internal/repository/openflare_observability_store.go;其中的 memory store 测试替身迁到logstore/memory_store_test.go(保留NewMemoryAccessLogStore等价物供 repository 测试)。 - Step 4: 编译 + 测试
go build ./internal/...;go test ./internal/repository/...修复引用。 - Step 5: 提交
git add internal/repository/logstore/ internal/repository/ && git commit -m "refactor(logstore): wrap ClickHouse analytics repo behind interface"
Task 6: repository 公开函数改委托 logstore
Files:
- Modify:
internal/repository/openflare_access_log.go(函数体改为logstore.Active(ctx)委托) - Modify:
internal/repository/openflare_observability.go(同上)
Interfaces:
-
Consumes:
logstore.Active、logstore.Store字段。 -
Produces: 保留原公开函数签名,行为不变(CH 激活时与现状一致)。
-
Step 1: 改写
openflare_access_log.go各函数
package repository
import (
"context"
"time"
"github.com/Rain-kl/Wavelet/internal/model"
"github.com/Rain-kl/Wavelet/internal/model/analytics" // 若类型别名仍需要
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
)
// ListOpenFlareAccessLogs lists access logs matching the query.
func ListOpenFlareAccessLogs(ctx context.Context, query model.OpenFlareAccessLogQuery) ([]*model.OpenFlareAccessLog, error) {
s, err := logstore.Active(ctx)
if err != nil {
return nil, err
}
return s.AccessLogs.List(ctx, query)
}
对同文件其余函数(ListOpenFlareAccessLogWAFIPAggregates、InsertOpenFlareAccessLogsBatch、CountOpenFlareAccessLogs、TrafficSummaryOpenFlareAccessLogs、RegionCountsOpenFlareAccessLogs、BucketAggregates*、CountBuckets*、BucketDimensions*、IPAggregates*、IPSummaries*、CountIPSummaries*、IPTrend*、ValueCounts*、NodeAggregates*、Delete*)逐一委托到 s.AccessLogs 对应方法;InsertOpenFlareAccessLogsBatch → s.AccessLogs.InsertBatch。保留行类型别名(openFlareAccessLogBucketAggregateRow 等)供调用方编译。
- Step 2: 改写
openflare_observability.go——InsertOpenFlareMetricSnapshot→s.Observability.InsertMetricSnapshot;ListMetricSnapshots*/Delete*同理;健康事件(ReconcileOpenFlareHealthEvents等主库表逻辑)保持原实现不动。 - Step 3: 编译 + 测试
go build ./internal/...、go test ./internal/repository/...(旧测试若引用 memory store 替换为 logstore 测试替身)。 - Step 4: 提交
git add internal/repository/ && git commit -m "refactor(repository): delegate log CRUD to logstore"
Task 7: import-lint 测试(代码级约束验收)
Files:
-
Create:
internal/repository/logstore/imports_test.go -
Step 1: 写测试
package logstore
import (
"os/exec"
"strings"
"testing"
)
// forbiddenImports 上层应用禁止直接触碰的底层日志实现。
var forbiddenImports = []string{
"github.com/Rain-kl/Wavelet/internal/repository/analytics",
}
// allowedInfraPersistence 允许 apps 引入的 infra/persistence 子包。
// batchwriter=批量写入框架;idgen=snowflake ID 生成工具(apps 合法使用,非日志后端访问)。
var allowedInfraPersistence = []string{
"github.com/Rain-kl/Wavelet/internal/infra/persistence/batchwriter",
"github.com/Rain-kl/Wavelet/internal/infra/persistence/idgen",
}
func TestAppsMustNotImportLogBackendDirectly(t *testing.T) {
t.Chdir("../../..")
out, err := exec.Command("go", "list", "-test", "-f", `{{.ImportPath}} {{join .Imports " "}}`, "./internal/apps/...").Output()
if err != nil {
t.Fatalf("go list: %v", err)
}
for _, line := range strings.Split(string(out), "\n") {
fields := strings.Fields(line)
if len(fields) == 0 {
continue
}
pkg := fields[0]
if !strings.HasPrefix(pkg, "github.com/Rain-kl/Wavelet/internal/apps") {
continue
}
for _, imp := range fields[1:] {
for _, forbidden := range forbiddenImports {
if imp == forbidden && !allowedAnalyticsDelegation[pkg] {
t.Errorf("%s must not import forbidden log backend %s", pkg, forbidden)
}
}
if strings.HasPrefix(imp, "github.com/Rain-kl/Wavelet/internal/infra/persistence/") {
allowed := false
for _, a := range allowedInfraPersistence {
if imp == a || strings.HasPrefix(imp, a+"/") {
allowed = true
break
}
}
if !allowed {
t.Errorf("%s must not import infra/persistence subpackage directly: %s", pkg, imp)
}
}
}
}
}
说明:
go list -deps在测试工作目录执行,先t.Chdir到仓库根(../../..)再运行,避免依赖go test的临时目录。若internal/apps/admin/logs等仍 import analyticsrepo,本测试失败——正好驱动 Task 9。
- Step 2: 运行
go test ./internal/repository/logstore/ -run TestAppsMustNotImportLogBackendDirectly -v——预期当前失败(列出违规包)。 - Step 3: 暂不提交——本测试在 apps 改造完成前保持 RED(预期失败列出违规包)。Task 9 完成 apps 改造、本测试转绿后,随 Task 9 一并提交(提交信息:
test(logstore): enforce apps must not import log backend directly)。
Task 8: 系统配置 key + 启动校验 + key 保护
Files:
- Modify:
internal/model/system_configs.go(新增 key 常量) - Modify:
internal/platform/bootstrap/bootstrap.go(Init加校验与 seed) - Modify:
internal/apps/admin/system_config/routers.go(受保护 key 拒绝修改) - Modify:
internal/apps/openflare/option/validate.go(同) - Create:
internal/platform/bootstrap/bootstrap_test.go(追加校验测试)
Interfaces:
-
Consumes:
config.Config.Database.Enabled、config.Config.ClickHouse.Enabled、repository.GetSystemConfigByKey、repository.UpdateSystemConfigFields。 -
Produces:
model.ConfigKeyLogDatabase = "log_database"、model.ConfigKeyLogDBMigration = "log_db_migration"、model.ConfigKeyLogRetentionDaysPostgres = "log_retention_days_postgres"、model.ConfigKeyLogRetentionDaysSQLite = "log_retention_days_sqlite"、model.ConfigKeyLogRetentionDaysClickHouse = "log_retention_days_clickhouse"。 -
Step 1: 新增 key 常量(system_configs.go)
// 日志数据库解耦
ConfigKeyLogDatabase = "log_database" // 当前日志主库:postgres|sqlite|clickhouse(仅迁移任务写入)
ConfigKeyLogDBMigration = "log_db_migration" // 迁移冻结标记:"migrating" 或空
ConfigKeyLogRetentionDaysPostgres = "log_retention_days_postgres" // PostgreSQL 日志保留天数
ConfigKeyLogRetentionDaysSQLite = "log_retention_days_sqlite" // SQLite 日志保留天数
ConfigKeyLogRetentionDaysClickHouse = "log_retention_days_clickhouse" // ClickHouse 日志保留天数
- Step 2: bootstrap 校验 + seed(bootstrap.go
Init内,initRuntimeOnce.Do开头)
// validateAndSeedLogDatabase 校验日志主库标记与运行配置的一致性,首次启动 seed。
func validateAndSeedLogDatabase(ctx context.Context) error {
cfg, err := repository.GetSystemConfigByKey(ctx, model.ConfigKeyLogDatabase)
if err != nil {
return fmt.Errorf("读取日志主库配置失败: %w", err)
}
current := cfg.Value
if current == "" {
// 首次启动 seed:CH 启用 → clickhouse;否则随主库。
current = "sqlite"
if config.Config.Database.Enabled {
current = "postgres"
}
if config.Config.ClickHouse.Enabled {
current = "clickhouse"
}
if err := repository.UpdateSystemConfigFields(ctx, &model.SystemConfig{Key: model.ConfigKeyLogDatabase}, map[string]any{"value": current}); err != nil {
return fmt.Errorf("初始化日志主库配置失败: %w", err)
}
return nil
}
switch current {
case "clickhouse":
if !config.Config.ClickHouse.Enabled {
return errors.New("当前日志主库为 ClickHouse 但 ClickHouse 未启用。请先重新启用 ClickHouse 配置并启动,在任务管理运行『切换日志数据库』迁移到 PostgreSQL/SQLite 后再禁用 ClickHouse")
}
case "postgres":
if !config.Config.Database.Enabled {
return errors.New("当前日志主库为 PostgreSQL 但 PostgreSQL 未启用(当前为 SQLite 主库)。请运行『切换日志数据库』迁回 SQLite 或启用 PostgreSQL")
}
case "sqlite":
if config.Config.Database.Enabled {
return errors.New("当前日志主库为 SQLite 但当前主库为 PostgreSQL。请运行『切换日志数据库』迁移到 PostgreSQL")
}
default:
return fmt.Errorf("未知的日志主库配置: %s", current)
}
return nil
}
在 Init 的 initRuntimeOnce.Do 内最先调用:if err := validateAndSeedLogDatabase(ctx); err != nil { logger.ErrorF(...); log.Fatalf(...) }(或按项目既有致命启动错误处理方式)。
- Step 3: key 保护(admin system-config 更新路径)
internal/apps/admin/system_config/routers.go 的 UpdateSystemConfig 与 internal/apps/openflare/option/validate.go 增加:
// protectedConfigKeys 仅允许内部(迁移任务/bootstrap)写入的 key。
var protectedConfigKeys = map[string]bool{
model.ConfigKeyLogDatabase: true,
model.ConfigKeyLogDBMigration: true,
}
func isProtectedConfigKey(key string) bool { return protectedConfigKeys[key] }
更新处理:命中保护 key 时返回业务错误(response.AbortBadRequest(c, "该配置项由系统任务管理,禁止手动修改")),且不写库。
- Step 4: 单测——
bootstrap_test.go三态校验(clickhouse 未启用 / postgres 但 sqlite 主库 / sqlite 但 postgres 主库)各自返回明确错误;seed 缺失时写入正确默认值。 - Step 5: 运行
go test ./internal/platform/bootstrap/ ./internal/model/ ./internal/apps/admin/system_config/PASS。 - Step 6: 提交
git add internal/model/system_configs.go internal/platform/bootstrap/ internal/apps/admin/system_config/ internal/apps/openflare/option/ && git commit -m "feat(config): log database marker, boot validation, and protected keys"
Task 9: apps 层改走 logstore(消除 import-lint 违规)
Files:
- Modify:
internal/apps/risk_control/logics.go、internal/apps/openflare/chwriter/writer.go - Modify:
internal/apps/openflare/tasks/database_cleanup.go(本任务只改 import;清理合并到 M2) - Modify:
internal/apps/openflare/observability/access_log_logics.go(仅解析 helper 保留 analyticsrepo 合法引用则不动;若违规则把ParseDeviceType/ParseBrowserName/ParseOSName迁到model/analytics或internal/util) - Modify:
internal/apps/admin/logs/routers.go、internal/apps/admin/status/clickhouse.go - Test:
internal/repository/logstore/imports_test.go(回归)
Interfaces:
-
Consumes:
logstore.Active、logstore.Migrating、logstore.ErrMigrating、logstore.SetAccessLogHooks/SetObservabilityHooks。 -
Step 1: chwriter flush func 改为 logstore
writer.go 中 5 处 analyticsrepo.BatchInsertNode* → logstore.Active(ctx).Observability/AccessLogs 对应 flush 方法(或包级 helper):
func flushNodeAccessLogs(ctx context.Context, rows []analyticsmodel.NodeAccessLog) error {
s, err := logstore.Active(ctx)
if err != nil {
return err
}
return s.AccessLogs.BatchInsertNodeAccessLogs(ctx, rows)
}
Init 内 if !config.Config.ClickHouse.Enabled { return } 改为 if logstore.Active(ctx) == nil ... 或直接始终初始化 writer(writer flush 走 logstore,激活库由 logstore 决定);wireModelInsertHooks 改为调用 logstore.SetAccessLogHooks/logstore.SetObservabilityHooks。
- Step 2: risk_control flush 与冻结
logics.go:flush func 中 analyticsrepo.BatchInsert → logstore.Active(ctx).UserAccessLogs.BatchInsert;InitLogWriter 的 CH 开关条件移除,改为由 logstore 激活库决定(PG/SQLite 也启用该 writer);middleware 入队前:
if logstore.Migrating(c.Request.Context()) {
logger.WarnF(c.Request.Context(), "[RiskControl] log DB migrating, skip audit log")
return // 不阻断业务请求
}
- Step 3: admin/logs 改走 logstore
routers.go 中 analyticsrepo.ListAccessLogs/CountAccessLogs/GetDailyTrend/GetBrowserDistribution/GetTopActiveUsers → logstore.Active(ctx).UserAccessLogs.*;config.Config.ClickHouse.Enabled || !db.ChConnReady() 的守卫改为按激活库判断(logstore.Active(ctx) 成功即可用),错误文案从「ClickHouse 存储服务未启用」改为「日志存储未启用」。
- Step 4: admin/status 端点骨架
clickhouse.go 改为读取 logstore.Active 与激活库名,返回统一结构(M3 Task 16 完成前端与完整字段):
type LogDatabaseStatus struct {
ActiveDatabase string `json:"active_database"`
Migration string `json:"migration"` // idle | migrating
RetentionDays map[string]int `json:"retention_days"`
AvailableTargets []string `json:"available_targets"`
}
CH 激活时保留 GetClickHouseOperationalStats 与 collectBatchWriterStats。
- Step 5: database_cleanup.go 临时保留 import 但标记 TODO(M2 Task 13 迁移)——若 import-lint 在 Task 7 已注册,本任务先让
database_cleanup.go改为经 repository 公开函数(其逻辑已走 logstore),并同步access_log_logics.go解析 helper(迁ParseBrowserName等为model/analytics纯函数,analyticsrepo 内部复用)。 - Step 6: 运行 import-lint 回归
go test ./internal/repository/logstore/ -run TestAppsMustNotImportLogBackendDirectly -v期望 PASS。 - Step 7: 全量编译
go build ./internal/...、go test ./internal/apps/...修复。 - Step 8: 提交
git add internal/apps/ && git commit -m "refactor(apps): route log reads/writes through logstore"
Task 10: bootstrap 装配 logstore
Files:
- Modify:
internal/platform/bootstrap/bootstrap.go - Modify:
internal/cmd/all.go、api.go、worker.go、root.go(如有必要)
Interfaces:
-
Consumes:
logstore.SetConfigReader、logstore.Init。 -
Produces: 运行期
logstore激活 store 可解析。 -
Step 1: 装配 config reader + Init
bootstrap.Init 的 initRuntimeOnce.Do 内、校验之后:
logstore.SetConfigReader(func(ctx context.Context, key string) (string, error) {
cfg, err := repository.GetSystemConfigByKey(ctx, key)
if err != nil {
return "", err
}
return cfg.Value, nil
})
logstore.Init(ctx)
- Step 2: worker 进程也需要 Init——确认
cmd/worker.go与cmd/all.go都调用bootstrap.Init(现 API 分支启动 writer;worker 迁移任务需能读配置与激活 store,logstore.Init必须在两种进程都执行)。 - Step 3: 测试
go test ./internal/platform/bootstrap/;go build ./cmd/...。 - Step 4: 提交
git add internal/platform/bootstrap/ internal/cmd/ && git commit -m "feat(bootstrap): wire logstore config reader and init"
M2:建表与清理
Task 10b: 小时级聚合读经 logstore(PG 实时计算 / CH 读 rollup 表)
Files:
- Modify:
internal/repository/logstore/logstore.go(ObservabilityStore增 3 个方法) - Modify:
internal/repository/logstore/postgres_store.go(PG 按小时从原始表实时聚合) - Modify:
internal/repository/logstore/clickhouse_store.go(委托 analyticsrepo rollup 读 + 现有 raw 兜底逻辑) - Modify:
internal/repository/openflare_observability.go(3 个ListOpenFlare*HourlySince改委托 logstore) - Modify:
internal/repository/logstore/imports_test.go(若internal/repository不再直接 import analyticsrepo,可移除其对allowedAnalyticsDelegation的豁免)
Interfaces:
-
Consumes: Task 3/4 GORM store、Task 5 CH store、
analyticsrepo.ListNodeTrafficHourly/ListAccessLogHourly/ListNodeMetricHourly及mergeNodeMetricHourlyPreferRollup/listNodeMetricHourlyFromRaw语义。 -
Produces:
ObservabilityStore.ListTrafficHourly(ctx, nodeID, since) ([]analyticsmodel.NodeTrafficHourly, error)、ListAccessLogHourly(...)、ListMetricHourly(...)。 -
Step 1: 接口加方法(logstore.go)
-
Step 2: CH 实现委托 analyticsrepo(rollup 表 + raw 兜底,逐行复制现有逻辑)
-
Step 3: PG 实现按小时实时聚合——
date_trunc('hour', logged_at/captured_at)分组(方言timeBucketSQL(col, 3600)复用),请求/错误/字节数与 CH rollup 同字段;ListMetricHourly用avg(cpu)/max-min 计数器近似同 CHmergeNodeMetricHourlyPreferRollup口径。 -
Step 4: repository 门面 3 个函数改委托 logstore;若门面不再 import analyticsrepo,收紧 lint 豁免。
-
Step 5: 测试——PG/SQLite 实时聚合与 CH rollup 口径一致性(sqlite 写原始行断言小时桶输出);CH 委托回归。
-
Step 6: 提交
git add internal/repository/ && git commit -m "feat(logstore): hourly rollup reads with PG real-time aggregation"
Task 11: goose 双方言建表迁移(6 张原始日志表)
Files:
- Create:
internal/infra/persistence/migrator/goose/postgres/202608080001_create_log_tables.sql - Create:
internal/infra/persistence/migrator/goose/sqlite/202608080001_create_log_tables.sql
Interfaces:
-
Consumes: database-migration 技能规则(双方言同版本号、无物理外键、默认值与 Go 零值一致)。
-
Produces: PG/SQLite 各 6 张日志表(
w_user_access_logs、of_node_access_logs、of_node_metric_snapshots、of_node_edge_health、of_node_obs_frps、of_node_obs_frpc)。 -
Step 1: PG 建表(含分区)
-- +goose Up
-- 节点访问日志:按月 RANGE 分区,复合主键 (id, logged_at) 满足分区键进唯一索引要求。
CREATE TABLE of_node_access_logs (
id BIGINT NOT NULL,
node_id VARCHAR(64) NOT NULL,
logged_at TIMESTAMPTZ NOT NULL,
remote_addr VARCHAR(128) NOT NULL DEFAULT '',
region VARCHAR(128) NOT NULL DEFAULT '',
host VARCHAR(255) NOT NULL DEFAULT '',
path VARCHAR(2048) NOT NULL DEFAULT '',
user_agent TEXT NOT NULL DEFAULT '',
cache_status VARCHAR(64) NOT NULL DEFAULT '',
status_code INTEGER NOT NULL DEFAULT 0,
bytes_sent BIGINT NOT NULL DEFAULT 0,
request_length BIGINT NOT NULL DEFAULT 0,
request_time_ms INTEGER NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (id, logged_at)
) PARTITION BY RANGE (logged_at);
CREATE INDEX idx_of_node_access_logs_node_id ON of_node_access_logs (node_id, logged_at DESC);
CREATE INDEX idx_of_node_access_logs_host ON of_node_access_logs (host, logged_at DESC);
CREATE INDEX idx_of_node_access_logs_remote_addr ON of_node_access_logs (remote_addr, logged_at DESC);
CREATE INDEX idx_of_node_access_logs_status_code ON of_node_access_logs (status_code, logged_at DESC);
-- 用户访问日志:按月分区。
CREATE TABLE w_user_access_logs (
id BIGINT NOT NULL,
user_id BIGINT NOT NULL DEFAULT 0,
path VARCHAR(2048) NOT NULL DEFAULT '',
method VARCHAR(16) NOT NULL DEFAULT '',
ip VARCHAR(128) NOT NULL DEFAULT '',
user_agent TEXT NOT NULL DEFAULT '',
headers TEXT NOT NULL DEFAULT '',
status INTEGER NOT NULL DEFAULT 0,
latency BIGINT NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (id, created_at)
) PARTITION BY RANGE (created_at);
CREATE INDEX idx_w_user_access_logs_user_id ON w_user_access_logs (user_id, created_at DESC);
-- 可观测 4 表:普通表 + 索引。
CREATE TABLE of_node_metric_snapshots (
id BIGINT NOT NULL PRIMARY KEY,
node_id VARCHAR(64) NOT NULL,
captured_at TIMESTAMPTZ NOT NULL,
cpu_usage_percent DOUBLE PRECISION NOT NULL DEFAULT 0,
memory_used_bytes BIGINT NOT NULL DEFAULT 0,
memory_total_bytes BIGINT NOT NULL DEFAULT 0,
storage_used_bytes BIGINT NOT NULL DEFAULT 0,
storage_total_bytes BIGINT NOT NULL DEFAULT 0,
disk_read_bytes BIGINT NOT NULL DEFAULT 0,
disk_write_bytes BIGINT NOT NULL DEFAULT 0,
network_rx_bytes BIGINT NOT NULL DEFAULT 0,
network_tx_bytes BIGINT NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX idx_of_node_metric_snapshots_node ON of_node_metric_snapshots (node_id, captured_at DESC);
CREATE TABLE of_node_edge_health (
id BIGINT NOT NULL PRIMARY KEY,
node_id VARCHAR(64) NOT NULL,
captured_at TIMESTAMPTZ NOT NULL,
status VARCHAR(64) NOT NULL DEFAULT '',
connections BIGINT NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX idx_of_node_edge_health_node ON of_node_edge_health (node_id, captured_at DESC);
CREATE TABLE of_node_obs_frps (
id BIGINT NOT NULL PRIMARY KEY,
node_id VARCHAR(64) NOT NULL,
captured_at TIMESTAMPTZ NOT NULL,
frps_connections INTEGER NOT NULL DEFAULT 0,
frps_proxy_count INTEGER NOT NULL DEFAULT 0,
frps_client_count INTEGER NOT NULL DEFAULT 0,
frps_proxies TEXT NOT NULL DEFAULT '',
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX idx_of_node_obs_frps_node ON of_node_obs_frps (node_id, captured_at DESC);
CREATE TABLE of_node_obs_frpc (
id BIGINT NOT NULL PRIMARY KEY,
node_id VARCHAR(64) NOT NULL,
captured_at TIMESTAMPTZ NOT NULL,
tunnel_status VARCHAR(16) NOT NULL DEFAULT '',
connected_relays_count INTEGER NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX idx_of_node_obs_frpc_node ON of_node_obs_frpc (node_id, captured_at DESC);
-- 分区预建:创建未来 3 个月与当前月分区(当月及下两个月)。
DO $$
DECLARE
d date;
BEGIN
FOR d IN SELECT generate_series(date_trunc('month', now())::date, (date_trunc('month', now()) + interval '2 months')::date, interval '1 month')::date
LOOP
EXECUTE format('CREATE TABLE IF NOT EXISTS of_node_access_logs_%s PARTITION OF of_node_access_logs FOR VALUES FROM (%L) TO (%L)',
to_char(d, 'YYYYMM'), d, d + interval '1 month');
EXECUTE format('CREATE TABLE IF NOT EXISTS w_user_access_logs_%s PARTITION OF w_user_access_logs FOR VALUES FROM (%L) TO (%L)',
to_char(d, 'YYYYMM'), d, d + interval '1 month');
END LOOP;
END $$;
-- +goose Down
DROP TABLE IF EXISTS w_user_access_logs;
DROP TABLE IF EXISTS of_node_access_logs;
DROP TABLE IF EXISTS of_node_metric_snapshots;
DROP TABLE IF EXISTS of_node_edge_health;
DROP TABLE IF EXISTS of_node_obs_frps;
DROP TABLE IF EXISTS of_node_obs_frpc;
- Step 2: SQLite 建表(普通表,同语义)
-- +goose Up
CREATE TABLE IF NOT EXISTS of_node_access_logs (
id INTEGER PRIMARY KEY,
node_id TEXT NOT NULL DEFAULT '',
logged_at DATETIME NOT NULL,
remote_addr TEXT NOT NULL DEFAULT '',
region TEXT NOT NULL DEFAULT '',
host TEXT NOT NULL DEFAULT '',
path TEXT NOT NULL DEFAULT '',
user_agent TEXT NOT NULL DEFAULT '',
cache_status TEXT NOT NULL DEFAULT '',
status_code INTEGER NOT NULL DEFAULT 0,
bytes_sent INTEGER NOT NULL DEFAULT 0,
request_length INTEGER NOT NULL DEFAULT 0,
request_time_ms INTEGER NOT NULL DEFAULT 0,
created_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP
);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_node ON of_node_access_logs (node_id, logged_at DESC);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_host ON of_node_access_logs (host, logged_at DESC);
CREATE INDEX IF NOT EXISTS idx_of_node_access_logs_remote_addr ON of_node_access_logs (remote_addr, logged_at DESC);
-- 其余 5 表同构(w_user_access_logs 主键 id;可观测表 id INTEGER PRIMARY KEY + (node_id, captured_at DESC) 索引)
-- +goose Down
DROP TABLE IF EXISTS of_node_access_logs;
DROP TABLE IF EXISTS w_user_access_logs;
DROP TABLE IF EXISTS of_node_metric_snapshots;
DROP TABLE IF EXISTS of_node_edge_health;
DROP TABLE IF EXISTS of_node_obs_frps;
DROP TABLE IF EXISTS of_node_obs_frpc;
- Step 3: 验证 goose
go test ./internal/infra/persistence/migrator(空库 Up 全量)。 - Step 4: 提交
git add internal/infra/persistence/migrator/goose/ && git commit -m "feat(migrate): create log tables in postgres and sqlite"
Task 12: 保留时间配置 + 旧 key 下线
Files:
- Create:
internal/infra/persistence/migrator/goose/postgres/202608080002_log_retention_configs.sql - Create:
internal/infra/persistence/migrator/goose/sqlite/202608080002_log_retention_configs.sql - Modify:
internal/model/system_configs.go(删除旧 key 常量或标记废弃) - Modify:
internal/testhelper/test_helper.go(seed 同步)
Interfaces:
-
Produces: 3 个 business 配置(默认 90);旧
database_auto_cleanup_enabled/database_auto_cleanup_retention_days从system_configs删除。 -
Step 1: PG 迁移
-- +goose Up
INSERT INTO system_configs (key, value, type, visibility, description, created_at, updated_at)
VALUES
('log_retention_days_postgres', '90', 'business', 0, 'PostgreSQL 日志保留天数(访问日志与可观测统一)', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP),
('log_retention_days_sqlite', '90', 'business', 0, 'SQLite 日志保留天数', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP),
('log_retention_days_clickhouse','90', 'business', 0, 'ClickHouse 日志保留天数', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP)
ON CONFLICT (key) DO NOTHING;
DELETE FROM system_configs WHERE key IN ('database_auto_cleanup_enabled', 'database_auto_cleanup_retention_days');
-- +goose Down
INSERT INTO system_configs (key, value, type, visibility, description, created_at, updated_at)
VALUES
('database_auto_cleanup_enabled', 'true', 'business', 0, '数据库自动清理开关', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP),
('database_auto_cleanup_retention_days', '30', 'business', 0, '数据库保留天数', CURRENT_TIMESTAMP, CURRENT_TIMESTAMP)
ON CONFLICT (key) DO NOTHING;
DELETE FROM system_configs WHERE key IN ('log_retention_days_postgres', 'log_retention_days_sqlite', 'log_retention_days_clickhouse');
- Step 2: SQLite 同版本号镜像(
INSERT OR IGNORE/DELETE,语义一致)。 - Step 3: model 常量更新——旧 key 常量删除;
validate.go中validateDatabaseCleanupOption替换为validateLogRetentionOption(3 个新 key,值 ≥1 整数)。 - Step 4: testhelper seed 同步——
seedDefaultConfigs增 3 个新 key、删旧 key(含公共 key 列表如有)。 - Step 5: 验证
go test ./internal/infra/persistence/migrator ./internal/apps/config ./internal/apps/admin/system_config ./internal/testhelper。 - Step 6: 提交
git add internal/ && git commit -m "feat(config): per-store log retention settings, drop legacy cleanup config"
Task 13: CleanupStore + system_cleanup 日志清理步骤 + PG 分区预建
含 Task 11 审查跟进:PG 分区表仅在建表迁移时预建当前+2 月;
CleanupExpired每次运行时必须先确保「当前月 + 未来 2 个月」的分区存在(幂等CREATE TABLE IF NOT EXISTS ... PARTITION OF),否则 3 个月后新写入会报 "no partition of relation found"。在CleanupStore(或 logstore 包内EnsurePartitions(ctx))实现,PG 方言执行、SQLite/CH 为 no-op;system_cleanup每日调用保证分区持续存在。
Files:
- Create:
internal/repository/logstore/cleanup.go - Modify:
internal/apps/upload/task/cleanup.go(追加日志清理步骤) - Create:
internal/repository/logstore/cleanup_test.go
Interfaces:
-
Consumes:
model.ConfigKeyLogRetentionDays*、logstore.Active。 -
Produces:
CleanupExpired(ctx) (*CleanupSummary, error)(repository 层入口,system_cleanup调用)。 -
Step 1: cleanup.go
package logstore
import (
"context"
"fmt"
"strconv"
"time"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
)
// CleanupSummary 汇总本次清理结果。
type CleanupSummary struct {
ActiveDatabase string `json:"active_database"`
RetentionDays int `json:"retention_days"`
Deleted int64 `json:"deleted"`
Tables []string `json:"tables"`
}
// retentionDaysForActive 按当前激活库读取保留天数(默认 90)。
func retentionDaysForActive(ctx context.Context) int {
key := model.ConfigKeyLogRetentionDaysPostgres
if dbName, _ := resolveDatabase(ctx); dbName == "sqlite" {
key = model.ConfigKeyLogRetentionDaysSQLite
} else if dbName == "clickhouse" {
key = model.ConfigKeyLogRetentionDaysClickHouse
}
v, err := getConfig(ctx, key)
if err != nil {
return 90
}
days, perr := strconv.Atoi(v)
if perr != nil || days <= 0 {
return 90
}
return days
}
// CleanupExpired 按当前激活库保留天数清理过期日志(每日由 system_cleanup 调用)。
func CleanupExpired(ctx context.Context) (*CleanupSummary, error) {
s, err := Active(ctx)
if err != nil {
return nil, err
}
days := retentionDaysForActive(ctx)
cutoff := time.Now().AddDate(0, 0, -days)
summary := &CleanupSummary{RetentionDays: days, Tables: []string{}}
summary.ActiveDatabase, _ = resolveDatabase(ctx)
if err := cleanupTable(ctx, s, "node_access_logs", func() (int64, error) {
return s.AccessLogs.DeleteBefore(ctx, cutoff)
}, summary); err != nil {
return nil, err
}
if err := cleanupTable(ctx, s, "metric_snapshots", func() (int64, error) {
return s.Observability.DeleteMetricSnapshotsBefore(ctx, cutoff)
}, summary); err != nil {
return nil, err
}
// edge_health / obs_frps / obs_frpc 同构
return summary, nil
}
func cleanupTable(ctx context.Context, s *Store, name string, fn func() (int64, error), summary *CleanupSummary) error {
n, err := fn()
if err != nil {
return fmt.Errorf("cleanup %s: %w", name, err)
}
summary.Deleted += n
summary.Tables = append(summary.Tables, name)
return nil
}
PG 实现优化(可选,首版用 DeleteBefore 即可):
DeleteBefore在 PG 分区表上命中logged_at分区键,按月 DROP 整分区后再 DELETE 不满月——M1 Task 3 的DeleteBefore已按logged_at < cutoff实现,满足正确性;后续再优化为 DROP PARTITION。CH 实现:DeleteNodeAccessLogsBefore已做 TTL materialize;保留天数变化时clickhouseLogStore.DeleteBefore增加ALTER TABLE ... MODIFY TTL(见 M4 优化项,可延后)。
- Step 2: system_cleanup 追加步骤(upload/task/cleanup.go)
在现有清理步骤之后追加:
task.AppendLog(ctx, "开始清理过期日志(按当前日志库保留天数)...")
summary, err := logstore.CleanupExpired(ctx)
if err != nil {
task.AppendLog(ctx, "清理过期日志失败: %v", err)
} else if summary.Deleted == 0 {
task.AppendLog(ctx, "没有需要清理的过期日志 (保留 %d 天)", summary.RetentionDays)
} else {
task.AppendLog(ctx, "日志清理完成:保留 %d 天,删除 %d 条", summary.RetentionDays, summary.Deleted)
}
(internal/apps/upload/task/cleanup.go import internal/repository/logstore——upload/task 属 apps 层,import logstore 合法。)
- Step 3: 单测(cleanup_test.go)——sqlite store 写入 40 天前/昨天各 1 条,
CleanupExpired用SetConfigReader注入log_retention_days_sqlite=30,断言 40 天前的被删、昨天的保留。 - Step 4: 运行
go test ./internal/repository/logstore/ ./internal/apps/upload/task/。 - Step 5: 提交
git add internal/repository/logstore/ internal/apps/upload/task/ && git commit -m "feat(cleanup): log retention cleanup in system_cleanup task"
Task 14: 下线 of_database_auto_cleanup
Files:
- Create:
internal/infra/persistence/migrator/goose/postgres/202608080003_drop_database_cleanup_schedule.sql、sqlite/202608080003_... - Modify:
internal/apps/openflare/async_tasks.go(删除DatabaseAutoCleanupTask/DatabaseAutoCleanupMeta/DatabaseAutoCleanupHandler) - Modify:
internal/infra/task/handlers/register.go(注销) - Modify:
internal/apps/openflare/tasks/database_cleanup.go(删除;清理能力已并入 system_cleanup)
Interfaces:
-
Consumes: Task 13 完成。
-
Produces:
of_database_auto_cleanup从 schedule 与任务注册中消失。 -
Step 1: goose 删 schedule
-- +goose Up
DELETE FROM w_schedules WHERE task_type = 'of_database_auto_cleanup';
-- +goose Down
INSERT INTO w_schedules (id, name, task_type, cron, payload, is_active, created_at, updated_at)
VALUES (102, 'OpenFlare 可观测数据自动清理', 'of_database_auto_cleanup', '0 3 * * *', '{}', TRUE, CURRENT_TIMESTAMP, CURRENT_TIMESTAMP)
ON CONFLICT (id) DO NOTHING;
- Step 2: 注销任务与删除文件——
register.go移除对应两行;async_tasks.go删除常量/元数据/Handler;删除tasks/database_cleanup.go。 - Step 3: 前端清理——搜索前端对
of_database_auto_cleanup/database_auto_cleanup_*引用并删除(任务页硬编码列表如有)。 - Step 4: 验证
go build ./internal/...、go test ./internal/infra/persistence/migrator ./internal/infra/task/。 - Step 5: 提交
git add internal/ frontend/ && git commit -m "chore(cleanup): decommission of_database_auto_cleanup task and schedule"
M3:迁移任务与展示
Task 15: 「切换日志数据库」任务 Handler
Files:
- Create:
internal/apps/openflare/tasks/log_db_switch.go - Create:
internal/apps/openflare/tasks/log_db_switch_test.go - Modify:
internal/apps/openflare/async_tasks.go(注册元数据) - Modify:
internal/infra/task/handlers/register.go(注册 Handler)
Interfaces:
-
Consumes:
logstore.Active/logstore.Migrating、repository.UpdateSystemConfigFields、model.ConfigKeyLogDatabase/ConfigKeyLogDBMigration、analyticsmodel.*、config.Config。 -
Produces: Asynq
openflare:log_db_switch,管理类型of_log_db_switch,参数target。 -
Step 1: 元数据(async_tasks.go)
// LogDBSwitchTask 切换日志数据库任务标识。
const (
LogDBSwitchTask = "openflare:log_db_switch"
TaskTypeLogDBSwitch = "of_log_db_switch"
)
var LogDBSwitchMeta = task.TaskMeta{
Type: TaskTypeLogDBSwitch,
AsynqTask: LogDBSwitchTask,
Name: "切换日志数据库",
Description: "复制迁移日志数据并在成功后切换日志主库(期间禁止日志写入)",
SupportsTime: false,
MaxRetry: task.DefaultMaxRetry,
Queue: task.QueueDefault,
Retryable: true,
Params: []task.TaskParam{
{Name: "target", Label: "目标日志库", Type: "string", Required: true,
Placeholder: "postgres|sqlite|clickhouse", Description: "迁移目标:postgres(主库为 PG 时)、sqlite(主库为 SQLite 时)或 clickhouse"},
},
}
- Step 2: Handler(log_db_switch.go)
package tasks
import (
"context"
"encoding/json"
"errors"
"fmt"
"time"
"github.com/Rain-kl/Wavelet/internal/infra/config"
"github.com/Rain-kl/Wavelet/internal/infra/task"
"github.com/Rain-kl/Wavelet/internal/model"
analyticsmodel "github.com/Rain-kl/Wavelet/internal/model/analytics"
"github.com/Rain-kl/Wavelet/internal/repository"
"github.com/Rain-kl/Wavelet/internal/repository/logstore"
"github.com/Rain-kl/Wavelet/pkg/logger"
)
const copyBatchSize = 1000
type logDBSwitchPayload struct {
Target string `json:"target"`
}
// LogDBSwitchHandler 切换日志数据库任务处理器。
type LogDBSwitchHandler struct{}
// ValidatePayload 校验并规范化参数。
func (h *LogDBSwitchHandler) ValidatePayload(payload []byte) ([]byte, error) {
var p logDBSwitchPayload
if err := json.Unmarshal(payload, &p); err != nil {
return nil, fmt.Errorf("参数解析失败: %w", err)
}
p.Target = normalizeTarget(p.Target)
if !validTarget(p.Target) {
return nil, fmt.Errorf("目标日志库不合法: %s", p.Target)
}
out, err := json.Marshal(p)
if err != nil {
return nil, err
}
return out, nil
}
func normalizeTarget(v string) string {
switch v {
case "postgres", "postgresql":
return "postgres"
case "sqlite", "sqlite3":
return "sqlite"
case "clickhouse", "ch":
return "clickhouse"
}
return v
}
func validTarget(v string) bool {
return v == "postgres" || v == "sqlite" || v == "clickhouse"
}
// Execute 执行迁移。
func (h *LogDBSwitchHandler) Execute(ctx context.Context, payload []byte) (*task.TaskResult, error) {
var p logDBSwitchPayload
if err := json.Unmarshal(payload, &p); err != nil {
return nil, fmt.Errorf("参数解析失败: %w", err)
}
p.Target = normalizeTarget(p.Target)
if err := validateSwitch(ctx, p.Target); err != nil {
return nil, err
}
source, _ := currentLogDatabase(ctx)
task.AppendLog(ctx, "开始切换日志数据库:%s -> %s", source, p.Target)
if err := setMigrationFlag(ctx, "migrating"); err != nil {
return nil, err
}
defer func() { _ = setMigrationFlag(ctx, "") }() // 失败也清除,保持源库可写
if err := drainLogWriters(ctx); err != nil {
return nil, fmt.Errorf("排空日志写入队列失败: %w", err)
}
src, err := logstore.Active(ctx)
if err != nil {
return nil, err
}
dst, err := buildTargetStore(ctx, p.Target)
if err != nil {
return nil, err
}
// 清空目标库日志表(幂等重试前提)。
if err := clearTargetLogTables(ctx, dst, p.Target); err != nil {
return nil, err
}
// 逐表复制。
if err := copyAccessLogs(ctx, src, dst); err != nil {
return nil, err
}
if err := copyUserAccessLogs(ctx, src, dst); err != nil {
return nil, err
}
if err := copyObservability(ctx, src, dst); err != nil {
return nil, err
}
// 翻转主库标记。
if err := flipLogDatabase(ctx, p.Target); err != nil {
return nil, err
}
task.AppendLog(ctx, "日志数据库已切换为 %s,写入恢复", p.Target)
return &task.TaskResult{Message: fmt.Sprintf("日志数据库已从 %s 切换为 %s", source, p.Target)}, nil
}
- Step 3: 辅助函数(同文件)
func validateSwitch(ctx context.Context, target string) error {
source, err := currentLogDatabase(ctx)
if err != nil {
return err
}
if source == target {
return errors.New("目标日志库与当前日志库相同,无需迁移")
}
switch target {
case "clickhouse":
if !config.Config.ClickHouse.Enabled {
return errors.New("ClickHouse 未启用,无法迁移到 ClickHouse")
}
case "postgres":
if !config.Config.Database.Enabled {
return errors.New("PostgreSQL 未启用(当前主库为 SQLite),无法迁移到 PostgreSQL")
}
case "sqlite":
if config.Config.Database.Enabled {
return errors.New("当前主库为 PostgreSQL,日志库不能设置为 SQLite")
}
}
return nil
}
func currentLogDatabase(ctx context.Context) (string, error) {
cfg, err := repository.GetSystemConfigByKey(ctx, model.ConfigKeyLogDatabase)
if err != nil {
return "", fmt.Errorf("读取日志主库失败: %w", err)
}
if cfg.Value == "" {
return "", errors.New("日志主库配置为空")
}
return cfg.Value, nil
}
func setMigrationFlag(ctx context.Context, v string) error {
// 必须用 SaveOrUpdateSystemConfig:UpdateSystemConfigFields 缺行时静默 no-op,
// 且不失效 RAM 配置缓存(TTL=-1 永不过期),会导致冻结/翻转不生效、进程间脑裂。
return repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyLogDBMigration, v)
}
func flipLogDatabase(ctx context.Context, target string) error {
return repository.SaveOrUpdateSystemConfig(ctx, model.ConfigKeyLogDatabase, target)
}
// buildTargetStore 构造目标库 Store(不经过 Active 缓存,直接 Build)。
func buildTargetStore(ctx context.Context, database string) (*logstore.Store, error) {
return logstore.Build(ctx, database)
}
func clearTargetLogTables(ctx context.Context, dst *logstore.Store, target string) error {
// 依次清空 6 张表:AccessLogs.DeleteAll、UserAccessLogs.DeleteAll、Observability.DeleteAll*(SQLite/PG 用 DeleteAll;CH 用 TRUNCATE 语义)。
if _, err := dst.AccessLogs.DeleteAll(ctx); err != nil {
return fmt.Errorf("清空目标访问日志失败: %w", err)
}
if _, err := dst.UserAccessLogs.DeleteAll(ctx); err != nil {
return fmt.Errorf("清空目标用户访问日志失败: %w", err)
}
for _, fn := range []func(context.Context) (int64, error){
dst.Observability.DeleteAllMetricSnapshots,
dst.Observability.DeleteAllEdgeHealth,
dst.Observability.DeleteAllNodeObservationFrps,
dst.Observability.DeleteAllNodeObservationFrpc,
} {
if _, err := fn(ctx); err != nil {
return err
}
}
return nil
}
// copyAccessLogs 从 src 复制节点访问日志到 dst。
func copyAccessLogs(ctx context.Context, src, dst *logstore.Store) error {
// 注意:迁移期间 src 已冻结,但复制读取不受冻结影响;每批按 id 升序扫描。
var lastID uint64
for {
rows, err := listNodeAccessLogsByID(ctx, src, lastID, copyBatchSize)
if err != nil {
return err
}
if len(rows) == 0 {
break
}
if err := dst.AccessLogs.BatchInsertNodeAccessLogs(ctx, rows); err != nil {
return fmt.Errorf("写入目标访问日志失败(批 %d): %w", lastID, err)
}
task.AppendLog(ctx, "已复制访问日志 %d 条(截至 id=%d)", len(rows), rows[len(rows)-1].ID)
lastID = rows[len(rows)-1].ID
if len(rows) < copyBatchSize {
break
}
}
return nil
}
ListForMigration已在 Task 2 接口定义:GORM 实现Where("id > ?", afterID).Order("id ASC").Limit(limit);CH 实现原生 SQLSELECT ... FROM of_node_access_logs WHERE id > ? ORDER BY id LIMIT ?。可观测 4 表的*ForMigration同理(按各自表名/模型)。
// copyObservability 复制 4 张可观测表。
func copyObservability(ctx context.Context, src, dst *logstore.Store) error {
for _, c := range []struct {
name string
read func(ctx context.Context, afterID uint64, limit int) (int, error)
}{
{"metric_snapshots", func(ctx context.Context, afterID uint64, limit int) (int, error) {
rows, err := src.Observability.ListMetricSnapshotsForMigration(ctx, afterID, limit)
if err != nil || len(rows) == 0 {
return len(rows), err
}
return len(rows), dst.Observability.BatchInsertNodeMetricSnapshots(ctx, rows)
}},
// edge_health / obs_frps / obs_frpc 同构,调用各自 ForMigration/BatchInsert 对。
} {
var lastID uint64
for {
n, err := c.read(ctx, lastID, copyBatchSize)
if err != nil {
return fmt.Errorf("复制 %s 失败: %w", c.name, err)
}
if n == 0 {
break
}
task.AppendLog(ctx, "已复制 %s %d 条", c.name, n)
if n < copyBatchSize {
break
}
lastID += uint64(n) // 近似游标;实现时改为每批最后一条 id 更精确
}
}
return nil
}
- Step 4: 注册——
register.go加task.RegisterHandler(openflare.LogDBSwitchTask, &openflare.LogDBSwitchHandler{})+task.RegisterTaskMeta(openflare.LogDBSwitchMeta)。 - Step 5: 单测(log_db_switch_test.go)——sqlite↔sqlite 模拟(源 store 写入 3 条,目标 store 空库),执行
copyAccessLogs断言 ID 保留、数量一致;validateSwitch各非法组合报错;ValidatePayload归一化。 - Step 6: 运行
go test ./internal/apps/openflare/tasks/ ./internal/infra/task/。 - Step 7: 提交
git add internal/apps/openflare/ internal/infra/task/ && git commit -m "feat(task): add switch log database migration task"
Task 16: 日志库状态端点
Files:
- Modify:
internal/apps/admin/status/clickhouse.go(改造为log-database状态端点,保留旧路径兼容或重命名 + 路由更新) - Modify:
internal/router/v1/admin.go(路由注册) - Modify:
internal/apps/admin/status/swagger注释
Interfaces:
-
Consumes:
logstore.Active、logstore.Migrating、repository.GetIntByKey(3 个保留配置)、config.Config。 -
Produces:
GET /api/v1/admin/status/log-database返回LogDatabaseStatus。 -
Step 1: 实现状态结构(改造 clickhouse.go)
// GetLogDatabaseStatus 返回当前日志库状态。
// @Summary 获取日志数据库状态
// @Description 返回当前日志主库、迁移状态、各库保留天数与合法迁移目标,需要管理员权限
// @Tags admin
// @Produce json
// @Security SessionCookie
// @Success 200 {object} response.Any{data=status.LogDatabaseStatus} "获取成功"
// @Failure 401 {object} response.Any "未登录"
// @Failure 403 {object} response.Any "无管理员权限"
// @Failure 500 {object} response.Any "内部错误"
// @Router /api/v1/admin/status/log-database [get]
func GetLogDatabaseStatus(c *gin.Context) {
ctx := c.Request.Context()
s, err := logstore.Active(ctx)
if err != nil {
response.AbortInternal(c, "日志存储初始化失败")
return
}
activeDB, _ := logstore.ActiveDatabase(ctx) // provider 增加 ActiveDatabase(ctx) 返回当前库名
migration := "idle"
if logstore.Migrating(ctx) {
migration = "migrating"
}
out := LogDatabaseStatus{
ActiveDatabase: activeDB,
Migration: migration,
RetentionDays: map[string]int{
"postgres": retentionOr(ctx, model.ConfigKeyLogRetentionDaysPostgres),
"sqlite": retentionOr(ctx, model.ConfigKeyLogRetentionDaysSQLite),
"clickhouse": retentionOr(ctx, model.ConfigKeyLogRetentionDaysClickHouse),
},
AvailableTargets: availableTargets(ctx),
}
if activeDB == "clickhouse" {
stats, err := analyticsrepo.GetClickHouseOperationalStats(ctx) // 经 logstore StatusStore 暴露
if err == nil {
stats.BatchWriters = collectBatchWriterStats()
out.ClickHouse = stats
}
}
c.JSON(http.StatusOK, response.OK(out))
}
logstore.ActiveDatabase(ctx)与logstore.Build(ctx, database)(Task 15 用到)需在 provider 增加并实现;analyticsrepo.GetClickHouseOperationalStats改为经logstore.StatusStore暴露,避免 admin/status import analyticsrepo(违反 import-lint)。
- Step 2: 路由——
internal/router/v1/admin.go将/status/clickhouse替换/新增为/status/log-database;旧路径保留 301 或删除(实现时选删除并同步前端)。 - Step 3: 单测——
logstore.ActiveDatabase/Build分支测试;availableTargets(当前=clickhouse → 主库;当前=主库 → clickhouse)。 - Step 4: swagger
make swagger。 - Step 5: 验证
go test ./internal/apps/admin/status/、go build ./internal/...。 - Step 6: 提交
git add internal/apps/admin/ internal/router/ && git commit -m "feat(status): log database status endpoint"
Task 17: 前端——任务参数、业务配置、状态展示
Files:
- Modify:
frontend/lib/services/admin/*(任务/状态类型,若需) - Modify:
frontend/components/common/settings/operation-tab.tsx或业务配置分组(「日志保留时间」) - Modify: 任务管理页组件(
frontend/.../tasks.tsx或等价文件)——展示当前日志主库 + 迁移状态 + 「切换日志数据库」参数下拉 - Modify: 状态页/仪表盘(日志库状态卡片)
Interfaces:
-
Consumes: 现有 Admin 任务派发 API、
/api/v1/admin/status/log-database、AdminService.updateSystemConfig。 -
Step 1: 业务配置分组——在
/admin/settings业务配置 Tab 新增「日志保留时间」:3 个Input type="number"(PG/SQLite/CH),保存调AdminService.updateSystemConfig,成功后 invalidate["admin","system-configs"],Sonner toast。 -
Step 2: 任务管理页——「切换日志数据库」出现在任务列表;参数
target下拉按状态端点available_targets渲染(显示「PostgreSQL(主库)」/「SQLite(主库)」/「ClickHouse」);任务卡片显示active_database与迁移状态徽标。 -
Step 3: 状态卡片——仪表盘或任务页展示当前日志主库、保留天数、迁移中提示。
-
Step 4: 验证
cd frontend && pnpm build(或pnpm lint)。 -
Step 5: 提交
git add frontend/ && git commit -m "feat(frontend): log database status, retention settings, and switch task UI"
M4:收尾与全量验证
Task 18: 全量验证、文档与 changelog
Files:
-
Modify:
docs/changelog/index.md([Unreleased]中文条目) -
Modify:
docs/design/(如需要,日志数据库解耦设计说明) -
全局验证
-
Step 1: 全量检查 运行:
go build ./...go test ./...make code-checkmake swagger(若 API 有变)make format- goose 三套空库 Up 验证(
go test ./internal/infra/persistence/migrator)
-
Step 2: changelog——在
docs/changelog/index.md的[Unreleased]增加合并条目:
- 日志存储解耦:新增日志存储抽象(`internal/repository/logstore`),ClickHouse 变为可选项,不启用时由 PostgreSQL/SQLite 承担全部日志功能;新增「切换日志数据库」任务支持 PostgreSQL/SQLite 与 ClickHouse 间数据迁移(迁移期间冻结日志写入,成功后自动切换主库并保留源数据);日志保留时间改为按存储库在业务配置中设置(`log_retention_days_*`),过期清理并入系统垃圾清理每日任务。
- Step 3: 设计文档归档——确认
docs/superpowers/specs/2026-08-08-log-database-decoupling-design.md与计划一致;实现偏差在 spec 或 changelog 标注。 - Step 4: 提交
git add docs/ && git commit -m "docs: log database decoupling changelog and design notes"
自检记录(writing-plans self-review)
- 规格覆盖:M1 Task 1-10 覆盖规格第 4 节(包结构/接口/约束/标记校验);M2 Task 11-14 覆盖第 5、6 节(表/优化/清理);M3 Task 15-17 覆盖第 7、8 节(迁移任务/API/前端);M4 Task 18 覆盖第 9 节(测试验证)与文档。
- 已知实现决策(由实现者按此执行,避免歧义):
logstore不 importinternal/repository(防循环);配置读取经 bootstrap 注入SetConfigReader。- 迁移复制按 id 升序扫描:
AccessLogStore.ListForMigration+ 可观测 4 个*ForMigration(Task 2 已定义),CH 与 GORM 各自实现;copyObservability用每批最后一条 id 作为下一批游标(实现时修正计划里lastID += n的近似写法)。 logstore.Build(ctx, database)导出供迁移任务构造目标 store;ActiveDatabase(ctx)供状态端点。- admin/status 不直接 import analyticsrepo——CH 运行指标经
logstore.StatusStore暴露。 - 解析 helper(
ParseBrowserName等)迁至model/analytics纯函数,apps 不再依赖 analyticsrepo。 - 迁移期间源库冻结由 logstore 各实现
ensureWritable统一保证;risk_control 审计中间件在冻结期跳过写日志但不阻断请求。 - 失败回退:
defer setMigrationFlag("")保证失败后源库恢复可写;重试时先清空目标再复制(幂等)。