mirror of
https://github.com/Rain-kl/OpenFlare.git
synced 2026-10-06 07:36:37 +08:00
docs(i18n): 恢复并补齐英文版 vitepress,README 默认改为英文
- README 默认英文:README.en.md → README.md(英文为默认),中文移至 README.zh-CN.md,语言切换链接同步 - 恢复被删除的 docs/en/ 英文文档(git 历史 cc5e53c5^),删除 4 篇已废弃文件 - 英文导航 config.ts 对齐中文结构(新增 Deployment/Changelog 侧栏,同步 Guide/Design 条目) - 翻译 15 篇中文新增文档:guide 5 篇(certificates/pages-usage/proxy-config/uptime-kuma/zone-domain-migration)+ design 10 篇(zone-design/cloudflare-pointing/waf-orchestration/origin-error-page/edge-cache-design/pages-design/logstore/kuma-design/login-captcha/observability 三篇) - en 首页更新(新增 Pages 特性、tagline 同步);changelog 英文入口指向中文版 - vitepress 构建验证:43 个英文页面全部渲染 注意:29 篇旧英文文档为恢复版,部分内容(如 deployment/server、reference/configuration)可能落后于中文,需后续逐篇同步
This commit is contained in:
@@ -0,0 +1,181 @@
|
||||
# Agent Design Document
|
||||
|
||||
You will learn: Agent design principles, core functional modules, interaction links with the Server, and how configuration applications are secured and made reliable through immutable version models and the three-stage disaster recovery rollback mechanism.
|
||||
|
||||
---
|
||||
|
||||
## Requirements Analysis
|
||||
|
||||
In distributed reverse proxy and edge security gateway scenarios, the Agent plays a central role in connecting the control plane (Server) and the data plane (OpenResty). Since the Agent runs on the user's actual node server, its design must adhere to the following core security and high-availability requirements:
|
||||
|
||||
1. **Active Pull (Pull Model) instead of Push**: The Server does not hold the SSH keys of the nodes, nor does it actively initiate inbound connections to the nodes. All control directives and configuration updates are actively pulled by the Agent via heartbeats or long-lived connections (WebSockets). This eliminates inbound firewall security risks on the node side and prevents control channels from being hijacked.
|
||||
2. **Minimal Intrusiveness**: The Agent runs as an independent Go binary process. It only interacts with the local OpenResty process through file-based configuration rewriting and signal notifications, without interfering with other system services on the node.
|
||||
3. **Robust Disaster Recovery & Self-Healing**: Since network jitter, disk exhaustion, or erroneous configurations can easily lead to configuration sync failures, the Agent must possess zero-dependency local rollback and self-healing capabilities, strictly preventing a single configuration error from causing a complete node outage.
|
||||
4. **Pure Data and State Landing**: The Agent is only responsible for executing file generation and control intentions rendered by the Server. It does not carry complex control plane duties like business logic validation or multi-tenant authorization, ensuring the node side remains highly efficient and lightweight.
|
||||
|
||||
---
|
||||
|
||||
## Core Capabilities
|
||||
|
||||
The Agent is composed of the following core sub-modules, cooperating to manage its complete lifecycle:
|
||||
|
||||
| Module Name | Directory | Responsibilities |
|
||||
| :--- | :--- | :--- |
|
||||
| **Config Sync** | `sync/` | Pulls full configuration packages, writes files, triggers reloads, and records and reports sync statuses. |
|
||||
| **Heartbeat** | `heartbeat/` | Periodically reports node health and resource metrics to the Server and retrieves the latest active version summary. |
|
||||
| **WebSocket** | `wsclient/` | Maintains a persistent connection with the Server, providing sub-second real-time configuration pushes and commands. |
|
||||
| **OpenResty Control** | `nginx/` | Executes Nginx config validation (`openresty -t`), rewrites, graceful reloads (`reload`), and process auto-start. |
|
||||
| **Local State Store** | `state/` | Persistently records local applied versions, error logs, and buffers unsent observability metrics. |
|
||||
| **Self-Updater** | `updater/` | Listens to Server self-update commands, securely pulls new binary versions, and completes in-place upgrades. |
|
||||
| **Observability** | `observability/` | Collects host CPU/memory/disk and Nginx performance metrics, processes access logs, and uploads them. |
|
||||
| **GeoIP Maintenance** | `geoipdata/` `geoipupdate/` | Maintains and updates the local GeoIP database periodically to support WAF country-level filtering. |
|
||||
|
||||
---
|
||||
|
||||
## Interaction Flows with Server
|
||||
|
||||
The Agent communicates with the control plane through **Token-based Auto-Registration** and a **Dual-channel Heartbeat/WebSocket** system during its lifecycle.
|
||||
|
||||
### 1. Auto-Registration Flow
|
||||
|
||||
If the Agent starts with an empty `access_token` in its local `agent.json`, but has a `discovery_token` configured, it triggers the auto-registration flow:
|
||||
1. The Agent sends a registration request to `/api/agent/register`, carrying a local hardware fingerprint, IP, and hostname.
|
||||
2. After validating the `discovery_token`, the Server generates a unique `NodeID` and a dedicated `AccessToken` (i.e., `agent_token`) in the database and returns them.
|
||||
3. The Agent writes the dedicated Token to its local configuration file, clears the one-time `discovery_token`, and uses the `AccessToken` for all subsequent authenticated communications.
|
||||
|
||||
### 2. Dual-Channel Heartbeat & Sync Mechanism
|
||||
|
||||
* **HTTP Polling (Fallback and Detection)**: The Agent sends POST heartbeat packets at configured `heartbeat_interval` intervals by default. It reports health metrics while retrieving the currently active configuration version summary (Version & Checksum).
|
||||
* **WebSocket Channel (Real-time Communication)**: Upon a successful HTTP heartbeat, the Agent automatically attempts to upgrade the connection to WebSocket (`/api/agent/ws`).
|
||||
* Once the WS connection is established, heartbeats and metrics reporting shift entirely to the WS pipeline, reducing network overhead.
|
||||
* When the Server publishes or activates a new version, it broadcasts a notification to the Agent via WS. The Agent triggers the synchronization flow **immediately** upon receiving the change event, achieving sub-second configuration deployment.
|
||||
* If the WS connection drops due to network issues, the Agent automatically falls back to HTTP polling and uses an exponential backoff algorithm to attempt rebuilding the WS channel.
|
||||
|
||||
### 3. Interaction Sequence Diagram
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
participant Agent as OpenFlare Agent
|
||||
participant OR as Local OpenResty
|
||||
participant Server as OpenFlare Server
|
||||
|
||||
Note over Agent: First Startup (No AccessToken)
|
||||
Agent->>Server: 1. Auto-registration request (carrying discovery_token)
|
||||
Server-->>Agent: 2. Issue NodeID & dedicated AccessToken (agent_token)
|
||||
Note over Agent: Store Token in local configuration file
|
||||
|
||||
rect rgb(240, 248, 255)
|
||||
Note over Agent, Server: HTTP Fallback & WebSocket Upgrade
|
||||
Agent->>Server: 3. Send HTTP Heartbeat (report system metrics & health)
|
||||
Server-->>Agent: 4. Return ActiveConfig summary & AgentSettings
|
||||
Agent->>Server: 5. Initiate WebSocket upgrade request (/api/agent/ws)
|
||||
Server-->>Agent: 6. Upgrade successful (persistent bi-directional channel)
|
||||
end
|
||||
|
||||
rect rgb(245, 245, 245)
|
||||
Note over Agent, Server: Real-time Configuration Publication
|
||||
Note over Server: Administrator clicks publish config in UI
|
||||
Server->>Agent: 7. Broadcast active config summary via WS (WSMessageTypeActiveConfig)
|
||||
Agent->>Server: 8. Request full configuration details (carrying target Version/Checksum)
|
||||
Server-->>Agent: 9. Return complete configuration snapshot (Nginx configs, certs, WAF rules, etc.)
|
||||
Note over Agent: Backup old files, write new config to local temp path
|
||||
Agent->>OR: 10. Execute config syntax validation (openresty -t)
|
||||
OR-->>Agent: 11. Return validation result (OK)
|
||||
Agent->>OR: 12. Send graceful reload signal (openresty -s reload)
|
||||
Agent->>Server: 13. Report application success status (Apply Log & ActiveVersion)
|
||||
end
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Control of OpenResty
|
||||
|
||||
The Agent implements end-to-end closed-loop control of the data plane OpenResty, including configuration rendering, syntax validation, graceful reloading, and exception state capturing:
|
||||
|
||||
### 1. Configuration Layout on Disk
|
||||
|
||||
Upon successful sync, the Agent writes configuration files to `/etc/nginx/openflare-lua/` (or the configured `LuaDir`) according to a strict physical structure:
|
||||
* `nginx.conf`: Main configuration file (replaces absolute path placeholders, configures performance parameters, shared dictionaries, and global server blocks).
|
||||
* `routes.conf`: Route configuration file (generated by the Agent, containing all website server blocks, certificate paths, cache settings, and rate limit directives).
|
||||
* `certs/`: Certificate storage directory (files named as `{cert_id}.crt` and `{cert_id}.key`).
|
||||
* `waf/` and `pow/`: Dedicated Lua runtime scripts required for WAF and CC mitigation.
|
||||
* `waf_config.json` and `waf_ip_groups.json`: Structured rules and IP databases required by the WAF filtering engine.
|
||||
|
||||
### 2. Refined Reload Operations
|
||||
|
||||
1. **Backup Current Config**: Before writing new files, the Agent copies the existing configuration files to a `.backup` directory, keeping a complete rollback snapshot.
|
||||
2. **Write and Replace Placeholders**: Writes the pulled templates, automatically replacing absolute path placeholders (e.g., `__OPENFLARE_LUA_DIR__`) with actual local execution paths.
|
||||
3. **Syntax Validation**: Calls `openresty -t -c <temp_nginx.conf>` to run a strict syntax test.
|
||||
4. **Graceful Reload**: If validation passes, the Agent moves the files to the official paths and executes `openresty -s reload`. If OpenResty is not running, it launches the process.
|
||||
5. **Exception Capture**: If validation or reload fails, the Agent intercepts the standard error output (stderr) and extracts the first 2000 characters of the detailed error log.
|
||||
|
||||
---
|
||||
|
||||
## Publishing & Config Application Model
|
||||
|
||||
OpenFlare discards the fragile mechanism of dynamically patching node configurations, instead using an **immutable configuration version publishing model**.
|
||||
|
||||
```text
|
||||
Edit rules -> Preview / View diff -> Publish -> Generate full configuration version -> Activate version -> Agent pulls -> Local application -> Report result
|
||||
```
|
||||
|
||||
### 1. Core Design Principles
|
||||
|
||||
* **Complete Publication**: Every publication compiles all enabled proxy routes, certificates, and global/custom WAF rules at once, generating a complete version package with a unique `checksum`.
|
||||
* **Version Format**: Uses the `YYYYMMDD-NNN` incremental format, ensuring version histories are intuitive and strictly monotonic.
|
||||
* **Global Single Active Version**: The system supports only one globally `active` configuration version at any given time. Rollbacks do not require reverse patching; they simply transition an older healthy version to the `active` state, and the Agent pulls and applies it.
|
||||
|
||||
### 2. Three-Stage Disaster Recovery & Rollback Mechanism
|
||||
|
||||
If the Agent fails to apply a configuration (or reload fails), it automatically triggers the following three-stage self-healing pipeline:
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
A[Config Application Failed] --> B[Stage 1: Attempt Local Backup Recovery]
|
||||
B -- Backup Exists --> C[Write Local Backup Files]
|
||||
C --> D[Run openresty -t Validation]
|
||||
D -- Validation OK --> E[Reload Old Configuration]
|
||||
D -- Validation Failed --> F[Proceed to Stage 2]
|
||||
B -- No Backup --> F[Stage 2: Write Built-in Safe Fallback Config]
|
||||
F --> G[Write fallback nginx.conf: Listen on Port 80 Only]
|
||||
G --> H[Enable stub_status health checks]
|
||||
G --> I[Return 503 for all other routes & block errors]
|
||||
G --> J[Attempt to launch OpenResty to maintain basic survival]
|
||||
J --> K[Proceed to Stage 3]
|
||||
E --> L[Report Apply Warning]
|
||||
K --> M[Block Local Repeated Application of Failed Version]
|
||||
M --> N[Report Apply Error with detailed logs]
|
||||
```
|
||||
|
||||
1. **Stage 1: Local Backup Rollback**
|
||||
* The Agent attempts to restore the main configuration, routes, and certificates from the `.backup` directory.
|
||||
* It runs `openresty -t` validation on the restored backup. If successful, it reloads and reports a `Warning` to the Server (Warning: failed to apply new version, automatically rolled back to the previous healthy version).
|
||||
2. **Stage 2: Built-in Safe Fallback Runtime**
|
||||
* If no local backup exists (e.g., first deployment failed) or if the rollback validation fails, the Agent activates the ultimate self-healing mechanism: writing a **built-in safe fallback configuration**.
|
||||
* **Fallback Configuration Specification**:
|
||||
* Listens only on port `80`, containing no real user reverse proxy routes.
|
||||
* The `/openflare/stub_status` endpoint returns a healthy response, while all other requests uniformly return a `503 Service Unavailable` status code with the fixed response body `OpenFlare: No Valid Configuration`.
|
||||
* It attempts to launch OpenResty with this minimal configuration. This keeps the Nginx process alive, preserving underlying health probes and metric endpoints, preventing containers/pods from being repeatedly killed and restarted by orchestration systems, while keeping sensitive routes secure.
|
||||
3. **Stage 3: Local Configuration Blocking**
|
||||
* The Agent records the failing configuration's `version + checksum` in its local state store blacklist.
|
||||
* Until the control plane activates a new configuration (resulting in a changed `checksum`), the Agent's heartbeat blocks repeated synchronization pulls of this erroneous version, preventing nodes from entering an infinite loop of "heartbeat -> pull failing config -> crash rollback".
|
||||
|
||||
### 3. WAF IP Group Asynchronous Runtime Synchronization
|
||||
|
||||
To prevent highly volatile IP blacklists from triggering frequent full config publications and Nginx reloads (which still incur minor CPU and connection overhead), WAF IP groups are synchronized via an **asynchronous differential sync design**:
|
||||
|
||||
* **Static Publication Snapshot**: The `waf_config.json` generated upon publication only contains the group ID reference mapping (i.e., `ip_whitelist_group_ids` / `ip_blacklist_group_ids`) and does not contain the actual list of IP addresses.
|
||||
* **Heartbeat Differential Check**: The Agent uploads its locally cached IP groups MD5 checksum map in its heartbeat.
|
||||
* **Differential Delivery**: The Server compares checksums and only delivers missing or modified IP groups, which are written directly to `waf_ip_groups.json` on the node without reload.
|
||||
* **WebSocket Real-time Push**: When an administrator updates an IP group, or a threat intelligence subscription successfully pulls, or a security rule triggers a temporary block, the Server immediately broadcasts the IP group update package via WebSocket. The Agent receives and applies it instantly **without Nginx reloads**.
|
||||
|
||||
---
|
||||
|
||||
## Design Constraints
|
||||
|
||||
To protect the security boundary of the data and control plane, Agent development must strictly comply with the following engineering constraints:
|
||||
|
||||
1. **Zero-Privilege Command Execution**: The Server is strictly prohibited from sending any arbitrary shell commands or scripts to the Agent (such as exec/eval). All system control operations (such as start, stop, reload, update) must be hardcoded inside the Agent binary.
|
||||
2. **Strict Token Filtering and Prefix Validation**: Agent requests to the Server must be prefixed with `/api/agent/` and must carry the `X-Agent-Token` header for signature or token verification.
|
||||
3. **Node Autonomy**: The Agent must support complete offline capabilities. During disconnected periods, the local OpenResty must rely on local configuration copies to keep reverse proxy services running normally.
|
||||
@@ -0,0 +1,223 @@
|
||||
# System Architecture
|
||||
|
||||
You will learn: The overall architecture of OpenFlare, the boundaries of responsibilities for Server, Agent, OpenResty, and Admin Frontend, and the request flow of a configuration publication from the admin dashboard to activation on a node.
|
||||
|
||||
OpenFlare consists of the Server, the Agent, the node-local OpenResty, and the Admin Frontend. The Server is the control plane, the Agent is the only controlled entry point on the node side, and OpenResty serves as the actual data plane. In intranet penetration scenarios, the Relay (frps manager) and OpenFlared (frpc manager) extend the data plane traffic path.
|
||||
|
||||
### Standard Reverse Proxy Traffic Path
|
||||
|
||||
```text
|
||||
Browser
|
||||
|
|
||||
| Management UI / API
|
||||
v
|
||||
OpenFlare Server (Gin + GORM + SQLite/PostgreSQL)
|
||||
|
|
||||
| Agent API / heartbeat / config pull
|
||||
v
|
||||
OpenFlare Agent
|
||||
|
|
||||
| write config / openresty -t / reload / rollback
|
||||
v
|
||||
OpenResty binary
|
||||
|
|
||||
| reverse proxy
|
||||
v
|
||||
Origin
|
||||
```
|
||||
|
||||
### Intranet Penetration Traffic Path
|
||||
|
||||
```text
|
||||
Browser
|
||||
|
|
||||
| HTTPS request
|
||||
v
|
||||
OpenResty (Agent, TLS/WAF) <-- TunnelRelay Node
|
||||
|
|
||||
| proxy_pass http://localhost:vhost_port (Host header preserved)
|
||||
v
|
||||
OpenFlareRelay (frps) <-- TunnelRelay Node, co-located with Agent
|
||||
|
|
||||
| frp tunnel protocol (HTTP Vhost routing by Host header)
|
||||
v
|
||||
OpenFlared (frpc) <-- Intranet Server
|
||||
|
|
||||
| HTTP/HTTPS forward
|
||||
v
|
||||
Internal Service (192.168.x.x)
|
||||
```
|
||||
|
||||
## Component Responsibilities
|
||||
|
||||
| Component | Responsibility |
|
||||
| --- | --- |
|
||||
| Server | Admin UI, Admin API, Agent/Relay/Client API, configuration rendering, version publishing, data storage, and aggregated queries. |
|
||||
| Agent | Registration, heartbeats, synchronization, file writing, validation, reload, rollback on failure, self-updating, and light metrics collection. |
|
||||
| OpenResty | Receives real traffic, executing WAF, PoW, authentication, and reverse proxying according to the configuration rendered by OpenFlare. |
|
||||
| OpenFlareRelay | Manages the lifecycle of the frps process, providing tunnel relay services and receiving frps configurations via heartbeat. |
|
||||
| OpenFlared | Manages frpc processes (can be multiple), connecting to the Relay and forwarding traffic to intranet services. |
|
||||
| Frontend | Manages pages for website configs, WAF, origins, certificates, nodes, tunnels, versions, users, settings, and observability. |
|
||||
|
||||
## Server
|
||||
|
||||
`openflare-server` is the single-control-plane monolith:
|
||||
|
||||
* Gin provides the HTTP services.
|
||||
* GORM accesses SQLite or PostgreSQL.
|
||||
* The existing login system provides Admin Session management.
|
||||
* Authentication sources support GitHub OAuth and standard OIDC logins with external account binding.
|
||||
* The Go Server hosts the `openflare-server/web` static build assets.
|
||||
|
||||
The Server does not directly SSH to nodes, nor does it modify node files online. It only stores control plane state, generates complete configuration versions, and lets nodes actively pull them via the Agent API.
|
||||
|
||||
## Agent
|
||||
|
||||
`openflare-agent` is a Go monolithic application:
|
||||
|
||||
* Runs as a single binary on the node side.
|
||||
* Reads or generates local node information on startup.
|
||||
* Performs periodic heartbeat check-ins to report status and retrieve active version summaries.
|
||||
* Upon discovering a new version, it pulls the configuration, backs up old files, writes new files, validates them, and reloads.
|
||||
* Automatically rolls back to restore operations if the application fails.
|
||||
* Maintains the local WAF GeoIP mmdb, writing the built-in library on startup and updating it periodically based on configuration.
|
||||
|
||||
The Agent executes validation, reload, startup, and restart uniformly via the path specified in `openresty_path`; if unconfigured, it defaults to calling `openresty`. During Docker deployments, the Agent image packages OpenResty and follows the same execution control logic.
|
||||
|
||||
The node IP is maintained by default through Agent registration and heartbeat reporting; if the administrator locks the node IP, the Server only updates running status, versions, and observability fields, and no longer accepts reports from the Agent to override the locked IP.
|
||||
|
||||
## Frontend
|
||||
|
||||
`openflare-server/web` is the official Next.js-based frontend:
|
||||
|
||||
* Next.js 15 App Router.
|
||||
* React 19.
|
||||
* TypeScript.
|
||||
* Tailwind CSS.
|
||||
* TanStack Query for server-side state.
|
||||
|
||||
The frontend uses static export mode (`output: 'export'`), which is then hosted by the Go Server using `embed.FS`. All API requests must go through `lib/api/` and process the `success/message/data` response structure.
|
||||
|
||||
The Server integrates the following security features:
|
||||
* CORS middleware: Cross-Origin Resource Sharing protection.
|
||||
* Rate limiting: Global and key API endpoint throttling.
|
||||
* Session management: Cookie/Redis-based session storage.
|
||||
|
||||
## Data & Request Flow
|
||||
|
||||
### Management Request Flow
|
||||
|
||||
```text
|
||||
Browser -> Frontend -> /api/* -> controller -> service -> model -> database
|
||||
```
|
||||
|
||||
Admin mutation APIs use `POST`, while read-only APIs use `GET`. Both success and failure responses return a clear `message`.
|
||||
|
||||
### Agent Sync Flow
|
||||
|
||||
```text
|
||||
Agent HTTP heartbeat -> Server returns active version summary
|
||||
Agent detects new version -> Pulls complete configuration details
|
||||
Agent writes main configuration / route configurations / certificates / Lua resources / WAF runtimes
|
||||
Agent runs OpenResty validation (openresty -t) and reload
|
||||
Agent reports application result
|
||||
```
|
||||
|
||||
### Relay Sync Flow
|
||||
|
||||
The Relay (OpenFlareRelay process) runs on the TunnelRelay node and shares the same `agent_token` with the Agent:
|
||||
|
||||
```text
|
||||
Relay HTTP heartbeat -> Server returns frps base configuration (bindPort, vhostHTTPPort, auth_token)
|
||||
Relay generates frps.toml and starts or updates the frps process
|
||||
Relay periodically reports frps health status and connection statistics
|
||||
Relay attempts WebSocket upgrade for real-time configuration pushes
|
||||
```
|
||||
|
||||
frps configurations are relatively static (ports, auth token), dispatched via heartbeats, and **not included in the versioned publishing flow**. The Relay must monitor the frps process and auto-recover it on failures. Authentication: `X-Agent-Token` + API path prefix `/api/relay/*`, distinguished by Server via `node_type = tunnel_relay`.
|
||||
|
||||
### OpenFlared Sync Flow
|
||||
|
||||
OpenFlared (client) runs inside the intranet server, using independent `tunnel_token` authentication:
|
||||
|
||||
```text
|
||||
Client HTTP heartbeat -> Server returns tunnel configuration version summary (version, checksum)
|
||||
Client detects new version -> Pulls complete tunnel route configuration (relay list + frpc proxy definitions)
|
||||
Client generates independent frpc.toml configuration files for each Relay
|
||||
Client starts a new frpc process for new Relays, or hot-reloads (frpc reload) existing ones
|
||||
Client reports application results (success/failure details)
|
||||
```
|
||||
|
||||
OpenFlared communicates with the Server via `/api/flared/*` using the `X-Tunnel-Token` header. Tunnel route configurations are versioned along with the publishing flow, ensuring all configuration changes are consistently published to both Agents and Clients via a single version number.
|
||||
|
||||
**WebSocket Upgrade Flow** (Optional, controlled via `AgentWebsocketUpgradeEnabled`):
|
||||
|
||||
When WebSocket upgrade is enabled:
|
||||
1. The Agent retrieves run configurations and settings via HTTP heartbeat.
|
||||
2. The Agent attempts to upgrade the connection to `GET /api/agent/ws` (WebSocket).
|
||||
3. Once the WS connection is established, periodic state reporting and real-time commands are carried over the WebSocket pipeline, minimizing latency.
|
||||
4. When the Server publishes or activates a version, it immediately broadcasts the active version summary to connected Agents, triggering the sync flow instantly.
|
||||
5. If the WebSocket disconnects or fails to establish, the Agent automatically falls back to HTTP heartbeats, ensuring high availability.
|
||||
|
||||
Through the `OpenRestyWebsocketEnabled` option, WebSocket reverse proxy support can be enabled or disabled at the OpenResty layer.
|
||||
|
||||
### Reverse Proxy Flow
|
||||
|
||||
```text
|
||||
Client -> OpenResty server block -> WAF Lua -> named upstream -> Origin
|
||||
```
|
||||
|
||||
Website configurations are the boundaries of reverse proxy aggregation. A single website configuration can bind multiple domains, sharing site-level rate limiting, reverse proxy, and cache settings.
|
||||
|
||||
WAF executes in the OpenResty `access_by_lua_file` phase. Rules originate from the `waf_config.json` carried in the currently active version; global rule groups take effect by default, and websites can overlay custom rule groups. `waf_config.json` only stores rule group references and IP group IDs; IP group members are synchronized independently by the Agent into `waf_ip_groups.json`, and the OpenResty Lua engine merges and evaluates them by reference ID.
|
||||
|
||||
WAF IP groups are managed by the Server. Manual IP groups store IP/CIDR lists directly; auto IP groups are evaluated by Server cron jobs reading request logs and applying Expr boolean rules; subscription IP groups are fetched by Server cron jobs from remote text or JSON sources. The Agent reports local IP group checksums in heartbeats, and the Server only returns mismatched IP groups. When an IP group is updated on the Server, a broadcast is sent via WebSocket to push changes, and the OpenResty Lua reads the local JSON file directly without querying the DB, request logs, or remote subscription sources.
|
||||
|
||||
## Core Objects
|
||||
|
||||
Current valid entities include:
|
||||
|
||||
* `proxy_routes`
|
||||
* `origins`
|
||||
* `config_versions`
|
||||
* `nodes`
|
||||
* `tunnels`
|
||||
* `auth_sources`
|
||||
* `external_accounts`
|
||||
* `node_system_profiles`
|
||||
* `apply_logs`
|
||||
* `tls_certificates`
|
||||
* `managed_domains`
|
||||
* `node_request_reports`
|
||||
* `node_access_logs`
|
||||
* `node_metric_snapshots`
|
||||
* `traffic_analytics_rollups`
|
||||
* `node_health_events`
|
||||
* `waf_rule_groups`
|
||||
* `waf_ip_groups`
|
||||
* `waf_rule_group_bindings`
|
||||
* `acme_accounts`
|
||||
* `dns_accounts`
|
||||
* `geoip_update_configs`
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
| Decision | Rationale |
|
||||
| --- | --- |
|
||||
| Full Config Versioning instead of Patches | Provides stable, verifiable boundaries for previewing, activating, history, and rollbacks. |
|
||||
| Pull Model (Agent-driven) | Server does not need SSH keys or inbound command ports, preventing control channel hijacking. Supports HTTP and WebSocket. |
|
||||
| Global Single Active Version | Reduces MVP complexity, ensuring all nodes are uniform by default. Supports previews, version history, and one-click rollback. |
|
||||
| Website Multi-Domain Aggregation | Enables sharing site-level policies across domains while supporting per-domain certificate binding. |
|
||||
| Server-side Observability Aggregation | Prevents UI-side temporary statistical calculations from producing inconsistent data metrics. |
|
||||
| Intranet Penetration based on frp | Reuses a mature tunnel protocol rather than custom implementations to minimize stability risks. frps Vhost routing aligns naturally with HTTP. |
|
||||
| Independent Binary for Relay/Client | Separation of concerns: Relay manages frps, Client manages frpc, allowing independent updates and deployments. |
|
||||
| Tunnel decoupled from Node system | Tunnel clients run internally, using completely different registration and authentication flows compared to edge nodes. |
|
||||
|
||||
## Recommended Reading for Contributors
|
||||
|
||||
Before modifying architectural code, please read:
|
||||
|
||||
1. [Product Boundaries](./index.md)
|
||||
2. [Agent & Publish Model](./agent-design.md)
|
||||
3. [Development Constraints](../../guideline/Constraints.md)
|
||||
4. [Repository Structure](./repository.md)
|
||||
@@ -0,0 +1,223 @@
|
||||
# Cloudflare DNS Pointing Design
|
||||
|
||||
## Goals
|
||||
|
||||
Point **ZoneDomains (explicit FQDNs)** in OpenFlare to edge node IPs quickly via the Cloudflare API, replacing manual A-record edits in the CF console. Users organize domains into **pointing groups**: each group configures a primary node, a backup node, and a default orange-cloud (proxied) policy; members can override orange-cloud individually. The system treats the database tables as the desired state and idempotently syncs remote DNS.
|
||||
|
||||
This module is an **optional integration capability**; it does not turn Zones into an authoritative DNS control plane. Zones still only handle root-domain boundaries, domains, certificates, and reverse-proxy associations; A record create/update/delete is driven by this module through Cloudflare.
|
||||
|
||||
## Scope and Phasing
|
||||
|
||||
### Phase 1 (this design's scope)
|
||||
|
||||
* Sidebar **Cloudflare** entry with a Token-ready gate
|
||||
* Connection config: import from an existing DNS account **or** standalone entry within the module (mixed sources), stored encrypted
|
||||
* Pointing group CRUD: primary node, backup node (reserved), group default orange-cloud
|
||||
* Member management: add/remove at the granularity of `zone_domain_id`; member-level orange-cloud
|
||||
* Sync: write each member as a **single A record** on Cloudflare → current active node IPv4
|
||||
* Triggers: manual sync, member add, node/orange-cloud change, node IP change enqueue
|
||||
* Async task batch sync; member sync status with readable errors
|
||||
|
||||
### Phase 2
|
||||
|
||||
* Agent heartbeat offline detection of primary node failure → `active_node` switches to backup → whole-group auto sync
|
||||
* Optional auto failback, failure notification push
|
||||
|
||||
### Explicitly Out of Scope (later or permanent)
|
||||
|
||||
* Multiple parallel Cloudflare accounts (one global connection config)
|
||||
* AAAA / multi-A load balancing / CNAME to node hostnames
|
||||
* Managing MX/TXT/Page Rules and other non-module A records
|
||||
* Non-Cloudflare DNS providers
|
||||
* Merging DNS record management into the Zone core model
|
||||
|
||||
## Relationship with Existing Capabilities
|
||||
|
||||
| Existing Capability | Relationship |
|
||||
| --- | --- |
|
||||
| `of_zones` / `of_zone_domains` | Provide the pointable FQDN list; this module only references `zone_domain_id` |
|
||||
| `of_nodes.ip` | Source of A record `content`; recommended to restrict to edge nodes with valid IPv4 |
|
||||
| `of_dns_accounts` + `sealSensitive` | ACME DNS-01 already supports Cloudflare Token; this module can **import** the same account or store a Token standalone |
|
||||
| lego Cloudflare provider | **Only** TXT/DNS-01; this module builds its own CF HTTP client for Zone/DNS Record APIs |
|
||||
|
||||
## Core Model
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
CF_CONNECTIONS ||--o| DNS_ACCOUNTS : optional_import
|
||||
CF_POINTING_GROUPS ||--o{ CF_POINTING_MEMBERS : contains
|
||||
ZONE_DOMAINS ||--o| CF_POINTING_MEMBERS : pointed_as
|
||||
NODES ||--o{ CF_POINTING_GROUPS : primary
|
||||
NODES ||--o{ CF_POINTING_GROUPS : backup
|
||||
NODES ||--o{ CF_POINTING_GROUPS : active
|
||||
|
||||
CF_CONNECTIONS {
|
||||
uint id PK
|
||||
string source
|
||||
uint dns_account_id
|
||||
string authorization
|
||||
string status
|
||||
time verified_at
|
||||
}
|
||||
CF_POINTING_GROUPS {
|
||||
uint id PK
|
||||
string name
|
||||
uint primary_node_id
|
||||
uint backup_node_id
|
||||
uint active_node_id
|
||||
bool default_proxied
|
||||
bool enabled
|
||||
}
|
||||
CF_POINTING_MEMBERS {
|
||||
uint id PK
|
||||
uint group_id
|
||||
uint zone_domain_id UK
|
||||
bool proxied
|
||||
string cf_zone_id
|
||||
string cf_record_id
|
||||
string desired_ip
|
||||
string sync_status
|
||||
string last_error
|
||||
time synced_at
|
||||
}
|
||||
```
|
||||
|
||||
### `of_cf_connections` (one valid connection globally)
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `source` | `dns_account` \| `standalone` |
|
||||
| `dns_account_id` | references `of_dns_accounts` (type=cloudflare) when `source=dns_account` |
|
||||
| `authorization` | encrypted storage when `source=standalone`, payload shape `{"api_token":"..."}`, consistent with DNS accounts; the API **never returns it** |
|
||||
| `status` / `verified_at` | connectivity check result and time |
|
||||
|
||||
**Token resolution:** `dns_account` → decrypt the associated account; `standalone` → decrypt this row. Associated account deleted or validation failed → module not ready, sync forbidden.
|
||||
|
||||
**Recommended permissions:** Cloudflare API Token with `Zone:Read`, `DNS:Edit`.
|
||||
|
||||
### `of_cf_pointing_groups`
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `name` | display name |
|
||||
| `primary_node_id` | primary node |
|
||||
| `backup_node_id` | backup (nullable; phase 1 stores only) |
|
||||
| `active_node_id` | currently effective node; equals primary in phase 1; rewritten by phase 2 failover |
|
||||
| `default_proxied` | group default orange-cloud; **only affects newly added members** |
|
||||
| `enabled` | whether to participate in sync |
|
||||
|
||||
Constraints: primary and backup must not be the same node; the node chosen as the active target must have a valid IPv4.
|
||||
|
||||
### `of_cf_pointing_members`
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `group_id` | owning group |
|
||||
| `zone_domain_id` | globally unique: one domain belongs to at most one group |
|
||||
| `proxied` | member orange-cloud (the only runtime basis) |
|
||||
| `cf_zone_id` / `cf_record_id` | Cloudflare cache for idempotent updates |
|
||||
| `desired_ip` / `sync_status` / `last_error` / `synced_at` | desired and sync state |
|
||||
|
||||
`sync_status`: `pending` \| `syncing` \| `ok` \| `error`.
|
||||
|
||||
No physical foreign keys; `zone_domain_id` unique index; query indexes on `group_id` etc.
|
||||
|
||||
## Orange-Cloud Priority
|
||||
|
||||
1. **Member `proxied`**: the only basis written to CF during sync.
|
||||
2. **Group `default_proxied`**: copied to `proxied` when a member is **added**.
|
||||
3. Later changes to the group default **do not rewrite** existing members.
|
||||
|
||||
## Sync Semantics
|
||||
|
||||
### Desired State
|
||||
|
||||
The OpenFlare DB tables are the Source of Truth. Each member expects:
|
||||
|
||||
| Item | Value |
|
||||
| --- | --- |
|
||||
| type | `A` |
|
||||
| name | the ZoneDomain's FQDN |
|
||||
| content | the group `active_node`'s IPv4 |
|
||||
| proxied | member `proxied` |
|
||||
| ttl | forced Auto by CF when orange-cloud is on; unified default (e.g. 300) when off |
|
||||
|
||||
Phase 1 does not write AAAA. Node IP not a valid IPv4 → that member is `error`.
|
||||
|
||||
### Triggers
|
||||
|
||||
| Trigger | Behavior |
|
||||
| --- | --- |
|
||||
| Manual sync (all / group / member) | reconcile |
|
||||
| Member added | initialize `proxied`, then enqueue sync |
|
||||
| Member removed / group deleted | delete the remote A managed by this module by default (configurable keep) |
|
||||
| Primary node / active / member `proxied` changed | re-sync the corresponding scope |
|
||||
| Node IP changed (heartbeat or manual) | enqueue members whose `active_node_id` points to that node |
|
||||
| Token not ready | refuse sync |
|
||||
|
||||
Phase 1 does not do scheduled full reconciliation.
|
||||
|
||||
### Reconcile (single member, idempotent)
|
||||
|
||||
1. Resolve the CF Zone by the FQDN's registrable root domain, cache `cf_zone_id`.
|
||||
2. With a `cf_record_id`, prefer Update; if stale, list by `name+type=A`.
|
||||
3. **0 records** → Create; **exactly 1** → take over and Update; **multiple** → fail and tell the user to clean up in CF.
|
||||
4. Write back `cf_record_id`, `desired_ip`, `sync_status`, `synced_at` / `last_error`.
|
||||
5. On rate limiting, retry with bounded backoff.
|
||||
|
||||
**Ownership:** only manage records cached by this module or taken over as "the only same-name A"; do not clear the Zone or touch other record types. After a user edits in the CF console, the next sync overwrites with the OpenFlare desired state.
|
||||
|
||||
### Execution Carrier
|
||||
|
||||
* Single record: can sync on the request path.
|
||||
* Whole group / per-node batch: Asynq tasks (`cloudflare:sync_member` / `sync_group` / `sync_by_node`), registered in `bootstrap`.
|
||||
* Per-member mutex to prevent concurrent double-writes.
|
||||
* Node IP change path delivers tasks **best-effort**, not blocking the heartbeat.
|
||||
|
||||
## API (Admin Panel)
|
||||
|
||||
Prefix: `/api/v1/d/cloudflare`, Session admin auth. Package: `internal/apps/openflare/cloudflare/`; routes: `internal/router/v1/openflare/register_cloudflare.go`.
|
||||
|
||||
| Resource | Method & Path |
|
||||
| --- | --- |
|
||||
| Connection | `GET/PUT /connection`, `POST /connection/verify`, `POST /connection/clear` |
|
||||
| Overview | `GET /overview` |
|
||||
| Groups | `GET/POST /groups`, `GET /groups/:id`, `POST /groups/:id/update|delete|sync` |
|
||||
| Members | `GET/POST /groups/:id/members`, `POST .../members/:memberId/update|remove|sync` |
|
||||
| Available domains | `GET /domains/available` |
|
||||
|
||||
* Success `response.OK`; failure `response.Abort*`; the Token is **never** returned in JSON.
|
||||
* Handlers separated from `logics.go`; the CF client is abstracted behind an interface for replaceability.
|
||||
|
||||
## Frontend
|
||||
|
||||
* Navigation: `frontend/lib/navigation/openflare-nav.ts` adds **Cloudflare** → `/cloudflare` (near Website Management / DNS Accounts).
|
||||
* Routes:
|
||||
* `/cloudflare`: overview; guide to configure when not ready
|
||||
* `/cloudflare/settings`: mixed Token config and test connection
|
||||
* `/cloudflare/groups`, `/cloudflare/groups/[id]`: list and detail (members, orange-cloud, sync)
|
||||
* Services: independent service under `frontend/lib/services/openflare/`, extending `BaseService`.
|
||||
* Pages follow the existing title-bar and component-split conventions; destructive actions need double confirmation.
|
||||
* Copy that must be visible: sync overwrites module-managed A records; multiple same-name A records need manual cleanup; removal deletes remote records by default; phase 1 has no automatic failover.
|
||||
|
||||
## Errors and Security
|
||||
|
||||
* User-visible copy is module-internal constants; internal errors log via `pkg/logger`.
|
||||
* Typical: token not configured, invalid token, node without IP, no CF Zone, multiple same-name A records, rate limiting.
|
||||
* Token is only decrypted server-side for use; responses and logs must never contain plaintext tokens.
|
||||
|
||||
## Data Migration
|
||||
|
||||
* goose both dialects (PG/SQLite) create the three tables; defaults match Go zero values.
|
||||
|
||||
## Key Decision Summary
|
||||
|
||||
| Decision | Conclusion |
|
||||
| --- | --- |
|
||||
| Module shape | standalone Cloudflare pointing module, not embedded Zone fields |
|
||||
| Token | mixed: imported from DNS account or encrypted standalone |
|
||||
| Domain granularity | ZoneDomain (FQDN) |
|
||||
| Record shape | single A → active node IPv4 |
|
||||
| Failover | phase 2; heartbeat offline; phase 1 only stores backup/active |
|
||||
| Orange-cloud | member-level effective; group default only initializes |
|
||||
| SoT | DB tables as desired state drive CF |
|
||||
@@ -0,0 +1,264 @@
|
||||
# Edge Cache Strategy Design
|
||||
|
||||
You will learn: how OpenFlare's edge `proxy_cache` aligns with the Cloudflare default loop between "should cache" and "should not cache": request eligibility (extension/policy) × response shareability (origin `Cache-Control` / `Expires` / `Set-Cookie`), and the differences from the previous over-strict request bypass.
|
||||
|
||||
This design is the productized chapter on "basic caching" in [System Architecture](./architecture.md); cache results in access logs are in [Observability Data Model §3.5.1](./observability-data-model.md).
|
||||
|
||||
---
|
||||
|
||||
## 1. Goals and Non-Goals
|
||||
|
||||
### 1.1 Goals
|
||||
|
||||
* **Close to CF default out of the box**: after enabling cache on a route, **only static extensions are cached by default** — HTML is not cached by default; **request session cookies / Authorization / client Cache-Control no longer cause a blanket BYPASS**.
|
||||
* **Cacheable content hits**: a logged-in user visiting `/_app/**/*.js` and other static assets can show `MISS` → `HIT`.
|
||||
* **Non-cacheable stays blocked**: policy not eligible (equivalent to CF `DYNAMIC`); origin `private` / `no-store`; responses with **`Set-Cookie` not stored** (aligned with CF OCC default); `all` is an advanced option with documented warnings.
|
||||
* **Default Edge TTL when no origin freshness**: aligned with CF's per-status default TTL (see §3.5).
|
||||
* **Consistent observability**: keep relying on `$upstream_cache_status` → three-state `cache_status` detail.
|
||||
* **Backward compatible**: legacy route `cache_policy=url` maps to `all`; policy enum and migration rules stay in [§5](#5-兼容与迁移).
|
||||
|
||||
### 1.2 Non-Goals (later iterations)
|
||||
|
||||
* Cache Rules expression engine
|
||||
* Forced Edge TTL ignoring origin `Cache-Control` (CF Cache Rules "Ignore cache-control")
|
||||
* Purge (by URL/prefix/site-wide)
|
||||
* Browser TTL rewriting, client `CF-Cache-Status` response header
|
||||
* Full RFC conditions: `Authorization` cached only when the response has `public`/`s-maxage`/`must-revalidate` (needs Lua; this iteration deletes the request-side bypass entirely, relying on policy + origin headers)
|
||||
* HEAD → GET conversion then cache
|
||||
* Hit-rate dashboard
|
||||
|
||||
---
|
||||
|
||||
## 2. Cloudflare Decision Loop (Alignment Baseline)
|
||||
|
||||
CF default is a **two-stage decision**, **not** "request has Cookie → don't cache".
|
||||
|
||||
### 2.1 Stage A — Eligible at Request Time
|
||||
|
||||
| Condition | CF Result |
|
||||
| --- | --- |
|
||||
| Non-GET | not cached by default |
|
||||
| Extension not in default cacheable table, no Rules forcing eligible | **`DYNAMIC`** (no cache lookup) |
|
||||
| Extension in default table, or Rules eligible | continue to Stage B |
|
||||
| **Request Cookie** | **no effect by default** |
|
||||
| Cache Rules Bypass | `DYNAMIC` |
|
||||
|
||||
CF's default cacheable extensions are keyed by **extension** rather than MIME; **HTML / JSON are not cached by default**.
|
||||
|
||||
### 2.2 Stage B — Response Storeable (OCC on, Free/Pro/Biz default)
|
||||
|
||||
| Condition | Result |
|
||||
| --- | --- |
|
||||
| `Cache-Control: no-store` / `private` | not stored |
|
||||
| `public` + `max-age>0`, or future `Expires` | cacheable |
|
||||
| No Cache-Control / Expires | still cacheable with per-status **default Edge TTL** (e.g. 200 → 120m) |
|
||||
| Response **`Set-Cookie`** (default cache level + OCC) | **not stored**, status tends toward **BYPASS** |
|
||||
| Request `Authorization` | cacheable only when the response also has `public` / `s-maxage` / `must-revalidate` (full condition simplified with Nginx this iteration, see §3.4) |
|
||||
|
||||
### 2.3 Status Semantics (vs. Observability)
|
||||
|
||||
| CF | Meaning | OpenFlare `cache_status` |
|
||||
| --- | --- | --- |
|
||||
| HIT / STALE / UPDATING / REVALIDATED | hit class | same-name or equivalent |
|
||||
| MISS / EXPIRED | fetch from origin | same-name |
|
||||
| BYPASS | eligible at request time, response not cacheable | `BYPASS` → UI "not cached" |
|
||||
| DYNAMIC | not eligible at request time | policy skip mostly `BYPASS` or empty → UI "not cached" |
|
||||
|
||||
---
|
||||
|
||||
## 3. Product Semantics
|
||||
|
||||
### 3.1 Two-Level Switch (unchanged)
|
||||
|
||||
* **Global** `openresty_cache_enabled`: generates `proxy_cache_path` etc.; when off, route-level cache directives are inert.
|
||||
* **Route** `cache_enabled`: whether to enable `proxy_cache` in that site's `location`.
|
||||
|
||||
Cache logic only runs when both are on.
|
||||
|
||||
### 3.2 Policy Enum
|
||||
|
||||
| `cache_policy` | Meaning | New Default | Legacy Compatibility |
|
||||
| --- | --- | --- | --- |
|
||||
| **`static`** | only eligible when URI matches **standard static extensions** | **yes** | — |
|
||||
| **`all`** | after method bypass, no path/extension restriction (advanced; risk similar to CF Cache Everything) | no | legacy `url` → `all` |
|
||||
| **`suffix`** | custom extension list (`cache_rules`) | no | kept |
|
||||
| **`path_prefix`** | custom path prefix | no | kept |
|
||||
| **`path_exact`** | custom exact path | no | kept |
|
||||
|
||||
Render layer: historical `url` is treated as `all`; API/UI only expose the enum above.
|
||||
|
||||
### 3.3 Standard Static Extensions (built-in)
|
||||
|
||||
Aligned with CF default "no HTML/JSON caching"; keeps modern frontend-friendly enhancements:
|
||||
|
||||
```text
|
||||
css js mjs map
|
||||
ico cur gif jpg jpeg png webp avif svg svgz
|
||||
ttf otf woff woff2 eot
|
||||
mp3 mp4 webm ogg flac
|
||||
wasm pdf
|
||||
zip 7z gz tar
|
||||
```
|
||||
|
||||
* **Excludes** `html` / `htm` / **`json`** (aligned with CF not caching JSON by default).
|
||||
* **Includes** `map` / `mjs` / `wasm` (deliberate enhancement for sourcemap / ES module / WASM hits).
|
||||
* Matching: `$uri` extension, case-insensitive:
|
||||
`if ($uri !~* \.(?:css|js|…)$) { set $openflare_skip_cache 1; }`
|
||||
|
||||
### 3.4 Request-Side Bypass (after CF alignment)
|
||||
|
||||
Only kept:
|
||||
|
||||
1. `$request_method != GET` (HEAD included, consistent with current network; no CF HEAD→GET)
|
||||
|
||||
**Removed** (previously over-strict, causing low hit rates):
|
||||
|
||||
* Session-cookie regex
|
||||
* `$http_authorization != ""`
|
||||
* request `$http_cache_control` matching `no-cache|no-store|private`
|
||||
|
||||
**How security still holds:**
|
||||
|
||||
| Threat | Gate |
|
||||
| --- | --- |
|
||||
| Accidentally caching HTML/API | default `static` extensions (no html/json) |
|
||||
| Personalized content | origin `private` / `no-store` (respected by Nginx) |
|
||||
| Response writes session | **`Set-Cookie` → not stored** (§3.6) |
|
||||
| `all` too broad | UI/doc warning: needs correct origin Cache-Control |
|
||||
| API with Bearer | rely on policy (don't use `all` for APIs) + origin headers; full Auth conditional caching is later |
|
||||
|
||||
### 3.5 Default Edge TTL (no origin freshness)
|
||||
|
||||
Aligned with CF's per-status default TTL without `Cache-Control`/`Expires`, emitted in cache-enabled locations:
|
||||
|
||||
| Status | TTL |
|
||||
| --- | --- |
|
||||
| 200, 206, 301 | 120m |
|
||||
| 302, 303 | 20m |
|
||||
| 404, 410 | 3m |
|
||||
|
||||
```nginx
|
||||
proxy_cache_valid 200 206 301 120m;
|
||||
proxy_cache_valid 302 303 20m;
|
||||
proxy_cache_valid 404 410 3m;
|
||||
```
|
||||
|
||||
* When the origin provides valid `Cache-Control` / `Expires`, the origin freshness wins (no `proxy_ignore_headers`).
|
||||
* **No** forced Edge TTL override ignoring origin headers.
|
||||
|
||||
### 3.6 Response Side: Set-Cookie Not Stored
|
||||
|
||||
Aligned with CF OCC default: an eligible request whose origin returns **`Set-Cookie`** is **not written** into `proxy_cache` (read path may still have MISS/BYPASS semantics).
|
||||
|
||||
```nginx
|
||||
proxy_no_cache $openflare_skip_cache $upstream_http_set_cookie;
|
||||
```
|
||||
|
||||
(`proxy_no_cache` multi-arg: any non-empty and non-`"0"` arg means no write.)
|
||||
|
||||
`proxy_cache_bypass` still only binds `$openflare_skip_cache` (request-side skip); the response side only affects **writes**, consistent with CF "eligible but response not cacheable".
|
||||
|
||||
### 3.7 Relationship with Origin Headers
|
||||
|
||||
* **Eligibility**: policy + method bypass.
|
||||
* **Store / duration**: origin `Cache-Control` / `Expires` + default `proxy_cache_valid` + Set-Cookie gate + global `inactive`.
|
||||
|
||||
---
|
||||
|
||||
## 4. Rendering and Data Flow
|
||||
|
||||
```text
|
||||
Global cache_enabled?
|
||||
│ no → no proxy_cache_* generated
|
||||
▼ yes
|
||||
Route cache_enabled?
|
||||
│ no → location without proxy_cache
|
||||
▼ yes
|
||||
set $openflare_skip_cache 0
|
||||
→ non-GET → set 1
|
||||
→ policy if (static/all/suffix/…) → may set 1
|
||||
proxy_cache openflare_cache
|
||||
proxy_cache_methods GET
|
||||
proxy_cache_bypass $openflare_skip_cache
|
||||
proxy_no_cache $openflare_skip_cache $upstream_http_set_cookie
|
||||
proxy_cache_valid …
|
||||
→
|
||||
access.log cache_status=$upstream_cache_status
|
||||
```
|
||||
|
||||
### 4.1 Policy → Nginx Conditions
|
||||
|
||||
| Policy | Extra Condition |
|
||||
| --- | --- |
|
||||
| `static` | `$uri` not matching built-in extension table → skip |
|
||||
| `all` | no extra path condition |
|
||||
| `suffix` | not matching `cache_rules` extensions → skip |
|
||||
| `path_prefix` / `path_exact` | same as current implementation |
|
||||
|
||||
### 4.2 Code Areas Involved
|
||||
|
||||
| Area | Path |
|
||||
| --- | --- |
|
||||
| Rendering | `pkg/render/openresty/render.go` (bypass, Set-Cookie, `proxy_cache_valid`, extension constants) |
|
||||
| Validation | `internal/apps/openflare/proxy_route/helpers.go` |
|
||||
| Model/defaults | creating a route defaults `cache_policy=static`; `url`→`all` on read/write |
|
||||
| Snapshot | `config_version` snapshot normalization |
|
||||
| UI | `proxy-routes/detail/components/cache-section.tsx` |
|
||||
|
||||
---
|
||||
|
||||
## 5. Compatibility and Migration
|
||||
|
||||
| Data | Handling |
|
||||
| --- | --- |
|
||||
| `cache_policy=''` or `url` in DB (and cache enabled) | read / snapshot / render → **`all`** |
|
||||
| API write with enabled and empty policy | normalized to **`all`**; UI new-create with cache on **explicitly submits** `static` |
|
||||
| New routes | default **`static`** when cache enabled |
|
||||
| Bypass behavior change | **breaking vs. old implementation**: cookie/auth traffic goes from "not cached" to cacheable HIT; requires **republishing node configs** |
|
||||
| Default extensions | **remove `json`** from the table; sites relying on caching `*.json` can use custom `suffix` or `all` |
|
||||
|
||||
**Release note:** document this alignment with the CF default model; hit rate expected to rise; `all` and wrong origin headers need ops self-check.
|
||||
|
||||
---
|
||||
|
||||
## 6. UI Copy Points (Cache Tab)
|
||||
|
||||
* After enabling cache, default: **standard static assets** (summary extensions, **excluding HTML/JSON**; including map/mjs etc.).
|
||||
* Options: standard static / all cacheable GET (advanced) / custom suffix / path prefix / exact path.
|
||||
* CF-aligned notes:
|
||||
* login cookies are **not** separately skipped from caching;
|
||||
* origin `private` / `no-store` / response **`Set-Cookie`** are not written to the edge cache;
|
||||
* default Edge TTL used when no origin cache headers.
|
||||
* **Advanced `all`**: warn "similar to Cache Everything; personalized pages must declare private/no-store from the origin".
|
||||
* Global Performance cache master switch must be on.
|
||||
|
||||
---
|
||||
|
||||
## 7. Decision Matrix (Avoid Missed Judgments)
|
||||
|
||||
| Scenario | CF | OpenFlare (this design) |
|
||||
| --- | --- | --- |
|
||||
| GET static + session Cookie + origin public max-age | HIT | HIT |
|
||||
| GET HTML + static policy | DYNAMIC | policy skip → not cached |
|
||||
| GET + all + origin private | not stored | not stored |
|
||||
| GET static + response Set-Cookie | BYPASS (OCC) | not stored |
|
||||
| GET + Authorization + static public | conditional cache | cacheable (simplified; rely on origin not marking sensitive APIs public) |
|
||||
| GET + no-CC 200 static | default 120m | `proxy_cache_valid` 120m |
|
||||
| DevTools Disable cache (request no-cache) | edge may still HIT by default | edge may still HIT by default |
|
||||
| POST | not cached | non-GET skip |
|
||||
|
||||
---
|
||||
|
||||
## 8. Decision Record
|
||||
|
||||
| Decision | Choice | Reason |
|
||||
| --- | --- | --- |
|
||||
| Request Cookie bypass | **removed** | aligned with CF; restore static hit rate for logged-in users |
|
||||
| Request Authorization / Cache-Control bypass | **removed** | aligned with CF request-eligibility model; response gate as backstop |
|
||||
| Set-Cookie | **bind to proxy_no_cache** | aligned with CF OCC "response Set-Cookie not stored" |
|
||||
| Default Edge TTL | **per-status proxy_cache_valid** | aligned with CF default TTL when headerless, avoiding "never stored" |
|
||||
| Remove json from default table | **yes** | aligned with CF not caching JSON by default |
|
||||
| Keep map/mjs/wasm | **yes** | useful hits for modern frontend, deliberate enhancement |
|
||||
| Default cacheable scope | cache-on defaults to `static` | benchmarked to CF, reduces HTML/API mis-caching |
|
||||
| Legacy `url` | maps to `all` | doesn't narrow existing behavior |
|
||||
| Full Auth conditions / Purge / Rules | later | close the default loop first, then extend |
|
||||
@@ -0,0 +1,204 @@
|
||||
# Product Boundaries
|
||||
|
||||
You will learn: What OpenFlare is, what problems it solves, who the target audience is, what current stable features are available, and which design boundaries cannot be bypassed during implementation.
|
||||
|
||||
OpenFlare is a self-hosted OpenResty control plane designed for single-team or single-organization internal operations. It solves the problems of decentralized management of reverse proxy configurations, node synchronization, certificate hosting, configuration publication and rollback, and basic observability.
|
||||
|
||||
## Project Positioning
|
||||
|
||||
OpenFlare is suitable for teams that need to centrally manage multiple OpenResty proxy nodes:
|
||||
|
||||
* Wanting to maintain reverse proxy website configurations using a management dashboard.
|
||||
* Wanting every configuration change to have a complete version history, preview, activation, and rollback support.
|
||||
* Wanting nodes to actively synchronize configurations, rather than the control plane SSHing into nodes to execute commands.
|
||||
* Wanting to manage TLS certificates, domain assets, node statuses, and basic access analytics in a single system.
|
||||
|
||||
OpenFlare is currently not positioned as a general-purpose logging platform, service mesh, Kubernetes Ingress Controller, or multi-tenant cloud platform.
|
||||
|
||||
## Current Capabilities
|
||||
|
||||
| Capability | Description |
|
||||
| --- | --- |
|
||||
| Reverse Proxy Rules | Uses website configuration as the aggregation boundary, supporting multiple domains and origin settings. |
|
||||
| Website-level Config | One rule corresponds to one website, which can bind one or more domains and share site-level configurations. |
|
||||
| Origin Management | Maintains a lightweight origin directory and allows websites to save renderable origin snapshots. |
|
||||
| Config Versioning | Supports previews, publishing, activation, immutable history, and rollbacks. |
|
||||
| Agent Sync | Supports registration, heartbeats, synchronization, application result reporting, and self-updating. |
|
||||
| OpenResty Hosting | Manages main config templates, performance parameters, cache parameters, and Lua resources. |
|
||||
| HTTPS/TLS | Hosts certificate and domain assets, binding certificates on a per-domain basis. |
|
||||
| WAF | Maintains IP/CIDR block blacklists/whitelists, IP groups, and country-level geographic access controls at both global and site-specific levels. |
|
||||
| Basic Observability | Aggregates node requests, resource snapshots, health events, and access analytics. |
|
||||
| Node Management | Manages node status, token systems, and deployment/update lifecycles. |
|
||||
| Admin UI | Next.js-based official management dashboard. |
|
||||
| Auth Source Login | Supports configuring GitHub OAuth and standard OIDC login portals, allowing third-party accounts to bind to existing local users. |
|
||||
| Intranet Penetration | Securely exposes intranet HTTP services to the public internet using TunnelRelay nodes and the OpenFlared client, reusing the Agent's HTTPS/WAF capabilities. |
|
||||
|
||||
Default Working Model:
|
||||
|
||||
* All nodes consume the same globally activated configuration version.
|
||||
* The Server stores configurations and state, and does not directly SSH to manage nodes.
|
||||
* The Agent is the only controlled entry point on the node side.
|
||||
* TunnelRelay nodes run both the Agent (OpenResty) and the Relay (frps manager) to provide intranet penetration relays.
|
||||
* The OpenFlared client runs inside the intranet, managing the frpc process to connect to the Relay and forward traffic to intranet services.
|
||||
|
||||
## Typical Use Cases
|
||||
|
||||
| Scenario | Description |
|
||||
| --- | --- |
|
||||
| Unified Entrance | Exposes multiple internal HTTP services via a unified domain and TLS certificate. |
|
||||
| Multi-Node Sync | Multiple OpenResty nodes consume the same active configuration version. |
|
||||
| Change Review | View previews or diffs before publishing, keeping an immutable history post-publish. |
|
||||
| Rapid Rollback | Re-activate an older version, letting the Agent pull and apply it. |
|
||||
| Certificate Hosting | Bind TLS certificates to different domains under the same website. |
|
||||
| Observability | Check node health status, aggregated requests, traffic analytics, and health events. |
|
||||
| Intranet Penetration | Exposes intranet HTTP services that are not directly reachable from the public internet using Tunnels, benefiting from HTTPS, WAF, and all other protections. |
|
||||
|
||||
## Website Configuration Constraints
|
||||
|
||||
`proxy_routes` is the aggregate object for "website configurations". One record corresponds to one website, which can bind one or more domains and share a set of site-level configurations.
|
||||
|
||||
Constraints:
|
||||
|
||||
* `proxy_routes.site_name` is the unique business identifier of the website.
|
||||
* `proxy_routes.domains` must contain at least one domain, and `domains[0]` is treated as the primary domain.
|
||||
* Any domain can globally belong to only one `proxy_routes`.
|
||||
* Site-level rate limits, reverse proxies, and caching configurations are shared by the site, with no per-domain differences allowed within the same website.
|
||||
* HTTPS allows binding certificates on a per-domain basis within the same site.
|
||||
|
||||
## Origin & Upstream Constraints
|
||||
|
||||
`origins` serve the reuse of the origin directory, storing only the origin address, display name, and remarks, without carrying protocols, ports, paths, weights, or health check policies. `proxy_routes` can optionally associate with an `origins` record, but the rule internally still saves a complete upstream snapshot for rendering.
|
||||
|
||||
Upstream Constraints:
|
||||
|
||||
* `proxy_routes` must contain at least one upstream address (for direct type `direct`), or be associated with a Tunnel (for intranet penetration type `tunnel`).
|
||||
* Multi-upstream load balancing is uniformly rendered into a named `upstream` with keepalive enabled.
|
||||
* A single upstream is allowed to carry a base path or query, which is appended in `proxy_pass`. Multi-upstream is strictly limited to pure `scheme://host[:port]` structures, and all upstreams in the same rule must use the same protocol.
|
||||
* `proxy_routes.origin_host` is an optional field used to override the `Host` header during back-to-source requests.
|
||||
* All direct upstream addresses must be valid `http://` or `https://` URLs.
|
||||
* Intranet penetration upstreams must associate with a valid `tunnel_id` and specify the intranet target address and protocol.
|
||||
|
||||
## Intranet Penetration Constraints
|
||||
|
||||
OpenFlare implements intranet penetration through TunnelRelay nodes and the OpenFlared client, built on top of frp (Fast Reverse Proxy).
|
||||
|
||||
### Node & Component Model
|
||||
|
||||
**Node Types**:
|
||||
|
||||
* `nodes.node_type` distinguishes the node type: `edge_node` (edge node, default) and `tunnel_relay` (tunnel relay).
|
||||
* TunnelRelay nodes run both the Agent (OpenResty) and the Relay (frps manager) concurrently, sharing the same `agent_token`.
|
||||
- The Agent is responsible for HTTPS termination, WAF protection, caching, and rate limiting.
|
||||
- The Relay manages the frps process, providing tunnel relay services for intranet clients.
|
||||
* TunnelRelay nodes introduce new fields: `node_type`, `relay_bind_port` (frpc connection port, default 7000), `relay_vhost_http_port` (HTTP Vhost port, default 8080), `relay_auth_token` (automatically generated), `relay_status`, etc.
|
||||
|
||||
**Tunnel Client**:
|
||||
|
||||
* The `tunnels` table independently stores intranet penetration client registration info and is decoupled from the `nodes` system.
|
||||
* Each Tunnel has a unique `tunnel_id` (format `tun-<32hex>`) and `tunnel_token` (client authentication credential).
|
||||
* The OpenFlared client runs inside the intranet, is not exposed to the public internet, uses `tunnel_token` for authentication, and communicates with the Server via `/api/flared/*` endpoints.
|
||||
* An OpenFlared client can connect to multiple Relays simultaneously for high availability.
|
||||
|
||||
### Upstream Type Expansion
|
||||
|
||||
The upstream configuration of `proxy_routes` is divided into two types, distinguished by the `upstream_type` field:
|
||||
|
||||
* **Direct Upstream (`direct`, default)**: Forwards traffic directly to the origin address, behaving exactly like the existing mechanism.
|
||||
* **Intranet Penetration Upstream (`tunnel`)**: Forwards traffic to the intranet service via a TunnelRelay node.
|
||||
- Must specify `tunnel_id` (associated with the `tunnels` table).
|
||||
- Must specify `tunnel_target_addr` (intranet target address, e.g., `192.168.1.100:8080`) and `tunnel_target_protocol` (`http` or `https`).
|
||||
- During publication, the Server automatically replaces the upstream address with `http://127.0.0.1:{relay_vhost_http_port}`.
|
||||
|
||||
### Traffic Paths & Protocols
|
||||
|
||||
**Complete Data Plane Traffic Path**:
|
||||
|
||||
```
|
||||
Browser → OpenResty (Agent, TLS/WAF) [TunnelRelay Node]
|
||||
↓
|
||||
frps (Relay, HTTP Vhost Routing) [TunnelRelay Node, 127.0.0.1:{vhost_port}]
|
||||
↓
|
||||
frp Tunnel Protocol (Host Header Routing)
|
||||
↓
|
||||
frpc (Client, Multi-process) [Intranet Server]
|
||||
↓
|
||||
Intranet Service (192.168.x.x:port)
|
||||
```
|
||||
|
||||
**Key Features**:
|
||||
|
||||
* frps uses the HTTP Vhost single-port reuse mechanism; all HTTP tunnels share one `vhost_port`, automatically routed to the corresponding frpc based on the Host header.
|
||||
* The Agent preserves the original `Host` header, which frps uses to match the virtual host.
|
||||
* Each tunnel corresponds to a single `proxy_routes` and can bind multiple domains.
|
||||
* The OpenFlared client manages an independent frpc process for each connected Relay, transmitting multiple HTTP proxy definitions via a single frp tunnel.
|
||||
|
||||
### Configuration Sync Model
|
||||
|
||||
The publication process generates two types of configuration version data simultaneously, linked by a single `config_version` version number:
|
||||
|
||||
* **Agent-side Config**: OpenResty main configuration + route configurations + WAF rules. If a tunnel upstream is included, it is automatically rendered as a `http://127.0.0.1:{vhost_port}` upstream.
|
||||
* **Tunnel-side Config**: Relay list + frpc proxy definitions. Versioned alongside the publishing process; changes are hot-reloaded using `frpc reload` first.
|
||||
* **Relay Config**: Dispatched via heartbeat responses, relatively static, and not included in the versioned publishing flow.
|
||||
|
||||
### Tunnel Design Constraints
|
||||
|
||||
* Only HTTP protocol tunnel traffic is supported (keeping TCP/UDP tunnels extensible); separate TCP/UDP port allocation is not supported for now.
|
||||
* The DNS for domains using Tunnel upstreams should resolve to the designated TunnelRelay node.
|
||||
* frp binaries (v0.61+) are packaged and provided by the system deployment script or container images.
|
||||
|
||||
## HTTPS Constraints
|
||||
|
||||
`proxy_routes.domain_cert_ids` is used to record the domain-certificate bindings parallel to `domains`; a value of `0` means the domain does not have HTTPS enabled and stays HTTP-only.
|
||||
|
||||
During rendering:
|
||||
|
||||
* Domains with certificates are grouped by certificate and output as independent `443 ssl` `server` blocks.
|
||||
* Domains without certificates bound must not be automatically routed to HTTPS.
|
||||
* All domains in `proxy_routes.domains` must be kept in the same site configuration to avoid being split across version snapshots.
|
||||
|
||||
## WAF Constraints
|
||||
|
||||
WAF centers around rule groups. The system provides a single global rule group (applied to all sites by default), on top of which websites can overlay multiple custom rule groups.
|
||||
|
||||
Core Capabilities:
|
||||
|
||||
* Supports individual IP / CIDR block whitelists and blacklists.
|
||||
* Supports IP group references (including manual, automatic Expr calculated, and URL subscribed IP groups).
|
||||
* Supports GeoIP-based country/region level admission filtering.
|
||||
* Supports custom interception responses for rule groups (custom status codes and interception HTML pages, default is `418`).
|
||||
|
||||
IP Group & Judgment Constraints:
|
||||
|
||||
* **Runtime Decoupling**: The WAF runtime only reads local JSON files and does not access the Server database; configuration versions only store referenced IP group IDs. IP group members are synchronized via MD5 checksum differences and WebSocket push notifications, achieving hot activation without reloading Nginx.
|
||||
* **Built-in Expr Rules**:
|
||||
* High-frequency 404 scanning block: `request_count > 100 && status_404_ratio >= 0.8`
|
||||
* Malicious IP direct probe: `ip_host_count > 50 && ip_host_ratio > 0.5`
|
||||
* **Decision Priority**: The whitelist has absolute priority. If it does not match the whitelist, the blacklist funnel is triggered (global rule group first, custom groups matched in ascending ID order).
|
||||
* GeoIP resolution depends on the local MaxMind database; if GeoIP is anomalous, region rules are automatically ignored and must not disrupt the availability of IP rules and the main reverse proxy chain.
|
||||
|
||||
## Authentication Source Constraints
|
||||
|
||||
`auth_sources` uniformly supports `github` and `oidc` login configurations. `external_accounts` stores bindings between third-party accounts and local users. Logic for first-time third-party login:
|
||||
|
||||
* If already bound, directly authorize login; if there is an active local session, automatically bind.
|
||||
* If unbound and registration is enabled, automatically create a local account; if registration is closed, require the user to provide an existing local username and password to establish the association.
|
||||
|
||||
## Version & Observability Constraints
|
||||
|
||||
* `config_versions` must save the complete snapshot, rendering result, and `checksum`.
|
||||
* Globally, only one version can be active at a time.
|
||||
* Rollback is achieved by re-activating an older version.
|
||||
* `nodes` only carry control plane state and low-frequency summaries; they do not carry high-frequency observability facts.
|
||||
* Metrics, trends, and access analytics prioritize server-side aggregation rather than client-side temporary statistics.
|
||||
* Access detail logs are only retained within a controlled time window, not evolving into a general logging platform.
|
||||
|
||||
## Documentation Maintenance Principles
|
||||
|
||||
* Update this document when the product range or system boundaries change.
|
||||
* Update [System Architecture](./architecture.md) when the system structure or module responsibilities change.
|
||||
* Update [Agent & Publish Model](./agent-design.md) when the publishing, synchronization, rollback, or Agent model changes.
|
||||
* Update [Development Constraints](../../guideline/Constraints.md) when developer constraints, code specifications, or API conventions change.
|
||||
* Update README and [Deployment Instructions](../../deployment/deployment.md) when deployment methods change.
|
||||
* Update [Configurations Reference](../reference/configuration.md) when configuration items change.
|
||||
* Completed phases should no longer be backfilled as "version plans".
|
||||
* Before starting a new phase, complement the design first, then proceed to implementation.
|
||||
@@ -0,0 +1,109 @@
|
||||
# Uptime Kuma Sync Design
|
||||
|
||||
You will learn: the design background of the OpenFlare × Uptime Kuma monitoring integration, the control-flow design based on the Socket.IO protocol, the anti-pollution model centered on tag isolation, and the differential incremental sync state machine.
|
||||
|
||||
---
|
||||
|
||||
## Requirements Analysis
|
||||
|
||||
In a multi-node gateway architecture, monitoring system state and reverse proxy route state are usually disconnected:
|
||||
1. **High entry overhead**: every time the gateway control plane adds or decommissions a site, the admin must re-configure the corresponding probe address and alert policy in the monitoring system (e.g. Uptime Kuma).
|
||||
2. **Data inconsistency**: when a proxy route domain changes or switches to HTTPS, monitoring parameters are easily left un-updated, causing false positives or missed alerts.
|
||||
3. **Environment pollution risk**: a full "delete-recreate" sync in monitoring would wipe historical statistics and SLA curves, and would also affect other monitor tasks the user configured manually on the instance that are unrelated to the gateway.
|
||||
|
||||
To address these, OpenFlare introduces a **Uptime Kuma auto-monitoring sync mechanism** based on the client/server model, achieving strongly consistent, low-overhead, zero-pollution synchronization between gateway site route definitions and the availability monitoring system.
|
||||
|
||||
---
|
||||
|
||||
## Core Architecture
|
||||
|
||||
The Uptime Kuma sync subsystem runs entirely in the **Server control plane** background scheduler.
|
||||
|
||||
```text
|
||||
[ OpenFlare Control Plane / DB ] [ Uptime Kuma Instance ]
|
||||
│ │
|
||||
1. Scheduled Cron trigger (Job) │
|
||||
│ │
|
||||
2. Read proxy routes & options config │
|
||||
│ │
|
||||
3. Connect to Socket.IO <──── 4. Socket.IO handshake & login ────┤
|
||||
│ │
|
||||
├────── 5. Validate / create "OpenFlare" tag ──►│
|
||||
├────── 6. Compare site attrs vs Kuma monitor list ─►│
|
||||
│ │
|
||||
└────── 7. Execute differential ops (add / edit / delete) ─►│
|
||||
```
|
||||
|
||||
The sync subsystem does not pass through the data-plane Agent nodes; the Server talks directly to Uptime Kuma's exposed Socket.IO endpoint. This reduces edge node network overhead and keeps auth credentials (Kuma username/password) safely inside the control plane.
|
||||
|
||||
---
|
||||
|
||||
## Tag Isolation and Anti-Pollution Design
|
||||
|
||||
To run safely in a shared Uptime Kuma instance without disturbing manually created monitors, a **dedicated tag isolation mechanism** is used:
|
||||
|
||||
1. **`OpenFlare`-specific tag**:
|
||||
* On first connect, the sync routine calls `getTags` to fetch all tags in the instance.
|
||||
* It checks whether a tag named `OpenFlare` exists (default color indigo `#4f46e5`). If not, it creates it automatically via the `addTag` API.
|
||||
2. **Filtered scope**:
|
||||
* After fetching Uptime Kuma's monitor list (`monitorList`), the sync task only keeps monitors **tagged with `OpenFlare`**.
|
||||
* All modification comparisons (`editMonitor`) and offline cleanups (`deleteMonitor`) operate **only within this filtered subset**. Any monitor not bound with the `OpenFlare` tag is "invisible" to the sync routine — perfect anti-pollution isolation.
|
||||
|
||||
---
|
||||
|
||||
## Differential Sync State Machine
|
||||
|
||||
On each run, the sync routine computes a diff between OpenFlare's local config and Uptime Kuma's data, then executes different Socket.IO events based on the comparison:
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> 检查站点状态与监控范围
|
||||
|
||||
state "检查监控范围" as Scope {
|
||||
[*] --> 校验站点是否启用并且在 Scope 内
|
||||
校验站点是否启用并且在 Scope 内 --> 在Scope内 : 是
|
||||
校验站点是否启用并且在 Scope 内 --> 不在Scope内 : 否
|
||||
}
|
||||
|
||||
不在Scope内 --> 检查Kuma中是否存在同名且带标签的监控
|
||||
检查Kuma中是否存在同名且带标签的监控 --> 执行清理 : 存在
|
||||
检查Kuma中是否存在同名且带标签的监控 --> 忽略 : 不存在
|
||||
|
||||
在Scope内 --> 检查Kuma中是否存在同名监控
|
||||
|
||||
state "比对属性" as Compare {
|
||||
[*] --> 检查是否存在
|
||||
检查是否存在 --> 新建监控项 : 否
|
||||
检查是否存在 --> 比对元数据 : 是
|
||||
比对元数据 --> 属性一致 : 匹配
|
||||
比对元数据 --> 属性不一致 : 不匹配
|
||||
}
|
||||
|
||||
新建监控项 --> 发送add指令并绑定Tag
|
||||
属性不一致 --> 发送editMonitor指令
|
||||
属性一致 --> 忽略
|
||||
|
||||
执行清理 --> 发送deleteMonitor指令
|
||||
忽略 --> [*]
|
||||
```
|
||||
|
||||
### 1. Monitor URL Normalization
|
||||
A site route in OpenFlare can configure multiple domains; the sync routine automatically extracts the primary domain and assembles a standard `http://` or `https://` prefix based on whether HTTPS is enabled.
|
||||
|
||||
### 2. Compared Attribute Set
|
||||
If a same-named, tagged monitor already exists, the sync routine compares the following 5 key fields against the current gateway global option. Any mismatch triggers an update:
|
||||
* **URL**: `Url`
|
||||
* **Probe interval**: `Interval` (default 60s)
|
||||
* **Max retries**: `MaxRetries`
|
||||
* **Retry interval**: `RetryInterval` (default 60s)
|
||||
* **Request timeout**: `Timeout` (default 48s)
|
||||
|
||||
---
|
||||
|
||||
## Scheduler and High-Concurrency Protection
|
||||
|
||||
1. **Cron-based single-thread execution**:
|
||||
* The Server periodically (every 1 minute) probes via a background Cron Job whether the configured sync interval (`UptimeKumaSyncInterval`) is reached.
|
||||
* The task uses mutex locking internally. If a previous sync request is still running due to network latency, the next schedule is skipped automatically, preventing concurrent Socket.IO connections from DDOS-ing the Uptime Kuma instance.
|
||||
2. **WebSocket state listening**:
|
||||
* The sync routine uses Socket.IO's event listener; after the connection is established, it only proceeds to the differential algorithm once the full `monitorList` event list push is received, avoiding monitor deletion caused by incomplete data loading.
|
||||
@@ -0,0 +1,127 @@
|
||||
# Login CAPTCHA Integration (Cap)
|
||||
|
||||
This document describes the design of introducing **Cap** — an open-source CAPTCHA solution based on Proof-of-Work (PoW) and invisible browser fingerprint features — into the OpenFlare control plane, to protect the login API against brute-force attacks and credential-stuffing by crawlers.
|
||||
|
||||
---
|
||||
|
||||
## 1. Business Background and Product Scope
|
||||
|
||||
### Background and Pain Points
|
||||
The OpenFlare login endpoint `/api/v1/user/login` lacks user-dimension protection; attackers can use proxy pools to perform credential stuffing and brute-force attacks on high-privilege accounts (such as `root`). At the same time, standard visual CAPTCHAs are unfriendly to login-page UX and accessibility.
|
||||
|
||||
### Product Scope and Technology Choice
|
||||
* **Technology choice**: Cap (a Proof-of-Work-driven, invisible, image-free CAPTCHA solution).
|
||||
- **Core principle**: the client (Widget/page) obtains a proof-of-work (PoW) challenge from the server, computes the solution in the browser background, and sends the answer back. The server verifies the answer to complete human-machine verification.
|
||||
- **Advantages**: invisible, image-free, no dependency on external third-party API nodes (private), tiny package size.
|
||||
* **Integration scope**: the control-plane Server login API (`/api/v1/user/login`) and the frontend login page.
|
||||
* **Config granularity**: admins can toggle the CAPTCHA on/off anytime via the console Option table (`cap_login_enabled`).
|
||||
|
||||
---
|
||||
|
||||
## 2. System Architecture and Interaction Sequence
|
||||
|
||||
### 2.1 Module Responsibilities
|
||||
1. **Frontend**:
|
||||
* Introduces the `cap-widget` (React 19 custom element) on the login page.
|
||||
* On form submit, accompanies the submission with the `cap-token` solved by the Widget.
|
||||
2. **Server (control-plane backend)**:
|
||||
* Exposes `POST /api/cap/challenge` to distribute the PoW challenge and a signed JWT token to the client.
|
||||
* Exposes `POST /api/cap/redeem` to verify the submitted PoW solution and issue a login credential (Redeem Token) with an expiry time.
|
||||
* Stores the Redeem Token and its expiry in the in-memory/Redis cache.
|
||||
* In `POST /api/v1/user/login`, when CAPTCHA protection is enabled, first validates and consumes (single-use) the corresponding `cap-token`.
|
||||
|
||||
### 2.2 Verification Flow Sequence Diagram
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
autonumber
|
||||
actor User as User
|
||||
participant Browser as Browser (Frontend Web)
|
||||
participant Server as OpenFlare Server (Backend)
|
||||
participant Cache as Memory/Redis Cache
|
||||
|
||||
User->>Browser: Open login page
|
||||
Browser->>Server: POST /api/cap/challenge (get challenge)
|
||||
Server->>Browser: Return {challenge, token, expires} (JWT format)
|
||||
Note over Browser: Widget computes the PoW challenge in background (WASM/Worker)
|
||||
Browser->>Server: POST /api/cap/redeem (submit solutions + token)
|
||||
alt PoW solution valid
|
||||
Server->>Cache: Store Redeem Token (tokenKey:expires)
|
||||
Server->>Browser: Return {success: true, token} (i.e. cap-token)
|
||||
else validation failed
|
||||
Server->>Browser: Return {success: false, reason}
|
||||
end
|
||||
User->>Browser: Enter account/password, click login
|
||||
Browser->>Server: POST /api/v1/user/login (with X-Cap-Token in HTTP header)
|
||||
alt CapLoginEnabled = true
|
||||
Server->>Server: Middleware (CapAuth) validates and consumes X-Cap-Token
|
||||
alt token valid, not expired, not consumed
|
||||
Server->>Server: c.Next() -> normal login logic (Bcrypt password check)
|
||||
Server->>Browser: Return login success (Session Cookie)
|
||||
else token invalid or already consumed
|
||||
Server->>Browser: Intercept and return CAPTCHA error (401 Unauthorized)
|
||||
end
|
||||
else CapLoginEnabled = false
|
||||
Server->>Server: c.Next() -> normal login logic
|
||||
end
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Core APIs and Data Model
|
||||
|
||||
### 3.1 API Definitions
|
||||
|
||||
#### 1. Get Challenge (POST /api/cap/challenge)
|
||||
* **Method**: `POST`
|
||||
* **Auth**: public
|
||||
* **Response payload** (unified API envelope, `data` is the business payload):
|
||||
```json
|
||||
{
|
||||
"error_msg": "",
|
||||
"data": {
|
||||
"challenge": {
|
||||
"c": 1,
|
||||
"s": 32,
|
||||
"d": 4
|
||||
},
|
||||
"token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...",
|
||||
"expires": 1717660800000
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### 2. Redeem Challenge (POST /api/cap/redeem)
|
||||
* **Method**: `POST`
|
||||
* **Request payload**:
|
||||
```json
|
||||
{
|
||||
"token": "challenge_jwt_token_here",
|
||||
"solutions": [12345, 67890, 54321]
|
||||
}
|
||||
```
|
||||
* **Response payload (success)**:
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"token": "random_id:ver_token",
|
||||
"expires": 1717661000000
|
||||
}
|
||||
```
|
||||
|
||||
#### 3. Login API (POST /api/v1/user/login)
|
||||
* **Request payload unchanged**:
|
||||
```json
|
||||
{
|
||||
"username": "root",
|
||||
"password": "your_password"
|
||||
}
|
||||
```
|
||||
* **CAPTCHA carrier**: placed in the HTTP Request Header `X-Cap-Token`.
|
||||
|
||||
---
|
||||
|
||||
## 4. Replay Attack Protection and Security Trade-offs
|
||||
1. **JWT temporary state binding**: the challenge is signed into the JWT payload at generation time, including an expiry limit (10 minutes).
|
||||
2. **Replay interception (nonce consumption)**: when the client calls `/redeem` to submit the solution, the backend marks the JWT signature as used in the cache. Re-submitting the same solution package returns `already_redeemed`.
|
||||
3. **Redeem single-use (one-time invalidation)**: when the client logs in and submits the `cap-token`, the backend immediately deletes the key from the cache after validating it, preventing attackers from extracting historical valid `cap-token`s for login replay.
|
||||
4. **Seamless verification**: by tuning parameters like `c` (challenge count) and `d` (difficulty), you balance solve time against anti-crawler strength; users solve silently in the background without interrupting the login flow.
|
||||
@@ -0,0 +1,86 @@
|
||||
# Log Store Decoupling
|
||||
|
||||
You will learn: which tables are log-purpose, why they must not be pinned to ClickHouse, and which code path a new log table must follow.
|
||||
|
||||
Observability fields and the reporting protocol are still governed by [Observability Protocol & Tables](./observability-data-model.md); this document only defines **where data is stored and how to switch databases**.
|
||||
|
||||
---
|
||||
|
||||
## 1. Goals
|
||||
|
||||
* **ClickHouse optional**: when not enabled, PostgreSQL (or SQLite when the primary DB is off) fully takes over writes, queries, aggregation, and cleanup.
|
||||
* **Upper layers don't touch the underlying DB**: apps only face `internal/repository/logstore` (or the `repository` facade). `repository/analytics` and `db.ChConn` / `db.ChDB` are only used by logstore's ClickHouse implementation.
|
||||
* **Switchable**: 「Switch Log Database」in Task Management copies data between PostgreSQL/SQLite and ClickHouse and flips the primary; writes are frozen during migration, the switch only happens on success, and source data is not deleted.
|
||||
|
||||
---
|
||||
|
||||
## 2. What Counts as a Log Table
|
||||
|
||||
A table enters logstore only if it meets all of:
|
||||
|
||||
* Append-only writes, almost no row updates
|
||||
* Query or aggregate by time, deletable by retention days
|
||||
* Must still support writes and queries when ClickHouse is off
|
||||
* Does not participate in transactional consistency for websites / nodes / certificates, etc.
|
||||
|
||||
**Don't** make these log tables: Zones, nodes, config versions, task executions, upload metadata. These go through the business primary DB `repository`.
|
||||
|
||||
Current log domains:
|
||||
|
||||
| Domain | Interface | Tables |
|
||||
| --- | --- | --- |
|
||||
| Node access logs | `AccessLogStore` | `of_node_access_logs` |
|
||||
| Observability time series | `ObservabilityStore` | `of_node_metric_snapshots` / `of_node_edge_health` / `of_node_obs_frps` / `of_node_obs_frpc` |
|
||||
| User access audit | `UserAccessLogStore` | `w_user_access_logs` |
|
||||
|
||||
Hourly materialized views on ClickHouse (e.g. `of_access_log_hourly`) only serve CH query acceleration. PostgreSQL / SQLite **do not** build isomorphic aggregation tables; queries aggregate in real time from raw logs.
|
||||
|
||||
---
|
||||
|
||||
## 3. Layering
|
||||
|
||||
| Layer | Path | Responsibility |
|
||||
| --- | --- | --- |
|
||||
| Abstraction | `internal/repository/logstore` | Interfaces + `Active` / `BuildForMigration`; selects implementation by `log_database` |
|
||||
| CH implementation | `logstore/clickhouse_store.go` | Delegates to `repository/analytics` (native batch + existing aggregation SQL) |
|
||||
| Primary DB implementation | `logstore/postgres_store.go` | PostgreSQL (high-frequency tables partitioned monthly) and SQLite (plain tables) share GORM |
|
||||
| Model | `internal/model/analytics` | Entities and batch SQL, no IO |
|
||||
| Enqueue | `chwriter` / `risk_control` + `batchwriter` | `FlushFunc` calls `logstore.Active`; node logs / observability enqueue via hooks |
|
||||
| Constraint | `logstore/imports_test.go` | apps are forbidden from importing `repository/analytics` |
|
||||
|
||||
`log_database` has only two legal states: **follow the business primary DB** (`postgres` or `sqlite`) or **`clickhouse`**. "Primary PostgreSQL + log SQLite" does not exist. `log_database` / `log_db_migration` are protected and cannot be changed from the admin panel.
|
||||
|
||||
At startup: `log_database=clickhouse` but ClickHouse not enabled → startup is refused; you must re-enable ClickHouse, switch back to the primary DB, and only then turn it off.
|
||||
|
||||
---
|
||||
|
||||
## 4. Switch Protocol
|
||||
|
||||
Task type `of_log_db_switch` (admin name 「Switch Log Database」), parameter `target`.
|
||||
|
||||
1. Validate the target is legal and not the current DB.
|
||||
2. Write `log_db_migration=migrating`, drain in-flight batchwriter (`Drain`, not `Stop` writer). Writes return a clear error afterward (HTTP 503), not queued backlog.
|
||||
3. Clear the target log tables, then copy by id in pages; call `EnsurePartitions` on the PostgreSQL target before copying.
|
||||
4. Only on full success write `log_database=target` and clear the migration marker; on failure clear the marker and writes continue on the source DB.
|
||||
5. Source data is not deleted; re-clear the target before retry to guarantee idempotency.
|
||||
|
||||
Don't invent another switch protocol, and don't connect `analyticsrepo` directly inside tasks.
|
||||
|
||||
---
|
||||
|
||||
## 5. Adding a New Log Table
|
||||
|
||||
Column names must be identical across the three goose migrations (ClickHouse / PostgreSQL / SQLite). Key points:
|
||||
|
||||
* High-frequency tables: CH uses `MergeTree` + `toYYYYMM`; PG uses `PARTITION BY RANGE(time column)` with the partition key in the primary key; SQLite uses a plain table + indexes.
|
||||
* IDs use snowflake `uint64`, preserved as-is during migration.
|
||||
* Writes go through a dedicated `batchwriter`; flush calls `logstore.Active`, not `analyticsrepo.BatchInsert`.
|
||||
* The switch task's `copy*` must cover the new table; cleanup uses existing `log_retention_days_*` or `metric_retention_days`, don't use the wrong TTL.
|
||||
|
||||
Runtime config: [Configuration Reference · Log Storage](../reference/configuration.md#8-日志存储log-database).
|
||||
|
||||
---
|
||||
|
||||
## 6. Related Docs
|
||||
|
||||
* Observability fields and reporting protocol: [Observability Protocol & Tables](./observability-data-model.md)
|
||||
@@ -0,0 +1,769 @@
|
||||
# Agent Reporting Protocol and Observability Data Model
|
||||
|
||||
You will learn: the **data structures** of the refactored Agent heartbeat/WS reports, how the Server **parses and writes** them, and the **target table structures** in ClickHouse / relational DBs.
|
||||
**No protocol compatibility layer**: Agents upgrade via destroy-and-recreate or binary replacement; old fields are not parsed, old buffers are discarded wholesale.
|
||||
|
||||
This design is the **protocol & storage chapter** of [Edge Observability & Business Traffic Stats Refactor](./observability-design.md); implement against the fields and DDL in this document.
|
||||
|
||||
**First read the transport overview and examples:** [Observability Transport Model](./observability-transport-model.md).
|
||||
|
||||
---
|
||||
|
||||
## 1. Design Goals
|
||||
|
||||
| Goal | Description |
|
||||
| --- | --- |
|
||||
| Agent reports only facts | details + host readings + edge health instant state; no business pre-aggregation |
|
||||
| One business detail table | access logs are the only L1 write path |
|
||||
| Aggregation in DB/control plane | hourly summaries come from ClickHouse MV or queries; the Agent never writes summary tables |
|
||||
| No field overlap | `bytes_sent` = data provided; NIC `network_*` = host; no business `openresty_tx` anymore |
|
||||
| Evolvable | new fields optional; missing numeric values default to 0; removed legacy protocol fields are not parsed |
|
||||
|
||||
---
|
||||
|
||||
## 2. Layering and Write Overview
|
||||
|
||||
```text
|
||||
Agent NodePayload (v2)
|
||||
│
|
||||
┌───────────────┼───────────────┐
|
||||
▼ ▼ ▼
|
||||
access_logs host_metrics edge_health
|
||||
(L1 details) (L3 readings) (L2 instant)
|
||||
│ │ │
|
||||
▼ ▼ ▼
|
||||
of_node_access_logs of_node_metric_ of_node_edge_health
|
||||
│ snapshots │
|
||||
│ │ │
|
||||
▼ ▼ │
|
||||
of_access_log_hourly of_node_metric_ │
|
||||
(MV, Server side) capacity_hourly (MV) │
|
||||
│ │ │
|
||||
└─────── admin aggregation API ────┘
|
||||
|
||||
Relational DB (PostgreSQL/SQLite): node latest state, Profile, health events (not a detail lake)
|
||||
```
|
||||
|
||||
| Layer | Meaning | Agent Report Block | ClickHouse Fact Table |
|
||||
| --- | --- | --- | --- |
|
||||
| L1 | business delivery | `access_logs` | `of_node_access_logs` |
|
||||
| L2 | edge health | `edge_health` | `of_node_edge_health` |
|
||||
| L3 | host capacity | `host_metrics` | `of_node_metric_snapshots` |
|
||||
|
||||
---
|
||||
|
||||
## 3. Agent Report Data Structures (protocol v2)
|
||||
|
||||
### 3.1 Top-Level `NodePayload`
|
||||
|
||||
Transport: HTTP heartbeat body and WebSocket `status` messages share the same structure.
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n_xxx",
|
||||
"name": "edge-1",
|
||||
"ip": "1.2.3.4",
|
||||
"version": "3.3.0",
|
||||
"ext_version": "",
|
||||
"current_version": "cfg-checksum-or-version",
|
||||
"last_error": "",
|
||||
"profile": { },
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ],
|
||||
"buffered": [ ],
|
||||
"health_events": [ ],
|
||||
"waf_ip_group_checksums": { "1": "md5..." }
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `schema_version` | int | suggested | fixed to `2` (this design) |
|
||||
| `node_id` | string | ✅ | node ID |
|
||||
| `name` | string | ✅ | display name |
|
||||
| `ip` | string | ✅ | reporting IP |
|
||||
| `version` / `ext_version` | string | ✅ | Agent version |
|
||||
| `current_version` | string | | locally active config version summary |
|
||||
| `last_error` | string | | latest sync/runtime error, nullable |
|
||||
| `openresty_status` | string | ✅ (when OpenResty present) | **latest health-state authoritative field** → written to PG node table |
|
||||
| `openresty_message` | string | | **latest health-description authoritative field** → written to PG node table (**not into CH**) |
|
||||
| `profile` | object | | host overview, report on change (may throttle) |
|
||||
| `host_metrics` | object | suggested each beat | L3 resource snapshot |
|
||||
| `edge_health` | object | suggested each beat | L2 connection time series + status aligned with top level |
|
||||
| `access_logs` | array | | this beat's incremental access details |
|
||||
| `buffered` | array | | offline backfill fact batches (see §3.6) |
|
||||
| `health_events` | array | | edge health events |
|
||||
| `waf_ip_group_checksums` | map | | for differential sync, not an observability lake |
|
||||
|
||||
**Removed, Server no longer parses (no compatibility layer):**
|
||||
|
||||
| Old Field | Disposition |
|
||||
| --- | --- |
|
||||
| `traffic_report` | not in the protocol; not stored |
|
||||
| `openresty_observation` | not present; connections/status go through `edge_health` |
|
||||
| `snapshot` | not present; only `host_metrics` |
|
||||
| `buffered_observability` | not present; only `buffered` |
|
||||
|
||||
### 3.2 `profile` — Host Overview (low frequency)
|
||||
|
||||
Maps to relational `of_node_system_profiles` (or an existing equivalent), **not into the ClickHouse detail lake**.
|
||||
|
||||
```json
|
||||
{
|
||||
"hostname": "edge-1",
|
||||
"os_name": "linux",
|
||||
"os_version": "...",
|
||||
"kernel_version": "...",
|
||||
"architecture": "amd64",
|
||||
"cpu_model": "...",
|
||||
"cpu_cores": 8,
|
||||
"total_memory_bytes": 16106127360,
|
||||
"total_disk_bytes": 107374182400,
|
||||
"uptime_seconds": 864000,
|
||||
"reported_at_unix": 1720000000
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Semantics |
|
||||
| --- | --- |
|
||||
| hardware/OS description fields | factual readings |
|
||||
| `reported_at_unix` | Agent collection time (UTC seconds) |
|
||||
|
||||
### 3.3 `host_metrics` — Host Capacity (L3)
|
||||
|
||||
**All readings, no 24h business totals.**
|
||||
NIC/disk bytes are **kernel cumulative counter raw values** (monotonically increasing, may reset on restart); CPU is an instant percentage; memory/disk usage is current usage.
|
||||
|
||||
```json
|
||||
{
|
||||
"captured_at_unix": 1720000000,
|
||||
"cpu_usage_percent": 12.5,
|
||||
"memory_used_bytes": 4294967296,
|
||||
"memory_total_bytes": 16106127360,
|
||||
"storage_used_bytes": 50000000000,
|
||||
"storage_total_bytes": 107374182400,
|
||||
"disk_read_bytes": 9000000000,
|
||||
"disk_write_bytes": 12000000000,
|
||||
"network_rx_bytes": 500000000000,
|
||||
"network_tx_bytes": 800000000000
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Semantics | How Server Uses It |
|
||||
| --- | --- | --- | --- |
|
||||
| `captured_at_unix` | int64 | sampling time | `captured_at` |
|
||||
| `cpu_usage_percent` | float | instant CPU% | store directly; average for trends |
|
||||
| `memory_*` / `storage_*` | int64 | current used/total | store directly; compute usage rate |
|
||||
| `disk_read_bytes` / `disk_write_bytes` | int64 | **cumulative** IO bytes | store raw; adjacent deltas at query time |
|
||||
| `network_rx_bytes` / `network_tx_bytes` | int64 | **cumulative** NIC bytes | store raw; adjacent deltas at query time → "host NIC in/outbound" |
|
||||
|
||||
> The Agent is **forbidden** from replacing cumulative values with "this period's delta" before reporting (otherwise Server deltas would be wrong).
|
||||
|
||||
### 3.4 `edge_health` — OpenResty Edge Health (L2)
|
||||
|
||||
**Instant state only, no business throughput.**
|
||||
|
||||
```json
|
||||
{
|
||||
"captured_at_unix": 1720000000,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 42
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Semantics |
|
||||
| --- | --- | --- |
|
||||
| `status` | string | `healthy` / `unhealthy` / `unknown` (must match top-level `openresty_status`) |
|
||||
| `message` | string | status description (may be reported; **only backfills PG latest state, not into CH**) |
|
||||
| `connections` | int64 | stub_status Active connections |
|
||||
|
||||
#### Health-State Authoritative Sources (converged)
|
||||
|
||||
| Data | Authoritative Storage | Description |
|
||||
| --- | --- | --- |
|
||||
| **Current** OpenResty health + description | **PG node table** `openresty_status` / `openresty_message` | UI badges, lists, alerts use this |
|
||||
| **Time series** health status + connections | **CH** `of_node_edge_health` (`status`, `connections`) | connection curves / health history; **no message column** |
|
||||
| Agent report | top-level status/message + `edge_health` | Server normalizes both statuses aligned; message **only written to PG** |
|
||||
|
||||
So: "is it unhealthy now" → read PG; "connections over the past 24h" → read CH.
|
||||
|
||||
### 3.5 `access_logs[]` — Access Details (L1, single business truth)
|
||||
|
||||
Agent: tail access.log → parse JSON lines → report fields as-is (path may be truncated).
|
||||
|
||||
```json
|
||||
{
|
||||
"logged_at_unix": 1720000001,
|
||||
"remote_addr": "203.0.113.10",
|
||||
"host": "www.example.com",
|
||||
"path": "/api/v1/ping",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 1024,
|
||||
"request_length": 128,
|
||||
"request_time_ms": 15,
|
||||
"user_agent": "Mozilla/5.0 ...",
|
||||
"cache_status": "HIT"
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Required | Source (OpenResty) | Business Meaning |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| `logged_at_unix` | int64 | ✅ | parse `$time_iso8601` | request completion time |
|
||||
| `remote_addr` | string | ✅ | `$remote_addr` | client IP → UV |
|
||||
| `host` | string | ✅ | `$host` | domain → Zone ownership |
|
||||
| `path` | string | ✅ | `$request_uri`, Agent may truncate | path |
|
||||
| `status_code` | int | ✅ | `$status` | status code |
|
||||
| `bytes_sent` | int64 | ✅ | **`$body_bytes_sent`** | **data provided** (response body) |
|
||||
| `request_length` | int64 | suggested | `$request_length` | **data received** |
|
||||
| `request_time_ms` | int64 | optional | `$request_time * 1000` | latency; default 0 |
|
||||
| `user_agent` | string | suggested | `$http_user_agent` | UA; may truncate on store |
|
||||
| `cache_status` | string | suggested | **`$upstream_cache_status`** | edge cache result (see §3.5.1) |
|
||||
|
||||
**Explicitly not reported by Agent (written by Server):**
|
||||
|
||||
* `region` / country: GeoIP resolved at insert time
|
||||
* `id` / `created_at`: Server-generated
|
||||
* `node_id`: from payload / auth context
|
||||
|
||||
**Explicitly not reported:**
|
||||
|
||||
* `upstream_addr` / origin address / `origin_fetched`: no origin-endpoint tracking; "did it fetch from origin" is only derived from `cache_status` at the control plane (§3.5.1)
|
||||
|
||||
### 3.5.1 `cache_status` — Cache Hit and Origin Fetch (detail-first)
|
||||
|
||||
**Goal (phase 1):** access log details/list can show "cache hit / origin fetch / no cache used".
|
||||
**Caliber:** only store the raw OpenResty `$upstream_cache_status`; **no upstream address reported**.
|
||||
|
||||
#### Raw Values (stored)
|
||||
|
||||
| Value | Meaning (OpenResty) |
|
||||
| --- | --- |
|
||||
| `HIT` | cache hit |
|
||||
| `MISS` | miss, fetched from upstream |
|
||||
| `BYPASS` | cache skipped (e.g. method/cookie/policy caused `$openflare_skip_cache`) |
|
||||
| `EXPIRED` | expired then origin fetch |
|
||||
| `STALE` | served stale |
|
||||
| `UPDATING` | background updating, may return old cache |
|
||||
| `REVALIDATED` | revalidated, still used cache |
|
||||
| `-` or empty | didn't pass through `proxy_cache` (e.g. Pages local static, non-proxy location) |
|
||||
|
||||
#### UI Three-State Derivation (not stored)
|
||||
|
||||
Control-plane display uses derived enum `cache_outcome`, **not written to CH**:
|
||||
|
||||
| Three-State | Condition (`cache_status`) | Suggested List Label |
|
||||
| --- | --- | --- |
|
||||
| **Cache hit** | `HIT` / `STALE` / `REVALIDATED` / `UPDATING` | hit |
|
||||
| **Origin fetch** | `MISS` / `EXPIRED` | origin |
|
||||
| **No cache used** | `BYPASS` / `-` / `""` | not cached |
|
||||
|
||||
Details can show both the three-state and the raw `cache_status`.
|
||||
|
||||
#### Boundaries
|
||||
|
||||
* Pages static / locations without `proxy_cache`: mostly empty or `-` → **no cache used**, must not be labeled "hit".
|
||||
* Detail pages show cache state; hit-rate dashboards and hourly dimensions can extend from the same column.
|
||||
|
||||
**Per-heartbeat count suggestion:**
|
||||
|
||||
* Soft cap e.g. 2000 lines/beat; overflow goes into `buffered` next batch, **forbidden** to compress into a TrafficReport in the Agent.
|
||||
|
||||
### 3.6 `buffered[]` — Offline Backfill (facts only)
|
||||
|
||||
```json
|
||||
{
|
||||
"captured_at_unix": 1719999900,
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ]
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Description |
|
||||
| --- | --- |
|
||||
| `captured_at_unix` | batch collection/buffer time, used for ack and dedup window |
|
||||
| `host_metrics` / `edge_health` / `access_logs` | same structures as the main payload; empty blocks may be omitted |
|
||||
|
||||
**Forbidden** to carry `traffic_report` or rx/tx throughput in buffered.
|
||||
|
||||
### 3.7 `health_events[]`
|
||||
|
||||
```json
|
||||
{
|
||||
"event_type": "openresty_unhealthy",
|
||||
"severity": "critical",
|
||||
"message": "...",
|
||||
"triggered_at_unix": 1720000000,
|
||||
"metadata": { }
|
||||
}
|
||||
```
|
||||
|
||||
Written to the relational health-event table (existing model suffices), not into the access log lake.
|
||||
|
||||
### 3.8 Go Protocol Structures
|
||||
|
||||
```go
|
||||
// pkg/protocol/agent.go (current implementation)
|
||||
|
||||
type NodePayload struct {
|
||||
SchemaVersion int `json:"schema_version,omitempty"`
|
||||
NodeID string `json:"node_id"`
|
||||
Name string `json:"name"`
|
||||
IP string `json:"ip"`
|
||||
Version string `json:"version"`
|
||||
ExtVersion string `json:"ext_version"`
|
||||
CurrentVersion string `json:"current_version"`
|
||||
LastError string `json:"last_error"`
|
||||
OpenrestyStatus string `json:"openresty_status"` // PG latest-state authority
|
||||
OpenrestyMessage string `json:"openresty_message"` // PG latest-state authority; not into CH
|
||||
Profile *NodeSystemProfile `json:"profile,omitempty"`
|
||||
HostMetrics *NodeHostMetrics `json:"host_metrics,omitempty"`
|
||||
EdgeHealth *NodeEdgeHealth `json:"edge_health,omitempty"`
|
||||
AccessLogs []NodeAccessLog `json:"access_logs,omitempty"`
|
||||
Buffered []BufferedFacts `json:"buffered,omitempty"`
|
||||
HealthEvents []NodeHealthEvent `json:"health_events"`
|
||||
WAFIPGroupChecksums map[string]string `json:"waf_ip_group_checksums,omitempty"`
|
||||
}
|
||||
|
||||
type NodeHostMetrics struct {
|
||||
CapturedAtUnix int64 `json:"captured_at_unix"`
|
||||
CPUUsagePercent float64 `json:"cpu_usage_percent"`
|
||||
MemoryUsedBytes int64 `json:"memory_used_bytes"`
|
||||
MemoryTotalBytes int64 `json:"memory_total_bytes"`
|
||||
StorageUsedBytes int64 `json:"storage_used_bytes"`
|
||||
StorageTotalBytes int64 `json:"storage_total_bytes"`
|
||||
DiskReadBytes int64 `json:"disk_read_bytes"`
|
||||
DiskWriteBytes int64 `json:"disk_write_bytes"`
|
||||
NetworkRxBytes int64 `json:"network_rx_bytes"`
|
||||
NetworkTxBytes int64 `json:"network_tx_bytes"`
|
||||
}
|
||||
|
||||
type NodeEdgeHealth struct {
|
||||
CapturedAtUnix int64 `json:"captured_at_unix"`
|
||||
Status string `json:"status"`
|
||||
Message string `json:"message"`
|
||||
Connections int64 `json:"connections"`
|
||||
}
|
||||
|
||||
type NodeAccessLog struct {
|
||||
LoggedAtUnix int64 `json:"logged_at_unix"`
|
||||
RemoteAddr string `json:"remote_addr"`
|
||||
Host string `json:"host"`
|
||||
Path string `json:"path"`
|
||||
UserAgent string `json:"user_agent,omitempty"`
|
||||
CacheStatus string `json:"cache_status,omitempty"` // $upstream_cache_status
|
||||
StatusCode int `json:"status_code"`
|
||||
BytesSent int64 `json:"bytes_sent"` // body_bytes_sent, data provided
|
||||
RequestLength int64 `json:"request_length"` // data received
|
||||
RequestTimeMs int64 `json:"request_time_ms"` // optional
|
||||
}
|
||||
|
||||
type BufferedFacts struct {
|
||||
CapturedAtUnix int64 `json:"captured_at_unix"`
|
||||
HostMetrics *NodeHostMetrics `json:"host_metrics,omitempty"`
|
||||
EdgeHealth *NodeEdgeHealth `json:"edge_health,omitempty"`
|
||||
AccessLogs []NodeAccessLog `json:"access_logs,omitempty"`
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Server Parsing and Storage Flow
|
||||
|
||||
### 4.1 Entry Points
|
||||
|
||||
* HTTP: `POST /api/v1/agent/...` heartbeat (existing path)
|
||||
* WebSocket: `type=status` payload = `NodePayload`
|
||||
* Auth: `X-Agent-Token` → binds `node_id` (payload.node_id must match the token node)
|
||||
|
||||
### 4.2 Processing Pipeline (single payload)
|
||||
|
||||
```text
|
||||
1. Deserialize NodePayload
|
||||
2. Normalize
|
||||
- schema_version < 2:
|
||||
host_metrics ← snapshot
|
||||
edge_health.status ← openresty_status
|
||||
edge_health.connections ← openresty_observation.connections (if present)
|
||||
traffic_report → drop
|
||||
openresty_observation.rx/tx → drop
|
||||
buffered ← buffered_observability
|
||||
- path truncated again, status clamped, negative bytes → 0
|
||||
3. Relational transaction (node latest state)
|
||||
- update node online time, IP, version, edge_health.status/message
|
||||
- upsert profile (if present)
|
||||
- insert health_events (if present)
|
||||
4. ClickHouse async batch (log on failure; doesn't block heartbeat response config delivery)
|
||||
a. access_logs + buffered[].access_logs
|
||||
→ fill region (GeoIP)
|
||||
→ assign snowflake id
|
||||
→ BatchInsert of_node_access_logs
|
||||
b. host_metrics + buffered[].host_metrics
|
||||
→ of_node_metric_snapshots
|
||||
c. edge_health + buffered[].edge_health
|
||||
→ of_node_edge_health (connections + status snapshot only, optional)
|
||||
5. Return heartbeat response (settings / active_config / waf diff)
|
||||
6. If using buffer ack: confirm by the list of buffered.captured_at_unix
|
||||
```
|
||||
|
||||
### 4.3 Normalization Rules (hard constraints)
|
||||
|
||||
| Rule | Behavior |
|
||||
| --- | --- |
|
||||
| `logged_at` ahead of now+5m | clamp to now or drop the line (pick one in implementation and unit-test it) |
|
||||
| `logged_at` older than now−TTL | still writable; relies on table TTL cleanup |
|
||||
| empty `host` | allowed, aggregated into "unassigned" |
|
||||
| `bytes_sent` / `request_length` < 0 | set to 0 |
|
||||
| single batch access_logs > N | truncate and alert-metric it (or only queue into buffer), never switch to pre-aggregation |
|
||||
| duplicate backfill | CH tolerates a few duplicate rows; queries approximate with sum (no forced exact dedup) |
|
||||
|
||||
### 4.4 Field Mapping Table (report → table)
|
||||
|
||||
| Report Path | Target Storage | Columns |
|
||||
| --- | --- | --- |
|
||||
| `access_logs[]` | CH `of_node_access_logs` | see §5.1 |
|
||||
| `host_metrics` | CH `of_node_metric_snapshots` | see §5.2 |
|
||||
| `edge_health` | CH `of_node_edge_health` + PG node latest state | see §5.3 / §5.6 |
|
||||
| `profile` | PG `of_node_system_profiles` | existing columns |
|
||||
| `health_events` | PG health-event table | existing model |
|
||||
| `waf_ip_group_checksums` | not into observability tables | sync logic |
|
||||
| `traffic_report` (legacy) | **not written** | — |
|
||||
| `openresty_rx/tx` (legacy) | **not written** | — |
|
||||
|
||||
### 4.5 Query Side (no new "business outbound" column)
|
||||
|
||||
| Product Metric | SQL Semantics (sketch) |
|
||||
| --- | --- |
|
||||
| Data provided | `sum(bytes_sent)` |
|
||||
| Data received | `sum(request_length)` |
|
||||
| Request count | `count()` |
|
||||
| UV | `uniqExact(remote_addr)` |
|
||||
| 5xx | `countIf(status_code >= 500)` |
|
||||
| By domain/status/region | `GROUP BY host / status_code / region` |
|
||||
| Host NIC outbound | non-negative delta over `network_tx_bytes` per node in time order, then sum |
|
||||
| OpenResty connections | `of_node_edge_health.connections` latest or average |
|
||||
|
||||
---
|
||||
|
||||
## 5. Table Structures (DDL)
|
||||
|
||||
> Engine and TTL tend to match production: access logs 90 days, metrics 30 days.
|
||||
> `id` uses control-plane Snowflake/unique UInt64.
|
||||
|
||||
### 5.1 L1 Fact Table: `of_node_access_logs`
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS of_node_access_logs
|
||||
(
|
||||
id UInt64,
|
||||
node_id String,
|
||||
logged_at DateTime64(3, 'UTC'),
|
||||
remote_addr String,
|
||||
region String, -- Server GeoIP writes, Agent doesn't send
|
||||
host String,
|
||||
path String,
|
||||
user_agent String DEFAULT '', -- $http_user_agent
|
||||
cache_status String DEFAULT '', -- $upstream_cache_status
|
||||
status_code Int32,
|
||||
bytes_sent UInt64, -- data provided (body)
|
||||
request_length UInt64 DEFAULT 0, -- data received
|
||||
request_time_ms UInt32 DEFAULT 0, -- optional
|
||||
created_at DateTime64(3, 'UTC')
|
||||
)
|
||||
ENGINE = MergeTree()
|
||||
PARTITION BY toYYYYMM(logged_at)
|
||||
ORDER BY (node_id, logged_at, host, status_code, remote_addr)
|
||||
TTL toDateTime(logged_at) + INTERVAL 90 DAY
|
||||
SETTINGS index_granularity = 8192;
|
||||
```
|
||||
|
||||
| Column | Type | Source |
|
||||
| --- | --- | --- |
|
||||
| `id` | UInt64 | Server |
|
||||
| `node_id` | String | auth/payload |
|
||||
| `logged_at` | DateTime64(3) | `logged_at_unix` |
|
||||
| `remote_addr` | String | report |
|
||||
| `region` | String | Server GeoIP |
|
||||
| `host` | String | report |
|
||||
| `path` | String | report |
|
||||
| `user_agent` | String | report (nullable) |
|
||||
| `cache_status` | String | report (nullable) → **cache status** |
|
||||
| `status_code` | Int32 | report |
|
||||
| `bytes_sent` | UInt64 | report → **data provided** |
|
||||
| `request_length` | UInt64 | report → **data received** |
|
||||
| `request_time_ms` | UInt32 | report optional |
|
||||
| `created_at` | DateTime64(3) | Server now |
|
||||
|
||||
**Migration:** the current table already has `bytes_sent` / `request_length` / `request_time_ms` / `user_agent`; cache status adds:
|
||||
|
||||
```sql
|
||||
ALTER TABLE of_node_access_logs
|
||||
ADD COLUMN IF NOT EXISTS cache_status String DEFAULT '';
|
||||
```
|
||||
|
||||
### 5.2 L1 Hourly Rollup (Server-side MV)
|
||||
|
||||
**Agent forbidden to write.** Serves dashboard/node 24h fast queries of request count, error count, bytes.
|
||||
|
||||
**Implemented choice: `SummingMergeTree` + no UV column.**
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS of_access_log_hourly
|
||||
(
|
||||
node_id String,
|
||||
hour DateTime('UTC'),
|
||||
host String,
|
||||
request_count UInt64,
|
||||
error_count UInt64,
|
||||
bytes_sent UInt64,
|
||||
request_length UInt64
|
||||
)
|
||||
ENGINE = SummingMergeTree()
|
||||
PARTITION BY toYYYYMM(hour)
|
||||
ORDER BY (node_id, hour, host)
|
||||
TTL hour + INTERVAL 90 DAY;
|
||||
|
||||
CREATE MATERIALIZED VIEW IF NOT EXISTS of_access_log_hourly_mv
|
||||
TO of_access_log_hourly
|
||||
AS
|
||||
SELECT
|
||||
node_id,
|
||||
toStartOfHour(logged_at) AS hour,
|
||||
host,
|
||||
toUInt64(count()) AS request_count,
|
||||
toUInt64(countIf(status_code >= 500)) AS error_count,
|
||||
sum(bytes_sent) AS bytes_sent,
|
||||
sum(request_length) AS request_length
|
||||
FROM of_node_access_logs
|
||||
GROUP BY node_id, hour, host;
|
||||
```
|
||||
|
||||
Historical hours (details stored before the MV existed) need a one-time backfill, see migration `202607180003_backfill_access_log_hourly.sql` (ANTI JOIN to prevent duplicates).
|
||||
|
||||
#### UV Policy (must follow)
|
||||
|
||||
| Scenario | Data Source | Algorithm | Notes |
|
||||
| --- | --- | --- | --- |
|
||||
| **Window total UV** (dashboard totals, node cards, Zone totals) | `of_node_access_logs` details | `uniqExact(remote_addr)` (`TrafficSummary` / node aggregation) | **single authority**; never sum hourly UV |
|
||||
| **24h trend line request/error/bytes** | `of_access_log_hourly` preferred, fall back to detail buckets | `sum(request_count)` etc. | hourly path **doesn't fill** `unique_visitor_count` (always 0) |
|
||||
| **24h trend per-hour UV** | detail bucket path only | in-bucket `uniqExact` | when using hourly, UI should show empty/0 or hide the UV series; **forbidden** to `sum(UV)` over hourly rows |
|
||||
|
||||
**Why hourly doesn't store UV:**
|
||||
|
||||
1. `SummingMergeTree` can only safely merge addable counts; `uniqExact` across parts needs `AggregatingMergeTree` + state, heavier to implement and query.
|
||||
2. Even if hourly UV were stored, **summing over multi-hour windows severely overestimates** (the same IP is counted once per hour).
|
||||
3. Product "24h unique visitors" only recognizes whole-window `uniqExact`; the trend chart's main series are requests/errors/bytes — per-hour UV is not a primary metric.
|
||||
|
||||
### 5.3 L3 Fact Table: `of_node_metric_snapshots` (kept, semantics clarified)
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS of_node_metric_snapshots
|
||||
(
|
||||
id UInt64,
|
||||
node_id String,
|
||||
captured_at DateTime64(3, 'UTC'),
|
||||
cpu_usage_percent Float64,
|
||||
memory_used_bytes Int64,
|
||||
memory_total_bytes Int64,
|
||||
storage_used_bytes Int64,
|
||||
storage_total_bytes Int64,
|
||||
disk_read_bytes Int64, -- cumulative raw
|
||||
disk_write_bytes Int64,
|
||||
network_rx_bytes Int64, -- cumulative raw → host NIC inbound
|
||||
network_tx_bytes Int64, -- cumulative raw → host NIC outbound
|
||||
created_at DateTime64(3, 'UTC')
|
||||
)
|
||||
ENGINE = MergeTree()
|
||||
PARTITION BY toYYYYMM(captured_at)
|
||||
ORDER BY (node_id, captured_at, id)
|
||||
TTL toDateTime(captured_at) + INTERVAL 30 DAY
|
||||
SETTINGS index_granularity = 8192;
|
||||
```
|
||||
|
||||
Columns match production; **docs and API must label `network_*` as host NIC cumulative values**.
|
||||
|
||||
### 5.4 L3 Hourly Rollup: `of_node_metric_capacity_hourly` (kept)
|
||||
|
||||
Existing min/max used for cumulative-counter hourly increment approximation + CPU/memory averages. Logic unchanged:
|
||||
|
||||
* `network_tx_max - network_tx_min` ≈ that hour's host outbound
|
||||
* **must not** be used for "data provided"
|
||||
|
||||
### 5.5 L2 Fact Table: `of_node_edge_health` (new, replaces throughput-style openresty table)
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS of_node_edge_health
|
||||
(
|
||||
id UInt64,
|
||||
node_id String,
|
||||
captured_at DateTime64(3, 'UTC'),
|
||||
status LowCardinality(String), -- healthy / unhealthy / unknown
|
||||
connections Int64,
|
||||
created_at DateTime64(3, 'UTC')
|
||||
)
|
||||
ENGINE = MergeTree()
|
||||
PARTITION BY toYYYYMM(captured_at)
|
||||
ORDER BY (node_id, captured_at, id)
|
||||
TTL toDateTime(captured_at) + INTERVAL 30 DAY
|
||||
SETTINGS index_granularity = 8192;
|
||||
```
|
||||
|
||||
| Column | Description |
|
||||
| --- | --- |
|
||||
| `status` | instant health (same source as PG current state; for time series, not the sole UI authority) |
|
||||
| `connections` | current connection count |
|
||||
|
||||
**No** `message` column (description only in PG latest state).
|
||||
**No** `openresty_rx_bytes` / `openresty_tx_bytes`.
|
||||
|
||||
### 5.6 Relational DB (node latest state, not an analytics lake)
|
||||
|
||||
Separate from the observability lake, keeping "latest one":
|
||||
|
||||
| Table (logical name) | Purpose | Key Columns |
|
||||
| --- | --- | --- |
|
||||
| `of_nodes` (or current node table) | online, version, IP | `last_seen_at`, `openresty_status`, `openresty_message`, `agent_version` |
|
||||
| `of_node_system_profiles` | profile upsert | hostname, cpu_cores, total_memory_bytes, ... |
|
||||
| health-event table | `health_events` | event_type, severity, message, triggered_at |
|
||||
|
||||
> Actual physical table names follow the repo's existing GORM models; this design doesn't force renames, only forces **business throughput no longer written into node tables**.
|
||||
|
||||
### 5.7 Deprecated Tables (stop writing → delete after TTL)
|
||||
|
||||
| Table | Reason | Replacement |
|
||||
| --- | --- | --- |
|
||||
| `of_node_request_reports` | Agent pre-aggregation | `of_node_access_logs` + hourly |
|
||||
| `of_node_traffic_hourly` + MV | depends on request_reports | `of_access_log_hourly` |
|
||||
| `of_node_obs_openresty` | contains business rx/tx | `of_node_edge_health` |
|
||||
| `of_node_openresty_hourly` + MV | business throughput deltas | `of_access_log_hourly` bytes_* |
|
||||
|
||||
Relay-specific `of_node_obs_frps` / `of_node_obs_frpc` **kept** (not this Agent's main path, but same CH observability).
|
||||
|
||||
---
|
||||
|
||||
## 6. Table-Protocol Cross-Reference
|
||||
|
||||
| Product Concept | Protocol Field | Table.Column | Aggregation |
|
||||
| --- | --- | --- | --- |
|
||||
| Data provided | `access_logs[].bytes_sent` | `of_node_access_logs.bytes_sent` | `sum` |
|
||||
| Data received | `access_logs[].request_length` | `...request_length` | `sum` |
|
||||
| Request count | row count | — | `count` |
|
||||
| UV (window total) | `remote_addr` | same details | `uniqExact` (**forbidden** to sum hourly UV) |
|
||||
| Top domains | `host` | same | `group by` |
|
||||
| Status distribution | `status_code` | same | `group by` |
|
||||
| Source region | — | `region` (Server) | `group by` |
|
||||
| Host NIC outbound | `host_metrics.network_tx_bytes` | `of_node_metric_snapshots.network_tx_bytes` | time-series delta |
|
||||
| Host NIC inbound | `network_rx_bytes` | same | delta |
|
||||
| Disk read/write | `disk_*_bytes` | same | delta |
|
||||
| CPU/memory | instant fields | same | avg |
|
||||
| OpenResty connections | `edge_health.connections` | `of_node_edge_health.connections` | latest/avg |
|
||||
| OpenResty health | `edge_health.status` | node table + optional CH | latest |
|
||||
|
||||
**Mappings that no longer exist:**
|
||||
|
||||
| Old Concept | Old Field | Disposition |
|
||||
| --- | --- | --- |
|
||||
| OpenResty outbound | `openresty_tx_bytes` | removed; use data provided |
|
||||
| OpenResty inbound | `openresty_rx_bytes` | removed; use data received |
|
||||
| Window request report | `traffic_report` | removed |
|
||||
|
||||
---
|
||||
|
||||
## 7. OpenResty Log Format (aligned with details)
|
||||
|
||||
Target `log_format` (ensures the `bytes_sent` key = body; includes UA and cache status):
|
||||
|
||||
```nginx
|
||||
log_format openflare_json escape=json
|
||||
'{"ts":"$time_iso8601","host":"$host","path":"$request_uri",'
|
||||
'"remote_addr":"$remote_addr","status":$status,'
|
||||
'"request_time":$request_time,'
|
||||
'"bytes_sent":$body_bytes_sent,"request_length":$request_length,'
|
||||
'"user_agent":"$http_user_agent",'
|
||||
'"cache_status":"$upstream_cache_status"}';
|
||||
```
|
||||
|
||||
Agent parsing:
|
||||
|
||||
* `ts` → `logged_at_unix`
|
||||
* `bytes_sent` → protocol `bytes_sent` (provided)
|
||||
* `request_length` → protocol `request_length`
|
||||
* `request_time` → optional `request_time_ms = round(sec * 1000)`
|
||||
* `user_agent` → protocol `user_agent`
|
||||
* `cache_status` → protocol `cache_status` (passed through as-is, no three-state compression)
|
||||
|
||||
---
|
||||
|
||||
## 8. Upgrade Strategy (no compatibility layer)
|
||||
|
||||
| Item | Strategy |
|
||||
| --- | --- |
|
||||
| Agent upgrade | **destroy-and-recreate** preferred; **binary replacement** allowed |
|
||||
| Protocol | only schema v2 fields; legacy JSON fields not parsed |
|
||||
| Local observability buffer | if still in old format (containing `snapshot` / `openresty_observation` / `traffic_report`) or corrupt → **delete the file wholesale**, rebuild at runtime |
|
||||
| Read path | business APIs **only read** access_logs (and hourly); current health reads PG; connection series reads CH edge_health |
|
||||
| Old Agents | must upgrade; the control plane provides no v1 dual-read path |
|
||||
|
||||
---
|
||||
|
||||
## 9. Example: Storage Result of One Heartbeat
|
||||
|
||||
**Agent report (excerpt):**
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n1",
|
||||
"host_metrics": {
|
||||
"captured_at_unix": 1720000000,
|
||||
"cpu_usage_percent": 10,
|
||||
"memory_used_bytes": 1,
|
||||
"memory_total_bytes": 2,
|
||||
"storage_used_bytes": 3,
|
||||
"storage_total_bytes": 4,
|
||||
"disk_read_bytes": 100,
|
||||
"disk_write_bytes": 200,
|
||||
"network_rx_bytes": 1000,
|
||||
"network_tx_bytes": 2000
|
||||
},
|
||||
"edge_health": {
|
||||
"captured_at_unix": 1720000000,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 5
|
||||
},
|
||||
"access_logs": [
|
||||
{
|
||||
"logged_at_unix": 1720000001,
|
||||
"remote_addr": "1.1.1.1",
|
||||
"host": "a.example.com",
|
||||
"path": "/",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 500,
|
||||
"request_length": 80
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Written:**
|
||||
|
||||
1. PG node latest state: `openresty_status` / `openresty_message` (if reported)
|
||||
2. `of_node_metric_snapshots` 1 row (network_tx=2000 cumulative)
|
||||
3. `of_node_edge_health` 1 row (status + connections=5; **no message**)
|
||||
4. `of_node_access_logs` 1 row (bytes_sent=500, request_length=80, region filled by Server)
|
||||
5. MV asynchronously counts into `of_access_log_hourly`
|
||||
|
||||
**Query 24h data provided:** `sum(bytes_sent)` → at least 500 (plus history)
|
||||
**Query host outbound:** delta over snapshots; **no forced equality** with 500.
|
||||
|
||||
---
|
||||
|
||||
## 10. Revision History
|
||||
|
||||
| Date | Notes |
|
||||
| --- | --- |
|
||||
| 2026-07-17 | initial draft: protocol v2, Server storage pipeline, CH/relational target table structures and deprecated table list |
|
||||
@@ -0,0 +1,552 @@
|
||||
# Edge Observability and Business Traffic Statistics Refactor
|
||||
|
||||
You will learn: the problems this refactor solves (dashboard "OpenResty outbound" vs "Zone data provided" inconsistency, field and aggregation redundancy), and how the target architecture makes **the Agent report only facts and the Server interpret facts**, with access logs as the single source of truth for business traffic.
|
||||
|
||||
---
|
||||
|
||||
## 1. Goals
|
||||
|
||||
### 1.1 Problems to Solve
|
||||
|
||||
1. **Dual sources of truth**: business throughput comes from both access-log aggregation and OpenResty observability deltas, and the numbers never match long-term.
|
||||
2. **Agent over-computes**: the edge pre-aggregates `TrafficReport`, throughput accumulation, and the control plane aggregates again — semantics are hard to evolve and reconcile.
|
||||
3. **Field semantic overlap**: "OpenResty outbound" and "data provided" are the same business problem for users, but the system uses two field sets and two pipelines.
|
||||
4. **Instant vs cumulative mixed**: 60-second window counts are treated as process cumulative values for 24h deltas, causing severe underestimation.
|
||||
5. **UI induces wrong comparisons**: the dashboard and Zone page use similar "traffic/data" wording without declaring scope and caliber differences.
|
||||
|
||||
### 1.2 Refactor Goals
|
||||
|
||||
| Goal | Description |
|
||||
| --- | --- |
|
||||
| **Single business truth** | request count, data provided, UV, status distribution, Top domains etc. **only** derived from access logs (and Server-side rollups) |
|
||||
| **Agent reports only facts** | detail logs + machine readings + health snapshots; **business UV/TopN/24h totals pre-aggregation is forbidden** |
|
||||
| **Field convergence** | one business concept maps to one authoritative field; machine NIC and business delivery strictly separated by name |
|
||||
| **Reconcilable** | global "data provided" ≈ sum of per-Zone "data provided" (difference only from unbound/unknown Hosts) |
|
||||
| **Evolvable** | changing time windows, TopN, ownership rules only changes the Server, not the Agent |
|
||||
|
||||
### 1.3 Non-Goals (outside this design)
|
||||
|
||||
* Building a general log platform, full-log long-term archive, or search product.
|
||||
* Replacing ClickHouse / removing the analytics DB dependency.
|
||||
* Reworking Relay / OpenFlared host metric collection (principles align, but not in this round's protocol main path).
|
||||
* Real-time streaming alert engine, APM tracing (the OpenTelemetry server side already exists and is orthogonal to this business traffic model).
|
||||
|
||||
---
|
||||
|
||||
## 2. Scope and Constraints
|
||||
|
||||
### 2.1 Product Constraints (inherited)
|
||||
|
||||
* Single-tenant, single globally active config; observability introduces no multi-tenant billing isolation.
|
||||
* Access logs and time-series observability use the switchable log primary DB (ClickHouse by default; switchable to PostgreSQL/SQLite), see [Log Store Decoupling](./logstore.md).
|
||||
* Agent has no inbound control, Pull model; during offline periods local OpenResty keeps serving, and observability can buffer locally and backfill.
|
||||
|
||||
### 2.2 Engineering Constraints
|
||||
|
||||
* Agent stays lightweight: parse log lines, read `/proc`, health checks; no business analysis.
|
||||
* Control-plane API errors still use the unified envelope and `response.Abort*`.
|
||||
* Access log field changes must update both the OpenResty `log_format` and the Agent parser simultaneously; Agent and control plane ship at the same version, no legacy protocol parsing.
|
||||
|
||||
---
|
||||
|
||||
## 3. Design Principles
|
||||
|
||||
### Principle P1: Agent Reports Facts, Server Interprets Facts
|
||||
|
||||
```text
|
||||
Agent = collection + reliable delivery (raw / near-raw)
|
||||
Server = storage + aggregation + ownership + trends + reconciliation
|
||||
```
|
||||
|
||||
**Allowed edge processing (collection)**
|
||||
|
||||
* Parsing JSON access.log lines into structured fields
|
||||
* path length caps, dropping invalid lines, skipping observability-port's own requests
|
||||
* Reading NIC/CPU/memory counters as **raw values**
|
||||
* Batching, compression, offline buffering and retries
|
||||
|
||||
**Forbidden edge processing (business computation)**
|
||||
|
||||
* UV / Top domains / status histograms / window request_count as authoritative metrics
|
||||
* Maintaining "business in/out cumulative" for the dashboard
|
||||
* Zone / domain ownership stats, country distribution (country can be resolved at Server insert time)
|
||||
|
||||
### Principle P2: Single Truth for Business Traffic = Access Logs
|
||||
|
||||
| Business Question | Single Answer |
|
||||
| --- | --- |
|
||||
| How much data was provided | `sum(bytes_sent)` |
|
||||
| How many requests | `count()` |
|
||||
| How many unique visitors | `uniqExact(remote_addr)` (or product-defined hashing) |
|
||||
| Status codes / Top domains | `group by` on logs |
|
||||
|
||||
### Principle P3: Three Metric Layers Never Mixed
|
||||
|
||||
| Layer | Name | Purpose | Typical Fields |
|
||||
| --- | --- | --- | --- |
|
||||
| L1 Business delivery | Business Traffic | user & Zone reconciliation, dashboard business trends | access log |
|
||||
| L2 Edge health | Edge Health | is OpenResty alive, current connections | status, connections |
|
||||
| L3 Host capacity | Host Capacity | capacity planning, is the machine saturated | CPU, memory, disk, **NIC** |
|
||||
|
||||
Never name L3 NIC or L2 instant counts as "data provided"; never draw L1 and L3 on the same summary card without labeling semantics.
|
||||
|
||||
### Principle P4: One Business Concept, One Field
|
||||
|
||||
* **Data provided** ≡ response body delivered ≡ "OpenResty outbound (business meaning)" in legacy copy → **keep only `bytes_sent` aggregation**
|
||||
* **Data received** (optional) ≡ request-side volume → log `request_length` aggregation
|
||||
* **Host outbound** ≡ `network_tx` delta, copy must include "host/NIC"
|
||||
|
||||
---
|
||||
|
||||
## 4. Pre-Refactor Problems (Baseline)
|
||||
|
||||
### 4.1 Pre-Refactor Data Flow (redundant)
|
||||
|
||||
```text
|
||||
One HTTP request
|
||||
│
|
||||
├─ access.log line
|
||||
│ → Agent tail → AccessLogs[]
|
||||
│ → CH of_node_access_logs
|
||||
│ → Zone "data provided" ✅
|
||||
│
|
||||
├─ Lua shared dict window/cumulative counts
|
||||
│ → /openflare/observability
|
||||
│ → TrafficReport + OpenrestyObservation(rx/tx)
|
||||
│ → CH request_reports / obs_openresty
|
||||
│ → dashboard "OpenResty in/outbound" ❌ easily inconsistent with Zone
|
||||
│
|
||||
├─ second access.log aggregation (fallback when observability endpoint fails)
|
||||
│ → yet another TrafficReport / throughput
|
||||
│
|
||||
└─ host network_rx/tx
|
||||
→ Snapshot → "host" curve in network trends
|
||||
```
|
||||
|
||||
### 4.2 Field Overlap
|
||||
|
||||
| User Perception | System Field A | System Field B | Problem |
|
||||
| --- | --- | --- | --- |
|
||||
| Outbound / provided | `openresty_tx_bytes` | `bytes_sent` | duplicate business semantics |
|
||||
| Inbound | `openresty_rx_bytes` | `request_length` (log) | duplicate business semantics |
|
||||
| Request count | `TrafficReport.request_count` | `count(access_logs)` | duplicate aggregation, window easily double-counted |
|
||||
| Outbound (machine) | `network_tx_bytes` | (no business equivalent) | should be named separately, never reconciled with business |
|
||||
|
||||
### 4.3 Typical Failure Modes
|
||||
|
||||
1. Window counts treated as cumulative deltas → 24h business throughput severely underestimated.
|
||||
2. Hourly rollup `max−min` broken for resetting counters.
|
||||
3. Zone uses logs, dashboard uses observability → users think the system is wrong.
|
||||
4. Changing caliber requires syncing Lua, Agent state accumulation, Server deltas, and frontend copy.
|
||||
|
||||
---
|
||||
|
||||
## 5. Target Architecture
|
||||
|
||||
### 5.1 Target Data Flow
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph edge [Edge Node]
|
||||
OR[OpenResty]
|
||||
LOG[access.log]
|
||||
PROC[host /proc and disk]
|
||||
STUB[stub_status connections]
|
||||
AG[Agent]
|
||||
OR -->|log_format writes line| LOG
|
||||
LOG -->|tail incremental details only| AG
|
||||
PROC -->|reading snapshots| AG
|
||||
STUB -->|instant connections| AG
|
||||
OR -->|health probe| AG
|
||||
end
|
||||
|
||||
subgraph server [Control-Plane Server]
|
||||
HB[Heartbeat / WS receive]
|
||||
CH[(ClickHouse)]
|
||||
AGG[Aggregation query layer]
|
||||
API[Admin API]
|
||||
HB --> CH
|
||||
CH --> AGG
|
||||
AGG --> API
|
||||
end
|
||||
|
||||
subgraph ui [Admin Panel]
|
||||
DASH[Dashboard: global business trends]
|
||||
ZONE[Zone: filter by domain]
|
||||
NODE[Node: host resources + health]
|
||||
end
|
||||
|
||||
AG -->|AccessLogs + HostSnapshot + Health| HB
|
||||
API --> DASH
|
||||
API --> ZONE
|
||||
API --> NODE
|
||||
```
|
||||
|
||||
### 5.2 Responsibility Matrix
|
||||
|
||||
| Capability | Agent | Server | Frontend |
|
||||
| --- | --- | --- | --- |
|
||||
| Write access.log | OpenResty | — | — |
|
||||
| Read and report details | ✅ | store | — |
|
||||
| sum/count/uniq/TopN | ❌ | ✅ | display |
|
||||
| Zone domain filtering | ❌ | ✅ | select Zone |
|
||||
| Host CPU/memory/NIC | read raw values and report | delta/average | node/dashboard resource area |
|
||||
| OpenResty connections | read instant and report | latest value | node health |
|
||||
| Business 24h in/outbound | ❌ | log aggregation | uniformly called "data provided/received" |
|
||||
|
||||
---
|
||||
|
||||
## 6. Metrics and Field Model
|
||||
|
||||
### 6.1 Authoritative Field Table (target)
|
||||
|
||||
#### L1 Business Delivery (from access logs)
|
||||
|
||||
| Concept | Storage Field | Aggregation | Display Name |
|
||||
| --- | --- | --- | --- |
|
||||
| Request time | `logged_at` | window filter | — |
|
||||
| Node | `node_id` | group | — |
|
||||
| Client IP | `remote_addr` | `uniq` → UV | Unique visitors |
|
||||
| Host | `host` | group / Zone mapping | Domain |
|
||||
| Path | `path` | optional | — |
|
||||
| Status code | `status_code` | group | Status distribution |
|
||||
| **Data provided** | **`bytes_sent`** | **`sum`** | **Data provided** |
|
||||
| **Data received** | **`request_length`** | **`sum`** | **Data received** (optional display) |
|
||||
| Region | `region` (resolved & written by Server) | group | Source region |
|
||||
|
||||
> Note: the JSON key in the OpenResty `log_format` may keep the name `bytes_sent`; the value must come from **`$body_bytes_sent`** (consistent with production), representing response body delivered, i.e. "data provided".
|
||||
|
||||
#### L2 Edge Health (instant; no 24h business totals)
|
||||
|
||||
| Concept | Field | Description |
|
||||
| --- | --- | --- |
|
||||
| OpenResty health | `openresty_status` / message | existing |
|
||||
| Current connections | `openresty_connections` | stub_status |
|
||||
| (optional) rough recent-window QPS | node detail "right now" only, **never** authoritative 24h totals | if implemented must be labeled "instant" |
|
||||
|
||||
#### L3 Host Capacity
|
||||
|
||||
| Concept | Field | Display Name |
|
||||
| --- | --- | --- |
|
||||
| CPU / memory / disk usage | `host_metrics` | keep |
|
||||
| NIC cumulative bytes | `network_rx_bytes` / `network_tx_bytes` | **Host NIC in/outbound** |
|
||||
| Disk IO cumulative | `disk_read_bytes` / `disk_write_bytes` | Disk read/write |
|
||||
|
||||
### 6.2 Removed Fields (no compatibility layer)
|
||||
|
||||
| Original Field | Disposition | Reason |
|
||||
| --- | --- | --- |
|
||||
| `openresty_tx_bytes` / `openresty_rx_bytes` | **removed** | business bytes follow access logs |
|
||||
| `TrafficReport` and TopN/window UV | **removed** | edge pre-aggregation |
|
||||
| Agent state business lifetime accumulators | removed | violates P1 |
|
||||
| Lua shared dict business throughput/window request counts | removed | not the delivery main path |
|
||||
|
||||
### 6.3 Naming Reference (frontend copy enforced)
|
||||
|
||||
| Forbidden Copy | Correct Copy | Data Source |
|
||||
| --- | --- | --- |
|
||||
| OpenResty outbound (business volume) | **Data provided** | `sum(bytes_sent)` |
|
||||
| OpenResty inbound (business volume) | **Data received** | `sum(request_length)` |
|
||||
| Network outbound (unspecified) | **Host NIC outbound** | `network_tx` delta |
|
||||
| Two cards: data provided vs outbound | **keep only one business card** | logs |
|
||||
|
||||
---
|
||||
|
||||
## 7. Agent Design
|
||||
|
||||
### 7.1 Heartbeat Payload (target protocol)
|
||||
|
||||
Keep and strengthen:
|
||||
|
||||
```text
|
||||
NodePayload
|
||||
identity / version / openresty_status / openresty_message # latest state → PG
|
||||
profile # host overview (low frequency)
|
||||
host_metrics # L3 resource readings (incl. NIC cumulative raw values)
|
||||
edge_health # L2: status + connections (CH time series; message not in CH)
|
||||
access_logs[] # L1 details (main path)
|
||||
health_events[]
|
||||
buffered[] # buffered facts above, not reports
|
||||
waf_ip_group_checksums
|
||||
```
|
||||
|
||||
Removed from the protocol (no compatibility layer):
|
||||
|
||||
```text
|
||||
traffic_report
|
||||
openresty_observation
|
||||
snapshot / buffered_observability aliases
|
||||
```
|
||||
|
||||
### 7.2 Access Log Reporting Requirements
|
||||
|
||||
Each detail at minimum contains:
|
||||
|
||||
| Field | Required | Note |
|
||||
| --- | --- | --- |
|
||||
| `logged_at_unix` | ✅ | request completion time |
|
||||
| `remote_addr` | ✅ | UV |
|
||||
| `host` | ✅ | Zone mapping |
|
||||
| `path` | ✅ | may be truncated |
|
||||
| `status_code` | ✅ | |
|
||||
| `bytes_sent` | ✅ | body bytes, data provided |
|
||||
| `request_length` | ✅ | data received |
|
||||
|
||||
Agent responsibilities:
|
||||
|
||||
1. Tail `access.log` by offset (reset offset on truncation/rotation, **only report new lines still present in the file**).
|
||||
2. Parse into structured form, batch into heartbeat / WS.
|
||||
3. Offline writes to local buffer, backfill by window once connected.
|
||||
4. **No sum/count/uniq on details.**
|
||||
|
||||
### 7.3 Host Snapshot
|
||||
|
||||
* Keep reporting NIC/disk **cumulative counter raw values** (not business pre-aggregation).
|
||||
* Server does non-negative deltas between adjacent samples → host trends.
|
||||
* This is unrelated to "data provided"; the UI must display it in a separate section.
|
||||
|
||||
### 7.4 OpenResty Local Observability
|
||||
|
||||
Converged state:
|
||||
|
||||
* Keep: health checks, `stub_status` current connections.
|
||||
* The main path no longer relies on `log.lua` shared dict business counts; `/openflare/observability` only returns health and connection snapshots, not business report sources.
|
||||
|
||||
### 7.5 Relationship with the Agent Design Doc
|
||||
|
||||
This design strengthens "pure data landing" in [Agent & Publish Model](./agent-design.md):
|
||||
|
||||
* Config and certificates: land and report applied state.
|
||||
* Observability: only carry facts, not business conclusions.
|
||||
|
||||
---
|
||||
|
||||
## 8. Server Design
|
||||
|
||||
### 8.1 Storage
|
||||
|
||||
| Input | Table | Description |
|
||||
| --- | --- | --- |
|
||||
| `access_logs[]` | `of_node_access_logs` | authoritative business details |
|
||||
| `host_metrics` | `of_node_metric_snapshots` | L3; NIC/disk cumulative |
|
||||
| `openresty_status` / `openresty_message` | **PG node table** | L2 **latest-state authority** (message only here) |
|
||||
| `edge_health` | `of_node_edge_health` | L2 time series: status + connections (**no message**) |
|
||||
|
||||
GeoIP: continue resolving `remote_addr` → `region` in the Server insert path, not in the Agent.
|
||||
|
||||
### 8.2 Aggregation Layer (unified)
|
||||
|
||||
All business trends and Zone stats share the same query semantics:
|
||||
|
||||
```text
|
||||
filter: logged_at ∈ [since, until]
|
||||
optional: node_id / host IN (...)
|
||||
metrics:
|
||||
request_count = count()
|
||||
unique_visitors = uniqExact(remote_addr)
|
||||
bytes_provided = sum(bytes_sent) -- data provided
|
||||
bytes_received = sum(request_length) -- data received
|
||||
series folded by hour/bucket
|
||||
distributions by status_code / host / region
|
||||
```
|
||||
|
||||
Implementation locations:
|
||||
|
||||
* Zone: `GET .../zones/:id/stats` (existing, align field naming)
|
||||
* Dashboard: overview traffic / business network trends **switch to the same aggregation** (global, no host filter or Top filter)
|
||||
* Node detail: business volume = the same aggregation filtered by that `node_id`; host NIC still uses metric deltas
|
||||
|
||||
### 8.3 Derived Rollups (optional performance path)
|
||||
|
||||
When detail queries over all nodes for 24h are too heavy, allow **Server-side** materialized views:
|
||||
|
||||
```text
|
||||
of_access_log_hourly
|
||||
(hour, node_id, host, request_count, bytes_sent, bytes_received, ...)
|
||||
```
|
||||
|
||||
Constraints:
|
||||
|
||||
* Derived only by CH from `of_node_access_logs`; **Agent is forbidden from writing this table directly**.
|
||||
* Zone / dashboard prefer reading the rollup, falling back to details (similar to the existing metric hourly policy).
|
||||
|
||||
### 8.4 Decommissioned Analytics Paths
|
||||
|
||||
| Path | After Migration |
|
||||
| --- | --- |
|
||||
| `BuildNetworkTrendPoints` delta on openresty_rx/tx | deleted, or keep only `network_*` host curves |
|
||||
| `of_node_obs_openresty` throughput fields | stop writing; drop table or shrink columns after TTL expiry |
|
||||
| `of_node_request_reports` + traffic hourly | business trends no longer depend on it; table can be deprecated wholesale |
|
||||
| Dashboard compact openresty_tx series | change to bytes_provided series |
|
||||
|
||||
---
|
||||
|
||||
## 9. API and Frontend
|
||||
|
||||
### 9.1 Semantically Unified Response Fields
|
||||
|
||||
Business stats APIs should uniformly use:
|
||||
|
||||
```json
|
||||
{
|
||||
"request_count": 0,
|
||||
"unique_visitors": 0,
|
||||
"bytes_provided": 0,
|
||||
"bytes_received": 0,
|
||||
"series": [
|
||||
{
|
||||
"bucket_started_at": "...",
|
||||
"request_count": 0,
|
||||
"unique_visitors": 0,
|
||||
"bytes_provided": 0,
|
||||
"bytes_received": 0
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
API business byte fields use `bytes_provided` / `bytes_received` (access-log aggregation); no more openresty throughput aliases.
|
||||
|
||||
### 9.2 Dashboard
|
||||
|
||||
* **Business area**: request trend, data provided, data received (optional), status codes, Top domains, source regions — all L1.
|
||||
* **Resource area**: CPU/memory, **host NIC**, disk IO — all L3.
|
||||
* **Forbidden**: showing "OpenResty in/outbound" in the business area as a metric reconciled with Zone.
|
||||
|
||||
Suggest splitting or retitling "24-hour network and disk trends":
|
||||
|
||||
* "24-hour business traffic" → `bytes_provided` / `bytes_received` / requests
|
||||
* "24-hour host network and disk" → `network_*` / `disk_*`
|
||||
|
||||
### 9.3 Zone `/websites/:id`
|
||||
|
||||
* Keep cards like "total data provided".
|
||||
* Data and the dashboard business area use **the same aggregation function**, only `hosts = zone domain list`.
|
||||
* Docs and UI may note: the global dashboard includes all Hosts; this page is only this Zone.
|
||||
|
||||
### 9.4 Node Detail
|
||||
|
||||
* Business throughput: that node's `sum(bytes_sent)` etc.
|
||||
* OpenResty: health + current connections.
|
||||
* NIC: clearly "host".
|
||||
|
||||
---
|
||||
|
||||
## 10. OpenResty and Log Format
|
||||
|
||||
### 10.1 Keep
|
||||
|
||||
Existing JSON `log_format` core fields:
|
||||
|
||||
```text
|
||||
ts, host, path, remote_addr, status, request_time,
|
||||
bytes_sent (= $body_bytes_sent), request_length
|
||||
```
|
||||
|
||||
### 10.2 Changes
|
||||
|
||||
* No longer rely on log-phase writes of business shared dict counts as control-plane input.
|
||||
* Observability-port requests continue not writing business stats (or `access_log off`).
|
||||
|
||||
### 10.3 Agent Parsing
|
||||
|
||||
* Protocol `NodeAccessLog` adds `request_length`.
|
||||
* Legacy log lines missing fields default to 0, not blocking the whole batch.
|
||||
|
||||
---
|
||||
|
||||
## 11. Upgrade and Migration (no compatibility layer)
|
||||
|
||||
### 11.1 Phase Review (shipped)
|
||||
|
||||
| Phase | Content |
|
||||
| --- | --- |
|
||||
| **M1–M5** | read path switches to access logs; protocol v2; stop pre-aggregation; edge_health + access_log_hourly; drop old tables and API compat fields |
|
||||
|
||||
### 11.2 Upgrade Strategy
|
||||
|
||||
* **Agent: destroy-and-recreate preferred**; binary replacement allowed.
|
||||
* On binary replacement: the local old observability buffer (including `snapshot` / `openresty_observation` / `traffic_report`) is **deleted wholesale**, rebuilt after running.
|
||||
* Server **does not** parse v1 fields, **does not** dual-read request_reports / openresty throughput.
|
||||
* Detail-missing periods: business charts are empty or partial; **never** impersonate data provided with NIC or removed openresty throughput.
|
||||
|
||||
### 11.3 Data Backfill
|
||||
|
||||
* Historical "data provided" follows access logs.
|
||||
* Before `of_access_log_hourly` is created, history is backfilled with goose SQL (ANTI JOIN to prevent duplicates).
|
||||
|
||||
### 11.4 Health-State Authority
|
||||
|
||||
* **Current state**: PG `openresty_status` / `openresty_message`.
|
||||
* **Time series**: log primary DB `of_node_edge_health` (status + connections; no message).
|
||||
|
||||
### 11.5 UV
|
||||
|
||||
* **Whole-window unique visitors**: `uniqExact(remote_addr)` (dashboard totals, Zone totals).
|
||||
* **Bucketed UV** (Zone curves): per-bucket uniq, **not summable across buckets**; UI must note it.
|
||||
* **Hourly trend path**: don't plot / fill per-hour UV (hourly table has no UV).
|
||||
|
||||
---
|
||||
|
||||
## 12. Storage and Capacity
|
||||
|
||||
* Business trends rely on details or hourly rollups; watch `of_node_access_logs` TTL and sampling.
|
||||
* If details are too large: prefer **Server-side rollup** rather than restoring Agent pre-aggregation.
|
||||
* For high-cardinality path scenarios, limit detail path length (existing); aggregation doesn't do global Top over full paths by default.
|
||||
|
||||
---
|
||||
|
||||
## 13. Risks and Trade-offs
|
||||
|
||||
| Risk | Mitigation |
|
||||
| --- | --- |
|
||||
| Large detail volume makes CH and heartbeat heavy | batching, compression, sampling policy evaluation; Server rollup; limit per-batch count |
|
||||
| Brief log loss lowers business volume | local buffer and rotation handling; monitor access log collection lag |
|
||||
| Users still compare "NIC outbound" with "data provided" | UI sections and copy enforce the "host" prefix |
|
||||
| Old Agents stay online long-term | **no compatibility layer**; Agents must be upgraded/rebuilt |
|
||||
|
||||
**Why not keep Agent pre-aggregation as an optimization?**
|
||||
|
||||
* Saving bandwidth re-splits the truth, drifts calibers, and repeats this problem.
|
||||
* Optimization belongs in Server derived tables and queries, not edge business computation.
|
||||
|
||||
---
|
||||
|
||||
## 14. Key Decision Summary
|
||||
|
||||
| Decision | Choice | Rejected Alternative |
|
||||
| --- | --- | --- |
|
||||
| Business traffic truth | access logs | OpenResty dict / TrafficReport |
|
||||
| Agent role | report only facts | edge UV/TopN/throughput accumulation |
|
||||
| "Outbound" vs "provided" | merged into data provided | dual fields and dual pipelines long-term |
|
||||
| NIC traffic | independent L3, separate copy | reconciled side-by-side with business outbound |
|
||||
| Performance | CH rollup | Agent pre-aggregation |
|
||||
| Migration | switch read path first, then slim Agent | drop details first, rely on pre-aggregation |
|
||||
|
||||
---
|
||||
|
||||
## 15. Doc and Code Mapping
|
||||
|
||||
| Area | Main Paths |
|
||||
| --- | --- |
|
||||
| Protocol | `pkg/protocol/agent.go` |
|
||||
| Agent collection | `internal/apps/agent/observability/`, `heartbeat/` |
|
||||
| OpenResty logging and Lua | `pkg/render/openresty/`, `internal/apps/agent/nginx/observability_assets.go` |
|
||||
| Server storage | `internal/apps/openflare/agent/observability.go` |
|
||||
| Log aggregation | `internal/repository/analytics/node_access_log*.go`, `internal/apps/openflare/zone/stats.go` |
|
||||
| Dashboard | `internal/apps/openflare/dashboard/`, `internal/apps/openflare/observability/analytics.go` |
|
||||
| Frontend | `frontend/app/(main)/page.tsx`, `components/dashboard/*`, `websites/.../zone-overview.tsx` |
|
||||
|
||||
**Recommended reading order:**
|
||||
|
||||
1. **[Observability Transport Model](./observability-transport-model.md)** (latest: what to send, where collected from, frequency, sample JSON)
|
||||
2. [Agent Reporting Protocol and Observability Data Model](./observability-data-model.md) (protocol fields and DDL)
|
||||
|
||||
---
|
||||
|
||||
## 16. Revision History
|
||||
|
||||
| Date | Notes |
|
||||
| --- | --- |
|
||||
| 2026-07-17 | initial draft: target architecture and migration phases for dual truth, Agent pre-aggregation, field redundancy |
|
||||
| 2026-07-17 | added protocol/table-structure chapter links `observability-data-model.md` |
|
||||
@@ -0,0 +1,502 @@
|
||||
# Edge Observability Transport Model (current target version)
|
||||
|
||||
> **This document is the latest authoritative description of "how Agent ↔ Server observability data is transmitted".**
|
||||
> After reading you should be able to answer: what is sent, where it is collected from, how often, how the Server stores it, and where product metrics are queried from.
|
||||
> Protocol fields and DDL details: [Observability Reporting Protocol & Data Model](./observability-data-model.md); background: [Edge Observability & Business Traffic Stats](./observability-design.md).
|
||||
|
||||
---
|
||||
|
||||
## 0. Remember the Three Layers First
|
||||
|
||||
| Layer | Question Answered | Single Data Source | Product Examples |
|
||||
| --- | --- | --- | --- |
|
||||
| **L1 Business delivery** | How much data provided? How many requests? | **access.log details** | data provided, request count, UV, status codes, Top domains |
|
||||
| **L2 Edge health** | Is OpenResty alive? Current connections? | **local `/openflare/observability`** | node health, current connections |
|
||||
| **L3 Host capacity** | How are CPU/memory/disk/NIC? | **OS readings** | capacity trends, host NIC |
|
||||
|
||||
**The three layers are never reconciled against each other.**
|
||||
"Data provided" ≠ "current connections" ≠ "host NIC outbound".
|
||||
|
||||
---
|
||||
|
||||
## 1. Overview: Who Collects, Who Reports, Who Aggregates
|
||||
|
||||
```text
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Edge Node │
|
||||
│ │
|
||||
│ Visitor request ──► OpenResty │
|
||||
│ │ │
|
||||
│ ├─ access.log (one line per request) ←── L1 collection point │
|
||||
│ │ │
|
||||
│ └─ connection state (maintained in-process) │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ GET /openflare/observability ←── L2 reads snapshot │
|
||||
│ (no log scanning, no business recomputation) │
|
||||
│ │
|
||||
│ OS /proc etc. ──────────────────────────── L3 reads snapshot │
|
||||
│ │
|
||||
│ ┌────────── Agent ──────────┐ │
|
||||
│ │ default: one NodePayload per 3s │ │
|
||||
│ │ · tail access.log incremental │ │
|
||||
│ │ · GET local observability │ │
|
||||
│ │ · read host_metrics │ │
|
||||
│ └────────────┬──────────────┘ │
|
||||
└─────────────────────────────│──────────────────────────────────┘
|
||||
│ HTTP heartbeat or WebSocket status
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Server (control plane) │
|
||||
│ · details → ClickHouse of_node_access_logs │
|
||||
│ · health → node latest state + of_node_edge_health │
|
||||
│ · host → of_node_metric_snapshots │
|
||||
│ · business trends / Zone stats = sum/count/uniq over access_logs only │
|
||||
└─────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
| Role | Does | Doesn't |
|
||||
| --- | --- | --- |
|
||||
| OpenResty | writes access.log; maintains connection counts | no direct reporting to the control plane |
|
||||
| Agent | **collects facts and reports them** | **does not compute** UV/TopN/24h data provided |
|
||||
| Server | stores + **aggregates/interpreters** | does not trust edge business pre-summaries |
|
||||
|
||||
---
|
||||
|
||||
## 2. Collection Frequency (defaults)
|
||||
|
||||
| Action | Default Frequency | Config |
|
||||
| --- | --- | --- |
|
||||
| Agent → Server report | **every 3 seconds** a full payload | `heartbeat_interval` / control-plane `agent_heartbeat_interval` (ms, default `3000`) |
|
||||
| Tail access.log when packing | **with report** (new lines since last report) | same |
|
||||
| GET `/openflare/observability` when packing | **with report** (reads **current** connection snapshot) | same |
|
||||
| Read host metrics when packing | **with report** | same |
|
||||
| OpenResty writes access.log | **1 line at each request end** | unrelated to heartbeat |
|
||||
| Connection counts update in-process | **on connection change** (kernel-maintained) | unrelated to heartbeat |
|
||||
| Offline replay window | keep ~**60 minutes** by default | `observability_replay_minutes` |
|
||||
| Node offline detection | ~**60s** without a successful heartbeat | `node_offline_threshold` (default `60000` ms) |
|
||||
|
||||
**Notes:**
|
||||
|
||||
- The Agent has **no separate "sampling clock"**; **sampling points = report points** (default 3s).
|
||||
- access.log is "per-request continuous writes"; the Agent only **moves incremental lines** periodically.
|
||||
- `/openflare/observability` is **not** "business stats start being counted when called"; for connections it **reads Nginx's existing instantaneous values**.
|
||||
|
||||
Transport channels:
|
||||
|
||||
- **HTTP heartbeat**: POST the full payload at the interval.
|
||||
- **WebSocket**: after connecting, sends `status` messages at the same interval (same content shape); HTTP heartbeat is not double-sent then.
|
||||
|
||||
---
|
||||
|
||||
## 3. Agent → Server Packet (NodePayload v2)
|
||||
|
||||
### 3.1 Structure Skeleton
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n_01hxyz",
|
||||
"name": "edge-shanghai-1",
|
||||
"ip": "203.0.113.10",
|
||||
"version": "3.4.0",
|
||||
"ext_version": "",
|
||||
"current_version": "20260718-abc",
|
||||
"last_error": "",
|
||||
"profile": { },
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ],
|
||||
"buffered": [ ],
|
||||
"health_events": [ ],
|
||||
"waf_ip_group_checksums": { }
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Layer | Meaning |
|
||||
| --- | --- | --- |
|
||||
| identity/version/last_error | control | who the node is, what version it runs |
|
||||
| `profile` | low-frequency overview | hostname, core count, etc. (report on change) |
|
||||
| `access_logs` | **L1** | access detail increments |
|
||||
| `edge_health` | **L2** | OpenResty health + current connections |
|
||||
| `host_metrics` | **L3** | CPU/memory/disk/NIC readings |
|
||||
| `buffered` | backfill | batches of facts accumulated while offline |
|
||||
| `health_events` | events | e.g. openresty_unhealthy |
|
||||
| `waf_ip_group_checksums` | sync | not an observability lake |
|
||||
|
||||
**Removed from the protocol (no compatibility layer; old Agents must upgrade):**
|
||||
|
||||
- `traffic_report`
|
||||
- `openresty_observation` (incl. rx/tx)
|
||||
- `snapshot` / `buffered_observability`
|
||||
- business-meaning openresty throughput fields
|
||||
|
||||
---
|
||||
|
||||
## 4. L1 Business: access_logs
|
||||
|
||||
### 4.1 Where Collection Comes From
|
||||
|
||||
| Step | Location | Description |
|
||||
| --- | --- | --- |
|
||||
| 1 | OpenResty `log_format openflare_json` | writes one JSON line per request to `access_log_path` |
|
||||
| 2 | Agent **tails increments** by file offset | new lines between two heartbeats |
|
||||
| 3 | parse and put into `access_logs[]` | overlong paths may be truncated; **no sum/count** |
|
||||
|
||||
Log format (OpenResty variables):
|
||||
|
||||
```text
|
||||
ts ← $time_iso8601
|
||||
host ← $host
|
||||
path ← $request_uri
|
||||
remote_addr ← $remote_addr
|
||||
status ← $status
|
||||
request_time ← $request_time
|
||||
bytes_sent ← $body_bytes_sent 【data provided = response body bytes】
|
||||
request_length← $request_length 【data received】
|
||||
user_agent ← $http_user_agent
|
||||
cache_status ← $upstream_cache_status 【cache status; UI can derive hit/origin/un-cached】
|
||||
```
|
||||
|
||||
Observability-port requests **don't write** business access.log (separate server with `access_log off`).
|
||||
|
||||
### 4.2 Report Example
|
||||
|
||||
```json
|
||||
"access_logs": [
|
||||
{
|
||||
"logged_at_unix": 1721289601,
|
||||
"remote_addr": "198.51.100.20",
|
||||
"host": "www.example.com",
|
||||
"path": "/api/v1/ping",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 1024,
|
||||
"request_length": 128,
|
||||
"request_time_ms": 15,
|
||||
"user_agent": "curl/8.0",
|
||||
"cache_status": "MISS"
|
||||
},
|
||||
{
|
||||
"logged_at_unix": 1721289602,
|
||||
"remote_addr": "198.51.100.21",
|
||||
"host": "www.example.com",
|
||||
"path": "/index.html",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 8192,
|
||||
"request_length": 300,
|
||||
"request_time_ms": 8,
|
||||
"user_agent": "Mozilla/5.0",
|
||||
"cache_status": "HIT"
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
| Field | Explanation |
|
||||
| --- | --- |
|
||||
| `bytes_sent` | **data provided** (single request); global/Zone totals = Server `sum` |
|
||||
| `request_length` | **data received** (single request) |
|
||||
| `logged_at_unix` | request completion time (business timeline) |
|
||||
| `host` | used for Zone domain filtering |
|
||||
| `cache_status` | `$upstream_cache_status` as-is; detail/list can derive three states (hit/origin/un-cached); **no upstream address reported** |
|
||||
| no `region` | **written by Server at insert time** via GeoIP |
|
||||
|
||||
### 4.3 How the Server Uses It (product metrics)
|
||||
|
||||
| Product Metric | Algorithm (L1 only) |
|
||||
| --- | --- |
|
||||
| Data provided | `sum(bytes_sent)` |
|
||||
| Data received | `sum(request_length)` |
|
||||
| Request count | `count()` |
|
||||
| UV | `uniqExact(remote_addr)` |
|
||||
| Status distribution | `group by status_code` |
|
||||
| Top domains | `group by host` |
|
||||
| Zone page | same + `host IN (that Zone's domains)` |
|
||||
| Dashboard business area | same, global or Top-filtered |
|
||||
|
||||
Stored in: `of_node_access_logs` (optional Server-side `of_access_log_hourly` acceleration, **Agent never writes it**).
|
||||
|
||||
### 4.4 Report Frequency
|
||||
|
||||
```text
|
||||
Request happens ──immediately──► write access.log
|
||||
Agent every 3s ──moves──► new lines in those 3s (possibly 0, possibly many)
|
||||
Server ──immediately/batched──► CH
|
||||
```
|
||||
|
||||
Business volume correctness does **not** depend on 3s alignment; 3s only affects "detail arrival latency at the control plane" and per-packet line count.
|
||||
|
||||
---
|
||||
|
||||
## 5. L2 Health: edge_health and `/openflare/observability`
|
||||
|
||||
### 5.1 Local Monitoring Endpoint
|
||||
|
||||
**Data collection endpoint:**
|
||||
|
||||
```http
|
||||
GET http://127.0.0.1:{openresty_observability_port}/openflare/observability
|
||||
```
|
||||
|
||||
Default port: **18081** (`openresty_observability_port`).
|
||||
|
||||
**Responsibility:** answers "how is OpenResty right now", **not** "how much business data was provided".
|
||||
|
||||
#### Response Example
|
||||
|
||||
```json
|
||||
{
|
||||
"ok": true,
|
||||
"captured_at_unix": 1721289600,
|
||||
"connections": {
|
||||
"active": 42,
|
||||
"reading": 0,
|
||||
"writing": 1,
|
||||
"waiting": 41
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Instant? | Source | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `ok` | this probe | returns 200 → true | liveness |
|
||||
| `captured_at_unix` | sampling time | `ngx.time()` | aligned with report |
|
||||
| `connections.active` | **instant** | Nginx connection state (original stub_status Active) | current active connections |
|
||||
| `reading` / `writing` / `waiting` | **instant** | same, subdivided | optional but recommended |
|
||||
|
||||
**Not returned (removed):**
|
||||
|
||||
| Old Field | Reason |
|
||||
| --- | --- |
|
||||
| `request_count` / `error_count` / UV / status_codes / top_domains | business window summaries, now from access log |
|
||||
| `openresty_rx_bytes` / `openresty_tx_bytes` | duplicates data provided/received and error-prone |
|
||||
| `source_countries` | never implemented; countries go through Server GeoIP |
|
||||
| `server.accepts/handled/requests` | process cumulative counters, easily confused with business requests; not on the main path |
|
||||
|
||||
**`/openflare/stub_status`:** kept; `/openflare/observability` internally reads that endpoint to assemble the connection-count JSON, and the Agent health check also probes it directly.
|
||||
|
||||
### 5.2 Collection Mechanism (read snapshot)
|
||||
|
||||
```text
|
||||
Nginx maintains Active connections etc. on connect/disconnect
|
||||
│
|
||||
Agent GET /openflare/observability
|
||||
│
|
||||
only reads "current values" and returns JSON
|
||||
```
|
||||
|
||||
- No access.log scanning, no 60-second business averages.
|
||||
- Returns an **instant gauge snapshot**.
|
||||
|
||||
### 5.3 Report Example (packed into NodePayload)
|
||||
|
||||
```json
|
||||
"edge_health": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 42
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Source |
|
||||
| --- | --- |
|
||||
| `status` / `message` | Agent health probe (config validation/process etc., may work with the observability endpoint's `ok`); must align with top-level `openresty_status` / `openresty_message` |
|
||||
| `connections` | observability endpoint `connections.active` |
|
||||
|
||||
**Storage split (authoritative sources):**
|
||||
|
||||
| Content | Written To |
|
||||
| --- | --- |
|
||||
| latest `status` + `message` | **PG node table** (UI / list / alerts) |
|
||||
| time-series `status` + `connections` | **CH `of_node_edge_health`** (**no message**) |
|
||||
|
||||
---
|
||||
|
||||
## 6. L3 Host: host_metrics
|
||||
|
||||
### 6.1 Where Collection Comes From
|
||||
|
||||
The Agent reads the local machine (e.g. `/proc`, disk stats), **once per packet**.
|
||||
|
||||
| Field | Semantics | Description |
|
||||
| --- | --- | --- |
|
||||
| `cpu_usage_percent` | instant | current CPU% |
|
||||
| `memory_*` / `storage_*` | instant used/total | usage rates computed at Server or display layer |
|
||||
| `disk_read_bytes` / `disk_write_bytes` | **cumulative counter** | kernel cumulative IO |
|
||||
| `network_rx_bytes` / `network_tx_bytes` | **cumulative counter** | **host NIC**, not data provided |
|
||||
|
||||
### 6.2 Report Example
|
||||
|
||||
```json
|
||||
"host_metrics": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"cpu_usage_percent": 12.5,
|
||||
"memory_used_bytes": 4294967296,
|
||||
"memory_total_bytes": 16106127360,
|
||||
"storage_used_bytes": 50000000000,
|
||||
"storage_total_bytes": 107374182400,
|
||||
"disk_read_bytes": 9000000000,
|
||||
"disk_write_bytes": 12000000000,
|
||||
"network_rx_bytes": 500000000000,
|
||||
"network_tx_bytes": 800000000000
|
||||
}
|
||||
```
|
||||
|
||||
### 6.3 How the Server Handles Cumulative Fields
|
||||
|
||||
```text
|
||||
store raw-value time series
|
||||
when displaying "NIC outbound over this period":
|
||||
delta = current - previous
|
||||
if delta < 0 → treat as restart/counter reset, record this segment's increment as 0, continue from new baseline
|
||||
if delta >= 0 → record into that period's increment
|
||||
```
|
||||
|
||||
- The Agent **reports raw values**, never computes 24h totals at the edge.
|
||||
- **Forbidden** to `sum` cumulative raw values as business volume.
|
||||
- Copy must be **"host NIC"**, never "data provided / OpenResty outbound".
|
||||
|
||||
Stored in: `of_node_metric_snapshots` (optional capacity hourly MV).
|
||||
|
||||
---
|
||||
|
||||
## 7. One Complete Report Example
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n_01hxyz",
|
||||
"name": "edge-shanghai-1",
|
||||
"ip": "203.0.113.10",
|
||||
"version": "3.4.0",
|
||||
"ext_version": "",
|
||||
"current_version": "20260718-abc",
|
||||
"last_error": "",
|
||||
"host_metrics": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"cpu_usage_percent": 12.5,
|
||||
"memory_used_bytes": 4294967296,
|
||||
"memory_total_bytes": 16106127360,
|
||||
"storage_used_bytes": 50000000000,
|
||||
"storage_total_bytes": 107374182400,
|
||||
"disk_read_bytes": 9000000000,
|
||||
"disk_write_bytes": 12000000000,
|
||||
"network_rx_bytes": 500000000000,
|
||||
"network_tx_bytes": 800000000000
|
||||
},
|
||||
"edge_health": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 42
|
||||
},
|
||||
"access_logs": [
|
||||
{
|
||||
"logged_at_unix": 1721289595,
|
||||
"remote_addr": "198.51.100.20",
|
||||
"host": "www.example.com",
|
||||
"path": "/",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 4096,
|
||||
"request_length": 200,
|
||||
"request_time_ms": 12
|
||||
}
|
||||
],
|
||||
"buffered": [],
|
||||
"health_events": [],
|
||||
"waf_ip_group_checksums": {
|
||||
"1": "d41d8cd98f00b204e9800998ecf8427e"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Server storage sketch:**
|
||||
|
||||
| payload block | written to |
|
||||
| --- | --- |
|
||||
| `access_logs[0]` | one CH row, `bytes_sent=4096`, `region` filled by GeoIP |
|
||||
| `edge_health` | node `openresty_status=healthy`, connections=42 |
|
||||
| `host_metrics` | one CH metric row with cumulative/instant fields |
|
||||
|
||||
**Product query sketch (24h):**
|
||||
|
||||
- data provided = `sum(bytes_sent)` over that node's (or global) logs
|
||||
- current connections = latest `edge_health.connections`
|
||||
- host NIC outbound = sum of non-negative `network_tx` deltas over metrics
|
||||
|
||||
The three numbers **need not be equal**.
|
||||
|
||||
---
|
||||
|
||||
## 8. Offline Backfill `buffered`
|
||||
|
||||
When reporting fails, the Agent caches **the same kind of facts** locally by window (default ~60 minutes), then packs them into `buffered[]` after recovery:
|
||||
|
||||
```json
|
||||
"buffered": [
|
||||
{
|
||||
"captured_at_unix": 1721289500,
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ]
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
- Only facts, no legacy TrafficReport.
|
||||
- Server processing logic is identical to the main fields.
|
||||
|
||||
---
|
||||
|
||||
## 9. End-to-End Timeline (default 3s)
|
||||
|
||||
```text
|
||||
t=0.0s visitor request completes → writes one access.log line; connection count may change
|
||||
t=0.1s another request → another log line
|
||||
…
|
||||
t=3s Agent heartbeat:
|
||||
· reads 2 access_logs lines
|
||||
· GET observability → connections=42
|
||||
· reads host_metrics
|
||||
· sends to Server
|
||||
t=3s+ Server stores; dashboard/Zone queries aggregate logs
|
||||
t=6s next round…
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 10. Old Model Comparison
|
||||
|
||||
| Old Approach | New Model |
|
||||
| --- | --- |
|
||||
| Lua dict 60s window request_count + Agent 10s pull + Server sum | **removed**; request count = log count |
|
||||
| openresty_tx as "outbound" | **removed**; data provided = `sum(bytes_sent)` |
|
||||
| Two endpoints observability + stub_status | data collection unified through observability; stub_status kept as liveness and internal read endpoint |
|
||||
| TrafficReport pre-aggregation | **removed**; no such path in protocol or API |
|
||||
| Business and NIC both called "traffic" | **separate copy, separate APIs, separate tables** |
|
||||
| health status/message | **PG latest-state authority**; CH only status+connection time series |
|
||||
|
||||
---
|
||||
|
||||
## 11. Config and Implementation Index
|
||||
|
||||
| Item | Location/Key |
|
||||
| --- | --- |
|
||||
| Heartbeat interval | Agent `heartbeat_interval`; control plane `agent_heartbeat_interval` (default 3000ms) |
|
||||
| Offline threshold | control plane `node_offline_threshold` (default 60000ms) |
|
||||
| Observability port | `openresty_observability_port` (default 18081) |
|
||||
| access.log path | `access_log_path` |
|
||||
| Replay minutes | `observability_replay_minutes` (default 60) |
|
||||
| Protocol types | `pkg/protocol/agent.go` (evolves to v2 when landing) |
|
||||
| Table DDL | [observability-data-model.md](./observability-data-model.md) |
|
||||
|
||||
---
|
||||
|
||||
## 12. Revision History
|
||||
|
||||
| Date | Notes |
|
||||
| --- | --- |
|
||||
| 2026-07-18 | initial draft: single-page "latest transport model" — three layers, frequency, sample JSON, collection sources, old-model comparison |
|
||||
| 2026-07-18 | default report interval 3s; offline threshold 60s; replay window 60 minutes |
|
||||
| 2026-07-18 | M5: edge_health table, access_log_hourly, deprecate request_reports/obs_openresty throughput tables |
|
||||
| 2026-07-18 | no compatibility layer: removed "may ignore during compat period" wording; health message only in PG, no message in CH |
|
||||
@@ -0,0 +1,204 @@
|
||||
# Origin Error Page Design
|
||||
|
||||
You will learn: when an origin or the gateway returns a specified error status code, how OpenFlare replaces the pass-through response with a globally configurable page; how the config enters the immutable config version; and how the edge OpenResty keeps the real HTTP status code while displaying it in the page.
|
||||
|
||||
This design is the productized complement of the reverse proxy traffic path in [System Architecture](./architecture.md); the config release model is in [Agent & Publish Model](./agent-design.md).
|
||||
|
||||
---
|
||||
|
||||
## 1. Goals and Non-Goals
|
||||
|
||||
### 1.1 Goals
|
||||
|
||||
* **Interceptable**: for a user-configured status code set, replace the previously pass-through origin/Nginx default error response with a unified HTML.
|
||||
* **Disableable**: when the global switch is off, behavior matches today (pass-through / Nginx default page).
|
||||
* **Visible by default**: enabled by default, default status code tag `500-599`, default minimal OpenFlare error page.
|
||||
* **Customizable**: admins can edit the full HTML online; empty HTML means the built-in default template.
|
||||
* **Status passthrough**: the HTTP response `status` keeps the original error code (e.g. 502, 522); the page body shows the same value via `{{status}}`.
|
||||
* **Globally unified**: a single config under sidebar「Website Management → Response Pages」shared by all reverse proxy routes.
|
||||
* **Consistent with release**: the config persists via Option, enters the config version snapshot, and is distributed with release/rollback.
|
||||
|
||||
### 1.2 Non-Goals
|
||||
|
||||
* Per-route / per-Zone error page overrides
|
||||
* Hosting error pages via file upload (online HTML only)
|
||||
* Modifying WAF / PoW / rate-limit's own response pages (unless the user adds those status codes to the list)
|
||||
* Pages static route error pages
|
||||
* Multi-language error pages, brand asset CDN
|
||||
|
||||
---
|
||||
|
||||
## 2. Product Behavior
|
||||
|
||||
### 2.1 When to Replace
|
||||
|
||||
| Condition | Behavior |
|
||||
| --- | --- |
|
||||
| Switch on and the response status falls in the expanded set | return custom/default HTML, **status unchanged** |
|
||||
| Switch on with GET-only enabled, non-GET request returns a matching status | pass through the origin's raw response, no replacement |
|
||||
| Switch off | no `error_page` directives generated, pass through |
|
||||
| Status not in the set | no replacement |
|
||||
| Pages upstream routes | this feature is not applied |
|
||||
| Origin returns 2xx/3xx/4xx successfully (not configured) | no replacement |
|
||||
|
||||
In all-methods mode, `proxy_intercept_errors on` is enabled on the reverse proxy `location`, so **origin-returned** matching 5xx etc. are also intercepted, not just gateway-local 502s; GET-only mode switches to Lua header/body filters that only replace GET response bodies.
|
||||
|
||||
### 2.2 Status Code Tag Syntax
|
||||
|
||||
Each Tags Input entry:
|
||||
|
||||
| Form | Example | Meaning |
|
||||
| --- | --- | --- |
|
||||
| Single code | `522` | only that code |
|
||||
| Closed range | `500-599` | expand including endpoints |
|
||||
|
||||
* Valid range: single codes and range endpoints must be in **400–599**; `lo ≤ hi`.
|
||||
* Default tag list: `["500-599"]`.
|
||||
* Persist the **raw tags** (JSON array string); expand, dedupe, and sort at render time.
|
||||
* If the expanded result is empty while enabled → save rejected.
|
||||
* Invalid tags → save rejected with a readable error.
|
||||
|
||||
### 2.3 Page Placeholders
|
||||
|
||||
| Placeholder | Meaning |
|
||||
| --- | --- |
|
||||
| `{{status}}` | the current response status code (consistent with the HTTP status) |
|
||||
| `{{host}}` | request Host |
|
||||
|
||||
Both custom HTML and the default template support these placeholders; replaced at the edge at runtime. Unused placeholders may be omitted from the template.
|
||||
|
||||
### 2.4 Default Page
|
||||
|
||||
Built-in minimal white-background OpenFlare default page: large pass-through status code, short English description, Host, and a brand footer. Supports `{{status}}` / `{{host}}`; the frontend can load prebuilt styles from the built-in template catalog on the edit page.
|
||||
|
||||
---
|
||||
|
||||
## 3. Config Model
|
||||
|
||||
### 3.1 Option Keys (`w_system_configs` / OpenFlare Option API)
|
||||
|
||||
| Key | Type | Default | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `origin_error_page_enabled` | bool string | `true` | master switch |
|
||||
| `origin_error_page_status_codes` | JSON string array | `["500-599"]` | raw tags |
|
||||
| `origin_error_page_html` | text | `""` | empty = built-in default; max **256 KiB** |
|
||||
| `origin_error_page_get_only` | bool string | `false` | replace error pages only for GET; other methods pass through |
|
||||
|
||||
Reuses APIs:
|
||||
|
||||
* `GET /api/v1/d/option`
|
||||
* `POST /api/v1/d/option/update-batch`
|
||||
|
||||
No new resource routes. goose migration writes the seed; constants defined in the `internal/model` config key area.
|
||||
|
||||
### 3.2 Validation (update-batch)
|
||||
|
||||
1. `enabled`: parseable as bool.
|
||||
2. `status_codes`: valid JSON array; each entry `^\d{3}$` or `^\d{3}-\d{3}$`; expanded values all in 400–599; non-empty when enabled.
|
||||
3. `html`: length ≤ 256 KiB (bytes); empty allowed.
|
||||
4. Parse/expand logic is a **pure function** shared by the API and `pkg/render/openresty` to avoid semantic forks.
|
||||
|
||||
No XSS sanitization on HTML: it's an admin global ops config consistent with public edge display; docs warn not to embed untrusted third-party scripts.
|
||||
|
||||
### 3.3 Config Version Snapshot
|
||||
|
||||
`ConfigSnapshot` adds fields:
|
||||
|
||||
```text
|
||||
OriginErrorPageEnabled bool
|
||||
OriginErrorPageStatusCodes []string // raw tags
|
||||
OriginErrorPageHTML string // empty => renderer uses built-in default
|
||||
OriginErrorPageGetOnly bool
|
||||
```
|
||||
|
||||
Read from Option when building the snapshot; the Agent only consumes the snapshot, never reading the control-plane DB directly.
|
||||
|
||||
---
|
||||
|
||||
## 4. Edge Rendering
|
||||
|
||||
### 4.1 Content Generated When Enabled
|
||||
|
||||
1. **SupportFile**: error page template (e.g. `error_pages/origin_error.html.tmpl`), content is the custom HTML or built-in default, keeping `{{status}}` / `{{host}}`.
|
||||
2. **Each reverse proxy server** (HTTP/HTTPS proxy; excluding Pages):
|
||||
|
||||
```nginx
|
||||
proxy_intercept_errors on;
|
||||
error_page <expanded codes...> @__openflare_origin_error;
|
||||
|
||||
location @__openflare_origin_error {
|
||||
default_type text/html;
|
||||
charset utf-8;
|
||||
content_by_lua_block {
|
||||
# read template, replace {{status}} / {{host}}, output body
|
||||
# ngx.status keeps the original error code
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 4.2 Runtime Replacement
|
||||
|
||||
Use a **named location with `content_by_lua_block`** to read the template and replace placeholders — the status is **not** baked into a static file (status differs per request). GET-only mode uses `header_filter_by_lua_block` + `body_filter_by_lua_block` inside the reverse proxy location to replace only GET response bodies; non-GET requests pass through.
|
||||
|
||||
Never rewrite the error page to HTTP 200.
|
||||
|
||||
### 4.3 When Disabled
|
||||
|
||||
Do not output `proxy_intercept_errors`, `error_page`, the internal location, or the corresponding SupportFile (or the file may be written but unreferenced). GET-only mode also omits the Lua filters.
|
||||
|
||||
### 4.4 Interaction with Cache / Stale
|
||||
|
||||
If global `proxy_cache_use_stale` returns stale cache for some error codes, **successful stale responses never enter `error_page`**. The error page is only shown when the client actually receives an error status in the configured list. Behavior depends on existing cache directives; this feature does not change stale policy.
|
||||
|
||||
---
|
||||
|
||||
## 5. Frontend
|
||||
|
||||
### 5.1 Entry
|
||||
|
||||
* Sidebar「Website Management → Response Pages」: Error Page tab (`/responses`), edit page `/responses/error-page/edit`, preview page `/responses/error-page/preview`.
|
||||
|
||||
### 5.2 Page Structure
|
||||
|
||||
* Header note: takes effect after releasing via「Version Release」.
|
||||
* **Switch + Tags Input** (shadcn-extension Tags Input: `@/components/ui/tags-input`): status code tags.
|
||||
* **HTML editor area** +「Load default template」「Restore default (clear)」+ placeholder docs.
|
||||
* **Client-side preview**: replace with sample `status=502`, `host=example.com` and preview in sandbox/iframe.
|
||||
* Save: `OptionService.updateBatch`; permissions same as the performance tuning page (admin).
|
||||
|
||||
### 5.3 Component Dependencies
|
||||
|
||||
Tags Input and the HTML editor reuse existing shadcn/ui components, consistent with the existing UI style.
|
||||
|
||||
---
|
||||
|
||||
## 6. Data Flow
|
||||
|
||||
```text
|
||||
Admin /responses (Error Page tab)
|
||||
→ Option update-batch (validate tags & HTML)
|
||||
→ w_system_configs
|
||||
|
||||
Release config version
|
||||
→ snapshot writes OriginErrorPage*
|
||||
→ render OpenResty conf + SupportFile
|
||||
→ Agent pulls and reloads
|
||||
|
||||
Visitor requests a proxied domain
|
||||
→ origin/gateway produces a matching status code
|
||||
→ error_page → named location
|
||||
→ replace placeholders, keep original status, return HTML
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. Decision Record
|
||||
|
||||
| Decision | Choice | Reason |
|
||||
| --- | --- | --- |
|
||||
| Config scope | global | product requirement; simple implementation and ops |
|
||||
| Storage | Option + config version | consistent with performance tuning, rollbackable |
|
||||
| Status input | tags: single code and range | default whole 5xx, but can name 522 |
|
||||
| Response status | keep original | correct for monitoring/SEO/client semantics |
|
||||
| Runtime replacement | internal + lightweight template replacement | status differs per request |
|
||||
| Customization | online HTML | flexible without a file-upload chain |
|
||||
@@ -0,0 +1,252 @@
|
||||
# Pages Static Hosting Design
|
||||
|
||||
You will learn: the architecture design of OpenFlare Pages static hosting, the immutable deployment and secure extraction flow, the OpenResty static serving and API reverse proxy config rendering, and the cooperative workflow between the control plane and the Agent.
|
||||
|
||||
---
|
||||
|
||||
## Requirements Analysis
|
||||
|
||||
In modern web operations, besides reverse proxying dynamic applications, deploying and hosting static frontend sites (SPA apps built with React/Vue, or static generator output like Hugo/VitePress) is extremely common. Traditional approaches suffer from:
|
||||
1. **Release disconnected from proxy config**: after uploading frontend build artifacts to the Nginx host, you still need to modify the Nginx vhost config manually or via other scripts — error-prone and without version control.
|
||||
2. **Multi-node distribution is hard**: with multiple edge nodes managed by the control plane, syncing static files to all nodes consistently requires complex sync scripts (e.g. rsync).
|
||||
3. **Rollback lacks consistency**: once a new frontend package fails or has serious defects, you must restore both the static files and the proxy rules — atomic rollback is hard.
|
||||
|
||||
To solve these, OpenFlare introduces **Pages static hosting**, inspired by Cloudflare Pages. It brings "pre-built artifact import" and "website proxy rule config" into the same control plane, leveraging OpenFlare's pull-based cooperative architecture to achieve eventual convergence across multiple Agents via immutable deployments, single-node atomic switching, and periodic reconciliation, with fast rollback.
|
||||
|
||||
---
|
||||
|
||||
## Core Features
|
||||
|
||||
The Pages static hosting subsystem includes:
|
||||
* **Pre-built artifact deployment**: upload a static resource archive directly, or save a Remote URL or public GitHub Release asset source for a project. External sources are only accessed by the Server; on successful sync they uniformly create or reuse an immutable deployment and activate it atomically.
|
||||
* **Immutable deployment snapshots**: each local upload creates a new candidate deployment; persistent-source sync creates or reuses a deployment by source identity/revision and activates it. All deployments have a unique ID and whole-package SHA-256, support keeping the most recent N historical versions per system config, and can be rolled back anytime.
|
||||
* **Check and auto-update**: GitHub latest can be checked periodically per project; by default it only hints at available updates; only after an admin explicitly enables it does it auto-sync and publish by the exact revision found.
|
||||
* **SPA Fallback**: supports fallback routing for single-page apps; when a static file isn't found, requests redirect to the entry file.
|
||||
* **Built-in API reverse proxy**: enables API proxying within Pages rules with one click, eliminating cross-origin issues by forwarding requests to a designated backend.
|
||||
* **Secure package validation and extraction**: built-in path-traversal defense, symlink-hijack protection, file size/count limits, and configurable upload package size control to keep nodes physically safe.
|
||||
* **Configurable limits**: admins can adjust the "deployment package size limit" and "historical deployment retention" in ops settings.
|
||||
|
||||
### Deployment Sources
|
||||
|
||||
Projects currently support three source views: manual, Remote URL, and GitHub Release. No source record means manual; switching or deleting a source doesn't delete historical deployments or change the current active deployment. Remote URL only allows manual "Sync and Publish"; GitHub Release supports manual check/sync for latest/tag, and only latest can opt into scheduled checking and auto-update.
|
||||
|
||||
A source is mutable config; a deployment is an immutable fact. Source config and runtime cursor/state/lease are stored separately; deployments only save the security provenance snapshot at creation time. All artifacts reuse the "download or receive artifact → real-byte and entry validation → `upload.Ingest` → deployment" artifact pipeline: manual uploads stop at candidate, waiting for explicit admin activation; persistent-source sync creates-or-loads and atomically activates in the same business transaction.
|
||||
|
||||
The admin project detail is organized as "current production deployment → deployment source → deployment history".
|
||||
|
||||
---
|
||||
|
||||
## Pages Static Hosting Architecture
|
||||
|
||||
Pages hosting is logically split into a **Control Plane** and a **Data Plane**.
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
%% data flow
|
||||
Browser[1. Browser / Visitor] -->|HTTPS request / traffic| OpenResty[2. OpenResty / WAF]
|
||||
OpenResty -->|1. static serving try_files| StaticFiles[3. Edge node local static dir current]
|
||||
OpenResty -->|2. forward API proxy| BackEnd[4. Backend API service]
|
||||
|
||||
%% control flow & heartbeat
|
||||
Admin[Admin / CI] -->|upload or configure source| Server[OpenFlare Server control plane]
|
||||
Providers[Remote / GitHub Provider] -->|restricted artifact candidate| Server
|
||||
Scanner[internal scanner / action task] -->|check & auto sync| Server
|
||||
Server <-->|Agent API / Heartbeat| Agent[openflare-agent process]
|
||||
Server -.->|unified upload.Ingest| UploadStore[(platform upload backend)]
|
||||
|
||||
Agent -->|1. discover new version| Server
|
||||
Agent -->|2. download deployment package| Server
|
||||
Agent -->|3. validate, extract, atomic switch| StaticFiles
|
||||
```
|
||||
|
||||
* **Control Plane**: the Server receives local uploads or fetches Remote/GitHub pre-built artifacts via restricted providers; action tasks and the internal scanner handle checking, syncing, and auto-updates. All artifacts pass unified inspection and `upload.Ingest` into the platform storage backend; manual uploads create a new candidate, persistent-source sync creates-or-loads a deployment and activates it atomically. Config release only compiles the stable project anchor and static serving metadata.
|
||||
* **Data Plane**: the Agent discovers Pages projects referenced by config during heartbeat/WS reconciliation, pulls the project's currently active package via a dedicated API, and performs validation and extraction. OpenResty serves static files locally; the Agent doesn't know whether the artifact came from upload, Remote, or GitHub.
|
||||
|
||||
---
|
||||
|
||||
## Data Model and Metadata Design
|
||||
|
||||
### 1. Core DB Entities
|
||||
* **Pages project (`of_pages_projects`)**:
|
||||
* Records the business name, Slug (URL-friendly), enabled state, static serving root dir (RootDir, nullable), entry filename (EntryFile, default `index.html`), SPA Fallback settings, and API reverse proxy config (APIProxyPath, APIProxyPass, APIProxyRewrite).
|
||||
* **Source config (`of_pages_project_sources`)**:
|
||||
* At most one mutable source config per project, distinguishing Remote URL and GitHub Release by `source_type`. `config_version` fences stale tasks; the full Remote URL lives only in the config table — never in responses, logs, task payloads, or deployment provenance. V2 does not promise DB column encryption.
|
||||
* **Source runtime (`of_pages_project_source_runtime`)**:
|
||||
* 1:1 with the source, storing ETag, seen/applied revision, last check/sync, next check, errors, and lease. State is fixed to `idle | checking | update_available | syncing | failed | attention`; queued/completed state is carried by `TaskExecution`.
|
||||
* **Pages deployment (`of_pages_deployments`)**:
|
||||
* Records immutable deployment facts: project-incremental deployment number, whole-package SHA-256, `upload_id`, file count/total bytes, creator, and nullable source identity/revision, source security snapshot, and trigger. `artifact_path` is only a legacy-compat field and is no longer the storage truth for new deployments.
|
||||
* **Deployment file manifest (`of_pages_deployment_files`)**:
|
||||
* Stores the full regular-file path list and actual byte counts per deployment for console display and statistics.
|
||||
* No longer computes content hashes per file; integrity is guaranteed by the **whole-package** SHA-256 (`of_pages_deployments.checksum`), verified by the Agent when pulling.
|
||||
* Control-plane inspection reads the archive via file handles, streaming each regular file body and checking declared vs actual size — no whole-package `ReadFile` into memory, no per-file disk write for hashing.
|
||||
|
||||
### 2. Route Association and Snapshot
|
||||
`proxy_routes` rules associate with a Pages project via `upstream_type = "pages"` and `pages_project_id`. A route may join the release flow only when its type is `pages` and the project has an activated deployment.
|
||||
The version snapshot emitted at release includes `snapshotPagesDeployment`:
|
||||
```json
|
||||
{
|
||||
"project_id": 1,
|
||||
"project_slug": "my-spa-app",
|
||||
"deployment_id": 12,
|
||||
"deployment_number": 3,
|
||||
"checksum": "a7b3c2...",
|
||||
"entry_file": "index.html",
|
||||
"spa_fallback_enabled": true,
|
||||
"spa_fallback_path": "/index.html",
|
||||
"api_proxy_enabled": true,
|
||||
"api_proxy_path": "/api",
|
||||
"api_proxy_pass": "http://api.internal:8000",
|
||||
"api_proxy_rewrite": "/api/(.*) /$1",
|
||||
"local_root": "__OPENFLARE_PAGES_DIR__/projects/1/current"
|
||||
}
|
||||
```
|
||||
|
||||
### 3. Dual-Track Relationship with the Main Config Version (project anchor + latest pull)
|
||||
* The **main config version** and **Pages deployments** are two independent version systems.
|
||||
* The stable anchor of a Pages route in the main config is **`pages_project_id` (project ID)**, not a deployment ID.
|
||||
* The OpenResty `root` uses the project-level path `__OPENFLARE_PAGES_DIR__/projects/{project_id}/current`; the path stays unchanged on activation switch, so swapping packages never requires republishing the main config.
|
||||
* The Agent requests the "latest active package" per project (like `github/release/latest`):
|
||||
* `GET /api/v1/agent/pages/projects/:project_id/latest/hash`
|
||||
* `GET /api/v1/agent/pages/projects/:project_id/latest/package`
|
||||
* The control plane returns the deployment ID, hash, package size, and expanded manifest metadata for the project's **currently active deployment**. The Agent uses the deployment ID and other latest metadata to detect pointer races during download, but the stable anchor of the main config and local dir remains the project ID.
|
||||
* Therefore: switching the active deployment within a project **does not require publishing the main config**; the Agent polls the latest hash during periodic reconciliation, downloads on change, and switches `current`.
|
||||
* The `pages_deployment` field in the snapshot still records release-time metadata (entry file, SPA/API proxy, etc.) but does not lock the Agent's package version.
|
||||
|
||||
---
|
||||
|
||||
## Server (Control Plane) Responsibilities and Lifecycle
|
||||
|
||||
### 1. Deployment Package Security Validation and Analysis
|
||||
To protect the server from untrusted artifacts, the control plane applies the same strict validation to local uploads and all external sources:
|
||||
* **Format support**: `zip`, `tar.gz` / `tgz`, `tar.xz` / `txz`, `tar.bz2` / `tbz2`, `tar`, `7z`.
|
||||
* **Size limits**: archive size is controlled by system config `pages_max_package_size_mb` (default 100 MiB, range 1–2048); expanded single-file and total limits are "package size × 4" with a floor of 100 MiB. Inspection always streams regular file bodies, checking declared vs actual size and enforcing limits on actual values.
|
||||
* **Count limit**: at most 1,000 static files per package.
|
||||
* **Symlink blocking**: any symlink detected while walking the archive immediately errors and rejects the upload, defending against symlink-hijack attacks.
|
||||
* **Path traversal defense**: every archive file path is `Clean`ed and checked for `..` or leading `/`, defending against directory-traversal writes to sensitive system paths.
|
||||
* **Entry file validation**: the project's entry file (e.g. `index.html`, possibly under `project.RootDir`) must exist in the package, otherwise the upload is rejected.
|
||||
* **Common root prefix stripping**: many packaging tools add a redundant top folder as a common root prefix; the control plane auto-detects and safely strips it.
|
||||
* **Whole-package integrity**: SHA-256 is computed once over the archive bytes at upload/import and written to the deployment record; the Agent reconciles against the whole-package hash after pulling. No per-file content hashes.
|
||||
* **Actual size recheck**: `InspectOptions.VerifySizes` is kept only for compatibility; current inspection always reads regular file bodies, checks declared values, and accumulates actual sizes, but still does not compute per-file content hashes.
|
||||
* **History retention**: system config `pages_max_history_count` (default 20; 0 = unlimited) trims after successful deployment. Semantics: **each project keeps at most N deployments**; the currently active deployment is always kept, remaining slots fill from newest to oldest by deployment ID. With `history_count=1`, manual uploads temporarily keep both the active and the newest candidate; the next upload replaces the old candidate; after the candidate activates, the strict limit resumes. Exceeding non-active deployments and their file manifests are deleted; the corresponding upload record is soft-deleted idempotently via platform primitives — Pages never physically deletes blobs that may be shared by dedup. If trimming fails after a successful deployment, it only logs and doesn't roll back activation; concurrent operations may temporarily exceed N and converge on later trims. Main config version rollback does not depend on old Pages packages (see the dual-track section above).
|
||||
|
||||
### 2. Deployment Package Storage Planning
|
||||
The control plane stores local, Remote, and GitHub artifacts into the configured local/S3 backend via the unified upload framework (`upload.Ingest`), recording `upload_id` and the file manifest in the DB. **Large static packages never enter `config_versions` records or any config push channel**, keeping control-plane data sync lightweight.
|
||||
|
||||
### 3. Source Check, Auto-Update, and Upload Compensation
|
||||
|
||||
* `openflare:pages_source_action` executes admin check/sync or scanner-dispatched exact-revision sync; payloads never carry URL, Token, ETag, or lease tokens. Manual sync only accepts real user actors; auto sync only accepts the system actor with the `scheduled_auto_update` trigger.
|
||||
* `openflare:pages_source_scan` is a fixed `*/5 * * * *` internal-only TaskHandler accepting only `{}`; it never appears in generic task types or the schedule management UI. Each round runs "recover expired leases → compensate orphan uploads → scan due sources".
|
||||
* The scanner sorts stably by `next_check_at, source_id`, serially checking at most 20 GitHub latest sources per batch; ETag/304 still advances the check time; 403/429 record the status code and the actual backoff deadline; a single source failure doesn't block subsequent sources.
|
||||
* On finding an update, the seen cursor is always saved first. Only with `auto_update_enabled=true` and a normal `update_available` state does it dispatch sync with the exact revision found this check; `attention`, Remote, and fixed tags never auto-publish. Manually activating another deployment fences in-flight tasks and disables auto.
|
||||
* Orphan compensation checks at most 100 upload records per round that have been quarantined for at least 2 hours, requiring a system owner, Pages retention type, V2 marker, and no deployment references. Candidates are re-checked in the `project → source → runtime → upload` lock order and only soft-deleted via the upload framework with stat updates — never physically deleting blobs possibly shared by dedup.
|
||||
|
||||
---
|
||||
|
||||
## Agent (Data Landing) Responsibilities and Self-Healing
|
||||
|
||||
The Agent runs on each edge proxy node: on first applying config referencing a Pages project, and on subsequent periodic latest reconciliation, it "atomically" pulls the currently active static assets to the node.
|
||||
|
||||
### 1. Pull Latest Per Project
|
||||
1. The Agent parses routes with `UpstreamType == "pages"` from the active main config and collects the stable anchor **`pages_project_id`**.
|
||||
2. For each project it calls `GET /api/v1/agent/pages/projects/:project_id/latest/hash` to get the control plane's currently active package hash (a latest pointer).
|
||||
3. If the local `projects/{project_id}/releases/{hash}` isn't ready, it streams `.../latest/package` to a temp file, enforcing real response limits and SHA-256; after download it **requests the hash again** to avoid activation-switch races, retrying a bounded number of times on mismatch.
|
||||
4. The request carries the node's `X-Agent-Token`.
|
||||
|
||||
### 2. Secure Extraction, Atomic Switch, Keep Only Latest
|
||||
1. Absolute package cap is 2 GiB; the downloaded content's SHA-256 must match the latest hash from the post-download re-query; the whole package never enters `[]byte`.
|
||||
2. Extract into a random staging dir `projects/{project_id}/releases/.{hash}-<random>.tmp` (zip / tar.* / 7z supported), rejecting path traversal, links, and special files. The Agent obeys both Server metadata limits and local absolute limits: at most 1,000 files, single file and total at most 8 GiB.
|
||||
3. After extraction, walk the actual file tree and precisely recheck file count and total bytes against the Server metadata; mismatch → refuse to switch.
|
||||
4. Write `.openflare-pages.json`, then rename to `releases/{hash}`.
|
||||
5. **Atomic switch** `projects/{project_id}/current` to the new release (symlink preferred, copy on failure).
|
||||
6. **Only after the new package is ready and current has switched successfully**, delete other `releases/*` (including `.tmp`) under the project — **historical deployment packages are never kept on the edge**. Each project always keeps exactly one latest content per node.
|
||||
7. Multi-project reconciliation **isolates failures**: a single project failure logs and continues with others, finally aggregating errors.
|
||||
|
||||
---
|
||||
|
||||
## OpenResty (Static Serving and Proxy) Config Rendering
|
||||
|
||||
For Pages-hosted sites, the control plane renders the corresponding `server` block, replacing the regular proxy route's `proxy_pass`.
|
||||
|
||||
### 1. Static Serving Directive Rendering
|
||||
* **`root` and `index`**:
|
||||
The Server points `root` at the project-level placeholder path `__OPENFLARE_PAGES_DIR__/projects/{project_id}/current` (optionally appending `RootDir`). Activation switching only changes directory contents, not the path, so swapping packages never requires republishing the main config.
|
||||
```nginx
|
||||
server {
|
||||
listen 80;
|
||||
server_name myapp.example.com;
|
||||
|
||||
root "/var/lib/openflare/pages/projects/3/current";
|
||||
index "index.html";
|
||||
...
|
||||
}
|
||||
```
|
||||
|
||||
### 2. try_files and SPA Fallback
|
||||
* **SPA Fallback disabled (default)**:
|
||||
only match physically existing files, otherwise strict 404:
|
||||
```nginx
|
||||
location / {
|
||||
try_files $uri $uri/ =404;
|
||||
}
|
||||
```
|
||||
* **SPA Fallback enabled**:
|
||||
if the requested file doesn't exist, redirect to the project's configured entry fallback (usually `/index.html`):
|
||||
```nginx
|
||||
location / {
|
||||
try_files $uri $uri/ /index.html;
|
||||
}
|
||||
```
|
||||
|
||||
### 3. API Reverse Proxy and Rewrite Rendering
|
||||
When a static frontend needs backend API access without cross-origin issues, enable the API proxy. The OpenResty renderer nests a dedicated API `location` branch inside the static `server` block:
|
||||
```nginx
|
||||
server {
|
||||
listen 80;
|
||||
server_name myapp.example.com;
|
||||
...
|
||||
# API proxy path match
|
||||
location /api {
|
||||
# apply rewrite rules when configured
|
||||
rewrite ^/api/(.*)$ /v1/$1 break;
|
||||
rewrite ^/api$ / break;
|
||||
|
||||
proxy_pass http://api.internal:8000;
|
||||
proxy_http_version 1.1;
|
||||
proxy_set_header Host $http_host;
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
|
||||
proxy_set_header X-Forwarded-Proto $scheme;
|
||||
proxy_set_header Upgrade $http_upgrade;
|
||||
proxy_set_header Connection $connection_upgrade;
|
||||
}
|
||||
|
||||
location / {
|
||||
try_files $uri $uri/ /index.html;
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Interaction Logic and Sync Flow
|
||||
|
||||
A full pre-built artifact import and activation lifecycle looks like this. Binding a project for the first time requires publishing the main config; subsequent active deployment changes converge independently via the project latest:
|
||||
|
||||
```text
|
||||
[Admin / scanner] [Server control plane] [Agent] [OpenResty]
|
||||
| | | |
|
||||
|-- manual upload ----->|-- inspect / Ingest ---->| |
|
||||
| |-- create candidate | |
|
||||
|-- explicit activate -->|-- switch active | |
|
||||
| | | |
|
||||
|-- source sync ------->|-- inspect / Ingest | |
|
||||
| |-- create/load + atomic activation | |
|
||||
| | | |
|
||||
|-- first bind & publish|-- broadcast project anchor -->|-- write/reload route -->|
|
||||
| | | |
|
||||
|-- later activate/sync/rollback -->|-- active latest changed -->| |
|
||||
| |<-- latest metadata reconciliation ----| |
|
||||
| |--- stream package -------------------->| |
|
||||
| | |-- validate, extract, recheck --|
|
||||
| | |-- atomic switch current ------>|
|
||||
```
|
||||
@@ -0,0 +1,130 @@
|
||||
# Intranet Penetration Tunnel Design Document
|
||||
|
||||
You will learn: The architectural design of the OpenFlare intranet penetration tunnel, the internal principles of the dual-ended control components (Relay and Client), their interaction logics, and the communication flows for the data plane and control plane.
|
||||
|
||||
---
|
||||
|
||||
## Requirements Analysis
|
||||
|
||||
In typical web application hosting scenarios, many origin servers (Origin Servers) are deployed in local intranet environments (such as local development machines, LAN servers, or firewalled private clusters). These servers typically suffer from:
|
||||
1. **No Public IP**: Cannot be directly accessed by public internet traffic.
|
||||
2. **Security Compliance Restrictions**: Creating port mappings (NAT) on border routers is strictly prohibited by security policies.
|
||||
3. **Dynamic IP Changes**: Traditional DDNS solutions exhibit high latency and are highly unstable.
|
||||
|
||||
To allow internal origin servers to seamlessly integrate into the OpenFlare global data gateway, benefiting from premium features like WAF geographic protection and TLS certificate hosting, OpenFlare designed an end-to-end solution based on a **reverse relay penetration tunnel**. In this architecture, public edge nodes act as reverse proxy entrances and traffic relays, while the intranet side only needs to initiate secure outbound connections to achieve secure and stable reverse penetration of public traffic to internal origin servers.
|
||||
|
||||
---
|
||||
|
||||
## Core Capabilities
|
||||
|
||||
The intranet penetration tunnel subsystem includes the following core capabilities:
|
||||
|
||||
* **Dynamic Relay Node Management**: The control plane dynamically dispatches relay services (frps), distributing service ports and authentication tokens dynamically.
|
||||
* **Multi-Tunnel Reverse Proxy Mapping**: Supports mapping multiple internal web ports on a single intranet client, binding multiple domain routes to corresponding relay nodes.
|
||||
* **Independent Process Lifecycle Control**: Both the relay and client are independent daemon processes written in Go, responsible for spawning, monitoring, self-healing, and hot-upgrading the underlying frp engine.
|
||||
* **Token-based Independent Authentication**: The relay uses `agent_token` for authorization, whereas the intranet client uses its dedicated `tunnel_token`, enforcing isolation of permissions and routing boundaries.
|
||||
* **Validation & Incremental Hot Reload**: Config files are rewritten and processes are gracefully reloaded only when tunnel bindings, certificates, or Relay topologies change, reducing runtime overhead.
|
||||
|
||||
---
|
||||
|
||||
## Intranet Penetration & Tunnel Architecture
|
||||
|
||||
The intranet penetration subsystem is integrated on top of the mature and high-performance `frp` tunnel protocol, divided into the **Control Plane** and the **Data Plane**.
|
||||
|
||||
```mermaid
|
||||
graph TD
|
||||
%% Data Flow
|
||||
Browser[1. Browser / Visitor] -->|HTTPS Request| Agent[2. OpenResty / Agent]
|
||||
Agent -->|Local proxy_pass| RelayFrps[3. OpenFlare Relay / frps]
|
||||
RelayFrps -->|Encrypted Tunnel Protocol| FlaredFrpc[4. OpenFlared / frpc]
|
||||
FlaredFrpc -->|Forward Local Request| LocalOrigin[5. Intranet Origin 192.168.x.x]
|
||||
|
||||
%% Control Flow & Heartbeats
|
||||
Server[OpenFlare Server Control Plane] <-->|Relay API / Heartbeat| RelayManager[openflare-relay process]
|
||||
Server <-->|Client API / Heartbeat| ClientManager[openflared process]
|
||||
|
||||
RelayManager -.->|Control Process & Config| RelayFrps
|
||||
ClientManager -.->|Control Multi-Relay Processes| FlaredFrpc
|
||||
|
||||
style Browser fill:#f9f,stroke:#333,stroke-width:2px
|
||||
style LocalOrigin fill:#9f9,stroke:#333,stroke-width:2px
|
||||
style Server fill:#f96,stroke:#333,stroke-width:2px
|
||||
```
|
||||
|
||||
* **Control Plane**: The Server maintains the database state. The `openflare-relay` process on relay nodes and the `openflared` process on intranet servers synchronize tunnel configurations via HTTP heartbeats and long-lived WebSocket connections.
|
||||
* **Data Plane**: Public traffic enters the public edge Agent (OpenResty), where the TLS handshake, HTTPS termination, and WAF filtering are executed. It is then forwarded via `proxy_pass` to the co-located `openflare-relay (frps)` on the loopback address. `frps` encapsulates the HTTP requests into the encrypted TCP tunnel and sends them down to the intranet `openflared (frpc)`. Finally, `frpc` unpacks the requests and forwards them to the actual intranet origin service.
|
||||
|
||||
---
|
||||
|
||||
## Relay (Server-side) Design
|
||||
|
||||
`openflare-relay` is a relay manager deployed on the public edge, running on nodes of type `tunnel_relay`.
|
||||
|
||||
### 1. Core Architecture & Logic
|
||||
* **Process Daemon**: The Relay process embeds the `frps` binary, spawning the `frps -c frps.toml` subprocess via `exec.Command` and using goroutines to asynchronously listen to its exit status. If `frps` exits unexpectedly, it automatically restarts using an exponential backoff policy.
|
||||
* **Dynamic Configuration Rendering**: Periodically synchronizes status with the control plane via HTTP heartbeats to retrieve the active `RelayConfig`, including:
|
||||
* `bindPort`: The public control port that frps listens to for incoming intranet frpc connections.
|
||||
* `vhostHTTPPort`: The virtual host HTTP listening port where the Agent's proxy_pass points.
|
||||
* `authToken`: The security credential used during the client connection handshake.
|
||||
* `webServer`: Enables the frps dashboard API, which the Relay queries to collect active tunnel counts and traffic metrics.
|
||||
* **Status Reporting**: In each heartbeat cycle, the Relay reports the active connections, registered clients, individual proxy tunnel statuses, and Relay version back to the Server.
|
||||
|
||||
---
|
||||
|
||||
## Openflared (Client-side) Design
|
||||
|
||||
`openflared` is the client manager running inside the user's intranet server, authenticated using a dedicated `tunnel_token`.
|
||||
|
||||
### 1. Core Design Mechanisms
|
||||
* **Multi-Relay Support (Multiplexing)**:
|
||||
To guarantee high availability and geographical proximity, the control plane may schedule the client to connect to multiple public Relays. `openflared` parses the list of Relays dispatched in the `TunnelConfig`, generating dedicated configurations (`frpc_<relay_node_id>.toml`) and allocating distinct cancelable contexts for each Relay process locally.
|
||||
* **Independent Subprocess Monitoring**:
|
||||
`openflared` maintains a local `processes` map to manage the lifecycles of individual `frpc` subprocesses. When the control plane adds or removes Relays, the client incrementally spawns new processes or gracefully shuts down obsolete ones without affecting other functioning tunnels.
|
||||
* **Dynamic TOML Generation**:
|
||||
When rendering TOML configs for each Relay, the client iterates over the Proxies list, writing each intranet service's `LocalAddr`, `LocalPort`, and bound `CustomDomains` into standard `[[proxies]]` blocks.
|
||||
|
||||
---
|
||||
|
||||
## Interaction Logic & Traffic Model
|
||||
|
||||
The intranet penetration subsystem implements consistent version control and status feedback loops.
|
||||
|
||||
### 1. Control Plane Publishing & Sync Flow
|
||||
|
||||
```text
|
||||
Admin modifies tunnel/intranet port mappings -> Click Publish -> Generate new Tunnel version & Checksum
|
||||
|
|
||||
v (Push or Heartbeat Pull)
|
||||
+-----------------------------------------------------------------------+-----------------------------------------------------------------------+
|
||||
| |
|
||||
v (Relay Side) v (Client Side)
|
||||
openflare-relay heartbeat detects frps port/Token change openflared heartbeat detects tunnel_version change
|
||||
Re-render local frps.toml Request full proxy configuration details
|
||||
Kill and restart the frps process Re-render frpc_<relay_id>.toml configs
|
||||
Report health status as healthy Restart or hot-reload changed frpc processes
|
||||
Report application results (Apply Success/Error)
|
||||
```
|
||||
|
||||
1. **Versioned Controls**: All intranet tunnel routes and mapping relationships are version-controlled, dispatching a unique `version` and `checksum` to ensure clients do not repeatedly write files or trigger redundant reloads.
|
||||
2. **Closed-Loop Application Feedback**: After applying new configurations, the client reports the application result in the next heartbeat. If the intranet port is unreachable or certificate bindings fail, the client intercepts the stdout/stderr of the subprocess to report `LastError` to the Server, providing administrators with transparent error details.
|
||||
|
||||
### 2. Data Plane Traffic Model
|
||||
1. **Public Entrance (Agent)**:
|
||||
```nginx
|
||||
server {
|
||||
listen 443 ssl;
|
||||
server_name intranet.example.com;
|
||||
# ... TLS certificates & WAF filtering ...
|
||||
location / {
|
||||
proxy_pass http://127.0.0.1:18080; # Points to local frps vhost port
|
||||
proxy_set_header Host $host; # Must preserve the original Host header, which frps relies on to route requests
|
||||
proxy_set_header X-Real-IP $remote_addr;
|
||||
}
|
||||
}
|
||||
```
|
||||
2. **Relay Node (frps)**:
|
||||
`frps` listens to the Vhost port `18080`. When an HTTP request arrives, it extracts `Host: intranet.example.com` from the request headers and searches its active registered tunnel registry to locate the matching encrypted TCP connection (initiated by the intranet frpc).
|
||||
3. **Encrypted Tunnel Transmission (TCP)**:
|
||||
`frps` encapsulates the HTTP request into the custom TCP tunnel protocol and transmits it down to the intranet `frpc` client.
|
||||
4. **Intranet Client Distribution (frpc)**:
|
||||
The `frpc` instance managed by `openflared` receives the payload, resolves it according to local settings (`localIP = "127.0.0.1"`, `localPort = 8080`), initiates a local TCP connection to forward the request to the intranet web service, and returns the response back through the tunnel to the public viewer.
|
||||
@@ -0,0 +1,132 @@
|
||||
# WAF Design Document
|
||||
|
||||
You will learn: The core architecture of the OpenFlare edge Web Application Firewall (WAF), the dynamic IP group asynchronous differential sync model, the high-performance OpenResty Lua caching scheme, and the complete request filtering and decision logic.
|
||||
|
||||
---
|
||||
|
||||
## Requirements Analysis
|
||||
|
||||
In public internet environments, web applications face a wide variety of security threats (such as scanner profiling, api scraping, malicious botnets targeted at specific regions, ransomware, and CC attacks). Allowing malicious requests to pass directly to the origin server (Origin Server) results in:
|
||||
1. **Origin Server Overload**: High-frequency database queries and intensive CPU computations easily exhaust server resources.
|
||||
2. **Sensitive API Abuse**: APIs like login, registration, and SMS verification codes can be maliciously exploited, leading to financial and computational losses.
|
||||
3. **Data Exposure Risks**: Malicious common vulnerability probing actions are not intercepted proactively.
|
||||
|
||||
Therefore, OpenFlare needs to build a **high-performance, resiliently scalable WAF filtering engine** at the frontmost data plane layer (OpenResty). This engine is capable of executing deep filtering on malicious requests at the edge layer closest to users with sub-millisecond overhead. This relieves pressure on origin servers and provides core security capabilities like CC protection (PoW challenge), IP whitelisting/blacklisting, and region-level interception.
|
||||
|
||||
---
|
||||
|
||||
## Core Capabilities
|
||||
|
||||
OpenFlare WAF includes the following core protection dimensions:
|
||||
|
||||
* **IP Interception (IP Whitelist/Blacklist)**: Supports filtering by single IP or CIDR block, and aggregating tens of thousands of IPs into IP groups for highly efficient matching.
|
||||
* **Geographical Whitelist/Blacklist (GeoIP Limit)**: Integrates MaxMind databases to support precise admission controls based on countries and provinces/regions.
|
||||
* **Custom Interception Responses**: Supports custom block status codes (e.g., 403, 418) and personalized HTML block pages for different filtering rules.
|
||||
* **Human-Machine Challenge (PoW CC Protection)**: Supports seamless client-side PoW challenges, calculating Hash collisions to prevent automated scripts and botnets from hitting endpoints concurrently.
|
||||
|
||||
---
|
||||
|
||||
## IP Group Design & Dynamic Asynchronous Sync
|
||||
|
||||
IP groups are the core containers for highly efficient IP whitelisting and blacklisting. OpenFlare classifies IP groups into three types based on their update frequencies and source channels:
|
||||
|
||||
### 1. IP Group Types
|
||||
* **Manual**: Manually input by administrators in the control panel. Primarily used for static trusted IPs or long-term blocks.
|
||||
* **Subscription**: Configured with remote text feeds (one IP/CIDR per line) or standard JSON subscription URLs. Server-side cron jobs periodically fetch and parse the remote subscription sources. Primarily used for integrating open-source threat intelligence feeds, cloud provider IP ranges, etc.
|
||||
* **Automatic**: **The most resilient dynamic protection channel**. Control plane scanning jobs read access logs from all nodes, performing aggregation and analysis based on configured Expr rules (e.g., "requesting the `/api/login` endpoint over 50 times with a 401 status code in 5 minutes"). Once matched, the source IP is automatically added to a temporary block list for a specified duration.
|
||||
|
||||
### 2. Asynchronous Differential Sync Design (No Nginx Reload)
|
||||
In traditional Nginx WAF designs, IP blacklist updates typically require writing configurations and executing reloads. If malicious IP blocks occur at high frequencies (seconds or minutes), frequent reloads force Nginx to constantly spawn new worker processes and tear down old ones, severely degrading performance.
|
||||
|
||||
OpenFlare adopts a **dynamic IP group asynchronous differential sync design**:
|
||||
|
||||
```text
|
||||
WAF IP member updates (Manual/Subscription/Auto-trigger)
|
||||
|
|
||||
v
|
||||
Server updates the database and calculates the new MD5 Checksum of the IP group
|
||||
|
|
||||
+----------------------------------------+
|
||||
| (WebSocket Real-time Broadcast) | (Heartbeat Fallback Comparison)
|
||||
v v
|
||||
Server immediately pushes complete members Agent heartbeats report the local IP groups
|
||||
of modified groups to all Agents checksum mapping table
|
||||
| |
|
||||
| v
|
||||
| Server detects Checksum mismatch and dispatches
|
||||
v the modified IP group members
|
||||
Agent receives member data and writes it as JSON to local disk: waf_ip_groups.json
|
||||
|
|
||||
v (Lua Memory Awareness)
|
||||
OpenResty Lua engine detects file changes via MD5 checksum in seconds and hot-updates its memory,
|
||||
completely bypassing Nginx process reloads.
|
||||
```
|
||||
|
||||
Through this architecture, the persistence and activation of tens of thousands of highly volatile dynamic blacklist IPs **require absolutely no Nginx reloads**, maximally protecting the high-concurrency throughput of the gateway.
|
||||
|
||||
---
|
||||
|
||||
## Rule Groups & Site Bindings
|
||||
|
||||
* **WAF Rule Group**: The smallest logical collection of WAF filtering policies. A single rule group can contain IP whitelists/blacklists, IP group references, regional restrictions, and CC protection configurations.
|
||||
* **Global Rule Group**: When a rule group is marked as `is_global = true`, it takes effect on **all website routes** hosted on the node by default.
|
||||
* **Site Binding**: Website routes (`proxy_routes`) can bind one or more non-global rule groups. During request validation, WAF evaluates the union of `Global Rule Group + Bound Rule Groups`.
|
||||
|
||||
---
|
||||
|
||||
## Implementation Details & High-Performance Caching
|
||||
|
||||
WAF is triggered in the OpenResty `access_by_lua` phase, implemented primarily through Lua files and local JSON configurations.
|
||||
|
||||
### 1. Physical Structures
|
||||
* `waf_config.json`: Contains metadata for all rule groups, geographic country/region limits, and website-to-rule-group bindings.
|
||||
* `waf_ip_groups.json`: Contains all synchronized IP groups and their corresponding IP lists.
|
||||
* `waf/runtime.lua`: The actual runtime engine responsible for WAF rule comparison.
|
||||
* `waf/check.lua`: The entry point for the access layer, handling packages inclusion and triggering `check()`.
|
||||
|
||||
### 2. Shared Memory Dictionary (ngx.shared) High-Performance Cache Design
|
||||
Reading JSON files from the disk and decoding them upon every incoming web request would make disk I/O a severe performance bottleneck.
|
||||
|
||||
OpenFlare leverages the **OpenResty Shared Memory Dictionary (ngx.shared.openflare_waf_config)** to implement a two-level caching mechanism:
|
||||
|
||||
1. **Zero File I/O Path**:
|
||||
In Lua, every time `check()` executes, it first computes the MD5 hash of the local JSON file using `ngx.md5` (which takes virtually zero time since the file is cached in the OS Page Cache).
|
||||
2. **Hash Comparison & Hot Loading**:
|
||||
It compares this against the cached hash key (`_config_hash`) stored in the shared memory dictionary.
|
||||
* **If the hash is unchanged**: It reads the pre-decoded Lua Table configuration stored directly in shared memory. The entire verification runs purely in **shared memory**, completing in **microseconds**.
|
||||
* **If the hash is mismatched**: Indicating that the Agent has just updated the WAF rules or IP groups on the disk, the Lua engine automatically reads the disk file, decodes it via `cjson.decode`, writes the decoded data and the new MD5 hash into shared memory, and makes it seamlessly readable by all subsequent worker processes.
|
||||
|
||||
---
|
||||
|
||||
## Application Flow & Decision Judgment Control Logic
|
||||
|
||||
When an HTTP/HTTPS request arrives at OpenResty, WAF evaluates and intercepts it step-by-step in the `access` phase according to the funnel decision chain below:
|
||||
|
||||
### 1. WAF Decision Flowchart
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
A[Request enters access phase] --> B[Get Site Name of current request]
|
||||
B --> C[Load all active rule groups bound to this Site in shared memory]
|
||||
C --> D{Matches IP whitelist or Whitelist IP group?}
|
||||
D -- Yes (Matched) --> E[Pass request - ALLOW]
|
||||
D -- No --> F{Matches country/region whitelist?}
|
||||
F -- Yes (Matched) --> E
|
||||
F -- No --> G{Matches IP blacklist or Blacklist IP group?}
|
||||
G -- Yes (Matched) --> H[Block request - BLOCK]
|
||||
G -- No --> I{Matches country/region blacklist?}
|
||||
I -- Yes (Matched) --> H
|
||||
I -- No --> J{Is CC PoW verification enabled?}
|
||||
J -- Yes --> K[Transfer to CC Protection module]
|
||||
J -- No --> L[No security risks, pass normally]
|
||||
|
||||
H --> M[Exit and return custom status code and block page HTML configured in the rule group]
|
||||
```
|
||||
|
||||
### 2. Decision Step Details
|
||||
1. **Whitelist Precedence**:
|
||||
To prevent false positives and guarantee smooth passage of core back-to-source traffic (such as search engine spiders, CDN back-to-source IPs, and office egresses), WAF **prioritizes matching IP whitelists and regional whitelists**. Once a whitelist matches, it immediately bypasses all subsequent blacklist checks and CC challenges.
|
||||
2. **Blacklist Aggressive Block**:
|
||||
If a request is not captured by the whitelist evaluation, it enters the blacklist funnel. Once the source IP matches an IP blacklist, a referenced blacklist IP group, or lies within a prohibited country/region, the Lua engine immediately marks `ngx.ctx.openflare_waf_blocked` as `true`.
|
||||
3. **Response Output**:
|
||||
Upon hitting the blacklist, Lua extracts the `block_status_code` (defaults to 418 or 403) and `block_response_body` (interception HTML page) configured in the matching rule group. It outputs the response body via `ngx.say()` and gracefully terminates the request using `ngx.exit(status)` to prevent the request from passing upstream.
|
||||
@@ -0,0 +1,118 @@
|
||||
# WAF Orchestration Rule Design
|
||||
|
||||
This document defines the target architecture, data model, execution semantics, release model, and migration boundaries for reworking OpenFlare WAF from a fixed decision chain into a visual directed acyclic graph (DAG). IP group sources and membership computation still follow [WAF Design](./waf-design.md); this document only changes how rules are composed and executed.
|
||||
|
||||
## Goals and Boundaries
|
||||
|
||||
When a user adds a WAF rule, they only enter a name. The Server immediately creates a legal default graph `start → pass`, and the frontend enters a standalone orchestration page based on React Flow. Users build policies by adding processing units, configuring nodes, and connecting branches — no more filling in fixed-order allow/block lists and PoW forms.
|
||||
|
||||
Supported nodes:
|
||||
|
||||
| Node | Count Constraint | Inputs | Outputs | Config |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| Start | exactly one per graph | none | `next` | none |
|
||||
| Pass | exactly one per graph | one or more | none | none |
|
||||
| Block | multiple allowed | one or more | none | HTTP status code, HTML response body |
|
||||
| IP match | multiple allowed | one or more | `true`, `false` | IP, CIDR, IP group ID |
|
||||
| Geo match | multiple allowed | one or more | `true`, `false` | country code, region code |
|
||||
| UA check | multiple allowed | one or more | `true`, `false` | UA required, browser/OS allowlist with and/or, block crawlers/abnormal UAs (excluding crawlers)/custom regex |
|
||||
| Security | multiple allowed | one or more | `true`, `false` | basic signature detection (path traversal/file inclusion on by default; SQL/XSS/command injection/SSRF/upload/XXE/CRLF toggleable); any enabled rule hit → false |
|
||||
| PoW | multiple allowed | one or more | `next` | algorithm, difficulty, session TTL, challenge TTL |
|
||||
|
||||
IP match, geo match, UA check, and security do not distinguish allowlist vs blocklist. `true` only means the request passed that node's judgment; `false` only means it did not; the business meaning of allow vs block is entirely determined by the wiring. UA check evaluation order: require UA → block crawlers/abnormal UAs → allowlist match. Security performs signature matching on the request Path/Query/Header/Cookie/Body (bounded). After PoW verification succeeds, execution continues along `next`; when incomplete, the challenge page takes over the request and no `false` branch is produced.
|
||||
|
||||
Loops, script nodes, arbitrary expression nodes, subgraph calls, and cross-rule jumps are not implemented.
|
||||
|
||||
## Control-Plane Architecture
|
||||
|
||||
The rule graph uses a dual model separating control-plane edit state and data-plane runtime state:
|
||||
|
||||
1. The React Flow editor submits a versioned graph JSON containing node IDs, node types, display names, coordinates, typed configs, and edges.
|
||||
2. The Server performs authoritative validation of the whole graph; on success it saves the graph in a single transaction and increments the revision number.
|
||||
3. On config release, the Server validates all enabled rules again, compiles the graph into a compact runtime DAG without UI fields (coordinates, labels, etc.), and collects the referenced IP group IDs.
|
||||
4. The Agent atomically writes the full release snapshot and reloads OpenResty. New Workers only load and parse the rule JSON once at startup.
|
||||
5. The request hot path only traverses the immutable in-memory runtime graph in the Worker — no file reads, no checksum computation, no JSON parsing.
|
||||
|
||||
The edit-state JSON uses an explicit `schema_version`. Node configs use per-node-type structures; unconstrained key-value objects bypassing Server validation are not allowed. Initial safety limits: 128 nodes, 256 edges, and 256 KiB of edit-state JSON per rule; these limits are enforced by both the API and the release compiler.
|
||||
|
||||
## Graph Structure Constraints
|
||||
|
||||
Rule save and release must satisfy all constraints:
|
||||
|
||||
* The graph is a DAG; self-loops and arbitrary cycles are forbidden.
|
||||
* Exactly one start node and one pass node exist; multiple block nodes are allowed.
|
||||
* The start node has no incoming edges and exactly one `next` outlet; pass and block nodes have no outlets.
|
||||
* IP match, geo match, UA check, and security must each connect their `true` and `false` outlets once; PoW's `next` must connect once.
|
||||
* No dangling outlets except terminal nodes; every non-start node has at least one incoming edge.
|
||||
* All nodes must be reachable from start, and every executable node must be able to reach pass or block.
|
||||
* Edge source ports must belong to the source node type; the same source port must not connect to multiple targets.
|
||||
* Node IDs are unique within the graph, edge IDs are unique within the graph, and all referenced nodes must exist.
|
||||
* Node configs must pass per-type field, range, reference-existence, and size validation.
|
||||
|
||||
The frontend provides instant validation and connection restrictions to improve UX, but the Server is the only authoritative validator. When deleting a node, the frontend synchronously removes related edges and marks the rule unsaved; saving is forbidden until the graph is legal again.
|
||||
|
||||
## Multi-Rule Execution Semantics
|
||||
|
||||
A route can bind multiple custom rules. The binding is an ordered list following this order:
|
||||
|
||||
1. Enabled global rules always execute first and do not participate in route-side ordering.
|
||||
2. Enabled rules bound to the route execute in binding order.
|
||||
3. When the current rule reaches a block node, it immediately outputs that node's configured response and terminates the request.
|
||||
4. When the current rule reaches a pass node, that only means the current rule finished; if more rules remain, execution continues.
|
||||
5. Only after all rules reach a pass node is the request truly allowed to proceed into the OpenResty/origin chain.
|
||||
|
||||
The runtime graph is fully validated before release. If the Lua executor still encounters an unknown node, unknown port, missing target, or exceeds the node-step limit, it logs a rate-limited error and blocks the request, preventing a broken security config from accidentally allowing traffic.
|
||||
|
||||
## IP Group Memory Refresh
|
||||
|
||||
Rule topology only takes effect on release + OpenResty reload; IP group membership can still be updated independently by manual, subscription, or auto tasks without release or reload.
|
||||
|
||||
IP groups use a two-level cache of coordinating worker, shared snapshot, and worker-local objects:
|
||||
|
||||
1. Requests always read the IP group object in the current worker's memory — no file access or shared-dict JSON parsing.
|
||||
2. Every 5 seconds only one worker holding a shared lock reads the lightweight checksum file.
|
||||
3. If the checksum is unchanged, it ends immediately without reading the full `waf_ip_groups.json`.
|
||||
4. On checksum change, the coordinating worker reads and validates the full JSON once, writes the raw snapshot to a dedicated 64 MiB `ngx.shared.openflare_waf_ip_groups` keyed by checksum, then updates the commit pointer.
|
||||
5. Other workers detecting a shared version change fetch the snapshot from shared memory, parse it, and atomically replace their local object — no repeated disk reads.
|
||||
6. On refresh failure, keep using the previous valid object, log a rate-limited error, and retry next cycle.
|
||||
|
||||
The Agent must atomically replace the IP group JSON first, then atomically update the checksum last, so workers never recognize a half-written file as a new version. Server release/sync and Agent disk writes jointly enforce the 20 MiB aggregate snapshot limit; shared dict uses a safe write that never force-evicts old keys, keeping the current and previous immutable snapshots on failure.
|
||||
|
||||
## API and Editor
|
||||
|
||||
The create API only accepts a rule name and returns the rule detail with the default graph. Rule metadata, graph save, and route binding use separate operations, so toggling enabled state or binding does not overwrite the canvas.
|
||||
|
||||
The graph detail includes `revision`. Save requests submit `revision + graph`; the Server only updates and increments the revision when it matches. On mismatch it returns a conflict, the frontend prompts a reload, and silent overwrites of another page's changes are forbidden. The route binding API accepts an ordered array of rule IDs.
|
||||
|
||||
The React Flow editor page uses a full-width canvas and a fixed right property panel:
|
||||
|
||||
* Top bar: back, rule name, enabled state, validation state, and save.
|
||||
* The canvas uses compact height with a smaller initial fit scale; zoom, pan, box-select, delete, auto-layout, MiniMap/Controls and other necessary navigation are supported. Node dragging is handled by React Flow's local controlled state in real time; coordinates are written back to the edit graph only after the drag ends.
|
||||
* "Add processing unit" offers IP match, geo match, UA check, security, PoW, and block; start and pass are provided by the default graph and cannot be deleted or duplicated.
|
||||
* Selecting a normal node or edge allows deletion via the canvas delete button or Delete/Backspace; deleting a node synchronously removes associated edges.
|
||||
* The right property panel is hidden by default, shown only when a node is selected; it collapses when clicking an edge or blank canvas.
|
||||
* Geo-match properties use the full country and ISO 3166-2 first-level administrative division data; country options show both localized names and codes; administrative divisions support search by country name, division name, or code to avoid rendering thousands of options at once.
|
||||
* Leaving the page with unsaved changes must prompt; save conflicts and Server validation errors should locate the relevant node or edge.
|
||||
|
||||
The WAF list shows rule name, enabled state, node count, bound route count, and update time. The new-rule dialog only has the name field and navigates to the orchestration page immediately on success.
|
||||
|
||||
## Persistence and Migration
|
||||
|
||||
Rule records add a versioned graph JSON and a revision number; binding records add execution order. The graph is saved as a single aggregate (not split into node/edge tables) to keep edit operations transactional and let new node types avoid frequent DB schema extensions.
|
||||
|
||||
When upgrading existing installs:
|
||||
|
||||
* Keep rule names, global flags, enabled state, and route bindings.
|
||||
* All rule graphs reset to `start → pass`; legacy IP/geo lists, PoW, or block-response configs are not migrated.
|
||||
* Existing bindings are written into the order field in a stable sequence; global rules remain fixed in front.
|
||||
* Once the new graph and runtime stabilize, remove the legacy rule fields, fixed-order compile logic, and old frontend forms — do not maintain dual executors long-term.
|
||||
|
||||
This migration stops the old protection config from taking effect; the upgrade notes must prominently tell admins to re-orchestrate rules before releasing the next version.
|
||||
|
||||
## Release, Failure, and Rollback
|
||||
|
||||
Rule graphs only take effect on config release. If validation or compilation fails before release, the release is refused and the current active version stays unchanged. If Agent write, OpenResty config check, or reload fails, the apply flow fails and restores the previous valid released version.
|
||||
|
||||
New workers only accept complete, parseable runtime rule configs. During an OpenResty graceful reload, old workers keep the old in-memory graph and new workers use the new graph, so requests never observe a half-updated state.
|
||||
|
||||
When the geo database is unavailable, geo match returns `false` with a rate-limited warning, preserving current behavior. On IP group refresh failure, the old in-memory snapshot is kept. An incomplete PoW is taken over by the challenge module, not treated as an execution error; PoW node config is first written to OpenResty shared memory with a short-lived key, then passed to the internal challenge handler via explicit `ngx.exec` parameters — it cannot rely on internal redirects preserving `ngx.ctx` or implicitly inherited request params. Empty rule bindings in the release snapshot must be encoded as JSON empty arrays; at runtime, legacy `null` optional arrays in old snapshots are treated as empty arrays, so `cjson`'s `ngx.null` userdata never breaks the request.
|
||||
@@ -0,0 +1,93 @@
|
||||
# Zone & Domain Resource Design
|
||||
|
||||
## Goals
|
||||
|
||||
Refactor "websites" into a Zone management experience keyed by registrable root domains. A Zone like `example.com` is a stable management boundary; users enter the Zone through a stable ID path to view and maintain its explicitly declared domains, the reverse proxy routes and certificates bound to those domains, and route-level WAF, Pages and other capabilities.
|
||||
|
||||
This design replaces the concept, tables, and APIs of `managed_domains`. The Zone core does **not** include authoritative DNS record management; to point ZoneDomain A records at edge nodes, use the optional module [Cloudflare DNS Pointing](./cloudflare-pointing.md).
|
||||
|
||||
## Scope and Constraints
|
||||
|
||||
* Zone root domains are resolved with the Public Suffix List, e.g. `api.example.co.uk` belongs to `example.co.uk`.
|
||||
* URLs use IDs: list at `/websites`, detail at `/websites/:zoneId`; domains are not used as URL parameters.
|
||||
* Zone domains must be explicit FQDNs; `*.example.com` is not allowed. TLS certificates may still contain wildcard SANs and cover explicit Zone domains.
|
||||
* A Zone domain is associated with at most one reverse proxy route; a route may associate with multiple Zone domains, thus sharing the same upstream, cache, rate limit, WAF and Pages config across Zones.
|
||||
* The Zone model itself adds no DNS records, edge functions, preview subdomains, or tenant isolation. Creating/updating external DNS A records is handled by the separate Cloudflare pointing module and does not change the Zone / ZoneDomain table responsibilities.
|
||||
|
||||
## Core Model
|
||||
|
||||
```mermaid
|
||||
erDiagram
|
||||
ZONES ||--o{ ZONE_DOMAINS : contains
|
||||
PROXY_ROUTES ||--o{ ZONE_DOMAINS : serves
|
||||
TLS_CERTIFICATES ||--o{ ZONE_DOMAINS : secures
|
||||
PROXY_ROUTES ||--o{ WAF_RULE_GROUP_BINDINGS : applies
|
||||
PAGES_PROJECTS ||--o{ PROXY_ROUTES : backs
|
||||
|
||||
ZONES {
|
||||
uint id PK
|
||||
string domain UK
|
||||
}
|
||||
ZONE_DOMAINS {
|
||||
uint id PK
|
||||
uint zone_id
|
||||
uint proxy_route_id
|
||||
string domain UK
|
||||
uint cert_id
|
||||
}
|
||||
```
|
||||
|
||||
### `of_zones`
|
||||
|
||||
Stores the root domain, created time, and updated time. Root domains are globally unique and cannot be modified in place after creation; to change one, create a new Zone and migrate the domains. Before deleting a Zone, all of its Zone domains must be cleared first.
|
||||
|
||||
### `of_zone_domains`
|
||||
|
||||
Stores `zone_id`, explicit `domain`, nullable `proxy_route_id`, nullable `cert_id`, and timestamps. `domain` is globally unique; all relationship fields are indexed but no physical foreign keys are created. `proxy_route_id` may be null to host historical domains that have a certificate prepared but no reverse proxy configured yet.
|
||||
|
||||
`of_proxy_routes` gradually removes the domain/certificate redundancy columns `domain`, `domains`, `cert_id`, `cert_ids`, and `domain_cert_ids`. Routes must no longer specify any TLS certificate; the route name `site_name` becomes the stable human-readable identifier, and the compiler reads `server_name` and its `cert_id` from the associated Zone domains. This gives each explicit domain a single certificate source.
|
||||
|
||||
## Business and API
|
||||
|
||||
New Zone resources in the admin panel:
|
||||
|
||||
* `GET/POST /api/v1/d/zones`
|
||||
* `GET/POST /api/v1/d/zones/:id/update`
|
||||
* `POST /api/v1/d/zones/:id/delete`
|
||||
* `POST /api/v1/d/zones/:id/domains` (list returned via overview)
|
||||
* `POST /api/v1/d/zones/:id/domains/:domainID/update`
|
||||
* `POST /api/v1/d/zones/:id/domains/:domainID/delete`
|
||||
* `GET /api/v1/d/zones/:id/overview`
|
||||
|
||||
Reverse proxy route create/update requests switch to `zone_domain_ids` and no longer submit `domains`, `cert_id`, `cert_ids`, or `domain_cert_ids`. The server validates domain ownership, global uniqueness, and certificate SAN coverage in a transaction; failures are returned uniformly via `response.Abort*`. Deleting a Zone domain bound to a route requires unbinding or deleting the route first; deleting a Zone that still has domains must be rejected.
|
||||
|
||||
WAF, Pages, upstream, and release versions remain part of `proxy_routes`. The Zone overview only aggregates the route state associated with its domains and does not copy or redefine those configs.
|
||||
|
||||
## Frontend Experience
|
||||
|
||||
`/websites` shows only Zone root domains with configured domain count, route count and status, plus search, create, and action menus. Clicking enters `/websites/:zoneId`.
|
||||
|
||||
The detail page includes:
|
||||
|
||||
* Overview: domain, route, and valid certificate statistics; domain—route—certificate summary; route-level WAF and Pages summary.
|
||||
* Domains: a list of explicit FQDNs, certificate selection, and associated routes; wildcard domains are not shown or accepted.
|
||||
* Routes: routes filtered to the current Zone, linking to existing route details.
|
||||
* Certificates: certificates actually referenced by the current Zone's domains.
|
||||
* Settings: Zone notes and a protected delete operation.
|
||||
|
||||
When adding a route, select from Zone domains; users can also register domains in the Zone first, then bind a route. The global reverse proxy route entry remains but uses the same Zone domain selector.
|
||||
|
||||
## Data Migration
|
||||
|
||||
This rework ships in two release phases to avoid SQL using a wrong "last two labels" rule for multi-level public suffixes. Operation details: [Zone Domain Migration and Release Acceptance](../guide/zone-domain-migration.md).
|
||||
|
||||
1. **Phase 1 DDL**: PostgreSQL and SQLite goose create `of_zones` / `of_zone_domains` at the same version; `of_managed_domains` and route redundancy columns are temporarily kept.
|
||||
2. **Data import (automatic)**: at Server startup `migrator.Migrate()` first applies goose SQL up to `202607120002`, then automatically imports legacy route domains / `managed_domains` (registering root domains via `publicsuffix` parsing, writing `cert_id` and `proxy_route_id`), then continues with the remaining SQL. Conflicts fail startup; fixing and restarting retries idempotently. No manual command needed.
|
||||
3. **Code switch**: control-plane APIs, config snapshots, rendering, and frontend all use Zone domains as the single source; route writes only use `zone_domain_ids`.
|
||||
4. **Phase 2 cleanup**: goose SQL `202607130001_drop_legacy_route_domain_columns` drops `of_managed_domains` and the `of_proxy_routes` redundancy columns. Down only restores an empty dev-DB structure and does not backfill historical data.
|
||||
|
||||
### Runtime Model Boundaries
|
||||
|
||||
* Persistence: domains and certificates exist only in `of_zone_domains`; `of_proxy_routes` only stores route policy (upstream, cache, rate limit, WAF binding keys, etc.).
|
||||
* Rendering: the config snapshot assembles temporary `Domains` / `DomainCertIDs` in memory for OpenResty rendering and does not write back to the database.
|
||||
* Structure migration only uses `internal/infra/persistence/migrator/goose/{postgres,sqlite}/*.sql`; legacy domains are imported automatically at startup, and after phase 2 the old columns no longer exist so it is a no-op.
|
||||
Reference in New Issue
Block a user