mirror of
https://github.com/Rain-kl/OpenFlare.git
synced 2026-10-11 01:36:37 +08:00
docs(i18n): 恢复并补齐英文版 vitepress,README 默认改为英文
- README 默认英文:README.en.md → README.md(英文为默认),中文移至 README.zh-CN.md,语言切换链接同步 - 恢复被删除的 docs/en/ 英文文档(git 历史 cc5e53c5^),删除 4 篇已废弃文件 - 英文导航 config.ts 对齐中文结构(新增 Deployment/Changelog 侧栏,同步 Guide/Design 条目) - 翻译 15 篇中文新增文档:guide 5 篇(certificates/pages-usage/proxy-config/uptime-kuma/zone-domain-migration)+ design 10 篇(zone-design/cloudflare-pointing/waf-orchestration/origin-error-page/edge-cache-design/pages-design/logstore/kuma-design/login-captcha/observability 三篇) - en 首页更新(新增 Pages 特性、tagline 同步);changelog 英文入口指向中文版 - vitepress 构建验证:43 个英文页面全部渲染 注意:29 篇旧英文文档为恢复版,部分内容(如 deployment/server、reference/configuration)可能落后于中文,需后续逐篇同步
This commit is contained in:
@@ -0,0 +1,502 @@
|
||||
# Edge Observability Transport Model (current target version)
|
||||
|
||||
> **This document is the latest authoritative description of "how Agent ↔ Server observability data is transmitted".**
|
||||
> After reading you should be able to answer: what is sent, where it is collected from, how often, how the Server stores it, and where product metrics are queried from.
|
||||
> Protocol fields and DDL details: [Observability Reporting Protocol & Data Model](./observability-data-model.md); background: [Edge Observability & Business Traffic Stats](./observability-design.md).
|
||||
|
||||
---
|
||||
|
||||
## 0. Remember the Three Layers First
|
||||
|
||||
| Layer | Question Answered | Single Data Source | Product Examples |
|
||||
| --- | --- | --- | --- |
|
||||
| **L1 Business delivery** | How much data provided? How many requests? | **access.log details** | data provided, request count, UV, status codes, Top domains |
|
||||
| **L2 Edge health** | Is OpenResty alive? Current connections? | **local `/openflare/observability`** | node health, current connections |
|
||||
| **L3 Host capacity** | How are CPU/memory/disk/NIC? | **OS readings** | capacity trends, host NIC |
|
||||
|
||||
**The three layers are never reconciled against each other.**
|
||||
"Data provided" ≠ "current connections" ≠ "host NIC outbound".
|
||||
|
||||
---
|
||||
|
||||
## 1. Overview: Who Collects, Who Reports, Who Aggregates
|
||||
|
||||
```text
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Edge Node │
|
||||
│ │
|
||||
│ Visitor request ──► OpenResty │
|
||||
│ │ │
|
||||
│ ├─ access.log (one line per request) ←── L1 collection point │
|
||||
│ │ │
|
||||
│ └─ connection state (maintained in-process) │
|
||||
│ │ │
|
||||
│ ▼ │
|
||||
│ GET /openflare/observability ←── L2 reads snapshot │
|
||||
│ (no log scanning, no business recomputation) │
|
||||
│ │
|
||||
│ OS /proc etc. ──────────────────────────── L3 reads snapshot │
|
||||
│ │
|
||||
│ ┌────────── Agent ──────────┐ │
|
||||
│ │ default: one NodePayload per 3s │ │
|
||||
│ │ · tail access.log incremental │ │
|
||||
│ │ · GET local observability │ │
|
||||
│ │ · read host_metrics │ │
|
||||
│ └────────────┬──────────────┘ │
|
||||
└─────────────────────────────│──────────────────────────────────┘
|
||||
│ HTTP heartbeat or WebSocket status
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────┐
|
||||
│ Server (control plane) │
|
||||
│ · details → ClickHouse of_node_access_logs │
|
||||
│ · health → node latest state + of_node_edge_health │
|
||||
│ · host → of_node_metric_snapshots │
|
||||
│ · business trends / Zone stats = sum/count/uniq over access_logs only │
|
||||
└─────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
| Role | Does | Doesn't |
|
||||
| --- | --- | --- |
|
||||
| OpenResty | writes access.log; maintains connection counts | no direct reporting to the control plane |
|
||||
| Agent | **collects facts and reports them** | **does not compute** UV/TopN/24h data provided |
|
||||
| Server | stores + **aggregates/interpreters** | does not trust edge business pre-summaries |
|
||||
|
||||
---
|
||||
|
||||
## 2. Collection Frequency (defaults)
|
||||
|
||||
| Action | Default Frequency | Config |
|
||||
| --- | --- | --- |
|
||||
| Agent → Server report | **every 3 seconds** a full payload | `heartbeat_interval` / control-plane `agent_heartbeat_interval` (ms, default `3000`) |
|
||||
| Tail access.log when packing | **with report** (new lines since last report) | same |
|
||||
| GET `/openflare/observability` when packing | **with report** (reads **current** connection snapshot) | same |
|
||||
| Read host metrics when packing | **with report** | same |
|
||||
| OpenResty writes access.log | **1 line at each request end** | unrelated to heartbeat |
|
||||
| Connection counts update in-process | **on connection change** (kernel-maintained) | unrelated to heartbeat |
|
||||
| Offline replay window | keep ~**60 minutes** by default | `observability_replay_minutes` |
|
||||
| Node offline detection | ~**60s** without a successful heartbeat | `node_offline_threshold` (default `60000` ms) |
|
||||
|
||||
**Notes:**
|
||||
|
||||
- The Agent has **no separate "sampling clock"**; **sampling points = report points** (default 3s).
|
||||
- access.log is "per-request continuous writes"; the Agent only **moves incremental lines** periodically.
|
||||
- `/openflare/observability` is **not** "business stats start being counted when called"; for connections it **reads Nginx's existing instantaneous values**.
|
||||
|
||||
Transport channels:
|
||||
|
||||
- **HTTP heartbeat**: POST the full payload at the interval.
|
||||
- **WebSocket**: after connecting, sends `status` messages at the same interval (same content shape); HTTP heartbeat is not double-sent then.
|
||||
|
||||
---
|
||||
|
||||
## 3. Agent → Server Packet (NodePayload v2)
|
||||
|
||||
### 3.1 Structure Skeleton
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n_01hxyz",
|
||||
"name": "edge-shanghai-1",
|
||||
"ip": "203.0.113.10",
|
||||
"version": "3.4.0",
|
||||
"ext_version": "",
|
||||
"current_version": "20260718-abc",
|
||||
"last_error": "",
|
||||
"profile": { },
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ],
|
||||
"buffered": [ ],
|
||||
"health_events": [ ],
|
||||
"waf_ip_group_checksums": { }
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Layer | Meaning |
|
||||
| --- | --- | --- |
|
||||
| identity/version/last_error | control | who the node is, what version it runs |
|
||||
| `profile` | low-frequency overview | hostname, core count, etc. (report on change) |
|
||||
| `access_logs` | **L1** | access detail increments |
|
||||
| `edge_health` | **L2** | OpenResty health + current connections |
|
||||
| `host_metrics` | **L3** | CPU/memory/disk/NIC readings |
|
||||
| `buffered` | backfill | batches of facts accumulated while offline |
|
||||
| `health_events` | events | e.g. openresty_unhealthy |
|
||||
| `waf_ip_group_checksums` | sync | not an observability lake |
|
||||
|
||||
**Removed from the protocol (no compatibility layer; old Agents must upgrade):**
|
||||
|
||||
- `traffic_report`
|
||||
- `openresty_observation` (incl. rx/tx)
|
||||
- `snapshot` / `buffered_observability`
|
||||
- business-meaning openresty throughput fields
|
||||
|
||||
---
|
||||
|
||||
## 4. L1 Business: access_logs
|
||||
|
||||
### 4.1 Where Collection Comes From
|
||||
|
||||
| Step | Location | Description |
|
||||
| --- | --- | --- |
|
||||
| 1 | OpenResty `log_format openflare_json` | writes one JSON line per request to `access_log_path` |
|
||||
| 2 | Agent **tails increments** by file offset | new lines between two heartbeats |
|
||||
| 3 | parse and put into `access_logs[]` | overlong paths may be truncated; **no sum/count** |
|
||||
|
||||
Log format (OpenResty variables):
|
||||
|
||||
```text
|
||||
ts ← $time_iso8601
|
||||
host ← $host
|
||||
path ← $request_uri
|
||||
remote_addr ← $remote_addr
|
||||
status ← $status
|
||||
request_time ← $request_time
|
||||
bytes_sent ← $body_bytes_sent 【data provided = response body bytes】
|
||||
request_length← $request_length 【data received】
|
||||
user_agent ← $http_user_agent
|
||||
cache_status ← $upstream_cache_status 【cache status; UI can derive hit/origin/un-cached】
|
||||
```
|
||||
|
||||
Observability-port requests **don't write** business access.log (separate server with `access_log off`).
|
||||
|
||||
### 4.2 Report Example
|
||||
|
||||
```json
|
||||
"access_logs": [
|
||||
{
|
||||
"logged_at_unix": 1721289601,
|
||||
"remote_addr": "198.51.100.20",
|
||||
"host": "www.example.com",
|
||||
"path": "/api/v1/ping",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 1024,
|
||||
"request_length": 128,
|
||||
"request_time_ms": 15,
|
||||
"user_agent": "curl/8.0",
|
||||
"cache_status": "MISS"
|
||||
},
|
||||
{
|
||||
"logged_at_unix": 1721289602,
|
||||
"remote_addr": "198.51.100.21",
|
||||
"host": "www.example.com",
|
||||
"path": "/index.html",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 8192,
|
||||
"request_length": 300,
|
||||
"request_time_ms": 8,
|
||||
"user_agent": "Mozilla/5.0",
|
||||
"cache_status": "HIT"
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
| Field | Explanation |
|
||||
| --- | --- |
|
||||
| `bytes_sent` | **data provided** (single request); global/Zone totals = Server `sum` |
|
||||
| `request_length` | **data received** (single request) |
|
||||
| `logged_at_unix` | request completion time (business timeline) |
|
||||
| `host` | used for Zone domain filtering |
|
||||
| `cache_status` | `$upstream_cache_status` as-is; detail/list can derive three states (hit/origin/un-cached); **no upstream address reported** |
|
||||
| no `region` | **written by Server at insert time** via GeoIP |
|
||||
|
||||
### 4.3 How the Server Uses It (product metrics)
|
||||
|
||||
| Product Metric | Algorithm (L1 only) |
|
||||
| --- | --- |
|
||||
| Data provided | `sum(bytes_sent)` |
|
||||
| Data received | `sum(request_length)` |
|
||||
| Request count | `count()` |
|
||||
| UV | `uniqExact(remote_addr)` |
|
||||
| Status distribution | `group by status_code` |
|
||||
| Top domains | `group by host` |
|
||||
| Zone page | same + `host IN (that Zone's domains)` |
|
||||
| Dashboard business area | same, global or Top-filtered |
|
||||
|
||||
Stored in: `of_node_access_logs` (optional Server-side `of_access_log_hourly` acceleration, **Agent never writes it**).
|
||||
|
||||
### 4.4 Report Frequency
|
||||
|
||||
```text
|
||||
Request happens ──immediately──► write access.log
|
||||
Agent every 3s ──moves──► new lines in those 3s (possibly 0, possibly many)
|
||||
Server ──immediately/batched──► CH
|
||||
```
|
||||
|
||||
Business volume correctness does **not** depend on 3s alignment; 3s only affects "detail arrival latency at the control plane" and per-packet line count.
|
||||
|
||||
---
|
||||
|
||||
## 5. L2 Health: edge_health and `/openflare/observability`
|
||||
|
||||
### 5.1 Local Monitoring Endpoint
|
||||
|
||||
**Data collection endpoint:**
|
||||
|
||||
```http
|
||||
GET http://127.0.0.1:{openresty_observability_port}/openflare/observability
|
||||
```
|
||||
|
||||
Default port: **18081** (`openresty_observability_port`).
|
||||
|
||||
**Responsibility:** answers "how is OpenResty right now", **not** "how much business data was provided".
|
||||
|
||||
#### Response Example
|
||||
|
||||
```json
|
||||
{
|
||||
"ok": true,
|
||||
"captured_at_unix": 1721289600,
|
||||
"connections": {
|
||||
"active": 42,
|
||||
"reading": 0,
|
||||
"writing": 1,
|
||||
"waiting": 41
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Instant? | Source | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `ok` | this probe | returns 200 → true | liveness |
|
||||
| `captured_at_unix` | sampling time | `ngx.time()` | aligned with report |
|
||||
| `connections.active` | **instant** | Nginx connection state (original stub_status Active) | current active connections |
|
||||
| `reading` / `writing` / `waiting` | **instant** | same, subdivided | optional but recommended |
|
||||
|
||||
**Not returned (removed):**
|
||||
|
||||
| Old Field | Reason |
|
||||
| --- | --- |
|
||||
| `request_count` / `error_count` / UV / status_codes / top_domains | business window summaries, now from access log |
|
||||
| `openresty_rx_bytes` / `openresty_tx_bytes` | duplicates data provided/received and error-prone |
|
||||
| `source_countries` | never implemented; countries go through Server GeoIP |
|
||||
| `server.accepts/handled/requests` | process cumulative counters, easily confused with business requests; not on the main path |
|
||||
|
||||
**`/openflare/stub_status`:** kept; `/openflare/observability` internally reads that endpoint to assemble the connection-count JSON, and the Agent health check also probes it directly.
|
||||
|
||||
### 5.2 Collection Mechanism (read snapshot)
|
||||
|
||||
```text
|
||||
Nginx maintains Active connections etc. on connect/disconnect
|
||||
│
|
||||
Agent GET /openflare/observability
|
||||
│
|
||||
only reads "current values" and returns JSON
|
||||
```
|
||||
|
||||
- No access.log scanning, no 60-second business averages.
|
||||
- Returns an **instant gauge snapshot**.
|
||||
|
||||
### 5.3 Report Example (packed into NodePayload)
|
||||
|
||||
```json
|
||||
"edge_health": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 42
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Source |
|
||||
| --- | --- |
|
||||
| `status` / `message` | Agent health probe (config validation/process etc., may work with the observability endpoint's `ok`); must align with top-level `openresty_status` / `openresty_message` |
|
||||
| `connections` | observability endpoint `connections.active` |
|
||||
|
||||
**Storage split (authoritative sources):**
|
||||
|
||||
| Content | Written To |
|
||||
| --- | --- |
|
||||
| latest `status` + `message` | **PG node table** (UI / list / alerts) |
|
||||
| time-series `status` + `connections` | **CH `of_node_edge_health`** (**no message**) |
|
||||
|
||||
---
|
||||
|
||||
## 6. L3 Host: host_metrics
|
||||
|
||||
### 6.1 Where Collection Comes From
|
||||
|
||||
The Agent reads the local machine (e.g. `/proc`, disk stats), **once per packet**.
|
||||
|
||||
| Field | Semantics | Description |
|
||||
| --- | --- | --- |
|
||||
| `cpu_usage_percent` | instant | current CPU% |
|
||||
| `memory_*` / `storage_*` | instant used/total | usage rates computed at Server or display layer |
|
||||
| `disk_read_bytes` / `disk_write_bytes` | **cumulative counter** | kernel cumulative IO |
|
||||
| `network_rx_bytes` / `network_tx_bytes` | **cumulative counter** | **host NIC**, not data provided |
|
||||
|
||||
### 6.2 Report Example
|
||||
|
||||
```json
|
||||
"host_metrics": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"cpu_usage_percent": 12.5,
|
||||
"memory_used_bytes": 4294967296,
|
||||
"memory_total_bytes": 16106127360,
|
||||
"storage_used_bytes": 50000000000,
|
||||
"storage_total_bytes": 107374182400,
|
||||
"disk_read_bytes": 9000000000,
|
||||
"disk_write_bytes": 12000000000,
|
||||
"network_rx_bytes": 500000000000,
|
||||
"network_tx_bytes": 800000000000
|
||||
}
|
||||
```
|
||||
|
||||
### 6.3 How the Server Handles Cumulative Fields
|
||||
|
||||
```text
|
||||
store raw-value time series
|
||||
when displaying "NIC outbound over this period":
|
||||
delta = current - previous
|
||||
if delta < 0 → treat as restart/counter reset, record this segment's increment as 0, continue from new baseline
|
||||
if delta >= 0 → record into that period's increment
|
||||
```
|
||||
|
||||
- The Agent **reports raw values**, never computes 24h totals at the edge.
|
||||
- **Forbidden** to `sum` cumulative raw values as business volume.
|
||||
- Copy must be **"host NIC"**, never "data provided / OpenResty outbound".
|
||||
|
||||
Stored in: `of_node_metric_snapshots` (optional capacity hourly MV).
|
||||
|
||||
---
|
||||
|
||||
## 7. One Complete Report Example
|
||||
|
||||
```json
|
||||
{
|
||||
"schema_version": 2,
|
||||
"node_id": "n_01hxyz",
|
||||
"name": "edge-shanghai-1",
|
||||
"ip": "203.0.113.10",
|
||||
"version": "3.4.0",
|
||||
"ext_version": "",
|
||||
"current_version": "20260718-abc",
|
||||
"last_error": "",
|
||||
"host_metrics": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"cpu_usage_percent": 12.5,
|
||||
"memory_used_bytes": 4294967296,
|
||||
"memory_total_bytes": 16106127360,
|
||||
"storage_used_bytes": 50000000000,
|
||||
"storage_total_bytes": 107374182400,
|
||||
"disk_read_bytes": 9000000000,
|
||||
"disk_write_bytes": 12000000000,
|
||||
"network_rx_bytes": 500000000000,
|
||||
"network_tx_bytes": 800000000000
|
||||
},
|
||||
"edge_health": {
|
||||
"captured_at_unix": 1721289600,
|
||||
"status": "healthy",
|
||||
"message": "",
|
||||
"connections": 42
|
||||
},
|
||||
"access_logs": [
|
||||
{
|
||||
"logged_at_unix": 1721289595,
|
||||
"remote_addr": "198.51.100.20",
|
||||
"host": "www.example.com",
|
||||
"path": "/",
|
||||
"status_code": 200,
|
||||
"bytes_sent": 4096,
|
||||
"request_length": 200,
|
||||
"request_time_ms": 12
|
||||
}
|
||||
],
|
||||
"buffered": [],
|
||||
"health_events": [],
|
||||
"waf_ip_group_checksums": {
|
||||
"1": "d41d8cd98f00b204e9800998ecf8427e"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Server storage sketch:**
|
||||
|
||||
| payload block | written to |
|
||||
| --- | --- |
|
||||
| `access_logs[0]` | one CH row, `bytes_sent=4096`, `region` filled by GeoIP |
|
||||
| `edge_health` | node `openresty_status=healthy`, connections=42 |
|
||||
| `host_metrics` | one CH metric row with cumulative/instant fields |
|
||||
|
||||
**Product query sketch (24h):**
|
||||
|
||||
- data provided = `sum(bytes_sent)` over that node's (or global) logs
|
||||
- current connections = latest `edge_health.connections`
|
||||
- host NIC outbound = sum of non-negative `network_tx` deltas over metrics
|
||||
|
||||
The three numbers **need not be equal**.
|
||||
|
||||
---
|
||||
|
||||
## 8. Offline Backfill `buffered`
|
||||
|
||||
When reporting fails, the Agent caches **the same kind of facts** locally by window (default ~60 minutes), then packs them into `buffered[]` after recovery:
|
||||
|
||||
```json
|
||||
"buffered": [
|
||||
{
|
||||
"captured_at_unix": 1721289500,
|
||||
"host_metrics": { },
|
||||
"edge_health": { },
|
||||
"access_logs": [ ]
|
||||
}
|
||||
]
|
||||
```
|
||||
|
||||
- Only facts, no legacy TrafficReport.
|
||||
- Server processing logic is identical to the main fields.
|
||||
|
||||
---
|
||||
|
||||
## 9. End-to-End Timeline (default 3s)
|
||||
|
||||
```text
|
||||
t=0.0s visitor request completes → writes one access.log line; connection count may change
|
||||
t=0.1s another request → another log line
|
||||
…
|
||||
t=3s Agent heartbeat:
|
||||
· reads 2 access_logs lines
|
||||
· GET observability → connections=42
|
||||
· reads host_metrics
|
||||
· sends to Server
|
||||
t=3s+ Server stores; dashboard/Zone queries aggregate logs
|
||||
t=6s next round…
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 10. Old Model Comparison
|
||||
|
||||
| Old Approach | New Model |
|
||||
| --- | --- |
|
||||
| Lua dict 60s window request_count + Agent 10s pull + Server sum | **removed**; request count = log count |
|
||||
| openresty_tx as "outbound" | **removed**; data provided = `sum(bytes_sent)` |
|
||||
| Two endpoints observability + stub_status | data collection unified through observability; stub_status kept as liveness and internal read endpoint |
|
||||
| TrafficReport pre-aggregation | **removed**; no such path in protocol or API |
|
||||
| Business and NIC both called "traffic" | **separate copy, separate APIs, separate tables** |
|
||||
| health status/message | **PG latest-state authority**; CH only status+connection time series |
|
||||
|
||||
---
|
||||
|
||||
## 11. Config and Implementation Index
|
||||
|
||||
| Item | Location/Key |
|
||||
| --- | --- |
|
||||
| Heartbeat interval | Agent `heartbeat_interval`; control plane `agent_heartbeat_interval` (default 3000ms) |
|
||||
| Offline threshold | control plane `node_offline_threshold` (default 60000ms) |
|
||||
| Observability port | `openresty_observability_port` (default 18081) |
|
||||
| access.log path | `access_log_path` |
|
||||
| Replay minutes | `observability_replay_minutes` (default 60) |
|
||||
| Protocol types | `pkg/protocol/agent.go` (evolves to v2 when landing) |
|
||||
| Table DDL | [observability-data-model.md](./observability-data-model.md) |
|
||||
|
||||
---
|
||||
|
||||
## 12. Revision History
|
||||
|
||||
| Date | Notes |
|
||||
| --- | --- |
|
||||
| 2026-07-18 | initial draft: single-page "latest transport model" — three layers, frequency, sample JSON, collection sources, old-model comparison |
|
||||
| 2026-07-18 | default report interval 3s; offline threshold 60s; replay window 60 minutes |
|
||||
| 2026-07-18 | M5: edge_health table, access_log_hourly, deprecate request_reports/obs_openresty throughput tables |
|
||||
| 2026-07-18 | no compatibility layer: removed "may ignore during compat period" wording; health message only in PG, no message in CH |
|
||||
Reference in New Issue
Block a user