Files
OpenFlare/docs/en/design/observability-transport-model.md
T
ryan 454542c1d0 docs(i18n): 恢复并补齐英文版 vitepress,README 默认改为英文
- README 默认英文:README.en.md → README.md(英文为默认),中文移至 README.zh-CN.md,语言切换链接同步
- 恢复被删除的 docs/en/ 英文文档(git 历史 cc5e53c5^),删除 4 篇已废弃文件
- 英文导航 config.ts 对齐中文结构(新增 Deployment/Changelog 侧栏,同步 Guide/Design 条目)
- 翻译 15 篇中文新增文档:guide 5 篇(certificates/pages-usage/proxy-config/uptime-kuma/zone-domain-migration)+ design 10 篇(zone-design/cloudflare-pointing/waf-orchestration/origin-error-page/edge-cache-design/pages-design/logstore/kuma-design/login-captcha/observability 三篇)
- en 首页更新(新增 Pages 特性、tagline 同步);changelog 英文入口指向中文版
- vitepress 构建验证:43 个英文页面全部渲染

注意:29 篇旧英文文档为恢复版,部分内容(如 deployment/server、reference/configuration)可能落后于中文,需后续逐篇同步
2026-08-16 23:18:29 +08:00

19 KiB

Edge Observability Transport Model (current target version)

This document is the latest authoritative description of "how Agent ↔ Server observability data is transmitted".
After reading you should be able to answer: what is sent, where it is collected from, how often, how the Server stores it, and where product metrics are queried from.
Protocol fields and DDL details: Observability Reporting Protocol & Data Model; background: Edge Observability & Business Traffic Stats.


0. Remember the Three Layers First

Layer Question Answered Single Data Source Product Examples
L1 Business delivery How much data provided? How many requests? access.log details data provided, request count, UV, status codes, Top domains
L2 Edge health Is OpenResty alive? Current connections? local /openflare/observability node health, current connections
L3 Host capacity How are CPU/memory/disk/NIC? OS readings capacity trends, host NIC

The three layers are never reconciled against each other.
"Data provided" ≠ "current connections" ≠ "host NIC outbound".


1. Overview: Who Collects, Who Reports, Who Aggregates

┌─────────────────────────────────────────────────────────────┐
│ Edge Node                                                    │
│                                                               │
│  Visitor request ──► OpenResty                                 │
│                 │                                              │
│                 ├─ access.log (one line per request)  ←── L1 collection point │
│                 │                                              │
│                 └─ connection state (maintained in-process)    │
│                        │                                       │
│                        ▼                                       │
│              GET /openflare/observability  ←── L2 reads snapshot │
│              (no log scanning, no business recomputation)      │
│                                                               │
│  OS /proc etc.  ──────────────────────────── L3 reads snapshot │
│                                                               │
│              ┌────────── Agent ──────────┐                     │
│              │ default: one NodePayload per 3s  │              │
│              │  · tail access.log incremental │               │
│              │  · GET local observability     │               │
│              │  · read host_metrics           │               │
│              └────────────┬──────────────┘                     │
└─────────────────────────────│──────────────────────────────────┘
                              │ HTTP heartbeat or WebSocket status
                              ▼
┌─────────────────────────────────────────────────────────────┐
│ Server (control plane)                                       │
│  · details → ClickHouse of_node_access_logs                  │
│  · health → node latest state + of_node_edge_health          │
│  · host → of_node_metric_snapshots                           │
│  · business trends / Zone stats = sum/count/uniq over access_logs only │
└─────────────────────────────────────────────────────────────┘
Role Does Doesn't
OpenResty writes access.log; maintains connection counts no direct reporting to the control plane
Agent collects facts and reports them does not compute UV/TopN/24h data provided
Server stores + aggregates/interpreters does not trust edge business pre-summaries

2. Collection Frequency (defaults)

Action Default Frequency Config
Agent → Server report every 3 seconds a full payload heartbeat_interval / control-plane agent_heartbeat_interval (ms, default 3000)
Tail access.log when packing with report (new lines since last report) same
GET /openflare/observability when packing with report (reads current connection snapshot) same
Read host metrics when packing with report same
OpenResty writes access.log 1 line at each request end unrelated to heartbeat
Connection counts update in-process on connection change (kernel-maintained) unrelated to heartbeat
Offline replay window keep ~60 minutes by default observability_replay_minutes
Node offline detection ~60s without a successful heartbeat node_offline_threshold (default 60000 ms)

Notes:

  • The Agent has no separate "sampling clock"; sampling points = report points (default 3s).
  • access.log is "per-request continuous writes"; the Agent only moves incremental lines periodically.
  • /openflare/observability is not "business stats start being counted when called"; for connections it reads Nginx's existing instantaneous values.

Transport channels:

  • HTTP heartbeat: POST the full payload at the interval.
  • WebSocket: after connecting, sends status messages at the same interval (same content shape); HTTP heartbeat is not double-sent then.

3. Agent → Server Packet (NodePayload v2)

3.1 Structure Skeleton

{
  "schema_version": 2,
  "node_id": "n_01hxyz",
  "name": "edge-shanghai-1",
  "ip": "203.0.113.10",
  "version": "3.4.0",
  "ext_version": "",
  "current_version": "20260718-abc",
  "last_error": "",
  "profile": { },
  "host_metrics": { },
  "edge_health": { },
  "access_logs": [ ],
  "buffered": [ ],
  "health_events": [ ],
  "waf_ip_group_checksums": { }
}
Field Layer Meaning
identity/version/last_error control who the node is, what version it runs
profile low-frequency overview hostname, core count, etc. (report on change)
access_logs L1 access detail increments
edge_health L2 OpenResty health + current connections
host_metrics L3 CPU/memory/disk/NIC readings
buffered backfill batches of facts accumulated while offline
health_events events e.g. openresty_unhealthy
waf_ip_group_checksums sync not an observability lake

Removed from the protocol (no compatibility layer; old Agents must upgrade):

  • traffic_report
  • openresty_observation (incl. rx/tx)
  • snapshot / buffered_observability
  • business-meaning openresty throughput fields

4. L1 Business: access_logs

4.1 Where Collection Comes From

Step Location Description
1 OpenResty log_format openflare_json writes one JSON line per request to access_log_path
2 Agent tails increments by file offset new lines between two heartbeats
3 parse and put into access_logs[] overlong paths may be truncated; no sum/count

Log format (OpenResty variables):

ts            ← $time_iso8601
host          ← $host
path          ← $request_uri
remote_addr   ← $remote_addr
status        ← $status
request_time  ← $request_time
bytes_sent    ← $body_bytes_sent     【data provided = response body bytes】
request_length← $request_length      【data received】
user_agent    ← $http_user_agent
cache_status  ← $upstream_cache_status  【cache status; UI can derive hit/origin/un-cached】

Observability-port requests don't write business access.log (separate server with access_log off).

4.2 Report Example

"access_logs": [
  {
    "logged_at_unix": 1721289601,
    "remote_addr": "198.51.100.20",
    "host": "www.example.com",
    "path": "/api/v1/ping",
    "status_code": 200,
    "bytes_sent": 1024,
    "request_length": 128,
    "request_time_ms": 15,
    "user_agent": "curl/8.0",
    "cache_status": "MISS"
  },
  {
    "logged_at_unix": 1721289602,
    "remote_addr": "198.51.100.21",
    "host": "www.example.com",
    "path": "/index.html",
    "status_code": 200,
    "bytes_sent": 8192,
    "request_length": 300,
    "request_time_ms": 8,
    "user_agent": "Mozilla/5.0",
    "cache_status": "HIT"
  }
]
Field Explanation
bytes_sent data provided (single request); global/Zone totals = Server sum
request_length data received (single request)
logged_at_unix request completion time (business timeline)
host used for Zone domain filtering
cache_status $upstream_cache_status as-is; detail/list can derive three states (hit/origin/un-cached); no upstream address reported
no region written by Server at insert time via GeoIP

4.3 How the Server Uses It (product metrics)

Product Metric Algorithm (L1 only)
Data provided sum(bytes_sent)
Data received sum(request_length)
Request count count()
UV uniqExact(remote_addr)
Status distribution group by status_code
Top domains group by host
Zone page same + host IN (that Zone's domains)
Dashboard business area same, global or Top-filtered

Stored in: of_node_access_logs (optional Server-side of_access_log_hourly acceleration, Agent never writes it).

4.4 Report Frequency

Request happens ──immediately──► write access.log
Agent every 3s ──moves──► new lines in those 3s (possibly 0, possibly many)
Server ──immediately/batched──► CH

Business volume correctness does not depend on 3s alignment; 3s only affects "detail arrival latency at the control plane" and per-packet line count.


5. L2 Health: edge_health and /openflare/observability

5.1 Local Monitoring Endpoint

Data collection endpoint:

GET http://127.0.0.1:{openresty_observability_port}/openflare/observability

Default port: 18081 (openresty_observability_port).

Responsibility: answers "how is OpenResty right now", not "how much business data was provided".

Response Example

{
  "ok": true,
  "captured_at_unix": 1721289600,
  "connections": {
    "active": 42,
    "reading": 0,
    "writing": 1,
    "waiting": 41
  }
}
Field Instant? Source Description
ok this probe returns 200 → true liveness
captured_at_unix sampling time ngx.time() aligned with report
connections.active instant Nginx connection state (original stub_status Active) current active connections
reading / writing / waiting instant same, subdivided optional but recommended

Not returned (removed):

Old Field Reason
request_count / error_count / UV / status_codes / top_domains business window summaries, now from access log
openresty_rx_bytes / openresty_tx_bytes duplicates data provided/received and error-prone
source_countries never implemented; countries go through Server GeoIP
server.accepts/handled/requests process cumulative counters, easily confused with business requests; not on the main path

/openflare/stub_status: kept; /openflare/observability internally reads that endpoint to assemble the connection-count JSON, and the Agent health check also probes it directly.

5.2 Collection Mechanism (read snapshot)

Nginx maintains Active connections etc. on connect/disconnect
        │
Agent GET /openflare/observability
        │
only reads "current values" and returns JSON
  • No access.log scanning, no 60-second business averages.
  • Returns an instant gauge snapshot.

5.3 Report Example (packed into NodePayload)

"edge_health": {
  "captured_at_unix": 1721289600,
  "status": "healthy",
  "message": "",
  "connections": 42
}
Field Source
status / message Agent health probe (config validation/process etc., may work with the observability endpoint's ok); must align with top-level openresty_status / openresty_message
connections observability endpoint connections.active

Storage split (authoritative sources):

Content Written To
latest status + message PG node table (UI / list / alerts)
time-series status + connections CH of_node_edge_health (no message)

6. L3 Host: host_metrics

6.1 Where Collection Comes From

The Agent reads the local machine (e.g. /proc, disk stats), once per packet.

Field Semantics Description
cpu_usage_percent instant current CPU%
memory_* / storage_* instant used/total usage rates computed at Server or display layer
disk_read_bytes / disk_write_bytes cumulative counter kernel cumulative IO
network_rx_bytes / network_tx_bytes cumulative counter host NIC, not data provided

6.2 Report Example

"host_metrics": {
  "captured_at_unix": 1721289600,
  "cpu_usage_percent": 12.5,
  "memory_used_bytes": 4294967296,
  "memory_total_bytes": 16106127360,
  "storage_used_bytes": 50000000000,
  "storage_total_bytes": 107374182400,
  "disk_read_bytes": 9000000000,
  "disk_write_bytes": 12000000000,
  "network_rx_bytes": 500000000000,
  "network_tx_bytes": 800000000000
}

6.3 How the Server Handles Cumulative Fields

store raw-value time series
when displaying "NIC outbound over this period":
  delta = current - previous
  if delta < 0 → treat as restart/counter reset, record this segment's increment as 0, continue from new baseline
  if delta >= 0 → record into that period's increment
  • The Agent reports raw values, never computes 24h totals at the edge.
  • Forbidden to sum cumulative raw values as business volume.
  • Copy must be "host NIC", never "data provided / OpenResty outbound".

Stored in: of_node_metric_snapshots (optional capacity hourly MV).


7. One Complete Report Example

{
  "schema_version": 2,
  "node_id": "n_01hxyz",
  "name": "edge-shanghai-1",
  "ip": "203.0.113.10",
  "version": "3.4.0",
  "ext_version": "",
  "current_version": "20260718-abc",
  "last_error": "",
  "host_metrics": {
    "captured_at_unix": 1721289600,
    "cpu_usage_percent": 12.5,
    "memory_used_bytes": 4294967296,
    "memory_total_bytes": 16106127360,
    "storage_used_bytes": 50000000000,
    "storage_total_bytes": 107374182400,
    "disk_read_bytes": 9000000000,
    "disk_write_bytes": 12000000000,
    "network_rx_bytes": 500000000000,
    "network_tx_bytes": 800000000000
  },
  "edge_health": {
    "captured_at_unix": 1721289600,
    "status": "healthy",
    "message": "",
    "connections": 42
  },
  "access_logs": [
    {
      "logged_at_unix": 1721289595,
      "remote_addr": "198.51.100.20",
      "host": "www.example.com",
      "path": "/",
      "status_code": 200,
      "bytes_sent": 4096,
      "request_length": 200,
      "request_time_ms": 12
    }
  ],
  "buffered": [],
  "health_events": [],
  "waf_ip_group_checksums": {
    "1": "d41d8cd98f00b204e9800998ecf8427e"
  }
}

Server storage sketch:

payload block written to
access_logs[0] one CH row, bytes_sent=4096, region filled by GeoIP
edge_health node openresty_status=healthy, connections=42
host_metrics one CH metric row with cumulative/instant fields

Product query sketch (24h):

  • data provided = sum(bytes_sent) over that node's (or global) logs
  • current connections = latest edge_health.connections
  • host NIC outbound = sum of non-negative network_tx deltas over metrics

The three numbers need not be equal.


8. Offline Backfill buffered

When reporting fails, the Agent caches the same kind of facts locally by window (default ~60 minutes), then packs them into buffered[] after recovery:

"buffered": [
  {
    "captured_at_unix": 1721289500,
    "host_metrics": { },
    "edge_health": { },
    "access_logs": [ ]
  }
]
  • Only facts, no legacy TrafficReport.
  • Server processing logic is identical to the main fields.

9. End-to-End Timeline (default 3s)

t=0.0s   visitor request completes → writes one access.log line; connection count may change
t=0.1s   another request → another log line
…
t=3s     Agent heartbeat:
           · reads 2 access_logs lines
           · GET observability → connections=42
           · reads host_metrics
           · sends to Server
t=3s+    Server stores; dashboard/Zone queries aggregate logs
t=6s     next round…

10. Old Model Comparison

Old Approach New Model
Lua dict 60s window request_count + Agent 10s pull + Server sum removed; request count = log count
openresty_tx as "outbound" removed; data provided = sum(bytes_sent)
Two endpoints observability + stub_status data collection unified through observability; stub_status kept as liveness and internal read endpoint
TrafficReport pre-aggregation removed; no such path in protocol or API
Business and NIC both called "traffic" separate copy, separate APIs, separate tables
health status/message PG latest-state authority; CH only status+connection time series

11. Config and Implementation Index

Item Location/Key
Heartbeat interval Agent heartbeat_interval; control plane agent_heartbeat_interval (default 3000ms)
Offline threshold control plane node_offline_threshold (default 60000ms)
Observability port openresty_observability_port (default 18081)
access.log path access_log_path
Replay minutes observability_replay_minutes (default 60)
Protocol types pkg/protocol/agent.go (evolves to v2 when landing)
Table DDL observability-data-model.md

12. Revision History

Date Notes
2026-07-18 initial draft: single-page "latest transport model" — three layers, frequency, sample JSON, collection sources, old-model comparison
2026-07-18 default report interval 3s; offline threshold 60s; replay window 60 minutes
2026-07-18 M5: edge_health table, access_log_hourly, deprecate request_reports/obs_openresty throughput tables
2026-07-18 no compatibility layer: removed "may ignore during compat period" wording; health message only in PG, no message in CH