diff --git a/docs/reading-module-design.md b/docs/reading-module-design.md new file mode 100644 index 0000000..7706d7c --- /dev/null +++ b/docs/reading-module-design.md @@ -0,0 +1,743 @@ +# 阅读模块(Reading Module)设计与实施方案 + +> 状态:设计稿,待评审 +> 目标版本:v0.2.0(分期落地,见第 9 节) +> 关联现有子系统:媒体库 / 网盘存储 / 权限体系 / 任务队列 + +--- + +## 1. 目标与范围 + +### 1.1 已确认的产品决策 + +| 维度 | 决策 | +| --- | --- | +| 内容类型 | **电子书 + 漫画统一书架**(EPUB / TXT / PDF / MOBI 与 CBZ / CBR / 图片文件夹) | +| 书源 | **独立书库**(不复用影视媒体库)+ **网盘直链阅读** | +| 首版范围 | **完整版**:多用户书库权限 + 阅读统计 | +| 阅读形态 | **滚动流式与分页翻页双模式**,用户可切换并持久化偏好 | + +### 1.2 明确的非目标 + +- **不接入 Emby/Jellyfin 协议。** Emby 的 `Items` / `Views` / `PlaybackInfo` 语义围绕音视频构建,没有书籍章节与阅读进度的对应概念。强行映射会污染 `internal/service/emby_*.go` 与 `internal/handler/emby_*.go` 的既有兼容层,收益极低。阅读能力只通过 MeBox 自己的 Web UI 提供。 +- **不复用 `model.Library` / `LibraryRoot`。** `Library.Type` 的取值域是 `movie/tv/anime/music`,且被海报墙轮播(`CarouselEnabled`)、自动整理管线、Emby 视图、首页预览等链路消费。把书库塞进去会导致这些链路需要到处加 `type != "book"` 判断。 +- 首版不做:听书 TTS、在线书源(笔趣阁类)、社交分享、跨设备同步批注冲突合并。 + +--- + +## 2. 总体架构 + +### 2.1 分层落位 + +完全沿用现有分层,不引入新模式: + +``` +web/src/pages/Books*.tsx ← 页面 +web/src/components/Book*.tsx ← 阅读器与书架组件 +web/src/api/books.ts ← axios 封装(仿 web/src/api/library.ts) + ↓ /api/books/* +internal/handler/books*.go ← 反序列化 + 权限校验 + 响应 +internal/service/book_*.go ← 业务策略(扫描、解析、进度、统计) +internal/repository/book_*.go ← 纯持久化 +internal/model/book.go ← GORM 模型,注册进 model.AllModels() +``` + +新增路由注册走 `internal/handler/routes_authenticated_features.go` 的既有范式,新增一个 `registerAuthedBookRoutes(authed, svc)`,在 `registerAuthenticatedRoutes` 链上挂载。`service.Container` 与 `repository.Container` 各追加一个字段。 + +### 2.2 与现有能力的复用点 + +| 现有部件 | 复用方式 | +| --- | --- | +| `service.StreamService.ServeFile`(`internal/service/stream_file.go`) | 已用 `http.ServeContent` 处理 HEAD / Range / If-Modified-Since,**PDF 与原始文件流直接照搬这条路径** | +| `cloud.Provider.Resolve(ctx, fileRef) (*DirectLink, error)`(`internal/service/cloud/cloud.go`) | 网盘书源的直链解析入口,`DirectLink.Proxy` 决定 302 还是反代 | +| `model.StorageConfig`(`internal/model/storage_assistant.go`) | 直接复用为网盘书源的账号凭据载体,**不新建凭据表** | +| `service.ImageProxy`(`internal/service/image_proxy*.go`) | 漫画页与封面的磁盘缓存 + 远程拉取 + 缩放,复用其缓存目录与命名思路 | +| `service.PruneImageCache`(`internal/service/cache_cleanup.go`) | 现成的「按总大小做 LRU 淘汰」助手,书籍缓存淘汰直接复用它 | +| `service/scheduler_local_jobs.go` | 本地定时任务的挂载点,书籍缓存清理与每日统计汇总都注册在这里 | +| `config.CacheConfig`(`internal/config/types.go`) | 已有 `CacheDir` / `MaxDiskUsageMB` / `TTLHours` / `AutoCleanup` / `CleanupIntervalMin`,书籍缓存容量配置直接挂进去 | +| `service.FileManager`(`internal/service/filemanager.go`) | 本地书源目录浏览,前端复用 `LocalDirBrowserDialog.tsx` | +| `service.Scheduler` | 书库定时扫描(默认关闭,管理员可开) | +| `model.UserPermission` | 新增阅读权限位,见第 7 节 | +| `helper.Go` / `Container.stopCtx` | 后台扫描任务的生命周期管理 | + +--- + +## 3. 数据模型 + +新增文件 `internal/model/book.go`,并在 `internal/model/model.go` 的 `AllModels()` 中追加。所有表继承 `model.Base`(UUID 主键 + 时间戳 + 软删除)。 + +**表名约定**:`internal/model` 全包**没有任何 `TableName()` 覆盖**,一律使用 GORM 默认复数化(例如 `PlaybackHistory` → `playback_histories`,可从 `internal/database/schema_migration.go` 的裸 SQL 印证)。新表沿用该约定,不引入例外。因此模型命名要保证复数化结果干净: + +| 模型 | 表名 | +| --- | --- | +| `Book` | `books` | +| `BookLibrary` | `book_libraries` | +| `BookSource` | `book_sources` | +| `BookChapter` | `book_chapters` | +| `BookProgress` | `book_progresses` | +| `BookAnnotation` | `book_annotations` | +| `BookFavorite` | `book_favorites` | +| `BookReadingSession` | `book_reading_sessions` | +| `BookDailyStat` | `book_daily_stats` | + +(刻意用 `BookDailyStat` 而不是 `BookStatDaily`——后者复数化会得到 `book_stat_dailies`。) + +### 3.1 书库与书源 + +```go +// BookLibrary 是独立于影视媒体库的书库。 +type BookLibrary struct { + Base + Name string `gorm:"size:128;not null" json:"name"` + Kind string `gorm:"size:16;not null;default:mixed" json:"kind"` // ebook / comic / mixed + CoverURL string `gorm:"size:1024" json:"cover_url,omitempty"` + Enabled bool `gorm:"default:true" json:"enabled"` + SortOrder int `gorm:"index;default:0" json:"sort_order"` + LastScanAt *time.Time `json:"last_scan_at,omitempty"` + ScanStatus string `gorm:"size:16;default:idle" json:"scan_status"` // idle / scanning / error + ScanMessage string `gorm:"size:512" json:"scan_message,omitempty"` +} + +// BookSource 是书库下的一条挂载来源:本地目录或网盘路径。 +type BookSource struct { + Base + LibraryID string `gorm:"index;size:36;not null" json:"library_id"` + Name string `gorm:"size:128" json:"name,omitempty"` + StorageKind string `gorm:"size:16;not null;default:local" json:"storage_kind"` // local / cloud + Path string `gorm:"size:1024;not null" json:"path"` // 本地绝对路径 / 网盘内路径 + StorageConfigID string `gorm:"index;size:36" json:"storage_config_id,omitempty"` // 复用 model.StorageConfig + Depth int `gorm:"default:3" json:"depth"` // 扫描递归深度上限 + Enabled bool `gorm:"default:true" json:"enabled"` + SortOrder int `gorm:"default:0" json:"sort_order"` +} +``` + +`StorageKind = cloud` 时,`StorageConfigID` 指向一条 `StorageConfig`(`Type` ∈ `cloud115 / clouddrive2 / openlist / emby_remote`)。凭据解密沿用 `service.CryptoService`。 + +### 3.2 书籍与章节 + +```go +type Book struct { + Base + LibraryID string `gorm:"index;size:36;not null" json:"library_id"` + SourceID string `gorm:"uniqueIndex:uniq_book_source_path,priority:1;index;size:36;not null" json:"source_id"` + // SourcePath 在本地源是绝对路径,在网盘源是「网盘内路径」,两者都用 + // (source_id, source_path) 做唯一键,天然隔离两个 ID 空间。 + SourcePath string `gorm:"uniqueIndex:uniq_book_source_path,priority:2;size:1024;not null" json:"source_path"` + SourceRef string `gorm:"size:256" json:"source_ref,omitempty"` // 网盘 file id / pickcode + Title string `gorm:"size:512;not null" json:"title"` + Author string `gorm:"size:256;index" json:"author,omitempty"` + SeriesName string `gorm:"size:256;index" json:"series_name,omitempty"` + Volume int `json:"volume"` + Format string `gorm:"size:16;not null" json:"format"` // epub/txt/pdf/mobi/cbz/cbr/folder + MediaKind string `gorm:"size:16;not null;default:ebook" json:"media_kind"` // ebook / comic + SizeBytes int64 `json:"size_bytes"` + FileHash string `gorm:"index;size:64" json:"file_hash,omitempty"` // 大小+首尾采样,去重 + CoverURL string `gorm:"size:1024" json:"cover_url,omitempty"` + Description string `gorm:"type:text" json:"description,omitempty"` + Language string `gorm:"size:32" json:"language,omitempty"` + Tags string `gorm:"type:text" json:"tags,omitempty"` // 逗号分隔 + ChapterCount int `json:"chapter_count"` + WordCount int64 `json:"word_count"` + PageCount int `json:"page_count"` // 漫画总页数 / PDF 页数 + ParseStatus string `gorm:"size:16;default:pending" json:"parse_status"` // pending/ok/failed + ParseMessage string `gorm:"size:512" json:"parse_message,omitempty"` + NSFW bool `gorm:"default:false" json:"nsfw"` + AddedAt time.Time `json:"added_at"` +} +``` + +**唯一键说明**:`SourcePath` 上的 `uniqueIndex` 需与 `SourceID` 组成复合键(`uniq_book_source_path`,priority 1 = `source_id`)。同一本书被两个书源包含时允许重复入库,这是符合预期的(用户可能故意如此)。 + +```go +// BookChapter 只存索引,不存正文(见 3.4 的取舍)。 +type BookChapter struct { + Base + BookID string `gorm:"index:idx_book_chapter,priority:1;size:36;not null" json:"book_id"` + Index int `gorm:"index:idx_book_chapter,priority:2" json:"index"` + Title string `gorm:"size:512" json:"title"` + Level int `gorm:"default:1" json:"level"` // 目录嵌套层级,1 = 顶级 + // 电子书定位:二选一 + Href string `gorm:"size:1024" json:"href,omitempty"` // EPUB zip 内条目路径 + StartOffset int64 `json:"start_offset"` // TXT 字节区间 + EndOffset int64 `json:"end_offset"` + // 漫画/PDF 定位 + PageStart int `json:"page_start"` + PageEnd int `json:"page_end"` + CharCount int `json:"char_count"` +} +``` + +### 3.3 进度、批注、收藏、统计 + +```go +// BookProgress 每个用户每本书一行(复合唯一键,仿 model.PlaybackHistory 的 uniq_user_history 模式)。 +type BookProgress struct { + Base + UserID string `gorm:"uniqueIndex:uniq_user_book,priority:1;size:36;not null" json:"user_id"` + BookID string `gorm:"uniqueIndex:uniq_user_book,priority:2;size:36;not null" json:"book_id"` + ChapterIndex int `gorm:"default:0" json:"chapter_index"` + ChapterTitle string `gorm:"size:512" json:"chapter_title,omitempty"` + CharOffset int `json:"char_offset"` // 章内字符偏移(电子书) + PageIndex int `json:"page_index"` // 页码(漫画 / PDF) + Percent float64 `json:"percent"` // 全书百分比,书架进度条展示用 + ScrollRatio float64 `json:"scroll_ratio"` // 章内滚动比例,跨端还原更精确 + ReaderMode string `gorm:"size:16;default:scroll" json:"reader_mode"` // scroll / paged + Finished bool `gorm:"default:false" json:"finished"` + TotalSeconds int64 `json:"total_seconds"` + LastReadAt time.Time `gorm:"index" json:"last_read_at"` +} + +type BookAnnotation struct { + Base + UserID string `gorm:"index:idx_book_anno,priority:1;size:36;not null" json:"user_id"` + BookID string `gorm:"index:idx_book_anno,priority:2;size:36;not null" json:"book_id"` + ChapterIndex int `json:"chapter_index"` + Type string `gorm:"size:16;not null" json:"type"` // bookmark / highlight / note + StartOffset int `json:"start_offset"` + EndOffset int `json:"end_offset"` + SelectedText string `gorm:"size:2048" json:"selected_text,omitempty"` + Note string `gorm:"type:text" json:"note,omitempty"` + Color string `gorm:"size:16" json:"color,omitempty"` +} + +type BookFavorite struct { + Base + UserID string `gorm:"uniqueIndex:uniq_user_book_fav,priority:1;size:36;not null" json:"user_id"` + BookID string `gorm:"uniqueIndex:uniq_user_book_fav,priority:2;size:36;not null" json:"book_id"` +} + +// BookReadingSession 由前端心跳驱动,服务端按小时聚合,避免行数爆炸。 +type BookReadingSession struct { + Base + UserID string `gorm:"index:idx_book_stat,priority:1;size:36;not null" json:"user_id"` + BookID string `gorm:"index;size:36;not null" json:"book_id"` + BucketStart time.Time `gorm:"index:idx_book_stat,priority:2" json:"bucket_start"` // 截断到小时 + Seconds int64 `json:"seconds"` + CharsRead int64 `json:"chars_read"` + PagesRead int `json:"pages_read"` +} + +// BookDailyStat 每日汇总,供热力图与「年度阅读报告」查询,避免实时扫 session 表。 +type BookDailyStat struct { + Base + UserID string `gorm:"uniqueIndex:uniq_user_book_daily,priority:1;size:36;not null" json:"user_id"` + Day string `gorm:"uniqueIndex:uniq_user_book_daily,priority:2;size:10;not null" json:"day"` // YYYY-MM-DD + Seconds int64 `json:"seconds"` + Chars int64 `json:"chars"` + Pages int `json:"pages"` + Books int `json:"books"` // 当日有阅读记录的书数 +} +``` + +### 3.4 关键取舍:正文不入库 + +**决策:DB 只存章节索引(偏移量 / zip 内路径 / 页码区间),正文按需从源文件读取。** + +理由: +1. 网文 TXT 常见 5–50MB,漫画单册 100–800MB。入库会让 SQLite 单文件膨胀到数十 GB,直接冲击 `docker-compose.simple.yml` 的「单文件数据库好备份」定位,也会拖慢全库 VACUUM / 备份 / 数据库迁移(`internal/service/database_admin.go`)。 +2. 源文件本来就是权威副本,重复存储没有收益。 +3. EPUB 与 CBZ 本质上都是 zip,**随机读取 zip 内单个条目成本极低**(读中央目录 + 解压目标条目),不需要把整本解压落盘。 + +代价是每次打开章节都要读源文件。缓解手段: +- 本地源:`os.Open` + `io.SectionReader`,代价可忽略。 +- 网盘源:见 4.3 的本地缓存策略,且对已缓存的章节走本地。 + +### 3.5 用户级字段(挂在 `model.User` 上) + +沿用 `AllowedLibraryIDs` 的 JSON-in-text 模式(见 `internal/model/user.go`),**不复用影视库字段**,避免两个 ID 空间交叉: + +```go +// 追加到 model.User +ReaderSettings string `gorm:"type:text" json:"-"` // 阅读器偏好 JSON +AllowedBookLibraryIDs string `gorm:"type:text" json:"-"` // 空 = 不限制 +AllowedBookLibraryList []string `gorm:"-" json:"allowed_book_library_ids,omitempty"` +``` + +`ReaderSettings` 结构(前端读写,服务端仅透传与长度校验): + +```json +{ + "mode": "scroll|paged", + "fontSize": 18, + "lineHeight": 1.8, + "fontFamily": "serif|sans|custom", + "contentWidth": 720, + "theme": "light|sepia|dark|black", + "pageAnimation": "slide|fade|none", + "comicLayout": "single|double|auto", + "comicDirection": "ltr|rtl", + "hideScrollbar": true +} +``` + +放在 `User` 行内(而非新表)的理由:与 `PlayerVolume` / `DanmakuFontSize` 等既有播放器偏好一致,读取时随用户信息一并返回,无需额外查询。 + +--- + +## 4. 书源与内容读取管线 + +### 4.1 扫描流程 + +``` +POST /api/books/libraries/:id/scan + → BookScannerService.ScanLibrary(ctx, libraryID) + 1. 置 ScanStatus=scanning,通过 SSEHub 广播进度(复用 service.SSEHub) + 2. 遍历启用的 BookSource + - local: filepath.WalkDir,按扩展名白名单过滤,超过 Depth 停止递归 + - cloud: cloud.New(cfg.Type, cfg, client).List(ctx, dirID) 递归列目录 + 3. 对每个候选文件调 BookParser.ParseMeta(reader) 拿元信息 + 目录 + 4. Upsert 到 books / book_chapters(source_id + source_path 为幂等键) + 5. 源上已消失的书标记软删除(与影视库扫描语义保持一致) + 6. 置 ScanStatus=idle,记录 LastScanAt +``` + +扩展名白名单:`.epub .txt .pdf .mobi .azw3 .cbz .cbr .zip .rar`(`.zip/.rar` 仅当目录内全是图片时按漫画处理,否则跳过,防止误吞压缩包)。 + +并发:复用 `internal/service` 现有的 worker 池写法,默认 2–4 并发解析(解析要读文件,IO 密集)。 + +### 4.2 各格式解析策略 + +| 格式 | 元信息 | 章节 / 页 | 正文读取 | +| --- | --- | --- | --- | +| **EPUB** | zip → `META-INF/container.xml` → OPF → `dc:title/dc:creator/dc:language/dc:description`;封面取 OPF `meta[name=cover]` 指向项,退化到 `guide` | 按 spine 顺序,标题取每个 XHTML 的 `