Skip to content

fix(catalog): 慢网不再因为 8s 元数据超时被判成离线(Mod 列表拿不到) - #81

Merged
std-microblock merged 1 commit into
std-microblock:masterfrom
std-external:fix/slow-network-catalog
Oct 8, 2026
Merged

std-microblock merged 1 commit into
std-microblock:masterfrom
std-external:fix/slow-network-catalog

Conversation

@std-external

Copy link
Copy Markdown
Contributor

现象

新版本给「下载 Mod 元数据」加了一个 8 秒的等待上限后,网速慢的用户在 8 秒内拿不到 Mod 列表(catalog):

  • 启动加载时直接进入「模组目录离线模式」横幅(App.tsx 的 catalog-offline),只能手动点「重新连接」;
  • 本地没有缓存时(首次使用 / 清过缓存),管理页的 Mod 列表拿不到元数据,等于空白。

根因

src-tauri/src/everest.rs(master 913c980):

  1. fetch_catalog_interruptible() 在第 139 行用 wait_for_catalog(receiver, Duration::from_secs(8), ...) 等待,而底层的 fetch_raw_catalog() 自己还有 20 秒的 ureq 总超时(第 124 行)。8 秒到了以后,wait_for_catalog 直接 bail!("Mod catalog download timed out; continuing offline"),正在跑的 HTTP 请求被丢弃(注释明确写了「a skipped request must never publish a late response」),也就是说慢网再等 1 秒本来就能成功的响应被扔掉了。
  2. 这个超时被当成网络故障:load_catalog() 的 Err 分支(第 260-269 行)会 CATALOG_OFFLINE.store(true),永久锁死离线模式——此后所有取 catalog 的调用都走 offline_catalog(),直到用户手动点「重新连接」(或 set_catalog_offline(false))才解除。没有缓存的用户因此整局都看不到元数据,也不会自动恢复。
  3. 20 秒的 HTTP 总超时本身对慢网也偏紧。实测 https://celeste.weg.fan/api/v2/mod/list:gzip 后 1.0 MB,解压后 5.8 MB(6265 个 Mod)。按 ureq 的总超时算,20 秒要求持续 ~290 KB/s 的解压速率,8 秒更是要求 ~700 KB/s;慢网用户每次都失败,也就永远拿不到新数据。

改动

后端 src-tauri/src/everest.rs

  • 8 秒只当「先不等了」,不当失败:wait_for_catalog() 改为返回三态 CatalogWait { Finished / StillDownloading / Skipped };超时返回 StillDownloading,fetch_catalog_interruptible() 用 CatalogDownloadPending 错误类型表达「还在下,不是坏网」。
  • 迟到但有效的响应不再被丢弃:worker 线程在成功时把响应写进磁盘缓存(save_raw_cache),前提是请求没被取消(request_is_still_wanted():没进离线模式、generation 没变)。这样 8 秒后转入后台下载的请求一落地,下一次读 catalog(或前端的重试)就能拿到新数据,下次启动也直接命中缓存。
  • 不再把「还在下」当成离线:load_catalog() 的错误分支里,只有真正的失败才会 CATALOG_OFFLINE.store(true);CatalogDownloadPending 直接返回手头的缓存(offline_catalog()),不锁离线模式。取消/跳过(用户点「跳过下载,离线继续」、切离线、重新连接)仍然什么都不写、什么都不 latch。
  • HTTP 总超时 20s → 60s(CATALOG_FETCH_TIMEOUT):首屏已经不阻塞在这个请求上,超时只用来兜底「服务器不吐数据」,所以慢网该有的时间就给它。worker 失败时也会补 latch 离线模式(原来只有等在调用方的那个线程会 latch,慢网情况下调用方早就返回了),避免对不可达的服务器反复重试。
  • 新增 is_catalog_downloading() 及 Tauri 命令 is_mod_catalog_downloading:前端可以知道「后台还在下」,据此等它落地后自动刷新,而不是把它当成离线。

前端

  • 新增纯逻辑模块 src/catalogDownload.ts:waitForCatalogDownload() 轮询上述状态直到下载结束(3s × 20 次,和 60s 的后端超时对齐),可注入依赖、便于单测。
  • src/api/modCatalog.ts:loadModCatalog() 失败时,若是「后台还在下」且没进离线模式,则等下载结束后重试一次(原来的 8 秒超时会让管理页/搜索页直接拿到空目录)。已缓存快路径(新鲜的 catalog / catalogPromise)完全没变。
  • src/context/modManage.tsx:启动加载完成后,如果 catalog 还在后台下载,等它落地再 reloadMods() 一次,Mod 行的版本/更新信息会自动补上,不需要用户点「重新连接」。

测试

  • src-tauri/src/everest.rs 单测:更新 wait_for_catalog 的 4 个既有用例(超时不再是错误),新增「慢下载不能切成离线」「只有真失败才 latch 离线」「迟到的响应只在请求仍有效时才发布」。
  • 新增 src/celemod-ui/src/catalogDownload.test.ts(node:test,和仓库里其它前端测试一致):无下载立刻返回、慢下载等到落地、超预算后放弃而不是永久挂住。

验证

# Rust(需要 nightly,CI 的构建工具链就是 nightly)
cd src-tauri && cargo test --lib everest

# 前端
cd src/celemod-ui && node --import tsx --test src/catalogDownload.test.ts
cd src/celemod-ui && npx tsc --noEmit

结果:everest 8 个用例全过;catalogDownload.test.ts 3/3;tsc --noEmit 无错误。

慢网行为(人工推演,供 review 参考):

场景 修改前 修改后
有旧缓存 + 列表 12s 才下完 8s 放弃 → 弹「离线模式」横幅,缓存里的旧数据被标成 stale,永不刷新 8s 先用缓存渲染(不弹横幅),后台继续下,落地后自动刷新
无缓存 + 列表 12s 才下完 8s 报错 + latch 离线,管理页空白,必须手动「重新连接」 8s 时只是首屏先渲染,下载落地后前端重试一次即可拿到完整列表
服务器真的不可达 8s(或首个错误)后 latch 离线 不变:真失败才 latch 离线,避免反复重试

不改下载链路(src-tauri/src/ureq.rs)的任何超时——本次只处理元数据这条。

来源:Celeste Coders 群里的反馈(「新版本给下载 mod 元数据加了 8s 的超时,有的人网慢 8s 下不完」)。

The 8s wait in fetch_catalog_interruptible treated a merely slow
connection as an offline server:

- the response that was still on its way was dropped, even though the
  HTTP request itself may run for its own 20s timeout
- load_catalog latched CATALOG_OFFLINE on that timeout, so every later
  catalog read served the cache (or failed) until the user manually
  pressed "reconnect"

Keep the 8s deadline as "do not block the first screen any longer"
instead of a failure:

- wait_for_catalog now reports Finished/StillDownloading/Skipped
- a download that outlives the wait keeps running and publishes its
  response to the on-disk cache when it lands (unless the request was
  cancelled, in which case nothing is written)
- load_catalog only latches offline mode for real failures, and serves
  the cached catalog meanwhile
- the catalog request timeout is 60s: the Mod list decompresses to
  ~5.8MB (6265 Mods), so 20s required ~290KB/s and 8s ~700KB/s
- new is_mod_catalog_downloading command lets the UI wait for a running
  download; loadModCatalog retries once after it settles and the startup
  flow refreshes the Mod list, so slow users end up with the metadata
  instead of an empty list
@std-microblock
std-microblock merged commit 14dae83 into std-microblock:master Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants