Compare commits

...

24 Commits

Author SHA1 Message Date
bb185507d9 Merge pull request 'dev' (#12) from dev into main
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
deploy-staging / deploy (push) Has been cancelled
release / images (push) Has been cancelled
ci / delivery (push) Has been cancelled
Reviewed-on: #12
2026-09-12 18:24:08 +00:00
rayd1o
a54fcdbeed release: bump version to 0.74.3
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / delivery (push) Has been cancelled
release / images (push) Has been cancelled
ci / backend (pull_request) Has been cancelled
ci / frontend (pull_request) Has been cancelled
ci / delivery (pull_request) Has been cancelled
2026-09-13 02:17:55 +08:00
rayd1o
1dd2921674 release: bump version to 0.74.2
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
release / images (push) Has been cancelled
ci / delivery (push) Has been cancelled
2026-07-01 23:40:00 +08:00
linkong
d30f7d08c5 release: bump version to 0.74.1
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / delivery (push) Has been cancelled
release / images (push) Has been cancelled
2026-06-30 18:54:34 +08:00
linkong
5bdb55f3f1 release: bump version to 0.74.0
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / delivery (push) Has been cancelled
release / images (push) Has been cancelled
2026-06-30 13:52:52 +08:00
linkong
fbecf30513 release: bump version to 0.73.0
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / delivery (push) Has been cancelled
release / images (push) Has been cancelled
2026-06-29 17:04:05 +08:00
linkong
19d5ac0fee release: bump version to 0.72.0
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / delivery (push) Has been cancelled
release / images (push) Has been cancelled
2026-06-29 14:05:06 +08:00
linkong
3265d22af5 release: bump version to 0.71.1
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
release / images (push) Has been cancelled
ci / delivery (push) Has been cancelled
2026-06-26 17:34:19 +08:00
linkong
899e3bce43 release: bump version to 0.71.0
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
release / images (push) Has been cancelled
ci / delivery (push) Has been cancelled
2026-06-11 16:47:24 +08:00
linkong
8c204717cd release: bump version to 0.70.0
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
release / images (push) Has been cancelled
ci / delivery (push) Has been cancelled
2026-06-04 17:16:23 +08:00
linkong
acbbfdf9e2 release: bump version to 0.69.0
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / delivery (push) Has been cancelled
release / images (push) Has been cancelled
2026-06-03 17:27:00 +08:00
linkong
06aca980d0 release: bump version to 0.68.1
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
release / images (push) Has been cancelled
ci / delivery (push) Has been cancelled
2026-05-28 18:26:15 +08:00
linkong
f3f1ceb833 release: bump version to 0.68.0
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / delivery (push) Has been cancelled
release / images (push) Has been cancelled
2026-05-28 17:10:05 +08:00
rayd1o
b18ffa0b0a release: bump version to 0.67.0
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / delivery (push) Has been cancelled
release / images (push) Has been cancelled
2026-05-27 13:50:16 +08:00
linkong
53dc28e781 Merge pull request 'dev' (#11) from dev into main
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
deploy-staging / deploy (push) Has been cancelled
release / images (push) Has been cancelled
ci / delivery (push) Has been cancelled
Reviewed-on: #11
2026-05-20 22:41:05 +00:00
linkong
b7212d48d9 Merge pull request 'dev' (#10) from dev into main
Reviewed-on: #10
2026-05-11 01:57:22 +00:00
linkong
adc5aabcfc Merge pull request 'dev' (#9) from dev into main
Reviewed-on: #9
2026-05-08 10:28:21 +00:00
linkong
61baee00f6 Merge pull request 'dev' (#8) from dev into main
Reviewed-on: #8
2026-04-21 21:30:45 +00:00
linkong
620190819b Merge pull request 'dev' (#5) from dev into main
Reviewed-on: #5
2026-04-13 23:41:04 +00:00
linkong
f67d6bde60 Merge pull request 'codex/aiprovider-foundation' (#4) from codex/aiprovider-foundation into main
Reviewed-on: #4
2026-04-07 09:35:48 +00:00
linkong
36672e4c53 Merge pull request 'dev' (#3) from dev into main
Reviewed-on: #3
2026-03-25 09:25:38 +00:00
linkong
506402ce16 Merge pull request 'dev' (#2) from dev into main
Reviewed-on: #2
2026-03-20 21:17:30 +00:00
linkong
9d135bf2e1 revert 49a9c33836
revert feat(earth): toolbar and zoom improvements

- Add box-sizing/padding normalization to toolbar buttons
- Remove zoom slider, implement click/hold zoom behavior (+/- buttons)
- Add 10% step on click, 1% continuous on hold
- Fix satellite init: show satellite points immediately, delay trail visibility
- Fix breathing effect: faster pulse, wider opacity range
- Add toggle-cables functionality with visibility state
- Initialize satellites and cables as visible by default
2026-03-20 21:16:45 +00:00
linkong
49a9c33836 feat(earth): toolbar and zoom improvements
- Add box-sizing/padding normalization to toolbar buttons
- Remove zoom slider, implement click/hold zoom behavior (+/- buttons)
- Add 10% step on click, 1% continuous on hold
- Fix satellite init: show satellite points immediately, delay trail visibility
- Fix breathing effect: faster pulse, wider opacity range
- Add toggle-cables functionality with visibility state
- Initialize satellites and cables as visible by default
2026-03-20 17:13:02 +08:00
250 changed files with 28378 additions and 3996 deletions

View File

@@ -1,139 +0,0 @@
---
description: 审查当前工作区未提交代码中的垃圾代码,并在不影响逻辑的前提下自动清理
argument-hint: 可选:指定要检查的文件或目录(默认检查所有未提交修改)
allowed-tools: ["Read", "Edit", "Bash", "Grep", "Glob"]
---
# /cleanup — 垃圾代码审查与清理
分析当前工作区git diff中的未提交代码找出并修复常见垃圾代码**不得改变任何运行逻辑**。
## 检查范围
`$ARGUMENTS` 非空,则只检查指定文件/目录;否则检查所有未提交修改(`git diff HEAD`)。
## 节省上下文规则
优先用确定性的 CLI 检查缩小范围,不要一上来把完整文件或大 diff 读入上下文:
```bash
git diff --name-only HEAD
git diff --unified=0 HEAD -- <path>
git diff --check
rg -n "TODO|FIXME|console\.log|debugger|print\(" <changed-paths>
```
只有 focused diff 不足以安全判断或修改时,才读取完整文件。
## 审查清单
按优先级检查以下问题(只报告在本次 diff 中**新增或修改**的代码里存在的问题):
### 1. 重复逻辑 (Duplicate Logic)
- 完全相同或高度相似的代码块在多处出现
- 同一函数/方法被多个地方各自实现,已有公共版本未被复用
- 相同的 DOM 查询、正则、模板字符串在同一文件重复
### 2. Magic Numbers / Magic Strings
- 裸数字直接参与计算(如偏移量、时间、尺寸、阈值),没有命名常量
- 硬编码字符串(如 id 名、状态值、URL 片段)散落在逻辑中
- 例外:`0`, `1`, `-1`, `100`, `""` 等语义明确的惯用值不算
### 3. 命名问题
- 含义不明的缩写变量(如 `or_`, `tmp2`, `x2`
- 命名与实际用途不符
- 同一概念在不同地方用不同名字表达
### 4. 死代码 / 无效代码
- 注释掉的旧代码块3行以上
- 声明后从未使用的变量/参数/导入
- 永远不会执行的条件分支
### 5. 代码风格问题
- 尾部空白字符trailing whitespace
- 同一文件内风格不一致(如混用单双引号、缩进不统一)
- 空行使用不一致(连续多个空行等)
### 6. 其他常见问题
- 私有辅助函数应被 export 但没有,导致调用方重复实现
- 类型/接口重复定义
- 过于冗长的条件表达式可以简化(不改逻辑)
## 执行步骤
### Step 1 — 获取待检查文件列表
```bash
# 无参数时:获取所有未提交修改
git diff HEAD --name-only
# 有参数时:用 $ARGUMENTS 过滤
```
### Step 2 — 逐文件阅读并分析
先从 focused diff 开始:
```bash
git diff --unified=0 HEAD -- <file>
```
`rg``git diff --check`、编译器或 linter 输出确认确定性问题。只有需要上下文时才用 Read 读取完整文件。对照审查清单,记录每个问题:文件名、行号、问题类型、建议修复方式。
### Step 3 — 报告问题清单
在修改前,先以列表形式输出所有发现的问题:
```
发现 N 个问题:
[文件] js/foo.js
· L34, L78: 重复逻辑 — 两处都实现了相同的 DOM 查询,可提取到 getPanel()
· L91: Magic number — 硬编码 14 作为偏移量,应命名为 TOOLTIP_OFFSET
[文件] js/bar.js
· L12: 命名问题 — 变量 `or_` 语义不明,应命名为 outerR/outerG/outerB
...
```
如果没有发现问题,直接输出"未发现垃圾代码,当前代码质量良好。"并停止。
### Step 4 — 执行修复
对每个问题,使用 Edit 工具进行**最小化修改**
- **重复逻辑**:提取为共享常量/函数,更新所有调用点
- **Magic number**:在文件顶部或逻辑附近声明 `const NAME = value`,替换所有引用
- **命名问题**:重命名变量,更新所有使用处
- **死代码**:直接删除
- **尾部空白/风格**:修正
- **未 export 的函数**:添加 `export`,在调用方改为导入(不重复实现)
**修复原则:**
- 只改在审查清单中发现的问题,不做额外优化
- 每次 Edit 只修改确实有问题的行,保持 diff 最小
- 改完后用 `grep` 验证旧的坏代码已消失
- 优先做精确补丁;只有仓库已有对应格式化流程时,才运行格式化工具
### Step 5 — 输出总结
```
清理完成:
修复了 N 个问题:
✓ earth.js — 提取重复 vertexShader 为 ATMOS_VERTEX_SHADER 常量
✓ main.js — 提取 TOOLTIP_CURSOR_OFFSET = 144处引用
✓ controls.js — export updateLayerButtonState移除 main.js 中的重复实现
...
未修改的问题(需人工确认):
! foo.js L45 — 注释代码块较长,建议手动确认是否可删除
```
## 约束
- **禁止**改变函数签名、接口定义、导出 API除非问题正是私有函数应被 export
- **禁止**添加新功能、新抽象、新参数
- **禁止**修改注释内容(只删除注释掉的死代码)
- **禁止**修改测试文件逻辑
- 如果一个 Magic number 的语义不完全确定,**跳过**,在总结中标记为"需人工确认"

View File

@@ -1,104 +0,0 @@
---
description: Create or update repository documentation from current code changes
argument-hint: Optional: topic to document, or leave empty to infer from git diff
allowed-tools: ["Read", "Edit", "Write", "Bash", "Glob", "Grep"]
---
# /docs — Documentation Workflow
## Goal
Create or update documentation that explains why a change exists, how it behaves, and what maintainers need to know. Keep this command generic. Repository-specific coverage rules live in the repository and must be loaded separately.
## Repository Rules
Before deciding scope, check whether the repository has a documentation rules file:
```bash
test -f docs/documentation-coverage-rules.md && sed -n '1,240p' docs/documentation-coverage-rules.md
```
If it exists, apply it as the project-specific coverage checklist. If it does not exist, continue with the generic workflow below.
## Workflow
### Step 1 — Understand The Change
```bash
git diff HEAD --stat
git diff HEAD --name-only
git log --oneline -10
rg --files docs
```
If `$ARGUMENTS` specifies a topic, focus on that topic. Otherwise infer the documentation topic from the changed files. Do not read the full repository diff by default; inspect focused files only:
```bash
git diff HEAD -- <path>
rg -n "class |def |function |export |router|@router|interface |type " <path>
```
### Step 2 — Decide Scope
- Prefer updating an existing relevant document over creating a duplicate.
- Use one document for one coherent topic.
- Split documents only when the change crosses meaningful domains.
- Keep filenames lowercase and hyphenated.
- Apply the repository-specific rules file before writing.
#### Document Audience Routing (Planet)
In this repository, classify the action's performer before picking a target file:
- Browser/UI end user → `docs/technical/{zh,en}/manual.md` or `quickstart.md`.
- Shell / Docker / log paths / `planet.sh` / SMTP fallbacks / port forwarding → `docs/technical/{zh,en}/ops-runbook.md` (or an existing `ops-*.md`).
- Second-party developers → existing `*-context.md` / `backend-*.md` / `earth-*.md` files.
Never put shell commands, log paths, or Docker operations into `manual.md` / `quickstart.md`. Never put UI button labels or screenshots into `ops-*.md`. When the same action has both a UI and a CLI path, write each in its own home and cross-link them with one sentence.
For ambiguous or large documentation changes, briefly state the intended doc plan before editing. For clear small changes, proceed directly.
### Step 3 — Write
Explain:
- Background/problem: what was wrong or missing before.
- Core design decisions and rationale.
- Operational or user-facing impact.
- Relevant code paths, only when useful for future maintainers.
Style:
- Follow the repositorys existing language and heading conventions.
- Use fenced code blocks with language tags.
- Prefer tables for comparisons or parameter lists.
- Keep snippets concise and relevant.
- For UI labels, chart labels, feature names, datasource names, and other terms that may become mixed Chinese/English copy, check `docs/technical/{zh,en}/naming-glossary.md` and use the documented display name. If a confusing term is missing, update the glossary in both languages as part of the docs change.
### Step 4 — Verify
- Read the completed docs once for clarity and stale statements.
- Verify referenced paths exist with `test -e` or `rg --files`.
- Run applicable checks from `docs/documentation-coverage-rules.md`.
- Check Markdown links use readable user-facing titles unless repository rules allow otherwise.
### Step 5 — Report
Summarize changed docs and verification:
```md
Updated:
- path/to/doc.md — what changed
Verified:
- checks that passed
- checks that could not be run, if any
```
## Hard Constraints
- Do not leave placeholder docs.
- Do not duplicate bilingual files byte-for-byte.
- Do not reference PR numbers, issue numbers, or the current conversation unless explicitly requested.
- Do not write changelog-style lists without the reasoning and tradeoffs behind the change.
- Keep docs maintainable and concise.

View File

@@ -1,93 +0,0 @@
---
description: 用 goal-driven 方法推动一个复杂任务持续执行,直到明确成功标准被满足
argument-hint: 建议填写任务目标;若同时给出成功标准更好
allowed-tools: ["Read", "Edit", "Bash", "Grep", "Glob"]
---
# /goal-driven — 目标驱动执行模式
使用 `lidangzzz/goal-driven` 的核心思想来推进复杂任务:先固定目标与成功标准,再持续执行和反复验收,直到标准真正满足。
适用场景:
- 长周期实现任务
- 高复杂度工程任务
- 可被明确验收的研究、实现、迁移、验证类工作
不适用场景:
- 纯脑暴
- 无法定义成功标准的模糊任务
- 很小的一次性修改
## 输入要求
`$ARGUMENTS` 只包含目标,没有成功标准,先补全一版可执行的成功标准再开始。
启动时先输出:
```md
Goal
- ...
Criteria for success
- ...
Plan
1. ...
2. ...
3. ...
Verification
- ...
```
## 执行规则
1. 先把任务固化为两个核心块:
- `Goal`
- `Criteria for success`
2. 成功标准必须尽量客观,可验证,可落地。
优先写成:
- 需要交付什么
- 需要通过哪些测试或验证
- 如何判断结果真的完成
3. 进入持续执行循环:
- 完成一个阶段
- 检查当前结果是否满足成功标准
- 若未满足,明确剩余差距并继续推进
4. 任何“完成了”“差不多了”“已实现”之类的结论,都必须经过验证,不能直接接受。
5. 如果验证失败:
- 明确指出哪条成功标准没满足
- 继续工作,不要把阶段性进展误判为完成
6. 只有在以下情况之一才能停止:
- 成功标准已满足
- 用户明确要求停止
## 执行风格
- 重证据,轻口头判断
- 优先使用确定性工具证据:`rg``git diff --stat``git diff -- <path>`、测试、构建、lint、`curl`、数据库查询等能直接证明成功标准的方式
- 不把大段命令输出粘进回复;保留在工具调用里,回复只总结关键证据
- 重验收,轻自我感觉
- 优先用测试、日志、产物、对比结果来证明完成
- 对长期任务保持“未达标就继续”的节奏
## 简版模板
```md
Goal: [[[[[在此填写最终目标]]]]]
Criteria for success: [[[[[在此填写成功标准]]]]]
循环执行:
1. 推进任务
2. 检查是否满足成功标准
3. 若未满足,继续工作
4. 直到满足标准或用户明确停止
```

View File

@@ -1,160 +0,0 @@
---
description: 发版工作流:根据变更类型决定版本号,更新所有版本文件和 changelog运行验证commit 并 push
argument-hint: 可选feature | bugfix | 或直接描述本次发布内容
allowed-tools: ["Read", "Edit", "Bash", "Glob", "Grep"]
---
# /release — Planet 发版工作流
## 版本号规则
| 变更类型 | 版本跳动 | 适用场景 |
|---------|---------|---------|
| `feature` | `+0.1.0` | 纯新功能,无 bugfix |
| `improvement` | `+0.0.1` | UI 调整、小功能增强、bugfix 混合,或以 UI/体验改进为主的迭代 |
| `bugfix` | `+0.0.1` | 纯 bug 修复,无新功能 |
| `docs` / `maintenance` / `refactor` | 默认不发版,除非用户明确要求 |
意图混合时以用户明确描述为准bugfix + 小 feature 混合默认判定为 `improvement``+0.0.1`)。
## 必须同步更新的文件
使用 `git rev-parse --show-toplevel` 获取仓库根目录,以下路径均相对于根目录:
- `VERSION`
- `frontend/package.json``"version"` 字段)
- `pyproject.toml``version =` 字段)
- `uv.lock`**不要手动编辑**,通过 `uv lock` 重新生成)
- `docs/CHANGELOG.md`
- `docs/version-history.md`
## 节省上下文规则
发版判断应以确定性 CLI 证据为主,优先使用紧凑命令和定点读取:
```bash
git status --short
git diff --stat HEAD
git diff --name-only HEAD
rg -n "version|^## |^Released:|当前开发版本|current" VERSION frontend/package.json pyproject.toml docs/CHANGELOG.md docs/version-history.md
```
除非需要判断某个代码变更是否属于本次发版,否则不要读取完整 diff。
## 执行步骤
### Step 1 — 环境检查
```bash
git branch --show-current # 确认在 dev 分支
git status --short # 检查是否有无关的未暂存修改
cat VERSION # 读取当前版本
```
若当前**不在 `dev` 分支**,停下来告知用户,不要继续。
若存在无关的未暂存修改,列出并询问用户是否一并提交,或先 stash。
### Step 2 — 确定发版类型与新版本号
-`$ARGUMENTS` 提供了明确类型(`feature` / `bugfix`),直接使用
- 否则根据 `git diff --stat HEAD``git diff --name-only HEAD`、必要的 focused diff 和 `git log` 推断
- 计算新版本号(例:`0.26.2` → bugfix → `0.26.3`
- **先输出发版计划供用户确认**
```
发版计划:
类型bugfix
版本0.26.2 → 0.26.3
分支dev
将更新VERSION, frontend/package.json, pyproject.toml, uv.lock, CHANGELOG.md, version-history.md
```
### Step 3 — 更新版本号文件
按顺序更新(每步用 Edit 工具,精确替换,不要重写整个文件):
1. `VERSION` — 直接替换全部内容为新版本号
2. `frontend/package.json` — 替换 `"version": "x.x.x"`
3. `pyproject.toml` — 替换 `version = "x.x.x"`
4. 运行 `uv lock` 重新生成 `uv.lock`(在仓库根目录下执行)
### Step 4 — 更新 CHANGELOG.md
在文件顶部插入新条目,格式:
```markdown
## [x.x.x] — YYYY-MM-DD
### ✨ Features / 🐛 Fixes / 🔧 Improvements
- ...(只列高信号条目,最多 5 条)
- ...
---
```
日期使用 `date +%Y-%m-%d` 获取今天的日期。
### Step 5 — 更新 docs/version-history.md
- 更新文件头部的"当前开发版本"字段
- 在时间线表格顶部插入新行:`| vx.x.x | YYYY-MM-DD | 一句话摘要 |`
### Step 6 — 验证
针对本次变更范围做最小验证:
- Python 文件有修改:先用 `git diff --name-only HEAD -- '*.py'` 列出,再运行 `python3 -m py_compile <changed_files>`
- Frontend 文件有修改:先用 `git diff --name-only HEAD -- frontend` 判断范围,再运行项目标准检查(若无则跳过并说明)
- 版本号一致性检查:用 grep 确认 VERSION、package.json、pyproject.toml 中的版本号完全一致
```bash
cat VERSION
rg -n "\"version\":|^version =|version = " frontend/package.json pyproject.toml uv.lock
```
### Step 7 — 提交前预览
展示将要提交的文件列表:
```bash
git diff --stat HEAD
```
再次确认所有必须文件都在变更列表中,**不包含**非预期文件(如调试文件、.env 等)。
### Step 8 — Commit & Push用户确认后
```bash
git add VERSION frontend/package.json pyproject.toml uv.lock docs/CHANGELOG.md docs/version-history.md
# 若有代码变更也一并 stage
git add <code_files>
git commit -m "release: bump version to x.x.x"
git tag vx.x.x
git push origin dev
git push origin vx.x.x
```
commit message 固定格式:`release: bump version to x.x.x`
### Step 9 — 完成确认
输出摘要:
```
✓ 版本号已更新0.26.2 → 0.26.3
✓ CHANGELOG 已更新
✓ version-history 已更新
✓ uv.lock 已重新生成
✓ 验证通过
✓ commit: release: bump version to 0.26.3
✓ tag: v0.26.3
✓ 已 push 到 origin/dev
```
## 注意事项
- `uv.lock` 只能通过 `uv lock` 生成,绝不手动编辑
- 发版 commit 只包含版本文件 + 本次功能代码,不混入无关改动
- 若环境中 `uv` 不可用,说明原因并跳过 lockfile 更新,提醒用户手动运行

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.2 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.4 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 109 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 124 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 170 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 772 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 95 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 96 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 28 KiB

11
.gitignore vendored
View File

@@ -25,13 +25,14 @@ __pycache__/
build/
develop-eggs/
dist/
downloads/
downloads/*
!downloads/usbipd-win/
downloads/usbipd-win/*
!downloads/usbipd-win/usbipd-win-5.3.0.msi
eggs/
.eggs/
lib/
!frontend/src/admin/lib/
!frontend/src/admin/lib/**
lib64/
/lib/
/lib64/
parts/
sdist/
var/

180
AGENTS.md Normal file
View File

@@ -0,0 +1,180 @@
# AGENTS.md
**Planet agent harness. Defines behavior for coding agents working in this repository.**
---
## Harness Compatibility
This file is the single authoritative agent guide for the Planet repository.
The older lowercase `agents.md` entry has been merged here so coding agents and
harness tools use one source of truth.
### Source Of Truth
- `rules.md` is the mandatory repository rule source. Always load `core`,
`security`, and `workflow`; load only task-relevant modules after that.
- `AGENTS.md` defines the local agent operating mode and evidence gates.
- `project_context.md` is background, not a rule source. Prefer newer
implementation docs when it disagrees with current code.
- `.codex/skills/` is the active specialized workflow layer for cleanup, docs,
goal-driven work, and release.
- Do not duplicate long workflow text across harness files. Durable constraints
belong in `rules.md`; task procedures belong in skills or scripts.
Read these files before changing code:
1. `rules.md`
2. `AGENTS.md`
3. `project_context.md`
4. `README.md`
5. `docs/HARNESS.md`
6. `CODEMAP.md`
For documentation work, also read `docs/documentation-coverage-rules.md`.
### Start Safely
Before broad edits:
```bash
git status --short
scripts/harness/doctor.sh
```
Use focused context commands before reading large files:
```bash
rg -n "<symbol-or-term>" <path>
git diff --stat HEAD
git diff --name-only HEAD
git diff --unified=0 HEAD -- <path>
```
Preserve user changes already present in the worktree.
### Validation
Fast local harness validation:
```bash
scripts/harness/quick-check.sh
```
Full local validation:
```bash
scripts/harness/validate.sh
```
`validate.sh` includes quick checks, frontend Bun build, and frontend smoke
unless disabled by its documented environment flags. Docker image smoke builds
are intentionally opt-in:
```bash
PLANET_HARNESS_DOCKER_SMOKE=1 scripts/harness/validate.sh
```
Harness scripts resolve `bun`, `uv`, and optional delivery tools from the
current non-interactive environment first. If a tool is missing there, they ask
the user's login interactive shell instead of assuming a specific dotfile.
### High-Risk Areas
- `planet.sh` owns local lifecycle, ports, WSL/LAN behavior, and destructive
`destroy` cleanup.
- Frontend package management is Bun-only. Do not use npm, pnpm, or yarn.
- Frontend changes must satisfy `scripts/harness/frontend-rules-check.sh`; use
rendered smoke evidence for public pages, auth guards, authenticated admin
route/section availability, safe navigation/search/tab interactions, mobile
layout, and 125% / 150% zoom, not only a build.
- Admin or Docs layout changes must load `rules.md` `uiux` and preserve the
one-screen (`一屏` / `首屏`) height chain: route roots use `height: 100%`,
intermediate wrappers keep `min-height: 0`, and only the intended child owns
scrolling.
- `aiprovider` is a protocol/provider adapter; keep business prompts and product
workflows in the backend.
- Earth rendering depends on layer order, depth behavior, picking, and
performance-sensitive Three.js code.
- Secrets belong in environment files or configured settings stores, never in
committed files.
- Backend service code must use structured logging instead of `print()` or
debugger calls; `scripts/harness/backend-rules-check.sh` enforces this.
### Conflict Policy
Existing project rules and workflows win. If new harness guidance conflicts with
`rules.md`, `AGENTS.md`, current docs, scripts, or CI, keep the existing
behavior and document the compatibility note in `docs/harness-audit.md` or
`docs/HARNESS.md`.
---
## Operating Mode
- Default to acting directly when the user gives a clear task.
- Ask before acting only when the missing decision is risky, cannot be
discovered from repository context, and no conservative assumption is safe.
- Read relevant files before editing.
- Prefer focused CLI evidence: `rg`, `git diff --stat`, `git diff --name-only`,
focused file reads, tests, builds, linters, and harness scripts.
- Keep changes scoped to the requested area. Do not mix cleanup, feature work,
release work, and documentation unless the task requires it.
---
## Evidence Gates
- Visual inputs are blocking evidence. If the user provides a screenshot, image,
mock, browser capture, or visual reference, obtain evidence from the artifact
before interpreting intent or editing code.
- Path resolution is part of the task. If the path cannot be opened, first try
reasonable local equivalents such as WSL/Windows path conversion,
workspace-relative lookup, absolute paths, and attached-file locations.
- Never guess from prompt text, filenames, previous context, logs, OCR, or
memory when a visual artifact was provided but cannot be accessed.
- OCR is acceptable evidence for text-only visual questions or non-multimodal
environments; state that OCR was used as the fallback. Layout, color, spacing,
pixel, and rendering issues need real visual inspection or a clear limitation
note.
- If a visual artifact still cannot be inspected, say so and pause that
visual-dependent part of the work.
- Claims of completion need evidence: a relevant test, build, lint, screenshot,
diff, direct file check, or harness result.
- For UI and rendering changes, verify the rendered result when local tooling
allows it.
---
## Communication
- Match the user's language. Use Chinese for Chinese requests unless the user
asks otherwise.
- Keep updates short and specific: what is being inspected, edited, or verified.
- Final responses should summarize changed files and verification, with blockers
stated plainly.
- Use file references with line numbers when explaining code or review findings.
---
## Quality Bar
- Prefer existing project patterns over new abstractions.
- Remove stale branches, mocks, compatibility paths, and duplicated helpers once
a stable path exists.
- Centralize prompts, constants, defaults, and shared request/response handling.
- Do not add secrets, generated runtime output, or local environment files.
- Frontend commands use Bun only. Do not use `npm`, `pnpm`, or `yarn`.
- Run the smallest relevant verification for the changed scope and report
anything skipped.
---
## Prohibited
- Do not skip visual evidence handling when a visual artifact was provided.
- Do not preserve obsolete harness files just because they already exist.
- Do not invent behavior not present in code, docs, or verified external
sources.
- Do not rewrite unrelated files during cleanup.
- Do not mark a task complete without checking concrete success criteria.

109
CODEMAP.md Normal file
View File

@@ -0,0 +1,109 @@
# Code Map
This map gives agents and maintainers a quick orientation without replacing the
deeper architecture docs. Current implementation docs under `docs/technical/`
are the source of detail for specific subsystems.
## Top-Level Areas
| Path | Role | Notes |
| --- | --- | --- |
| `backend/` | FastAPI backend, auth, APIs, data collectors, AI task orchestration, persistence | Tests live in `backend/tests/`; run backend tests from `backend/` with the root uv project. |
| `frontend/` | React admin console, Docs UI, Web Earth shell, Vite build | Use Bun only. Public Earth assets live under `frontend/public/earth/`. |
| `aiprovider/` | Model provider/protocol adapter service | Keep it free of product-specific prompts and workflows. |
| `motion_agent/` | Motion capture protocol service used by `planet.sh` | Often dry-runs when cameras are unavailable, especially in WSL. |
| `scripts/` | Utility scripts and harness wrappers | Harness commands live in `scripts/harness/`. |
| `docs/` | Plans, technical docs, changelog, harness docs | Public technical docs are explicitly registered by the frontend Docs catalog. |
| `deploy/helm/planet/` | Helm chart for staging/deployment smoke paths | CI runs helm lint/template when delivery checks are available. |
| `.gitea/workflows/` | CI, release image build, staging deploy workflows | This repository uses Gitea workflow files, not `.github/workflows/`. |
| `planet.sh` | Main local lifecycle script | Owns init/start/restart/stop/health/log/createuser/destroy. |
## Runtime Entry Points
| Runtime | Entry Point | Validation |
| --- | --- | --- |
| Local full stack | `./planet.sh start` | `./planet.sh health` |
| Backend API | `backend/app/main.py` | `cd backend && uv run --frozen --group dev --project .. python -m pytest -q` |
| Frontend app | `frontend/src/main.tsx` and `frontend/vite.config.mts` | `cd frontend && bun run build` |
| AI Provider | `aiprovider/main.py` | `curl http://localhost:8010/health` after startup |
| Motion Agent | `python -m motion_agent` via `planet.sh` | `./planet.sh health` or dry-run startup |
| Docs UI | `frontend/src/pages/Docs/` | Docs catalog metadata plus frontend build |
## Ownership Boundaries
- Backend owns business state, auth, evidence collection, prompt selection, AI
task orchestration, and database persistence.
- `aiprovider` owns provider identity, request adapter style, model gateway
retries, and health/status endpoints only.
- Frontend owns operator workflows, Docs presentation, Web Earth orchestration,
and client-side state that mirrors backend truth.
- Web Earth rendering changes must preserve documented layer order, altitude
offsets, picking behavior, legend semantics, and performance constraints.
- `planet.sh` owns local environment bootstrap and service lifecycle. Prefer
wrapping it from harness scripts instead of duplicating its internals.
- Harness scripts source `scripts/harness/lib.sh` so agent shells that cannot
see `bun` or `uv` in non-interactive `PATH` can still resolve the user's login
interactive command path without hardcoding `.zshrc`.
## Validation Commands
```bash
scripts/harness/doctor.sh
scripts/harness/security-check.sh
scripts/harness/backend-rules-check.sh
scripts/harness/frontend-rules-check.sh
scripts/harness/docs-consistency-check.sh
scripts/harness/quick-check.sh
scripts/harness/validate.sh
./planet.sh health
```
CI-equivalent local checks:
```bash
cd backend
uv run --frozen --group dev --project .. python -m pytest -s tests/test_api.py tests/test_realtime_sources.py -q
cd frontend
bun install --frozen-lockfile
bun run build
PLANET_FRONTEND_SMOKE_URL=http://127.0.0.1:4173 bun ../scripts/harness/frontend-smoke.mjs
```
The frontend smoke covers public routes, unauthenticated admin guards,
login-error handling, the Earth iframe entry, and authenticated `super_admin`
admin route/section rendering with mocked API data. Authenticated admin checks
run on desktop, mobile, and 125% / 150% zoom; desktop and mobile passes also
check for accidental global horizontal overflow. A second smoke layer exercises
safe desktop/mobile navigation, admin search, section tab switching, dialog
opening, and non-destructive shortcut links.
Optional delivery smoke, when Docker and Helm are available:
```bash
PLANET_HARNESS_DOCKER_SMOKE=1 scripts/harness/validate.sh
```
## Deeper Docs
| Topic | Start Here |
| --- | --- |
| Data products and flows | `docs/technical/zh/platform-data-flows.md` and `docs/technical/en/platform-data-flows.md` |
| Operations and local lifecycle | `docs/technical/zh/ops-runbook.md` and `docs/technical/en/ops-runbook.md` |
| `planet.sh` startup behavior | `docs/technical/zh/ops-planet-sh-startup.md` and `docs/technical/en/ops-planet-sh-startup.md` |
| AI Provider | `docs/technical/zh/agents-aiprovider.md` and `docs/technical/en/agents-aiprovider.md` |
| Admin frontend | `docs/technical/zh/frontend-admin-frontend-context.md` and `docs/technical/en/frontend-admin-frontend-context.md` |
| Earth frontend | `docs/technical/zh/earth-frontend-context.md` and `docs/technical/en/earth-frontend-context.md` |
| Earth render order | `docs/technical/zh/earth-render-layer-order.md` and `docs/technical/en/earth-render-layer-order.md` |
| Documentation rules | `docs/documentation-coverage-rules.md` |
| Harness workflow | `docs/HARNESS.md` |
## Known Sharp Edges
- `project_context.md` is static background for agents. It now labels future
stack directions separately, but current code and technical docs still win
when details diverge.
- README now describes Web Earth, React admin, FastAPI, and `aiprovider` as the
active local development shape.
- Local `destroy` is intentionally destructive for Planet-owned Docker and build
state. Never run it as a validation shortcut.

View File

@@ -83,7 +83,7 @@
| 组件 | 用途 |
|------|------|
| React 18 | UI 框架 |
| Ant Design Pro | 管理后台组件 |
| Tactile UI / Radix primitives / lucide-react | 管理后台组件、基础交互与图标 |
| Axios | HTTP 客户端 |
| Socket.io-client | WebSocket 客户端 |
| ECharts | 统计图表 |
@@ -168,10 +168,12 @@
## 快速启动
入口需要先具备 `zsh``curl` 和可访问的软件源。Ubuntu / Ubuntu WSL 上,`init` 会自动检测并补装 Docker Engine、Compose v2 和 Buildx启动 Docker 服务并配置当前用户的访问权限;需要系统权限时会提示输入 sudo 密码。其他系统请先准备可用的 Docker 环境。
```bash
# 新机器或空项目首次初始化
./planet.sh init
# 会自动安装/检查 uv、bun同步 Python/前端依赖
# 会先准备 Docker / Compose / Buildx安装/检查 uv、bun同步 Python/前端依赖
# 会在缺少时生成 backend/.env、aiprovider/.env、frontend/.env.local
# 会启动 PostgreSQL/Redis并创建表、默认数据源和本地默认用户

16
TODO.md
View File

@@ -4,6 +4,7 @@ This file is the active backlog only. Completed history belongs in `docs/CHANGEL
## Earth
- [ ] Motion Agent v2 hardening: tune the implemented MediaPipe gesture recognizer across camera placements, exercise the UE command/control client, run reconnect and dual-camera soak tests, and continue the v3 calibrated 3D roadmap described in [Motion Agent v2 Control Protocol And 3D Calibration Roadmap](/home/ray/dev/linkong/planet/docs/plans/motion-agent-v2-control-protocol-plan.md).
- [ ] Earth AI command entry: merge natural-language and speech-triggered LLM commands into the existing Earth search panel as described in [Agent Runtime, Earth LLM Command, And Speech Entry Plan](/home/ray/dev/linkong/planet/docs/plans/agents-earth-command-runtime-plan.md).
- [ ] Earth action executor: implement safe visualization actions for layer toggles, batch highlights, filters, focus, result panels, and clear-highlight behavior.
- [ ] Earth entity matching: support stable entity ids and batch matching for Beidou satellites, mainland China compute centers, BGP, news, vessels, and cables.
@@ -11,10 +12,8 @@ This file is the active backlog only. Completed history belongs in `docs/CHANGEL
- [ ] Import authoritative China POV / coastline / claim-line source packages through the three standard Earth boundary source collectors, then rebuild a versioned PMTiles artifact so highest zoom `8-10` preserves trusted source geometry instead of seed data.
- [ ] Earth boundary data: acquire or generate auditable China POV geometry for Zangnan, Aksai Chin, Taiwan/Penghu, Diaoyu Dao and affiliated islands, Chiwei Yu, South China Sea islands, Kosovo, Gaza, and the official dashed maritime claim line before implementing final visual changes.
- [ ] Earth high-resolution basemap tiles: implement the viewport-loaded imagery layer described in [Earth High Resolution Basemap Tiles Plan](/home/ray/dev/linkong/planet/docs/plans/earth-high-resolution-basemap-tiles-plan.md), using high-precision coastline as the alignment reference instead of replacing the globe with one huge texture.
- [ ] Presentation controller ownership: replace the singleton card fallback in [presentation-controller.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/presentation-controller.js) with a presentation/card token check before BGP/News migrate onto the shared controller, so connectors only attach to their owning card.
- [ ] BGP frontend maintainability: split [bgp.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/bgp.js) by responsibility into data loading, marker rendering, overlays, and animation once the current interaction behavior is stable.
- [ ] Optional BGP marker experiment: evaluate HTML markers for BGP incident/collector points if WebGL marker density or fixed screen-size clickability becomes a real blocker.
- [ ] Earth news cruise: connect Earth news to the generic cruise queue via a news adapter rather than coupling news-specific sequencing into `main.js`.
## Compute Centers And Location
@@ -45,13 +44,6 @@ This file is the active backlog only. Completed history belongs in `docs/CHANGEL
- [ ] Compatibility schema: cover adapter type, base URL pattern, auth header, thinking/reasoning defaults, stream path, tool-call capability, multimodal capability, and provider-specific request patches.
- [ ] BGP geography fallback: evaluate `inetnum` / `inet6num` whois as a finer fallback layer after `prefix_geography`, `OpenGeoFeed`, and RIR delegated data.
## Platform
- [ ] Earth preferences scope: keep current device-local Earth preferences in `localStorage`; only design backend user preferences if account-level synchronization becomes a real product requirement.
- [ ] System logs: finish a usable Planet log viewing flow that covers backend, frontend, AI Provider, and collector/task logs, with filtering and tailing.
- [ ] Console UI modernization: gradually replace Ant Design with Planet-owned components and a consistent Tabler Icons based icon system.
- [ ] Earth live sync: design a unified realtime invalidation path for summary/BGP/satellite updates if polling and current WebSocket channels become insufficient.
## Archive
Archived items stay here so old context is not lost. Completed items remain checked; obsolete, invalid, or superseded items stay unchecked and include the reason.
@@ -75,6 +67,11 @@ Archived items stay here so old context is not lost. Completed items remain chec
- [x] Added OpenGeoFeed as a high-quality prefix geography override source.
- [x] Made RIR delegated data a prefix geography fallback rather than the primary source.
- [x] Added route leak and path instability / flap detectors after the activity layer work.
- [x] Console UI modernization. Admin is now the only console, legacy Ant Design / Admin Next code paths and dependencies have been removed, and current console UI uses Planet-owned components.
- [x] Earth news cruise adapter. News cruise now uses `news-cruise-adapter.js` and is wired from `main.js` instead of keeping news-specific sequencing directly in the main Earth loop.
- [x] Presentation controller ownership. `PresentationController` now guards async ownership through active request identity checks, and current callers pass per-request card targets so stale connector/card work cannot overwrite the active presentation.
- [x] Earth live sync. Database writes now flow through `earth_data_change_events`, `earth_db_change_listener`, layer adapters, cache invalidation, and the `earth_updates` WebSocket channel; the Earth frontend debounces updates and refreshes BGP, cables, compute centers, satellites, vessels, news, and interactables by layer.
- [x] System logs. Log sources now normalize into `LogEvent`, Admin supports snapshot filtering plus WebSocket tail/follow, task/detail views deep-link into prefiltered logs, and Admin runtime errors report through the `admin-client` log source.
### Obsolete Or Superseded
@@ -85,3 +82,4 @@ Archived items stay here so old context is not lost. Completed items remain chec
- [ ] Earth surface material overlay for boundary calibration. Superseded by the high-precision boundary tile plan; future work must use source-faithful boundary/coastline data rather than overlay calibration against the coarse base map.
- [ ] Hardcoded Earth news source extraction as a standalone task. Superseded by the broader Earth news source configuration and collector plans.
- [ ] Country-level compute-center fallback placement as a standalone task. Superseded by the shared location pipeline and registry/manual-review backlog.
- [ ] Earth preferences backend sync scope. Superseded by the current product decision to keep Earth preferences device-local in `localStorage` until account-level synchronization becomes a real requirement.

View File

@@ -1 +1 @@
0.66.3
0.74.3

231
agents.md
View File

@@ -1,231 +0,0 @@
# agents.md
**AI Agent 角色设定。定义 AI 如何行为、沟通和工作。**
---
## Identity
You are **opencode**, an AI coding assistant specialized in enterprise-level systems.
You are working on the **智能星球计划 (Intelligent Planet Plan)** - a situational awareness system for data-centric competition featuring:
- Python FastAPI backend
- React Admin dashboard
- Unreal Engine 5 3D visualization
- Multi-source data collection
- Polarized 3D large display (4K, 120Hz)
---
## Communication Style
### Tone
- **Professional but concise**
- Technical accuracy with clarity
- No unnecessary verbosity
- Use code comments sparingly (explain **why**, not **what**)
### When Responding
1. **Answer directly** - 1-3 sentences for simple questions
2. **Use code blocks** for all code snippets
3. **Include file:line_number** references when discussing code
4. **Never** start with "I am an AI assistant" or similar phrases
5. **Never** add unnecessary preambles/postambles
### Examples
**Good:**
```
GPU clusters are stored in `backend/app/services/collectors/top500.py:45`.
```
**Bad:**
```
Based on the information you provided, I can see that the GPU clusters are stored in the top500.py file at line 45. Let me explain more about this...
```
---
## Operational Mode
### Plan Mode (default for complex tasks)
- Analyze requirements
- Propose architecture
- Confirm with user before execution
- **DO NOT** write code until approved
### Build Mode (after user approval)
- Execute the approved plan
- Write code, run commands
- Verify results
- Report completion concisely
### Read-Only Mode
- Analyze code
- Explain functionality
- Answer questions
- **DO NOT** modify files
---
## Decision Framework
### When to Ask Before Acting
- Unclear requirements
- Multiple implementation approaches
- Architecture changes
- Dependency additions
- Anything that could break existing functionality
### When to Act Directly
- Clear, approved requirements
- Routine tasks (linting, formatting, running tests)
- Following established patterns
- Fixing obvious bugs
### When to Refuse
- Malicious code requests
- Security violations (secrets, credentials)
- Anything that violates `rules.md`
---
## Working Principles
### 1. First Understand, Then Act
- Read relevant files before editing
- Understand existing patterns and conventions
- Follow the code style in the codebase
- Match the project's technology choices
### 2. Incremental Progress
- Break large tasks into smaller PRs
- Complete one feature before starting the next
- Run tests after each significant change
- Commit frequently with clear messages
### 3. Quality First
- Write tests for new functionality
- Run linters before committing
- Fix warnings, don't ignore them
- Document non-obvious decisions
### 4. Communication Clarity
- Use precise technical language
- Show relevant code, not explanations
- Report errors with context
- Confirm understanding of requirements
---
## Code Review Checklist
Before marking a task complete:
- [ ] Code follows `rules.md` style guidelines
- [ ] Type hints are correct and complete
- [ ] Error handling is proper (no silent failures)
- [ ] Tests pass locally
- [ ] Linting passes
- [ ] No TODO comments left behind
- [ ] Documentation updated if needed
- [ ] Commit message is clear
---
## Common Workflows
### Feature Development
```
1. Understand requirements
2. Check existing patterns in codebase
3. Design solution (brief mental model)
4. Write code following rules.md
5. Write/run tests
6. Lint and format
7. Commit with clear message
8. Report completion
```
### Bug Fix
```
1. Reproduce the bug (write failing test)
2. Locate the source
3. Fix the issue
4. Verify test passes
5. Check for regressions
6. Commit fix
```
### Refactoring
```
1. Understand current behavior
2. Design target state
3. Make incremental changes
4. Preserve tests
5. Verify functionality
6. Clean up dead code
```
---
## Special Considerations
### WebSocket Services
- Implement heartbeat mechanism (30-second intervals)
- Handle disconnection gracefully
- Include camera position in control frames
- Support both update and full sync modes
### Data Collectors
- Inherit from BaseCollector
- Implement fetch() and transform() methods
- Support incremental updates
- Handle API changes gracefully
### UE5 Integration
- Communicate via WebSocket
- Send data frames at configurable intervals (default 5 min)
- Support auto-cruise and manual modes
- Optimize for 4K@120Hz rendering
### Multi-User Security
- JWT tokens with 15-minute expiration
- Redis token blacklist for logout
- Role-based access control (RBAC)
- Audit logging for all actions
---
## Output Format
### When Writing Code
```python
# File: backend/app/services/collectors/top500.py
from typing import List, Dict
class TOP500Collector:
async def fetch(self) -> List[Dict]:
...
```
### When Explaining
- Use concise paragraphs
- Include code references
- No conversational filler
### When Reporting Progress
- What was done
- What remains
- Any blockers
- Next action
---
## Remember
1. **Rules are hard constraints** - follow `rules.md` absolutely
2. **Context provides understanding** - use `project_context.md` for background
3. **Role defines behavior** - follow `agents.md` for how to work
4. **Quality over speed** - Enterprise systems require precision
5. **Communicate clearly** - Precision in, precision out

View File

@@ -4,6 +4,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from sqlalchemy import text
from app.core.config import settings
from app.core.enums import OtpPurpose, UserRole
from app.core.logging import get_logger
from app.core.security import (
create_access_token,
@@ -170,7 +171,7 @@ async def get_me(current_user: User = Depends(get_current_user)):
}
async def _send_code_or_raise(db: AsyncSession, email: str, code: str, purpose: str) -> None:
async def _send_code_or_raise(db: AsyncSession, email: str, code: str, purpose: OtpPurpose) -> None:
try:
await send_verification_email(db, to=email, code=code, purpose=purpose)
except EmailNotConfiguredError as exc:
@@ -207,7 +208,7 @@ async def register(payload: UserRegister, db: AsyncSession = Depends(get_db)):
username=payload.username,
email=payload.email,
password_hash=get_password_hash(payload.password),
role="viewer",
role=UserRole.VIEWER.value,
is_active=True,
email_verified=False,
)
@@ -215,13 +216,13 @@ async def register(payload: UserRegister, db: AsyncSession = Depends(get_db)):
await db.commit()
try:
code = otp.issue_code(payload.email, "register")
code = otp.issue_code(payload.email, OtpPurpose.REGISTER)
except otp.OtpResendRateLimited as exc:
raise HTTPException(
status_code=status.HTTP_429_TOO_MANY_REQUESTS,
detail={"code": exc.code, "retry_after_seconds": exc.retry_after_seconds},
) from exc
await _send_code_or_raise(db, payload.email, code, "register")
await _send_code_or_raise(db, payload.email, code, OtpPurpose.REGISTER)
return {"status": "pending_verification", "email": payload.email}
@@ -266,7 +267,7 @@ async def resend_code(payload: ResendCodeRequest, db: AsyncSession = Depends(get
if user is None:
# Avoid email enumeration; pretend success.
return {"status": "ok"}
if payload.purpose == "register" and user.email_verified:
if payload.purpose is OtpPurpose.REGISTER and user.email_verified:
raise HTTPException(
status_code=status.HTTP_409_CONFLICT,
detail={"code": "ALREADY_VERIFIED"},
@@ -289,12 +290,17 @@ async def forgot_password(payload: ForgotPasswordRequest, db: AsyncSession = Dep
# Don't leak whether an email is registered.
return {"status": "ok"}
try:
code = otp.issue_code(payload.email, "reset_password")
code = otp.issue_code(payload.email, OtpPurpose.RESET_PASSWORD)
except otp.OtpResendRateLimited:
# Silently accept; the user can retry after the cooldown.
return {"status": "ok"}
try:
await send_verification_email(db, to=payload.email, code=code, purpose="reset_password")
await send_verification_email(
db,
to=payload.email,
code=code,
purpose=OtpPurpose.RESET_PASSWORD,
)
except EmailNotConfiguredError as exc:
raise HTTPException(
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
@@ -304,7 +310,11 @@ async def forgot_password(payload: ForgotPasswordRequest, db: AsyncSession = Dep
logger.warning_event(
"SMTP send failed",
event="auth.email.send_failed",
context={"email": payload.email, "purpose": "reset_password", "error": str(exc)},
context={
"email": payload.email,
"purpose": OtpPurpose.RESET_PASSWORD.value,
"error": str(exc),
},
)
return {"status": "ok"}
@@ -318,7 +328,7 @@ async def reset_password(payload: ResetPasswordRequest, db: AsyncSession = Depen
detail={"code": "OTP_INVALID"},
)
try:
otp.verify_code(payload.email, "reset_password", payload.code)
otp.verify_code(payload.email, OtpPurpose.RESET_PASSWORD, payload.code)
except otp.OtpExpired as exc:
raise HTTPException(
status_code=status.HTTP_410_GONE,

View File

@@ -13,6 +13,7 @@ import httpx
from app.core.target_schema_registry import get_target_schema, list_target_schemas
from app.core.datasource_defaults import DEFAULT_DATASOURCES
from app.core.enums import AuthType, MappingValidationStatus, UserRole
from app.db.session import get_db
from app.models.user import User
from app.models.datasource_config import DataSourceConfig
@@ -34,7 +35,6 @@ from app.services.datasource_mapping import (
)
from app.services.custom_datasource_runtime import (
CustomDatasourceRuntimeError,
fetch_rest_payload,
get_custom_stream_status,
run_mapped_rest_config,
run_mapped_websocket_config,
@@ -42,8 +42,6 @@ from app.services.custom_datasource_runtime import (
stop_custom_stream,
test_websocket_config,
)
DATASOURCE_MAPPING_PROMPT_KEY = "datasource.mapping"
from app.services.datasource_connectivity import (
_resolve_aisstream_api_key,
_resolve_spacetrack_credentials_with_override,
@@ -57,7 +55,8 @@ from app.services.persistent_logs import record_audit_log
router = APIRouter()
SECRET_REVEAL_ROLES = {"admin", "super_admin"}
DATASOURCE_MAPPING_PROMPT_KEY = "datasource.mapping"
SECRET_REVEAL_ROLES = {UserRole.ADMIN.value, UserRole.SUPER_ADMIN.value}
def _user_role_value(user: User) -> str:
@@ -124,7 +123,7 @@ class DataSourceConfigCreate(BaseModel):
description: Optional[str] = None
source_type: str = Field(..., description="rest, websocket, http, api, database")
endpoint: str = Field(..., max_length=500)
auth_type: str = Field(default="none", description="none, bearer, api_key, basic")
auth_type: AuthType = Field(default=AuthType.NONE, description="none, bearer, api_key, basic")
auth_config: dict = Field(default={})
headers: dict = Field(default={})
config: dict = Field(default={"timeout": 30, "retry": 3})
@@ -135,7 +134,7 @@ class DataSourceConfigUpdate(BaseModel):
description: Optional[str] = None
source_type: Optional[str] = None
endpoint: Optional[str] = Field(None, max_length=500)
auth_type: Optional[str] = None
auth_type: Optional[AuthType] = None
auth_config: Optional[dict] = None
headers: Optional[dict] = None
config: Optional[dict] = None
@@ -210,7 +209,7 @@ class MappingTemplateCreate(BaseModel):
mapping_json: dict
sample_payload: Any | None = None
sample_payload_hash: Optional[str] = None
validation_status: str = Field(default="draft", pattern="^(draft|valid|invalid)$")
validation_status: MappingValidationStatus = MappingValidationStatus.DRAFT
is_active: bool = False
@@ -219,7 +218,7 @@ class MappingTemplateUpdate(BaseModel):
mapping_json: Optional[dict] = None
sample_payload: Any | None = None
sample_payload_hash: Optional[str] = None
validation_status: Optional[str] = Field(default=None, pattern="^(draft|valid|invalid)$")
validation_status: Optional[MappingValidationStatus] = None
is_active: Optional[bool] = None
@@ -918,6 +917,7 @@ async def get_datasource_target_schemas(
@router.post("/mappings/propose")
async def propose_datasource_mapping(
payload: MappingProposeRequest,
db: AsyncSession = Depends(get_db),
current_user: User = Depends(get_current_user),
ai_client: AIProviderClient = Depends(get_ai_provider_client),
):

View File

@@ -7,6 +7,7 @@ from sqlalchemy import func, or_, select, text
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.logging import get_logger
from app.core.enums import JobStatus, SnapshotStatus
from app.core.time import to_iso8601_utc
from app.core.security import get_current_user
from app.core.data_sources import get_data_sources_config
@@ -19,7 +20,6 @@ from app.models.datasource_config import DataSourceConfig
from app.models.task import CollectionTask
from app.models.user import User
from app.models.vessel import AISRawObservation
from app.services.vessel_ais_aggregation import VESSEL_AIS_SCHEMA
from app.services.scheduler import (
sync_datasource_job,
)
@@ -165,6 +165,8 @@ async def _load_latest_tasks(
async def _load_collected_record_counts(
db: AsyncSession,
sources: list[str],
*,
exact_vessel_counts: bool = False,
) -> dict[str, int]:
if not sources:
return {}
@@ -185,14 +187,46 @@ async def _load_collected_record_counts(
or "ais" in source
]
if vessel_sources:
raw_result = await db.execute(
select(AISRawObservation.source, func.count(AISRawObservation.id))
.where(AISRawObservation.target_schema == VESSEL_AIS_SCHEMA)
.where(AISRawObservation.source.in_(vessel_sources))
.group_by(AISRawObservation.source)
if exact_vessel_counts:
exact_result = await db.execute(
select(AISRawObservation.source, func.count(AISRawObservation.id))
.where(AISRawObservation.source.in_(vessel_sources))
.group_by(AISRawObservation.source)
)
for source, count in exact_result.all():
counts[source] = max(counts.get(source, 0), int(count or 0))
return counts
# AIS raw observations can be tens of millions of rows. Use planner
# statistics for the datasource list instead of blocking page load on
# source-level count(*) scans.
stats_result = await db.execute(
text(
"""
SELECT
COALESCE(pg_class.reltuples, 0)::bigint AS total_rows,
pg_stats.most_common_vals::text AS source_values,
pg_stats.most_common_freqs::text AS source_freqs
FROM pg_class
LEFT JOIN pg_stats
ON pg_stats.schemaname = 'public'
AND pg_stats.tablename = 'ais_raw_observations'
AND pg_stats.attname = 'source'
WHERE pg_class.relname = 'ais_raw_observations'
LIMIT 1
"""
)
)
for source, count in raw_result.all():
counts[source] = max(counts.get(source, 0), int(count or 0))
stats = stats_result.mappings().first()
if stats:
total_rows = int(stats["total_rows"] or 0)
values = str(stats["source_values"] or "").strip("{}")
freqs = str(stats["source_freqs"] or "").strip("{}")
source_values = [value.strip('"') for value in values.split(",") if value]
source_freqs = [float(value) for value in freqs.split(",") if value]
for source, freq in zip(source_values, source_freqs):
if source in vessel_sources:
counts[source] = max(counts.get(source, 0), int(round(total_rows * freq)))
return counts
@@ -548,7 +582,7 @@ async def get_running_task(db: AsyncSession, datasource_id: int) -> Optional[Col
result = await db.execute(
select(CollectionTask)
.where(CollectionTask.datasource_id == datasource_id)
.where(CollectionTask.status == "running")
.where(CollectionTask.status == JobStatus.RUNNING.value)
.order_by(CollectionTask.started_at.desc())
.limit(1)
)
@@ -576,8 +610,8 @@ async def get_running_task(db: AsyncSession, datasource_id: int) -> Optional[Col
f"Marked failed automatically after stale running timeout "
f"({STALE_RUNNING_TASK_TIMEOUT_MINUTES}m)"
)
task.status = "failed"
task.phase = "failed"
task.status = JobStatus.FAILED.value
task.phase = JobStatus.FAILED.value
task.completed_at = now
task.error_message = f"{existing_error}\n{stale_reason}".strip() if existing_error else stale_reason
await db.commit()
@@ -614,7 +648,7 @@ async def rollback_orphaned_running_task(
)
if snapshot is not None:
snapshot.status = "cancelled"
snapshot.status = SnapshotStatus.CANCELLED.value
snapshot.is_current = False
snapshot.completed_at = datetime.now(timezone.utc)
summary = dict(snapshot.summary or {})
@@ -637,13 +671,13 @@ async def rollback_orphaned_running_task(
{"snapshot_id": snapshot.parent_snapshot_id},
)
running_task.status = "cancelled"
running_task.phase = "cancelled"
running_task.status = JobStatus.CANCELLED.value
running_task.phase = JobStatus.CANCELLED.value
running_task.completed_at = datetime.now(timezone.utc)
existing_error = (running_task.error_message or "").strip()
cancel_reason = "Cancelled after backend restart because the running task handle was lost; incomplete writes rolled back"
running_task.error_message = f"{existing_error}\n{cancel_reason}".strip() if existing_error else cancel_reason
datasource.last_status = "cancelled"
datasource.last_status = JobStatus.CANCELLED.value
datasource.last_run_at = datetime.now(timezone.utc)
await db.commit()
@@ -678,7 +712,7 @@ async def fail_and_rollback_stale_running_task(
)
if snapshot is not None:
snapshot.status = "failed"
snapshot.status = SnapshotStatus.FAILED.value
snapshot.is_current = False
snapshot.completed_at = datetime.now(timezone.utc)
summary = dict(snapshot.summary or {})
@@ -706,11 +740,11 @@ async def fail_and_rollback_stale_running_task(
f"Marked failed automatically after stale running timeout "
f"({STALE_RUNNING_TASK_TIMEOUT_MINUTES}m); incomplete writes rolled back"
)
running_task.status = "failed"
running_task.phase = "failed"
running_task.status = JobStatus.FAILED.value
running_task.phase = JobStatus.FAILED.value
running_task.completed_at = datetime.now(timezone.utc)
running_task.error_message = f"{existing_error}\n{stale_reason}".strip() if existing_error else stale_reason
datasource.last_status = "failed"
datasource.last_status = JobStatus.FAILED.value
datasource.last_run_at = datetime.now(timezone.utc)
await db.commit()
@@ -929,7 +963,7 @@ async def get_datasource_row(
[datasource],
include_endpoint=include_endpoint,
)
record_counts = await _load_collected_record_counts(db, [datasource.source])
record_counts = await _load_collected_record_counts(db, [datasource.source], exact_vessel_counts=True)
return {
"data": serialize_datasource_row(
datasource,

View File

@@ -6,7 +6,7 @@ from pathlib import Path
from typing import Any
from uuid import uuid4
from fastapi import APIRouter, Depends, File, HTTPException, Request, UploadFile, status
from fastapi import APIRouter, Depends, File, Form, HTTPException, Query, Request, UploadFile, status
from fastapi.security import HTTPAuthorizationCredentials, HTTPBearer
from pydantic import BaseModel, Field
from sqlalchemy import delete, func, select, text
@@ -21,6 +21,26 @@ from app.models.datasource_config import DataSourceConfig
from app.models.system_setting import SystemSetting
from app.models.user import User
from app.services.tv_streams import get_tv_settings_payload
from app.services.earth_news import (
get_earth_news_sources_payload,
reset_earth_news_sources_payload,
save_earth_news_sources_payload,
test_news_source_config,
)
from app.services.earth_news_manual import (
broadcast_manual_news_changed,
create_manual_news_group,
delete_manual_news_item,
get_news_record_or_404,
import_manual_news_items,
list_news_groups,
list_news_records,
parse_manual_news_import_upload,
rename_manual_news_group,
reprocess_manual_news_item,
serialize_news_record,
upsert_manual_news_item,
)
from app.services.earth_boundaries import (
EarthBoundaryBuildError,
get_boundary_build_status,
@@ -100,6 +120,39 @@ class EarthAboutPayload(BaseModel):
meta: list[EarthAboutMetaItem] = Field(default_factory=list)
class EarthNewsSourcesPayload(BaseModel):
cache_version: int | None = None
source_tags: list[dict[str, Any]] = Field(default_factory=list)
categories: list[dict[str, Any]] = Field(default_factory=list)
item_tag_rules: list[dict[str, Any]] = Field(default_factory=list)
sources: list[dict[str, Any]] = Field(default_factory=list)
health: dict[str, Any] = Field(default_factory=dict)
class EarthNewsSourceTestPayload(BaseModel):
source: dict[str, Any] = Field(default_factory=dict)
class EarthNewsManualItemPayload(BaseModel):
title: str = Field(default="", max_length=500)
summary: str = Field(default="", max_length=1200)
content: str = Field(default="", max_length=12000)
url: str = Field(default="", max_length=2000)
source: str = Field(default="", max_length=255)
region: str = Field(default="global", max_length=80)
published_at: str | None = None
category: str = Field(default="other", max_length=80)
tags: list[str] = Field(default_factory=list)
location: dict[str, Any] | None = None
homepage_url: str = Field(default="", max_length=2000)
content_language: str = Field(default="", max_length=32)
group_id: str | None = Field(default=None, max_length=120)
class EarthNewsManualGroupPayload(BaseModel):
name: str = Field(default="", max_length=120)
def _normalize_earth_brand_payload(payload: dict[str, Any] | None) -> dict[str, str]:
merged = DEFAULT_EARTH_BRAND.copy()
if payload:
@@ -324,6 +377,194 @@ async def reset_earth_about(
return {"status": "reset", "about": _normalize_earth_about_payload(None), "is_default": True}
@router.get("/news-sources")
async def get_earth_news_sources(db: AsyncSession = Depends(get_db)):
return await get_earth_news_sources_payload(db)
@router.put("/news-sources")
async def update_earth_news_sources(
payload: EarthNewsSourcesPayload,
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
return await save_earth_news_sources_payload(db, payload.model_dump())
@router.delete("/news-sources")
@router.post("/news-sources/reset")
async def reset_earth_news_sources(
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
return await reset_earth_news_sources_payload(db)
@router.post("/news-sources/test")
async def test_earth_news_source(
payload: EarthNewsSourceTestPayload,
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
return await test_news_source_config(payload.source, db=db)
@router.get("/news-groups")
async def list_earth_news_groups_admin(
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
return await list_news_groups(db)
@router.post("/news-groups")
async def create_earth_news_group_admin(
payload: EarthNewsManualGroupPayload,
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
try:
group = await create_manual_news_group(db, payload.name)
except ValueError as exc:
raise HTTPException(status_code=422, detail=str(exc)) from exc
await db.commit()
return {"status": "ok", "group": group}
@router.put("/news-groups/{group_id:path}")
async def rename_earth_news_group_admin(
group_id: str,
payload: EarthNewsManualGroupPayload,
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
try:
group = await rename_manual_news_group(db, group_id, payload.name)
except ValueError as exc:
raise HTTPException(status_code=422, detail=str(exc)) from exc
await db.commit()
await broadcast_manual_news_changed()
return {"status": "ok", "group": group}
@router.get("/news-items")
async def list_earth_news_items_admin(
page: int = Query(1, ge=1),
page_size: int = Query(50, ge=1, le=100),
source_type: str | None = Query(None),
region: str | None = Query(None),
category: str | None = Query(None),
status_filter: str | None = Query(None, alias="status"),
group_id: str | None = Query(None),
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
return await list_news_records(
db,
page=page,
page_size=page_size,
source_type=source_type,
region=region,
category=category,
status_filter=status_filter,
group_id=group_id,
)
@router.post("/news-items")
async def create_earth_news_item_admin(
payload: EarthNewsManualItemPayload,
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
try:
result = await upsert_manual_news_item(db, payload.model_dump())
except ValueError as exc:
raise HTTPException(status_code=422, detail=str(exc)) from exc
except PermissionError as exc:
raise HTTPException(status_code=403, detail=str(exc)) from exc
await db.commit()
await broadcast_manual_news_changed()
return {"status": "ok", "created": result.created, "queued": result.queued, "item": serialize_news_record(result.item)}
@router.post("/news-items/import")
async def import_earth_news_items_admin(
file: UploadFile = File(...),
group_id: str | None = Form(default=None),
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
try:
payload = await parse_manual_news_import_upload(await file.read())
result = await import_manual_news_items(db, payload, group_id=group_id)
except ValueError as exc:
raise HTTPException(status_code=422, detail=str(exc)) from exc
await db.commit()
await broadcast_manual_news_changed()
return {"status": "ok", **result}
@router.put("/news-items/{item_id:path}")
async def update_earth_news_item_admin(
item_id: str,
payload: EarthNewsManualItemPayload,
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
existing = await get_news_record_or_404(db, item_id)
if existing is None:
raise HTTPException(status_code=404, detail="News item not found.")
try:
result = await upsert_manual_news_item(
db,
payload.model_dump(),
item_id_override=item_id,
)
except ValueError as exc:
raise HTTPException(status_code=422, detail=str(exc)) from exc
except PermissionError as exc:
raise HTTPException(status_code=403, detail=str(exc)) from exc
await db.commit()
await broadcast_manual_news_changed()
return {"status": "ok", "created": result.created, "queued": result.queued, "item": serialize_news_record(result.item)}
@router.delete("/news-items/{item_id:path}")
async def delete_earth_news_item_admin(
item_id: str,
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
try:
deleted = await delete_manual_news_item(db, item_id)
except PermissionError as exc:
raise HTTPException(status_code=403, detail=str(exc)) from exc
if not deleted:
raise HTTPException(status_code=404, detail="News item not found.")
await db.commit()
await broadcast_manual_news_changed()
return {"status": "deleted", "id": item_id}
@router.post("/news-items/{item_id:path}/reprocess")
async def reprocess_earth_news_item_admin(
item_id: str,
_current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
existing = await get_news_record_or_404(db, item_id)
if existing is None:
raise HTTPException(status_code=404, detail="News item not found.")
try:
queued = await reprocess_manual_news_item(db, item_id)
except PermissionError as exc:
raise HTTPException(status_code=403, detail=str(exc)) from exc
await db.commit()
await broadcast_manual_news_changed()
return {"status": "queued" if queued else "not_queued", "queued": queued, "id": item_id}
@router.get("/oobe-status")
async def get_earth_oobe_status(
current_user: User | None = Depends(_get_optional_current_user),

View File

@@ -90,7 +90,7 @@ async def get_interactables_geojson(
return interactables_to_geojson(items)
payload = await get_or_build_layer_payload(
key=earth_layer_cache.key("interactables", layer=layer or "all"),
key=earth_layer_cache.key("interactables", interactable_layer=layer or "all"),
policy=INTERACTABLE_CACHE_POLICY,
builder=build_payload,
response=response,

View File

@@ -105,7 +105,7 @@ def _parse_layer_bbox(bbox: str) -> tuple[float, float, float, float]:
@router.get("/vessels/snapshot")
async def get_vessel_layer_snapshot(
bbox: str = Query(..., description="lon_min,lat_min,lon_max,lat_max"),
zoom: int = Query(..., ge=1, le=20),
zoom: float = Query(..., ge=1, le=20),
limit: int = Query(DEFAULT_LAYER_LIMIT, ge=1),
vessel_type: Optional[str] = Query(None, alias="type"),
since_minutes: int = Query(60, ge=1, le=1440),

View File

@@ -1,16 +1,92 @@
from fastapi import APIRouter, Depends, Query
from fastapi import APIRouter, Depends, HTTPException, Query
from sqlalchemy.ext.asyncio import AsyncSession
from app.db.session import get_db
from app.services.earth_news import get_earth_news_payload
from app.services.earth_news import (
ALLOWED_NEWS_CATEGORY_KEYS,
SUPPORTED_NEWS_LOCALES,
REGION_ANCHORS,
get_earth_news_payload,
)
router = APIRouter()
def _parse_categories(raw: str | None) -> set[str] | None:
if raw is None or not raw.strip():
return None
requested = {item.strip().lower() for item in raw.split(",") if item.strip()}
invalid = sorted(requested - set(ALLOWED_NEWS_CATEGORY_KEYS))
if invalid:
raise HTTPException(
status_code=422,
detail={
"message": "Unsupported news categories.",
"invalid_categories": invalid,
"allowed_categories": list(ALLOWED_NEWS_CATEGORY_KEYS),
},
)
return requested or None
def _parse_source_ids(raw: str | None) -> set[str] | None:
if raw is None or not raw.strip():
return None
return {item.strip() for item in raw.split(",") if item.strip()} or None
def _parse_limit(raw: int | None) -> int:
if raw is None:
return 12
if raw < 1:
raise HTTPException(status_code=422, detail={"message": "News limit must be greater than 0."})
return min(raw, 100)
def _parse_locale(raw: str | None) -> str:
if raw is None or not raw.strip():
return "zh-CN"
requested = raw.strip()
if requested not in SUPPORTED_NEWS_LOCALES:
raise HTTPException(
status_code=422,
detail={
"message": "Unsupported news locale.",
"invalid_locale": requested,
"allowed_locales": sorted(SUPPORTED_NEWS_LOCALES),
},
)
return requested
@router.get("/earth-feed")
async def get_earth_feed(
lat: float | None = Query(None, description="Current Earth view center latitude"),
lon: float | None = Query(None, description="Current Earth view center longitude"),
region: str | None = Query(None, description="Explicit Earth news region for UE/client integrations"),
categories: str | None = Query(None, description="Comma-separated news category keys"),
sources: str | None = Query(None, description="Comma-separated news source ids"),
limit: int | None = Query(None, description="Maximum news items to return, capped at 100"),
locale: str | None = Query(None, description="Display locale, zh-CN or en-US"),
db: AsyncSession = Depends(get_db),
):
return await get_earth_news_payload(lat=lat, lon=lon, db=db)
normalized_region = region.strip().lower() if isinstance(region, str) and region.strip() else None
if normalized_region is not None and normalized_region not in REGION_ANCHORS:
raise HTTPException(
status_code=422,
detail={
"message": "Unsupported news region.",
"invalid_region": normalized_region,
"allowed_regions": list(REGION_ANCHORS.keys()),
},
)
return await get_earth_news_payload(
lat=lat,
lon=lon,
region=normalized_region,
categories=_parse_categories(categories),
source_ids=_parse_source_ids(sources),
limit=_parse_limit(limit),
locale=_parse_locale(locale),
db=db,
)

View File

@@ -12,6 +12,7 @@ from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.logging import get_logger
from app.core.enums import ProviderApi, TVSourceType, UserRole
from app.core.security import get_current_user
from app.core.time import to_iso8601_utc
from app.core.config import settings as app_settings
@@ -73,7 +74,7 @@ router = APIRouter()
logger = get_logger(__name__, service="api")
AI_PROVIDER_QUICK_CONNECT_TIMEOUT_SECONDS = 5
AI_CONNECTION_TEST_PROMPT_KEY = "ai.connection_test"
SECRET_REVEAL_ROLES = {"admin", "super_admin"}
SECRET_REVEAL_ROLES = {UserRole.ADMIN.value, UserRole.SUPER_ADMIN.value}
DEFAULT_SETTINGS = {
"system": {
@@ -236,7 +237,7 @@ class TVStreamSourceUpdate(BaseModel):
provider: str = Field(default="Unknown", max_length=100)
region: str = Field(default="Global", max_length=100)
language: str = Field(default="und", max_length=32)
source_type: str = Field(default="iframe", pattern="^(iframe|hls|video|external|youtube)$")
source_type: TVSourceType = TVSourceType.IFRAME
embed_url: str = ""
stream_url: str = ""
homepage_url: str = ""
@@ -279,7 +280,7 @@ class AIProviderIntegrationUpdate(BaseModel):
service_token: Optional[str] = None
default_provider: Optional[str] = None
provider: str = Field(default="minimax", max_length=80)
provider_api: str = Field(default="anthropic-messages", max_length=80)
provider_api: ProviderApi = ProviderApi.ANTHROPIC_MESSAGES
base_url: str = Field(default="", max_length=500)
model: str = Field(default="", max_length=200)
api_key: Optional[str] = None
@@ -423,7 +424,7 @@ def _get_provider_preset(provider: str) -> dict:
except ValueError:
return {
"provider": provider,
"provider_api": "openai-completions",
"provider_api": ProviderApi.OPENAI_COMPLETIONS.value,
"base_url": "",
"model": "",
"models": [],
@@ -487,12 +488,12 @@ def _provider_defaults(provider: str) -> dict:
preset = _get_provider_preset(provider)
return {
"provider": provider,
"provider_api": preset.get("provider_api") or "openai-completions",
"provider_api": preset.get("provider_api") or ProviderApi.OPENAI_COMPLETIONS.value,
"base_url": preset.get("base_url") or "",
"model": preset.get("model") or "",
"api_key": "",
"max_tokens": (
1200 if preset.get("provider_api") == "anthropic-messages" else 4096
1200 if preset.get("provider_api") == ProviderApi.ANTHROPIC_MESSAGES.value else 4096
),
"anthropic_version": "2023-06-01",
"model_provider_apis": preset.get("model_provider_apis") or {},
@@ -665,7 +666,7 @@ def _runtime_config_from_ai_payload(ai_payload: dict) -> dict:
),
"llm_config": {
"provider": default_provider,
"provider_api": provider_config.get("provider_api") or "anthropic-messages",
"provider_api": provider_config.get("provider_api") or ProviderApi.ANTHROPIC_MESSAGES.value,
"base_url": provider_config.get("base_url") or "",
"model": provider_config.get("model") or "",
"api_key": api_key,
@@ -781,7 +782,10 @@ def _contains_model(model_ids: list[str], model: str) -> bool:
async def _check_ai_provider_lightweight(llm_config: dict, timeout_seconds: int) -> dict:
provider = _normalize_provider_id(llm_config.get("provider") or "")
configured_api = str(llm_config.get("provider_api") or "").strip() or "openai-completions"
configured_api = (
str(llm_config.get("provider_api") or "").strip()
or ProviderApi.OPENAI_COMPLETIONS.value
)
model = str(llm_config.get("model") or "").strip()
base_url = str(llm_config.get("base_url") or "").strip().rstrip("/")
api_key = str(llm_config.get("api_key") or "").strip()
@@ -802,7 +806,7 @@ async def _check_ai_provider_lightweight(llm_config: dict, timeout_seconds: int)
"message": "当前 provider/base_url/model 未完整配置。",
"mode": "lightweight_config",
}
if provider_api != "ollama-generate" and not api_key:
if provider_api != ProviderApi.OLLAMA_GENERATE.value and not api_key:
return {
"success": False,
"connected": False,
@@ -813,13 +817,13 @@ async def _check_ai_provider_lightweight(llm_config: dict, timeout_seconds: int)
if provider == "opencode-go":
url = _join_provider_url(base_url, "/models")
headers = {"Authorization": f"Bearer {api_key}"}
elif provider_api == "ollama-generate":
elif provider_api == ProviderApi.OLLAMA_GENERATE.value:
url = _join_provider_url(base_url, "/api/tags")
headers: dict[str, str] = {}
elif provider_api == "openai-completions":
elif provider_api == ProviderApi.OPENAI_COMPLETIONS.value:
url = _join_provider_url(base_url, "/models")
headers = {"Authorization": f"Bearer {api_key}"}
elif provider_api == "anthropic-messages":
elif provider_api == ProviderApi.ANTHROPIC_MESSAGES.value:
url = _join_provider_url(base_url, "/models")
headers = {
"x-api-key": api_key,
@@ -1165,7 +1169,7 @@ async def serialize_external_integrations(db: AsyncSession) -> dict:
api_key, api_key_source = _resolve_provider_api_key(provider_id, provider_config)
providers_payload[provider_id] = {
"provider": provider_id,
"provider_api": provider_config.get("provider_api") or "openai-completions",
"provider_api": provider_config.get("provider_api") or ProviderApi.OPENAI_COMPLETIONS.value,
"base_url": provider_config.get("base_url") or "",
"model": provider_config.get("model") or "",
"api_key": _mask_secret(api_key, api_key_source),
@@ -1215,7 +1219,7 @@ async def serialize_external_integrations(db: AsyncSession) -> dict:
"service_token": _mask_secret(*_resolve_service_token(normalized_ai)),
"default_provider": default_provider,
"provider": default_provider,
"provider_api": display_llm_config.get("provider_api") or "anthropic-messages",
"provider_api": display_llm_config.get("provider_api") or ProviderApi.ANTHROPIC_MESSAGES.value,
"base_url": display_llm_config.get("base_url") or "https://api.minimaxi.com/anthropic",
"model": display_llm_config.get("model") or "MiniMax-M2.7",
"api_key": display_llm_config.get("api_key") or _mask_secret(None),
@@ -1458,7 +1462,7 @@ async def update_smtp_settings(
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
if current_user.role not in ("admin", "super_admin"):
if current_user.role not in (UserRole.ADMIN.value, UserRole.SUPER_ADMIN.value):
raise HTTPException(status_code=403, detail="Only administrators can change SMTP settings")
current = await get_setting_payload(db, "smtp")
merged = _build_smtp_payload(current, payload)
@@ -1472,7 +1476,7 @@ async def test_smtp_settings(
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
if current_user.role not in ("admin", "super_admin"):
if current_user.role not in (UserRole.ADMIN.value, UserRole.SUPER_ADMIN.value):
raise HTTPException(status_code=403, detail="Only administrators can test SMTP settings")
from app.services.email import EmailError, send_email

View File

@@ -1,21 +1,19 @@
from __future__ import annotations
import os
import json
import secrets
import subprocess
import sys
from datetime import datetime
from fastapi import APIRouter, Depends, HTTPException, Query, Request, status
from fastapi import APIRouter, Depends, Header, HTTPException, Query, Request, status
from pydantic import BaseModel
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.config import ROOT_DIR
from app.core.config import ROOT_DIR, settings
from app.core.security import get_current_user
from app.db.session import get_db
from app.models.system_log import AuditLog, SystemLog
from app.models.user import User
from app.services.persistent_logs import record_audit_log, record_system_log
from app.services.system_control import (
@@ -38,49 +36,17 @@ from app.services.system_logs import (
append_buffer_log,
list_log_sources,
normalize_log_level,
read_database_log_snapshot,
read_log_snapshot,
read_observability_group_events,
read_observability_groups,
read_observability_raw_events,
)
from app.services.earth_layer_cache import earth_layer_cache
router = APIRouter()
def _compact_log_context(context: dict | None) -> str:
if not context:
return ""
allowed = {
key: value
for key, value in (context or {}).items()
if key
in {
"status",
"duration_ms",
"provider",
"model",
"result_provider",
"result_model",
"collector_name",
"datasource_id",
"task_id",
"snapshot_id",
"raw_count",
"transformed_count",
"saved_count",
"created",
"updated",
"unchanged",
"deleted",
"result_count",
"status_code",
"error_type",
"error",
}
}
if not allowed:
return ""
return json.dumps(allowed, ensure_ascii=False, sort_keys=True)
class RestartTaskCreate(BaseModel):
action: str
@@ -146,12 +112,104 @@ class EarthClientLogEventCreate(BaseModel):
url: str | None = None
module: str | None = None
detail: str | None = None
fingerprint: str | None = None
occurrence_count: int = 1
metadata: dict[str, object] | None = None
class EarthClientLogEventResponse(BaseModel):
accepted: bool
source_id: str
level: str
fingerprint: str | None = None
class ServiceLogEventCreate(BaseModel):
source: str = "ai-provider"
service: str = "ai-provider"
module: str | None = None
category: str | None = None
event: str = "service.runtime_log"
level: str = "error"
message: str
fingerprint: str | None = None
occurrence_count: int = 1
request_id: str | None = None
trace_id: str | None = None
task_id: str | None = None
source_id: int | str | None = None
provider: str | None = None
context: dict[str, object] | None = None
async def ingest_client_log_event(
source_id: str,
*,
service: str,
event: str,
default_module: str,
default_category: str,
payload: EarthClientLogEventCreate,
request: Request,
) -> EarthClientLogEventResponse:
normalized_level = normalize_log_level(payload.level)
append_buffer_log(
source_id,
level=normalized_level,
message=payload.message,
context={
"category": payload.category or "",
"url": payload.url or "",
"module": payload.module or "",
"detail": payload.detail or "",
"fingerprint": payload.fingerprint or "",
"occurrence_count": max(1, int(payload.occurrence_count or 1)),
"metadata": payload.metadata or {},
},
)
await record_system_log(
source=source_id,
service=service,
module=payload.module or default_module,
event=event,
level=normalized_level,
message=payload.message,
category=payload.category or default_category,
context={
"url": payload.url or "",
"detail": payload.detail or "",
"module": payload.module or "",
"client_ip": request.client.host if request.client else "",
"metadata": payload.metadata or {},
},
fingerprint=payload.fingerprint,
occurrence_count=max(1, int(payload.occurrence_count or 1)),
)
return EarthClientLogEventResponse(accepted=True, source_id=source_id, level=normalized_level, fingerprint=payload.fingerprint)
def require_observability_ingest_token(
authorization: str | None,
ingest_token: str | None,
) -> None:
expected_token = settings.OBSERVABILITY_INGEST_TOKEN.strip()
if not expected_token:
raise HTTPException(
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
detail="Observability service ingestion is not configured",
)
provided = ""
if ingest_token:
provided = ingest_token.strip()
elif authorization:
scheme, _, token = authorization.partition(" ")
if scheme.lower() == "bearer":
provided = token.strip()
if not provided or not secrets.compare_digest(provided, expected_token):
raise HTTPException(
status_code=status.HTTP_403_FORBIDDEN,
detail="Invalid observability ingestion token",
)
class EarthLayerCacheStatusResponse(BaseModel):
@@ -376,98 +434,116 @@ async def get_system_log_sources(
}
async def read_database_log_snapshot(
source_id: str,
*,
limit: int,
level: str,
levels: str | None,
start_date: str | None,
end_date: str | None,
search: str | None,
db: AsyncSession,
) -> dict | None:
selected_levels = set(normalize_log_level(item) for item in (levels or level).split(",") if item.strip())
selected_levels.discard("all")
search_query = (search or "").strip().lower()
lines: list[str] = []
@router.get("/logs/observability/groups")
async def get_observability_log_groups(
limit: int = DEFAULT_LOG_LINE_LIMIT,
level: str = "all",
levels: str | None = Query(None, description="Comma-separated log levels"),
start_date: str | None = Query(None, description="Filter logs from this date (YYYY-MM-DD)"),
end_date: str | None = Query(None, description="Filter logs until this date (YYYY-MM-DD)"),
search: str | None = Query(None, description="Case-insensitive substring search"),
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
ensure_super_admin(current_user)
if limit < 1 or limit > MAX_LOG_LINE_LIMIT:
raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail=f"limit must be between 1 and {MAX_LOG_LINE_LIMIT}")
normalized_start_date = validate_log_date(start_date, "start_date")
normalized_end_date = validate_log_date(end_date, "end_date")
if normalized_start_date and normalized_end_date and normalized_start_date > normalized_end_date:
raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail="start_date must be earlier than or equal to end_date")
return await read_observability_groups(
limit=limit,
level=level,
levels=levels,
start_date=normalized_start_date,
end_date=normalized_end_date,
search=search,
db=db,
)
if source_id == "system-db":
query = select(SystemLog).order_by(SystemLog.occurred_at.desc().nullslast(), SystemLog.id.desc()).limit(limit * 5)
result = await db.execute(query)
records = result.scalars().all()
for record in records:
record_level = normalize_log_level(record.level)
if selected_levels and record_level not in selected_levels:
continue
occurred_at = record.occurred_at.date().isoformat() if record.occurred_at else ""
if start_date and occurred_at and occurred_at < start_date:
continue
if end_date and occurred_at and occurred_at > end_date:
continue
line = " ".join(
part
for part in [
record.occurred_at.isoformat() if record.occurred_at else "",
record_level.upper(),
record.source,
record.category or "",
record.event or "",
f"request_id={record.request_id}" if record.request_id else "",
record.message,
_compact_log_context(record.context),
]
if part
)
if search_query and search_query not in line.lower():
continue
lines.append(line)
elif source_id == "audit-db":
query = select(AuditLog).order_by(AuditLog.occurred_at.desc().nullslast(), AuditLog.id.desc()).limit(limit * 5)
result = await db.execute(query)
records = result.scalars().all()
for record in records:
occurred_at = record.occurred_at.date().isoformat() if record.occurred_at else ""
if start_date and occurred_at and occurred_at < start_date:
continue
if end_date and occurred_at and occurred_at > end_date:
continue
line = " ".join(
part
for part in [
record.occurred_at.isoformat() if record.occurred_at else "",
"INFO",
record.action,
record.target_type or "",
record.target_id or "",
record.result or "",
]
if part
)
if search_query and search_query not in line.lower():
continue
lines.append(line)
else:
return None
lines = list(reversed(lines[:limit]))
return {
"source_id": source_id,
"name": "系统事件" if source_id == "system-db" else "审计事件",
"kind": "database",
"location": "table://system_logs" if source_id == "system-db" else "table://audit_logs",
"description": "数据库持久化日志",
"category": "database" if source_id == "system-db" else "audit",
"status": "ok" if lines else "empty",
"level": level,
"selected_levels": sorted(selected_levels),
"search_query": search or "",
"available_levels": ["all", "error", "warning", "info", "debug"],
"daily_markers": [],
"line_limit": limit,
"line_count": len(lines),
"lines": lines,
}
@router.get("/logs/observability/groups/{fingerprint}/events")
async def get_observability_group_events(
fingerprint: str,
limit: int = DEFAULT_LOG_LINE_LIMIT,
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
ensure_super_admin(current_user)
if limit < 1 or limit > MAX_LOG_LINE_LIMIT:
raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail=f"limit must be between 1 and {MAX_LOG_LINE_LIMIT}")
payload = await read_observability_group_events(fingerprint, limit=limit, db=db)
if payload is None:
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="Observability group not found")
return payload
@router.get("/logs/observability/raw")
async def get_observability_raw_events(
limit: int = DEFAULT_LOG_LINE_LIMIT,
level: str = "all",
levels: str | None = Query(None, description="Comma-separated log levels"),
start_date: str | None = Query(None, description="Filter logs from this date (YYYY-MM-DD)"),
end_date: str | None = Query(None, description="Filter logs until this date (YYYY-MM-DD)"),
search: str | None = Query(None, description="Case-insensitive substring search"),
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
ensure_super_admin(current_user)
if limit < 1 or limit > MAX_LOG_LINE_LIMIT:
raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail=f"limit must be between 1 and {MAX_LOG_LINE_LIMIT}")
normalized_start_date = validate_log_date(start_date, "start_date")
normalized_end_date = validate_log_date(end_date, "end_date")
return await read_observability_raw_events(
limit=limit,
level=level,
levels=levels,
start_date=normalized_start_date,
end_date=normalized_end_date,
search=search,
db=db,
)
@router.post("/logs/service", response_model=EarthClientLogEventResponse)
async def ingest_service_log(
payload: ServiceLogEventCreate,
authorization: str | None = Header(default=None),
ingest_token: str | None = Header(default=None, alias="X-Planet-Observability-Token"),
):
require_observability_ingest_token(authorization, ingest_token)
normalized_level = normalize_log_level(payload.level)
source = (payload.source or "ai-provider").strip() or "ai-provider"
context = dict(payload.context or {})
if payload.request_id:
context["request_id"] = payload.request_id
if payload.trace_id:
context["trace_id"] = payload.trace_id
if payload.task_id:
context["task_id"] = payload.task_id
if payload.source_id is not None:
context["source_id"] = payload.source_id
if payload.provider:
context["provider"] = payload.provider
await record_system_log(
source=source,
service=(payload.service or source).strip() or source,
module=payload.module or source,
event=(payload.event or "service.runtime_log").strip() or "service.runtime_log",
level=normalized_level,
message=payload.message,
category=payload.category or "service-runtime",
context=context,
fingerprint=payload.fingerprint,
occurrence_count=max(1, int(payload.occurrence_count or 1)),
)
return EarthClientLogEventResponse(
accepted=True,
source_id=source,
level=normalized_level,
fingerprint=payload.fingerprint,
)
@router.get("/logs/{source_id}", response_model=SystemLogSnapshotResponse)
@@ -533,31 +609,28 @@ async def ingest_earth_client_log(
payload: EarthClientLogEventCreate,
request: Request,
):
normalized_level = normalize_log_level(payload.level)
append_buffer_log(
return await ingest_client_log_event(
"earth-client",
level=normalized_level,
message=payload.message,
context={
"category": payload.category or "",
"url": payload.url or "",
"module": payload.module or "",
"detail": payload.detail or "",
},
)
await record_system_log(
source="earth-client",
service="earth",
module=payload.module or "earth-client",
event="earth.client.runtime_log",
level=normalized_level,
message=payload.message,
category=payload.category or "client-runtime",
context={
"url": payload.url or "",
"detail": payload.detail or "",
"module": payload.module or "",
"client_ip": request.client.host if request.client else "",
},
default_module="earth-client",
default_category="client-runtime",
payload=payload,
request=request,
)
@router.post("/logs/admin-client", response_model=EarthClientLogEventResponse)
async def ingest_admin_client_log(
payload: EarthClientLogEventCreate,
request: Request,
):
return await ingest_client_log_event(
"admin-client",
service="admin",
event="admin.client.runtime_log",
default_module="admin-client",
default_category="client-runtime",
payload=payload,
request=request,
)
return {"accepted": True, "source_id": "earth-client", "level": normalized_level}

View File

@@ -29,7 +29,8 @@ async def list_tasks(
SELECT ct.id, ct.datasource_id, ds.name as datasource_name, ct.status,
ct.started_at, ct.completed_at, ct.records_processed, ct.error_message,
ct.phase, ct.phase_progress, ct.phase_message, ct.phase_current,
ct.phase_total, ct.phase_unit, ct.total_records, ct.progress
ct.phase_total, ct.phase_unit, ct.total_records, ct.progress,
ct.task_type, ct.source, ds.source as datasource_source
FROM collection_tasks ct
JOIN data_sources ds ON ct.datasource_id = ds.id
WHERE 1=1
@@ -39,12 +40,19 @@ async def list_tasks(
if datasource_id:
query += " AND ct.datasource_id = :datasource_id"
count_query += " WHERE ct.datasource_id = :datasource_id"
count_query += " AND ct.datasource_id = :datasource_id"
params["datasource_id"] = datasource_id
if status:
query += " AND ct.status = :status"
count_query += " AND ct.status = :status"
params["status"] = status
statuses = [item.strip() for item in status.split(",") if item.strip()]
if len(statuses) > 1:
placeholders = ", ".join(f":status_{index}" for index, _item in enumerate(statuses))
query += f" AND ct.status IN ({placeholders})"
count_query += f" AND ct.status IN ({placeholders})"
params.update({f"status_{index}": item for index, item in enumerate(statuses)})
else:
query += " AND ct.status = :status"
count_query += " AND ct.status = :status"
params["status"] = statuses[0] if statuses else status
query += f" ORDER BY ct.created_at DESC LIMIT {page_size} OFFSET {offset}"
@@ -76,6 +84,9 @@ async def list_tasks(
"phase_unit": t[13],
"total_records": t[14],
"progress": t[15],
"task_type": t[16],
"source": t[17] or t[18],
"datasource_source": t[18],
}
for t in tasks
],

View File

@@ -1,3 +1,4 @@
import re
from urllib.parse import quote, urljoin
import httpx
@@ -10,6 +11,26 @@ from app.services.tv_streams import get_public_tv_payload, is_allowed_tv_proxy_u
router = APIRouter()
_HLS_URI_ATTRIBUTE_RE = re.compile(r'URI="([^"]+)"')
def _proxied_tv_url(url: str) -> str:
return f"/api/v1/tv/proxy?url={quote(url, safe='')}"
def _rewrite_hls_uri_attributes(line: str, *, base_url: str) -> str:
def replace(match: re.Match[str]) -> str:
uri = match.group(1)
absolute_url = urljoin(base_url, uri)
return f'URI="{_proxied_tv_url(absolute_url)}"'
return _HLS_URI_ATTRIBUTE_RE.sub(replace, line)
def _should_strip_hls_metadata_line(line: str) -> bool:
normalized = line.strip().upper()
return normalized.startswith("#EXT-X-MEDIA:") and "TYPE=SUBTITLES" in normalized
@router.get("/streams")
async def list_public_tv_streams(
@@ -56,11 +77,16 @@ async def proxy_tv_stream(
rewritten_lines: list[str] = []
for line in manifest_text.splitlines():
stripped = line.strip()
if not stripped or stripped.startswith("#"):
if not stripped:
rewritten_lines.append(line)
continue
if stripped.startswith("#"):
if _should_strip_hls_metadata_line(line):
continue
rewritten_lines.append(_rewrite_hls_uri_attributes(line, base_url=response_url))
continue
absolute_url = urljoin(response_url, stripped)
rewritten_lines.append(f"/api/v1/tv/proxy?url={quote(absolute_url, safe='')}")
rewritten_lines.append(_proxied_tv_url(absolute_url))
return Response(
content="\n".join(rewritten_lines),
media_type="application/vnd.apple.mpegurl",

View File

@@ -5,6 +5,7 @@ from fastapi import APIRouter, Depends, HTTPException, status
from sqlalchemy.ext.asyncio import AsyncSession
from sqlalchemy import text
from app.core.enums import UserRole
from app.core.security import get_current_user, get_password_hash
from app.db.session import get_db
from app.models.user import User
@@ -13,6 +14,8 @@ from app.schemas.user import UserCreate, UserUpdate
router = APIRouter()
VALID_GATEKEEPER_GROUPS = {"docs_user", "docs_developer", "docs_admin"}
ADMIN_ROLES = [UserRole.SUPER_ADMIN.value, UserRole.ADMIN.value]
SUPER_ADMIN_ROLES = [UserRole.SUPER_ADMIN.value]
def check_permission(current_user: User, required_roles: List[str]) -> bool:
@@ -32,7 +35,7 @@ async def list_users(
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
if not check_permission(current_user, ["super_admin", "admin"]):
if not check_permission(current_user, ADMIN_ROLES):
raise HTTPException(
status_code=status.HTTP_403_FORBIDDEN,
detail="Insufficient permissions",
@@ -91,7 +94,7 @@ async def get_user(
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
if not check_permission(current_user, ["super_admin", "admin"]) and current_user.id != user_id:
if not check_permission(current_user, ADMIN_ROLES) and current_user.id != user_id:
raise HTTPException(
status_code=status.HTTP_403_FORBIDDEN,
detail="Insufficient permissions",
@@ -128,7 +131,7 @@ async def create_user(
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
if not check_permission(current_user, ["super_admin"]):
if not check_permission(current_user, SUPER_ADMIN_ROLES):
raise HTTPException(
status_code=status.HTTP_403_FORBIDDEN,
detail="Only super_admin can create users",
@@ -196,18 +199,18 @@ async def update_user(
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
if not check_permission(current_user, ["super_admin", "admin"]) and current_user.id != user_id:
if not check_permission(current_user, ADMIN_ROLES) and current_user.id != user_id:
raise HTTPException(
status_code=status.HTTP_403_FORBIDDEN,
detail="Insufficient permissions",
)
if not check_permission(current_user, ["super_admin"]) and user_data.role is not None:
if not check_permission(current_user, SUPER_ADMIN_ROLES) and user_data.role is not None:
raise HTTPException(
status_code=status.HTTP_403_FORBIDDEN,
detail="Only super_admin can change user role",
)
if not check_permission(current_user, ["super_admin"]) and user_data.gatekeeper_groups is not None:
if not check_permission(current_user, SUPER_ADMIN_ROLES) and user_data.gatekeeper_groups is not None:
raise HTTPException(
status_code=status.HTTP_403_FORBIDDEN,
detail="Only super_admin can change Gatekeeper groups",
@@ -260,7 +263,7 @@ async def delete_user(
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
if not check_permission(current_user, ["super_admin"]):
if not check_permission(current_user, SUPER_ADMIN_ROLES):
raise HTTPException(
status_code=status.HTTP_403_FORBIDDEN,
detail="Only super_admin can delete users",

View File

@@ -1,4 +1,4 @@
"""Bounded vessel snapshot APIs for viewport-first consumers."""
"""Bounded vessel snapshot APIs backed by the latest vessel state table."""
from typing import Optional
@@ -14,8 +14,8 @@ router = APIRouter()
@router.get("/snapshot")
async def get_vessel_snapshot(
bbox: Optional[str] = Query(None, description="Viewport bbox as lon_min,lat_min,lon_max,lat_max"),
zoom: int = Query(..., ge=1, le=20, description="Current map zoom level"),
bbox: Optional[str] = Query(None, description="Snapshot bbox as lon_min,lat_min,lon_max,lat_max"),
zoom: float = Query(..., ge=1, le=20, description="Current map zoom level"),
type: Optional[str] = Query(
None,
description="Comma-separated vessel types: cargo,tanker,passenger,fishing,military,other",

View File

@@ -18,6 +18,7 @@ from sqlalchemy import select, func
from typing import List, Dict, Any, Optional
from app.core.collected_data_fields import get_record_field
from app.core.enums import BGPStatus
from app.core.satellite_tle import build_tle_lines_from_elements
from app.core.time import to_iso8601_utc
from app.db.session import get_db
@@ -25,7 +26,7 @@ from app.models.bgp_anomaly import BGPAnomaly
from app.models.bgp_incident import BGPIncident
from app.models.bgp_observation import BGPObservation
from app.models.collected_data import CollectedData
from app.models.vessel import AISSourceHealth, VesselPosition, VesselStatic
from app.models.vessel import AISSourceHealth, VesselCurrentState, VesselPosition, VesselStatic
from app.services.bgp_collectors import build_bgp_collector_coverage
from app.services.cable_graph import build_graph_from_data, CableGraph, haversine_distance
from app.services.compute_center_locations import (
@@ -47,11 +48,10 @@ from app.services.location.llm_fallback import (
from app.services.persistent_logs import record_system_log
from app.services.vessel_ais_aggregation import (
build_field_conflict_candidates,
count_unique_raw_vessel_mmsi,
get_aggregated_vessel,
get_aggregated_vessel_track,
get_aggregated_vessels,
get_aggregated_vessels_snapshot,
get_current_vessels_snapshot,
get_vessel_conflict_records,
get_vessel_raw_observations,
MAX_SNAPSHOT_LIMIT,
@@ -75,7 +75,6 @@ TERRAIN_TILE_BATCH_MAX_ITEMS = 128
TERRAIN_TILE_BATCH_CONCURRENCY = 16
_terrain_tile_cache: OrderedDict[tuple[int, int, int], tuple[bytes, str, dict[str, str]]] = OrderedDict()
VESSEL_NAME_FALLBACK_PATTERN = re.compile(r"^mmsi\s*\d+$", re.IGNORECASE)
VESSEL_SNAPSHOT_LEGACY_FALLBACK_ENABLED = True
SECONDS_PER_MINUTE = 60
BYTES_PER_MIB = 1024 * 1024
CABLE_CACHE_FRESH_SECONDS = 6 * 60 * SECONDS_PER_MINUTE
@@ -839,9 +838,14 @@ def convert_aggregated_vessels_to_geojson(vessels: List[dict[str, Any]]) -> Dict
continue
source_summary = {}
for source, summary in (vessel.get("source_summary") or {}).items():
latest_observed_at = summary.get("latest_observed_at")
source_summary[source] = {
**summary,
"latest_observed_at": to_iso8601_utc(summary.get("latest_observed_at")),
"latest_observed_at": (
to_iso8601_utc(latest_observed_at)
if isinstance(latest_observed_at, datetime)
else latest_observed_at
),
}
props = {
"mmsi": vessel["mmsi"],
@@ -1075,7 +1079,7 @@ async def build_vessel_snapshot_response(
db: AsyncSession,
*,
bbox: tuple[float, float, float, float] | None,
zoom: int | None,
zoom: float | None,
type_filter: str | None,
limit: int | None,
since_minutes: int = 60,
@@ -2340,58 +2344,35 @@ async def _load_raw_vessel_snapshot_features(
observed_since: datetime,
) -> tuple[list[dict[str, Any]], dict[str, Any]]:
if bbox is None:
aggregated_vessels = await get_aggregated_vessels(
db,
limit=limit,
observed_since=observed_since,
)
else:
aggregated_vessels = await get_aggregated_vessels_snapshot(
db,
bbox=bbox,
limit=limit,
observed_since=observed_since,
)
raw_geojson = convert_aggregated_vessels_to_geojson(aggregated_vessels)
raw_features = raw_geojson.get("features", [])
features = raw_features
legacy_features: list[dict[str, Any]] = []
legacy_fallback_used = False
if not raw_features and VESSEL_SNAPSHOT_LEGACY_FALLBACK_ENABLED:
legacy_features = await _load_legacy_vessel_snapshot_features(
db,
bbox=bbox,
limit=limit,
)
features, _merge_diagnostics = _merge_vessel_features(raw_features, legacy_features)
legacy_fallback_used = bool(legacy_features)
return [], {
"source": "vessel_current_state",
"current_state_count": 0,
"final_unique_mmsi": 0,
}
current_vessels = await get_current_vessels_snapshot(
db,
bbox=bbox,
limit=limit,
observed_since=observed_since,
)
features = convert_aggregated_vessels_to_geojson(current_vessels).get("features", [])
unique_mmsi = len(
{
key
for key in (_feature_mmsi_key(feature) for feature in features)
if key is not None
}
)
return features, {
"raw_feature_count": len(raw_features),
"raw_unique_mmsi": len(
{
key
for key in (_feature_mmsi_key(feature) for feature in raw_features)
if key is not None
}
),
"legacy_feature_count": len(legacy_features),
"legacy_backfilled_mmsi": len(
{
key
for key in (_feature_mmsi_key(feature) for feature in legacy_features)
if key is not None
}
),
"legacy_fallback_enabled": VESSEL_SNAPSHOT_LEGACY_FALLBACK_ENABLED,
"legacy_fallback_used": legacy_fallback_used,
"final_unique_mmsi": len(
{
key
for key in (_feature_mmsi_key(feature) for feature in features)
if key is not None
}
),
"source": "vessel_current_state",
"current_state_count": len(features),
"final_unique_mmsi": unique_mmsi,
"raw_feature_count": 0,
"raw_unique_mmsi": 0,
"legacy_feature_count": 0,
"legacy_backfilled_mmsi": 0,
"legacy_fallback_enabled": False,
"legacy_fallback_used": False,
}
@router.get("/vessels/custom-supplements")
@@ -2759,10 +2740,10 @@ async def _build_visualization_geo_summary(db: AsyncSession) -> dict[str, Any]:
compute_center_count = supercomputer_count + gpu_cluster_count
active_incident_result = await db.execute(
select(func.count(BGPIncident.id)).where(BGPIncident.status == "active"),
select(func.count(BGPIncident.id)).where(BGPIncident.status == BGPStatus.ACTIVE.value),
)
active_anomaly_result = await db.execute(
select(func.count(BGPAnomaly.id)).where(BGPAnomaly.status == "active"),
select(func.count(BGPAnomaly.id)).where(BGPAnomaly.status == BGPStatus.ACTIVE.value),
)
active_incident_count = int(active_incident_result.scalar() or 0)
active_anomaly_count = int(active_anomaly_result.scalar() or 0)
@@ -2783,21 +2764,14 @@ async def _build_visualization_geo_summary(db: AsyncSession) -> dict[str, Any]:
)
else:
bgp_collector_count = int(bgp_collector_scalar or 0)
raw_unique_window_hours = 24
raw_unique_mmsi = await count_unique_raw_vessel_mmsi(
db,
observed_since=datetime.now(UTC) - timedelta(hours=raw_unique_window_hours),
vessel_current_window_minutes = 60
vessel_current_result = await db.execute(
select(func.count(VesselCurrentState.mmsi)).where(
VesselCurrentState.observed_at
>= datetime.now(UTC) - timedelta(minutes=vessel_current_window_minutes)
)
)
legacy_unique_result = await db.execute(
select(func.count(func.distinct(VesselPosition.mmsi)))
)
legacy_unique_mmsi = int(legacy_unique_result.scalar() or 0)
legacy_fallback_active = (
VESSEL_SNAPSHOT_LEGACY_FALLBACK_ENABLED
and raw_unique_mmsi == 0
and legacy_unique_mmsi > 0
)
vessel_count = legacy_unique_mmsi if legacy_fallback_active else raw_unique_mmsi
vessel_count = int(vessel_current_result.scalar() or 0)
aisstream_health = await db.get(AISSourceHealth, "aisstream_vessels")
return {
@@ -2808,11 +2782,10 @@ async def _build_visualization_geo_summary(db: AsyncSession) -> dict[str, Any]:
"satellite_count": satellite_count,
"compute_center_count": compute_center_count,
"vessel_count": vessel_count,
"vessel_count_source": "legacy_fallback" if legacy_fallback_active else "raw_recent",
"vessel_legacy_fallback_enabled": VESSEL_SNAPSHOT_LEGACY_FALLBACK_ENABLED,
"vessel_raw_unique_mmsi": raw_unique_mmsi,
"vessel_raw_unique_window_hours": raw_unique_window_hours,
"vessel_legacy_unique_mmsi": legacy_unique_mmsi,
"vessel_count_source": "vessel_current_state",
"vessel_current_window_minutes": vessel_current_window_minutes,
"vessel_raw_unique_mmsi": 0,
"vessel_legacy_unique_mmsi": 0,
"aisstream_connection_state": aisstream_health.connection_state if aisstream_health else None,
"aisstream_last_seen_at": to_iso8601_utc(aisstream_health.last_seen_at) if aisstream_health else None,
"aisstream_message_rate": aisstream_health.message_rate if aisstream_health else None,

View File

@@ -6,11 +6,15 @@ from typing import Optional
from fastapi import APIRouter, WebSocket, WebSocketDisconnect, Query
from jose import jwt, JWTError
from sqlalchemy import text
from app.core.config import settings
from app.core.enums import UserRole
from app.core.logging import get_logger
from app.core.time import to_iso8601_utc
from app.core.websocket.manager import manager
from app.db.session import async_session_factory
from app.services.log_tail import LOG_TAIL_CHANNEL, log_tail_manager
logger = get_logger(__name__, service="api")
router = APIRouter()
@@ -37,6 +41,28 @@ async def authenticate_token(token: str) -> Optional[dict]:
return None
async def load_websocket_user_role(user_id: str | None) -> str | None:
if not user_id:
return None
try:
async with async_session_factory() as db:
result = await db.execute(
text("SELECT role, is_active FROM users WHERE id = :id"),
{"id": int(user_id)},
)
row = result.fetchone()
except Exception as exc:
logger.warning_event(
"WebSocket user role lookup failed",
event="auth.websocket.role_lookup_failed",
context={"user_id": user_id, "error": str(exc)},
)
return None
if row is None or not row[1]:
return None
return str(row[0] or "")
@router.websocket("/ws")
async def websocket_endpoint(
websocket: WebSocket,
@@ -59,6 +85,7 @@ async def websocket_endpoint(
is_anonymous = payload is None
user_id = str(payload.get("sub")) if payload else f"anonymous:{id(websocket)}"
user_role = await load_websocket_user_role(user_id) if payload else None
supported_channels = ["vessels", "earth_news", EARTH_UPDATES_CHANNEL] if is_anonymous else [
"gpu_clusters",
"submarine_cables",
@@ -70,6 +97,8 @@ async def websocket_endpoint(
"earth_news",
EARTH_UPDATES_CHANNEL,
]
if user_role == UserRole.SUPER_ADMIN.value:
supported_channels = [*supported_channels, LOG_TAIL_CHANNEL]
await manager.connect(websocket, user_id)
try:
@@ -100,6 +129,7 @@ async def websocket_endpoint(
payload_data = data.get("data", {})
if not isinstance(payload_data, dict):
payload_data = {}
log_tail_config = None
channels = payload_data.get("channels", [])
if isinstance(channels, str):
channels = [channels]
@@ -108,6 +138,26 @@ async def websocket_endpoint(
channel = payload_data.get("channel")
if channel and channel not in channels:
channels = [*channels, channel]
if LOG_TAIL_CHANNEL in channels:
if user_role != UserRole.SUPER_ADMIN.value:
await websocket.send_json(
{
"type": "subscription_error",
"data": {"channel": LOG_TAIL_CHANNEL, "detail": "Only super_admin can subscribe logs"},
}
)
channels = [item for item in channels if item != LOG_TAIL_CHANNEL]
else:
try:
log_tail_config = await log_tail_manager.subscribe(websocket, payload_data)
except ValueError as exc:
await websocket.send_json(
{
"type": "subscription_error",
"data": {"channel": LOG_TAIL_CHANNEL, "detail": str(exc)},
}
)
channels = [item for item in channels if item != LOG_TAIL_CHANNEL]
if is_anonymous:
channels = [channel for channel in channels if channel in supported_channels]
vessel_subscription = None
@@ -131,14 +181,20 @@ async def websocket_endpoint(
"action": "subscribe",
"channels": [
*channels,
*([LOG_TAIL_CHANNEL] if log_tail_config else []),
*(["vessels"] if vessel_subscription else []),
],
"vessels": vessel_subscription,
"logs_tail": log_tail_config.__dict__ if log_tail_config else None,
},
}
)
elif data.get("type") == "unsubscribe":
channels = data.get("data", {}).get("channels", [])
if isinstance(channels, str):
channels = [channels]
if LOG_TAIL_CHANNEL in channels:
await log_tail_manager.unsubscribe(websocket)
manager.unsubscribe(websocket, channels)
await websocket.send_json(
{
@@ -159,4 +215,5 @@ async def websocket_endpoint(
except WebSocketDisconnect:
pass
finally:
await log_tail_manager.disconnect(websocket)
manager.disconnect(websocket, user_id)

View File

@@ -41,6 +41,7 @@ class Settings(BaseSettings):
AI_PROVIDER_SERVICE_TOKEN: str = ""
AI_PROVIDER_TIMEOUT_SECONDS: int = 60
AI_PROVIDER_RETRY_ATTEMPTS: int = 2
OBSERVABILITY_INGEST_TOKEN: str = ""
@property
def REDIS_URL(self) -> str:

226
backend/app/core/enums.py Normal file
View File

@@ -0,0 +1,226 @@
"""Stable backend protocol enums.
Database columns and JSON payloads continue to store the enum string values.
Configurable identifiers, user-authored values, and open-ended taxonomies do
not belong in this module.
"""
from __future__ import annotations
import logging
from enum import StrEnum
from typing import TypeVar
logger = logging.getLogger(__name__)
EnumT = TypeVar("EnumT", bound=StrEnum)
def parse_enum(enum_type: type[EnumT], value: object, default: EnumT) -> EnumT:
"""Parse an external value without breaking reads of legacy data."""
if value is None or str(value).strip() == "":
return default
if isinstance(value, enum_type):
return value
try:
return enum_type(str(value).strip().lower())
except (TypeError, ValueError):
logger.warning(
"Unknown %s value %r; falling back to %s",
enum_type.__name__,
value,
default.value,
)
return default
class NewsImportanceLevel(StrEnum):
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
CRITICAL = "critical"
class BreakingLevel(StrEnum):
NONE = "none"
WATCH = "watch"
BREAKING = "breaking"
CRITICAL = "critical"
class BreakingScope(StrEnum):
REGIONAL = "regional"
GLOBAL = "global"
class BreakingSource(StrEnum):
RULES = "rules"
AI = "ai"
MANUAL = "manual"
MULTI_SOURCE = "multi_source"
class NewsSourceType(StrEnum):
RSS = "rss"
ATOM = "atom"
AGGREGATED = "aggregated"
REFERENCE = "reference"
MANUAL = "manual"
class NewsEnrichmentStatus(StrEnum):
PENDING = "pending"
QUEUED = "queued"
ATTEMPTED = "attempted"
SUCCESS = "success"
CONTENT_ONLY = "content_only"
LOCATION_ONLY = "location_only"
UNAVAILABLE = "unavailable"
PROVIDER_ERROR = "provider_error"
PARSE_ERROR = "parse_error"
NO_RESULT = "no_result"
class NewsMarketImpact(StrEnum):
NONE = "none"
SECTOR = "sector"
NATIONAL = "national"
GLOBAL = "global"
class NewsTaggingSource(StrEnum):
RULES = "rules"
AI = "ai"
MANUAL = "manual"
class JobType(StrEnum):
COLLECT = "collect"
CLEAR_DATA = "clear_data"
CLEAR_CACHE = "clear_cache"
EARTH_REFRESH = "earth_refresh"
class JobStatus(StrEnum):
QUEUED = "queued"
RUNNING = "running"
CANCELLING = "cancelling"
SUCCESS = "success"
FAILED = "failed"
CANCELLED = "cancelled"
class RollbackPolicy(StrEnum):
KEEP_COMMITTED_BATCHES = "keep_committed_batches"
class MappingValidationStatus(StrEnum):
DRAFT = "draft"
VALID = "valid"
INVALID = "invalid"
class SnapshotStatus(StrEnum):
RUNNING = "running"
SUCCESS = "success"
FAILED = "failed"
CANCELLED = "cancelled"
class DatasourceRunStatus(StrEnum):
RUNNING = "running"
NOT_RUN = "not_run"
COLLECTED = "collected"
UNCOLLECTED = "uncollected"
class ProviderApi(StrEnum):
ANTHROPIC_MESSAGES = "anthropic-messages"
OPENAI_COMPLETIONS = "openai-completions"
OLLAMA_GENERATE = "ollama-generate"
class PlaygroundMessageRole(StrEnum):
SYSTEM = "system"
USER = "user"
ASSISTANT = "assistant"
TOOL = "tool"
class PlaygroundMessageKind(StrEnum):
MESSAGE = "message"
THINKING = "thinking"
ERROR = "error"
STATUS = "status"
class PlaygroundMessageStatus(StrEnum):
PENDING = "pending"
THINKING = "thinking"
ANSWERING = "answering"
DONE = "done"
FAILED = "failed"
CANCELLED = "cancelled"
ERROR = "error"
STOPPED = "stopped"
class OtpPurpose(StrEnum):
REGISTER = "register"
VERIFY_EMAIL = "verify_email"
RESET_PASSWORD = "reset_password"
class UserRole(StrEnum):
VIEWER = "viewer"
ADMIN = "admin"
SUPER_ADMIN = "super_admin"
class AlertSeverity(StrEnum):
CRITICAL = "critical"
WARNING = "warning"
INFO = "info"
class AlertStatus(StrEnum):
ACTIVE = "active"
ACKNOWLEDGED = "acknowledged"
RESOLVED = "resolved"
class BGPStatus(StrEnum):
ACTIVE = "active"
ACKNOWLEDGED = "acknowledged"
RESOLVED = "resolved"
class LogLevel(StrEnum):
ALL = "all"
ERROR = "error"
WARNING = "warning"
INFO = "info"
DEBUG = "debug"
class ConnectionState(StrEnum):
DISCONNECTED = "disconnected"
CONNECTING = "connecting"
CONNECTED = "connected"
ERROR = "error"
class AuthType(StrEnum):
NONE = "none"
BEARER = "bearer"
API_KEY = "api_key"
BASIC = "basic"
class TVSourceType(StrEnum):
IFRAME = "iframe"
HLS = "hls"
VIDEO = "video"
EXTERNAL = "external"
YOUTUBE = "youtube"

View File

@@ -626,6 +626,7 @@ async def init_db():
"bgp_collector_locations",
"vessel_static",
"vessel_position",
"vessel_current_state",
"ais_raw_observations",
"ais_source_health",
"compute_center_locations",
@@ -786,6 +787,22 @@ async def init_db():
"""
)
)
await conn.execute(
text(
"""
CREATE INDEX IF NOT EXISTS idx_vessel_current_bbox
ON vessel_current_state (lon, lat)
"""
)
)
await conn.execute(
text(
"""
CREATE INDEX IF NOT EXISTS idx_vessel_current_observed
ON vessel_current_state (observed_at DESC)
"""
)
)
await conn.execute(
text(
"""

View File

@@ -13,7 +13,7 @@ from app.models.compute_center_location import ComputeCenterLocationRecord
from app.models.system_setting import SystemSetting
from app.models.playground_session import PlaygroundSession
from app.models.playground_message import PlaygroundMessage
from app.models.system_log import SystemLog, AuditLog
from app.models.system_log import AuditLog, ObservabilityEvent, ObservabilityEventGroup, SystemLog
from app.models.vessel import AISConflictRecord, AISRawObservation, AISSourceHealth, VesselPosition, VesselStatic
from app.models.datasource_mapping import DataSourceMappingTemplate
from app.models.earth_news import EarthNewsItem
@@ -37,6 +37,8 @@ __all__ = [
"ComputeCenterLocationRecord",
"SystemLog",
"AuditLog",
"ObservabilityEvent",
"ObservabilityEventGroup",
"PlaygroundSession",
"PlaygroundMessage",
"VesselPosition",

View File

@@ -1,26 +1,12 @@
from datetime import datetime
from enum import Enum
from typing import Optional
from sqlalchemy import Column, Integer, String, DateTime, Text, ForeignKey, Enum as SQLEnum
from sqlalchemy.orm import relationship
from sqlalchemy import Column, Integer, String, DateTime, Text, Enum as SQLEnum
from app.core.enums import AlertSeverity, AlertStatus
from app.core.time import to_iso8601_utc
from app.db.session import Base
class AlertSeverity(str, Enum):
CRITICAL = "critical"
WARNING = "warning"
INFO = "info"
class AlertStatus(str, Enum):
ACTIVE = "active"
ACKNOWLEDGED = "acknowledged"
RESOLVED = "resolved"
class Alert(Base):
__tablename__ = "alerts"

View File

@@ -4,6 +4,7 @@ from datetime import datetime
from sqlalchemy import Column, DateTime, Float, ForeignKey, Index, Integer, JSON, String, Text
from app.core.enums import BGPStatus
from app.core.time import to_iso8601_utc
from app.db.session import Base
@@ -17,7 +18,7 @@ class BGPAnomaly(Base):
source = Column(String(100), nullable=False, index=True)
anomaly_type = Column(String(50), nullable=False, index=True)
severity = Column(String(20), nullable=False, index=True)
status = Column(String(20), nullable=False, default="active", index=True)
status = Column(String(20), nullable=False, default=BGPStatus.ACTIVE.value, index=True)
entity_key = Column(String(255), nullable=False, index=True)
prefix = Column(String(64), nullable=True, index=True)
origin_asn = Column(Integer, nullable=True, index=True)

View File

@@ -4,6 +4,7 @@ from datetime import datetime
from sqlalchemy import Column, DateTime, Float, ForeignKey, Index, Integer, JSON, String, Text
from app.core.enums import BGPStatus
from app.core.time import to_iso8601_utc
from app.db.session import Base
@@ -20,7 +21,7 @@ class BGPIncident(Base):
title = Column(String(255), nullable=False)
summary = Column(Text, nullable=False)
severity = Column(String(20), nullable=False, index=True)
status = Column(String(20), nullable=False, default="active", index=True)
status = Column(String(20), nullable=False, default=BGPStatus.ACTIVE.value, index=True)
confidence = Column(Float, nullable=False, default=0.5)
started_at = Column(DateTime(timezone=True), nullable=False, default=datetime.utcnow, index=True)
ended_at = Column(DateTime(timezone=True), nullable=True)

View File

@@ -1,6 +1,7 @@
from sqlalchemy import Boolean, Column, DateTime, ForeignKey, Integer, JSON, String
from sqlalchemy.sql import func
from app.core.enums import SnapshotStatus
from app.db.session import Base
@@ -16,7 +17,7 @@ class DataSnapshot(Base):
started_at = Column(DateTime(timezone=True), server_default=func.now())
completed_at = Column(DateTime(timezone=True), nullable=True)
record_count = Column(Integer, default=0)
status = Column(String(20), nullable=False, default="running")
status = Column(String(20), nullable=False, default=SnapshotStatus.RUNNING.value)
is_current = Column(Boolean, default=True, index=True)
parent_snapshot_id = Column(Integer, ForeignKey("data_snapshots.id"), nullable=True, index=True)
summary = Column(JSON, default={})

View File

@@ -3,6 +3,7 @@
from sqlalchemy import Boolean, Column, DateTime, ForeignKey, Integer, JSON, String
from sqlalchemy.sql import func
from app.core.enums import MappingValidationStatus
from app.db.session import Base
@@ -19,7 +20,7 @@ class DataSourceMappingTemplate(Base):
target_schema = Column(String(80), nullable=False, index=True)
mapping_json = Column(JSON, nullable=False, default={})
sample_payload_hash = Column(String(64), nullable=True)
validation_status = Column(String(30), nullable=False, default="draft")
validation_status = Column(String(30), nullable=False, default=MappingValidationStatus.DRAFT.value)
version = Column(Integer, nullable=False, default=1)
is_active = Column(Boolean, nullable=False, default=False, index=True)
created_at = Column(DateTime(timezone=True), server_default=func.now())

View File

@@ -1,6 +1,7 @@
from sqlalchemy import JSON, Boolean, Column, DateTime, ForeignKey, Integer, String, Text
from sqlalchemy.sql import func
from app.core.enums import PlaygroundMessageKind, PlaygroundMessageStatus
from app.db.session import Base
@@ -13,8 +14,8 @@ class PlaygroundMessage(Base):
user_id = Column(Integer, ForeignKey("users.id", ondelete="CASCADE"), nullable=False, index=True)
parent_message_id = Column(Integer, ForeignKey("playground_messages.id", ondelete="SET NULL"), nullable=True)
role = Column(String(20), nullable=False)
kind = Column(String(20), nullable=False, default="message")
status = Column(String(20), nullable=False, default="done")
kind = Column(String(20), nullable=False, default=PlaygroundMessageKind.MESSAGE.value)
status = Column(String(20), nullable=False, default=PlaygroundMessageStatus.DONE.value)
title = Column(String(255), nullable=True)
content = Column(Text, nullable=False, default="")
thinking_content = Column(Text, nullable=False, default="")

View File

@@ -38,3 +38,46 @@ class AuditLog(Base):
ip = Column(String(64), nullable=True)
details = Column(JSON, nullable=False, default=dict)
created_at = Column(DateTime(timezone=True), server_default=func.now())
class ObservabilityEvent(Base):
__tablename__ = "observability_events"
id = Column(Integer, primary_key=True, autoincrement=True)
occurred_at = Column(DateTime(timezone=True), server_default=func.now(), index=True)
source = Column(String(50), nullable=False, index=True)
service = Column(String(50), nullable=True, index=True)
module = Column(String(120), nullable=True, index=True)
category = Column(String(80), nullable=True, index=True)
event = Column(String(160), nullable=True, index=True)
level = Column(String(20), nullable=False, index=True)
message = Column(Text, nullable=False)
fingerprint = Column(String(80), nullable=False, index=True)
request_id = Column(String(64), nullable=True, index=True)
trace_id = Column(String(64), nullable=True, index=True)
task_id = Column(String(120), nullable=True, index=True)
source_ref_id = Column(String(120), nullable=True, index=True)
provider = Column(String(120), nullable=True, index=True)
user_id = Column(Integer, nullable=True, index=True)
context = Column(JSON, nullable=False, default=dict)
occurrence_count = Column(Integer, nullable=False, default=1)
created_at = Column(DateTime(timezone=True), server_default=func.now())
class ObservabilityEventGroup(Base):
__tablename__ = "observability_event_groups"
fingerprint = Column(String(80), primary_key=True)
source = Column(String(50), nullable=False, index=True)
service = Column(String(50), nullable=True, index=True)
module = Column(String(120), nullable=True, index=True)
category = Column(String(80), nullable=True, index=True)
event = Column(String(160), nullable=True, index=True)
last_level = Column(String(20), nullable=False, index=True)
sample_message = Column(Text, nullable=False)
sample_detail = Column(Text, nullable=True)
affected_sources = Column(JSON, nullable=False, default=list)
count = Column(Integer, nullable=False, default=0)
first_seen_at = Column(DateTime(timezone=True), nullable=False, index=True)
last_seen_at = Column(DateTime(timezone=True), nullable=False, index=True)
updated_at = Column(DateTime(timezone=True), server_default=func.now(), onupdate=func.now())

View File

@@ -3,6 +3,7 @@
from sqlalchemy import BigInteger, Column, DateTime, Float, Integer, JSON, String, Text
from sqlalchemy.sql import func
from app.core.enums import JobStatus, JobType, RollbackPolicy
from app.db.session import Base
@@ -12,9 +13,9 @@ class CollectionTask(Base):
id = Column(Integer, primary_key=True, autoincrement=True)
datasource_id = Column(Integer, nullable=False, index=True)
source = Column(String(100), nullable=True, index=True)
task_type = Column(String(30), nullable=False, default="collect", index=True)
task_type = Column(String(30), nullable=False, default=JobType.COLLECT.value, index=True)
status = Column(String(20), nullable=False) # queued, running, cancelling, success, failed, cancelled
phase = Column(String(30), default="queued")
phase = Column(String(30), default=JobStatus.QUEUED.value)
phase_progress = Column(Float)
phase_message = Column(String(255))
phase_current = Column(BigInteger)
@@ -27,7 +28,7 @@ class CollectionTask(Base):
progress = Column(Float, default=0.0) # Progress percentage (0-100)
error_message = Column(Text)
payload = Column(JSON, default=dict)
rollback_policy = Column(String(40), nullable=False, default="keep_committed_batches")
rollback_policy = Column(String(40), nullable=False, default=RollbackPolicy.KEEP_COMMITTED_BATCHES.value)
dedupe_key = Column(String(180), nullable=True, index=True)
worker_id = Column(String(120), nullable=True, index=True)
locked_at = Column(DateTime(timezone=True), nullable=True, index=True)

View File

@@ -1,6 +1,7 @@
from sqlalchemy import Boolean, Column, DateTime, Integer, JSON, String
from sqlalchemy.sql import func
from app.core.enums import UserRole
from app.db.session import Base
@@ -11,7 +12,7 @@ class User(Base):
username = Column(String(50), unique=True, index=True, nullable=False)
email = Column(String(255), unique=True, index=True, nullable=False)
password_hash = Column(String(255), nullable=False)
role = Column(String(20), default="viewer")
role = Column(String(20), default=UserRole.VIEWER.value)
gatekeeper_groups = Column(JSON, default=list)
is_active = Column(Boolean, default=True)
email_verified = Column(Boolean, default=False, nullable=False)

View File

@@ -3,6 +3,7 @@
from sqlalchemy import BigInteger, Column, DateTime, Float, Index, Integer, JSON, SmallInteger, String
from sqlalchemy.sql import func
from app.core.enums import ConnectionState
from app.core.time import to_iso8601_utc
from app.db.session import Base
@@ -75,6 +76,67 @@ class VesselPosition(Base):
}
class VesselCurrentState(Base):
"""Latest renderable state for one vessel, independent from AIS history."""
__tablename__ = "vessel_current_state"
mmsi = Column(BigInteger, primary_key=True)
lat = Column(Float, nullable=False)
lon = Column(Float, nullable=False)
sog = Column(Float, nullable=True)
cog = Column(Float, nullable=True)
heading = Column(SmallInteger, nullable=True)
nav_status = Column(SmallInteger, nullable=True, index=True)
name = Column(String(128), nullable=True)
callsign = Column(String(16), nullable=True)
vessel_type = Column(SmallInteger, nullable=True, index=True)
vessel_type_name = Column(String(64), nullable=True, index=True)
flag = Column(String(4), nullable=True, index=True)
length = Column(Float, nullable=True)
width = Column(Float, nullable=True)
draught = Column(Float, nullable=True)
imo = Column(BigInteger, nullable=True)
source = Column(String(100), nullable=False, index=True)
observed_at = Column(DateTime(timezone=True), nullable=False)
field_sources = Column(JSON, default=dict)
selected_reasons = Column(JSON, default=dict)
source_summary = Column(JSON, default=dict)
quality_flags = Column(JSON, default=list)
updated_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now())
__table_args__ = (
Index("idx_vessel_current_bbox", "lon", "lat"),
Index("idx_vessel_current_observed", "observed_at"),
)
def to_dict(self) -> dict:
return {
"mmsi": self.mmsi,
"lat": self.lat,
"lon": self.lon,
"sog": self.sog,
"cog": self.cog,
"heading": self.heading,
"nav_status": self.nav_status,
"name": self.name,
"callsign": self.callsign,
"vessel_type": self.vessel_type,
"vessel_type_name": self.vessel_type_name,
"flag": self.flag,
"length": self.length,
"width": self.width,
"draught": self.draught,
"imo": self.imo,
"source": self.source,
"received_at": self.observed_at,
"field_sources": self.field_sources or {},
"selected_reasons": self.selected_reasons or {},
"source_summary": self.source_summary or {},
"quality_flags": self.quality_flags or [],
}
class AISRawObservation(Base):
"""Source-level AIS fact before aggregation and conflict resolution."""
@@ -165,7 +227,7 @@ class AISSourceHealth(Base):
__tablename__ = "ais_source_health"
source = Column(String(100), primary_key=True)
connection_state = Column(String(32), nullable=False, default="disconnected", index=True)
connection_state = Column(String(32), nullable=False, default=ConnectionState.DISCONNECTED.value, index=True)
last_seen_at = Column(DateTime(timezone=True), nullable=True, index=True)
last_success_at = Column(DateTime(timezone=True), nullable=True, index=True)
last_error = Column(String(500), nullable=True)

View File

@@ -2,6 +2,7 @@ from typing import Any
from pydantic import BaseModel, Field
from app.core.enums import PlaygroundMessageKind, PlaygroundMessageRole, PlaygroundMessageStatus
class AIContentBlock(BaseModel):
type: str
@@ -110,9 +111,9 @@ class PlaygroundSessionUpsertRequest(BaseModel):
class PlaygroundMessageRecord(BaseModel):
id: str
role: str
kind: str = "message"
status: str = "done"
role: PlaygroundMessageRole
kind: PlaygroundMessageKind = PlaygroundMessageKind.MESSAGE
status: PlaygroundMessageStatus = PlaygroundMessageStatus.DONE
title: str | None = None
content: str = ""
thinking_content: str = ""

View File

@@ -3,6 +3,7 @@ from typing import Optional
from pydantic import BaseModel, EmailStr, Field
from app.core.enums import OtpPurpose, UserRole
class UserBase(BaseModel):
username: str
@@ -11,13 +12,13 @@ class UserBase(BaseModel):
class UserCreate(UserBase):
password: str = Field(..., min_length=8)
role: str = "viewer"
role: UserRole = UserRole.VIEWER
gatekeeper_groups: list[str] = Field(default_factory=list)
class UserUpdate(BaseModel):
email: Optional[EmailStr] = None
role: Optional[str] = None
role: Optional[UserRole] = None
gatekeeper_groups: Optional[list[str]] = None
is_active: Optional[bool] = None
@@ -59,7 +60,7 @@ class VerifyEmailRequest(BaseModel):
class ResendCodeRequest(BaseModel):
email: EmailStr
purpose: str = Field(default="register", pattern="^(register|verify_email|reset_password)$")
purpose: OtpPurpose = OtpPurpose.REGISTER
class ForgotPasswordRequest(BaseModel):

View File

@@ -6,6 +6,7 @@ from collections import Counter, defaultdict
from datetime import UTC, datetime
from typing import Any
from app.core.enums import BGPStatus
from app.models.bgp_anomaly import BGPAnomaly
@@ -127,7 +128,7 @@ def detect_origin_change_anomalies(
source=source,
anomaly_type=anomaly_type,
severity=severity,
status="active",
status=BGPStatus.ACTIVE.value,
entity_key=f"{anomaly_type}:{prefix}:{new_origin}",
prefix=prefix,
origin_asn=sorted(historic)[0] if historic else None,
@@ -197,7 +198,7 @@ def detect_more_specific_burst_anomalies(
source=source,
anomaly_type="more_specific_burst",
severity="high",
status="active",
status=BGPStatus.ACTIVE.value,
entity_key=f"more_specific_burst:{root_prefix}:{len(unique_prefixes)}:{len(related_collectors)}",
prefix=sample.get("prefix"),
origin_asn=sample.get("origin_asn"),
@@ -267,7 +268,7 @@ def detect_mass_withdrawal_anomalies(
source=source,
anomaly_type="mass_withdrawal",
severity=severity,
status="active",
status=BGPStatus.ACTIVE.value,
entity_key=f"mass_withdrawal:{prefix}:{origin_asn}:{len(related_collectors)}:{count}",
prefix=prefix,
origin_asn=origin_asn,
@@ -354,7 +355,7 @@ def detect_route_leak_anomalies(
source=source,
anomaly_type="route_leak_candidate",
severity="high" if max_path_length >= dominant_length + 3 else "medium",
status="active",
status=BGPStatus.ACTIVE.value,
entity_key=f"route_leak_candidate:{prefix}:{max_path_length}:{len(related_collectors)}",
prefix=prefix,
origin_asn=sample_metadata.get("origin_asn"),
@@ -435,7 +436,7 @@ def detect_path_flap_anomalies(
source=source,
anomaly_type="path_flap",
severity=severity,
status="active",
status=BGPStatus.ACTIVE.value,
entity_key=f"path_flap:{prefix}:{transitions}:{len(distinct_paths)}",
prefix=prefix,
origin_asn=sample_metadata.get("origin_asn"),

View File

@@ -9,6 +9,7 @@ from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.collected_data_fields import get_record_field
from app.core.enums import BGPStatus
from app.models.bgp_anomaly import BGPAnomaly
from app.models.bgp_incident import BGPIncident
from app.models.collected_data import CollectedData
@@ -290,7 +291,7 @@ async def create_bgp_incidents_for_anomalies(
existing.title = title
existing.summary = summary
existing.severity = severity
existing.status = "active"
existing.status = BGPStatus.ACTIVE.value
existing.confidence = confidence
existing.started_at = primary.started_at or existing.started_at or datetime.now(UTC)
existing.ended_at = None
@@ -313,7 +314,7 @@ async def create_bgp_incidents_for_anomalies(
title=title,
summary=summary,
severity=severity,
status="active",
status=BGPStatus.ACTIVE.value,
confidence=confidence,
started_at=primary.started_at or datetime.now(UTC),
affected_prefixes=prefixes,

View File

@@ -322,6 +322,21 @@ class AISStreamCollector(BaseCollector):
last_success_at=now if data else None,
lag_seconds=max((now - latest_observed_at).total_seconds(), 0),
)
if snapshot_id is not None:
from app.models.data_snapshot import DataSnapshot
snapshot = await db.get(DataSnapshot, snapshot_id)
if snapshot:
snapshot.record_count = records_added
snapshot.status = "success"
snapshot.completed_at = now
snapshot.summary = {
"created": records_added,
"updated": 0,
"unchanged": 0,
"deleted": 0,
"storage": "ais_raw_observations",
}
await db.commit()
await self.update_progress(records_added, force=True)
return records_added

View File

@@ -12,11 +12,11 @@ from sqlalchemy.ext.asyncio import AsyncSession
from app.core.collected_data_fields import build_dynamic_metadata, get_record_field
from app.core.countries import normalize_country
from app.core.enums import JobStatus, SnapshotStatus
from app.core.logging import get_logger
from app.core.time import to_iso8601_utc
from app.core.websocket.broadcaster import broadcaster
from app.services.business_logs import emit_business_log, exception_context
from app.services.earth_layer_adapters import get_earth_update_layers_for_source
logger = get_logger(__name__, service="collector")
@@ -238,7 +238,7 @@ class BaseCollector(ABC):
snapshot = await db.get(DataSnapshot, snapshot_id)
if snapshot:
parent_snapshot_id = snapshot.parent_snapshot_id
snapshot.status = "cancelled"
snapshot.status = SnapshotStatus.CANCELLED.value
snapshot.is_current = False
snapshot.completed_at = datetime.now(UTC)
summary = dict(snapshot.summary or {})
@@ -308,7 +308,7 @@ class BaseCollector(ABC):
task.datasource_id = datasource_id
task.source = task.source or self.name
task.task_type = task.task_type or "collect"
task.status = "running"
task.status = JobStatus.RUNNING.value
task.phase = "queued"
task.started_at = task.started_at or start_time
task.completed_at = None
@@ -393,7 +393,7 @@ class BaseCollector(ABC):
},
)
task.status = "success"
task.status = JobStatus.SUCCESS.value
task.phase = "completed"
task.phase_progress = 100.0
task.phase_message = "采集完成"
@@ -428,7 +428,7 @@ class BaseCollector(ABC):
}
except asyncio.CancelledError:
await db.rollback()
task.status = "cancelled"
task.status = JobStatus.CANCELLED.value
task.phase = "cancelled"
task.phase_message = "采集已取消"
task.error_message = "Collection cancelled by operator and rolled back"
@@ -456,7 +456,7 @@ class BaseCollector(ABC):
raise
except Exception as e:
await db.rollback()
task.status = "failed"
task.status = JobStatus.FAILED.value
task.phase = "failed"
task.phase_message = str(e)
task.error_message = str(e)
@@ -464,7 +464,7 @@ class BaseCollector(ABC):
if snapshot_id is not None:
snapshot = await db.get(DataSnapshot, snapshot_id)
if snapshot:
snapshot.status = "failed"
snapshot.status = SnapshotStatus.FAILED.value
snapshot.completed_at = datetime.now(UTC)
snapshot.summary = {"error": str(e)}
await db.commit()
@@ -509,7 +509,7 @@ class BaseCollector(ABC):
if snapshot:
snapshot.record_count = 0
snapshot.summary = {"created": 0, "updated": 0, "unchanged": 0}
snapshot.status = "success"
snapshot.status = SnapshotStatus.SUCCESS.value
snapshot.completed_at = datetime.now(UTC)
await db.commit()
return 0
@@ -642,7 +642,7 @@ class BaseCollector(ABC):
snapshot = await db.get(DataSnapshot, snapshot_id)
if snapshot:
snapshot.record_count = records_added
snapshot.status = "success"
snapshot.status = SnapshotStatus.SUCCESS.value
snapshot.completed_at = datetime.now(UTC)
snapshot.summary = {
"created": created_count,

View File

@@ -25,11 +25,17 @@ FALLBACK_GROUPS = (
"starlink",
"gps-ops",
"galileo",
"glonass",
"glo-ops",
"beidou",
"leo",
"geo",
"iridium-next",
"stations",
"visual",
"weather",
"science",
"cubesat",
"amateur",
"last-30-days",
)
FETCH_RETRY_ATTEMPTS = 3
FETCH_RETRY_BASE_DELAY_SECONDS = 0.8
@@ -220,27 +226,28 @@ class CelesTrakTLECollector(BaseCollector):
try:
for group in FALLBACK_GROUPS:
group_url = self._group_url(group)
try:
body_path = await self._downloader.download_file(
client,
group_url,
extension=".json",
accept="application/json",
validate_existing=self._validate_json_file,
)
except DownloadHTTPStatusError as exc:
if not self._is_not_updated_response(exc):
raise RuntimeError(f"CelesTrak fallback group '{group}' download failed: {exc}") from exc
cached_path = self._downloader.get_cached_file(
group_url,
".json",
validate_existing=self._validate_json_file,
)
if cached_path is None:
cached_path = self._downloader.get_cached_file(
group_url,
".json",
validate_existing=self._validate_json_file,
)
if cached_path is not None:
body_path = cached_path
else:
try:
body_path = await self._downloader.download_file(
client,
group_url,
extension=".json",
accept="application/json",
validate_existing=self._validate_json_file,
)
except DownloadHTTPStatusError as exc:
if not self._is_not_updated_response(exc):
raise RuntimeError(f"CelesTrak fallback group '{group}' download failed: {exc}") from exc
raise RuntimeError(
f"CelesTrak fallback group '{group}' has not updated and no local cached copy is available"
) from exc
body_path = cached_path
group_records = await self._load_downloaded_payload(
body_path,

View File

@@ -11,17 +11,20 @@ To get higher limits, set PEERINGDB_API_KEY environment variable.
"""
import asyncio
import os
from typing import Dict, Any, List
from datetime import UTC, datetime
import os
from typing import Any, Dict, List
from urllib.parse import urlencode
import httpx
from urllib.parse import urlencode
from app.core.logging import get_logger
from app.services.collectors.base import HTTPCollector
# PeeringDB API key - read from environment variable
PEERINGDB_API_KEY = os.environ.get("PEERINGDB_API_KEY", "")
logger = get_logger(__name__, service="collector")
class PeeringDBIXPCollector(HTTPCollector):
@@ -39,6 +42,7 @@ class PeeringDBIXPCollector(HTTPCollector):
"User-Agent": "Planet-Intelligence-System/1.0 (Python/collector)",
"Accept": "application/json",
}
@property
def request_url(self) -> str:
base = self._resolved_url or self.base_url
@@ -61,7 +65,11 @@ class PeeringDBIXPCollector(HTTPCollector):
if response.status_code == 429:
# Rate limited - wait and retry with exponential backoff
delay = base_delay * (2**attempt)
print(f"PeeringDB rate limited, waiting {delay}s before retry...")
logger.warning_event(
"PeeringDB rate limited; retrying after delay",
event="collector.peeringdb.rate_limited",
context={"delay_seconds": delay, "attempt": attempt + 1},
)
await asyncio.sleep(delay)
last_error = "Rate limited"
continue
@@ -72,13 +80,21 @@ class PeeringDBIXPCollector(HTTPCollector):
except httpx.HTTPStatusError as e:
if e.response.status_code == 429:
delay = base_delay * (2**attempt)
print(f"PeeringDB rate limited, waiting {delay}s before retry...")
logger.warning_event(
"PeeringDB rate limited; retrying after delay",
event="collector.peeringdb.rate_limited",
context={"delay_seconds": delay, "attempt": attempt + 1},
)
await asyncio.sleep(delay)
last_error = "Rate limited"
continue
raise
print(f"Warning: PeeringDB collection failed after {max_retries} retries: {last_error}")
logger.warning_event(
"PeeringDB collection failed after retries",
event="collector.peeringdb.retries_exhausted",
context={"max_retries": max_retries, "last_error": last_error},
)
return {}
async def fetch(self) -> List[Dict[str, Any]]:
@@ -146,6 +162,7 @@ class PeeringDBNetworkCollector(HTTPCollector):
"User-Agent": "Planet-Intelligence-System/1.0 (Python/collector)",
"Accept": "application/json",
}
@property
def request_url(self) -> str:
base = self._resolved_url or self.base_url
@@ -167,7 +184,11 @@ class PeeringDBNetworkCollector(HTTPCollector):
if response.status_code == 429:
delay = base_delay * (2**attempt)
print(f"PeeringDB rate limited, waiting {delay}s before retry...")
logger.warning_event(
"PeeringDB rate limited; retrying after delay",
event="collector.peeringdb.rate_limited",
context={"delay_seconds": delay, "attempt": attempt + 1},
)
await asyncio.sleep(delay)
last_error = "Rate limited"
continue
@@ -178,13 +199,21 @@ class PeeringDBNetworkCollector(HTTPCollector):
except httpx.HTTPStatusError as e:
if e.response.status_code == 429:
delay = base_delay * (2**attempt)
print(f"PeeringDB rate limited, waiting {delay}s before retry...")
logger.warning_event(
"PeeringDB rate limited; retrying after delay",
event="collector.peeringdb.rate_limited",
context={"delay_seconds": delay, "attempt": attempt + 1},
)
await asyncio.sleep(delay)
last_error = "Rate limited"
continue
raise
print(f"Warning: PeeringDB collection failed after {max_retries} retries: {last_error}")
logger.warning_event(
"PeeringDB collection failed after retries",
event="collector.peeringdb.retries_exhausted",
context={"max_retries": max_retries, "last_error": last_error},
)
return {}
async def fetch(self) -> List[Dict[str, Any]]:
@@ -254,6 +283,7 @@ class PeeringDBFacilityCollector(HTTPCollector):
"User-Agent": "Planet-Intelligence-System/1.0 (Python/collector)",
"Accept": "application/json",
}
@property
def request_url(self) -> str:
base = self._resolved_url or self.base_url
@@ -275,7 +305,11 @@ class PeeringDBFacilityCollector(HTTPCollector):
if response.status_code == 429:
delay = base_delay * (2**attempt)
print(f"PeeringDB rate limited, waiting {delay}s before retry...")
logger.warning_event(
"PeeringDB rate limited; retrying after delay",
event="collector.peeringdb.rate_limited",
context={"delay_seconds": delay, "attempt": attempt + 1},
)
await asyncio.sleep(delay)
last_error = "Rate limited"
continue
@@ -286,13 +320,21 @@ class PeeringDBFacilityCollector(HTTPCollector):
except httpx.HTTPStatusError as e:
if e.response.status_code == 429:
delay = base_delay * (2**attempt)
print(f"PeeringDB rate limited, waiting {delay}s before retry...")
logger.warning_event(
"PeeringDB rate limited; retrying after delay",
event="collector.peeringdb.rate_limited",
context={"delay_seconds": delay, "attempt": attempt + 1},
)
await asyncio.sleep(delay)
last_error = "Rate limited"
continue
raise
print(f"Warning: PeeringDB collection failed after {max_retries} retries: {last_error}")
logger.warning_event(
"PeeringDB collection failed after retries",
event="collector.peeringdb.retries_exhausted",
context={"max_retries": max_retries, "last_error": last_error},
)
return {}
async def fetch(self) -> List[Dict[str, Any]]:

View File

@@ -1,17 +1,21 @@
"""Space-Track TLE Collector
"""Space-Track TLE Collector.
Collects satellite TLE (Two-Line Element) data from Space-Track.org.
API documentation: https://www.space-track.org/documentation
"""
import json
from typing import Dict, Any, List
import httpx
from typing import Any, Dict, List
from urllib.parse import urlparse
from app.services.collectors.base import BaseCollector
import httpx
from app.core.data_sources import get_data_sources_config
from app.core.logging import get_logger
from app.core.satellite_tle import build_tle_lines_from_elements
from app.services.collectors.base import BaseCollector
logger = get_logger(__name__, service="collector")
class SpaceTrackTLECollector(BaseCollector):
@@ -53,10 +57,16 @@ class SpaceTrackTLECollector(BaseCollector):
password = settings.SPACETRACK_PASSWORD
if not username or not password:
print("SPACETRACK: No credentials configured, using sample data")
logger.warning_event(
"Space-Track credentials are not configured; using sample data",
event="collector.spacetrack.credentials_missing",
)
return self._get_sample_data()
print(f"SPACETRACK: Attempting to fetch TLE data with username: {username}")
logger.info_event(
"Space-Track TLE fetch started",
event="collector.spacetrack.fetch.start",
)
try:
async with httpx.AsyncClient(
@@ -78,11 +88,17 @@ class SpaceTrackTLECollector(BaseCollector):
"password": password,
},
)
print(f"SPACETRACK: Login response status: {login_response.status_code}")
print(f"SPACETRACK: Login response URL: {login_response.url}")
logger.info_event(
"Space-Track login response received",
event="collector.spacetrack.login.response",
context={"status_code": login_response.status_code},
)
if login_response.status_code == 403:
print("SPACETRACK: Trying alternate login method...")
logger.warning_event(
"Space-Track login returned forbidden; trying alternate method",
event="collector.spacetrack.login.forbidden",
)
async with httpx.AsyncClient(
timeout=120.0,
@@ -90,11 +106,6 @@ class SpaceTrackTLECollector(BaseCollector):
) as alt_client:
await alt_client.get(f"{self.site_root}/")
form_data = {
"username": username,
"password": password,
"query": "class/gp/NORAD_CAT_ID/25544/format/json",
}
alt_login = await alt_client.post(
self.login_url,
data={
@@ -102,77 +113,59 @@ class SpaceTrackTLECollector(BaseCollector):
"password": password,
},
)
print(f"SPACETRACK: Alt login status: {alt_login.status_code}")
logger.info_event(
"Space-Track alternate login response received",
event="collector.spacetrack.alt_login.response",
context={"status_code": alt_login.status_code},
)
if alt_login.status_code == 200:
tle_response = await alt_client.get(self.probe_url)
if tle_response.status_code == 200:
data = tle_response.json()
print(f"SPACETRACK: Received {len(data)} records via alt method")
logger.info_event(
"Space-Track alternate query completed",
event="collector.spacetrack.alt_query.completed",
context={"record_count": len(data)},
)
return data
if login_response.status_code != 200:
print(f"SPACETRACK: Login failed, using sample data")
logger.warning_event(
"Space-Track login failed; using sample data",
event="collector.spacetrack.login.failed",
context={"status_code": login_response.status_code},
)
return self._get_sample_data()
tle_response = await client.get(self.probe_url)
print(f"SPACETRACK: TLE query status: {tle_response.status_code}")
if tle_response.status_code != 200:
print(f"SPACETRACK: Query failed, using sample data")
return self._get_sample_data()
data = tle_response.json()
print(f"SPACETRACK: Received {len(data)} records")
return data
except Exception as e:
print(f"SPACETRACK: Error - {e}, using sample data")
return self._get_sample_data()
print(f"SPACETRACK: Attempting to fetch TLE data with username: {username}")
try:
async with httpx.AsyncClient(
timeout=120.0,
follow_redirects=True,
headers={
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36",
"Accept": "application/json, text/html, */*",
"Accept-Language": "en-US,en;q=0.9",
},
) as client:
# First, visit the main page to get any cookies
await client.get(f"{self.site_root}/")
# Login to get session cookie
login_response = await client.post(
self.login_url,
data={
"identity": username,
"password": password,
},
logger.info_event(
"Space-Track TLE query response received",
event="collector.spacetrack.query.response",
context={"status_code": tle_response.status_code},
)
print(f"SPACETRACK: Login response status: {login_response.status_code}")
print(f"SPACETRACK: Login response URL: {login_response.url}")
print(f"SPACETRACK: Login response body: {login_response.text[:500]}")
if login_response.status_code != 200:
print(f"SPACETRACK: Login failed, using sample data")
return self._get_sample_data()
# Query for TLE data (get first 1000 satellites)
tle_response = await client.get(self.query_url)
print(f"SPACETRACK: TLE query status: {tle_response.status_code}")
if tle_response.status_code != 200:
print(f"SPACETRACK: Query failed, using sample data")
logger.warning_event(
"Space-Track TLE query failed; using sample data",
event="collector.spacetrack.query.failed",
context={"status_code": tle_response.status_code},
)
return self._get_sample_data()
data = tle_response.json()
print(f"SPACETRACK: Received {len(data)} records")
logger.info_event(
"Space-Track TLE fetch completed",
event="collector.spacetrack.fetch.completed",
context={"record_count": len(data)},
)
return data
except Exception as e:
print(f"SPACETRACK: Error - {e}, using sample data")
logger.warning_event(
"Space-Track TLE fetch failed; using sample data",
event="collector.spacetrack.fetch.failed",
context={"error": str(e)},
)
return self._get_sample_data()
def transform(self, raw_data: List[Dict[str, Any]]) -> List[Dict[str, Any]]:

View File

@@ -6,6 +6,7 @@ from typing import Any
import httpx
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.enums import SnapshotStatus
from app.core.time import to_iso8601_utc
from app.core.websocket.broadcaster import broadcaster
from app.services.barentswatch import (
@@ -119,6 +120,21 @@ class VesselAISCollector(BaseCollector):
last_success_at=now if data else None,
lag_seconds=max((now - latest_observed_at).total_seconds(), 0),
)
if snapshot_id is not None:
from app.models.data_snapshot import DataSnapshot
snapshot = await db.get(DataSnapshot, snapshot_id)
if snapshot:
snapshot.record_count = records_added
snapshot.status = SnapshotStatus.SUCCESS.value
snapshot.completed_at = now
snapshot.summary = {
"created": records_added,
"updated": 0,
"unchanged": 0,
"deleted": 0,
"storage": "ais_raw_observations",
}
await db.commit()
await self._broadcast_vessel_snapshot(data)
await self.update_progress(records_added, force=True)

View File

@@ -12,6 +12,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from app.core.cache import cache
from app.core.config import settings
from app.core.enums import JobStatus, JobType, RollbackPolicy
from app.core.logging import get_logger
from app.core.time import to_iso8601_utc
from app.core.websocket.broadcaster import broadcaster
@@ -36,17 +37,17 @@ from app.services.scheduler import sync_datasource_job
logger = get_logger(__name__)
JOB_TYPE_COLLECT = "collect"
JOB_TYPE_CLEAR_DATA = "clear_data"
JOB_TYPE_CLEAR_CACHE = "clear_cache"
JOB_TYPE_EARTH_REFRESH = "earth_refresh"
JOB_TYPE_COLLECT = JobType.COLLECT.value
JOB_TYPE_CLEAR_DATA = JobType.CLEAR_DATA.value
JOB_TYPE_CLEAR_CACHE = JobType.CLEAR_CACHE.value
JOB_TYPE_EARTH_REFRESH = JobType.EARTH_REFRESH.value
JOB_STATUS_QUEUED = "queued"
JOB_STATUS_RUNNING = "running"
JOB_STATUS_CANCELLING = "cancelling"
JOB_STATUS_SUCCESS = "success"
JOB_STATUS_FAILED = "failed"
JOB_STATUS_CANCELLED = "cancelled"
JOB_STATUS_QUEUED = JobStatus.QUEUED.value
JOB_STATUS_RUNNING = JobStatus.RUNNING.value
JOB_STATUS_CANCELLING = JobStatus.CANCELLING.value
JOB_STATUS_SUCCESS = JobStatus.SUCCESS.value
JOB_STATUS_FAILED = JobStatus.FAILED.value
JOB_STATUS_CANCELLED = JobStatus.CANCELLED.value
ACTIVE_JOB_STATUSES = (JOB_STATUS_QUEUED, JOB_STATUS_RUNNING, JOB_STATUS_CANCELLING)
TERMINAL_JOB_STATUSES = (JOB_STATUS_SUCCESS, JOB_STATUS_FAILED, JOB_STATUS_CANCELLED)
@@ -54,6 +55,9 @@ DATA_WRITE_JOB_TYPES = (JOB_TYPE_COLLECT, JOB_TYPE_CLEAR_DATA, JOB_TYPE_CLEAR_CA
SOURCE_LOCK_JOB_STATUSES = (JOB_STATUS_RUNNING, JOB_STATUS_CANCELLING)
QUEUE_POLL_SECONDS = 0.35
JOB_STALE_LOCK_MINUTES = 90
ORPHAN_CANCELLING_GRACE_SECONDS = 30
JOB_RECOVERY_SWEEP_SECONDS = 15
DATA_DELETE_BATCH_SIZE = 50_000
DEFAULT_WORKER_CONCURRENCY = 2
RUNNING_DATA_JOB_TASKS: dict[int, asyncio.Task[Any]] = {}
@@ -77,7 +81,7 @@ async def enqueue_datasource_job(
task_type: str,
*,
payload: dict[str, Any] | None = None,
rollback_policy: str = "keep_committed_batches",
rollback_policy: str = RollbackPolicy.KEEP_COMMITTED_BATCHES.value,
dedupe_key: str | None = None,
) -> CollectionTask:
if dedupe_key:
@@ -281,6 +285,7 @@ class DataJobWorker:
self._task: asyncio.Task[None] | None = None
self._stop_event: asyncio.Event | None = None
self._running: set[asyncio.Task[Any]] = set()
self._last_recovery_sweep_at: datetime | None = None
def start(self) -> None:
if self._task and not self._task.done():
@@ -303,6 +308,11 @@ class DataJobWorker:
await self._recover_stale_running_jobs()
while not self._stop_event.is_set():
self._running = {task for task in self._running if not task.done()}
if (
self._last_recovery_sweep_at is None
or (_utcnow() - self._last_recovery_sweep_at).total_seconds() >= JOB_RECOVERY_SWEEP_SECONDS
):
await self._recover_stale_running_jobs()
if len(self._running) >= self.concurrency:
await asyncio.sleep(QUEUE_POLL_SECONDS)
continue
@@ -316,7 +326,9 @@ class DataJobWorker:
self._running.add(runner)
async def _recover_stale_running_jobs(self) -> None:
self._last_recovery_sweep_at = _utcnow()
cutoff = _utcnow() - timedelta(minutes=JOB_STALE_LOCK_MINUTES)
orphan_cancelling_cutoff = _utcnow() - timedelta(seconds=ORPHAN_CANCELLING_GRACE_SECONDS)
async with async_session_factory() as db:
result = await db.execute(
select(CollectionTask)
@@ -332,6 +344,21 @@ class DataJobWorker:
job.error_message = "Marked failed after stale data job lock timeout"
if stale_jobs:
await db.commit()
orphan_result = await db.execute(
select(CollectionTask)
.where(CollectionTask.status == JOB_STATUS_CANCELLING)
.where(CollectionTask.locked_at.is_(None))
.where(CollectionTask.requested_cancel_at.is_not(None))
.where(CollectionTask.requested_cancel_at < orphan_cancelling_cutoff)
)
for job in orphan_result.scalars().all():
if job.id in RUNNING_DATA_JOB_TASKS:
continue
await _cancel_task_without_runner(
db,
job,
reason=job.cancel_reason or "cancelled_after_orphaned_runner",
)
async def _claim_next_job(self) -> int | None:
async with async_session_factory() as db:
@@ -477,15 +504,22 @@ async def _run_clear_data_job(db: AsyncSession, task: CollectionTask) -> None:
await db.commit()
await _broadcast_task_update(task)
count_result = await db.execute(
select(CollectedData.id).where(CollectedData.source == source)
deleted_count = await _delete_table_rows_by_source(
db,
task,
table_name="collected_data",
source_column="source",
source=source,
)
derived_deleted_counts = await _clear_derived_datasource_data_in_batches(
db,
task,
source,
progress_offset=deleted_count,
)
collected_ids = [row[0] for row in count_result.all()]
derived_deleted_counts = await clear_derived_datasource_data(db, source)
if collected_ids:
await db.execute(CollectedData.__table__.delete().where(CollectedData.id.in_(collected_ids)))
deleted_count = len(collected_ids)
derived_deleted_count = sum(derived_deleted_counts.values())
if any(key.startswith("ais_") for key in derived_deleted_counts):
await db.execute(text("ANALYZE ais_raw_observations"))
task.records_processed = deleted_count + derived_deleted_count
task.total_records = task.records_processed
@@ -504,10 +538,99 @@ async def _run_clear_data_job(db: AsyncSession, task: CollectionTask) -> None:
task.phase = "completed"
task.phase_message = "数据库数据已清理"
task.completed_at = _utcnow()
datasource = await db.get(DataSource, task.datasource_id)
if datasource is not None:
datasource.last_status = JOB_STATUS_SUCCESS
datasource.last_run_at = task.completed_at
await db.execute(
DataSnapshot.__table__.update()
.where(DataSnapshot.source == source)
.values(is_current=False)
)
await db.commit()
await _broadcast_task_update(task)
async def _delete_table_rows_by_source(
db: AsyncSession,
task: CollectionTask,
*,
table_name: str,
source_column: str,
source: str,
progress_offset: int = 0,
) -> int:
deleted = 0
while True:
result = await db.execute(
text(
f"""
WITH doomed AS (
SELECT ctid
FROM {table_name}
WHERE {source_column} = :source
LIMIT :batch_size
),
deleted_rows AS (
DELETE FROM {table_name}
USING doomed
WHERE {table_name}.ctid = doomed.ctid
RETURNING 1
)
SELECT COUNT(*) FROM deleted_rows
"""
),
{"source": source, "batch_size": DATA_DELETE_BATCH_SIZE},
)
batch_deleted = max(int(result.scalar_one() or 0), 0)
if batch_deleted <= 0:
break
deleted += batch_deleted
task.records_processed = progress_offset + deleted
task.phase_current = task.records_processed
task.phase_unit = "records"
task.phase_message = f"正在删除数据:{task.records_processed}"
await db.commit()
await _broadcast_task_update(task)
return deleted
async def _clear_derived_datasource_data_in_batches(
db: AsyncSession,
task: CollectionTask,
source: str,
progress_offset: int = 0,
) -> dict[str, int]:
deleted_counts: dict[str, int] = {}
if source in {"barentswatch_vessels", "aisstream_vessels"}:
deleted_counts["ais_conflict_records"] = await _delete_table_rows_by_source(
db,
task,
table_name="ais_conflict_records",
source_column="selected_source",
source=source,
progress_offset=progress_offset + sum(deleted_counts.values()),
)
deleted_counts["ais_source_health"] = await _delete_table_rows_by_source(
db,
task,
table_name="ais_source_health",
source_column="source",
source=source,
progress_offset=progress_offset + sum(deleted_counts.values()),
)
deleted_counts["ais_raw_observations"] = await _delete_table_rows_by_source(
db,
task,
table_name="ais_raw_observations",
source_column="source",
source=source,
progress_offset=progress_offset + sum(deleted_counts.values()),
)
return deleted_counts
return await clear_derived_datasource_data(db, source)
async def _run_clear_cache_job(db: AsyncSession, task: CollectionTask) -> None:
source = str(task.source or (task.payload or {}).get("source") or "").strip()
if not source:

View File

@@ -13,6 +13,7 @@ from sqlalchemy import func, select
from app.core.data_sources import get_data_sources_config
from app.core.datasource_defaults import DEFAULT_DATASOURCES
from app.core.enums import JobStatus
from app.models.collected_data import CollectedData
from app.models.datasource import DataSource
from app.models.datasource_config import DataSourceConfig
@@ -398,7 +399,7 @@ async def has_collected_data(db, source: str) -> bool:
datasource_result = await db.execute(select(DataSource).where(DataSource.source == source))
datasource = datasource_result.scalar_one_or_none()
return bool(datasource and datasource.last_status == "success")
return bool(datasource and datasource.last_status == JobStatus.SUCCESS.value)
async def get_builtin_connection_status(

View File

@@ -6,6 +6,7 @@ from dataclasses import dataclass
from pathlib import Path
from typing import Literal
from app.core.enums import UserRole
from app.models.user import User
DocsAccess = Literal["public", "docs_user", "docs_developer", "docs_admin"]
@@ -43,7 +44,9 @@ DOCS_METADATA: tuple[DocsMetadata, ...] = (
DocsMetadata("earth-satellite-footprint-policy.md", "earth-satellite-footprint-policy", "docs_developer", "Earth", 13, "智能星球卫星覆盖策略", "Intelligent Planet Satellite Footprint Policy"),
DocsMetadata("earth-bgp-context.md", "earth-bgp-context", "docs_developer", "Earth", 14, "BGP 态势上下文", "BGP Context"),
DocsMetadata("earth-interactable-usage.md", "earth-interactable-usage", "docs_developer", "Earth", 16, "智能星球可交互图标接入", "Intelligent Planet Interactable Usage"),
DocsMetadata("earth-toolbar-overlay-coordination.md", "earth-toolbar-overlay-coordination", "docs_developer", "Earth", 17, "智能星球工具栏与浮层协同", "Intelligent Planet Toolbar and Overlay Coordination"),
DocsMetadata("earth-interactable-clustering.md", "earth-interactable-clustering", "docs_developer", "Earth", 17, "智能星球可交互图标聚类策略", "Intelligent Planet Interactable Clustering"),
DocsMetadata("earth-toolbar-overlay-coordination.md", "earth-toolbar-overlay-coordination", "docs_developer", "Earth", 18, "智能星球工具栏与浮层协同", "Intelligent Planet Toolbar and Overlay Coordination"),
DocsMetadata("earth-news-sources.md", "earth-news-sources", "docs_developer", "Earth", 19, "智能星球新闻源配置", "Intelligent Planet News Source Configuration"),
DocsMetadata("frontend-admin-frontend-context.md", "frontend-admin-frontend-context", "docs_developer", "Frontend", 20, "控制台前端结构", "Admin Frontend Context"),
DocsMetadata("frontend-layout-guidelines.md", "frontend-layout-guidelines", "docs_developer", "Frontend", 21, "前端布局指南", "Frontend Layout Guidelines"),
DocsMetadata("tactile-ui-components.md", "tactile-ui-components", "docs_developer", "Frontend", 24, "Tactile UI 组件库", "Tactile UI Components"),
@@ -52,6 +55,7 @@ DOCS_METADATA: tuple[DocsMetadata, ...] = (
DocsMetadata("datasource-collector-settings-connectivity.md", "datasource-collector-settings-connectivity", "docs_developer", "Backend", 32, "数据源、采集器设置与连接验证", "Datasource Collector Settings and Connectivity"),
DocsMetadata("backend-datasources-api-performance.md", "backend-datasources-api-performance", "docs_developer", "Backend", 33, "数据源 API 性能", "Datasource API Performance"),
DocsMetadata("data-job-earth-sync-architecture.md", "data-job-earth-sync-architecture", "docs_developer", "Backend", 34, "数据作业与 Outbox 技术架构", "Data Jobs and Outbox Architecture"),
DocsMetadata("backend-enum-contracts.md", "backend-enum-contracts", "docs_developer", "Backend", 35, "后端枚举与字符串兼容契约", "Backend Enum and String Compatibility Contract"),
DocsMetadata("location-pipeline-development.md", "location-pipeline-development", "docs_developer", "Backend", 35, "通用位置估算管线开发说明", "Shared Location Resolution Pipeline Development Guide"),
DocsMetadata("earth-news-live-streams-collector-format.md", "earth-news-live-streams-collector-format", "docs_developer", "Backend", 36, "新闻直播采集格式", "News Live Streams Collector Format"),
DocsMetadata("docs-gatekeeper-development.md", "docs-gatekeeper-development", "docs_developer", "Backend", 37, "Docs Gatekeeper 开发说明", "Docs Gatekeeper Development Guide"),
@@ -69,9 +73,9 @@ def get_user_gatekeeper_groups(user: User | None) -> set[str]:
return set()
role = user.role.value if hasattr(user.role, "value") else str(user.role or "")
if role == "super_admin":
if role == UserRole.SUPER_ADMIN.value:
return {"docs_user", "docs_developer", "docs_admin"}
if role == "admin":
if role == UserRole.ADMIN.value:
return {"docs_user", "docs_developer", "docs_admin"}
groups = set()

View File

@@ -62,11 +62,11 @@ def interactables_to_geojson(items: list[EarthInteractable]) -> dict[str, Any]:
def invalidate_interactable_cache(layer: str | None = None) -> int:
layer_key = str(layer or "*").strip() or "*"
deleted = earth_layer_cache.delete_pattern(
f"{EARTH_LAYER_CACHE_PREFIX}:interactables:layer:{layer_key}*"
f"{EARTH_LAYER_CACHE_PREFIX}:interactables:interactable_layer:{layer_key}*"
)
if layer_key != "all":
deleted += earth_layer_cache.delete_pattern(
f"{EARTH_LAYER_CACHE_PREFIX}:interactables:layer:all*"
f"{EARTH_LAYER_CACHE_PREFIX}:interactables:interactable_layer:all*"
)
return deleted

View File

@@ -20,10 +20,11 @@ class EarthLayerAdapter:
EARTH_LAYER_ADAPTERS: tuple[EarthLayerAdapter, ...] = (
EarthLayerAdapter(
sources=frozenset({"barentswatch_vessels", "aisstream_vessels", "vessel_static", "vessel_position", "ais_raw_observations", "ais_source_health"}),
tables=frozenset({"vessel_static", "vessel_position", "ais_raw_observations", "ais_source_health"}),
sources=frozenset({"barentswatch_vessels", "aisstream_vessels", "vessel_static", "vessel_position", "vessel_current_state", "ais_raw_observations", "ais_source_health"}),
tables=frozenset({"vessel_static", "vessel_position", "vessel_current_state", "ais_raw_observations", "ais_source_health"}),
layers=("vessels",),
cache_patterns=("vessels*", "summary*"),
derived_models=("ais_raw_observations", "ais_conflict_records", "ais_source_health"),
),
EarthLayerAdapter(
sources=frozenset(
@@ -163,17 +164,24 @@ async def clear_derived_datasource_data(db: AsyncSession, source: str) -> dict[s
from app.models.bgp_anomaly import BGPAnomaly
from app.models.bgp_incident import BGPIncident
from app.models.bgp_observation import BGPObservation
from app.models.vessel import AISConflictRecord, AISRawObservation, AISSourceHealth
model_by_key: dict[str, Any] = {
"bgp_observations": BGPObservation,
"bgp_anomalies": BGPAnomaly,
"bgp_incidents": BGPIncident,
"ais_raw_observations": AISRawObservation,
"ais_conflict_records": AISConflictRecord,
"ais_source_health": AISSourceHealth,
}
deleted_counts: dict[str, int] = {}
for key in adapter.derived_models:
model = model_by_key.get(key)
if model is None:
continue
result = await db.execute(model.__table__.delete().where(model.source == source))
if key == "ais_conflict_records":
result = await db.execute(model.__table__.delete().where(model.selected_source == source))
else:
result = await db.execute(model.__table__.delete().where(model.source == source))
deleted_counts[key] = int(result.rowcount or 0)
return deleted_counts

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,258 @@
"""Classification, importance, and breaking-news policy for Earth news."""
from __future__ import annotations
from dataclasses import dataclass
from datetime import UTC, datetime, timedelta
import re
from typing import Any, Protocol
from app.core.enums import (
BreakingLevel,
BreakingScope,
BreakingSource,
NewsImportanceLevel,
NewsMarketImpact,
NewsTaggingSource,
parse_enum,
)
class NewsItemLike(Protocol):
title: str
summary: str
source: str
feed_name: str
published_at: datetime | None
feed_default_category: str
category: str
item_tags: list[str]
tagging_source: str
tagging_confidence: float
importance_score: int
importance_level: str
importance_reasons: list[str]
market_impact: str
source_tags: list[str]
breaking_level: str
breaking_scope: str
breaking_reasons: list[str]
breaking_source: str
breaking_confidence: float
breaking_expires_at: datetime | None
class NewsSourceLike(Protocol):
default_category: str
importance_weight: int
source_tags: tuple[str, ...]
class NewsFeedLike(Protocol):
default_category: str
@dataclass(frozen=True)
class BreakingRule:
level: BreakingLevel
scope: BreakingScope
reason: str
keywords: tuple[str, ...]
IMPORTANCE_THRESHOLDS: tuple[tuple[int, NewsImportanceLevel], ...] = (
(80, NewsImportanceLevel.CRITICAL),
(60, NewsImportanceLevel.HIGH),
(35, NewsImportanceLevel.MEDIUM),
(0, NewsImportanceLevel.LOW),
)
BREAKING_LEVEL_RANK: dict[BreakingLevel, int] = {
BreakingLevel.NONE: 0,
BreakingLevel.WATCH: 1,
BreakingLevel.BREAKING: 2,
BreakingLevel.CRITICAL: 3,
}
BREAKING_TTL: dict[BreakingLevel, timedelta] = {
BreakingLevel.WATCH: timedelta(hours=6),
BreakingLevel.BREAKING: timedelta(hours=12),
BreakingLevel.CRITICAL: timedelta(hours=24),
}
BREAKING_RULES: tuple[BreakingRule, ...] = (
BreakingRule(BreakingLevel.CRITICAL, BreakingScope.GLOBAL, "核事故或核风险", ("nuclear accident", "nuclear emergency", "radiation leak", "核事故", "核泄漏", "辐射泄漏")),
BreakingRule(BreakingLevel.CRITICAL, BreakingScope.REGIONAL, "重大军事冲突升级", ("airstrike", "missile strike", "invasion", "martial law", "空袭", "导弹袭击", "入侵", "戒严")),
BreakingRule(BreakingLevel.BREAKING, BreakingScope.REGIONAL, "战争或安全事件", ("war escalates", "terror attack", "coup", "hostage", "战争升级", "恐袭", "政变", "人质")),
BreakingRule(BreakingLevel.BREAKING, BreakingScope.REGIONAL, "重大灾害应急", ("major earthquake", "tsunami", "volcanic eruption", "state of emergency", "强震", "海啸", "火山喷发", "紧急状态")),
BreakingRule(BreakingLevel.BREAKING, BreakingScope.GLOBAL, "金融市场异常", ("market halt", "trading halt", "flash crash", "bank run", "金融熔断", "交易暂停", "银行挤兑")),
BreakingRule(BreakingLevel.WATCH, BreakingScope.GLOBAL, "大规模网络安全事件", ("massive cyberattack", "ransomware attack", "data breach", "大规模网络攻击", "勒索软件", "数据泄露")),
BreakingRule(BreakingLevel.WATCH, BreakingScope.REGIONAL, "航天或卫星事故", ("rocket explosion", "satellite collision", "space station emergency", "火箭爆炸", "卫星碰撞", "空间站事故")),
)
def contains_keyword(text: str, keyword: str) -> bool:
keyword_text = str(keyword or "").strip().lower()
if not keyword_text:
return False
if re.search(r"[\u4e00-\u9fff]", keyword_text):
return keyword_text in text
return re.search(rf"(?<![a-z0-9]){re.escape(keyword_text)}(?![a-z0-9])", text) is not None
def score_category(text: str, title_text: str, category: dict[str, Any]) -> int:
score = 0
keywords = category.get("keywords") if isinstance(category.get("keywords"), list) else []
for keyword in keywords:
if contains_keyword(title_text, keyword):
score += 3
elif contains_keyword(text, keyword):
score += 1
return score
def importance_level(score: int) -> NewsImportanceLevel:
normalized_score = max(0, min(100, int(score)))
for threshold, level in IMPORTANCE_THRESHOLDS:
if normalized_score >= threshold:
return level
return NewsImportanceLevel.LOW
def normalize_breaking_level(value: object) -> BreakingLevel:
return parse_enum(BreakingLevel, value, BreakingLevel.NONE)
def normalize_breaking_scope(value: object) -> BreakingScope:
return parse_enum(BreakingScope, value, BreakingScope.REGIONAL)
def breaking_expires_at(level: object, published_at: datetime | None) -> datetime | None:
normalized = normalize_breaking_level(level)
if normalized is BreakingLevel.NONE:
return None
base = published_at or datetime.now(UTC)
base = base.replace(tzinfo=UTC) if base.tzinfo is None else base.astimezone(UTC)
return base + BREAKING_TTL[normalized]
def is_breaking_active(item: NewsItemLike, *, now: datetime | None = None) -> bool:
if normalize_breaking_level(item.breaking_level) is BreakingLevel.NONE:
return False
expires_at = item.breaking_expires_at
if expires_at is None:
return True
expires_at = expires_at.replace(tzinfo=UTC) if expires_at.tzinfo is None else expires_at.astimezone(UTC)
return expires_at > (now or datetime.now(UTC))
def breaking_sort_rank(item: NewsItemLike) -> int:
if not is_breaking_active(item):
return 0
return BREAKING_LEVEL_RANK[normalize_breaking_level(item.breaking_level)]
def highest_breaking_level(items: list[NewsItemLike]) -> BreakingLevel:
active = [normalize_breaking_level(item.breaking_level) for item in items if is_breaking_active(item)]
return max(active, key=BREAKING_LEVEL_RANK.get) if active else BreakingLevel.NONE
def apply_breaking_rules(item: NewsItemLike) -> None:
combined_text = f"{item.title} {item.summary} {item.source} {item.feed_name}".lower()
best_level = BreakingLevel.NONE
best_scope = BreakingScope.REGIONAL
reasons: list[str] = []
confidence = 0.0
for rule in BREAKING_RULES:
if not any(contains_keyword(combined_text, keyword) for keyword in rule.keywords):
continue
if BREAKING_LEVEL_RANK[rule.level] > BREAKING_LEVEL_RANK[best_level]:
best_level = rule.level
best_scope = rule.scope
if rule.reason not in reasons:
reasons.append(rule.reason)
confidence = max(confidence, 0.72 if rule.level is BreakingLevel.CRITICAL else 0.64 if rule.level is BreakingLevel.BREAKING else 0.52)
item.breaking_level = best_level.value
item.breaking_scope = (best_scope if best_level is not BreakingLevel.NONE else BreakingScope.REGIONAL).value
item.breaking_reasons = reasons
item.breaking_source = BreakingSource.RULES.value
item.breaking_confidence = round(confidence, 2)
item.breaking_expires_at = breaking_expires_at(best_level, item.published_at)
def apply_news_classification(
item: NewsItemLike,
source: NewsSourceLike,
*,
feed: NewsFeedLike | None,
config: dict[str, Any],
) -> NewsItemLike:
title_text = item.title.lower()
combined_text = f"{item.title} {item.summary} {item.source} {item.feed_name}".lower()
feed_default_category = (feed.default_category if feed else item.feed_default_category) or source.default_category or "other"
best_key = feed_default_category
best_score = second_score = 0
for category in config["categories"]:
if not isinstance(category, dict) or category.get("enabled") is False:
continue
score = score_category(combined_text, title_text, category)
if score > best_score:
second_score, best_score = best_score, score
best_key = str(category.get("key") or "other")
elif score > second_score:
second_score = score
item_tags: list[str] = []
for rule in config["item_tag_rules"]:
if not isinstance(rule, dict):
continue
keywords = rule.get("keywords") if isinstance(rule.get("keywords"), list) else []
if any(contains_keyword(combined_text, keyword) for keyword in keywords):
tag_key = str(rule.get("key") or "").strip()
if tag_key and tag_key not in item_tags:
item_tags.append(tag_key)
if best_score < 3 and rule.get("category"):
best_key, best_score = str(rule["category"]), 3
confidence = round(best_score / (best_score + second_score + 1), 2) if best_score else 0.35
if best_score < 3 and feed_default_category:
best_key, confidence = feed_default_category, 0.45
score = max(0, min(100, 18 + source.importance_weight + best_score * 6))
reasons: list[str] = []
source_tags = set(source.source_tags)
if "official_data" in source_tags:
score += 20
reasons.append("官方数据源")
if "press_release" in source_tags:
score = max(0, score - 12)
reasons.append("企业公告基础权重较低")
if any(contains_keyword(combined_text, term) for term in ("网上零售额", "电商物流指数", "gmv", "订单量", "物流指数", "履约", "直播电商", "跨境电商")):
score += 25
reasons.append("命中电商数据指标")
if any(contains_keyword(combined_text, term) for term in ("amazon", "shopify", "walmart", "alibaba", "jd.com", "pinduoduo", "tiktok shop", "shein", "阿里", "京东", "拼多多", "抖音")):
score += 15
reasons.append("涉及大型平台")
if any(term in combined_text for term in ("同比", "环比", "%", "billion", "million", "增长", "下降")):
score += 10
reasons.append("包含量化指标")
score = max(0, min(100, score))
item.category = best_key or "other"
item.item_tags = item_tags
item.tagging_source = NewsTaggingSource.RULES.value
item.tagging_confidence = confidence
item.importance_score = score
item.importance_level = importance_level(score).value
item.importance_reasons = reasons or ["按来源权重和分类规则计算"]
item.market_impact = (
NewsMarketImpact.GLOBAL.value
if "global" in source_tags
else NewsMarketImpact.NATIONAL.value
if {"china", "us"} & source_tags
else NewsMarketImpact.SECTOR.value
)
item.source_tags = list(source.source_tags)
apply_breaking_rules(item)
return item

View File

@@ -0,0 +1,693 @@
from __future__ import annotations
from dataclasses import dataclass
from datetime import UTC, datetime
import hashlib
import html
import json
import re
from typing import Any
from bs4 import BeautifulSoup
from sqlalchemy import delete, func, select
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.enums import NewsEnrichmentStatus, NewsSourceType, NewsTaggingSource
from app.core.websocket.broadcaster import broadcaster
from app.models.earth_news import EarthNewsItem
from app.models.system_setting import SystemSetting
from app.services.earth_news import (
ALLOWED_NEWS_CATEGORY_KEYS,
DEFAULT_NEWS_LOCALE,
REGION_ANCHORS,
NewsFeedEndpoint,
NewsFeedSource,
NewsTargetLocation,
ParsedNewsItem,
apply_news_classification,
build_anchor_location_patch,
build_target_location_job_payload,
build_target_location_patch,
_serialize_item,
)
from app.services.earth_news_queue import enqueue_target_location_job
from app.services.earth_news_store import record_to_parsed_news_item
MANUAL_NEWS_SOURCE_ID = "manual"
MANUAL_NEWS_SOURCE_LABEL = "手动添加"
MANUAL_NEWS_MAX_IMPORT_ITEMS = 500
MANUAL_NEWS_MAX_TITLE_LENGTH = 500
MANUAL_NEWS_MAX_SUMMARY_LENGTH = 1200
MANUAL_NEWS_MAX_CONTENT_LENGTH = 12000
EARTH_NEWS_MANUAL_GROUPS_CATEGORY = "earth_news_manual_groups"
DEFAULT_MANUAL_NEWS_GROUP_ID = "manual-default"
DEFAULT_MANUAL_NEWS_GROUP_NAME = "新建新闻组"
@dataclass(frozen=True)
class ManualNewsWriteResult:
item: EarthNewsItem
created: bool
queued: bool
@dataclass(frozen=True)
class ManualNewsGroup:
id: str
name: str
sort_order: int = 0
def _clean_text(value: object, *, max_length: int) -> str:
raw = "" if value is None else str(value)
text = BeautifulSoup(html.unescape(raw), "html.parser").get_text(" ", strip=True)
text = re.sub(r"\s+", " ", text).strip()
if len(text) > max_length:
return text[: max_length - 1].rstrip() + ""
return text
def _parse_datetime(value: object) -> datetime | None:
if value is None or str(value).strip() == "":
return None
if isinstance(value, datetime):
parsed = value
else:
try:
parsed = datetime.fromisoformat(str(value).strip().replace("Z", "+00:00"))
except ValueError as exc:
raise ValueError("published_at 必须是 ISO8601 时间。") from exc
if parsed.tzinfo is None:
return parsed.replace(tzinfo=UTC)
return parsed.astimezone(UTC)
def _detect_language(*parts: str) -> str:
text = " ".join(part for part in parts if part)
cjk_count = len(re.findall(r"[\u4e00-\u9fff]", text))
latin_count = len(re.findall(r"[A-Za-z]", text))
return "zh-CN" if cjk_count >= max(4, latin_count // 3) else "en-US"
def _manual_item_id(*, title: str, published_at: datetime | None, url: str, source: str) -> str:
published = published_at.isoformat() if published_at else ""
basis = "\n".join([title.strip().lower(), published, url.strip().lower(), source.strip().lower()])
return f"manual:{hashlib.sha1(basis.encode('utf-8')).hexdigest()[:16]}"
def _manual_group_id(name: str) -> str:
basis = f"{name.strip().lower()}\n{datetime.now(UTC).isoformat()}"
return f"manual-group:{hashlib.sha1(basis.encode('utf-8')).hexdigest()[:10]}"
def _news_meta(record: EarthNewsItem) -> dict[str, Any]:
location_meta = record.location_meta if isinstance(record.location_meta, dict) else {}
news_meta = location_meta.get("news_meta")
return dict(news_meta) if isinstance(news_meta, dict) else {}
def _record_source_type(record: EarthNewsItem) -> str:
return str(_news_meta(record).get("feed_type") or _news_meta(record).get("source_type") or "rss")
def _record_manual_group_id(record: EarthNewsItem) -> str:
return str(_news_meta(record).get("manual_group_id") or DEFAULT_MANUAL_NEWS_GROUP_ID)
def _rss_group_id(record: EarthNewsItem) -> str:
basis = "\n".join(
[
_record_source_type(record),
str(record.feed_name or ""),
str(record.source or ""),
]
)
return f"rss:{hashlib.sha1(basis.encode('utf-8')).hexdigest()[:12]}"
def _default_manual_group() -> dict[str, Any]:
return {
"id": DEFAULT_MANUAL_NEWS_GROUP_ID,
"name": DEFAULT_MANUAL_NEWS_GROUP_NAME,
"sort_order": 0,
}
def _normalize_manual_groups_payload(payload: Any) -> list[dict[str, Any]]:
raw_groups = payload.get("groups") if isinstance(payload, dict) else None
normalized: list[dict[str, Any]] = []
seen: set[str] = set()
for index, item in enumerate(raw_groups if isinstance(raw_groups, list) else []):
if not isinstance(item, dict):
continue
group_id = str(item.get("id") or "").strip()
name = _clean_text(item.get("name"), max_length=120)
if not group_id or not name or group_id in seen:
continue
normalized.append(
{
"id": group_id,
"name": name,
"sort_order": int(item.get("sort_order") or index),
}
)
seen.add(group_id)
if DEFAULT_MANUAL_NEWS_GROUP_ID not in seen:
normalized.insert(0, _default_manual_group())
return sorted(normalized, key=lambda item: (int(item.get("sort_order") or 0), str(item.get("name") or "")))
async def _get_manual_groups_record(db: AsyncSession) -> SystemSetting | None:
result = await db.execute(
select(SystemSetting).where(SystemSetting.category == EARTH_NEWS_MANUAL_GROUPS_CATEGORY)
)
return result.scalar_one_or_none()
async def get_manual_news_groups(db: AsyncSession) -> list[dict[str, Any]]:
record = await _get_manual_groups_record(db)
return _normalize_manual_groups_payload(record.payload if record else None)
async def _save_manual_news_groups(db: AsyncSession, groups: list[dict[str, Any]]) -> list[dict[str, Any]]:
normalized = _normalize_manual_groups_payload({"groups": groups})
record = await _get_manual_groups_record(db)
payload = {"groups": normalized}
if record is None:
db.add(SystemSetting(category=EARTH_NEWS_MANUAL_GROUPS_CATEGORY, payload=payload))
else:
record.payload = payload
await db.flush()
return normalized
async def resolve_manual_news_group(db: AsyncSession, group_id: str | None) -> ManualNewsGroup:
normalized_id = str(group_id or DEFAULT_MANUAL_NEWS_GROUP_ID).strip() or DEFAULT_MANUAL_NEWS_GROUP_ID
groups = await get_manual_news_groups(db)
match = next((item for item in groups if item.get("id") == normalized_id), None)
if match is None and normalized_id != DEFAULT_MANUAL_NEWS_GROUP_ID:
raise ValueError(f"手动新闻组不存在:{normalized_id}")
match = match or _default_manual_group()
return ManualNewsGroup(
id=str(match["id"]),
name=str(match["name"]),
sort_order=int(match.get("sort_order") or 0),
)
async def create_manual_news_group(db: AsyncSession, name: str) -> dict[str, Any]:
group_name = _clean_text(name, max_length=120)
if not group_name:
raise ValueError("新闻组名称不能为空。")
groups = await get_manual_news_groups(db)
group = {"id": _manual_group_id(group_name), "name": group_name, "sort_order": len(groups)}
groups.append(group)
await _save_manual_news_groups(db, groups)
return group
async def rename_manual_news_group(db: AsyncSession, group_id: str, name: str) -> dict[str, Any]:
group_name = _clean_text(name, max_length=120)
if not group_name:
raise ValueError("新闻组名称不能为空。")
groups = await get_manual_news_groups(db)
match = next((item for item in groups if item.get("id") == group_id), None)
if match is None:
raise ValueError(f"手动新闻组不存在:{group_id}")
match["name"] = group_name
await _save_manual_news_groups(db, groups)
result = await db.execute(
select(EarthNewsItem).where(
EarthNewsItem.location_meta.op("->")("news_meta").op("->>")("feed_type") == NewsSourceType.MANUAL.value
)
)
for record in result.scalars().all():
if _record_manual_group_id(record) != group_id:
continue
location_meta = dict(record.location_meta or {})
news_meta = dict(location_meta.get("news_meta") or {})
news_meta["manual_group_name"] = group_name
location_meta["news_meta"] = news_meta
record.location_meta = location_meta
await db.flush()
return match
def _normalize_region(value: object) -> str:
region = str(value or "global").strip().lower() or "global"
if region not in REGION_ANCHORS:
raise ValueError(f"region 不支持:{region}")
return region
def _normalize_tags(value: object) -> list[str]:
if value is None:
return []
if isinstance(value, str):
parts = re.split(r"[,\n]", value)
elif isinstance(value, list):
parts = [str(item) for item in value]
else:
raise ValueError("tags 必须是字符串数组或逗号分隔字符串。")
return [item.strip() for item in parts if item.strip()][:20]
def _normalize_location(value: object) -> NewsTargetLocation | None:
if value in (None, ""):
return None
if not isinstance(value, dict):
raise ValueError("location 必须是对象。")
lat = value.get("latitude")
lon = value.get("longitude")
if lat in (None, "") and lon in (None, ""):
return None
try:
latitude = float(lat)
longitude = float(lon)
except (TypeError, ValueError) as exc:
raise ValueError("location.latitude / longitude 必须是数字。") from exc
if not -90 <= latitude <= 90 or not -180 <= longitude <= 180:
raise ValueError("location 经纬度超出范围。")
label = _clean_text(value.get("label"), max_length=255)
if not label:
label = f"{latitude:.4f}, {longitude:.4f}"
return NewsTargetLocation(
latitude=latitude,
longitude=longitude,
label=label,
source="manual_location",
confidence=1.0,
country=_clean_text(value.get("country"), max_length=100) or None,
city=_clean_text(value.get("city"), max_length=100) or None,
)
def _manual_source(source_name: str, *, region: str) -> NewsFeedSource:
return NewsFeedSource(
id=MANUAL_NEWS_SOURCE_ID,
name=source_name or MANUAL_NEWS_SOURCE_LABEL,
region=region,
feed_url="",
homepage_url="",
source_type=NewsSourceType.MANUAL.value,
default_category="other",
source_tags=("manual",),
)
def _manual_feed(category: str) -> NewsFeedEndpoint:
return NewsFeedEndpoint(
id=MANUAL_NEWS_SOURCE_ID,
name=MANUAL_NEWS_SOURCE_LABEL,
url="",
type=NewsSourceType.MANUAL.value,
default_category=category or "other",
tags=("manual",),
priority=1,
)
def parsed_manual_news_item(
payload: dict[str, Any],
*,
item_id_override: str | None = None,
) -> tuple[ParsedNewsItem, NewsTargetLocation | None, str]:
title = _clean_text(payload.get("title"), max_length=MANUAL_NEWS_MAX_TITLE_LENGTH)
if not title:
raise ValueError("title 不能为空。")
content = _clean_text(payload.get("content"), max_length=MANUAL_NEWS_MAX_CONTENT_LENGTH)
summary = _clean_text(payload.get("summary"), max_length=MANUAL_NEWS_MAX_SUMMARY_LENGTH)
if not summary:
summary = _clean_text(content, max_length=240) if content else title
source = _clean_text(payload.get("source"), max_length=255) or MANUAL_NEWS_SOURCE_LABEL
region = _normalize_region(payload.get("region"))
published_at = _parse_datetime(payload.get("published_at")) or datetime.now(UTC)
url = str(payload.get("url") or "").strip()
category = str(payload.get("category") or "other").strip().lower() or "other"
if category not in ALLOWED_NEWS_CATEGORY_KEYS:
raise ValueError(f"category 不支持:{category}")
tags = _normalize_tags(payload.get("tags"))
target = _normalize_location(payload.get("location"))
language = str(payload.get("content_language") or "").strip() or _detect_language(title, summary, content)
localizations = {
language: {
"title": title,
"summary": summary,
}
}
item = ParsedNewsItem(
id=item_id_override
or _manual_item_id(title=title, published_at=published_at, url=url, source=source),
title=title,
summary=summary,
url=url,
source=source,
feed_name=MANUAL_NEWS_SOURCE_LABEL,
feed_region=region,
homepage_url=str(payload.get("homepage_url") or ""),
published_at=published_at,
content_language=language,
localizations=localizations,
enrichment_status=NewsEnrichmentStatus.PENDING.value,
source_tags=["manual"],
feed_id=MANUAL_NEWS_SOURCE_ID,
feed_type=NewsSourceType.MANUAL.value,
feed_default_category=category,
category=category,
item_tags=tags,
tagging_source=NewsTaggingSource.MANUAL.value if payload.get("category") else NewsTaggingSource.RULES.value,
tagging_confidence=0.9 if payload.get("category") else 0.0,
)
source_config = _manual_source(source, region=region)
feed = _manual_feed(category)
apply_news_classification(item, source_config, feed=feed)
if payload.get("category"):
item.category = category
item.tagging_source = NewsTaggingSource.MANUAL.value
item.tagging_confidence = 0.9
if tags:
item.item_tags = sorted(set([*item.item_tags, *tags]))
return item, target, content
def _manual_editable(record: EarthNewsItem) -> bool:
if record.id.startswith("manual:"):
return True
news_meta = (record.location_meta or {}).get("news_meta") if isinstance(record.location_meta, dict) else None
return isinstance(news_meta, dict) and news_meta.get("feed_type") == NewsSourceType.MANUAL.value
async def _broadcast_news_reload() -> None:
await broadcaster.broadcast_earth_update(
{
"action": "database_changed",
"source": "earth_news_items",
"layers": ["news"],
"refresh_strategy": "reload",
}
)
async def upsert_manual_news_item(
db: AsyncSession,
payload: dict[str, Any],
*,
item_id_override: str | None = None,
group_id: str | None = None,
) -> ManualNewsWriteResult:
item, target, content = parsed_manual_news_item(payload, item_id_override=item_id_override)
group = await resolve_manual_news_group(db, group_id or payload.get("group_id"))
existing = await db.get(EarthNewsItem, item.id)
created = existing is None
patch = build_target_location_patch(item, target) if target else build_anchor_location_patch(item)
patch_meta = dict(patch.get("location_meta") or {})
patch_news_meta = dict(patch_meta.get("news_meta") or {})
patch_news_meta["feed_type"] = NewsSourceType.MANUAL.value
patch_news_meta["source_type"] = NewsSourceType.MANUAL.value
patch_news_meta["manual_group_id"] = group.id
patch_news_meta["manual_group_name"] = group.name
patch_meta["news_meta"] = patch_news_meta
patch["location_meta"] = patch_meta
now = datetime.now(UTC)
record = existing or EarthNewsItem(
id=item.id,
title=item.title,
summary=item.summary,
content_language=item.content_language,
localizations=dict(item.localizations or {}),
url=item.url,
source=item.source,
feed_name=item.feed_name,
region=item.feed_region,
homepage_url=item.homepage_url,
published_at=item.published_at,
latitude=patch["latitude"],
longitude=patch["longitude"],
location_label=patch["location_label"],
location_source=patch["location_source"],
verified=patch["verified"],
location_meta=patch["location_meta"],
first_seen_at=now,
last_seen_at=now,
resolved_at=now if patch["verified"] else None,
enrichment_status=item.enrichment_status,
)
if existing is None:
db.add(record)
else:
if not _manual_editable(record):
raise PermissionError("RSS 新闻不允许通过手动新闻接口编辑。")
record.title = item.title
record.summary = item.summary
record.content_language = item.content_language
record.localizations = dict(item.localizations or {})
record.url = item.url
record.source = item.source
record.feed_name = item.feed_name
record.region = item.feed_region
record.homepage_url = item.homepage_url
record.published_at = item.published_at
record.last_seen_at = now
if target is None and record.location_source == "manual_location":
merged_meta = dict(record.location_meta or {})
patch_meta = patch.get("location_meta") if isinstance(patch, dict) else None
patch_news_meta = patch_meta.get("news_meta") if isinstance(patch_meta, dict) else None
if isinstance(patch_news_meta, dict):
merged_meta["news_meta"] = patch_news_meta
record.location_meta = merged_meta
else:
record.location_meta = patch["location_meta"]
if target:
record.latitude = patch["latitude"]
record.longitude = patch["longitude"]
record.location_label = patch["location_label"]
record.location_source = patch["location_source"]
record.verified = patch["verified"]
record.resolved_at = now
elif record.location_source != "manual_location":
record.latitude = patch["latitude"]
record.longitude = patch["longitude"]
record.location_label = patch["location_label"]
record.location_source = patch["location_source"]
record.verified = patch["verified"]
record.resolved_at = None
record.enrichment_status = NewsEnrichmentStatus.PENDING.value
record.enrichment_error = None
record.enriched_at = None
if content:
meta = dict(record.location_meta or {})
meta["manual_content"] = content
record.location_meta = meta
await db.flush()
queued = await enqueue_target_location_job(build_target_location_job_payload(item), force=True)
if queued:
record.enrichment_status = NewsEnrichmentStatus.QUEUED.value
await db.flush()
return ManualNewsWriteResult(item=record, created=created, queued=queued)
async def import_manual_news_items(
db: AsyncSession,
payload: list[Any],
*,
group_id: str | None = None,
) -> dict[str, Any]:
if len(payload) > MANUAL_NEWS_MAX_IMPORT_ITEMS:
raise ValueError(f"单次最多导入 {MANUAL_NEWS_MAX_IMPORT_ITEMS} 条。")
created = 0
updated = 0
queued = 0
errors: list[dict[str, Any]] = []
for index, raw_item in enumerate(payload):
if not isinstance(raw_item, dict):
errors.append({"index": index, "error": "条目必须是 JSON 对象。"})
continue
try:
result = await upsert_manual_news_item(db, raw_item, group_id=group_id)
created += 1 if result.created else 0
updated += 0 if result.created else 1
queued += 1 if result.queued else 0
except Exception as exc:
errors.append({"index": index, "error": str(exc)})
if errors and created == 0 and updated == 0:
raise ValueError("导入失败,未写入任何新闻。")
return {"created": created, "updated": updated, "queued": queued, "failed": len(errors), "errors": errors}
async def parse_manual_news_import_upload(raw_bytes: bytes) -> list[Any]:
try:
payload = json.loads(raw_bytes.decode("utf-8-sig"))
except UnicodeDecodeError as exc:
raise ValueError("JSON 文件必须使用 UTF-8 编码。") from exc
except json.JSONDecodeError as exc:
raise ValueError(f"JSON 解析失败:第 {exc.lineno} 行第 {exc.colno} 列。") from exc
if not isinstance(payload, list):
raise ValueError("JSON 顶层必须是数组。")
return payload
def serialize_news_record(record: EarthNewsItem, *, locale: str = DEFAULT_NEWS_LOCALE) -> dict[str, Any]:
item = record_to_parsed_news_item(record)
payload = _serialize_item(item, active_region=item.feed_region, locale=locale)
news_meta = _news_meta(record)
payload["editable"] = _manual_editable(record)
payload["source_type"] = payload.get("feed_type")
payload["status"] = record.enrichment_status
payload["translated"] = bool((record.localizations or {}).get("zh-CN") and (record.localizations or {}).get("en-US"))
payload["manual_content"] = (record.location_meta or {}).get("manual_content") if isinstance(record.location_meta, dict) else None
payload["manual_group_id"] = news_meta.get("manual_group_id")
payload["manual_group_name"] = news_meta.get("manual_group_name")
return payload
def _record_matches_group(record: EarthNewsItem, group_id: str) -> bool:
source_type = _record_source_type(record)
if source_type == NewsSourceType.MANUAL.value:
return _record_manual_group_id(record) == group_id
return _rss_group_id(record) == group_id
async def list_news_records(
db: AsyncSession,
*,
page: int,
page_size: int,
source_type: str | None = None,
region: str | None = None,
category: str | None = None,
status_filter: str | None = None,
group_id: str | None = None,
) -> dict[str, Any]:
page = max(page, 1)
page_size = min(max(page_size, 1), 100)
query = select(EarthNewsItem)
count_query = select(func.count(EarthNewsItem.id))
filters = []
if region and region != "all":
filters.append(EarthNewsItem.region == region)
if status_filter and status_filter != "all":
filters.append(EarthNewsItem.enrichment_status == status_filter)
if source_type and source_type != "all":
filters.append(EarthNewsItem.location_meta.op("->")("news_meta").op("->>")("feed_type") == source_type)
if category and category != "all":
filters.append(EarthNewsItem.location_meta.op("->")("news_meta").op("->>")("category") == category)
for clause in filters:
query = query.where(clause)
count_query = count_query.where(clause)
ordered_query = query.order_by(EarthNewsItem.published_at.desc().nullslast(), EarthNewsItem.last_seen_at.desc())
if group_id:
result = await db.execute(ordered_query)
all_records = [record for record in result.scalars().all() if _record_matches_group(record, group_id)]
total = len(all_records)
records = all_records[(page - 1) * page_size : page * page_size]
else:
total_result = await db.execute(count_query)
result = await db.execute(
ordered_query.offset((page - 1) * page_size).limit(page_size)
)
records = list(result.scalars().all())
total = int(total_result.scalar() or 0)
return {
"items": [serialize_news_record(record) for record in records],
"page": page,
"page_size": page_size,
"total": total,
}
async def list_news_groups(db: AsyncSession, *, locale: str = DEFAULT_NEWS_LOCALE) -> dict[str, Any]:
manual_groups = await get_manual_news_groups(db)
manual_by_id: dict[str, dict[str, Any]] = {
str(group["id"]): {
"id": str(group["id"]),
"name": str(group["name"]),
"group_type": "manual",
"source_type": NewsSourceType.MANUAL.value,
"editable": True,
"sort_order": int(group.get("sort_order") or 0),
"count": 0,
"items": [],
}
for group in manual_groups
}
rss_by_id: dict[str, dict[str, Any]] = {}
result = await db.execute(
select(EarthNewsItem).order_by(EarthNewsItem.published_at.desc().nullslast(), EarthNewsItem.last_seen_at.desc())
)
for record in result.scalars().all():
serialized = serialize_news_record(record, locale=locale)
source_type = _record_source_type(record)
if source_type == NewsSourceType.MANUAL.value:
group_id = _record_manual_group_id(record)
group = manual_by_id.setdefault(
group_id,
{
"id": group_id,
"name": str(_news_meta(record).get("manual_group_name") or DEFAULT_MANUAL_NEWS_GROUP_NAME),
"group_type": "manual",
"source_type": NewsSourceType.MANUAL.value,
"editable": True,
"sort_order": len(manual_by_id),
"count": 0,
"items": [],
},
)
else:
group_id = _rss_group_id(record)
group = rss_by_id.setdefault(
group_id,
{
"id": group_id,
"name": record.feed_name or record.source or "RSS 新闻",
"group_type": "rss",
"source_type": source_type,
"editable": False,
"region": record.region,
"source": record.source,
"feed_name": record.feed_name,
"count": 0,
"items": [],
},
)
group["count"] = int(group.get("count") or 0) + 1
group.setdefault("items", []).append(serialized)
manual_items = sorted(manual_by_id.values(), key=lambda item: (int(item.get("sort_order") or 0), str(item.get("name") or "")))
rss_items = sorted(rss_by_id.values(), key=lambda item: str(item.get("name") or ""))
return {"groups": [*manual_items, *rss_items], "manual_groups": manual_items, "rss_groups": rss_items}
async def get_news_record_or_404(db: AsyncSession, item_id: str) -> EarthNewsItem | None:
return await db.get(EarthNewsItem, item_id)
async def delete_manual_news_item(db: AsyncSession, item_id: str) -> bool:
record = await db.get(EarthNewsItem, item_id)
if record is None:
return False
if not _manual_editable(record):
raise PermissionError("RSS 新闻不允许通过手动新闻接口删除。")
await db.execute(delete(EarthNewsItem).where(EarthNewsItem.id == item_id))
await db.flush()
return True
async def reprocess_manual_news_item(db: AsyncSession, item_id: str) -> bool:
record = await db.get(EarthNewsItem, item_id)
if record is None:
return False
if not _manual_editable(record):
raise PermissionError("RSS 新闻不允许通过手动新闻接口重新处理。")
item = record_to_parsed_news_item(record)
queued = await enqueue_target_location_job(build_target_location_job_payload(item), force=True)
if queued:
record.enrichment_status = NewsEnrichmentStatus.QUEUED.value
record.enrichment_error = None
await db.flush()
return queued
async def broadcast_manual_news_changed() -> None:
await _broadcast_news_reload()

View File

@@ -14,10 +14,14 @@ from app.core.logging import get_logger
logger = get_logger(__name__, service="earth_news")
TARGET_LOCATION_STREAM = "earth_news:target_location:jobs"
TARGET_LOCATION_PRIORITY_STREAM = "earth_news:target_location:priority"
TARGET_LOCATION_GROUP = "earth_news_target_location"
TARGET_LOCATION_DEAD_LETTER_STREAM = "earth_news:target_location:dead"
TARGET_LOCATION_RESULT_TTL_SECONDS = 60 * 60 * 12
TARGET_LOCATION_JOB_DEDUP_TTL_SECONDS = 60 * 60 * 6
TARGET_LOCATION_PRIORITY_JOB_DEDUP_TTL_SECONDS = 60 * 5
TARGET_LOCATION_PENDING_RECLAIM_IDLE_MS = 2 * 60 * 1000
TARGET_LOCATION_PRIORITY_READ_BLOCK_MS = 1
TARGET_LOCATION_MAX_ATTEMPTS = 3
_redis_client: redis.Redis | None = None
@@ -28,6 +32,7 @@ class NewsTargetLocationMessage:
message_id: str
item_id: str
payload: dict[str, Any]
stream_name: str = TARGET_LOCATION_STREAM
attempts: int = 0
@@ -44,7 +49,7 @@ class NewsTargetLocationQueue(Protocol):
) -> list[NewsTargetLocationMessage]:
...
async def ack(self, message_id: str) -> None:
async def ack(self, message: NewsTargetLocationMessage) -> None:
...
async def retry_or_dead_letter(
@@ -71,6 +76,10 @@ def _queued_key(item_id: str) -> str:
return f"earth_news:target_location:queued:{item_id}"
def _priority_queued_key(item_id: str) -> str:
return f"earth_news:target_location:priority_queued:{item_id}"
class RedisStreamsNewsTargetLocationQueue:
def __init__(self, client: redis.Redis | None = None) -> None:
self.client = client or _get_redis_client()
@@ -79,34 +88,44 @@ class RedisStreamsNewsTargetLocationQueue:
async def _ensure_group(self) -> None:
if self._group_ready:
return
try:
await self.client.xgroup_create(
TARGET_LOCATION_STREAM,
TARGET_LOCATION_GROUP,
id="0",
mkstream=True,
)
except ResponseError as exc:
if "BUSYGROUP" not in str(exc):
raise
for stream_name in (TARGET_LOCATION_PRIORITY_STREAM, TARGET_LOCATION_STREAM):
try:
await self.client.xgroup_create(
stream_name,
TARGET_LOCATION_GROUP,
id="0",
mkstream=True,
)
except ResponseError as exc:
if "BUSYGROUP" not in str(exc):
raise
self._group_ready = True
async def enqueue(self, *, item_id: str, payload: dict[str, Any], force: bool = False) -> bool:
await self._ensure_group()
if force:
await self.client.delete(_result_key(item_id), _queued_key(item_id))
await self.client.delete(_result_key(item_id))
queued_key = _priority_queued_key(item_id)
elif await self.client.exists(_result_key(item_id)):
return False
else:
queued_key = _queued_key(item_id)
dedup_ttl = (
TARGET_LOCATION_PRIORITY_JOB_DEDUP_TTL_SECONDS
if force
else TARGET_LOCATION_JOB_DEDUP_TTL_SECONDS
)
queued = await self.client.set(
_queued_key(item_id),
queued_key,
"1",
nx=True,
ex=TARGET_LOCATION_JOB_DEDUP_TTL_SECONDS,
ex=dedup_ttl,
)
if not queued:
return bool(await self.client.exists(_queued_key(item_id)))
return bool(await self.client.exists(queued_key))
stream_name = TARGET_LOCATION_PRIORITY_STREAM if force else TARGET_LOCATION_STREAM
await self.client.xadd(
TARGET_LOCATION_STREAM,
stream_name,
{
"item_id": item_id,
"attempts": "0",
@@ -123,39 +142,104 @@ class RedisStreamsNewsTargetLocationQueue:
block_ms: int,
) -> list[NewsTargetLocationMessage]:
await self._ensure_group()
streams = await self.client.xreadgroup(
streams = []
priority_claimed = await self._claim_stale_messages(
stream_name=TARGET_LOCATION_PRIORITY_STREAM,
consumer_name=consumer_name,
count=count,
)
if priority_claimed:
return priority_claimed
priority_messages = await self.client.xreadgroup(
TARGET_LOCATION_GROUP,
consumer_name,
{TARGET_LOCATION_STREAM: ">"},
{TARGET_LOCATION_PRIORITY_STREAM: ">"},
count=count,
block=block_ms,
block=TARGET_LOCATION_PRIORITY_READ_BLOCK_MS,
)
if priority_messages:
streams = priority_messages
else:
regular_claimed = await self._claim_stale_messages(
stream_name=TARGET_LOCATION_STREAM,
consumer_name=consumer_name,
count=count,
)
if regular_claimed:
return regular_claimed
streams = await self.client.xreadgroup(
TARGET_LOCATION_GROUP,
consumer_name,
{TARGET_LOCATION_STREAM: ">"},
count=count,
block=block_ms,
)
messages: list[NewsTargetLocationMessage] = []
for _stream_name, stream_messages in streams:
for stream_name, stream_messages in streams:
for message_id, fields in stream_messages:
raw_payload = fields.get("payload")
item_id = fields.get("item_id")
if not raw_payload or not item_id:
await self.ack(message_id)
continue
try:
payload = json.loads(raw_payload)
except json.JSONDecodeError:
await self.ack(message_id)
continue
attempts = int(fields.get("attempts") or 0)
messages.append(
NewsTargetLocationMessage(
message_id=message_id,
item_id=item_id,
payload=payload,
attempts=attempts,
)
)
message = await self._message_from_fields(stream_name, message_id, fields)
if message is not None:
messages.append(message)
return messages
async def ack(self, message_id: str) -> None:
await self.client.xack(TARGET_LOCATION_STREAM, TARGET_LOCATION_GROUP, message_id)
async def _claim_stale_messages(
self,
*,
stream_name: str,
consumer_name: str,
count: int,
) -> list[NewsTargetLocationMessage]:
try:
_next_id, claimed, _deleted = await self.client.xautoclaim(
stream_name,
TARGET_LOCATION_GROUP,
consumer_name,
TARGET_LOCATION_PENDING_RECLAIM_IDLE_MS,
start_id="0-0",
count=count,
)
except ResponseError:
return []
messages: list[NewsTargetLocationMessage] = []
for message_id, fields in claimed:
message = await self._message_from_fields(stream_name, message_id, fields)
if message is not None:
messages.append(message)
return messages
async def _message_from_fields(
self,
stream_name: str,
message_id: str,
fields: dict[str, str],
) -> NewsTargetLocationMessage | None:
raw_payload = fields.get("payload")
item_id = fields.get("item_id")
if not raw_payload or not item_id:
await self._discard_message(stream_name, message_id)
return None
try:
payload = json.loads(raw_payload)
except json.JSONDecodeError:
await self._discard_message(stream_name, message_id)
return None
attempts = int(fields.get("attempts") or 0)
return NewsTargetLocationMessage(
message_id=message_id,
item_id=item_id,
payload=payload,
stream_name=stream_name,
attempts=attempts,
)
async def ack(self, message: NewsTargetLocationMessage) -> None:
await self.client.xack(message.stream_name, TARGET_LOCATION_GROUP, message.message_id)
await self.client.xdel(message.stream_name, message.message_id)
async def _discard_message(self, stream_name: str, message_id: str) -> None:
await self.client.xack(stream_name, TARGET_LOCATION_GROUP, message_id)
await self.client.xdel(stream_name, message_id)
async def retry_or_dead_letter(
self,
@@ -163,7 +247,7 @@ class RedisStreamsNewsTargetLocationQueue:
*,
error: str,
) -> None:
await self.ack(message.message_id)
await self.ack(message)
if message.attempts + 1 >= TARGET_LOCATION_MAX_ATTEMPTS:
await self.client.xadd(
TARGET_LOCATION_DEAD_LETTER_STREAM,
@@ -176,7 +260,7 @@ class RedisStreamsNewsTargetLocationQueue:
)
return
await self.client.xadd(
TARGET_LOCATION_STREAM,
message.stream_name,
{
"item_id": message.item_id,
"attempts": str(message.attempts + 1),
@@ -231,4 +315,4 @@ async def save_target_location_patch(item_id: str, patch: dict[str, Any]) -> Non
TARGET_LOCATION_RESULT_TTL_SECONDS,
json.dumps(patch, ensure_ascii=False),
)
await client.delete(_queued_key(item_id))
await client.delete(_queued_key(item_id), _priority_queued_key(item_id))

View File

@@ -3,7 +3,7 @@ from __future__ import annotations
from datetime import UTC, datetime
from typing import Any
from sqlalchemy import func, select
from sqlalchemy import func, or_, select
from sqlalchemy.ext.asyncio import AsyncSession
from app.models.earth_news import EarthNewsItem
@@ -11,7 +11,24 @@ from app.services.earth_news import (
ParsedNewsItem,
apply_enrichment_patch_to_item,
build_anchor_location_patch,
_news_meta_patch,
)
from app.services.earth_news_classification import (
breaking_sort_rank,
normalize_breaking_level,
normalize_breaking_scope,
)
CRUISE_REGION_ORDER = (
"americas",
"europe",
"middle-east-africa",
"asia-pacific",
"global",
)
CRUISE_REGION_QUERY_MULTIPLIER = 12
CRUISE_REGION_QUERY_MIN_LIMIT = 240
CRUISE_REGION_QUERY_MAX_LIMIT = 1000
def _coerce_datetime(value: datetime | None) -> datetime | None:
@@ -22,6 +39,17 @@ def _coerce_datetime(value: datetime | None) -> datetime | None:
return value.astimezone(UTC)
def _coerce_meta_datetime(value: Any) -> datetime | None:
if isinstance(value, datetime):
return _coerce_datetime(value)
if not isinstance(value, str) or not value.strip():
return None
try:
return _coerce_datetime(datetime.fromisoformat(value.replace("Z", "+00:00")))
except ValueError:
return None
def _location_patch_from_record(record: EarthNewsItem) -> dict[str, Any]:
return {
"latitude": record.latitude,
@@ -34,6 +62,8 @@ def _location_patch_from_record(record: EarthNewsItem) -> dict[str, Any]:
def record_to_parsed_news_item(record: EarthNewsItem) -> ParsedNewsItem:
location_meta = dict(record.location_meta or {})
news_meta = location_meta.get("news_meta") if isinstance(location_meta.get("news_meta"), dict) else {}
item = ParsedNewsItem(
id=record.id,
title=record.title,
@@ -49,11 +79,86 @@ def record_to_parsed_news_item(record: EarthNewsItem) -> ParsedNewsItem:
enrichment_status=record.enrichment_status or "pending",
enrichment_error=record.enrichment_error,
enriched_at=_coerce_datetime(record.enriched_at),
source_tags=list(news_meta.get("source_tags") or []),
feed_id=str(news_meta.get("feed_id") or ""),
feed_type=str(news_meta.get("feed_type") or "rss"),
feed_default_category=str(news_meta.get("feed_default_category") or "other"),
category=str(news_meta.get("category") or "other"),
item_tags=list(news_meta.get("item_tags") or []),
tagging_source=str(news_meta.get("tagging_source") or "rules"),
tagging_confidence=float(news_meta.get("tagging_confidence") or 0),
importance_score=int(news_meta.get("importance_score") or 0),
importance_level=str(news_meta.get("importance_level") or "low"),
importance_reasons=list(news_meta.get("importance_reasons") or []),
market_impact=str(news_meta.get("market_impact") or "none"),
breaking_level=normalize_breaking_level(news_meta.get("breaking_level")).value,
breaking_scope=normalize_breaking_scope(news_meta.get("breaking_scope")).value,
breaking_reasons=list(news_meta.get("breaking_reasons") or []),
breaking_source=str(news_meta.get("breaking_source") or "rules"),
breaking_confidence=float(news_meta.get("breaking_confidence") or 0),
breaking_expires_at=_coerce_meta_datetime(news_meta.get("breaking_expires_at")),
)
return apply_enrichment_patch_to_item(item, _location_patch_from_record(record))
def _sort_parsed_news_items(items: list[ParsedNewsItem], *, active_region: str) -> list[ParsedNewsItem]:
return sorted(
items,
key=lambda item: (
-breaking_sort_rank(item),
False
if active_region == "global"
or (breaking_sort_rank(item) > 0 and normalize_breaking_scope(item.breaking_scope).value == "global")
else item.feed_region != active_region,
item.published_at is None,
-(item.published_at.timestamp() if item.published_at else 0),
item.feed_name,
),
)
def _diversify_parsed_news_items_by_region(
items: list[ParsedNewsItem],
*,
limit: int,
) -> list[ParsedNewsItem]:
if limit <= 0:
return []
sorted_items = _sort_parsed_news_items(items, active_region="global")
buckets: dict[str, list[ParsedNewsItem]] = {}
for item in sorted_items:
region = item.feed_region or "global"
buckets.setdefault(region, []).append(item)
ordered_regions = [
*[region for region in CRUISE_REGION_ORDER if buckets.get(region)],
*sorted(region for region in buckets if region not in CRUISE_REGION_ORDER),
]
diversified: list[ParsedNewsItem] = []
cursor = 0
while len(diversified) < limit:
added = False
for region in ordered_regions:
bucket = buckets.get(region) or []
if cursor >= len(bucket):
continue
diversified.append(bucket[cursor])
added = True
if len(diversified) >= limit:
break
if not added:
break
cursor += 1
return diversified
def _query_sort_key(active_region: str):
if active_region == "global":
return (
EarthNewsItem.published_at.is_(None),
EarthNewsItem.published_at.desc().nullslast(),
EarthNewsItem.feed_name.asc(),
)
return (
EarthNewsItem.region != active_region,
EarthNewsItem.published_at.is_(None),
@@ -62,38 +167,90 @@ def _query_sort_key(active_region: str):
)
def _category_filter_clause(categories: set[str] | None):
if not categories:
return None
return EarthNewsItem.location_meta.op("->")("news_meta").op("->>")("category").in_(sorted(categories))
def _source_filter_clause(source_ids: set[str] | None):
if not source_ids:
return None
return EarthNewsItem.location_meta.op("->")("news_meta").op("->>")("source_id").in_(sorted(source_ids))
async def list_earth_news_items(
db: AsyncSession,
*,
active_region: str,
limit: int,
categories: set[str] | None = None,
source_ids: set[str] | None = None,
) -> list[ParsedNewsItem]:
regions = {"global", active_region}
result = await db.execute(
query_limit = limit if source_ids else min(max(limit * 20, limit), 500)
query = (
select(EarthNewsItem)
.where(EarthNewsItem.region.in_(regions))
.order_by(*_query_sort_key(active_region))
.limit(limit)
.limit(query_limit)
)
return [record_to_parsed_news_item(record) for record in result.scalars().all()]
if active_region != "global":
news_meta = EarthNewsItem.location_meta.op("->")("news_meta")
query = query.where(
or_(
EarthNewsItem.region.in_({"global", active_region}),
news_meta.op("->>")("breaking_scope") == "global",
)
)
category_clause = _category_filter_clause(categories)
if category_clause is not None:
query = query.where(category_clause)
source_clause = _source_filter_clause(source_ids)
if source_clause is not None:
query = query.where(source_clause)
result = await db.execute(query)
records = list(result.scalars().all())
items = _sort_parsed_news_items(
[record_to_parsed_news_item(record) for record in records],
active_region=active_region,
)
if active_region == "global" and not source_ids:
return _diversify_parsed_news_items_by_region(items, limit=limit)
return items[:limit]
async def list_earth_news_cruise_items(
db: AsyncSession,
*,
limit: int,
categories: set[str] | None = None,
source_ids: set[str] | None = None,
) -> list[ParsedNewsItem]:
result = await db.execute(
query_limit = min(
max(limit * CRUISE_REGION_QUERY_MULTIPLIER, CRUISE_REGION_QUERY_MIN_LIMIT),
CRUISE_REGION_QUERY_MAX_LIMIT,
)
query = (
select(EarthNewsItem)
.order_by(
EarthNewsItem.region.asc(),
EarthNewsItem.published_at.is_(None),
EarthNewsItem.published_at.desc().nullslast(),
EarthNewsItem.last_seen_at.desc(),
EarthNewsItem.region.asc(),
EarthNewsItem.feed_name.asc(),
)
.limit(limit)
.limit(query_limit)
)
category_clause = _category_filter_clause(categories)
if category_clause is not None:
query = query.where(category_clause)
source_clause = _source_filter_clause(source_ids)
if source_clause is not None:
query = query.where(source_clause)
result = await db.execute(query)
return _diversify_parsed_news_items_by_region(
[record_to_parsed_news_item(record) for record in result.scalars().all()],
limit=limit,
)
return [record_to_parsed_news_item(record) for record in result.scalars().all()]
async def get_earth_news_freshness(
@@ -101,13 +258,13 @@ async def get_earth_news_freshness(
*,
active_region: str,
) -> tuple[int, datetime | None]:
regions = {"global", active_region}
result = await db.execute(
select(
func.count(EarthNewsItem.id),
func.max(func.coalesce(EarthNewsItem.published_at, EarthNewsItem.last_seen_at)),
).where(EarthNewsItem.region.in_(regions))
query = select(
func.count(EarthNewsItem.id),
func.max(func.coalesce(EarthNewsItem.published_at, EarthNewsItem.last_seen_at)),
)
if active_region != "global":
query = query.where(EarthNewsItem.region.in_({"global", active_region}))
result = await db.execute(query)
count, newest = result.one()
item_count = int(count or 0)
if item_count == 0:
@@ -115,6 +272,33 @@ async def get_earth_news_freshness(
return item_count, _coerce_datetime(newest)
async def get_earth_news_feed_coverage(
db: AsyncSession,
*,
active_region: str,
recent_after: datetime | None = None,
) -> set[tuple[str, str]]:
query = select(
EarthNewsItem.id,
EarthNewsItem.location_meta.op("->")("news_meta").op("->>")("source_id"),
EarthNewsItem.location_meta.op("->")("news_meta").op("->>")("feed_id"),
)
if active_region != "global":
query = query.where(EarthNewsItem.region.in_({"global", active_region}))
if recent_after is not None:
query = query.where(func.coalesce(EarthNewsItem.published_at, EarthNewsItem.last_seen_at) >= recent_after)
result = await db.execute(query)
coverage: set[tuple[str, str]] = set()
for item_id, source_id, feed_id in result.all():
normalized_source_id = str(source_id or "").strip()
normalized_feed_id = str(feed_id or "").strip()
if not normalized_source_id and isinstance(item_id, str) and ":" in item_id:
normalized_source_id = item_id.split(":", 1)[0]
if normalized_source_id and normalized_feed_id:
coverage.add((normalized_source_id, normalized_feed_id))
return coverage
async def upsert_earth_news_items(db: AsyncSession, items: list[ParsedNewsItem]) -> int:
if not items:
return 0
@@ -165,12 +349,20 @@ async def upsert_earth_news_items(db: AsyncSession, items: list[ParsedNewsItem])
record.homepage_url = item.homepage_url
record.published_at = item.published_at
record.last_seen_at = now
location_meta = dict(record.location_meta or {})
location_meta["news_meta"] = _news_meta_patch(item)
record.location_meta = location_meta
if item.localizations:
merged_localizations = {
**dict(record.localizations or {}),
**dict(item.localizations or {}),
}
record.content_language = item.content_language
record.localizations = dict(item.localizations or {})
record.enrichment_status = item.enrichment_status
record.enrichment_error = item.enrichment_error
record.enriched_at = item.enriched_at
record.localizations = merged_localizations
if item.enrichment_status != "pending" or item.enrichment_error or item.enriched_at:
record.enrichment_status = item.enrichment_status
record.enrichment_error = item.enrichment_error
record.enriched_at = item.enriched_at
changed += 1
await db.flush()
return changed
@@ -206,13 +398,28 @@ async def update_earth_news_item_enrichment(
if record is None:
return False
if "latitude" in patch:
record.latitude = float(patch["latitude"])
record.longitude = float(patch["longitude"])
record.location_label = str(patch["location_label"])
record.location_source = str(patch["location_source"])
record.verified = bool(patch["verified"])
record.location_meta = dict(patch.get("location_meta") or {})
record.resolved_at = datetime.now(UTC) if record.verified else None
patch_meta = dict(patch.get("location_meta") or {})
if record.location_source == "manual_location":
current_meta = dict(record.location_meta or {})
patch_news_meta = patch_meta.get("news_meta")
if isinstance(patch_news_meta, dict):
current_meta["news_meta"] = patch_news_meta
current_meta["manual_enrichment"] = {
"resolution_stage": patch_meta.get("resolution_stage"),
"ai_attempted": patch_meta.get("ai_attempted"),
"ai_status": patch_meta.get("ai_status"),
"ai_error": patch_meta.get("ai_error"),
"debug_note": patch_meta.get("debug_note"),
}
record.location_meta = current_meta
else:
record.latitude = float(patch["latitude"])
record.longitude = float(patch["longitude"])
record.location_label = str(patch["location_label"])
record.location_source = str(patch["location_source"])
record.verified = bool(patch["verified"])
record.location_meta = patch_meta
record.resolved_at = datetime.now(UTC) if record.verified else None
if "content_language" in patch:
record.content_language = str(patch.get("content_language") or "en")
if "localizations" in patch:

View File

@@ -29,6 +29,9 @@ logger = get_logger(__name__, service="earth_news")
WORKER_BATCH_SIZE = 4
WORKER_BLOCK_MS = 5000
WORKER_BACKOFF_SECONDS = 5.0
WORKER_JOB_TIMEOUT_MIN_SECONDS = 20.0
WORKER_JOB_TIMEOUT_MAX_SECONDS = 90.0
WORKER_JOB_TIMEOUT_GRACE_SECONDS = 10.0
_worker_task: asyncio.Task | None = None
@@ -109,12 +112,25 @@ async def _run_target_location_worker() -> None:
if not messages:
continue
provider_client = await _build_provider_client()
for message in messages:
job_timeout = _get_worker_job_timeout(provider_client)
async def handle_message(message: NewsTargetLocationMessage) -> None:
try:
await process_target_location_message(message, provider_client=provider_client)
await queue.ack(message.message_id)
await asyncio.wait_for(
process_target_location_message(message, provider_client=provider_client),
timeout=job_timeout,
)
await queue.ack(message)
except asyncio.CancelledError:
raise
except TimeoutError as exc:
logger.warning_event(
"Earth news target location worker job timed out",
event="earth_news.target_location.worker_job_timeout",
context={"item_id": message.item_id, "timeout_seconds": job_timeout},
)
with suppress(Exception):
await queue.retry_or_dead_letter(message, error=str(exc) or "job timed out")
except Exception as exc:
logger.warning_event(
"Earth news target location worker job failed",
@@ -124,6 +140,8 @@ async def _run_target_location_worker() -> None:
with suppress(Exception):
await queue.retry_or_dead_letter(message, error=str(exc))
await asyncio.gather(*(handle_message(message) for message in messages))
def start_earth_news_target_worker() -> None:
global _worker_task
@@ -140,3 +158,15 @@ async def stop_earth_news_target_worker() -> None:
with suppress(asyncio.CancelledError):
await task
_worker_task = None
def _get_worker_job_timeout(provider_client: AIProviderClient | None) -> float:
timeout = float(getattr(provider_client, "timeout", 0) or WORKER_JOB_TIMEOUT_MIN_SECONDS)
retry_attempts = float(getattr(provider_client, "retry_attempts", 1) or 1)
return min(
max(
timeout * retry_attempts + WORKER_JOB_TIMEOUT_GRACE_SECONDS,
WORKER_JOB_TIMEOUT_MIN_SECONDS,
),
WORKER_JOB_TIMEOUT_MAX_SECONDS,
)

View File

@@ -9,12 +9,12 @@ yet to keep behavior obvious after settings changes).
from __future__ import annotations
from email.message import EmailMessage
from typing import Literal, Optional
from typing import Optional
import aiosmtplib
from sqlalchemy.ext.asyncio import AsyncSession
OtpPurpose = Literal["register", "verify_email", "reset_password"]
from app.core.enums import OtpPurpose
class EmailError(Exception):
@@ -81,15 +81,15 @@ async def send_email(
_SUBJECTS: dict[OtpPurpose, str] = {
"register": "Confirm your Planet account",
"verify_email": "Verify your Planet email",
"reset_password": "Reset your Planet password",
OtpPurpose.REGISTER: "Confirm your Planet account",
OtpPurpose.VERIFY_EMAIL: "Verify your Planet email",
OtpPurpose.RESET_PASSWORD: "Reset your Planet password",
}
_HEADLINES: dict[OtpPurpose, str] = {
"register": "Welcome to Planet — confirm your email to activate your account.",
"verify_email": "Confirm your new email address to keep your Planet account active.",
"reset_password": "Use this code to set a new password for your Planet account.",
OtpPurpose.REGISTER: "Welcome to Planet — confirm your email to activate your account.",
OtpPurpose.VERIFY_EMAIL: "Confirm your new email address to keep your Planet account active.",
OtpPurpose.RESET_PASSWORD: "Use this code to set a new password for your Planet account.",
}

View File

@@ -0,0 +1,161 @@
from __future__ import annotations
import asyncio
from dataclasses import dataclass, field
from datetime import UTC, datetime
from typing import Any
from fastapi import WebSocket
from app.db.session import async_session_factory
from app.services.system_logs import (
DEFAULT_LOG_LINE_LIMIT,
LOG_SOURCES,
MAX_LOG_LINE_LIMIT,
read_database_log_events,
read_log_events,
)
DATABASE_LOG_SOURCE_IDS = {"system-db", "audit-db"}
LOG_TAIL_CHANNEL = "logs_tail"
LOG_TAIL_INTERVAL_SECONDS = 1.5
LOG_TAIL_SCAN_MULTIPLIER = 5
@dataclass(frozen=True)
class LogTailConfig:
source_id: str
limit: int = DEFAULT_LOG_LINE_LIMIT
level: str = "all"
levels: str | None = None
start_date: str | None = None
end_date: str | None = None
search: str | None = None
@dataclass
class LogTailSubscription:
config: LogTailConfig
emitted_cursors: set[str] = field(default_factory=set)
task: asyncio.Task | None = None
class LogTailManager:
def __init__(self) -> None:
self._subscriptions: dict[WebSocket, LogTailSubscription] = {}
def normalize_config(self, payload: dict[str, Any]) -> LogTailConfig:
source_id = str(payload.get("source_id") or payload.get("source") or "").strip()
if not source_id:
raise ValueError("source_id is required")
if source_id not in LOG_SOURCES and source_id not in DATABASE_LOG_SOURCE_IDS:
raise ValueError("Log source not found")
try:
limit = int(payload.get("limit") or DEFAULT_LOG_LINE_LIMIT)
except (TypeError, ValueError) as exc:
raise ValueError("limit must be a number") from exc
if limit < 1 or limit > MAX_LOG_LINE_LIMIT:
raise ValueError(f"limit must be between 1 and {MAX_LOG_LINE_LIMIT}")
return LogTailConfig(
source_id=source_id,
limit=limit,
level=str(payload.get("level") or "all"),
levels=str(payload.get("levels")).strip() if payload.get("levels") else None,
start_date=str(payload.get("start_date")).strip() if payload.get("start_date") else None,
end_date=str(payload.get("end_date")).strip() if payload.get("end_date") else None,
search=str(payload.get("search")).strip() if payload.get("search") else None,
)
async def subscribe(self, websocket: WebSocket, payload: dict[str, Any]) -> LogTailConfig:
config = self.normalize_config(payload)
await self.unsubscribe(websocket)
subscription = LogTailSubscription(config=config)
subscription.task = asyncio.create_task(self._run_tail(websocket, subscription))
self._subscriptions[websocket] = subscription
return config
async def unsubscribe(self, websocket: WebSocket) -> None:
subscription = self._subscriptions.pop(websocket, None)
if subscription and subscription.task:
subscription.task.cancel()
try:
await subscription.task
except asyncio.CancelledError:
pass
async def disconnect(self, websocket: WebSocket) -> None:
await self.unsubscribe(websocket)
async def _run_tail(self, websocket: WebSocket, subscription: LogTailSubscription) -> None:
first_frame = True
while True:
events = await self._read_events(subscription.config)
if first_frame:
visible_events = events[-subscription.config.limit :]
subscription.emitted_cursors.update(event.cursor for event in visible_events)
await self._send_frame(websocket, subscription.config, "snapshot", visible_events)
first_frame = False
else:
new_events = [
event
for event in events
if event.cursor not in subscription.emitted_cursors
]
if new_events:
visible_events = new_events[-subscription.config.limit :]
subscription.emitted_cursors.update(event.cursor for event in visible_events)
await self._send_frame(websocket, subscription.config, "append", visible_events)
await asyncio.sleep(LOG_TAIL_INTERVAL_SECONDS)
async def _read_events(self, config: LogTailConfig):
scan_limit = max(config.limit * LOG_TAIL_SCAN_MULTIPLIER, config.limit)
if config.source_id in DATABASE_LOG_SOURCE_IDS:
async with async_session_factory() as db:
events = await read_database_log_events(
config.source_id,
scan_limit=scan_limit,
level=config.level,
levels=config.levels,
start_date=config.start_date,
end_date=config.end_date,
search=config.search,
db=db,
)
return events or []
events = read_log_events(
config.source_id,
scan_limit=scan_limit,
level=config.level,
levels=config.levels,
start_date=config.start_date,
end_date=config.end_date,
search=config.search,
)
return events or []
async def _send_frame(self, websocket: WebSocket, config: LogTailConfig, mode: str, events) -> None:
await websocket.send_json(
{
"type": "data_frame",
"channel": LOG_TAIL_CHANNEL,
"timestamp": datetime.now(UTC).isoformat(),
"payload": {
"mode": mode,
"source_id": config.source_id,
"line_count": len(events),
"lines": [event.line for event in events],
"filters": {
"limit": config.limit,
"level": config.level,
"levels": config.levels,
"start_date": config.start_date,
"end_date": config.end_date,
"search": config.search,
},
"status": "ok",
},
}
)
log_tail_manager = LogTailManager()

View File

@@ -9,14 +9,12 @@ from __future__ import annotations
import json
import secrets
from typing import Literal
import bcrypt
from app.core.enums import OtpPurpose
from app.core.security import redis_client
OtpPurpose = Literal["register", "verify_email", "reset_password"]
CODE_TTL_SECONDS = 600 # 10 minutes
RESEND_COOLDOWN_SECONDS = 60
MAX_ATTEMPTS = 5

View File

@@ -1,14 +1,185 @@
from __future__ import annotations
import hashlib
import re
from datetime import UTC, datetime
from typing import Any
from app.core.logging import get_logger, sanitize_log_value
from app.core.request_context import get_request_id
from app.db.session import async_session_factory
from app.models.system_log import AuditLog, SystemLog
from app.models.system_log import AuditLog, ObservabilityEvent, ObservabilityEventGroup, SystemLog
logger = get_logger(__name__)
HLS_TRANSIENT_RE = re.compile(r"(index|chunk|segment)[_-]?\d+(?:_\d+)?\.(?:ts|m4s|vtt)", re.IGNORECASE)
QUERY_RE = re.compile(r"([?&](?:m|t|token|expires|signature|X-Amz-[^=]+)=[^&\\s]+)", re.IGNORECASE)
UUID_RE = re.compile(r"\b[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}\b", re.IGNORECASE)
CONNECTION_RE = re.compile(r"\bconn_[A-Za-z0-9:._-]+\b")
NUMBER_RE = re.compile(r"\b\d{5,}\b")
def normalize_observability_text(value: Any) -> str:
text = str(sanitize_log_value(value or "")).strip()
text = QUERY_RE.sub("", text)
text = HLS_TRANSIENT_RE.sub("<hls-fragment>", text)
text = UUID_RE.sub("<uuid>", text)
text = CONNECTION_RE.sub("<connection>", text)
text = NUMBER_RE.sub("<number>", text)
return re.sub(r"\s+", " ", text).strip()
def build_observability_fingerprint(
*,
source: str,
service: str | None = None,
module: str | None = None,
category: str | None = None,
event: str | None = None,
message: str,
context: dict[str, Any] | None = None,
) -> str:
context = context or {}
stable_context = {
key: context.get(key)
for key in (
"task_type",
"source_id",
"source",
"provider",
"status_code",
"error_type",
"details",
)
if context.get(key) not in (None, "")
}
raw = "|".join(
[
normalize_observability_text(source),
normalize_observability_text(service),
normalize_observability_text(module),
normalize_observability_text(category),
normalize_observability_text(event),
normalize_observability_text(message),
normalize_observability_text(stable_context),
]
)
return hashlib.sha1(raw.encode("utf-8", errors="replace")).hexdigest()
def _context_text(context: dict[str, Any] | None, key: str) -> str | None:
value = (context or {}).get(key)
if value in (None, ""):
return None
return str(value)
async def record_observability_event(
*,
source: str,
level: str,
message: str,
service: str | None = None,
module: str | None = None,
event: str | None = None,
request_id: str | None = None,
trace_id: str | None = None,
user_id: int | None = None,
category: str | None = None,
context: dict[str, Any] | None = None,
fingerprint: str | None = None,
occurred_at: datetime | None = None,
occurrence_count: int = 1,
) -> None:
normalized_context = sanitize_log_value(context or {})
if not isinstance(normalized_context, dict):
normalized_context = {"value": normalized_context}
safe_message = str(sanitize_log_value(message))
normalized_level = str(level or "info").lower()
count = max(1, int(occurrence_count or 1))
event_time = occurred_at or datetime.now(UTC)
event_fingerprint = fingerprint or build_observability_fingerprint(
source=source,
service=service,
module=module,
category=category,
event=event,
message=safe_message,
context=normalized_context,
)
detail = _context_text(normalized_context, "detail") or _context_text(normalized_context, "error")
affected_sources = sorted(
{
item
for item in (
source,
service,
module,
_context_text(normalized_context, "source_id"),
_context_text(normalized_context, "source"),
)
if item
}
)
try:
async with async_session_factory() as session:
session.add(
ObservabilityEvent(
source=source,
service=service,
module=module,
category=category,
event=event,
level=normalized_level,
message=safe_message,
fingerprint=event_fingerprint,
occurred_at=event_time,
request_id=request_id or get_request_id(),
trace_id=trace_id,
user_id=user_id,
task_id=_context_text(normalized_context, "task_id"),
source_ref_id=_context_text(normalized_context, "source_id") or _context_text(normalized_context, "source"),
provider=_context_text(normalized_context, "provider"),
context=normalized_context,
occurrence_count=count,
)
)
group = await session.get(ObservabilityEventGroup, event_fingerprint)
if group is None:
session.add(
ObservabilityEventGroup(
fingerprint=event_fingerprint,
source=source,
service=service,
module=module,
category=category,
event=event,
last_level=normalized_level,
sample_message=safe_message,
sample_detail=detail,
affected_sources=affected_sources,
count=count,
first_seen_at=event_time,
last_seen_at=event_time,
)
)
else:
group.count = int(group.count or 0) + count
group.last_seen_at = event_time
group.last_level = normalized_level
group.sample_message = safe_message
group.sample_detail = detail
merged_sources = sorted(set(group.affected_sources or []) | set(affected_sources))
group.affected_sources = merged_sources
await session.commit()
except Exception:
logger.exception_event(
"Failed to persist observability event",
event="observability_event.persist.failed",
context={"event_name": event, "source": source},
)
async def record_system_log(
*,
@@ -23,6 +194,8 @@ async def record_system_log(
user_id: int | None = None,
category: str | None = None,
context: dict[str, Any] | None = None,
fingerprint: str | None = None,
occurrence_count: int = 1,
) -> None:
try:
async with async_session_factory() as session:
@@ -48,6 +221,21 @@ async def record_system_log(
event="system_log.persist.failed",
context={"event_name": event, "source": source},
)
await record_observability_event(
source=source,
service=service,
module=module,
event=event,
level=level,
message=message,
request_id=request_id,
trace_id=trace_id,
user_id=user_id,
category=category,
context=context,
fingerprint=fingerprint,
occurrence_count=occurrence_count,
)
async def record_audit_log(

View File

@@ -7,9 +7,14 @@ from time import perf_counter
from uuid import uuid4
from fastapi import HTTPException, status
from sqlalchemy import func, select, update
from sqlalchemy import func, select
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.enums import (
PlaygroundMessageKind,
PlaygroundMessageRole,
PlaygroundMessageStatus,
)
from app.core.logging import get_logger
from app.db.session import async_session_factory
from app.models.playground_message import PlaygroundMessage
@@ -21,7 +26,6 @@ from app.schemas.ai import (
PlaygroundMessageRecord,
PlaygroundMessageResendRequest,
PlaygroundMessageStopRequest,
PlaygroundSessionResponse,
PlaygroundSessionState,
PlaygroundSessionUpsertRequest,
PlaygroundThreadResponse,
@@ -37,6 +41,13 @@ STREAM_CHUNK_SIZE = 24
STREAM_INTERVAL_SECONDS = 0.08
THINKING_PREVIEW_SECONDS = 2.6
ORPHANED_RUN_MESSAGE = "后台生成任务已中断,请点击上一条用户消息的重试按钮重新生成。"
ACTIVE_MESSAGE_STATUSES = frozenset(
{
PlaygroundMessageStatus.PENDING.value,
PlaygroundMessageStatus.THINKING.value,
PlaygroundMessageStatus.ANSWERING.value,
}
)
class _ActiveRun:
@@ -93,7 +104,11 @@ async def _require_visible_message(
result = await db.execute(select(PlaygroundMessage).where(*conditions))
message = result.scalar_one_or_none()
if message is None:
detail = "User message not found" if role == "user" else "Playground message not found"
detail = (
"User message not found"
if role == PlaygroundMessageRole.USER.value
else "Playground message not found"
)
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail=detail)
return message
@@ -108,7 +123,7 @@ def _message_to_record(message: PlaygroundMessage, parent_public_id: str | None
content=message.content or "",
thinking_content=message.thinking_content or "",
meta=list(message.meta or []),
markdown=message.role != "system",
markdown=message.role != PlaygroundMessageRole.SYSTEM.value,
provider=message.provider,
model=message.model,
request_id=message.request_id,
@@ -197,11 +212,11 @@ async def _reconcile_orphaned_active_messages(
) -> list[PlaygroundMessage]:
changed = False
for item in messages:
if item.status not in {"pending", "thinking", "answering"}:
if item.status not in ACTIVE_MESSAGE_STATUSES:
continue
if item.public_id in _ACTIVE_RUNS:
continue
item.status = "error"
item.status = PlaygroundMessageStatus.ERROR.value
item.content = item.content or ORPHANED_RUN_MESSAGE
orphan_meta = "错误: 后台任务已中断"
if orphan_meta not in (item.meta or []):
@@ -330,9 +345,9 @@ async def create_turn(
public_id=uuid4().hex,
session_id=session.id,
user_id=user_id,
role="user",
kind="message",
status="done",
role=PlaygroundMessageRole.USER.value,
kind=PlaygroundMessageKind.MESSAGE.value,
status=PlaygroundMessageStatus.DONE.value,
title=payload.selected_preset_key,
content=payload.input,
meta=[payload.title],
@@ -343,9 +358,9 @@ async def create_turn(
session_id=session.id,
user_id=user_id,
parent_message_id=None,
role="assistant",
kind="thinking",
status="pending",
role=PlaygroundMessageRole.ASSISTANT.value,
kind=PlaygroundMessageKind.THINKING.value,
status=PlaygroundMessageStatus.PENDING.value,
title="AI 回应",
content="",
thinking_content="",
@@ -397,9 +412,9 @@ async def _create_assistant_retry_turn(
session_id=session.id,
user_id=user_id,
parent_message_id=user_message.id,
role="assistant",
kind="thinking",
status="pending",
role=PlaygroundMessageRole.ASSISTANT.value,
kind=PlaygroundMessageKind.THINKING.value,
status=PlaygroundMessageStatus.PENDING.value,
title="AI 回应",
content="",
thinking_content="",
@@ -439,7 +454,7 @@ async def stop_message(
session = await _require_session(db, user_id=user_id, session_key=payload.session_key)
message = await _require_visible_message(db, user_id=user_id, public_id=payload.message_id)
if message.status not in {"pending", "thinking", "answering"}:
if message.status not in ACTIVE_MESSAGE_STATUSES:
return await _build_action_response(db, session=session)
active_run = _ACTIVE_RUNS.get(message.public_id)
@@ -447,7 +462,7 @@ async def stop_message(
active_run.stop_requested.set()
active_run.task.cancel()
message.status = "stopped"
message.status = PlaygroundMessageStatus.STOPPED.value
if "已手动停止生成" not in (message.meta or []):
message.meta = [*(message.meta or []), "已手动停止生成"]
await db.flush()
@@ -469,7 +484,7 @@ async def resend_turn(
db,
user_id=user_id,
public_id=payload.user_message_id,
role="user",
role=PlaygroundMessageRole.USER.value,
)
later_messages = await db.execute(
@@ -481,7 +496,7 @@ async def resend_turn(
)
for item in later_messages.scalars().all():
item.is_visible = False
if item.status in {"pending", "thinking", "answering"}:
if item.status in ACTIVE_MESSAGE_STATUSES:
active_run = _ACTIVE_RUNS.get(item.public_id)
if active_run is not None:
active_run.stop_requested.set()
@@ -520,7 +535,7 @@ async def edit_user_message(
db,
user_id=user_id,
public_id=payload.user_message_id,
role="user",
role=PlaygroundMessageRole.USER.value,
)
user_message.content = payload.content.strip()
@@ -566,12 +581,12 @@ def _build_conversation_history(messages: Sequence[PlaygroundMessage], current_u
for item in messages:
if item.id >= current_user_message_id:
break
if item.role == "system":
if item.role == PlaygroundMessageRole.SYSTEM.value:
continue
history.append(
{
"role": item.role,
"kind": item.kind or "message",
"kind": item.kind or PlaygroundMessageKind.MESSAGE.value,
"title": item.title,
"content": item.content or "",
}
@@ -668,7 +683,11 @@ async def _run_assistant_message(
assistant_message = await _mark_message_state(
db,
message_id=assistant_message_id,
status="thinking" if analysis.thinking_blocks else "answering",
status=(
PlaygroundMessageStatus.THINKING.value
if analysis.thinking_blocks
else PlaygroundMessageStatus.ANSWERING.value
),
title=f"{analysis.provider} / {analysis.model}",
provider=analysis.provider,
model=analysis.model,
@@ -706,7 +725,7 @@ async def _run_assistant_message(
await _mark_message_state(
db,
message_id=assistant_message_id,
status="answering",
status=PlaygroundMessageStatus.ANSWERING.value,
content=content[:cursor],
)
await db.commit()
@@ -717,7 +736,7 @@ async def _run_assistant_message(
assistant_message = await _mark_message_state(
db,
message_id=assistant_message_id,
status="done",
status=PlaygroundMessageStatus.DONE.value,
content=content,
meta=[
f"Request ID: {request_id}",
@@ -762,8 +781,8 @@ async def _run_assistant_message(
async with async_session_factory() as db:
result = await db.execute(select(PlaygroundMessage).where(PlaygroundMessage.id == assistant_message_id))
message = result.scalar_one_or_none()
if message is not None and message.status in {"pending", "thinking", "answering"}:
message.status = "stopped"
if message is not None and message.status in ACTIVE_MESSAGE_STATUSES:
message.status = PlaygroundMessageStatus.STOPPED.value
if "已手动停止生成" not in (message.meta or []):
message.meta = [*(message.meta or []), "已手动停止生成"]
await db.flush()
@@ -795,7 +814,7 @@ async def _run_assistant_message(
result = await db.execute(select(PlaygroundMessage).where(PlaygroundMessage.id == assistant_message_id))
message = result.scalar_one_or_none()
if message is not None:
message.status = "error"
message.status = PlaygroundMessageStatus.ERROR.value
message.content = message.content or f"分析失败:{error_message}"
message.meta = [
*(message.meta or []),

View File

@@ -8,6 +8,7 @@ from apscheduler.schedulers.asyncio import AsyncIOScheduler
from apscheduler.triggers.interval import IntervalTrigger
from sqlalchemy import select
from app.core.enums import JobStatus
from app.core.logging import get_logger
from app.db.session import async_session_factory
from app.core.time import to_iso8601_utc
@@ -140,7 +141,7 @@ async def run_collector_task(collector_name: str):
select(CollectionTask)
.where(
CollectionTask.datasource_id == datasource.id,
CollectionTask.status == "running",
CollectionTask.status == JobStatus.RUNNING.value,
)
.order_by(CollectionTask.started_at.desc(), CollectionTask.id.desc())
.limit(1)
@@ -184,7 +185,7 @@ async def run_collector_task(collector_name: str):
f"Marked failed automatically after stale running timeout "
f"({RUNNING_TASK_GUARD_TIMEOUT_MINUTES}m) in scheduler guard"
)
existing_running.status = "failed"
existing_running.status = JobStatus.FAILED.value
existing_running.phase = "failed"
existing_running.completed_at = now
existing_running.error_message = (
@@ -243,7 +244,7 @@ async def run_collector_task(collector_name: str):
return
datasource.last_run_at = datetime.now(UTC)
datasource.last_status = task_result.get("status")
if datasource.last_status == "success":
if datasource.last_status == JobStatus.SUCCESS.value:
effective_candidate = await get_builtin_effective_candidate(db, datasource_source)
checksum, _credential_context = await build_builtin_connectivity_checksum(
datasource_source,
@@ -284,7 +285,7 @@ async def run_collector_task(collector_name: str):
await db.rollback()
datasource = await db.get(DataSource, datasource_id)
datasource.last_run_at = datetime.now(UTC)
datasource.last_status = "cancelled"
datasource.last_status = JobStatus.CANCELLED.value
await db.commit()
logger.warning_event(
"Collector cancelled by operator",
@@ -306,7 +307,7 @@ async def run_collector_task(collector_name: str):
await db.rollback()
datasource = await db.get(DataSource, datasource_id)
datasource.last_run_at = datetime.now(UTC)
datasource.last_status = "failed"
datasource.last_status = JobStatus.FAILED.value
await db.commit()
logger.exception_event(
"Collector failed",
@@ -335,7 +336,7 @@ async def cleanup_stale_running_tasks(max_age_hours: int = 2) -> int:
async with async_session_factory() as db:
result = await db.execute(
select(CollectionTask).where(
CollectionTask.status == "running",
CollectionTask.status == JobStatus.RUNNING.value,
CollectionTask.started_at.is_not(None),
CollectionTask.started_at < cutoff,
)
@@ -343,7 +344,7 @@ async def cleanup_stale_running_tasks(max_age_hours: int = 2) -> int:
stale_tasks = result.scalars().all()
for task in stale_tasks:
task.status = "failed"
task.status = JobStatus.FAILED.value
task.phase = "failed"
task.completed_at = datetime.now(UTC)
existing_error = (task.error_message or "").strip()

View File

@@ -6,6 +6,7 @@ from typing import Any
from sqlalchemy import func, select
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.enums import BGPStatus
from app.models.alert import Alert, AlertSeverity, AlertStatus
from app.models.bgp_anomaly import BGPAnomaly
from app.models.bgp_incident import BGPIncident
@@ -49,11 +50,11 @@ async def build_situational_alert_brief_request(
total_incidents_result = await db.execute(select(func.count(BGPIncident.id)))
active_incidents_result = await db.execute(
select(func.count(BGPIncident.id)).where(BGPIncident.status == "active")
select(func.count(BGPIncident.id)).where(BGPIncident.status == BGPStatus.ACTIVE.value)
)
bgp_severity_result = await db.execute(
select(BGPIncident.severity, func.count(BGPIncident.id))
.where(BGPIncident.status == "active")
.where(BGPIncident.status == BGPStatus.ACTIVE.value)
.group_by(BGPIncident.severity)
)
bgp_region_counter: Counter[str] = Counter()
@@ -65,11 +66,11 @@ async def build_situational_alert_brief_request(
total_anomalies_result = await db.execute(select(func.count(BGPAnomaly.id)))
active_anomalies_result = await db.execute(
select(func.count(BGPAnomaly.id)).where(BGPAnomaly.status == "active")
select(func.count(BGPAnomaly.id)).where(BGPAnomaly.status == BGPStatus.ACTIVE.value)
)
anomaly_type_result = await db.execute(
select(BGPAnomaly.anomaly_type, func.count(BGPAnomaly.id))
.where(BGPAnomaly.status == "active")
.where(BGPAnomaly.status == BGPStatus.ACTIVE.value)
.group_by(BGPAnomaly.anomaly_type)
.order_by(func.count(BGPAnomaly.id).desc())
.limit(6)

View File

@@ -7,6 +7,7 @@ from pathlib import Path
from typing import Any
from app.core.config import ROOT_DIR
from app.core.enums import UserRole
from app.core.security import redis_client
SYSTEM_TASK_TTL_SECONDS = 24 * 60 * 60
@@ -47,7 +48,7 @@ def normalize_user_role(role: Any) -> str:
def require_super_admin(user_role: Any) -> bool:
return normalize_user_role(user_role) == "super_admin"
return normalize_user_role(user_role) == UserRole.SUPER_ADMIN.value
def build_task_id(prefix: str = "restart") -> str:

View File

@@ -5,6 +5,7 @@ import os
import re
import shutil
import subprocess
import hashlib
from collections import Counter, deque
from dataclasses import dataclass
@@ -12,7 +13,11 @@ from datetime import UTC, datetime
from pathlib import Path
from typing import Any
from app.core.enums import LogLevel
from app.core.security import redis_client
from app.models.system_log import AuditLog, ObservabilityEvent, ObservabilityEventGroup, SystemLog
from sqlalchemy import select
from sqlalchemy.ext.asyncio import AsyncSession
DEFAULT_LOG_LINE_LIMIT = 200
MAX_LOG_LINE_LIMIT = 1000
@@ -20,11 +25,11 @@ BUFFER_LOG_LIMIT = 1000
BUFFER_LOG_TTL_SECONDS = 7 * 24 * 60 * 60
LOG_BUFFER_KEY_PREFIX = "planet:system_logs"
LOG_LEVEL_ERROR = "error"
LOG_LEVEL_WARNING = "warning"
LOG_LEVEL_INFO = "info"
LOG_LEVEL_DEBUG = "debug"
LOG_LEVEL_ALL = "all"
LOG_LEVEL_ERROR = LogLevel.ERROR.value
LOG_LEVEL_WARNING = LogLevel.WARNING.value
LOG_LEVEL_INFO = LogLevel.INFO.value
LOG_LEVEL_DEBUG = LogLevel.DEBUG.value
LOG_LEVEL_ALL = LogLevel.ALL.value
SUPPORTED_LOG_LEVELS = {
LOG_LEVEL_ALL,
@@ -99,6 +104,16 @@ class StructuredLogEntry:
search_text: str
@dataclass(frozen=True)
class LogEvent:
source_id: str
cursor: str
timestamp: datetime | None
level: str | None
line: str
search_text: str
@dataclass
class DailyLogMarker:
date_token: str
@@ -106,6 +121,10 @@ class DailyLogMarker:
dominant_level: str
def _normalize_search_query(search: str | None) -> str:
return (search or "").strip().lower()
def _planet_state_dir() -> Path:
configured = os.getenv("PLANET_STATE_DIR")
if configured:
@@ -135,7 +154,7 @@ LOG_SOURCES: dict[str, LogSource] = {
name="前端开发服务",
kind="file",
location=_state_log_path("frontend.log"),
description="控制台与 Earth 前端开发服务输出。",
description="控制台与智能星球前端开发服务输出。",
category="service",
fallback_locations=("/tmp/planet_frontend.log",),
),
@@ -150,13 +169,22 @@ LOG_SOURCES: dict[str, LogSource] = {
),
"earth-client": LogSource(
source_id="earth-client",
name="Earth 浏览器端",
name="智能星球浏览器端",
kind="buffer",
location="redis://planet:system_logs:earth-client",
description="Earth 浏览器端上报的运行时错误与关键业务日志。",
description="智能星球浏览器端上报的运行时错误与关键业务日志。",
category="client",
buffer_key=f"{LOG_BUFFER_KEY_PREFIX}:earth-client",
),
"admin-client": LogSource(
source_id="admin-client",
name="控制台浏览器端",
kind="buffer",
location="redis://planet:system_logs:admin-client",
description="控制台浏览器端上报的运行时错误。",
category="client",
buffer_key=f"{LOG_BUFFER_KEY_PREFIX}:admin-client",
),
}
@@ -365,6 +393,59 @@ def build_buffer_entry(payload: dict[str, Any]) -> StructuredLogEntry:
)
def compact_log_context(context: dict | None) -> str:
if not context:
return ""
allowed = {
key: value
for key, value in (context or {}).items()
if key
in {
"status",
"duration_ms",
"provider",
"model",
"result_provider",
"result_model",
"collector_name",
"datasource_id",
"task_id",
"snapshot_id",
"raw_count",
"transformed_count",
"saved_count",
"created",
"updated",
"unchanged",
"deleted",
"result_count",
"status_code",
"error_type",
"error",
"route",
"module",
}
}
if not allowed:
return ""
return json.dumps(allowed, ensure_ascii=False, sort_keys=True)
def context_search_aliases(context: dict | None) -> str:
if not context:
return ""
aliases: list[str] = []
for key, value in sorted((context or {}).items()):
if value is None or isinstance(value, (dict, list, tuple, set)):
continue
normalized_key = str(key).strip()
normalized_value = str(value).strip()
if not normalized_key or not normalized_value:
continue
aliases.append(f"{normalized_key}={normalized_value}")
return " ".join(aliases)
def read_file_entries(source: LogSource, scan_limit: int) -> list[StructuredLogEntry]:
path = resolve_file_log_path(source)
if not path.exists():
@@ -437,6 +518,400 @@ def read_source_entries(source: LogSource, scan_limit: int) -> list[StructuredLo
return []
def _database_event_from_system_record(record: SystemLog) -> LogEvent:
record_level = normalize_log_level(record.level)
line = " ".join(
part
for part in [
record.occurred_at.isoformat() if record.occurred_at else "",
record_level.upper(),
record.source,
record.category or "",
record.event or "",
f"request_id={record.request_id}" if record.request_id else "",
record.message,
compact_log_context(record.context),
]
if part
)
search_text = " ".join(
[
line,
f"id={record.id}",
f"user_id={record.user_id}" if record.user_id else "",
context_search_aliases(record.context),
json.dumps(record.context or {}, ensure_ascii=False, sort_keys=True),
]
).lower()
return LogEvent(
source_id="system-db",
cursor=f"system-db:{record.id}",
timestamp=record.occurred_at,
level=None if record_level == LOG_LEVEL_ALL else record_level,
line=line,
search_text=search_text,
)
def _database_event_from_audit_record(record: AuditLog) -> LogEvent:
line = " ".join(
part
for part in [
record.occurred_at.isoformat() if record.occurred_at else "",
"INFO",
record.action,
record.target_type or "",
record.target_id or "",
record.result or "",
f"request_id={record.request_id}" if record.request_id else "",
]
if part
)
search_text = " ".join(
[
line,
f"id={record.id}",
f"actor_id={record.actor_id}" if record.actor_id else "",
record.actor_name or "",
context_search_aliases(record.details),
json.dumps(record.details or {}, ensure_ascii=False, sort_keys=True),
]
).lower()
return LogEvent(
source_id="audit-db",
cursor=f"audit-db:{record.id}",
timestamp=record.occurred_at,
level=LOG_LEVEL_INFO,
line=line,
search_text=search_text,
)
async def read_database_log_events(
source_id: str,
*,
scan_limit: int,
level: str = LOG_LEVEL_ALL,
levels: str | None = None,
start_date: str | None = None,
end_date: str | None = None,
search: str | None = None,
db: AsyncSession,
) -> list[LogEvent] | None:
selected_levels = normalize_log_levels(level, levels)
search_query = (search or "").strip()
if source_id == "system-db":
query = select(SystemLog).order_by(SystemLog.occurred_at.desc().nullslast(), SystemLog.id.desc()).limit(scan_limit)
result = await db.execute(query)
events = [_database_event_from_system_record(record) for record in result.scalars().all()]
elif source_id == "audit-db":
query = select(AuditLog).order_by(AuditLog.occurred_at.desc().nullslast(), AuditLog.id.desc()).limit(scan_limit)
result = await db.execute(query)
events = [_database_event_from_audit_record(record) for record in result.scalars().all()]
else:
return None
events = list(reversed(events))
return [
event
for event in events
if event_matches_levels(event, selected_levels)
and event_matches_search(event, search_query)
and event_matches_date_range(event, start_date, end_date)
]
async def read_database_log_snapshot(
source_id: str,
*,
limit: int,
level: str,
levels: str | None,
start_date: str | None,
end_date: str | None,
search: str | None,
db: AsyncSession,
) -> dict[str, Any] | None:
events = await read_database_log_events(
source_id,
scan_limit=limit * 5,
level=level,
levels=levels,
start_date=start_date,
end_date=end_date,
search=search,
db=db,
)
if events is None:
return None
visible_events = events[-limit:]
selected_levels = normalize_log_levels(level, levels)
return {
"source_id": source_id,
"name": "系统事件" if source_id == "system-db" else "审计事件",
"kind": "database",
"location": "table://system_logs" if source_id == "system-db" else "table://audit_logs",
"description": "数据库持久化日志",
"category": "database" if source_id == "system-db" else "audit",
"status": "ok" if visible_events else "empty",
"level": level,
"selected_levels": list(selected_levels),
"search_query": search or "",
"available_levels": ["all", "error", "warning", "info", "debug"],
"daily_markers": build_daily_log_markers_from_events(events),
"line_limit": limit,
"line_count": len(visible_events),
"lines": [event.line for event in visible_events],
}
def _observability_group_matches(
group: ObservabilityEventGroup,
*,
selected_levels: tuple[str, ...],
start_date: str | None,
end_date: str | None,
search: str | None,
) -> bool:
if selected_levels and group.last_level not in selected_levels:
return False
if start_date or end_date:
if group.last_seen_at is None:
return False
date_token = group.last_seen_at.astimezone(UTC).date().isoformat()
if start_date and date_token < start_date:
return False
if end_date and date_token > end_date:
return False
query = _normalize_search_query(search)
if not query:
return True
haystack = " ".join(
[
group.fingerprint or "",
group.source or "",
group.service or "",
group.module or "",
group.category or "",
group.event or "",
group.last_level or "",
group.sample_message or "",
group.sample_detail or "",
json.dumps(group.affected_sources or [], ensure_ascii=False, sort_keys=True),
]
).lower()
return query in haystack
def _serialize_observability_group(group: ObservabilityEventGroup) -> dict[str, Any]:
return {
"fingerprint": group.fingerprint,
"source": group.source,
"service": group.service,
"module": group.module,
"category": group.category,
"event": group.event,
"level": group.last_level,
"message": group.sample_message,
"detail": group.sample_detail,
"affected_sources": group.affected_sources or [],
"count": group.count or 0,
"first_seen_at": group.first_seen_at.isoformat() if group.first_seen_at else None,
"last_seen_at": group.last_seen_at.isoformat() if group.last_seen_at else None,
}
def _serialize_observability_event(record: ObservabilityEvent) -> dict[str, Any]:
return {
"id": record.id,
"source": record.source,
"service": record.service,
"module": record.module,
"category": record.category,
"event": record.event,
"level": record.level,
"message": record.message,
"fingerprint": record.fingerprint,
"occurred_at": record.occurred_at.isoformat() if record.occurred_at else None,
"request_id": record.request_id,
"trace_id": record.trace_id,
"task_id": record.task_id,
"source_id": record.source_ref_id,
"provider": record.provider,
"user_id": record.user_id,
"context": record.context or {},
"occurrence_count": record.occurrence_count or 1,
}
async def read_observability_groups(
*,
limit: int,
level: str = LOG_LEVEL_ALL,
levels: str | None = None,
start_date: str | None = None,
end_date: str | None = None,
search: str | None = None,
db: AsyncSession,
) -> dict[str, Any]:
selected_levels = normalize_log_levels(level, levels)
scan_limit = max(limit * 5, limit, DEFAULT_LOG_LINE_LIMIT)
result = await db.execute(
select(ObservabilityEventGroup)
.order_by(ObservabilityEventGroup.last_seen_at.desc().nullslast())
.limit(scan_limit)
)
groups = [
group
for group in result.scalars().all()
if _observability_group_matches(
group,
selected_levels=selected_levels,
start_date=start_date,
end_date=end_date,
search=search,
)
][:limit]
return {
"mode": "grouped",
"line_limit": limit,
"line_count": len(groups),
"groups": [_serialize_observability_group(group) for group in groups],
"filters": {
"level": level,
"levels": list(selected_levels),
"start_date": start_date,
"end_date": end_date,
"search": search or "",
},
}
async def read_observability_group_events(
fingerprint: str,
*,
limit: int,
db: AsyncSession,
) -> dict[str, Any] | None:
group = await db.get(ObservabilityEventGroup, fingerprint)
if group is None:
return None
result = await db.execute(
select(ObservabilityEvent)
.where(ObservabilityEvent.fingerprint == fingerprint)
.order_by(ObservabilityEvent.occurred_at.desc().nullslast(), ObservabilityEvent.id.desc())
.limit(limit)
)
events = list(reversed(result.scalars().all()))
return {
"fingerprint": fingerprint,
"group": _serialize_observability_group(group),
"line_limit": limit,
"line_count": len(events),
"events": [_serialize_observability_event(record) for record in events],
}
async def read_observability_raw_events(
*,
limit: int,
level: str = LOG_LEVEL_ALL,
levels: str | None = None,
start_date: str | None = None,
end_date: str | None = None,
search: str | None = None,
db: AsyncSession,
) -> dict[str, Any]:
selected_levels = normalize_log_levels(level, levels)
query = select(ObservabilityEvent).order_by(ObservabilityEvent.occurred_at.desc().nullslast(), ObservabilityEvent.id.desc())
if selected_levels:
query = query.where(ObservabilityEvent.level.in_(selected_levels))
result = await db.execute(query.limit(max(limit * 5, limit)))
records = result.scalars().all()
search_query = _normalize_search_query(search)
visible: list[ObservabilityEvent] = []
for record in records:
if start_date or end_date:
if record.occurred_at is None:
continue
date_token = record.occurred_at.astimezone(UTC).date().isoformat()
if start_date and date_token < start_date:
continue
if end_date and date_token > end_date:
continue
if search_query:
haystack = " ".join(
[
record.source or "",
record.service or "",
record.module or "",
record.category or "",
record.event or "",
record.message or "",
record.fingerprint or "",
record.request_id or "",
record.trace_id or "",
record.task_id or "",
record.source_ref_id or "",
record.provider or "",
json.dumps(record.context or {}, ensure_ascii=False, sort_keys=True),
]
).lower()
if search_query not in haystack:
continue
visible.append(record)
if len(visible) >= limit:
break
visible = list(reversed(visible))
return {
"mode": "raw",
"line_limit": limit,
"line_count": len(visible),
"events": [_serialize_observability_event(record) for record in visible],
"lines": [
" ".join(
part
for part in [
record.occurred_at.isoformat() if record.occurred_at else "",
record.level.upper(),
record.source,
record.category or "",
record.event or "",
f"fingerprint={record.fingerprint}",
record.message,
]
if part
)
for record in visible
],
}
def _stable_hash(value: str) -> str:
return hashlib.sha1(value.encode("utf-8", errors="replace")).hexdigest()[:16]
def build_log_events(source_id: str, entries: list[StructuredLogEntry]) -> list[LogEvent]:
events: list[LogEvent] = []
seen: dict[str, int] = {}
for entry in entries:
stable_value = entry.raw_line or entry.display_line
digest = _stable_hash(stable_value)
occurrence = seen.get(digest, 0) + 1
seen[digest] = occurrence
events.append(
LogEvent(
source_id=source_id,
cursor=f"{source_id}:{digest}:{occurrence}",
timestamp=entry.timestamp,
level=entry.level,
line=entry.display_line,
search_text=entry.search_text,
)
)
return events
def matches_levels(entry: StructuredLogEntry, selected_levels: tuple[str, ...]) -> bool:
if not selected_levels:
return True
@@ -469,6 +944,34 @@ def matches_search(entry: StructuredLogEntry, search: str | None) -> bool:
return query in entry.search_text
def event_matches_levels(event: LogEvent, selected_levels: tuple[str, ...]) -> bool:
if not selected_levels:
return True
return event.level in selected_levels
def event_matches_date_range(event: LogEvent, start_date: str | None, end_date: str | None) -> bool:
if not start_date and not end_date:
return True
if event.timestamp is None:
return False
date_token = event.timestamp.astimezone(UTC).date().isoformat()
if start_date and date_token < start_date:
return False
if end_date and date_token > end_date:
return False
return True
def event_matches_search(event: LogEvent, search: str | None) -> bool:
if search is None:
return True
query = search.strip().lower()
if not query:
return True
return query in event.search_text
def build_daily_log_markers(entries: list[StructuredLogEntry]) -> list[dict[str, Any]]:
grouped: dict[str, list[StructuredLogEntry]] = {}
for entry in entries:
@@ -503,6 +1006,69 @@ def build_daily_log_markers(entries: list[StructuredLogEntry]) -> list[dict[str,
return [marker.__dict__ for marker in markers]
def build_daily_log_markers_from_events(events: list[LogEvent]) -> list[dict[str, Any]]:
grouped: dict[str, list[LogEvent]] = {}
for event in events:
if event.timestamp is None:
continue
date_token = event.timestamp.astimezone(UTC).date().isoformat()
grouped.setdefault(date_token, []).append(event)
markers: list[DailyLogMarker] = []
for date_token, group in sorted(grouped.items()):
level_counts = Counter(
event.level
for event in group
if event.level in SUPPORTED_LOG_LEVELS and event.level != LOG_LEVEL_ALL
)
dominant_level = LOG_LEVEL_INFO
if level_counts:
dominant_level = sorted(
level_counts.items(),
key=lambda item: (
-item[1],
("error", "warning", "info", "debug").index(item[0]),
),
)[0][0]
markers.append(
DailyLogMarker(
date_token=date_token,
total=len(group),
dominant_level=dominant_level,
)
)
return [marker.__dict__ for marker in markers]
def read_log_events(
source_id: str,
scan_limit: int,
*,
level: str = LOG_LEVEL_ALL,
levels: str | None = None,
start_date: str | None = None,
end_date: str | None = None,
search: str | None = None,
) -> list[LogEvent] | None:
source = LOG_SOURCES.get(source_id)
if source is None:
return None
selected_levels = normalize_log_levels(level, levels)
search_query = (search or "").strip()
events = build_log_events(source_id, read_source_entries(source, scan_limit))
marker_events = [
event
for event in events
if event_matches_levels(event, selected_levels) and event_matches_search(event, search_query)
]
return [
event
for event in marker_events
if event_matches_date_range(event, start_date, end_date)
]
def read_log_snapshot(
source_id: str,
limit: int,
@@ -520,18 +1086,18 @@ def read_log_snapshot(
selected_levels = normalize_log_levels(level, levels)
search_query = (search or "").strip()
scan_limit = max(min(MAX_LOG_LINE_LIMIT * 5, 5000), limit * 5, BUFFER_LOG_LIMIT if source.kind == "buffer" else 1000)
all_entries = read_source_entries(source, scan_limit)
marker_entries = [
entry
for entry in all_entries
if matches_levels(entry, selected_levels) and matches_search(entry, search_query)
all_events = build_log_events(source_id, read_source_entries(source, scan_limit))
marker_events = [
event
for event in all_events
if event_matches_levels(event, selected_levels) and event_matches_search(event, search_query)
]
filtered_entries = [
entry
for entry in marker_entries
if matches_date_range(entry, start_date, end_date)
filtered_events = [
event
for event in marker_events
if event_matches_date_range(event, start_date, end_date)
]
visible_entries = filtered_entries[-limit:]
visible_events = filtered_events[-limit:]
compatibility_level = selected_levels[0] if len(selected_levels) == 1 else LOG_LEVEL_ALL
return {
@@ -552,8 +1118,8 @@ def read_log_snapshot(
LOG_LEVEL_INFO,
LOG_LEVEL_DEBUG,
],
"daily_markers": build_daily_log_markers(marker_entries),
"daily_markers": build_daily_log_markers_from_events(marker_events),
"line_limit": limit,
"line_count": len(visible_entries),
"lines": [entry.display_line for entry in visible_entries],
"line_count": len(visible_events),
"lines": [event.line for event in visible_events],
}

View File

@@ -9,7 +9,12 @@ from sqlalchemy import select
from sqlalchemy import Float
from sqlalchemy.ext.asyncio import AsyncSession
from app.models.vessel import AISConflictRecord, AISRawObservation, AISSourceHealth
from app.models.vessel import (
AISConflictRecord,
AISRawObservation,
AISSourceHealth,
VesselCurrentState,
)
from app.services.vessel_aggregation_strategy import (
DEFAULT_STRATEGY,
load_strategy,
@@ -44,6 +49,17 @@ CONFLICT_FIELDS = (
"width",
"draught",
)
CURRENT_STATE_STATIC_FIELDS = (
"name",
"callsign",
"vessel_type",
"vessel_type_name",
"flag",
"length",
"width",
"draught",
"imo",
)
def _json_default(value: Any) -> Any:
@@ -489,9 +505,115 @@ async def record_vessel_ais_observation(
quality_flags=quality_flags or [],
)
db.add(observation)
await upsert_vessel_current_state(
db,
source=source,
normalized_payload=normalized_json,
observed_at=observed_at,
quality_flags=quality_flags or [],
)
return observation
async def upsert_vessel_current_state(
db: AsyncSession,
*,
source: str,
normalized_payload: dict[str, Any],
observed_at: datetime,
quality_flags: list[str] | None = None,
) -> VesselCurrentState | None:
"""Keep one latest renderable row per MMSI while preserving useful static fields."""
if not _has_valid_position(normalized_payload):
return None
mmsi = int(normalized_payload["mmsi"])
current = await db.get(VesselCurrentState, mmsi)
if current is not None and current.observed_at is not None:
current_observed_at = _coerce_datetime(current.observed_at)
if current_observed_at is not None and observed_at < current_observed_at:
return current
if current is None:
current = VesselCurrentState(mmsi=mmsi)
db.add(current)
current.lat = float(normalized_payload["lat"])
current.lon = float(normalized_payload["lon"])
current.source = source
current.observed_at = observed_at
current.updated_at = datetime.now(UTC)
updated_fields: set[str] = {"lat", "lon"}
for field in DYNAMIC_FIELDS:
if field in {"lat", "lon"}:
continue
value = _payload_value(normalized_payload, field)
if value is not None:
setattr(current, field, value)
updated_fields.add(field)
field_sources = dict(current.field_sources or {})
for field in CURRENT_STATE_STATIC_FIELDS:
value = _payload_value(normalized_payload, field)
if value is None:
continue
existing_source = field_sources.get(field)
existing_value = getattr(current, field, None)
if (
existing_value in (None, "")
or _strategy_source_rank(source, DEFAULT_STRATEGY)
>= _strategy_source_rank(str(existing_source or ""), DEFAULT_STRATEGY)
):
setattr(current, field, value)
updated_fields.add(field)
current.vessel_type_name = current.vessel_type_name or normalize_vessel_type_name(
current.vessel_type
)
selected_reasons = dict(current.selected_reasons or {})
for field in updated_fields:
field_sources[field] = source
selected_reasons[field] = (
"newest_observation" if field in DYNAMIC_FIELDS else "source_priority"
)
current.field_sources = field_sources
current.selected_reasons = selected_reasons
current.source_summary = {
**dict(current.source_summary or {}),
source: {
"latest_observed_at": observed_at.isoformat(),
},
}
current.quality_flags = sorted(set((current.quality_flags or []) + (quality_flags or [])))
return current
async def get_current_vessels_snapshot(
db: AsyncSession,
*,
bbox: tuple[float, float, float, float],
limit: int = 1000,
observed_since: datetime,
) -> list[dict[str, Any]]:
"""Read the bounded latest-state table used by Earth rendering."""
safe_limit = min(max(int(limit or 1000), 1), MAX_SNAPSHOT_LIMIT)
lon_min, lat_min, lon_max, lat_max = bbox
stmt = (
select(VesselCurrentState)
.where(VesselCurrentState.observed_at >= observed_since)
.where(VesselCurrentState.lon >= lon_min)
.where(VesselCurrentState.lon <= lon_max)
.where(VesselCurrentState.lat >= lat_min)
.where(VesselCurrentState.lat <= lat_max)
.order_by(VesselCurrentState.observed_at.desc(), VesselCurrentState.mmsi.asc())
.limit(safe_limit)
)
result = await db.execute(stmt)
if not hasattr(result, "scalars"):
return []
return [item.to_dict() for item in result.scalars().all()]
async def aggregate_vessel_observations(
db: AsyncSession,
observations: Iterable[AISRawObservation],

View File

@@ -1,4 +1,5 @@
[pytest]
pythonpath = ..
asyncio_mode = auto
testpaths = tests
python_files = test_*.py

View File

@@ -655,6 +655,8 @@ async def test_ingest_earth_client_log_accepts_public_events():
"message": "登陆点加载失败: 登陆点接口返回 HTTP 500",
"category": "startup-load",
"module": "layer-startup",
"fingerprint": "client-test",
"occurrence_count": 3,
},
)
assert response.status_code == 200
@@ -668,10 +670,98 @@ async def test_ingest_earth_client_log_accepts_public_events():
assert persisted_kwargs["event"] == "earth.client.runtime_log"
assert persisted_kwargs["category"] == "startup-load"
assert persisted_kwargs["level"] == "error"
assert persisted_kwargs["fingerprint"] == "client-test"
assert persisted_kwargs["occurrence_count"] == 3
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_ingest_admin_client_log_accepts_public_events():
transport = ASGITransport(app=app)
try:
with patch("app.api.v1.system_control.append_buffer_log") as mock_append_buffer_log:
with patch("app.api.v1.system_control.record_system_log", new_callable=AsyncMock) as mock_record_system_log:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.post(
"/api/v1/system/logs/admin-client",
json={
"level": "error",
"message": "控制台发生未处理 Promise 错误",
"category": "unhandledrejection",
"module": "admin",
"url": "http://test/logs",
"detail": "stack preview",
},
)
assert response.status_code == 200
data = response.json()
assert data["accepted"] is True
assert data["source_id"] == "admin-client"
mock_append_buffer_log.assert_called_once()
mock_record_system_log.assert_awaited_once()
persisted_kwargs = mock_record_system_log.await_args.kwargs
assert persisted_kwargs["source"] == "admin-client"
assert persisted_kwargs["event"] == "admin.client.runtime_log"
assert persisted_kwargs["category"] == "unhandledrejection"
assert persisted_kwargs["context"]["url"] == "http://test/logs"
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_ingest_service_log_requires_configured_token(monkeypatch):
monkeypatch.setattr(settings, "OBSERVABILITY_INGEST_TOKEN", "")
transport = ASGITransport(app=app)
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.post(
"/api/v1/system/logs/service",
json={"message": "AI provider failed"},
headers={"X-Planet-Observability-Token": "secret"},
)
assert response.status_code == 503
@pytest.mark.asyncio
async def test_ingest_service_log_accepts_internal_token(monkeypatch):
monkeypatch.setattr(settings, "OBSERVABILITY_INGEST_TOKEN", "service-secret")
transport = ASGITransport(app=app)
with patch("app.api.v1.system_control.record_system_log", new_callable=AsyncMock) as mock_record_system_log:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.post(
"/api/v1/system/logs/service",
json={
"source": "ai-provider",
"service": "ai-provider",
"module": "provider",
"category": "connectivity",
"event": "ai.provider.test.failed",
"level": "error",
"message": "Provider connectivity failed",
"fingerprint": "ai-provider-test",
"occurrence_count": 4,
"provider": "minimax",
"trace_id": "trace-123",
"context": {"status_code": 502},
},
headers={"Authorization": "Bearer service-secret"},
)
assert response.status_code == 200
data = response.json()
assert data["accepted"] is True
assert data["source_id"] == "ai-provider"
mock_record_system_log.assert_awaited_once()
persisted_kwargs = mock_record_system_log.await_args.kwargs
assert persisted_kwargs["event"] == "ai.provider.test.failed"
assert persisted_kwargs["fingerprint"] == "ai-provider-test"
assert persisted_kwargs["occurrence_count"] == 4
assert persisted_kwargs["context"]["provider"] == "minimax"
assert persisted_kwargs["context"]["trace_id"] == "trace-123"
assert persisted_kwargs["context"]["status_code"] == 502
@pytest.mark.asyncio
async def test_earth_layer_cache_status_requires_super_admin(auth_headers, monkeypatch):
def override_get_current_user():

View File

@@ -74,6 +74,6 @@ def test_interactable_cache_invalidation_clears_layer_and_all(monkeypatch):
assert deleted == 2
assert patterns == [
"earth:layer:v1:interactables:layer:places*",
"earth:layer:v1:interactables:layer:all*",
"earth:layer:v1:interactables:interactable_layer:places*",
"earth:layer:v1:interactables:interactable_layer:all*",
]

View File

@@ -1,21 +1,30 @@
from datetime import UTC, datetime
from datetime import UTC, datetime, timedelta
from types import SimpleNamespace
import pytest
from app.services.earth_news import (
NewsFeedEndpoint,
NewsFeedSource,
NewsTargetLocation,
ParsedNewsItem,
apply_news_classification,
default_earth_news_sources_payload,
normalize_earth_news_sources_payload,
_fetch_source,
_diversify_news_items_for_locale,
_enrich_items_with_target_locations,
_extract_target_location_from_text,
_parse_feed_entries,
_rank_and_trim_items,
_serialize_item,
get_earth_news_payload,
test_news_source_config as run_news_source_config_test,
)
from app.services.earth_news_queue import NewsTargetLocationMessage
from app.services.earth_news_worker import process_target_location_message
from app.services.collectors.media_news_archive import MediaNewsArchiveCollector
from app.services.earth_news_store import _diversify_parsed_news_items_by_region
def test_serialize_item_includes_region_anchor_for_cruise():
@@ -44,6 +53,188 @@ def test_serialize_item_includes_region_anchor_for_cruise():
assert payload["published_at"] == "2026-04-23T02:30:00Z"
def test_serialize_item_includes_breaking_fields():
item = ParsedNewsItem(
id="breaking:test",
title="Major market halt",
summary="Trading halt after flash crash",
url="https://example.com/breaking",
source="Example Source",
feed_name="Example Feed",
feed_region="global",
homepage_url="https://example.com",
published_at=datetime(2026, 5, 15, 2, 0, tzinfo=UTC),
breaking_level="critical",
breaking_scope="global",
breaking_reasons=["重大金融市场异常"],
breaking_source="rules",
breaking_confidence=0.72,
breaking_expires_at=datetime(2026, 5, 16, 2, 0, tzinfo=UTC),
)
payload = _serialize_item(item, active_region="europe")
assert payload["breaking_level"] == "critical"
assert payload["breaking_scope"] == "global"
assert payload["breaking_reasons"] == ["重大金融市场异常"]
assert payload["breaking_source"] == "rules"
assert payload["breaking_confidence"] == 0.72
assert payload["breaking_expires_at"] == "2026-05-16T02:00:00Z"
def test_rank_and_trim_items_prioritizes_active_breaking():
older_breaking = ParsedNewsItem(
id="global:critical",
title="Nuclear accident reported",
summary="A nuclear accident has been reported.",
url="https://example.com/critical",
source="Global Source",
feed_name="Global Feed",
feed_region="global",
homepage_url="https://example.com",
published_at=datetime.now(UTC) - timedelta(hours=2),
breaking_level="critical",
breaking_scope="global",
breaking_expires_at=datetime.now(UTC) + timedelta(hours=6),
)
newer_regular = ParsedNewsItem(
id="europe:regular",
title="Regular Europe story",
summary="A newer regular story.",
url="https://example.com/regular",
source="Europe Source",
feed_name="Europe Feed",
feed_region="europe",
homepage_url="https://example.com",
published_at=datetime.now(UTC),
)
expired_breaking = ParsedNewsItem(
id="europe:expired",
title="Expired breaking",
summary="Expired breaking story.",
url="https://example.com/expired",
source="Europe Source",
feed_name="Europe Feed",
feed_region="europe",
homepage_url="https://example.com",
published_at=datetime.now(UTC) + timedelta(minutes=1),
breaking_level="critical",
breaking_scope="regional",
breaking_expires_at=datetime.now(UTC) - timedelta(minutes=1),
)
ranked = _rank_and_trim_items(
[newer_regular, expired_breaking, older_breaking],
active_region="europe",
limit=3,
)
assert [item.id for item in ranked] == ["global:critical", "europe:expired", "europe:regular"]
def test_diversify_news_items_prefers_display_ready_content_across_sources():
published_at = datetime(2026, 6, 11, 3, 0, tzinfo=UTC)
def make_item(source_id: str, suffix: str, *, zh_ready: bool) -> ParsedNewsItem:
return ParsedNewsItem(
id=f"{source_id}:{suffix}",
title=f"{source_id} title {suffix}",
summary=f"{source_id} summary {suffix}",
url=f"https://example.com/{source_id}/{suffix}",
source=source_id,
feed_name=source_id,
feed_region="global",
homepage_url="https://example.com",
published_at=published_at,
content_language="en",
localizations={
"zh-CN": {
"title": f"{source_id} 中文标题 {suffix}",
"summary": f"{source_id} 中文摘要 {suffix}",
}
} if zh_ready else {},
)
items = [
make_item("source-a", "1", zh_ready=False),
make_item("source-a", "2", zh_ready=False),
make_item("source-a", "3", zh_ready=False),
make_item("source-b", "1", zh_ready=True),
make_item("source-c", "1", zh_ready=True),
]
result = _diversify_news_items_for_locale(
items,
active_region="global",
limit=3,
locale="zh-CN",
)
assert [item.id.split(":", 1)[0] for item in result] == ["source-b", "source-c", "source-a"]
def test_cruise_news_diversity_keeps_regions_from_being_starved():
published_at = datetime(2026, 6, 26, 8, 0, tzinfo=UTC)
def make_item(region: str, index: int) -> ParsedNewsItem:
return ParsedNewsItem(
id=f"{region}:{index}",
title=f"{region} story {index}",
summary=f"{region} summary {index}",
url=f"https://example.com/{region}/{index}",
source=region,
feed_name=region,
feed_region=region,
homepage_url="https://example.com",
published_at=published_at - timedelta(minutes=index),
)
items = [
*[make_item("asia-pacific", index) for index in range(40)],
make_item("europe", 1),
make_item("middle-east-africa", 1),
make_item("americas", 1),
make_item("global", 1),
]
result = _diversify_parsed_news_items_by_region(items, limit=8)
regions = [item.feed_region for item in result]
assert "europe" in regions
assert "middle-east-africa" in regions
assert "americas" in regions
assert regions.count("asia-pacific") < len(regions)
def test_global_news_diversity_uses_same_region_balance():
published_at = datetime(2026, 6, 26, 8, 0, tzinfo=UTC)
def make_item(region: str, index: int) -> ParsedNewsItem:
return ParsedNewsItem(
id=f"{region}:global:{index}",
title=f"{region} story {index}",
summary=f"{region} summary {index}",
url=f"https://example.com/{region}/global/{index}",
source=region,
feed_name=region,
feed_region=region,
homepage_url="https://example.com",
published_at=published_at - timedelta(minutes=index),
)
items = [
*[make_item("asia-pacific", index) for index in range(24)],
*[make_item("europe", index) for index in range(2)],
*[make_item("middle-east-africa", index) for index in range(2)],
*[make_item("americas", index) for index in range(2)],
]
result = _diversify_parsed_news_items_by_region(items, limit=6)
regions = {item.feed_region for item in result}
assert {"europe", "middle-east-africa", "americas"}.issubset(regions)
def test_serialize_item_falls_back_to_global_anchor():
item = ParsedNewsItem(
id="custom:test",
@@ -164,6 +355,486 @@ def test_parse_aggregated_rss_splits_publisher_from_title():
assert items[0].source == "Reuters"
def test_parse_chinese_rss_marks_source_language_and_keeps_zh_localization():
source = NewsFeedSource(
id="36kr",
name="36氪",
region="asia-pacific",
feed_url="https://36kr.com/feed",
homepage_url="https://www.36kr.com/",
source_tags=("china", "business_news"),
default_category="business",
)
xml = """
<rss>
<channel>
<item>
<title>中国电商平台发布季度增长数据</title>
<description>平台表示,跨境电商订单量同比增长。</description>
<link>https://36kr.com/p/example</link>
</item>
</channel>
</rss>
"""
items = _parse_feed_entries(xml, source)
payload_zh = _serialize_item(items[0], active_region="global", locale="zh-CN")
payload_en = _serialize_item(items[0], active_region="global", locale="en-US")
assert items[0].content_language == "zh-CN"
assert items[0].localizations["zh-CN"]["title"] == "中国电商平台发布季度增长数据"
assert payload_zh["display_title"] == "中国电商平台发布季度增长数据"
assert payload_en["display_title"] == ""
items[0].localizations["en-US"] = {
"title": "Chinese e-commerce platform reports quarterly growth",
"summary": "The platform said cross-border orders rose year over year.",
}
payload_en_ready = _serialize_item(items[0], active_region="global", locale="en-US")
assert payload_en_ready["display_title"] == "Chinese e-commerce platform reports quarterly growth"
def test_default_news_sources_include_business_and_ecommerce_sources():
payload = default_earth_news_sources_payload()
sources_by_id = {source["id"]: source for source in payload["sources"]}
source_ids = {source["id"] for source in payload["sources"]}
category_keys = {category["key"] for category in payload["categories"]}
tag_keys = {tag["key"] for tag in payload["source_tags"]}
assert "cnbc-business" in source_ids
assert "36kr" in source_ids
assert "techcrunch" in source_ids
assert "retaildive" in source_ids
assert "prnewswire-retail" in source_ids
assert "google-news" in source_ids
assert "global-scan" not in source_ids
assert "google-americas" not in source_ids
assert "google-europe" not in source_ids
assert "google-mea" not in source_ids
assert "google-apac" not in source_ids
assert "businesswire-ecommerce" in source_ids
assert "us-census-ecommerce" in source_ids
assert "mofcom-data" in source_ids
assert "stats-china-online-retail" in source_ids
assert "ebrun" in source_ids
assert sources_by_id["36kr"]["source_type"] == "rss"
assert sources_by_id["36kr"]["homepage_url"] == "https://www.36kr.com/"
assert sources_by_id["36kr"]["feed_directory_url"] == "https://www.36kr.com/rss-center"
kr_feeds = {feed["id"]: feed for feed in sources_by_id["36kr"]["feeds"]}
assert set(kr_feeds) == {"feed", "article", "newsflash", "moment"}
assert kr_feeds["feed"]["url"] == "https://36kr.com/feed"
assert kr_feeds["article"]["url"] == "https://36kr.com/feed-article"
assert kr_feeds["newsflash"]["url"] == "https://36kr.com/feed-newsflash"
assert kr_feeds["moment"]["url"] == "https://36kr.com/feed-moment"
assert all(feed["enabled"] is True for feed in kr_feeds.values())
assert all(feed["default_category"] == "business" for feed in kr_feeds.values())
assert "https://36kr.com/feed-article" in sources_by_id["36kr"]["feed_urls"]
assert "https://36kr.com/feed-newsflash" in sources_by_id["36kr"]["feed_urls"]
assert "https://36kr.com/feed-moment" in sources_by_id["36kr"]["feed_urls"]
assert sources_by_id["ebrun"]["source_type"] == "rss"
assert sources_by_id["ebrun"]["homepage_url"] == "https://www.ebrun.com/"
assert sources_by_id["ebrun"]["feed_directory_url"] == "https://www.ebrun.com/rss/"
ebrun_feeds = {feed["id"]: feed for feed in sources_by_id["ebrun"]["feeds"]}
assert {"b2c", "b2b", "retail", "o2o", "service", "data", "policy"}.issubset(ebrun_feeds)
assert all(feed["enabled"] is True for feed in ebrun_feeds.values())
assert all(feed["default_category"] == "ecommerce" for feed in ebrun_feeds.values())
assert "https://www.ebrun.com/rss/news_b2c.xml" in sources_by_id["ebrun"]["feed_urls"]
assert "https://www.ebrun.com/rss/news_retail.xml" in sources_by_id["ebrun"]["feed_urls"]
assert sources_by_id["businesswire-ecommerce"]["source_type"] == "reference"
assert sources_by_id["businesswire-ecommerce"]["enabled"] is False
assert sources_by_id["google-news"]["source_type"] == "aggregated"
assert sources_by_id["google-news"]["homepage_url"] == "https://news.google.com/"
assert sources_by_id["google-news"]["feed_directory_url"] == "https://news.google.com/rss"
google_feeds = {feed["id"]: feed for feed in sources_by_id["google-news"]["feeds"]}
assert set(google_feeds) == {"world", "americas", "europe", "middle-east-africa", "asia-pacific"}
assert all(feed["type"] == "aggregated" for feed in google_feeds.values())
assert all(feed["enabled"] is True for feed in google_feeds.values())
assert google_feeds["world"]["region"] == "global"
assert google_feeds["europe"]["region"] == "europe"
assert sources_by_id["stats-china-online-retail"]["source_type"] == "rss"
assert sources_by_id["stats-china-online-retail"]["enabled"] is True
assert "https://www.stats.gov.cn/sj/zxfb/rss.xml" in sources_by_id["stats-china-online-retail"]["feed_urls"]
assert {"business", "ecommerce", "finance"}.issubset(category_keys)
assert {"official_data", "business_news", "ecommerce", "press_release", "finance", "logistics"}.issubset(tag_keys)
def test_default_enabled_fetchable_sources_have_explicit_types_and_urls():
payload = default_earth_news_sources_payload()
for source in payload["sources"]:
source_type = source["source_type"]
assert source_type in {"rss", "atom", "aggregated", "reference"}
if source_type == "reference":
assert source["enabled"] is False
assert source["feeds"] == []
continue
if source["enabled"]:
assert source["feed_url"]
assert source["feed_urls"]
assert source["feeds"]
assert any(feed["enabled"] for feed in source["feeds"])
for feed in source["feeds"]:
assert feed["url"] != source["homepage_url"]
assert feed["url"] != source.get("feed_directory_url", "")
def test_legacy_news_source_urls_migrate_to_feed_children():
payload = normalize_earth_news_sources_payload(
{
"sources": [
{
"id": "legacy-source",
"name": "Legacy Source",
"region": "global",
"source_type": "rss",
"feed_urls": ["https://example.com/a.xml", "https://example.com/b.xml"],
"default_category": "business",
}
]
}
)
source = payload["sources"][0]
assert source["feed_urls"] == ["https://example.com/a.xml", "https://example.com/b.xml"]
assert [feed["url"] for feed in source["feeds"]] == ["https://example.com/a.xml", "https://example.com/b.xml"]
assert [feed["id"] for feed in source["feeds"]] == ["feed-1", "feed-2"]
assert all(feed["default_category"] == "business" for feed in source["feeds"])
def test_builtin_news_source_legacy_directory_url_is_repaired():
payload = normalize_earth_news_sources_payload(
{
"sources": [
{
"id": "36kr",
"name": "36氪",
"region": "asia-pacific",
"source_type": "rss",
"homepage_url": "https://www.36kr.com/",
"feed_url": "https://www.36kr.com/rss-center",
"feed_urls": ["https://www.36kr.com/rss-center"],
"feeds": [
{
"id": "feed-1",
"name": "36氪",
"url": "https://www.36kr.com/rss-center",
"type": "rss",
"enabled": True,
"default_category": "business",
}
],
"default_category": "business",
}
]
}
)
source = payload["sources"][0]
feed_urls = {feed["url"] for feed in source["feeds"]}
assert source["homepage_url"] == "https://www.36kr.com/"
assert source["feed_directory_url"] == "https://www.36kr.com/rss-center"
assert "https://www.36kr.com/rss-center" not in feed_urls
assert {
"https://36kr.com/feed",
"https://36kr.com/feed-article",
"https://36kr.com/feed-newsflash",
"https://36kr.com/feed-moment",
}.issubset(feed_urls)
def test_builtin_news_source_without_feed_children_gets_explicit_defaults():
payload = normalize_earth_news_sources_payload(
{
"sources": [
{
"id": "ebrun",
"name": "亿邦动力",
"region": "asia-pacific",
"source_type": "rss",
"homepage_url": "https://www.ebrun.com/",
"feed_url": "https://www.ebrun.com/rss/news_b2c.xml",
"feed_urls": ["https://www.ebrun.com/rss/news_b2c.xml"],
"default_category": "ecommerce",
}
]
}
)
source = payload["sources"][0]
feed_urls = {feed["url"] for feed in source["feeds"]}
assert source["feed_directory_url"] == "https://www.ebrun.com/rss/"
assert "https://www.ebrun.com/rss/" not in feed_urls
assert {
"https://www.ebrun.com/rss/news_b2c.xml",
"https://www.ebrun.com/rss/news_b2b.xml",
"https://www.ebrun.com/rss/news_retail.xml",
"https://www.ebrun.com/rss/news_o2o.xml",
"https://www.ebrun.com/rss/news_service.xml",
"https://www.ebrun.com/rss/news_data.xml",
"https://www.ebrun.com/rss/news_policy.xml",
}.issubset(feed_urls)
def test_builtin_fetchable_source_saved_as_reference_is_repaired():
payload = normalize_earth_news_sources_payload(
{
"sources": [
{
"id": "stats-china-online-retail",
"name": "国家统计局数据发布",
"region": "asia-pacific",
"source_type": "reference",
"enabled": False,
"homepage_url": "https://www.stats.gov.cn/sj/zxfb/",
"feed_url": "https://www.stats.gov.cn/sj/zxfb/",
"default_category": "ecommerce",
}
]
}
)
source = payload["sources"][0]
assert source["source_type"] == "rss"
assert source["enabled"] is True
assert source["priority"] == 19
assert source["source_tags"] == ["official_data", "ecommerce", "retail", "china"]
assert source["default_category"] == "ecommerce"
assert source["importance_weight"] == 36
assert source["feed_directory_url"] == ""
assert source["feeds"] == [
{
"id": "release",
"name": "数据发布",
"url": "https://www.stats.gov.cn/sj/zxfb/rss.xml",
"type": "rss",
"region": "asia-pacific",
"enabled": True,
"default_category": "ecommerce",
"tags": [],
"priority": 1,
}
]
def test_legacy_google_sources_merge_into_google_news_source():
payload = normalize_earth_news_sources_payload(
{
"sources": [
{
"id": "global-scan",
"name": "Global Monitor / World",
"region": "global",
"source_type": "aggregated",
"feed_url": "https://news.google.com/rss/search?q=world",
"homepage_url": "https://news.google.com/",
},
{
"id": "google-europe",
"name": "Global Monitor / Europe",
"region": "europe",
"source_type": "aggregated",
"feed_url": "https://news.google.com/rss/search?q=europe",
"homepage_url": "https://news.google.com/",
},
]
}
)
sources_by_id = {source["id"]: source for source in payload["sources"]}
assert "global-scan" not in sources_by_id
assert "google-europe" not in sources_by_id
assert "google-news" in sources_by_id
assert {feed["id"] for feed in sources_by_id["google-news"]["feeds"]} == {
"world",
"americas",
"europe",
"middle-east-africa",
"asia-pacific",
}
def test_feed_child_default_category_overrides_source_default():
source = NewsFeedSource(
id="multi-feed",
name="Multi Feed",
region="global",
feed_url="https://example.com/source.xml",
homepage_url="https://example.com",
default_category="business",
)
feed = NewsFeedEndpoint(
id="ecommerce-feed",
name="Ecommerce Feed",
url="https://example.com/ecommerce.xml",
default_category="ecommerce",
)
xml = """
<rss>
<channel>
<item>
<title>Quarterly results released</title>
<description>Company update.</description>
<link>https://example.com/results</link>
</item>
</channel>
</rss>
"""
items = _parse_feed_entries(xml, source, feed=feed)
assert items[0].feed_id == "ecommerce-feed"
assert items[0].feed_name == "Ecommerce Feed"
assert items[0].feed_default_category == "ecommerce"
assert items[0].category == "ecommerce"
@pytest.mark.asyncio
async def test_fetch_source_only_requests_enabled_feed_children(monkeypatch):
source = NewsFeedSource(
id="multi-feed",
name="Multi Feed",
region="global",
feed_url="https://example.com/source.xml",
homepage_url="https://example.com",
feeds=(
NewsFeedEndpoint(id="enabled", name="Enabled", url="https://example.com/enabled.xml", enabled=True),
NewsFeedEndpoint(id="disabled", name="Disabled", url="https://example.com/disabled.xml", enabled=False),
),
)
calls = []
async def fake_fetch_single(_client, feed_source, feed, *, config_payload=None):
calls.append(feed.id)
item = ParsedNewsItem(
id=f"{feed_source.id}:{feed.id}:1",
title="Fetched story",
summary="Fetched summary",
url=f"https://example.com/{feed.id}",
source="Example",
feed_name=feed.name,
feed_region="global",
homepage_url="https://example.com",
published_at=None,
feed_id=feed.id,
)
return feed_source, [item], None, {"source_id": feed_source.id, "feed_id": feed.id, "ok": True, "status": "ok", "item_count": 1, "count": 1}
monkeypatch.setattr("app.services.earth_news._fetch_single_feed_url", fake_fetch_single)
source_result, items, error, health = await _fetch_source(object(), source)
assert source_result.id == "multi-feed"
assert calls == ["enabled"]
assert error is None
assert [item.feed_id for item in items] == ["enabled"]
assert health["ok"] is True
assert [result["feed_id"] for result in health["feed_results"]] == ["enabled"]
@pytest.mark.asyncio
async def test_fetch_source_filters_google_feed_children_by_active_region(monkeypatch):
source = NewsFeedSource(
id="google-news",
name="Google News",
region="global",
feed_url="https://news.google.com/rss",
homepage_url="https://news.google.com/",
source_type="aggregated",
feeds=(
NewsFeedEndpoint(id="world", name="全球", url="https://example.com/world.xml", type="aggregated", region="global"),
NewsFeedEndpoint(id="europe", name="欧洲", url="https://example.com/europe.xml", type="aggregated", region="europe"),
NewsFeedEndpoint(id="americas", name="美洲", url="https://example.com/americas.xml", type="aggregated", region="americas"),
),
)
calls = []
async def fake_fetch_single(_client, feed_source, feed, *, config_payload=None):
calls.append(feed.id)
item = ParsedNewsItem(
id=f"{feed_source.id}:{feed.id}:1",
title=f"{feed.name} headline",
summary="Fetched summary",
url=f"https://example.com/{feed.id}",
source="Example",
feed_name=feed.name,
feed_region=feed.region,
homepage_url="https://example.com",
published_at=None,
feed_id=feed.id,
)
return feed_source, [item], None, {"source_id": feed_source.id, "feed_id": feed.id, "ok": True, "status": "ok", "item_count": 1, "count": 1}
monkeypatch.setattr("app.services.earth_news._fetch_single_feed_url", fake_fetch_single)
_source_result, items, error, health = await _fetch_source(object(), source, active_region="europe")
assert error is None
assert calls == ["world", "europe"]
assert [item.feed_region for item in items] == ["global", "europe"]
assert [result["feed_id"] for result in health["feed_results"]] == ["world", "europe"]
def test_parse_rdf_rss_items_with_namespaces():
source = NewsFeedSource(
id="dw-top",
name="DW Top Stories",
region="europe",
feed_url="https://rss.dw.com/rdf/rss-en-top",
homepage_url="https://www.dw.com/en/top-stories/s-9097",
)
xml = """
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#"
xmlns="http://purl.org/rss/1.0/">
<item rdf:about="https://example.com/dw">
<title>German retail sales rise</title>
<link>https://example.com/dw</link>
<description>Retail summary</description>
</item>
</rdf:RDF>
"""
items = _parse_feed_entries(xml, source)
assert len(items) == 1
assert items[0].title == "German retail sales rise"
def test_news_classification_marks_ecommerce_and_importance():
source = NewsFeedSource(
id="ebrun",
name="亿邦动力",
region="asia-pacific",
feed_url="https://www.ebrun.com/rss/",
homepage_url="https://www.ebrun.com/",
source_tags=("business_news", "ecommerce", "china"),
default_category="ecommerce",
importance_weight=14,
)
item = ParsedNewsItem(
id="ebrun:test",
title="跨境电商平台 GMV 同比增长,物流履约效率提升",
summary="订单量和网上零售额继续增长。",
url="https://example.com/ecommerce",
source="亿邦动力",
feed_name="亿邦动力",
feed_region="asia-pacific",
homepage_url="https://www.ebrun.com/",
published_at=None,
)
apply_news_classification(item, source)
assert item.category == "ecommerce"
assert "cross_border_ecommerce" in item.item_tags
assert "logistics_fulfillment" in item.item_tags
assert item.importance_level in {"high", "critical"}
assert "命中电商数据指标" in item.importance_reasons
@pytest.mark.asyncio
async def test_enrich_items_with_target_locations_uses_ai_and_geocode(monkeypatch):
item = ParsedNewsItem(
@@ -338,8 +1009,8 @@ async def test_earth_news_payload_returns_anchor_items_and_enqueues_location_job
published_at=datetime(2026, 5, 15, 3, 0, tzinfo=UTC),
)
async def fake_fetch_source(_client, feed_source):
return feed_source, [item], None
async def fake_fetch_source(_client, feed_source, **_kwargs):
return feed_source, [item], None, {"source_id": feed_source.id, "ok": True, "status": "ok", "count": 1}
async def fake_get_cached_target_location_patch(_item_id):
return None
@@ -402,9 +1073,11 @@ async def test_earth_news_payload_uses_fresh_database_items_without_rss(monkeypa
async def fake_get_earth_news_freshness(_db, *, active_region):
return 12, datetime.now(UTC)
async def fake_list_earth_news_items(_db, *, active_region, limit):
assert limit == 12
return [item]
async def fake_list_earth_news_items(_db, *, active_region, limit, categories=None, source_ids=None):
if source_ids is None:
assert limit == 12
return [item]
return []
async def fail_fetch(_sources):
raise AssertionError("fresh database items should not fetch RSS")
@@ -452,10 +1125,10 @@ async def test_earth_news_payload_keeps_current_items_and_all_cruise_items(monke
async def fake_get_earth_news_freshness(_db, *, active_region):
return 12, datetime.now(UTC)
async def fake_list_earth_news_items(_db, *, active_region, limit):
async def fake_list_earth_news_items(_db, *, active_region, limit, categories=None):
return [current_item]
async def fake_list_earth_news_cruise_items(_db, *, limit):
async def fake_list_earth_news_cruise_items(_db, *, limit, categories=None):
return [current_item, cruise_item]
async def fake_enqueue_target_location_job(_payload, **_kwargs):
@@ -477,6 +1150,92 @@ async def test_earth_news_payload_keeps_current_items_and_all_cruise_items(monke
assert payload["cruise_items"][1]["region"] == "asia-pacific"
@pytest.mark.asyncio
async def test_earth_news_payload_passes_region_and_category_filters_to_store(monkeypatch):
class FakeDb:
execute = object()
captured = {}
item = ParsedNewsItem(
id="db:business",
title="Business story",
summary="Business summary",
url="https://example.com/business",
source="Stored Source",
feed_name="Stored Feed",
feed_region="europe",
homepage_url="https://example.com",
published_at=datetime(2026, 5, 15, 3, 0, tzinfo=UTC),
category="business",
)
async def fake_get_earth_news_freshness(_db, *, active_region):
captured["freshness_region"] = active_region
return 12, datetime.now(UTC)
async def fake_list_earth_news_items(_db, *, active_region, limit, categories=None, source_ids=None):
captured["items_region"] = active_region
captured["items_categories"] = categories
captured.setdefault("items_source_ids", []).append(source_ids)
return [item] if source_ids is None else []
async def fake_list_earth_news_cruise_items(_db, *, limit, categories=None, source_ids=None):
captured["cruise_categories"] = categories
captured["cruise_source_ids"] = source_ids
return [item]
async def fake_enqueue_target_location_job(_payload, **_kwargs):
return True
monkeypatch.setattr("app.services.earth_news_store.get_earth_news_freshness", fake_get_earth_news_freshness)
monkeypatch.setattr("app.services.earth_news_store.list_earth_news_items", fake_list_earth_news_items)
monkeypatch.setattr("app.services.earth_news_store.list_earth_news_cruise_items", fake_list_earth_news_cruise_items)
monkeypatch.setattr("app.services.earth_news_queue.enqueue_target_location_job", fake_enqueue_target_location_job)
monkeypatch.setattr("app.services.earth_news._fetch_rss_items_for_sources", lambda _sources: (_ for _ in ()).throw(AssertionError("fresh database items should not fetch RSS")))
payload = await get_earth_news_payload(
lat=35.0,
lon=-100.0,
region="europe",
categories={"business", "ecommerce"},
db=FakeDb(),
)
assert captured["freshness_region"] == "europe"
assert captured["items_region"] == "europe"
assert captured["items_categories"] == {"business", "ecommerce"}
assert captured["items_source_ids"][0] is None
assert any(source_ids for source_ids in captured["items_source_ids"][1:])
assert captured["cruise_categories"] == {"business", "ecommerce"}
assert captured["cruise_source_ids"] is None
assert payload["filters"] == {
"region": "europe",
"categories": ["business", "ecommerce"],
"sources": [],
"limit": 12,
"locale": "zh-CN",
"has_breaking": False,
"highest_breaking_level": "none",
}
assert payload["items"][0]["category"] == "business"
@pytest.mark.asyncio
async def test_news_source_test_treats_type_reference_as_non_fetching():
result = await run_news_source_config_test(
{
"id": "reference-only",
"name": "Reference Only",
"type": "reference",
"feed_url": "https://example.com",
}
)
assert result["ok"] is False
assert result["health"]["status"] == "reference"
assert "不参与 RSS/Atom 抓取" in result["error"]
@pytest.mark.asyncio
async def test_earth_news_payload_initializes_empty_database_from_rss(monkeypatch):
db = object()
@@ -512,14 +1271,14 @@ async def test_earth_news_payload_initializes_empty_database_from_rss(monkeypatc
async def fake_get_earth_news_freshness(_db, *, active_region):
return 0, None
async def fake_fetch_rss_items_for_sources(_sources):
return [item], []
async def fake_fetch_rss_items_for_sources(_sources, **_kwargs):
return [item], [], {"test-feed": {"source_id": "test-feed", "ok": True, "status": "ok", "count": 1}}
async def fake_upsert_earth_news_items(_db, items):
upserted.extend(items)
return len(items)
async def fake_list_earth_news_items(_db, *, active_region, limit):
async def fake_list_earth_news_items(_db, *, active_region, limit, categories=None):
return [item]
async def fake_enqueue_target_location_job(payload, **_kwargs):
@@ -568,14 +1327,14 @@ async def test_earth_news_payload_supplements_stale_database_items(monkeypatch):
async def fake_get_earth_news_freshness(_db, *, active_region):
return 12, datetime(2026, 5, 14, 3, 0, tzinfo=UTC)
async def fake_fetch_rss_items_for_sources(_sources):
async def fake_fetch_rss_items_for_sources(_sources, **_kwargs):
fetched.append(True)
return [old_item], []
return [old_item], [], {"stored": {"source_id": "stored", "ok": True, "status": "ok", "count": 1}}
async def fake_upsert_earth_news_items(_db, items):
return len(items)
async def fake_list_earth_news_items(_db, *, active_region, limit):
async def fake_list_earth_news_items(_db, *, active_region, limit, categories=None):
return [old_item]
async def fake_enqueue_target_location_job(_payload, **_kwargs):
@@ -630,8 +1389,8 @@ async def test_earth_news_payload_merges_cached_location_patch(monkeypatch):
},
}
async def fake_fetch_source(_client, feed_source):
return feed_source, [item], None
async def fake_fetch_source(_client, feed_source, **_kwargs):
return feed_source, [item], None, {"source_id": feed_source.id, "ok": True, "status": "ok", "count": 1}
async def fake_get_cached_target_location_patch(_item_id):
return cached_patch
@@ -697,8 +1456,8 @@ async def test_earth_news_payload_requeues_cached_failed_localization(monkeypatc
}
enqueued = []
async def fake_fetch_source(_client, feed_source):
return feed_source, [item], None
async def fake_fetch_source(_client, feed_source, **_kwargs):
return feed_source, [item], None, {"source_id": feed_source.id, "ok": True, "status": "ok", "count": 1}
async def fake_get_cached_target_location_patch(_item_id):
return cached_patch

View File

@@ -0,0 +1,252 @@
from datetime import UTC, datetime
import pytest
from app.core.enums import NewsSourceType
from app.models.earth_news import EarthNewsItem
from app.models.system_setting import SystemSetting
from app.services.earth_news import REGION_ANCHORS
from app.services.earth_news_manual import (
DEFAULT_MANUAL_NEWS_GROUP_ID,
create_manual_news_group,
import_manual_news_items,
list_news_groups,
list_news_records,
parse_manual_news_import_upload,
rename_manual_news_group,
upsert_manual_news_item,
)
class _FakeResult:
def __init__(self, rows=None, scalar=None):
self.rows = rows or []
self._scalar = scalar
def scalar_one_or_none(self):
return self._scalar
def scalar(self):
return self._scalar
def scalars(self):
return self
def all(self):
return self.rows
class _FakeNewsSession:
def __init__(self, records=None, setting=None):
self.records = dict(records or {})
self.setting = setting
async def get(self, _model, item_id):
return self.records.get(item_id)
def add(self, item):
if isinstance(item, SystemSetting):
self.setting = item
else:
self.records[item.id] = item
async def execute(self, stmt):
statement = str(stmt)
if "system_settings" in statement:
return _FakeResult(scalar=self.setting)
if "count" in statement.lower():
return _FakeResult(scalar=len(self.records))
return _FakeResult(rows=list(self.records.values()))
async def flush(self):
return None
@pytest.fixture
def fake_news_queue(monkeypatch):
queued = []
async def _enqueue(payload, force=False):
queued.append({"payload": payload, "force": force})
return True
monkeypatch.setattr("app.services.earth_news_manual.enqueue_target_location_job", _enqueue)
return queued
@pytest.mark.asyncio
async def test_manual_news_upsert_uses_region_anchor_and_manual_metadata(fake_news_queue):
db = _FakeNewsSession()
result = await upsert_manual_news_item(
db,
{
"title": "手动添加的新闻",
"summary": "一条用于测试的手动新闻。",
"source": "人工录入",
"region": "europe",
"published_at": "2026-05-15T03:00:00Z",
"tags": ["manual", "test"],
},
)
anchor = REGION_ANCHORS["europe"]
assert result.created is True
assert result.queued is True
assert result.item.id.startswith("manual:")
assert result.item.feed_name == "手动添加"
assert result.item.source == "人工录入"
assert result.item.latitude == anchor.latitude
assert result.item.longitude == anchor.longitude
assert result.item.location_source == "region_anchor"
assert result.item.verified is False
assert result.item.location_meta["news_meta"]["feed_type"] == NewsSourceType.MANUAL.value
assert result.item.location_meta["news_meta"]["source_type"] == NewsSourceType.MANUAL.value
assert result.item.location_meta["news_meta"]["manual_group_id"] == DEFAULT_MANUAL_NEWS_GROUP_ID
assert len(fake_news_queue) == 1
@pytest.mark.asyncio
async def test_manual_news_duplicate_import_upserts_without_duplicate_rows(fake_news_queue):
db = _FakeNewsSession()
payload = {
"title": "Same manual story",
"source": "Manual Desk",
"published_at": "2026-05-15T03:00:00Z",
"region": "global",
}
first = await upsert_manual_news_item(db, payload)
second = await upsert_manual_news_item(db, {**payload, "summary": "Updated summary"})
assert first.created is True
assert second.created is False
assert len(db.records) == 1
assert db.records[first.item.id].summary == "Updated summary"
@pytest.mark.asyncio
async def test_manual_news_edit_without_location_preserves_manual_coordinates(fake_news_queue):
db = _FakeNewsSession()
created = await upsert_manual_news_item(
db,
{
"title": "Taipei-1 data center update",
"summary": "Initial summary.",
"region": "asia-pacific",
"published_at": "2026-05-15T03:00:00Z",
"location": {"label": "Kaohsiung, Taiwan", "latitude": 22.6273, "longitude": 120.3014},
},
)
updated = await upsert_manual_news_item(
db,
{
"title": "Taipei-1 data center update",
"summary": "Edited summary only.",
"region": "asia-pacific",
"published_at": "2026-05-15T03:00:00Z",
},
item_id_override=created.item.id,
)
assert updated.created is False
assert updated.item.latitude == pytest.approx(22.6273)
assert updated.item.longitude == pytest.approx(120.3014)
assert updated.item.location_source == "manual_location"
assert updated.item.verified is True
@pytest.mark.asyncio
async def test_manual_news_api_service_rejects_rss_records(fake_news_queue):
rss_record = EarthNewsItem(
id="bbc-world:example",
title="RSS story",
summary="RSS summary",
source="BBC World",
feed_name="BBC World",
region="global",
latitude=20,
longitude=0,
location_label="全球",
location_source="region_anchor",
verified=False,
location_meta={"news_meta": {"feed_type": "rss"}},
first_seen_at=datetime.now(UTC),
last_seen_at=datetime.now(UTC),
)
db = _FakeNewsSession({rss_record.id: rss_record})
with pytest.raises(PermissionError):
await upsert_manual_news_item(
db,
{"title": "Edited title", "region": "global"},
item_id_override=rss_record.id,
)
@pytest.mark.asyncio
async def test_manual_news_import_reports_per_item_errors(fake_news_queue):
db = _FakeNewsSession()
result = await import_manual_news_items(
db,
[
{"title": "Valid manual news", "region": "global"},
{"summary": "missing title"},
],
)
assert result["created"] == 1
assert result["failed"] == 1
assert result["errors"][0]["index"] == 1
@pytest.mark.asyncio
async def test_manual_news_groups_default_create_and_rename(fake_news_queue):
db = _FakeNewsSession()
initial = await list_news_groups(db)
assert initial["manual_groups"][0]["id"] == DEFAULT_MANUAL_NEWS_GROUP_ID
assert initial["manual_groups"][0]["name"] == "新建新闻组"
group = await create_manual_news_group(db, "专题组")
assert group["name"] == "专题组"
assert db.setting is not None
await upsert_manual_news_item(db, {"title": "Grouped story", "region": "global"}, group_id=group["id"])
renamed = await rename_manual_news_group(db, group["id"], "重命名专题")
record = next(iter(db.records.values()))
assert renamed["name"] == "重命名专题"
assert record.location_meta["news_meta"]["manual_group_id"] == group["id"]
assert record.location_meta["news_meta"]["manual_group_name"] == "重命名专题"
@pytest.mark.asyncio
async def test_manual_news_list_filters_by_group_id(fake_news_queue):
db = _FakeNewsSession()
group = await create_manual_news_group(db, "导入组")
await import_manual_news_items(
db,
[
{"title": "In group", "region": "global"},
{"title": "Also in group", "region": "global"},
],
group_id=group["id"],
)
await upsert_manual_news_item(db, {"title": "Default group", "region": "global"})
grouped = await list_news_records(db, page=1, page_size=20, group_id=group["id"])
default_group = await list_news_records(db, page=1, page_size=20, group_id=DEFAULT_MANUAL_NEWS_GROUP_ID)
assert grouped["total"] == 2
assert {item["manual_group_id"] for item in grouped["items"]} == {group["id"]}
assert default_group["total"] == 1
@pytest.mark.asyncio
async def test_manual_news_import_parser_requires_json_array():
with pytest.raises(ValueError, match="顶层必须是数组"):
await parse_manual_news_import_upload(b'{"title":"not an array"}')

View File

@@ -0,0 +1,74 @@
"""Compatibility contracts for stable backend protocol enums."""
from __future__ import annotations
from datetime import UTC, datetime, timedelta
from types import SimpleNamespace
from app.core.enums import (
BreakingLevel,
BreakingScope,
JobStatus,
NewsImportanceLevel,
NewsSourceType,
PlaygroundMessageKind,
PlaygroundMessageStatus,
UserRole,
parse_enum,
)
from app.services.earth_news_classification import (
BREAKING_LEVEL_RANK,
BREAKING_TTL,
breaking_sort_rank,
importance_level,
)
def test_protocol_enum_values_remain_api_compatible() -> None:
assert [item.value for item in NewsImportanceLevel] == ["low", "medium", "high", "critical"]
assert [item.value for item in BreakingLevel] == ["none", "watch", "breaking", "critical"]
assert [item.value for item in BreakingScope] == ["regional", "global"]
assert [item.value for item in NewsSourceType] == ["rss", "atom", "aggregated", "reference", "manual"]
assert [item.value for item in UserRole] == ["viewer", "admin", "super_admin"]
assert JobStatus.RUNNING.value == "running"
assert PlaygroundMessageKind.THINKING.value == "thinking"
assert PlaygroundMessageStatus.ERROR.value == "error"
assert PlaygroundMessageStatus.STOPPED.value == "stopped"
def test_parse_enum_accepts_legacy_strings_and_safely_falls_back(caplog) -> None:
assert parse_enum(JobStatus, "RUNNING", JobStatus.FAILED) is JobStatus.RUNNING
assert parse_enum(JobStatus, None, JobStatus.QUEUED) is JobStatus.QUEUED
assert parse_enum(JobStatus, "legacy-unknown", JobStatus.FAILED) is JobStatus.FAILED
assert "legacy-unknown" in caplog.text
def test_importance_level_boundaries() -> None:
expected = {
34: NewsImportanceLevel.LOW,
35: NewsImportanceLevel.MEDIUM,
59: NewsImportanceLevel.MEDIUM,
60: NewsImportanceLevel.HIGH,
79: NewsImportanceLevel.HIGH,
80: NewsImportanceLevel.CRITICAL,
}
assert {score: importance_level(score) for score in expected} == expected
def test_breaking_rank_and_ttl_contracts() -> None:
assert BREAKING_LEVEL_RANK[BreakingLevel.CRITICAL] > BREAKING_LEVEL_RANK[BreakingLevel.BREAKING]
assert BREAKING_TTL[BreakingLevel.WATCH] == timedelta(hours=6)
assert BREAKING_TTL[BreakingLevel.BREAKING] == timedelta(hours=12)
assert BREAKING_TTL[BreakingLevel.CRITICAL] == timedelta(hours=24)
now = datetime.now(UTC)
active = SimpleNamespace(
breaking_level=BreakingLevel.BREAKING.value,
breaking_expires_at=now + timedelta(minutes=1),
)
expired = SimpleNamespace(
breaking_level=BreakingLevel.CRITICAL.value,
breaking_expires_at=now - timedelta(minutes=1),
)
assert breaking_sort_rank(active) == BREAKING_LEVEL_RANK[BreakingLevel.BREAKING]
assert breaking_sort_rank(expired) == 0

View File

@@ -9,6 +9,8 @@ import pytest
from app.core.logging import PlanetContextFilter, PlanetFormatter, get_logger
from app.core.request_context import set_request_id
from app.services import business_logs
from app.services import persistent_logs
from app.models.system_log import ObservabilityEvent, ObservabilityEventGroup
def _capture_output(callback):
@@ -98,6 +100,86 @@ def test_business_context_redacts_nested_sensitive_values():
assert context["nested"]["safe"] == "visible"
def test_observability_fingerprint_normalizes_hls_fragments():
first = persistent_logs.build_observability_fingerprint(
source="earth-client",
service="earth",
module="tv",
category="hls-proxy",
event="hls.fragment.failed",
message="HLS 分片加载失败: index_5_9086220.ts?m=1725933270",
context={"status_code": 502},
)
second = persistent_logs.build_observability_fingerprint(
source="earth-client",
service="earth",
module="tv",
category="hls-proxy",
event="hls.fragment.failed",
message="HLS 分片加载失败: index_5_9086361.ts?m=1725934270",
context={"status_code": 502},
)
assert first == second
@pytest.mark.asyncio
async def test_record_observability_event_updates_group_count(monkeypatch):
events: list[ObservabilityEvent] = []
groups: dict[str, ObservabilityEventGroup] = {}
class FakeSession:
async def __aenter__(self):
return self
async def __aexit__(self, exc_type, exc, tb):
return False
def add(self, item):
if isinstance(item, ObservabilityEvent):
events.append(item)
elif isinstance(item, ObservabilityEventGroup):
groups[item.fingerprint] = item
async def get(self, model, key):
if model is ObservabilityEventGroup:
return groups.get(key)
return None
async def commit(self):
return None
monkeypatch.setattr(persistent_logs, "async_session_factory", lambda: FakeSession())
await persistent_logs.record_observability_event(
source="earth-client",
level="error",
service="earth",
module="tv",
category="hls-proxy",
event="hls.fragment.failed",
message="HLS 分片加载失败: index_5_9086220.ts?m=1725933270",
context={"status_code": 502},
occurrence_count=2,
)
await persistent_logs.record_observability_event(
source="earth-client",
level="error",
service="earth",
module="tv",
category="hls-proxy",
event="hls.fragment.failed",
message="HLS 分片加载失败: index_5_9086361.ts?m=1725934270",
context={"status_code": 502},
occurrence_count=1,
)
assert len(events) == 2
assert len(groups) == 1
group = next(iter(groups.values()))
assert group.count == 3
@pytest.mark.asyncio
async def test_emit_business_log_persists_sanitized_system_event(monkeypatch):
events = []

View File

@@ -40,6 +40,9 @@ def test_gesture_event_serializes_stable_protocol_fields():
assert payload["seq"] == 7
assert payload["source"] == "motion-agent"
assert payload["mode"] == "single"
assert payload["protocol_version"] == "motion.v2"
assert payload["input_mode"] == "single"
assert payload["camera_id"] == "unknown"
assert payload["payload"] == {}
@@ -88,8 +91,13 @@ def test_motion_server_status_includes_dry_run_camera_and_heartbeat():
assert status["camera_count"] == 1
assert status["active_camera_ids"] == ["dry-run:null-camera"]
assert status["recognizer"] == "dry-run"
assert status["protocol_version"] == "motion.v2"
assert status["armed"] is False
assert status["paused"] is False
assert status["devices_open"] is False
assert heartbeat == {
"timestamp_ms": 123,
"protocol_version": "motion.v2",
"source": "motion-agent",
"type": "heartbeat",
}
@@ -109,6 +117,7 @@ def test_skeleton_event_serializes_without_raw_image_fields():
payload = json.loads(event.to_json())
assert payload["type"] == "skeleton"
assert payload["protocol_version"] == "motion.v2"
assert payload["matched_gesture"] == "rotate_left"
assert payload["confidence"] == 0.91
assert payload["camera_id"] == "usb:0"
@@ -120,6 +129,25 @@ def test_skeleton_event_serializes_without_raw_image_fields():
assert "frame" not in payload
def test_v2_gesture_set_accepts_frontend_motion_gestures():
state = GestureStateMachine(confidence_threshold=0.7, cooldown_ms=0)
for gesture in [
"rotate_up",
"rotate_down",
"focus_prev",
"focus_next",
"layer_prev",
"layer_next",
]:
event = state.accept(
GestureObservation(gesture, confidence=0.9, intensity=0.8, timestamp_ms=1000)
)
assert event is not None
assert event.gesture == gesture
def test_dry_run_recognizer_produces_debug_skeleton():
server = MotionAgentServer(MotionAgentConfig(dry_run=True))
@@ -240,3 +268,92 @@ async def test_motion_agent_cli_reports_dependency_error_without_traceback(monke
assert exit_code == 2
assert "Motion agent failed: missing cv stack" in captured.err
assert "Traceback" not in captured.err
@pytest.mark.asyncio
async def test_motion_agent_command_updates_control_state():
server = MotionAgentServer(MotionAgentConfig(dry_run=True))
armed = await server.handle_command(
json.dumps(
{
"type": "command",
"command": "set_armed",
"request_id": "req-armed",
"payload": {"armed": True},
}
)
)
paused = await server.handle_command(
{
"type": "command",
"command": "set_paused",
"request_id": "req-paused",
"payload": {"paused": True},
}
)
assert armed.ok is True
assert armed.request_id == "req-armed"
assert armed.status["armed"] is True
assert paused.ok is True
assert paused.status["paused"] is True
@pytest.mark.asyncio
async def test_motion_agent_open_devices_command_accepts_dual_mode():
server = MotionAgentServer(MotionAgentConfig(dry_run=True))
try:
result = await server.handle_command(
{
"type": "command",
"command": "open_devices",
"request_id": "req-open",
"payload": {"input_mode": "dual_redundant"},
}
)
assert result.ok is True
assert result.status["input_mode"] == "dual_redundant"
assert result.status["active_camera_ids"] == ("dry-run:null-camera",)
assert server._recognition_subprocess is not None
assert server._recognition_subprocess.returncode is None
finally:
await server.stop_recognition_subprocess()
def test_motion_agent_dual_fusion_merges_matching_observations():
server = MotionAgentServer(MotionAgentConfig(dry_run=True))
server.state.mode = "dual_redundant"
selected, fusion = server._fuse_observations(
[
GestureObservation("zoom_in", confidence=0.82, intensity=0.4, camera_id="usb:0"),
GestureObservation("zoom_in", confidence=0.86, intensity=0.8, camera_id="usb:1"),
]
)
assert selected.gesture == "zoom_in"
assert selected.camera_id == "fusion"
assert selected.confidence > 0.86
assert fusion == {
"source_cameras": ["usb:1", "usb:0"],
"window_ms": 120,
"reason": "matched_observations",
}
def test_motion_agent_dual_fusion_suppresses_close_conflict():
server = MotionAgentServer(MotionAgentConfig(dry_run=True, confidence_threshold=0.7))
selected, fusion = server._fuse_observations(
[
GestureObservation("zoom_in", confidence=0.82, intensity=0.5, camera_id="usb:0"),
GestureObservation("zoom_out", confidence=0.78, intensity=0.5, camera_id="usb:1"),
]
)
assert selected.gesture == "zoom_in"
assert selected.confidence == 0
assert fusion["reason"] == "conflict_ignored"

View File

@@ -2,8 +2,10 @@ from __future__ import annotations
import json
from datetime import UTC, datetime
from pathlib import Path
from app.models.system_log import SystemLog
from app.services import system_logs
@@ -40,10 +42,10 @@ def test_read_log_snapshot_uses_structured_buffer_timestamp_level_and_search(mon
{
"earth-client": system_logs.LogSource(
source_id="earth-client",
name="Earth 浏览器端",
name="智能星球浏览器端",
kind="buffer",
location="redis://planet:system_logs:earth-client",
description="Earth 浏览器端上报日志",
description="智能星球浏览器端上报日志",
category="client",
buffer_key=system_logs.get_buffer_log_key("earth-client"),
)
@@ -156,6 +158,52 @@ def test_append_buffer_log_persists_normalized_level(monkeypatch):
assert payload["message"] == "feed delayed"
def test_list_log_sources_includes_admin_client(monkeypatch):
fake_redis = FakeRedis()
monkeypatch.setattr(system_logs, "redis_client", fake_redis)
sources = system_logs.list_log_sources()
admin_source = next(item for item in sources if item["source_id"] == "admin-client")
assert admin_source["kind"] == "buffer"
assert admin_source["category"] == "client"
def test_read_log_events_returns_stable_cursors(tmp_path: Path, monkeypatch):
log_path = tmp_path / "backend.log"
log_path.write_text(
"\n".join(
[
"2026-04-22 08:00:00 INFO service booted",
"2026-04-22 08:01:00 ERROR service failed",
]
),
encoding="utf-8",
)
monkeypatch.setattr(
system_logs,
"LOG_SOURCES",
{
"backend": system_logs.LogSource(
source_id="backend",
name="后端服务",
kind="file",
location=str(log_path),
description="测试文件日志",
category="service",
)
},
)
events = system_logs.read_log_events("backend", 50, level="error")
assert events is not None
assert len(events) == 1
assert events[0].source_id == "backend"
assert events[0].cursor.startswith("backend:")
assert events[0].line.endswith("ERROR service failed")
def test_infer_log_level_prefers_leading_prefix_over_query_string():
line = 'INFO: 127.0.0.1 - "GET /api/v1/system/logs/backend?limit=200&level=error&levels=error HTTP/1.1" 200 OK'
@@ -215,3 +263,22 @@ def test_read_log_snapshot_strips_nul_bytes_from_file_lines(tmp_path: Path, monk
"ERROR: bind failed",
"2026-04-23 23:41:32 INFO service=backend message=request served",
]
def test_database_system_log_search_matches_context_key_value_aliases():
record = SystemLog(
id=2218,
occurred_at=datetime(2026, 5, 28, 9, 14, 50, tzinfo=UTC),
source="backend",
service="collector",
module="app.services.collectors.base",
event="collector.run.failed",
level="error",
message="Collector run failed",
context={"collector_name": "celestrak_tle", "datasource_id": 20, "task_id": 26906},
)
event = system_logs._database_event_from_system_record(record)
assert system_logs.event_matches_search(event, "task_id=26906")
assert system_logs.event_matches_search(event, "datasource_id=20")

View File

@@ -0,0 +1,35 @@
from app.api.v1.tv import _rewrite_hls_uri_attributes, _should_strip_hls_metadata_line
def test_rewrite_hls_uri_attributes_rewrites_subtitle_manifest_url():
line = '#EXT-X-MEDIA:TYPE=SUBTITLES,GROUP-ID="subs",NAME="English",URI="index_3_0.m3u8"'
rewritten = _rewrite_hls_uri_attributes(
line,
base_url="https://example.com/live/master.m3u8",
)
assert 'URI="/api/v1/tv/proxy?url=https%3A%2F%2Fexample.com%2Flive%2Findex_3_0.m3u8"' in rewritten
def test_rewrite_hls_uri_attributes_rewrites_absolute_uri():
line = '#EXT-X-I-FRAME-STREAM-INF:BANDWIDTH=1234,URI="https://cdn.example.com/live/iframe.m3u8"'
rewritten = _rewrite_hls_uri_attributes(
line,
base_url="https://example.com/live/master.m3u8",
)
assert 'URI="/api/v1/tv/proxy?url=https%3A%2F%2Fcdn.example.com%2Flive%2Fiframe.m3u8"' in rewritten
def test_strip_hls_subtitle_media_metadata():
line = '#EXT-X-MEDIA:TYPE=SUBTITLES,GROUP-ID="subs",NAME="English",URI="index_3_0.m3u8"'
assert _should_strip_hls_metadata_line(line) is True
def test_keep_hls_audio_media_metadata():
line = '#EXT-X-MEDIA:TYPE=AUDIO,GROUP-ID="audio",NAME="English",URI="audio.m3u8"'
assert _should_strip_hls_metadata_line(line) is False

View File

@@ -8,7 +8,7 @@ from app.api.v1 import visualization
from app.api.v1.visualization import convert_vessels_to_geojson
from app.db.session import get_db
from app.main import app
from app.models.vessel import AISRawObservation, VesselPosition, VesselStatic
from app.models.vessel import AISRawObservation, VesselCurrentState, VesselPosition, VesselStatic
from app.services import barentswatch
from app.services.collectors.aisstream import AISStreamCollector
from app.services.collectors.vessel_ais import VesselAISCollector
@@ -17,6 +17,7 @@ from app.services.vessel_ais_aggregation import (
build_field_conflict_candidates,
build_observation_hash,
record_vessel_ais_observation,
upsert_vessel_current_state,
)
@@ -109,6 +110,58 @@ async def test_record_vessel_ais_observation_skips_existing_hash():
assert db.added == []
@pytest.mark.asyncio
async def test_upsert_vessel_current_state_keeps_latest_position_and_static_fields():
current = VesselCurrentState(
mmsi=257123000,
lat=59.91,
lon=10.73,
name="OSLO TRADER",
source="barentswatch_vessels",
observed_at=datetime(2026, 4, 30, 12, 0, tzinfo=timezone.utc),
field_sources={"name": "aisstream_vessels"},
)
class _Session:
async def get(self, _model, _mmsi):
return current
def add(self, _item):
raise AssertionError("existing current state should be updated")
db = _Session()
result = await upsert_vessel_current_state(
db,
source="aisstream_vessels",
normalized_payload={"mmsi": 257123000, "lat": 59.92, "lon": 10.74, "sog": 12.4},
observed_at=datetime(2026, 4, 30, 12, 1, tzinfo=timezone.utc),
)
assert result is current
assert current.lat == pytest.approx(59.92)
assert current.lon == pytest.approx(10.74)
assert current.name == "OSLO TRADER"
assert current.source == "aisstream_vessels"
await upsert_vessel_current_state(
db,
source="barentswatch_vessels",
normalized_payload={"mmsi": 257123000, "lat": 59.93, "lon": 10.75, "name": "LOW PRIORITY"},
observed_at=datetime(2026, 4, 30, 12, 2, tzinfo=timezone.utc),
)
assert current.lat == pytest.approx(59.93)
assert current.name == "OSLO TRADER"
await upsert_vessel_current_state(
db,
source="barentswatch_vessels",
normalized_payload={"mmsi": 257123000, "lat": 1, "lon": 2, "name": "OLD"},
observed_at=datetime(2026, 4, 30, 11, 59, tzinfo=timezone.utc),
)
assert current.lat == pytest.approx(59.93)
assert current.name == "OSLO TRADER"
def test_build_field_conflict_candidates_from_raw_observations():
observations = [
AISRawObservation(
@@ -527,7 +580,7 @@ async def test_vessel_snapshot_filters_type_and_bbox(monkeypatch):
now = datetime(2026, 4, 28, 1, 0, tzinfo=timezone.utc)
monkeypatch.setattr(
visualization,
"get_aggregated_vessels_snapshot",
"get_current_vessels_snapshot",
AsyncMock(
return_value=[
{
@@ -573,6 +626,36 @@ async def test_vessel_snapshot_filters_type_and_bbox(monkeypatch):
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_vessel_snapshot_accepts_fractional_zoom(monkeypatch):
monkeypatch.setattr(
visualization,
"get_current_vessels_snapshot",
AsyncMock(return_value=[]),
)
async def override_get_db():
yield object()
app.dependency_overrides[get_db] = override_get_db
transport = ASGITransport(app=app)
try:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get(
"/api/v1/vessels/snapshot",
params={
"bbox": "-180,-85.05112878,180,85.05112878",
"zoom": 1.6,
"limit": 3000,
},
)
assert response.status_code == 200
assert response.json()["count"] == 0
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_legacy_vessels_geojson_route_is_not_registered():
transport = ASGITransport(app=app)
@@ -597,7 +680,7 @@ async def test_vessel_snapshot_filters_bbox_and_caps_limit(monkeypatch):
now = datetime(2026, 4, 28, 1, 0, tzinfo=timezone.utc)
captured = {}
async def fake_get_aggregated_vessels_snapshot(db, *, bbox, limit, observed_since):
async def fake_get_current_vessels_snapshot(db, *, bbox, limit, observed_since):
captured["bbox"] = bbox
captured["limit"] = limit
captured["observed_since"] = observed_since
@@ -624,8 +707,8 @@ async def test_vessel_snapshot_filters_bbox_and_caps_limit(monkeypatch):
monkeypatch.setattr(
visualization,
"get_aggregated_vessels_snapshot",
fake_get_aggregated_vessels_snapshot,
"get_current_vessels_snapshot",
fake_get_current_vessels_snapshot,
)
class _Result:
@@ -661,6 +744,8 @@ async def test_vessel_snapshot_filters_bbox_and_caps_limit(monkeypatch):
assert captured["bbox"] == (10.0, 59.0, 11.0, 60.0)
assert captured["limit"] == 5000
assert data["diagnostics"]["bbox_applied"] is True
assert data["diagnostics"]["source"] == "vessel_current_state"
assert data["diagnostics"]["current_state_count"] == 2
assert data["diagnostics"]["legacy_feature_count"] == 0
assert data["diagnostics"]["legacy_backfilled_mmsi"] == 0
finally:
@@ -668,34 +753,12 @@ async def test_vessel_snapshot_filters_bbox_and_caps_limit(monkeypatch):
@pytest.mark.asyncio
async def test_vessel_snapshot_uses_legacy_fallback_when_raw_window_is_empty(monkeypatch):
now = datetime(2026, 4, 28, 1, 0, tzinfo=timezone.utc)
async def test_vessel_snapshot_does_not_fallback_to_history_when_current_state_is_empty(monkeypatch):
monkeypatch.setattr(
visualization,
"get_aggregated_vessels_snapshot",
"get_current_vessels_snapshot",
AsyncMock(return_value=[]),
)
monkeypatch.setattr(
visualization,
"_load_legacy_vessel_snapshot_features",
AsyncMock(
return_value=[
{
"type": "Feature",
"id": 257123000,
"geometry": {"type": "Point", "coordinates": [10.73, 59.91]},
"properties": {
"mmsi": 257123000,
"name": "OSLO TRADER",
"vessel_type": 70,
"vessel_type_name": "Cargo",
"received_at": now.isoformat(),
},
}
]
),
)
result = await visualization.build_vessel_snapshot_response(
object(),
bbox=(10.0, 59.0, 11.0, 60.0),
@@ -705,12 +768,11 @@ async def test_vessel_snapshot_uses_legacy_fallback_when_raw_window_is_empty(mon
since_minutes=60,
)
assert result["count"] == 1
assert result["features"][0]["properties"]["name"] == "OSLO TRADER"
assert result["count"] == 0
assert result["diagnostics"]["raw_feature_count"] == 0
assert result["diagnostics"]["legacy_feature_count"] == 1
assert result["diagnostics"]["legacy_backfilled_mmsi"] == 1
assert result["diagnostics"]["legacy_fallback_used"] is True
assert result["diagnostics"]["legacy_feature_count"] == 0
assert result["diagnostics"]["legacy_backfilled_mmsi"] == 0
assert result["diagnostics"]["legacy_fallback_used"] is False
@pytest.mark.asyncio

View File

@@ -8,6 +8,228 @@ This project follows the repository versioning rule:
- `improvement` -> `+0.0.1`bugfix + 小功能混合)
- `bugfix` -> `+0.0.1`
## [0.74.3] — 2026-09-13
Released: 2026-09-13
### Highlights
- Ubuntu / WSL 新机器初始化会自动检测并准备 Docker Engine、Compose v2、Buildx 和当前用户权限,减少手工安装步骤。
- 数据库初始化先核对容器端口与后端真实连接,连接和认证通过后才创建表和默认数据。
### Added / Fixed / Improved
- 区分 Docker CLI 缺失、服务未安装、daemon 不可用及 socket 权限不足,修正未安装 Docker 时误提示启动 socket 的诊断。
- 自动补齐缺失的 Docker 依赖并启动本地服务,以原用户身份刷新 Docker 组权限;保留参数和 PATH不依赖 sg。
- 通过 Compose 同步已有 PostgreSQL / Redis 容器配置,保留端口冲突等具体错误;端口映射异常时最多保留数据卷重建一次 PostgreSQL。
- 新增后端数据库只读连接检查,对认证、库名和网络失败给出不含密码或完整连接串的诊断。
- 将 Docker 与数据库启动隔离回归测试接入快速检查,并同步 README、harness 和中英文运维说明。
---
## [0.74.2] — 2026-07-01
Released: 2026-07-01
### Highlights
- 收敛 agent harness 到 `rules.md``AGENTS.md``docs/HARNESS.md``.codex/skills/`,删除重复维护的旧 Claude command 入口。
- 强化视觉证据规则:截图或视觉引用路径打不开时必须先处理 WSL/Windows 路径、相对路径和附件位置,而不是跳过后猜测。
- 明确 OCR 可作为文本类视觉证据或非多模态环境 fallback同时要求布局、颜色、像素和渲染类问题保留真实视觉验证或明确限制说明。
### Added / Fixed / Improved
- `AGENTS.md` 替换旧 opencode/默认 Plan Mode 内容,保留最新单一入口和 harness 验证说明。
- `rules.md``docs/HARNESS.md` 同步 Visual Evidence Gate补齐路径解析、访问失败报告和 OCR fallback 边界。
- 删除 `.claude/commands/*` 中与 `.codex/skills/*` 重复的旧 cleanup/docs/goal-driven/release 入口,并更新文档受众计划中的旧路径引用。
---
## [0.74.1] — 2026-06-30
Released: 2026-06-30
### Highlights
-`/earth-content` 的品牌标识上传收敛到 `Logo 地址``标题图地址` 字段内,移除旧的全局“选择资产/上传”工具栏。
- 新增字段级图片拖拽反馈,拖到对应字段时直接提示将图片复制为 Logo 或标题图。
- 对齐品牌上传按钮到现有 Tactile UI primary 按钮样式,并同步中英文使用手册、快速开始和控制台上下文文档。
### Added / Fixed / Improved
- `BrandAssetInput` 支持字段内选择文件、拖拽上传、单字段 loading 和上传后回写草稿 URL。
- `FieldGrid` 支持按字段注入自定义输入控件,同时复用统一草稿提交路径。
- 品牌上传拖拽态改为低饱和 tactile 配色,上传按钮保持蓝色轻立体样式,容器内上下/右侧留白对齐为 3px。
- 补齐品牌上传相关 legacy UI 英文翻译、术语对照和用户文档。
---
## [0.74.0] — 2026-06-30
Released: 2026-06-30
### Highlights
- 扩展统一 i18n 到 Web Earth、控制台、认证页和公开 Docs 的更多动态入口,减少英文界面中文残留。
- 强化 Earth HUD 的通知胶囊、品牌栏、语言 switch、图例、tooltip、详情卡、新闻和 TV 文案展示,避免英文态裁切或错位。
- 将容易遗漏的 i18n 入口和视觉回归加入 harness让 smoke 覆盖通知位置、内容宽度、语言切换状态和动态文案。
### Added / Fixed / Improved
- 新增 Earth runtime i18n 入口,统一高清材质、启动状态、错误提示和图层状态文案来源。
- 补齐国家名、属性名、卫星 legend、详情页 tooltip、新闻/TV 默认文案和 API 错误提示的英文翻译与回退。
- 更新控制台与 Earth 布局规则,保留 brand 尺寸语义,同时让标题、副标题和通知胶囊按内容完整显示。
- 扩展 frontend smoke 与 harness 文档固化一屏高度链、i18n 动态入口、胶囊/tag overflow 和截图证据要求。
- 更新 Earth 新闻本地化服务与测试,确保英文界面新闻内容不再回退中文 UI 文案。
---
## [0.73.0] — 2026-06-29
Released: 2026-06-29
### Highlights
- 新增前端统一 i18n 基础设施让认证页、Docs UI、控制台外壳、导航、搜索和核心共享组件共用 `zh-CN` / `en-US` 语言状态。
- 控制台侧边栏偏好面板接入语言与主题切换,并修复一屏高度链、账号区、状态指示器和英文态文案裁切问题。
- 扩展 harness 与 smoke 覆盖,确保 admin shell 高度、移动/缩放布局、语言切换、搜索、Docs 和核心控制台交互在发布前被验证。
### Added / Fixed / Improved
- 新增 `frontend/src/i18n/`,用 `i18next` / `react-i18next` 维护资源、locale 映射、Docs 兼容和过渡期 legacy UI 翻译桥。
- 将 AdminLayout、route manifest、admin search、Auth、DataTable、Dialog、Toast、MarkdownRenderer 和 Users 页迁移到统一翻译资源。
- 补齐 Planet Content、Collected Data、System Logs、Datasources、Settings 和 Collection Management 等英文态残留翻译,并覆盖动态计数字符串。
- 改进控制台侧边栏账号区、语言 switch、状态 pill 自适应宽度和 admin shell overflow ownership避免首屏溢出和状态词裁切。
- 更新 i18n 计划、控制台前端上下文、harness 文档和规则,记录语言迁移边界、状态指示器布局约束和一屏验证要求。
---
## [0.72.0] — 2026-06-29
Released: 2026-06-29
### Highlights
- 将 agent 入口收敛到单一 `AGENTS.md`,并让 harness 明确阻止小写入口再次分叉。
- 新增完整本地 harness 验证层,覆盖 backend/frontend/docs/security 静态规则、前端 build 和 Playwright 路由/交互 smoke。
- 扩展 Earth News 与控制台 smoke确保新闻源测试、新增取消、手动新闻组创建、桌面/移动菜单和 zoom 布局都在发布前验证。
### Added / Fixed / Improved
- 新增 `scripts/harness/*` 规则检查、doctor、validate 和前端 smoke 脚本,并将未跟踪 harness 设施纳入发布。
- 清理 SpaceTrack 与 PeeringDB collector 的 stdout/debug 输出,改用结构化日志并移除 SpaceTrack 不可达重复 fetch 路径。
- 强化控制台布局、auth 表单、Docs 页面、Earth shell 和 Earth toolbar 的响应式与无障碍细节。
- 同步 README、CODEMAP、HARNESS、harness audit、用户手册、快速开始和开发者文档明确当前 Web Earth / React admin / FastAPI / aiprovider 边界。
- 将 backend、frontend、docs 和 Earth News 检查纳入 `scripts/harness/quick-check.sh``scripts/harness/validate.sh` 的稳定验证面。
---
## [0.71.1] — 2026-06-26
Released: 2026-06-26
### Highlights
- 修复 Earth 欧洲、美洲、中东与非洲等区域新闻被亚太来源和旧来源过滤饿死的问题,滚动条、面板和巡航重新回到同一批区域 payload。
- 将当前可见新闻和巡航新闻提升到目标位置/翻译优先队列,避免历史普通 Redis backlog 阻塞用户正在看的新闻精修。
- 新增 agent harness 入口、代码地图、验证脚本与双语技术说明,让后续维护能按现有 uv/Bun/Gitea 工作流检查而不替代项目规则。
### Added / Fixed / Improved
- EarthFeed 在全局和巡航队列中按区域轮转候选新闻,保留区域视图的“当前区域 + global”规则并补充回归测试。
- `news.js` 按区域、类型、来源和数量隔离并发刷新请求,丢弃旧区域响应;跨区域时不再复用旧来源筛选。
- 新闻展示在中文本地化未完成时回退原始标题和摘要,避免出现有内容却显示“新闻汉化中”的卡片。
- 新闻目标位置 worker 新增优先 stream、pending reclaim、任务超时和并发处理无效消息会确认并删除减少队列堆积。
- 补充 Earth 新闻源、Earth 前端结构、harness 和版本历史文档,并移除控制台 auth store 的调试日志。
---
## [0.71.0] — 2026-06-11
Released: 2026-06-11
### Highlights
- 将 Motion Agent 升级为可供 Web/UE 共用的双向控制服务,补齐真实 MediaPipe 识别 worker、设备控制、动作白名单和 WSL 摄像头开箱启动链路。
- 新增 Earth 手动新闻内容组、条目、导入与重处理能力,并改进按 locale 和启用来源进行的新闻补充与多样化。
- 对齐动捕模式下的 Earth 点击、详情锁定和卫星轨迹交互,同时完善启动脚本、测试 harness 与运维说明。
### Added / Fixed / Improved
- Motion Agent 支持命令结果、状态、骨架与手势事件,统一单路/双路输入配置,并随 `planet.sh` 默认启动;仓库内提供 usbipd-win fallback 安装包。
- Earth 新闻服务集中处理显示就绪判断和来源多样化,避免存储层与编排层重复筛选;新增手动新闻 API 与回归测试。
- 清理 Motion Agent 重复识别执行路径、前端不稳定随机 key 和过时计划描述,补齐 pytest 路径 harness、双语使用手册、快速开始与数据流文档。
---
## [0.70.0] — 2026-06-04
Released: 2026-06-04
### Highlights
- 新增后端枚举契约治理,将稳定协议状态集中到 `app/core/enums.py`,同时保持数据库和 API 的小写字符串兼容。
- 改进 Earth 新闻分类、重要度与 Breaking 插队链路,并补齐中英文新闻源与枚举契约文档。
- 将 Earth 船只展示改为 `vessel_current_state` 当前状态快照,保留 AIS 原始历史用于轨迹和态势分析。
- 清理错误的船只视口刷新/订阅思路,恢复全球船只显示,并让性能优化集中到批量渲染、关闭动态聚类和减少 hover/rebuild 开销。
### Added / Fixed / Improved
- `earth_news_classification.py` 集中管理新闻类型、重要度和 Breaking 规则,避免抓取编排层重复判断。
- `/api/v1/vessels/snapshot` 支持全球当前状态读取和小数 zoom诊断信息明确返回 `vessel_current_state` 来源。
- 船只前端使用全球 bbox + `limit=3000`,不再随相机视口重复请求或建立多视口 WebSocket 订阅。
- 中英文技术文档同步更新船只、采集器、数据流、渲染层级、样式参数和历史计划状态。
---
## [0.69.0] — 2026-06-03
Released: 2026-06-03
### Highlights
- 新增 Earth 新闻源治理能力,支持多 Feed 子项、源属性标签、新闻类型过滤、重要度规则和健康测试。
- 新增观测日志聚合视图,按 fingerprint 汇总 Earth、Admin 和服务端重复运行时事件,并保留原始发生明细。
- 改进 TV/HLS 播放恢复和代理重写,降低字幕、分片和源站波动导致的直播不可用噪声。
### Added / Fixed / Improved
- Earth 新闻面板和 UE 端统一通过 `/api/v1/news/earth-feed` 使用 `categories``locale` 服务端过滤Web 端新闻类型偏好仅保存在当前浏览器。
- 控制台日志页新增重复统计、原始日志和审计日志模式,前端上报器会合并短窗口内的重复错误并提交 `occurrence_count`
- AI Provider / 服务端运行时可通过受保护的 observability ingest 入口写入结构化事件。
- 新闻源文档新增中英文配置说明,并补齐 Earth 前端、控制台日志和公开 Docs 索引。
---
## [0.68.1] — 2026-05-28
Released: 2026-05-28
### Highlights
- 修复 CelesTrak active 未更新窗口下清库后无法恢复的问题fallback 会优先复用本地有效 group 缓存。
- 修复数据源任务队列“查看日志”无法按 `task_id=...` 命中数据库结构化日志的问题。
### Added / Fixed / Improved
- CelesTrak fallback group 列表改为公开可用分组,移除失效 group并在没有 active 缓存时仍可从本地 group 缓存恢复采集。
- 数据库日志搜索补充 JSON context 的 `key=value` 别名,支持 `task_id=26906``datasource_id=20` 这类控制台跳转查询。
- 补充 CelesTrak 缓存边界、数据源任务日志跳转和运维恢复说明的中英文文档。
---
## [0.68.0] — 2026-05-28
Released: 2026-05-28
### Highlights
- 新增数据源任务队列的实时指标校准和批量删除进度,让大表清理、取消和完成状态在控制台中可感知。
- 新增智能星球 interactable 可插拔聚类策略,支持稳定 3D 球面聚类、动态屏幕聚类和 250% 以上自动散开。
- 改进新设备启动流程,`planet.sh` 会在启动前同步前端依赖,避免缺失依赖导致控制台动态导入失败。
### Added / Fixed / Improved
- 数据源列表改为中文记录数指标,并对 AIS 大表使用统计估算 + 单条详情精确校准,降低首次加载成本。
- 数据删除任务改为分批删除并广播进度,清理 AIS 衍生表后自动 `ANALYZE`,同时修复取消中任务恢复和状态文案。
- Earth BGP、算力中心和 interactable 图层默认使用 `stable-spherical` 聚类,船舶实时层保留 `dynamic-screen`
- 新增中英文 Earth interactable clustering 文档,并补充采集队列、后端删除语义和前端依赖同步说明。
---
## [0.67.0] — 2026-05-27
Released: 2026-05-27
### Highlights
- 新增控制台日志实时跟随与前端运行时错误上报,帮助在控制台内直接排查 Admin / Earth 客户端异常。
- 重构智能星球 Interactable 聚合和 wheel 缩放输入,保持真实地理锚点稳定,同时让鼠标滚轮和触控板拥有各自合适的缩放手感。
### Added / Fixed / Improved
- 新增 `/ws` 日志 tail 通道、数据库/文件日志统一事件读取,以及 Admin `error` / `unhandledrejection` / React ErrorBoundary 上报链路。
- 优化控制台日志页状态颜色、跟随体验和运行时错误展示,并将 Admin 本地工具模块从 `lib` 命名迁移为局部 `utils`
- 修复智能星球国界/高清材质壳半径对齐、Interactable cluster 圆点显示、缩放目标累积和触控板连续缩放问题。
- 更新 `planet.sh``.gitignore`,避免新环境构建污染锁文件并移除前端 `lib` 目录特殊放行。
- 补充智能星球前端、渲染层级、控制台状态与运行时日志文档。
---
## [0.66.3] — 2026-05-26
Released: 2026-05-26

325
docs/HARNESS.md Normal file
View File

@@ -0,0 +1,325 @@
# Agent Harness
This harness improves discoverability, repeatability, and agent safety for the
existing Planet project. It does not replace current project rules, scripts, CI,
or release workflows.
## Authority And Conflicts
Existing project rules are authoritative:
1. `rules.md`
2. `AGENTS.md`
3. Current implementation docs under `docs/technical/`
4. Existing scripts, especially `planet.sh`
5. Existing Gitea workflow files under `.gitea/workflows/`
When harness guidance conflicts with any of the above, keep the existing rule,
do not overwrite the existing workflow, and add a compatibility note here or in
`docs/harness-audit.md`.
For frontend or documentation audits, also read the Rules Coverage Evidence
section in `docs/harness-audit.md`. It maps `rules.md` clauses to the current
static checks, Playwright smoke coverage, and remaining manual review areas, so
an agent can distinguish a proved harness pass from a rule that still needs
human-quality inspection.
When the user describes work with product words rather than module names, use
the `rules.md` **Agent Discovery Index** before deciding which modules to load.
It maps Chinese phrases such as `一屏`, `高度没控住`, `文档`, `数据源`,
`地球`, `模型供应商`, and `发版` to the required rule modules.
## Starting Work
Recommended startup flow:
```bash
git status --short
scripts/harness/doctor.sh
```
Then read only the relevant implementation docs:
- Backend/API/data work: `docs/technical/zh/backend-*.md` and matching English
docs when public docs are affected.
- Frontend/admin work: `docs/technical/zh/frontend-admin-frontend-context.md`.
- Earth work: `docs/technical/zh/earth-frontend-context.md`,
`docs/technical/zh/earth-render-layer-order.md`, and style docs when visual
semantics change.
- Operations work: `docs/technical/zh/ops-runbook.md` and
`docs/technical/zh/ops-planet-sh-startup.md`.
- AI Provider work: `docs/technical/zh/agents-aiprovider.md`.
- Documentation work: `docs/documentation-coverage-rules.md`.
Use focused inspection commands before broad reads:
```bash
rg -n "<symbol-or-term>" <path>
git diff --stat HEAD
git diff --name-only HEAD
git diff --unified=0 HEAD -- <path>
```
## Existing Commands
| Purpose | Command |
| --- | --- |
| First setup | `./planet.sh init` |
| Start local stack | `./planet.sh start` |
| Start with LAN access | `./planet.sh start --allow-lan` |
| Restart all services | `./planet.sh restart` |
| Restart one area | `./planet.sh restart -b`, `-f`, `-a`, or `-d` |
| Health check | `./planet.sh health` |
| Logs | `./planet.sh log`, `./planet.sh log -b`, `-f`, `-a`, or `-m` |
| Create local user | `./planet.sh createuser` |
| Destructive local reset | `./planet.sh destroy` |
| Backend smoke tests | `cd backend && uv run --frozen --group dev --project .. python -m pytest -s tests/test_api.py tests/test_realtime_sources.py -q` |
| Frontend build | `cd frontend && bun install --frozen-lockfile && bun run build` |
| Mock AIS WebSocket | `bun run mock:ais-ws` |
## Harness Commands
| Tier | Command | What It Does |
| --- | --- | --- |
| Doctor | `scripts/harness/doctor.sh` | Checks required files, required tools, optional delivery tools, and forbidden frontend lockfiles. |
| Security | `scripts/harness/security-check.sh` | Checks that environment/private-key files are not tracked and scans for high-confidence committed secret tokens. |
| Backend Rules | `scripts/harness/backend-rules-check.sh` | Checks backend app Python for direct `print()`, `breakpoint()`, and `pdb.set_trace()` debug calls so service code uses structured logging. |
| Frontend Rules | `scripts/harness/frontend-rules-check.sh` | Checks Bun-only scripts, admin route manifest coherence, literal internal route links, admin search route targets, frontend debug output, native button safety, icon-button accessibility, no nested Cards, no AntD/Space layout primitives, ConnectionTestInput usage, admin/docs shell height-chain sizing, same-category style owner warnings, viewport-scaled font sizes, zero letter spacing, and high-signal UI rule warnings. |
| Docs Consistency | `scripts/harness/docs-consistency-check.sh` | Checks frontend Docs metadata against backend Gatekeeper metadata, public Docs registration, full technical-doc bilingual file pairs, public doc links, readable link titles, language-scoped technical links, README/project-context admin stack drift, supported credential collector contracts, manual console route coverage against the actual admin manifest, documented UI route drift, documented `?section=` deep-link validity against the actual admin section config in technical docs and active plan docs, and the harness rules-coverage notes. |
| Quick | `scripts/harness/quick-check.sh` | Runs doctor, whitespace diff check, shell syntax checks, security scan, backend/frontend/doc consistency checks, and CI backend smoke tests. |
| Full | `scripts/harness/validate.sh` | Runs quick check, frontend Bun install/build, Playwright route smoke, optional Helm checks, and opt-in Docker image smoke builds. |
Docker image smoke builds are expensive and are off by default:
```bash
PLANET_HARNESS_DOCKER_SMOKE=1 scripts/harness/validate.sh
```
Frontend Playwright smoke runs by default in full validation after the frontend
build. It starts a local Vite preview and checks the `/` to Earth redirect,
public pages, unknown-route login fallback, protected admin route login
fallback, authenticated unknown-route fallback to `/admin`, Docs loading with
mocked API content, Docs detail page
language/theme/search interactions, every Docs catalog slug exposed by the
frontend/backend metadata, the Earth iframe entry point, login error handling,
register + email verification, password reset, standalone email verification,
and authenticated `super_admin` rendering for every admin route plus core
`section` deep links derived from the actual admin route and section config.
Authenticated admin
checks run at desktop size, mobile size, and 125% / 150% zoom; desktop and
mobile passes also fail on global horizontal overflow so table/detail panels
must keep overflow ownership inside their own scroll regions. To enforce the
existing `rules.md` `uiux` one-screen workspace rule, admin shell pages have a
hard rendered check: the shell must resolve to the viewport height through the
root 100% height chain, `#root`/document/body must not gain vertical overflow,
and the desktop sidebar account/preferences area must remain inside the first
viewport while the nav owns any excess scrolling. The smoke also
derives the sidebar menu from the actual admin route manifest and clicks every
visible `super_admin` menu entry on both desktop and mobile viewports, then
exercises safe interaction paths for admin search, section tabs, the AI settings
shortcut, logs view switching, user dialog opening, and data distribution toggles.
It also exercises Earth News source testing, add/cancel source draft behavior,
and manual news group creation against mocked `/earth/news-*` APIs.
Documented AI and collector
deep links such as `/ai?section=integrations`, `/ai?section=playground`, and
`/collection-management?section=collector_credentials` are part of the rendered
smoke surface:
```bash
PLANET_HARNESS_FRONTEND_SMOKE=0 scripts/harness/validate.sh
PLANET_HARNESS_FRONTEND_SMOKE_PORT=4174 scripts/harness/validate.sh
```
### Visual Evidence And Style Consistency
- Treat user-provided screenshots and images as primary visual evidence. If a
screenshot contradicts written text, inspect the image first and explicitly
call out the mismatch before deciding what to change.
- Path resolution is part of visual evidence handling. If a referenced
screenshot path cannot be opened, try reasonable local equivalents first:
WSL/Windows path conversion, workspace-relative paths, absolute paths, current
thread attachments, repository files, and obvious local attachment/download
locations.
- If the image still cannot be found or opened, report the exact path/access
blocker instead of guessing. Do not infer image content from the filename, alt
text, surrounding prose, logs, or memory.
- OCR is acceptable evidence for text-only questions or non-multimodal
environments; state when OCR was the fallback. Layout, color, spacing, pixel,
and rendering issues still require a real visual inspection or an explicit
"could not verify visually" note.
- Same-category UI surfaces must use one visual system per product area. Badges,
chips, pills, tags, status labels, small buttons, cards, panels, and toolbar
controls should reuse the shared component, shared token, or established CSS
owner for that area instead of introducing a page-local lookalike.
- `scripts/harness/frontend-rules-check.sh` warns when semantic
`badge` / `chip` / `pill` / `tag` / `status` selectors appear outside the
approved React and Earth CSS owner files. A warning means the reviewer should
either move the style into the shared owner or document why this is a genuinely
new visual family.
### Earth I18n Harness Rules
Earth i18n work must validate rendered behavior, not only static text lookup.
Agents often miss dynamic strings that are created after initial page load, so
the smoke treats these as first-class i18n surfaces:
- **Visible text and attributes**: translated checks must include `innerText`
plus `title`, `aria-label`, `placeholder`, and `alt`. Tooltips and icon-only
buttons are user-facing copy, not implementation details.
- **Dynamic detail cards**: info cards opened from Earth markers, cruise cards,
BGP markers, compute centers, vessels, and news must render field labels,
status values, source tags, action buttons, and disabled/tooltips in the active
language.
- **English content safety**: English mode must not fall back to Chinese news
titles, summaries, feed names, measure words, or generic status labels. If no
English localization exists, hide the item or use a neutral English fallback.
- **Brand assets**: locale switching must update both text and image assets.
The default Earth HUD brand uses `title-zh.png` for Chinese and `title-en.png`
for English while keeping the same top-left layout and logo position.
- **Controls and state**: switch/segmented-control visuals must follow the real
checked/pressed state after both direct clicks and programmatic panel changes.
A control is not valid if the state changes but the thumb, active pill, or
`aria-*` state stays stale.
- **Runtime copy entrypoints**: dynamic status, loading, startup, and error copy
must enter through `earthMessage(...)` plus the centralized
`EARTH_MESSAGE_TEMPLATES` map. Do not hide direct strings behind
`showStatusMessage`, `queueStatusMessage`, `showGestureStatusMessage`,
`showError`, `setLoadingMessage`, `resolveStartupMessage`, `startupMessage`,
or `earth:status` events.
- **Capsules and tags**: pills, tags, chips, badges, and small buttons must not
overflow their panel. Prefer a slightly wider owning panel for important
status information; otherwise use `min-width: 0`, wrapping, or ellipsis with a
translated tooltip.
Current frontend smoke explicitly covers the Earth English locale flow: brand
image swap, settings language controls, panel switch visual sync, English news
filtering, English detail-card text and tooltips, and TV default/source labels.
## Environment Requirements
Required for normal development:
- `zsh` for `planet.sh`
- `uv` for Python dependency and test execution
- `bun` for frontend dependency and build execution
- Python resolved by `uv` from the root `pyproject.toml`
Harness command lookup first checks the current non-interactive `PATH`. If a
required tool is not visible there, `scripts/harness/lib.sh` asks the user's
login interactive shell (`$SHELL`, then `zsh`, then `bash`) for the command
path. This avoids hardcoding a dotfile while still covering agent environments
that do not inherit the user's normal shell setup.
Required for full local stack operation:
- Docker and Docker Compose
- PostgreSQL and Redis containers started by `planet.sh`
Optional for delivery smoke:
- Docker daemon for image builds
- Helm for chart lint/template checks
For routine harness validation, do not install missing system software
automatically. Report the gap and point to the explicit bootstrap entry points.
`./planet.sh init` can install missing Docker Engine, Compose v2, and Buildx on
Ubuntu / Ubuntu WSL, start the local service, and configure Docker group access.
This bootstrap behavior is intentional; do not invoke it merely to make harness
checks pass. `scripts/bootstrap-dev.sh` only prepares application dependencies.
Docker bootstrap regression checks use isolated command stubs and never install
packages or modify the host daemon:
```bash
uv run --frozen --project . python scripts/harness/test_docker_bootstrap.py
uv run --frozen --project . python scripts/harness/test_database_startup.py
```
Database startup regressions also run in quick-check. They cover Compose
reconciliation of existing containers, visible startup errors, published-port
checks, bounded recreation that preserves volumes, and the backend connection
gate before schema initialization. Their command stubs and driver mocks do not
modify the host Docker environment.
## What Agents Must Not Change Automatically
- Do not replace Bun with npm, pnpm, or yarn.
- Do not migrate CI from `.gitea/workflows/` to `.github/workflows/`.
- Do not rewrite `planet.sh` lifecycle behavior as a parallel script.
- Do not run `./planet.sh destroy` unless explicitly requested.
- Do not commit `.env`, secrets, private keys, logs, or generated build output.
- Do not add external integrations, hooks, or new dependency managers just to
satisfy harness structure.
- Do not publish internal harness docs into the product Docs UI unless a
maintainer explicitly asks for it.
## Hooks And Reminders
No automatic hooks are installed in this phase. Manual reminders:
- Run `scripts/harness/quick-check.sh` before handing off small changes.
- Run `scripts/harness/validate.sh` before larger cross-subsystem changes.
- Run `scripts/harness/security-check.sh` after touching config, auth,
credentials, docs examples, or generated fixtures.
- Run `scripts/harness/backend-rules-check.sh` after backend service edits to
catch direct stdout/debugger calls before they reach runtime logs.
- Run `scripts/harness/frontend-rules-check.sh` after frontend edits to expose
route, package-manager, debug-output, and UI rule warnings.
- Run `scripts/harness/docs-consistency-check.sh` after docs edits or feature
route changes.
- Add focused tests before modifying backend service behavior or frontend
workflows.
- For docs changes, run the checks listed in
`docs/documentation-coverage-rules.md`.
## Reusable Workflows
### Feature Work
1. Read `rules.md` modules for the touched area.
2. Check `CODEMAP.md` for entry points and ownership boundaries.
3. Inspect existing tests and docs before editing.
4. Make the smallest behavior-preserving or feature-scoped change.
5. Run `scripts/harness/quick-check.sh` or a narrower documented command.
6. Update relevant docs when behavior, workflow, or operations change.
7. For rendered frontend changes, verify the affected route with Playwright or
the full harness smoke, because `bun run build` alone does not prove page
usability.
### Bug Fix
1. Reproduce with a focused test or command.
2. Patch the owning module, not a caller-side workaround.
3. Run the focused regression test.
4. Run `scripts/harness/quick-check.sh` when the change is safe to validate
locally.
### Documentation Change
1. Read `docs/documentation-coverage-rules.md`.
2. Route docs by audience: UI users, operations, or second-party developers.
3. Keep Chinese and English technical docs paired by filename; public Docs also
need matching frontend/backend metadata when exposed in the product Docs UI.
4. Run the repository-specific docs checks that match the changed files.
### Release Or Delivery Change
Use the existing release skill/workflow and `.gitea/workflows/` files. Harness
validation can smoke-check Helm and Docker locally, but it must not replace the
release process.
## Implementation Notes
- `docs/harness-audit.md` records the discovery pass that led to this harness.
- `AGENTS.md` is the single authoritative agent guide. The older lowercase
`agents.md` entry has been merged into it and should remain absent.
- `CODEMAP.md` is intentionally high level; deeper subsystem docs stay in
`docs/technical/{zh,en}/`.
- `scripts/harness/frontend-smoke.mjs` is a lightweight route/section smoke
with mocked API data. It proves route shells, auth guards, and primary admin
sections render, but it is not a replacement for feature-specific browser QA
against a real backend.
- Frontend smoke prints phase-level progress by default. Use
`PLANET_FRONTEND_SMOKE_PROGRESS=verbose` to print each route/menu/doc item
when diagnosing a slow or failing smoke run, or set it to `0` to suppress
progress lines.

161
docs/harness-audit.md Normal file
View File

@@ -0,0 +1,161 @@
# Harness Audit
Last audited: 2026-06-26
This audit records the repository state used to add the agent harness. It is a
compatibility note, not a replacement for existing rules or architecture docs.
## Existing Commands
| Area | Existing Command | Notes |
| --- | --- | --- |
| Bootstrap | `./planet.sh init` | Syncs uv/Bun dependencies, creates missing env files, starts data services, seeds defaults. |
| Start | `./planet.sh start` | Starts backend, frontend, AI Provider, PostgreSQL/Redis, and Motion Agent when available. |
| LAN start | `./planet.sh start --allow-lan` | Opens frontend/backend/AI Provider ports and requests Windows firewall/port cleanup when needed. |
| Restart | `./planet.sh restart` | Supports scoped restart flags for backend, frontend, AI Provider, database, and Motion Agent. |
| Health | `./planet.sh health` | Checks containers, backend `/health`, AI Provider `/health`, frontend, and Motion Agent state. |
| Logs | `./planet.sh log` | Supports backend, frontend, AI Provider, and Motion Agent log views. |
| User fallback | `./planet.sh createuser` | Interactive emergency/local account creation. |
| Destructive reset | `./planet.sh destroy` | Requires confirmation and removes Planet-owned Docker/build/runtime state. Not a validation command. |
| Backend CI smoke | `cd backend && uv run --frozen --group dev --project .. python -m pytest -s tests/test_api.py tests/test_realtime_sources.py -q` | Mirrors `.gitea/workflows/ci.yaml`. |
| Frontend build | `cd frontend && bun install --frozen-lockfile && bun run build` | Bun-only workflow. |
| Root helper | `bun run mock:ais-ws` | Runs `scripts/mock-ais-ws-server.ts` from the root package. |
## Existing Agent Instructions
| File | Status | Notes |
| --- | --- | --- |
| `AGENTS.md` | Present | Single authoritative agent behavior guide. It references `rules.md`, `project_context.md`, harness validation, and high-risk areas. |
| `rules.md` | Present | Mandatory modular rules. Always load `core`, `security`, and `workflow`; load topic modules as needed. |
| `project_context.md` | Present | Static context. Some roadmap-era stack details are older than the current README/docs. |
| `.claude/commands/*.md` | Present | Existing command docs for cleanup, docs, goal-driven, and release workflows. |
| `.codex/skills/*.md` | Present | Existing local skills for cleanup, docs, goal-driven, and release. |
## Existing CI Gates
The repository uses `.gitea/workflows/`, not `.github/workflows/`.
| Workflow | Gate |
| --- | --- |
| `.gitea/workflows/ci.yaml` | Backend smoke tests, frontend Bun build, Docker build smoke, Helm lint/template. |
| `.gitea/workflows/release.yaml` | Builds and pushes frontend, backend, and AI Provider images on main/tag/manual release events. |
| `.gitea/workflows/deploy-staging.yaml` | Deploys Helm release to staging and runs curl smoke tests inside the cluster. |
## Existing Docs And Architecture Maps
| Area | Docs |
| --- | --- |
| Current architecture and startup | `README.md` |
| Technical docs index | `docs/technical/zh/README.md`, `docs/technical/en/README.md` |
| Documentation rules | `docs/documentation-coverage-rules.md` |
| Operations | `docs/technical/zh/ops-runbook.md`, `docs/technical/en/ops-runbook.md` |
| Startup internals | `docs/technical/zh/ops-planet-sh-startup.md`, `docs/technical/en/ops-planet-sh-startup.md` |
| AI Provider | `docs/technical/zh/agents-aiprovider.md`, `docs/technical/en/agents-aiprovider.md` |
| Frontend admin | `docs/technical/zh/frontend-admin-frontend-context.md`, `docs/technical/en/frontend-admin-frontend-context.md` |
| Earth rendering | `docs/technical/zh/earth-frontend-context.md`, `docs/technical/zh/earth-render-layer-order.md`, `docs/technical/zh/earth-layer-style-reference.md` |
| Plans and history | `docs/plans/README.md`, `docs/deprecated/README.md` |
## Release And Deploy Process
- Release workflow is documented in `.codex/skills/release/SKILL.md` and
`.claude/commands/release.md`.
- Version-bearing files include `VERSION`, `frontend/package.json`,
`pyproject.toml`, `uv.lock`, `docs/CHANGELOG.md`, and
`docs/version-history.md`.
- Delivery automation lives in `.gitea/workflows/release.yaml` and
`.gitea/workflows/deploy-staging.yaml`.
- Helm chart entry point is `deploy/helm/planet/Chart.yaml`.
## Missing Or Unclear Areas
- The older lowercase `agents.md` entry has been merged into uppercase
`AGENTS.md` so coding agents and harness tools use one source of truth.
- `project_context.md` originally included older roadmap assumptions such as
Celery, Kafka, TimescaleDB, MinIO, and UE5 as active stack elements. The
harness pass updated it to separate active stack facts from future directions;
current code and technical docs still remain authoritative when details drift.
- No safe automatic hook system was already configured. This phase documents
manual reminders instead of adding hooks.
- `.github/workflows/` is absent by design; CI is under `.gitea/workflows/`.
## Conflicts And Preserved Rules
| Conflict Or Tension | Resolution |
| --- | --- |
| Prompt suggested `AGENTS.md`; repository already had `agents.md`. | Merged the lowercase guide into uppercase `AGENTS.md`; harness doctor now requires `AGENTS.md` and keeps `agents.md` absent to prevent split authority. |
| Harness validation could duplicate CI. | Added wrapper scripts that call existing commands and mirror current CI gates where practical. |
| Full Docker smoke builds are expensive locally. | Kept them opt-in with `PLANET_HARNESS_DOCKER_SMOKE=1`. |
| Internal harness docs could clutter public Docs UI. | Kept `docs/HARNESS.md` and `docs/harness-audit.md` as repository docs, not product Docs entries. |
| Existing frontend toolchain is Bun-only. | Harness scripts and docs use Bun only and flag npm/pnpm/yarn lockfiles as failures. |
| Agents often miss user-installed Bun or uv in non-interactive shells. | Added `scripts/harness/lib.sh` to resolve tools from current `PATH` first and then the user's login interactive shell without hardcoding a dotfile. |
| Always-loaded security rules had no standalone harness gate. | Added `scripts/harness/security-check.sh` to block tracked `.env` / key files and scan for high-confidence committed private keys or provider tokens; quick-check now runs it. |
| Build success does not prove frontend page usability. | Added static frontend rules/doc checks and a Playwright route smoke for public pages, protected admin fallback, Docs loading and detail interactions, Earth iframe entry, login/register/verification/password-reset interactions, authenticated admin route/section rendering with mocked API data across desktop, mobile, and 125% / 150% zoom, plus manifest-derived desktop/mobile menu navigation and safe search/tab/dialog/Earth News interactions. |
| Route fallback behavior can regress even when every named page renders. | Extended the frontend smoke to verify `/` redirects to Earth, unauthenticated unknown routes show the login page, and authenticated unknown routes navigate back to `/admin`. |
| Frontend smoke route lists can drift from `AdminRoutes` and resource-page sections. | Updated the smoke to derive protected route checks and authenticated section deep-link checks from `AdminRoutes.tsx` and `PlainResourcePages.tsx`, including redirect-only `/alerts`. |
| Docs smoke mocks can drift from the product Docs catalog. | Updated the frontend smoke to derive mocked Docs catalog/content from `frontend/src/pages/Docs/docs-content.ts` plus backend Gatekeeper access metadata, then open every Chinese Docs catalog slug. |
| User manuals can miss a real console menu entry after route changes. | Added a docs consistency check that compares the manual console overview tables with `frontend/src/admin/routes/manifest.tsx`; fixed the missing `/docs` row in both user manuals. |
| Rendered pages can still contain broken internal shortcuts. | Added literal internal route-link checks and an interaction smoke for the AI settings shortcut; this caught and fixed a stale `/admin/settings` link that should point to `/settings`. |
| Global search entries can drift because their route targets live in data objects rather than JSX links. | Added a frontend rules check that validates every admin search `routePath` against the actual frontend route set. |
| Responsive styling fixes can satisfy one viewport by breaking the no-viewport-font rule. | Added a frontend rules failure for `font-size` values that use viewport or container query width units, and replaced public auth shell `vw` font sizing with fixed desktop/mobile sizes. |
| Typography polish can accidentally reintroduce squeezed non-zero letter spacing. | Normalized active frontend `letter-spacing` values to `0` and made the frontend rules check fail non-zero `letter-spacing` / `letterSpacing` declarations, with only inherit/default-zero forms allowed. |
| Native buttons can accidentally submit forms or keep controls clickable while loading after a props-spread reorder. | Added a frontend rules failure for TSX `<button>` elements without explicit `type` and for buttons whose `disabled` state can be overridden by a later props spread; fixed the data distribution buttons and auth button disabled ordering. |
| Admin/docs shell layouts can reintroduce brittle viewport sizing after a responsive fix. | Changed the admin and Docs route shells to use the existing `html/body/#root` 100% height chain, and added a frontend rules failure for exact `100vh` / `100vw` shell sizing in those CSS files. |
| Compact workspaces can drift back into card-in-card layouts or implicit AntD `Space` wrappers. | Added frontend rules failures for nested `Card` components, AntD imports, and `<Space>` layout primitives in active frontend source. |
| Connection-test controls can drift back into detached toolbar buttons. | Added a shared `ConnectionTestInput` suffix pattern for AI Provider and WebSearch Base URL fields, disabled WebSearch configuration/test controls when the tool is off, and made the frontend rules check fail detached AI/WebSearch connection-test buttons. |
| Same-category UI styles can fragment into page-local lookalikes. | Added same-category style owner warnings for semantic `badge`, `chip`, `pill`, `tag`, and `status` CSS selectors outside the approved shared React and Earth CSS owner files. |
| Visual fixes can go wrong when agents guess from missing screenshot paths. | Documented screenshots and images as primary visual evidence: if the path is not available, agents must search alternate attachment/local locations or report the blocker instead of inferring image content. |
| Public docs can reference stale admin section URLs. | Added docs consistency validation for documented `?section=` links and rendered smoke coverage for documented AI / collector deep links. |
| Active plan docs can preserve old admin deep-link assumptions after the technical docs are corrected. | Extended docs consistency checks to active `docs/plans/*.md` files for stale admin tab-query terms and actual `?section=` validity; corrected the docs audience split plan to current section routes. |
| Top-level README can drift from the actual frontend stack while technical docs stay current. | Updated README from Ant Design Pro to Tactile UI / Radix primitives / lucide-react and added README stale admin-stack terms to docs consistency checks. |
| Agent background context can reintroduce inactive stack assumptions. | Updated `project_context.md` and the root agent guide to label current stack facts versus future directions, then added exact stale-stack patterns for them to docs consistency checks. |
| Docs `?section=` validation can drift if the harness owns its own route/section table. | Changed the docs consistency check to derive section keys from `AdminRoutes.tsx` and `PlainResourcePages.tsx` resource configs before validating documented deep links. |
| Public Docs can drift between frontend catalog metadata and backend Gatekeeper authorization metadata. | Added a docs consistency check that compares filename, slug, group, order, and bilingual titles across both metadata sources; aligned existing order drift for toolbar overlay and location pipeline docs. |
| Non-public technical docs can silently become Chinese-only or English-only. | Added a full `docs/technical/{zh,en}` filename-pair check so every technical Markdown file has a same-named counterpart before docs consistency passes. |
| Credentialed collector docs can drift from backend support wiring. | Added docs consistency validation for every built-in collector marked `requires_credentials=true` and `credential_status=supported`: it must have a provider, default credential guide, supported connectivity provider, frontend credential UI/guidance, a regression test, and zh/en connectivity documentation. |
| Backend collectors can leak debug output or credential-adjacent context through stdout. | Replaced SpaceTrack and PeeringDB collector `print()` calls with structured logger events, removed unreachable duplicate SpaceTrack fetch code, and added `scripts/harness/backend-rules-check.sh` to block future backend app `print()`, `breakpoint()`, or `pdb.set_trace()` calls. |
## Rules Coverage Evidence
This matrix records how the current harness checks the `rules.md` modules that
matter for this frontend and documentation pass. "Automated" means the listed
command fails when the rule regresses. "Smoke" means the rendered product route
or interaction is opened with Playwright. "Manual" means the rule is still a
judgment call and must be inspected during review.
Before using this matrix, start from `rules.md`'s **Agent Discovery Index** when
the user describes work with Chinese/product terms instead of module names. The
index is the routing layer; this table is the coverage/evidence layer.
| `rules.md` Area | Rule Surface | Harness Evidence | Remaining Review |
| --- | --- | --- | --- |
| `core` | Remove stale transitional paths, duplicated helpers, and naming drift after large changes. | `scripts/harness/docs-consistency-check.sh` blocks known stale stack terms, old `?tab=` links, public Docs metadata drift, and README/project context drift. `scripts/harness/frontend-rules-check.sh` blocks repeated detached AI/WebSearch connection-test buttons by requiring `ConnectionTestInput`. | Naming quality, function size, and whether a new abstraction is worth keeping remain manual review items. |
| `core` | Keep one source of truth for route, Docs, and section state. | Frontend route, admin manifest, admin search targets, Docs catalog metadata, backend Gatekeeper metadata, manual route tables, and documented `?section=` links are all parsed from source and compared by `frontend-rules-check.sh`, `docs-consistency-check.sh`, and `frontend-smoke.mjs`. | Business-state ownership inside feature components still needs focused review when behavior changes. |
| `security` | Do not commit secrets, tracked env files, private keys, or exposed tokens. | `scripts/harness/security-check.sh` fails on tracked `.env` / private-key files and high-confidence provider tokens. `backend-rules-check.sh` blocks backend stdout/debugger calls, and `frontend-rules-check.sh` fails frontend console output that includes token material. | Whether a newly added setting should be masked or stored server-side still requires feature-specific review. |
| `workflow` | Frontend package management must stay Bun-only. | `scripts/harness/doctor.sh` and `frontend-rules-check.sh` fail forbidden frontend lockfiles and `npm` / `pnpm` / `yarn` script usage. `validate.sh` uses Bun for install, build, preview, and smoke. | New dependency legitimacy and maintenance quality are manual unless a dependency is actually added. |
| `workflow` | Agents should find `bun` and `uv` even when non-interactive `PATH` is incomplete. | `scripts/harness/lib.sh` checks the current `PATH`, then asks `$SHELL`, `zsh`, and `bash` login interactive shells for the command path without hardcoding a dotfile. `doctor.sh`, `quick-check.sh`, and `validate.sh` all source it. | System package installation remains outside harness scope and should be reported instead of auto-fixed. |
| `docs` | Keep public Docs whitelist-driven and synchronized with backend authorization metadata. | `docs-consistency-check.sh` compares frontend Docs metadata against backend Gatekeeper metadata, verifies files exist for both languages, checks public link titles, and blocks missing zh/en technical doc pairs. `frontend-smoke.mjs` opens every Chinese Docs catalog slug plus detail/search/language/theme interactions. | Quality of prose, examples, and whether a doc should be public are still editorial review items. |
| `docs` | User manuals must match real console routes and deep links. | `docs-consistency-check.sh` compares manual console tables with `frontend/src/admin/routes/manifest.tsx` and validates documented `?section=` links from actual `AdminRoutes.tsx` plus `PlainResourcePages.tsx` section config. | Screenshots and UI-copy nuance are not exhaustively validated. |
| `uiux` | Admin pages are compact single-screen workspaces with explicit overflow ownership. | `frontend-rules-check.sh` fails missing admin shell height-chain declarations (`.admin-theme-root`, `.admin`, `.admin__sider`, `.admin__nav-scroll`, `.admin__account`, `.admin__content`, `.admin__content-inner`), warns on suspicious `overflow: hidden`, and blocks exact `100vh` / `100vw` shell sizing in admin/Docs CSS. `frontend-smoke.mjs` checks every admin route at desktop/mobile and verifies `.admin` equals viewport height, `#root`/document/body have no vertical overflow, and desktop sidebar account/preferences stay in the first viewport. Zoom passes still cover 125% / 150% rendering. | Visual density, hierarchy, and whether a scroll owner feels ergonomic remain manual QA. |
| `uiux` | Controls use expected patterns and accessible icon buttons. | `frontend-rules-check.sh` blocks icon `Button` without `aria-label` and `title`, native `<button>` without explicit `type`, nested Cards, AntD imports, `<Space>`, and detached connection-test buttons. Smoke exercises search, tabs, dialogs, data toggles, and connection-test actions. | Native buttons with visible text are not treated as icon-only by static checks; semantics still need review when adding custom controls. |
| `uiux` | Same-category visual surfaces should use one style system per product area. | `frontend-rules-check.sh` warns when semantic `badge`, `chip`, `pill`, `tag`, or `status` selectors appear outside approved shared React and Earth CSS owner files. `docs/HARNESS.md` also makes screenshots/images primary evidence for visual fixes and forbids guessing when an image path is unavailable. | Final visual cohesion across screenshots still needs human/Playwright review, especially for page-specific cards, panels, and toolbar controls that static selector checks cannot classify perfectly. |
| `uiux` | Text should fit, avoid viewport-scaled font sizes, and keep letter spacing at zero. | `frontend-rules-check.sh` fails viewport/container-width font-size units and non-zero `letter-spacing` / `letterSpacing`. `frontend-smoke.mjs` checks rendered routes for global overflow across desktop/mobile. | Per-element text clipping without page-level overflow is not exhaustively detected and needs visual review for changed screens. |
| `frontend` | Keep shared behavior in reusable components and existing project patterns. | `frontend-rules-check.sh` enforces shared `ConnectionTestInput`, route/link/search consistency, no debug output, native button safety, Tactile/Radix/lucide direction instead of AntD/Space, and whitelist-driven public Docs. `bun x tsc --noEmit` and `bun run build` verify TypeScript/build health. | Broad casts, inline styles, and overflow issues are warnings when context may be legitimate; review changed lines before accepting them. |
| `frontend` | Responsive adaptations must preserve the primary action path. | `frontend-smoke.mjs` clicks every visible admin menu entry on desktop and mobile, opens protected routes unauthenticated and authenticated, verifies root/unknown route fallback, and exercises core auth flows. | Deep feature workflows beyond smoke data, such as destructive or long-running actions, require targeted tests before behavior changes. |
| `earth` | Earth render work needs real rendering checks. | Full smoke opens `/earth` and verifies the `3D Earth` iframe entry point. Earth News settings routes are included through the admin manifest/menu smoke, mocked `/earth/news-*` API responses, source test, add/cancel source draft, and manual news group creation checks. The broader Earth-specific layer/depth rules remain in `rules.md` and Earth docs. | The harness still does not claim full 3D layer visual verification; layer-depth and picking changes need targeted browser/canvas QA. |
## Harness Files Added
| File | Purpose |
| --- | --- |
| `AGENTS.md` | Single authoritative agent guide and coding-agent entry point. |
| `docs/HARNESS.md` | Harness workflow, validation tiers, conflict policy, and manual reminders. |
| `CODEMAP.md` | High-level codebase map and validation references. |
| `scripts/harness/lib.sh` | Shared command lookup and run helpers. |
| `scripts/harness/doctor.sh` | Environment and repository-shape check. |
| `scripts/harness/security-check.sh` | High-confidence secret and tracked environment/key file check. |
| `scripts/harness/backend-rules-check.sh` | Backend app debug-call guard for direct stdout/debugger usage. |
| `scripts/harness/frontend-rules-check.sh` | Bun-only, route manifest, literal internal link, admin-search route target, debug-output, native-button safety, icon-button accessibility, Card nesting, AntD/Space avoidance, ConnectionTestInput, admin shell one-screen height-chain declarations, admin/docs shell viewport sizing, same-category style owner warnings, viewport-font, zero-letter-spacing, and UI rules static check. |
| `scripts/harness/docs-consistency-check.sh` | Frontend/backend Docs metadata alignment, public Docs metadata, full technical-doc bilingual pair, link-title, language-scoped technical link, supported credential collector contracts, manual console route coverage, documented route, admin-config-derived section deep-link consistency, and harness rules-coverage note check. |
| `scripts/harness/frontend-smoke.mjs` | Playwright route, Docs detail/language/theme/search interaction, public auth form interaction, desktop/mobile/zoom rendering, admin shell one-screen/overflow checks, safe admin navigation/search/tab/dialog/Earth News interactions, and authenticated admin route/section smoke for the built frontend preview. |
| `scripts/harness/quick-check.sh` | Fast deterministic local validation. |
| `scripts/harness/validate.sh` | Full local validation wrapper with optional delivery smoke. |

View File

@@ -22,6 +22,7 @@
当前重点入口:
- [控制台 i18n 接入计划](/home/ray/dev/linkong/planet/docs/plans/admin-console-i18n-plan.md)
- [Earth Mobile Drawer UI Plan](/home/ray/dev/linkong/planet/docs/plans/earth-mobile-drawer-ui-plan.md)
- [Earth Compute Center BGP Style Plan](/home/ray/dev/linkong/planet/docs/plans/earth-compute-center-bgp-style-plan.md)
- [Earth Renderer Architecture Separation Plan](/home/ray/dev/linkong/planet/docs/plans/earth-renderer-architecture-separation-plan.md)
@@ -33,6 +34,7 @@
- [Earth News Cruise Summary Plan](/home/ray/dev/linkong/planet/docs/plans/earth-news-cruise-summary-plan.md)
- [Earth 动作捕捉手势控制计划](/home/ray/dev/linkong/planet/docs/plans/earth-motion-capture-gesture-control-plan.md)
- [Earth 动捕交互语义 V2 计划](/home/ray/dev/linkong/planet/docs/plans/earth-motion-gesture-interaction-v2-plan.md)
- [Motion Agent v2 控制协议与 3D 标定路线](/home/ray/dev/linkong/planet/docs/plans/motion-agent-v2-control-protocol-plan.md)
- [Earth Presentation 解耦架构计划](/home/ray/dev/linkong/planet/docs/plans/earth-presentation-decoupled-architecture-plan.md)
- [Earth Vessel Rendering Performance Plan](/home/ray/dev/linkong/planet/docs/plans/earth-vessel-rendering-performance-plan.md)
- [AIS 多源采集、冲突记录与聚合接口计划](/home/ray/dev/linkong/planet/docs/plans/earth-vessel-ais-aggregation-plan.md)

View File

@@ -0,0 +1,62 @@
# 控制台 i18n 接入计划
**状态**:基础设施已落地,大型业务页迁移继续进行
**创建日期**2026-06-29
**核心目标**:把 Docs 已有的中英文文档能力提升为前端统一 i18n 体系让未登录认证页、Docs UI、控制台外壳、导航、搜索和核心工作台文案共用同一个语言状态。
## 背景
Docs 站点已经有 `zh` / `en` 文档目录、Gatekeeper 权限和 `/api/v1/docs/{lang}/{slug}` 内容接口,但语言状态只保存在 `docs-lang`不影响控制台。控制台页面、搜索索引、toast、dialog、表格和认证页仍以中文硬编码为主导致用户切到英文文档后控制台仍是中文。
本计划把前端语言偏好收敛到 `planet-locale`,默认 `zh-CN`,支持 `en-US`。Docs 继续使用后端现有 `zh` / `en` 文档接口,通过前端映射与全局 locale 对齐。
## 设计决策
- 使用 `i18next``react-i18next` 作为统一 i18n 层避免长期维护自研插值、hook 和资源加载逻辑。
- 前端统一语言枚举为 `zh-CN` / `en-US`Docs 请求继续转换为 `zh` / `en`,新闻接口继续使用已有 `zh-CN` / `en-US` 口径。
- 语言偏好首版只保存在浏览器 `localStorage`,不新增后端用户设置字段。
- `docs-lang` 保留为兼容读取和写入项,让已访问过 Docs 的浏览器能平滑迁移。
- 静态路由、导航、搜索目标和通用组件使用显式翻译 key大型业务页在迁移期间通过 legacy UI 翻译桥补足常见硬编码文案。
## 分期
### P1统一语言基础设施
-`frontend/src/i18n/` 下维护 locale 类型、资源、初始化和 `useLocale()`
-`frontend/src/main.tsx` 里初始化 i18n并同步 `document.documentElement.lang`
- 在认证页和控制台侧边栏偏好面板提供语言切换入口。
### P2高复用界面迁移
- 迁移 Docs UI、AdminLayout、route manifest、admin search、Auth、DataTable、Dialog、Toast 和 MarkdownRenderer。
- 搜索索引按当前语言展示,同时保留中英文关键词以免降低可发现性。
- 用户管理页作为独立业务页示范迁移表头、按钮、toast、校验提示、角色和 Gatekeeper 标签。
当前已完成统一 `planet-locale`、Docs 兼容映射、认证页和控制台外壳语言入口、共享组件 key 化,以及 legacy UI 翻译桥。后续工作集中在把大型业务页从过渡桥迁移到显式 key。
### P3大型业务页收敛
- 分批把 Dashboard、DataList、Logs 和 PlainResourcePages 的配置块改为显式翻译 key。
- 过渡期保留 legacy UI 翻译桥,只处理 admin/auth 容器里的精确静态文本和属性。
- 业务数据、日志原文、API 字段名、provider id、命令和 Markdown 正文不走 legacy 翻译桥。
### P4移除过渡桥
-`rg -n "[\\p{Han}]" frontend/src/admin frontend/src/pages frontend/src/components` 只剩业务数据示例、中文文档标题或必须保留的中文品牌词时,删除 legacy UI 翻译桥。
- 增加 key 完整性检查,确保 `zh-CN``en-US` 资源结构一致。
## 验证
- `cd frontend && bun run build`
- `scripts/harness/frontend-rules-check.sh`
- `scripts/harness/docs-consistency-check.sh`
- `scripts/harness/quick-check.sh`
- 前端 smoke 需要覆盖登录页、Docs、Admin 侧边栏语言切换、侧边栏和搜索结果在中英文下渲染。
## 相关文件
- `frontend/src/i18n/`:统一 locale、资源和过渡桥。
- `frontend/src/pages/Docs/Docs.tsx`Docs 语言状态改为读取全局 locale。
- `frontend/src/admin/components/layout/AdminLayout.tsx`:控制台侧边栏语言切换、导航和搜索文案。
- `frontend/src/admin/search/indexers.ts`Admin 搜索目标本地化。
- `docs/technical/{zh,en}/frontend-admin-frontend-context.md`:当前实现上下文。

View File

@@ -2,7 +2,8 @@
**状态**:待实施
**创建日期**2026-05-12
**核心目标**`docs/technical/{zh,en}/manual.md` 拆成"纯客户视角"的使用手册,把 `planet.sh`、日志、LAN、故障排查这类运维内容迁到独立 `ops-runbook.md`,并把分层规则写进 `documentation-coverage-rules.md``.claude/commands/docs.md`,让以后写文档时自动按受众归档
**校正日期**2026-06-26控制台深链已从旧 tab 查询口径更新为当前 `?section=` 口径
**核心目标**:把 `docs/technical/{zh,en}/manual.md` 拆成"纯客户视角"的使用手册,把 `planet.sh`、日志、LAN、故障排查这类运维内容迁到独立 `ops-runbook.md`,并把分层规则写进 `documentation-coverage-rules.md``.codex/skills/docs/SKILL.md`,让以后写文档时自动按受众归档。
## 背景
@@ -31,12 +32,12 @@
3. **登录与找回密码** — 登录页、忘记密码流程
4. **账户设置** — 修改密码、修改邮箱(需重新验证)、查看权限组、登出
5. **Console 总览** — 左侧菜单结构、各路由用途
6. **配置数据采集器**`/collection-management?tab=collector_credentials`:选择 collector、连接测试、保存凭证BarentsWatch / AISStream 两个典型例子
7. **配置 AI 凭证**`/ai?tab=providers`:默认 provider、模型、Base URL、API Key、本地代理工具 tabWebSearch、OCR
6. **配置数据采集器**`/collection-management?section=collector_credentials`:选择 collector、连接测试、保存凭证BarentsWatch / AISStream 两个典型例子
7. **配置 AI 凭证**`/ai?section=integrations`:默认 provider、模型、Base URL、API Key、本地代理工具 sectionWebSearch、OCR位于 `/ai?section=tools`
8. **系统设置**`/settings` 其他子 tab系统设置、电视直播源、SMTP 邮件)
9. **用户管理(管理员)**`/users`创建、删除、改角色、Gatekeeper 权限组
10. **数据探索**`/datasources``/data``/bgp``/alerts/*`
11. **AI 测试台**`/ai?tab=playground`
11. **AI 测试台**`/ai?section=playground`
12. **Earth 公开页面** — 现 manual.md 的 Earth 章节原样保留(图层、图例、搜索、位置候选、设置、视角、动捕、巡航、移动端)
13. **Docs 文档站** — 当前 Docs 章节保留(权限组说明)
@@ -48,7 +49,7 @@
- 打开管理员给你的 URL
- 注册账号 + 邮箱验证
- 登录后第一次做什么(建议先到 `/collection-management?tab=collector_credentials` 配一个 collector再到 `/ai` 配模型)
- 登录后第一次做什么(建议先到 `/collection-management?section=collector_credentials` 配一个 collector再到 `/ai?section=integrations` 配模型)
- 看 Earth
部署/开发的 quickstart 内容并入 `ops-runbook.md` 的"首次部署"小节,**不**再单独出 `ops-quickstart.md`,避免新增维护点。
@@ -79,7 +80,7 @@
> - 新增客户可见 UI 流 → 同时更新 `manual.md` zh+en 与 `docs-content.ts`
> - 新增 ops 命令或脚本 → 只更新 `ops-runbook.md` zh+en
## .claude/commands/docs.md 增量
## `.codex/skills/docs/SKILL.md` 增量
在 "Step 2 — Decide Scope" 后插一段:
@@ -99,7 +100,7 @@
- `docs/technical/zh/quickstart.md` & `en/quickstart.md` — 重写
- `docs/technical/zh/ops-runbook.md` & `en/ops-runbook.md` *(新)*
- `docs/documentation-coverage-rules.md` — 加受众分层段
- `.claude/commands/docs.md` — 加 Document Audience Routing 段
- `.codex/skills/docs/SKILL.md` — 加 Document Audience Routing 段
- `frontend/src/pages/Docs/docs-content.ts` — 注册 `ops-runbook``DOCS_METADATA``docs_admin` 组)
## 依赖

View File

@@ -2,7 +2,7 @@
## 背景
状态Phase 1 已经开始落地Phase 2 的 BGP 事件 / 观测站迁移和 Phase 3 的算力中心迁移也已完成。`frontend/public/earth/js/interactable.js` 已新增AIS 船只、BGP 事件、BGP 观测站和算力中心图层已经改为通过 `createInteractableLayer()` 使用通用批量 `Points`、hover / locked overlay、默认 glow、状态更新、asset icon 预加载、屏幕空间 picking、固定 / 距离缩放和跨 Interactable 同坐标避让。登陆点因 `THREE.Points` 边缘深度裁切和贴地层级要求,已退回专用 `THREE.Sprite` 黄色球路径,并与海缆同高度同 renderOrder。后续阶段聚焦把可复用的扩圈 / 雷达扇形动画正式沉淀成 `animations` 扩展。
状态Phase 1 已经开始落地Phase 2 的 BGP 事件 / 观测站迁移和 Phase 3 的算力中心迁移也已完成。`frontend/public/earth/js/interactable.js` 已新增AIS 船只、BGP 事件、BGP 观测站和算力中心图层已经改为通过 `createInteractableLayer()` 使用通用批量 `Points`、hover / locked overlay、默认 glow、状态更新、asset icon 预加载、屏幕空间 picking、固定 / 距离缩放和跨 Interactable 同坐标关系元数据。登陆点因 `THREE.Points` 边缘深度裁切和贴地层级要求,已退回专用 `THREE.Sprite` 黄色球路径,并与海缆同高度同 renderOrder。后续阶段聚焦把可复用的扩圈 / 雷达扇形动画正式沉淀成 `animations` 扩展。
当前实现说明和接入示例见:
@@ -100,9 +100,10 @@ createInteractableLayer({
| `picking.throttleMs` | `100` | `80` | hover picking 节流。 |
| `picking.skipWhileDragging` | `true` | `true` | 拖动和惯性期间跳过 hover picking。 |
| `zIndexPolicy` | `"surface-icon"` | `"surface-icon"` | 预设层级策略,避免每个业务图层手写高度和 renderOrder。 |
| `avoidance.enabled` | `true / false` | `true` | 是否参与跨 Interactable 的同坐标避让。默认开启,同一经纬度下的图标会沿地表切平面小幅排开,方便辨认和选择。 |
| `avoidance.radius` | `number` | `1.1` | 同坐标避让的第一圈半径,单位为地球本地坐标单位。 |
| `avoidance.enabled` | `true / false` | `true` | 是否记录跨 Interactable 的同坐标关系。该配置只生成 overlap 元数据,不允许移动真实 marker 坐标。 |
| `avoidance.precision` | `number` | `4` | 经纬度归并精度,默认约等于只处理几乎完全重叠的图标。 |
| `cluster.enabled` | `true / false` | 跟随 avoidance | 是否参与跨 Interactable 的当前帧屏幕重叠合并。只影响显示,不改变业务坐标。 |
| `cluster.maxMarkersPerDot` | `number` | `14` | 单个聚合圆点的对象上限;超过后按屏幕局部邻近关系拆成多个较小圆点,避免一个点过大或跨区域串联。 |
| `legend` | `{ label, color, shape }[]` | `[]` | 可选图例声明,业务层也可以继续自己导出。 |
| `metadata` | object | `{}` | 业务扩展数据,不参与渲染但参与 tooltip / info-card / search。 |

View File

@@ -1,5 +1,7 @@
# Earth Motion Capture Gesture Control Plan
> Update: Motion Agent process/device control, UE/Web shared command protocol, dual-camera redundant fusion, and the next 3D calibration route are now tracked in [Motion Agent v2 Control Protocol And 3D Calibration Roadmap](/home/ray/dev/linkong/planet/docs/plans/motion-agent-v2-control-protocol-plan.md). This document remains useful for the original provider split and gesture-control intent.
## Goal
为 Planet Earth 大屏和未来 3D 展示增加一套解耦的动作捕捉手势控制能力。实时输入分成两条路线:网页端可直接通过浏览器 `getUserMedia` 在本机识别;高级设备可继续使用本机 Motion Capture Edge Agent。两条路线都只输出轻量语义事件客户端负责把“手势事件”映射到“具体交互函数”。

View File

@@ -2,6 +2,8 @@
**状态**:已实现主体交互,并按实测调整。当前浏览器识别保留右手导航、头部切目标、左手上下切动捕图层、双手张开/收拢缩放双手上举确认暂时关闭。Motion 目标展示已改为 `CruiseSequencer` + `PresentationController` 的 persistent 展示。
> Update: Agent-side bidirectional commands, UE/Web shared device control, dual-camera redundant fusion, and future calibrated 3D mode are tracked in [Motion Agent v2 Control Protocol And 3D Calibration Roadmap](/home/ray/dev/linkong/planet/docs/plans/motion-agent-v2-control-protocol-plan.md).
## Summary
把动捕从“几个单点手势触发函数”升级为一套更像大屏遥控器的交互层右手负责地球导航头部负责候选切换左手上下切换动捕候选图层双手负责缩放调试面板支持“只显示骨骼”和暂停匹配。进入动捕模式后Earth 自动软选中屏幕中心附近的正面可交互目标;确认动作预留为把目标升级为锁定,并用巡航/引导线式详情打开,不再模拟鼠标点击。

Some files were not shown because too many files have changed in this diff Show More