Compare commits

...

35 Commits

Author SHA1 Message Date
linkong
2da25376bd release: bump version to 0.43.1 2026-04-28 16:21:33 +08:00
linkong
ac69d5d354 release: bump version to 0.43.0 2026-04-28 16:10:17 +08:00
rayd1o
1cd2dab0ee release: bump version to 0.42.2 2026-04-28 04:35:13 +08:00
rayd1o
42d019af36 release: bump version to 0.42.1 2026-04-28 04:29:44 +08:00
rayd1o
b4e8afb272 release: bump version to 0.42.0 2026-04-28 04:27:18 +08:00
rayd1o
eeee788530 release: bump version to 0.41.2 2026-04-27 23:23:23 +08:00
linkong
655e2a7d2d release: bump version to 0.41.1
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-27 16:31:34 +08:00
linkong
3ea99a9529 release: bump version to 0.41.0 2026-04-27 13:58:29 +08:00
rayd1o
f9c1334365 release: bump version to 0.40.5
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-26 05:03:30 +08:00
rayd1o
5f47ec1659 release: bump version to 0.40.4 2026-04-26 01:41:29 +08:00
rayd1o
229be0bced release: bump version to 0.40.3 2026-04-25 23:02:22 +08:00
linkong
50a417ca83 release: bump version to 0.40.2 2026-04-24 17:50:43 +08:00
linkong
e9464a9833 release: bump version to 0.40.1 2026-04-24 17:28:03 +08:00
linkong
86807f6af6 release: bump version to 0.40.0 2026-04-24 15:41:42 +08:00
rayd1o
8b8f7138c0 release: bump version to 0.39.0 2026-04-24 00:48:33 +08:00
linkong
d5f3784ffb release: bump version to 0.38.0 2026-04-23 17:57:35 +08:00
linkong
195a8bf71c release: bump version to 0.37.2 2026-04-23 11:12:49 +08:00
linkong
987c378f99 release: bump version to 0.37.1 2026-04-23 11:05:30 +08:00
rayd1o
67f82dc41c release: bump version to 0.37.0 2026-04-23 07:56:10 +08:00
rayd1o
abe04030fb release: bump version to 0.36.0 2026-04-22 23:42:10 +08:00
linkong
6a5f9f7ad4 release: bump version to 0.35.1 2026-04-22 18:04:54 +08:00
linkong
439a512148 docs: add earth mobile drawer UI plan and Claude Code/Codex toolchain
- Add earth-mobile-drawer-ui-plan documenting mobile drawer UX decisions
- Add goal-driven.md Claude Code command for autonomous task execution
- Add .codex/ config with OpenAI model definitions and goal-driven agent
- Add SKILL.md, openai.yaml, and prompt-template for Codex integration
2026-04-22 17:37:00 +08:00
linkong
f73fa1ea6d release: bump version to 0.35.0 2026-04-22 17:29:24 +08:00
linkong
5b623a6385 release: bump version to 0.34.0 2026-04-22 12:49:37 +08:00
rayd1o
0082cf3fbd release: bump version to 0.33.0 2026-04-22 05:28:54 +08:00
rayd1o
3ae4acdff8 release: bump version to 0.32.0 2026-04-22 04:41:39 +08:00
rayd1o
437efc848c release: bump version to 0.31.3 2026-04-22 03:52:09 +08:00
rayd1o
003a46ac30 release: bump version to 0.31.2 2026-04-21 23:50:35 +08:00
rayd1o
4b0be4cb76 release: bump version to 0.31.1 2026-04-21 22:49:39 +08:00
linkong
b7647379de release: bump version to 0.31.0 2026-04-21 18:35:40 +08:00
linkong
0f89372d71 release: bump version to 0.30.0 2026-04-21 12:28:04 +08:00
linkong
2b0d4cfc49 release: bump version to 0.29.2 2026-04-21 10:43:48 +08:00
rayd1o
e6d0332fba release: bump version to 0.29.1 2026-04-20 22:12:19 +08:00
linkong
fe45a99cbd release: bump version to 0.29.0 2026-04-20 17:43:59 +08:00
linkong
ae77b06c3c docs: expand earth celestial background implementation plan 2026-04-20 16:29:03 +08:00
203 changed files with 39909 additions and 2560 deletions

View File

@@ -12,6 +12,19 @@ allowed-tools: ["Read", "Edit", "Bash", "Grep", "Glob"]
`$ARGUMENTS` 非空,则只检查指定文件/目录;否则检查所有未提交修改(`git diff HEAD`)。
## 节省上下文规则
优先用确定性的 CLI 检查缩小范围,不要一上来把完整文件或大 diff 读入上下文:
```bash
git diff --name-only HEAD
git diff --unified=0 HEAD -- <path>
git diff --check
rg -n "TODO|FIXME|console\.log|debugger|print\(" <changed-paths>
```
只有 focused diff 不足以安全判断或修改时,才读取完整文件。
## 审查清单
按优先级检查以下问题(只报告在本次 diff 中**新增或修改**的代码里存在的问题):
@@ -59,8 +72,13 @@ git diff HEAD --name-only
### Step 2 — 逐文件阅读并分析
- 用 Read 工具读取完整文件(不只读 diff
- 对照审查清单,记录每个问题:文件名、行号、问题类型、建议修复方式
先从 focused diff 开始:
```bash
git diff --unified=0 HEAD -- <file>
```
`rg``git diff --check`、编译器或 linter 输出确认确定性问题。只有需要上下文时才用 Read 读取完整文件。对照审查清单,记录每个问题:文件名、行号、问题类型、建议修复方式。
### Step 3 — 报告问题清单
@@ -95,6 +113,7 @@ git diff HEAD --name-only
- 只改在审查清单中发现的问题,不做额外优化
- 每次 Edit 只修改确实有问题的行,保持 diff 最小
- 改完后用 `grep` 验证旧的坏代码已消失
- 优先做精确补丁;只有仓库已有对应格式化流程时,才运行格式化工具
### Step 5 — 输出总结

151
.claude/commands/docs.md Normal file
View File

@@ -0,0 +1,151 @@
---
description: 分析本次 git 变更,在 docs/technical/zh/ 中新建或更新对应的技术文档
argument-hint: 可选:指定要记录的主题,或留空自动从 git diff 推断
allowed-tools: ["Read", "Edit", "Write", "Bash", "Glob", "Grep"]
---
# /docs — 技术文档写入工作流
## 目标
根据当前 git 变更(或用户指定主题)在 `docs/technical/zh/` 中写入或更新技术文档,记录**为什么**这样做,而不只是记录做了什么。
## 执行步骤
### Step 1 — 理解变更范围
```bash
git diff HEAD --stat # 变更文件一览
git diff HEAD --name-only # 变更文件列表
git log --oneline -10 # 近期 commit 上下文
```
`$ARGUMENTS` 指定了主题,优先聚焦该主题;否则从文件列表和 diff stat 推断变更主题。不要默认读取完整仓库 diff只对决定文档主题所需的文件读取 focused diff
```bash
git diff HEAD -- <path>
rg -n "class |def |function |export |router|@router|interface |type " <path>
```
### Step 2 — 确认文档范围
分析变更,判断:
1. **应写几篇文档**:单一主题写一篇,跨领域变更可拆分(如后端性能优化 + 运维启动脚本分开写)
2. **是新建还是更新**:检查 `docs/technical/zh/` 中是否已有相关文档
3. **文档命名**:按 `领域-主题-副题.md` 格式,全小写,用连字符,如:
- `backend-datasources-api-performance.md`
- `ops-planet-sh-startup.md`
- `earth-bgp-context.md`
```bash
ls docs/technical/zh/ # 查看现有文档
```
**先输出写作计划供用户确认**(若变更明确且范围小,可直接执行):
```
文档计划:
新建docs/technical/zh/ops-planet-sh-startup.md — planet.sh 启动性能优化
更新docs/technical/zh/backend-datasources-api-performance.md — 补充并行化细节
```
### Step 3 — 写文档
遵循以下原则:
**记录 WHY不只记录 WHAT**
- 好:`将戳文件从 /tmp 移到 ~/.cache/planet/,因为 WSL 重启后 /tmp 被清空`
- 差:`修改了 AI_PROVIDER_BUILD_STAMP_FILE 的值`
**必须包含的内容**
- 背景/问题:改动之前存在什么问题,为什么要改
- 核心设计决策及其理由
- 关键代码片段(用 diff 或 before/after 展示)
- 相关文件列表
**格式要求**
- 使用 `##``###` 分级,不要超过三级
- 代码块注明语言python / bash / typescript / sql
- 表格用于对比多个选项或列出参数
- 中文写作,技术术语保留英文原文
- `docs/technical/zh/` 中的文档不得用英文原文占位;如果存在 `docs/technical/en/` 对应文件,禁止逐字复制成中文文件
- 中文文档内部链接应指向 `docs/technical/zh/...`,除非明确引用英文专属文档
**文档结构模板**
```markdown
# 标题(说明做了什么)
## 背景
为什么要做这个改动,改动前存在什么问题。
## 核心变更
### 子主题一
before/after 或决策说明 + 关键代码
### 子主题二
...
## 相关文件
- `path/to/file.py` — 简短说明
```
### Step 4 — 验证
- 读一遍写好的文档,确认逻辑清晰、代码片段无明显错误
-`rg --files``test -e` 确认文档中的文件路径在项目中真实存在,避免凭记忆判断:
- 检查中文文档没有误复制英文版:
```bash
python - <<'PY'
from pathlib import Path
same = []
for en in sorted(Path("docs/technical/en").glob("*.md")):
zh = Path("docs/technical/zh") / en.name
if zh.exists() and en.read_text() == zh.read_text():
same.append(en.name)
if same:
raise SystemExit("identical en/zh docs: " + ", ".join(same))
print("no identical en/zh docs")
PY
```
- 检查中文文档内部链接没有继续指向无语言目录:
```bash
rg -n "/home/ray/dev/linkong/planet/docs/technical/(?!zh|en)" docs/technical/zh --pcre2
```
```bash
# 对文档中提到的关键路径做快速验证
ls <mentioned_paths>
```
如需检查大量链接,优先用确定性提取:
```bash
rg -n "\]\(([^)]+)\)" docs/technical/zh/<doc>.md
```
### Step 5 — 完成确认
输出摘要:
```
✓ 新建docs/technical/zh/ops-planet-sh-startup.md约 xxx 字)
✓ 更新docs/technical/zh/backend-datasources-api-performance.md
```
## 注意事项
- 不要写流水账式的"改了 A、改了 B、改了 C",要写改动背后的约束和权衡
- 不要在文档中引用 PR 号、issue 号、或当前对话——这些会随时间失效
- 代码片段保持简洁,只保留说明问题的关键部分,省略无关样板代码
- 如果某个变更已有文档记录,优先在原文档中追加,而不是新建
- 文档是给未来的开发者看的,假设读者熟悉项目但不了解这次改动的背景

View File

@@ -0,0 +1,93 @@
---
description: 用 goal-driven 方法推动一个复杂任务持续执行,直到明确成功标准被满足
argument-hint: 建议填写任务目标;若同时给出成功标准更好
allowed-tools: ["Read", "Edit", "Bash", "Grep", "Glob"]
---
# /goal-driven — 目标驱动执行模式
使用 `lidangzzz/goal-driven` 的核心思想来推进复杂任务:先固定目标与成功标准,再持续执行和反复验收,直到标准真正满足。
适用场景:
- 长周期实现任务
- 高复杂度工程任务
- 可被明确验收的研究、实现、迁移、验证类工作
不适用场景:
- 纯脑暴
- 无法定义成功标准的模糊任务
- 很小的一次性修改
## 输入要求
`$ARGUMENTS` 只包含目标,没有成功标准,先补全一版可执行的成功标准再开始。
启动时先输出:
```md
Goal
- ...
Criteria for success
- ...
Plan
1. ...
2. ...
3. ...
Verification
- ...
```
## 执行规则
1. 先把任务固化为两个核心块:
- `Goal`
- `Criteria for success`
2. 成功标准必须尽量客观,可验证,可落地。
优先写成:
- 需要交付什么
- 需要通过哪些测试或验证
- 如何判断结果真的完成
3. 进入持续执行循环:
- 完成一个阶段
- 检查当前结果是否满足成功标准
- 若未满足,明确剩余差距并继续推进
4. 任何“完成了”“差不多了”“已实现”之类的结论,都必须经过验证,不能直接接受。
5. 如果验证失败:
- 明确指出哪条成功标准没满足
- 继续工作,不要把阶段性进展误判为完成
6. 只有在以下情况之一才能停止:
- 成功标准已满足
- 用户明确要求停止
## 执行风格
- 重证据,轻口头判断
- 优先使用确定性工具证据:`rg``git diff --stat``git diff -- <path>`、测试、构建、lint、`curl`、数据库查询等能直接证明成功标准的方式
- 不把大段命令输出粘进回复;保留在工具调用里,回复只总结关键证据
- 重验收,轻自我感觉
- 优先用测试、日志、产物、对比结果来证明完成
- 对长期任务保持“未达标就继续”的节奏
## 简版模板
```md
Goal: [[[[[在此填写最终目标]]]]]
Criteria for success: [[[[[在此填写成功标准]]]]]
循环执行:
1. 推进任务
2. 检查是否满足成功标准
3. 若未满足,继续工作
4. 直到满足标准或用户明确停止
```

View File

@@ -28,6 +28,19 @@ allowed-tools: ["Read", "Edit", "Bash", "Glob", "Grep"]
- `docs/CHANGELOG.md`
- `docs/version-history.md`
## 节省上下文规则
发版判断应以确定性 CLI 证据为主,优先使用紧凑命令和定点读取:
```bash
git status --short
git diff --stat HEAD
git diff --name-only HEAD
rg -n "version|^## |^Released:|当前开发版本|current" VERSION frontend/package.json pyproject.toml docs/CHANGELOG.md docs/version-history.md
```
除非需要判断某个代码变更是否属于本次发版,否则不要读取完整 diff。
## 执行步骤
### Step 1 — 环境检查
@@ -45,7 +58,7 @@ cat VERSION # 读取当前版本
### Step 2 — 确定发版类型与新版本号
-`$ARGUMENTS` 提供了明确类型(`feature` / `bugfix`),直接使用
- 否则根据当前 `git diff HEAD``git log` 推断
- 否则根据 `git diff --stat HEAD``git diff --name-only HEAD`、必要的 focused diff`git log` 推断
- 计算新版本号(例:`0.26.2` → bugfix → `0.26.3`
- **先输出发版计划供用户确认**
@@ -91,12 +104,13 @@ cat VERSION # 读取当前版本
针对本次变更范围做最小验证:
- Python 文件有修改:`python3 -m py_compile <changed_files>`
- Frontend 文件有修改:运行项目标准检查(若无则跳过并说明)
- Python 文件有修改:先用 `git diff --name-only HEAD -- '*.py'` 列出,再运行 `python3 -m py_compile <changed_files>`
- Frontend 文件有修改:先用 `git diff --name-only HEAD -- frontend` 判断范围,再运行项目标准检查(若无则跳过并说明)
- 版本号一致性检查:用 grep 确认 VERSION、package.json、pyproject.toml 中的版本号完全一致
```bash
grep -h "version" VERSION frontend/package.json pyproject.toml
cat VERSION
rg -n "\"version\":|^version =|version = " frontend/package.json pyproject.toml uv.lock
```
### Step 7 — 提交前预览

3
.codex/config.toml Normal file
View File

@@ -0,0 +1,3 @@
approval_policy = "never"
sandbox_mode = "danger-full-access"

View File

@@ -21,6 +21,19 @@ If the user specifies a file or directory, check only that. Otherwise check all
Only report issues present in **newly added or modified** lines of this diff — do not audit unchanged code.
## Token-Saving Rule
Prefer deterministic CLI checks before reading files into model context:
```bash
git diff --name-only HEAD
git diff --unified=0 HEAD -- <path>
git diff --check
rg -n "TODO|FIXME|console\.log|debugger|print\(" <changed-paths>
```
Read full files only when the focused diff does not provide enough surrounding context to make a safe edit.
## Checklist
### 1. Duplicate Logic
@@ -64,7 +77,13 @@ Filter to the user-specified path if one was provided.
### Step 2 — Read and analyze each file
Read the full file (not just the diff) with the Read tool. For each file, record every issue found: filename, line number, category, and suggested fix.
Start with focused diffs:
```bash
git diff --unified=0 HEAD -- <file>
```
Use `rg`, `git diff --check`, and compiler/linter output for deterministic findings. Read the full file only for files that need surrounding context. For each issue found, record filename, line number, category, and suggested fix.
### Step 3 — Report findings before touching anything
@@ -99,6 +118,7 @@ Principles:
- Only fix issues identified in the checklist — no extra improvements
- Keep each Edit as small as possible
- After fixing, verify the old bad pattern is gone with grep
- Prefer `apply_patch` for targeted edits; use formatters only when the repository already uses them for the touched file type
### Step 5 — Summary

117
.codex/skills/docs/SKILL.md Normal file
View File

@@ -0,0 +1,117 @@
---
name: docs
description: Analyze current Planet repo changes and create or update technical documentation under docs/technical/zh. Use when the user asks to write docs, update technical docs, summarize implementation changes into documentation, or port the Claude docs-codex workflow into Codex.
---
# Docs
Use this skill when the user asks to create or update Planet technical documentation, especially under `docs/technical/zh/`.
## Goal
Write or update technical docs that explain why a change exists, not only what files changed.
Default target directory:
- `docs/technical/zh/`
## Workflow
1. Gather change context:
```bash
git diff HEAD --stat
git diff HEAD --name-only
git log --oneline -10
ls docs/technical/zh/
```
If the user gives a specific topic, focus on that topic. Otherwise infer the documentation topic from the file list and diff stat. Do **not** read the full repository diff by default; inspect focused diffs only for the files that define the doc topic:
```bash
git diff HEAD -- <path>
rg -n "class |def |function |export |router|@router|interface |type " <path>
```
2. Decide document scope:
- Use one document for one coherent topic.
- Split documents when the changes cross meaningful domains, such as backend performance and ops startup behavior.
- Prefer updating an existing relevant doc over creating a duplicate.
- Name new files as lowercase hyphenated `domain-topic-detail.md`, for example:
- `backend-datasources-api-performance.md`
- `ops-planet-sh-startup.md`
- `earth-bgp-context.md`
3. Write the doc in Chinese:
- Write Chinese prose for `docs/technical/zh/`.
- Keep technical identifiers, API paths, config keys, code symbols, and standard product names in English where appropriate.
- Use `##` and `###` headings; avoid going deeper than three levels.
- Use fenced code blocks with language tags.
- Use tables when comparing options or listing parameters.
4. Required content:
- Background/problem: what was wrong before and why the change was needed.
- Core design decisions and rationale.
- Key code snippets, preferably before/after or focused excerpts.
- Related files and what each file contributes.
5. Verification:
- Read the completed doc and check that the reasoning is clear.
- Verify important referenced paths exist.
- Use `rg --files` or `test -e` for path existence instead of relying on memory.
- Run a quick duplicate-language check when editing bilingual docs:
```bash
python - <<'PY'
from pathlib import Path
same = []
for en in sorted(Path("docs/technical/en").glob("*.md")):
zh = Path("docs/technical/zh") / en.name
if zh.exists() and en.read_text() == zh.read_text():
same.append(en.name)
if same:
raise SystemExit("identical en/zh docs: " + ", ".join(same))
print("no identical en/zh docs")
PY
```
Also check that Chinese docs do not link to the old language-less technical docs path:
```bash
rg -n "/home/ray/dev/linkong/planet/docs/technical/(?!zh|en)" docs/technical/zh --pcre2
```
This command should return no matches.
If checking many links, prefer deterministic extraction:
```bash
rg -n "\]\(([^)]+)\)" docs/technical/zh/<doc>.md
```
## Hard Constraints
- A file under `docs/technical/zh/` must not be an English source file copied as a placeholder.
- Do not leave a Chinese doc with only an English title and English first-screen content.
- When an English counterpart exists in `docs/technical/en/`, never duplicate it byte-for-byte into `docs/technical/zh/`.
- Internal links inside `docs/technical/zh/` should point to `docs/technical/zh/...` for Chinese docs, unless intentionally linking to an English-only file.
- Do not reference PR numbers, issue numbers, or the current conversation.
- Do not write changelog-style lists like "changed A, changed B, changed C" without the constraints and tradeoffs behind those changes.
- Keep code snippets concise and relevant.
## Recommended Output
After editing, summarize:
```md
Updated:
- docs/technical/zh/example.md — what changed
Verified:
- no identical en/zh docs
- no language-less docs/technical links in zh docs
```

View File

@@ -0,0 +1,103 @@
---
name: goal-driven
description: Run a goal-driven execution loop for very large, long-horizon, rigorously verifiable tasks. Use when the user explicitly wants the lidangzzz/goal-driven method, a master-agent plus worker-agent style workflow, or a persistent loop that keeps working until concrete success criteria are satisfied.
---
# Goal-Driven
Use this skill when the user wants a strict goal-driven workflow for a hard task with:
- one clear end goal
- explicit success criteria
- repeated verification against those criteria
- continued execution until the criteria are actually met
This skill is adapted from `lidangzzz/goal-driven`, but trimmed for local skill use to avoid bloating context.
## When To Use
Use it for tasks like:
- compilers, interpreters, theorem-like proof work, deep refactors
- long-running system design or implementation work
- problems that are expensive and complex, but still objectively testable
Do not use it for:
- vague brainstorming without a success condition
- short one-shot edits
- tasks where "done" cannot be evaluated in a meaningful way
## Core Model
The workflow has two roles:
1. Master role
Defines the goal, defines the success criteria, audits progress, and decides whether the work is actually complete.
2. Worker role
Keeps advancing the task toward the goal. If a result is partial, stalled, or unverifiable, the worker continues.
In Codex, only use actual subagents when the user explicitly asks for delegation or subagent work and the platform supports it. Otherwise emulate the same loop locally: keep working, checkpointing, and re-verifying until the criteria are satisfied.
## Workflow
1. Normalize the task into two blocks:
- `Goal`
- `Criteria for success`
2. Make the criteria concrete and testable.
Good criteria usually include:
- required outputs
- required validations or tests
- edge cases or coverage thresholds
- what evidence proves completion
3. Break the work into milestones that can each produce evidence.
4. Execute the next milestone.
If subagents are explicitly allowed, the master may delegate bounded worker tasks.
If not, do the work locally but keep the master/worker mindset.
5. Whenever work pauses, stalls, or appears complete, audit against the criteria directly.
Check artifacts, tests, logs, diffs, metrics, or other real evidence.
6. If the criteria are not met, continue with a specific delta:
- what is still missing
- what evidence failed
- what the next worker pass must improve
7. Stop only when the criteria are met, or when the user explicitly stops the process.
## Operating Rules
- Prefer objective checks over self-reported completion.
- Prefer deterministic tool evidence over long model summaries: use `rg`, `git diff --stat`, targeted `git diff -- <path>`, tests, builds, linters, `curl`, or database queries when they can prove a criterion.
- Do not paste large command output into the conversation; summarize the evidence and keep raw output in tool calls.
- Do not confuse progress with completion.
- If the worker says "done", verify it.
- If verification fails, continue from the gap instead of restarting blindly.
- Keep the goal stable unless the user changes it.
- Tighten fuzzy criteria before sinking large amounts of effort.
## Recommended Response Shape
When starting a goal-driven task, structure the kickoff like this:
```md
Goal
- ...
Criteria for success
- ...
Current plan
1. ...
2. ...
3. ...
Verification
- What evidence will prove completion
```
For a reusable prompt template, read [references/prompt-template.md](references/prompt-template.md).

View File

@@ -0,0 +1,7 @@
interface:
display_name: "Goal-Driven"
short_description: "Drive complex work until explicit success criteria are met."
default_prompt: "Use $goal-driven to turn this task into a concrete goal, explicit success criteria, and a verification-driven execution loop."
policy:
allow_implicit_invocation: true

View File

@@ -0,0 +1,38 @@
# Goal-Driven Prompt Template
Use this when you want a reusable kickoff prompt for a master/worker execution loop.
```md
# Goal-Driven System
Goal: [[[[[DEFINE THE FINAL GOAL HERE]]]]]
Criteria for success: [[[[[DEFINE THE SUCCESS CRITERIA HERE]]]]]
You are the master agent.
Your job is to:
1. Keep the goal and criteria fixed.
2. Start worker execution toward the goal.
3. Audit any claimed progress against the criteria.
4. If the criteria are not met, continue the work with a precise next delta.
5. Stop only when the criteria are satisfied or the user explicitly stops the process.
Worker requirements:
1. Break the task into subproblems.
2. Keep producing concrete progress toward the goal.
3. Report evidence, not just claims.
4. Continue until the criteria are satisfied.
Master audit loop:
1. Check whether the worker is still making progress.
2. If the worker stalls or claims completion, verify against the criteria.
3. If verification fails, resume work from the remaining gap.
4. Repeat until the criteria are met.
```
## Notes
- Stronger criteria produce better results than stronger rhetoric.
- Prefer measurable checks such as tests, parity checks, generated artifacts, benchmarks, or reviewable outputs.
- If the environment does not support subagents, emulate the same loop locally.

View File

@@ -18,7 +18,7 @@ Do not use this skill for ordinary commits that are not being released.
## Versioning Rules
- `feature` -> bump `+0.1.0`
- `feature` -> bump minor and reset patch to `0` (`x.y.z``x.(y+1).0`; for example `0.41.2``0.42.0`)
- `bugfix` -> bump `+0.0.1`
- `docs`, `maintenance`, and `refactor` do not bump by default unless the user explicitly wants a release
@@ -35,6 +35,19 @@ Use `git rev-parse --show-toplevel` to get the repo root. All paths are relative
- `docs/CHANGELOG.md`
- `docs/version-history.md`
## Token-Saving Rule
Release work should be driven by deterministic CLI evidence. Prefer compact commands and targeted file reads:
```bash
git status --short
git diff --stat HEAD
git diff --name-only HEAD
rg -n "version|^## |^Released:|current" VERSION frontend/package.json pyproject.toml docs/CHANGELOG.md docs/version-history.md
```
Do not inspect full diffs unless deciding whether changed code belongs in the release.
## Workflow
### Step 1 — Environment check
@@ -52,8 +65,10 @@ If unrelated uncommitted changes exist, list them and ask the user whether to in
### Step 2 — Determine release type and next version
- If the user provided an explicit type (`feature` / `bugfix`), use it
- Otherwise infer from `git diff HEAD` and recent `git log`
- Compute the next version (e.g. `0.26.2` → bugfix → `0.26.3`)
- Otherwise infer from `git diff --stat HEAD`, `git diff --name-only HEAD`, focused diffs for changed code, and recent `git log`
- Compute the next version:
- `feature`: increment minor and reset patch to `0` (e.g. `0.41.2``0.42.0`)
- `bugfix`: increment patch only (e.g. `0.26.2``0.26.3`)
- **Show the release plan before making any changes:**
```
@@ -104,12 +119,13 @@ Get today's date with `date +%Y-%m-%d`.
Run the smallest relevant validation for the changes in scope:
- Python files changed: `python3 -m py_compile <changed_files>`
- Frontend files changed: run the project-standard check if available; otherwise skip and say so
- Python files changed: list changed Python files with `git diff --name-only HEAD -- '*.py'`, then run `python3 -m py_compile <changed_files>`
- Frontend files changed: list changed frontend files with `git diff --name-only HEAD -- frontend`, then run the project-standard check if available; otherwise skip and say so
- Version consistency: confirm VERSION, package.json, pyproject.toml, and uv.lock all show the same version
```bash
grep -h "version" VERSION frontend/package.json pyproject.toml
cat VERSION
rg -n "\"version\":|^version =|version = " frontend/package.json pyproject.toml uv.lock
```
### Step 7 — Pre-commit preview

124
README.md
View File

@@ -227,6 +227,120 @@ bun run build
启动服务后访问: `http://localhost:8000/docs`
## WSL / Windows 局域网访问
如果服务运行在 WSL 中,而你希望:
- Windows 本机浏览器访问开发服务
- 同一局域网内的手机或其他电脑访问开发服务
推荐按下面顺序排查和配置。
### 1. 在 WSL 中启动服务
```bash
./planet.sh start --allow-lan
```
这会让前端监听 `0.0.0.0:3000`,后端监听 `0.0.0.0:8000`
### 2. 先确认 WSL 内部服务正常
在 WSL 中执行:
```bash
curl http://localhost:3000
curl http://localhost:8000/health
ss -ltnp | grep -E ':3000|:8000'
```
预期:
- `3000` 返回前端 HTML
- `8000/health` 返回健康检查 JSON
- `ss` 中能看到 `0.0.0.0:3000``0.0.0.0:8000`
如果这一步不通,先不要继续做 Windows 转发。
### 3. 在 Windows 本机验证 localhost 直通
在 Windows PowerShell 中执行:
```powershell
curl http://localhost:3000
curl http://localhost:8000/health
```
在常见的 WSL2 开发环境下Windows 通常可以直接通过 `localhost` 访问 WSL 中的服务。
### 4. 如果需要让局域网设备访问,再做 Windows 端口转发
注意:下面的命令必须在“以管理员身份运行”的 PowerShell 中执行。
先把 Windows 对外网卡上的 `3000` / `8000` 转发到 Windows 本机 `127.0.0.1`
```powershell
netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=3000
netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=8000
netsh interface portproxy add v4tov4 listenaddress=0.0.0.0 listenport=3000 connectaddress=127.0.0.1 connectport=3000
netsh interface portproxy add v4tov4 listenaddress=0.0.0.0 listenport=8000 connectaddress=127.0.0.1 connectport=8000
```
再放行 Windows 防火墙:
```powershell
New-NetFirewallRule -DisplayName "WSL Planet 3000" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 3000
New-NetFirewallRule -DisplayName "WSL Planet 8000" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 8000
```
检查转发规则是否生效:
```powershell
netsh interface portproxy show all
```
预期能看到:
- `0.0.0.0:3000 -> 127.0.0.1:3000`
- `0.0.0.0:8000 -> 127.0.0.1:8000`
### 5. 查 Windows 局域网 IP并让其他设备访问
在 Windows PowerShell 中执行:
```powershell
ipconfig
```
找到当前联网网卡的 IPv4 地址,例如 `192.168.8.228`
局域网其他设备可访问:
- `http://<Windows局域网IP>:3000/earth`
- `http://<Windows局域网IP>:3000/admin`
例如:
- `http://192.168.8.228:3000/earth`
### 6. 常见现象与判断
- WSL 中 `curl localhost:3000` 能通,但 Windows 访问 `WSL 的局域网 IP:3000` 不通:这是正常现象之一,优先验证 Windows 的 `localhost:3000`
- Windows `localhost:3000` 能通,但局域网设备访问 `Windows 局域网 IP:3000` 不通:通常缺少 `portproxy` 或防火墙放行
- `whoami /groups``S-1-5-32-544` 显示 `deny only`:说明当前 PowerShell 不是提权管理员窗口
### 7. 本项目一次性验证顺序
建议固定按这个顺序验证:
1. WSL 中执行 `curl http://localhost:3000`
2. WSL 中执行 `curl http://localhost:8000/health`
3. Windows 中执行 `curl http://localhost:3000`
4. Windows 中执行 `curl http://localhost:8000/health`
5. 管理员 PowerShell 配置 `portproxy` 和防火墙
6. 用手机或其他电脑访问 `http://<Windows局域网IP>:3000/earth`
## 启动容错参数
`planet.sh` 现在为依赖安装、数据库、AI Provider 启动加入了有限次重试,并会在数据库与 `aiprovider` 启动后额外等待 Docker healthcheck。
@@ -328,11 +442,11 @@ AI_PROVIDER_SERVICE_TOKEN=change_me
详细文档:
- [docs/agents/aiprovider.md](/home/ray/dev/linkong/planet/docs/agents/aiprovider.md)
- [docs/technical/agents-aiprovider.md](/home/ray/dev/linkong/planet/docs/technical/agents-aiprovider.md)
- [aiprovider/README.md](/home/ray/dev/linkong/planet/aiprovider/README.md)
- [docs/frontend/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/frontend/frontend-layout-guidelines.md)
- [docs/frontend/ai-playground-development-plan.md](/home/ray/dev/linkong/planet/docs/frontend/ai-playground-development-plan.md)
- [docs/agents/situational-awareness-foundation-plan.md](/home/ray/dev/linkong/planet/docs/agents/situational-awareness-foundation-plan.md)
- [docs/technical/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/technical/frontend-layout-guidelines.md)
- [docs/plans/frontend-ai-playground-development-plan.md](/home/ray/dev/linkong/planet/docs/plans/frontend-ai-playground-development-plan.md)
- [docs/plans/agents-situational-awareness-foundation-plan.md](/home/ray/dev/linkong/planet/docs/plans/agents-situational-awareness-foundation-plan.md)
## 前端页面布局规范
@@ -346,7 +460,7 @@ AI_PROVIDER_SERVICE_TOKEN=change_me
当前推荐参考实现:
- [frontend/src/pages/BGP/BGP.tsx](/home/ray/dev/linkong/planet/frontend/src/pages/BGP/BGP.tsx)
- [docs/frontend/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/frontend/frontend-layout-guidelines.md)
- [docs/technical/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/technical/frontend-layout-guidelines.md)
## License

15
TODO.md
View File

@@ -20,3 +20,18 @@
- [x] 在 activity layer 之后继续补 `route leak``path instability / flap` detector
- [ ] 对 [frontend/public/earth/js/bgp.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/bgp.js) 做按职责拆分的小重构,拆成 data / markers / overlays / animation降低后续维护复杂度
- [ ] 可选优化(非必做):将 BGP incident/collector 标点改为 HTML marker参考 worldmonitor 的 `htmlElementsData` 思路),实现近乎固定屏幕尺寸与更高密度可点击性
- [ ] 保持 Earth 当前这批纯个人偏好设置继续走本地持久化:`旋转模式`、HUD 面板显示/隐藏、`地形透明度` 暂不升级到后端系统设置,避免把设备级偏好过早做成全局配置
- [ ] 如果后续明确需要“账号级同步 Earth 偏好”,再单独设计 `Earth user preferences`:优先按用户维度而不是全局系统设置保存,并规划 `localStorage -> backend` 的平滑迁移策略
- [ ] 为 Planet / Earth 补一个可用的日志查看系统:先明确前后端/AI Provider/采集任务的日志入口、最近日志聚合、筛选与 tail 能力,再决定是先做脚本级统一入口还是控制台内置日志面板
- [ ] 把 Earth 态势新闻源从 [earth_news.py](/home/ray/dev/linkong/planet/backend/app/services/earth_news.py) 的硬编码列表抽成可配置目录,优先保持当前“实时聚合”链路不变,只先解决新闻源不可配置的问题
- [ ] 为 Earth 态势新闻设计后续采集器化方案:明确新闻数据模型、去重策略、区域映射、过期清理和 Earth/AI 复用方式,再决定何时把新闻从实时抓取升级成正式 collector
- [ ] 为 Earth 地球表面增加一层与基础纹理对齐的材质/纹理 overlay并在同层叠加国界轮廓参考线要求国界线与底图稳定对齐且 hover 到国家轮廓时能高亮当前国家,便于校准地表和增强交互
- [ ] 把 Earth 新闻接入通用巡航队列:按新闻发生地和时间排序生成巡航目标,巡航聚焦到新闻事件时显示对应新闻卡片,并保持实现边界为“通用巡航层 + 新闻业务适配层”,不要再把新闻逻辑直接耦合回 `main.js` 状态机
- [ ] 为未知位置的算力中心建立分层坐标补全链路:优先 `精确坐标 > 站点/园区命中 > 城市 > 州/省 > 国家内主要算力城市 > 国家质心`,并把每次回退的 `confidence / reason / precision` 明确写进统一 GeoJSON
- [ ] 为算力中心补一份可维护的本地位置注册表,例如 `canonical_name / aliases / operator / country / region / city / lat / lon / confidence / source_note`,避免把地点知识长期硬编码在 `visualization.py`
- [ ] 增强 `epoch_ai_gpu` 和相关算力采集器的源页面解析:即使公开 API 不给坐标也继续尝试从详情页、HTML、内嵌 JSON、schema.org、OpenGraph、脚本变量和 PDF/新闻稿链接里抽地点线索
- [ ] 为未知位置算力中心增加外部富化策略评估:可选接入公开知识源或搜索兜底,只抓“站点名/园区名/城市名”级别线索,不直接抓经纬度结论,并把结果作为候选证据而不是真值
- [ ] 为算力中心建立 `operator / cluster name / facility alias` 归一化层,先解决 `xAI / Colossus / Memphis``OpenAI / Stargate``CoreWeave``Lambda``Crusoe` 这类同一对象多种写法导致的地点匹配失败
- [ ] 为估算位置增加更细的视觉和产品表达:除了问号角标,还要支持 tooltip/详情中的“估算依据”“精度级别”“最后核验时间”,并允许在设置中单独开关“仅看精确位置”
- [ ] 为国家级估算点设计更合理的落点策略:优先落在“该国主要算力/数据中心城市候选集”而不是几何质心,必要时同国多节点做稳定散列分配,避免大量节点堆在荒漠或海上
- [ ] 为未知位置算力中心建立人工校验工作流:支持导出待核验清单、记录人工确认结果,并把人工确认反哺到位置注册表,逐步减少问号点比例

View File

@@ -1 +1 @@
0.28.2
0.43.1

View File

@@ -1,6 +1,10 @@
FROM python:3.14-slim
ARG PYTHON_IMAGE=python:3.14-slim
ARG UV_IMAGE=ghcr.io/astral-sh/uv:latest
COPY --from=ghcr.io/astral-sh/uv:latest /uv /uvx /bin/
FROM ${UV_IMAGE} AS uv
FROM ${PYTHON_IMAGE}
COPY --from=uv /uv /uvx /bin/
WORKDIR /app

View File

@@ -37,8 +37,26 @@ def verify_service_token(x_provider_token: str | None = Header(default=None)) ->
)
def get_provider_service() -> ProviderService:
return ProviderService()
def get_provider_service(
x_ai_provider: str | None = Header(default=None),
x_ai_provider_api: str | None = Header(default=None),
x_ai_base_url: str | None = Header(default=None),
x_ai_api_key: str | None = Header(default=None),
x_ai_model: str | None = Header(default=None),
x_ai_max_tokens: str | None = Header(default=None),
x_ai_anthropic_version: str | None = Header(default=None),
) -> ProviderService:
overrides = {
"provider": x_ai_provider,
"provider_api": x_ai_provider_api,
"base_url": x_ai_base_url,
"api_key": x_ai_api_key,
"model": x_ai_model,
"anthropic_version": x_ai_anthropic_version,
}
if x_ai_max_tokens:
overrides["max_tokens"] = x_ai_max_tokens
return ProviderService({key: value for key, value in overrides.items() if value not in (None, "")})
@app.get("/health")

View File

@@ -46,19 +46,22 @@ def _resolve_provider_api(provider: str, configured_api: str) -> str:
class ProviderService:
def __init__(self) -> None:
self.provider = _normalize_provider(settings.AI_PROVIDER)
def __init__(self, overrides: dict[str, Any] | None = None) -> None:
overrides = overrides or {}
self.provider = _normalize_provider(overrides.get("provider") or settings.AI_PROVIDER)
self.provider_api = _resolve_provider_api(
self.provider,
_normalize_provider_api(settings.AI_PROVIDER_API),
_normalize_provider_api(overrides.get("provider_api") or settings.AI_PROVIDER_API),
)
self.base_url = settings.AI_BASE_URL.rstrip("/")
self.api_key = settings.AI_API_KEY
self.default_model = settings.AI_MODEL
self.base_url = str(overrides.get("base_url") or settings.AI_BASE_URL).rstrip("/")
self.api_key = str(overrides.get("api_key") or settings.AI_API_KEY)
self.default_model = str(overrides.get("model") or settings.AI_MODEL)
self.timeout = settings.AI_TIMEOUT_SECONDS
self.http_retry_attempts = max(settings.AI_HTTP_RETRY_ATTEMPTS, 1)
self.max_tokens = settings.AI_MAX_TOKENS
self.anthropic_version = settings.AI_ANTHROPIC_VERSION
self.max_tokens = int(overrides.get("max_tokens") or settings.AI_MAX_TOKENS)
self.anthropic_version = str(
overrides.get("anthropic_version") or settings.AI_ANTHROPIC_VERSION
)
self.system_prompt = settings.AI_ANALYSIS_SYSTEM_PROMPT
def get_status(self) -> AIProviderStatusResponse:

View File

@@ -1,6 +1,10 @@
FROM python:3.14-slim
ARG PYTHON_IMAGE=python:3.14-slim
ARG UV_IMAGE=ghcr.io/astral-sh/uv:latest
COPY --from=ghcr.io/astral-sh/uv:latest /uv /uvx /bin/
FROM ${UV_IMAGE} AS uv
FROM ${PYTHON_IMAGE}
COPY --from=uv /uv /uvx /bin/
WORKDIR /app

View File

@@ -1,20 +1,34 @@
"""DataSourceConfig API for user-defined data sources"""
from typing import Optional
from typing import Any, Optional
from datetime import datetime
import base64
import json
import re
from fastapi import APIRouter, Depends, HTTPException, status
from sqlalchemy import select, func
from sqlalchemy.ext.asyncio import AsyncSession
from pydantic import BaseModel, Field
import httpx
from app.core.target_schema_registry import get_target_schema, list_target_schemas
from app.db.session import get_db
from app.models.user import User
from app.models.datasource_config import DataSourceConfig
from app.models.datasource_mapping import DataSourceMappingTemplate
from app.core.security import get_current_user
from app.core.cache import cache
from app.core.time import to_iso8601_utc
from app.schemas.ai import SituationalAnalysisRequest
from app.services.ai_client import AIProviderClient, get_ai_provider_client
from app.services.datasource_mapping import (
MappingError,
build_heuristic_mapping,
execute_mapping,
persist_mapped_records,
redact_for_llm,
stable_payload_hash,
)
router = APIRouter()
@@ -59,6 +73,44 @@ class DataSourceConfigResponse(BaseModel):
from_attributes = True
class CustomSampleRequest(BaseModel):
datasource_config_id: Optional[int] = None
config: Optional[DataSourceConfigCreate] = None
limit_bytes: int = Field(default=200000, ge=1000, le=1000000)
class MappingProposeRequest(BaseModel):
sample_payload: Any
target_schema: str
use_ai: bool = True
class MappingPreviewRequest(BaseModel):
sample_payload: Any
target_schema: str
mapping_json: dict
limit: int = Field(default=20, ge=1, le=100)
class MappingTemplateCreate(BaseModel):
datasource_config_id: int
target_schema: str
mapping_json: dict
sample_payload: Any | None = None
sample_payload_hash: Optional[str] = None
validation_status: str = Field(default="draft", pattern="^(draft|valid|invalid)$")
is_active: bool = False
class MappingTemplateUpdate(BaseModel):
target_schema: Optional[str] = None
mapping_json: Optional[dict] = None
sample_payload: Any | None = None
sample_payload_hash: Optional[str] = None
validation_status: Optional[str] = Field(default=None, pattern="^(draft|valid|invalid)$")
is_active: Optional[bool] = None
async def test_endpoint(
endpoint: str,
auth_type: str,
@@ -96,6 +148,134 @@ async def test_endpoint(
}
def _build_request_headers(auth_type: str, auth_config: dict, headers: dict) -> dict[str, str]:
request_headers = {str(key): str(value) for key, value in (headers or {}).items()}
auth_type = str(auth_type or "none").lower()
auth_config = auth_config or {}
if auth_type == "bearer" and auth_config.get("token"):
request_headers["Authorization"] = f"Bearer {auth_config['token']}"
elif auth_type == "api_key" and auth_config.get("api_key"):
location = str(auth_config.get("in") or auth_config.get("location") or "header").lower()
if location != "query":
key_name = auth_config.get("key_name", "X-API-Key")
request_headers[str(key_name)] = str(auth_config["api_key"])
elif auth_type == "basic":
username = auth_config.get("username", "")
password = auth_config.get("password", "")
credentials = f"{username}:{password}"
encoded = base64.b64encode(credentials.encode()).decode()
request_headers["Authorization"] = f"Basic {encoded}"
return request_headers
def _build_query_params(auth_type: str, auth_config: dict, config: dict) -> dict[str, Any]:
params = {}
candidate = (config or {}).get("params") or (config or {}).get("query_params")
if isinstance(candidate, dict):
params.update(candidate)
auth_type = str(auth_type or "none").lower()
auth_config = auth_config or {}
if auth_type == "api_key" and auth_config.get("api_key"):
location = str(auth_config.get("in") or auth_config.get("location") or "header").lower()
if location == "query":
key_name = auth_config.get("key_name") or auth_config.get("param_name") or "api_key"
params[str(key_name)] = auth_config["api_key"]
return params
async def fetch_custom_sample_from_config(config: DataSourceConfig, limit_bytes: int) -> Any:
request_config = config.config or {}
method = str(request_config.get("method") or request_config.get("request_method") or "GET").upper()
if method not in {"GET", "POST"}:
raise HTTPException(status_code=400, detail="Only GET and POST sample requests are supported.")
headers = _build_request_headers(config.auth_type, config.auth_config or {}, config.headers or {})
params = _build_query_params(config.auth_type, config.auth_config or {}, request_config)
timeout = float(request_config.get("timeout", 30))
json_body = request_config.get("json_body")
if json_body is None and str(request_config.get("body_type") or "").lower() in {"json", ""}:
candidate = request_config.get("body")
if isinstance(candidate, (dict, list)):
json_body = candidate
async with httpx.AsyncClient(timeout=timeout, follow_redirects=True) as client:
response = await client.request(
method,
config.endpoint,
headers=headers,
params=params or None,
json=json_body,
)
response.raise_for_status()
content = response.content[:limit_bytes]
if "application/json" in response.headers.get("content-type", ""):
return json.loads(content.decode(response.encoding or "utf-8"))
return {"text": content.decode(response.encoding or "utf-8", errors="replace")}
def _parse_mapping_from_ai_text(content: str) -> dict[str, Any] | None:
if not content:
return None
candidates = [content]
fenced = re.findall(r"```(?:json)?\s*(\{.*?\})\s*```", content, flags=re.DOTALL)
candidates = fenced + candidates
for candidate in candidates:
try:
parsed = json.loads(candidate)
except json.JSONDecodeError:
continue
if isinstance(parsed, dict) and isinstance(parsed.get("fields"), dict):
return parsed
return None
async def _get_config_for_sample(
payload: CustomSampleRequest,
db: AsyncSession,
) -> DataSourceConfig:
if payload.datasource_config_id is not None:
result = await db.execute(
select(DataSourceConfig).where(DataSourceConfig.id == payload.datasource_config_id)
)
config = result.scalar_one_or_none()
if not config:
raise HTTPException(status_code=404, detail="Configuration not found")
return config
if payload.config is None:
raise HTTPException(status_code=400, detail="datasource_config_id or config is required")
config_data = payload.config
return DataSourceConfig(
name=config_data.name,
description=config_data.description,
source_type=config_data.source_type,
endpoint=config_data.endpoint,
auth_type=config_data.auth_type,
auth_config=config_data.auth_config,
headers=config_data.headers,
config=config_data.config,
)
def serialize_mapping_template(template: DataSourceMappingTemplate) -> dict[str, Any]:
return {
"id": template.id,
"datasource_config_id": template.datasource_config_id,
"target_schema": template.target_schema,
"mapping_json": template.mapping_json,
"sample_payload_hash": template.sample_payload_hash,
"validation_status": template.validation_status,
"version": template.version,
"is_active": template.is_active,
"created_at": to_iso8601_utc(template.created_at),
"updated_at": to_iso8601_utc(template.updated_at),
}
@router.get("/configs")
async def list_configs(
active_only: bool = False,
@@ -345,3 +525,311 @@ async def list_all_datasources(
)
return {"total": len(result), "data": result}
@router.post("/custom/sample")
async def fetch_custom_sample(
payload: CustomSampleRequest,
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
"""Fetch a sample payload for a saved or draft custom data source."""
config = await _get_config_for_sample(payload, db)
try:
sample = await fetch_custom_sample_from_config(config, payload.limit_bytes)
except httpx.HTTPStatusError as exc:
raise HTTPException(
status_code=exc.response.status_code,
detail=f"Sample request failed: HTTP {exc.response.status_code}",
) from exc
except httpx.HTTPError as exc:
raise HTTPException(status_code=502, detail=f"Sample request failed: {exc}") from exc
return {
"success": True,
"sample_payload": sample,
"sample_payload_hash": stable_payload_hash(sample),
"redacted_preview": redact_for_llm(sample),
}
@router.get("/target-schemas")
async def get_datasource_target_schemas(
current_user: User = Depends(get_current_user),
):
"""List target schemas available for custom datasource mapping."""
return {"data": list_target_schemas()}
@router.post("/mappings/propose")
async def propose_datasource_mapping(
payload: MappingProposeRequest,
current_user: User = Depends(get_current_user),
ai_client: AIProviderClient = Depends(get_ai_provider_client),
):
"""Generate a mapping draft for a sample payload and target schema."""
schema = get_target_schema(payload.target_schema)
redacted_sample = redact_for_llm(payload.sample_payload)
fallback_mapping = build_heuristic_mapping(redacted_sample, payload.target_schema)
ai_error: str | None = None
mapping = fallback_mapping
generated_by = "heuristic"
if payload.use_ai:
try:
response = await ai_client.analyze(
SituationalAnalysisRequest(
title=f"Generate datasource mapping for {schema.key}",
objective=(
"Return only JSON for a deterministic mapping DSL. "
"The JSON must contain source.items_path and fields. "
"Do not include prose or code."
),
context={
"target_schema": schema.to_dict(),
"sample_payload": redacted_sample,
"mapping_dsl_example": fallback_mapping,
},
observations=[
"Use JSONPath-like paths beginning with $.",
"Never generate executable code.",
"Use field types from the target schema.",
],
constraints=[
"Return a single JSON object.",
"Do not include credentials or secrets.",
"Mark uncertain optional fields with default null.",
],
)
)
parsed = _parse_mapping_from_ai_text(response.content)
if parsed:
mapping = parsed
generated_by = "ai_provider"
else:
ai_error = "AI provider did not return a valid mapping JSON object."
except HTTPException as exc:
ai_error = str(exc.detail)
mapping.setdefault("meta", {})
if isinstance(mapping["meta"], dict):
mapping["meta"].update(
{
"generated_by": generated_by,
"requires_review": True,
"ai_error": ai_error,
}
)
return {
"target_schema": schema.to_dict(),
"mapping_json": mapping,
"sample_payload_hash": stable_payload_hash(payload.sample_payload),
"redacted_sample_payload": redacted_sample,
}
@router.post("/mappings/preview")
async def preview_datasource_mapping(
payload: MappingPreviewRequest,
current_user: User = Depends(get_current_user),
):
"""Preview deterministic mapping output for a sample payload."""
try:
preview = execute_mapping(
payload.sample_payload,
payload.mapping_json,
payload.target_schema,
limit=payload.limit,
)
except (MappingError, ValueError) as exc:
raise HTTPException(status_code=400, detail=str(exc)) from exc
return {
"success": preview["failed_count"] == 0,
"preview": preview,
"sample_payload_hash": stable_payload_hash(payload.sample_payload),
}
@router.get("/mappings")
async def list_datasource_mappings(
datasource_config_id: Optional[int] = None,
active_only: bool = False,
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
"""List saved mapping templates."""
query = select(DataSourceMappingTemplate).order_by(
DataSourceMappingTemplate.datasource_config_id,
DataSourceMappingTemplate.version.desc(),
)
if datasource_config_id is not None:
query = query.where(DataSourceMappingTemplate.datasource_config_id == datasource_config_id)
if active_only:
query = query.where(DataSourceMappingTemplate.is_active.is_(True))
result = await db.execute(query)
mappings = result.scalars().all()
return {"total": len(mappings), "data": [serialize_mapping_template(item) for item in mappings]}
@router.post("/mappings")
async def create_datasource_mapping(
payload: MappingTemplateCreate,
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
"""Save a mapping template for a datasource config."""
get_target_schema(payload.target_schema)
datasource = await db.get(DataSourceConfig, payload.datasource_config_id)
if not datasource:
raise HTTPException(status_code=404, detail="Configuration not found")
if payload.sample_payload is not None:
try:
execute_mapping(payload.sample_payload, payload.mapping_json, payload.target_schema, limit=100)
except (MappingError, ValueError) as exc:
raise HTTPException(status_code=400, detail=f"Mapping validation failed: {exc}") from exc
result = await db.execute(
select(func.max(DataSourceMappingTemplate.version)).where(
DataSourceMappingTemplate.datasource_config_id == payload.datasource_config_id,
DataSourceMappingTemplate.target_schema == payload.target_schema,
)
)
next_version = int(result.scalar() or 0) + 1
if payload.is_active:
await db.execute(
DataSourceMappingTemplate.__table__.update()
.where(DataSourceMappingTemplate.datasource_config_id == payload.datasource_config_id)
.values(is_active=False)
)
template = DataSourceMappingTemplate(
datasource_config_id=payload.datasource_config_id,
target_schema=payload.target_schema,
mapping_json=payload.mapping_json,
sample_payload_hash=payload.sample_payload_hash
or (stable_payload_hash(payload.sample_payload) if payload.sample_payload is not None else None),
validation_status=payload.validation_status,
version=next_version,
is_active=payload.is_active,
)
db.add(template)
await db.commit()
await db.refresh(template)
return {"message": "Mapping template saved successfully", "data": serialize_mapping_template(template)}
@router.put("/mappings/{mapping_id}")
async def update_datasource_mapping(
mapping_id: int,
payload: MappingTemplateUpdate,
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
"""Update a mapping template in place."""
template = await db.get(DataSourceMappingTemplate, mapping_id)
if not template:
raise HTTPException(status_code=404, detail="Mapping template not found")
target_schema = payload.target_schema or template.target_schema
mapping_json = payload.mapping_json or template.mapping_json
get_target_schema(target_schema)
if payload.sample_payload is not None:
try:
execute_mapping(payload.sample_payload, mapping_json, target_schema, limit=100)
except (MappingError, ValueError) as exc:
raise HTTPException(status_code=400, detail=f"Mapping validation failed: {exc}") from exc
if payload.is_active is True:
await db.execute(
DataSourceMappingTemplate.__table__.update()
.where(DataSourceMappingTemplate.datasource_config_id == template.datasource_config_id)
.where(DataSourceMappingTemplate.id != template.id)
.values(is_active=False)
)
template.target_schema = target_schema
template.mapping_json = mapping_json
if payload.sample_payload_hash is not None:
template.sample_payload_hash = payload.sample_payload_hash
elif payload.sample_payload is not None:
template.sample_payload_hash = stable_payload_hash(payload.sample_payload)
if payload.validation_status is not None:
template.validation_status = payload.validation_status
if payload.is_active is not None:
template.is_active = payload.is_active
await db.commit()
await db.refresh(template)
return {"message": "Mapping template updated successfully", "data": serialize_mapping_template(template)}
@router.post("/{config_id}/run-mapped")
async def run_mapped_datasource(
config_id: int,
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
"""Run a saved custom datasource through its active deterministic mapping."""
datasource = await db.get(DataSourceConfig, config_id)
if not datasource:
raise HTTPException(status_code=404, detail="Configuration not found")
result = await db.execute(
select(DataSourceMappingTemplate)
.where(DataSourceMappingTemplate.datasource_config_id == config_id)
.where(DataSourceMappingTemplate.is_active.is_(True))
.order_by(DataSourceMappingTemplate.version.desc())
.limit(1)
)
mapping = result.scalar_one_or_none()
if not mapping:
raise HTTPException(status_code=404, detail="No active mapping template found")
try:
sample = await fetch_custom_sample_from_config(datasource, 5_000_000)
mapped = execute_mapping(sample, mapping.mapping_json, mapping.target_schema)
except httpx.HTTPStatusError as exc:
raise HTTPException(
status_code=exc.response.status_code,
detail=f"Datasource request failed: HTTP {exc.response.status_code}",
) from exc
except httpx.HTTPError as exc:
raise HTTPException(status_code=502, detail=f"Datasource request failed: {exc}") from exc
except (MappingError, ValueError) as exc:
raise HTTPException(status_code=400, detail=f"Mapping failed: {exc}") from exc
if mapped["failed_count"] > 0:
return {
"status": "failed",
"datasource_config_id": config_id,
"mapping_id": mapping.id,
"mapping_version": mapping.version,
"target_schema": mapping.target_schema,
"mapped_count": mapped["mapped_count"],
"failed_count": mapped["failed_count"],
"errors": mapped["errors"][:20],
}
written_count = await persist_mapped_records(
db,
datasource_name=datasource.name,
datasource_config_id=datasource.id,
target_schema=mapping.target_schema,
records=mapped["records"],
mapping_version=mapping.version,
)
return {
"status": "success",
"datasource_config_id": config_id,
"mapping_id": mapping.id,
"mapping_version": mapping.version,
"target_schema": mapping.target_schema,
"fetched_count": mapped["total_items"],
"mapped_count": mapped["mapped_count"],
"written_count": written_count,
}

View File

@@ -9,6 +9,7 @@ from sqlalchemy.ext.asyncio import AsyncSession
from app.core.time import to_iso8601_utc
from app.core.security import get_current_user
from app.core.data_sources import get_data_sources_config
from app.core.datasource_defaults import DEFAULT_DATASOURCES
from app.db.session import get_db
from app.models.collected_data import CollectedData
from app.models.data_snapshot import DataSnapshot
@@ -35,6 +36,17 @@ def format_frequency_label(minutes: int) -> str:
return f"{minutes}m"
def datasource_metadata(source: str) -> dict:
info = DEFAULT_DATASOURCES.get(source, {})
return {
"display_name": info.get("display_name") or info.get("name") or source,
"is_free": bool(info.get("is_free", True)),
"requires_credentials": bool(info.get("requires_credentials", False)),
"credential_provider": info.get("credential_provider"),
"credential_status": info.get("credential_status", "none"),
}
def is_due_for_collection(datasource: DataSource, now: datetime) -> bool:
if datasource.last_run_at is None:
return True
@@ -72,31 +84,6 @@ async def _load_latest_running_tasks(
return {task.datasource_id: task for task in result.scalars().all()}
async def _load_latest_completed_tasks(
db: AsyncSession,
datasource_ids: list[int],
) -> dict[int, CollectionTask]:
if not datasource_ids:
return {}
ranked_tasks = (
select(
CollectionTask.id.label("task_id"),
_task_rank_column(CollectionTask.completed_at),
)
.where(CollectionTask.datasource_id.in_(datasource_ids))
.where(CollectionTask.completed_at.isnot(None))
.where(CollectionTask.status.in_(("success", "failed", "cancelled")))
.subquery()
)
result = await db.execute(
select(CollectionTask)
.join(ranked_tasks, CollectionTask.id == ranked_tasks.c.task_id)
.where(ranked_tasks.c.row_num == 1)
)
return {task.datasource_id: task for task in result.scalars().all()}
async def _load_latest_task_ids(
db: AsyncSession,
datasource_ids: list[int],
@@ -123,21 +110,6 @@ async def _load_latest_task_ids(
return {datasource_id: task_id for datasource_id, task_id in result.all()}
async def _load_datasource_data_counts(
db: AsyncSession,
sources: list[str],
) -> dict[str, int]:
if not sources:
return {}
result = await db.execute(
select(CollectedData.source, func.count(CollectedData.id))
.where(CollectedData.source.in_(sources))
.group_by(CollectedData.source)
)
return {source: count for source, count in result.all()}
async def _load_datasource_endpoint_overrides(
db: AsyncSession,
sources: list[str],
@@ -161,7 +133,7 @@ async def _load_datasource_endpoint_overrides(
async def _load_datasource_list_context(
db: AsyncSession,
datasources: list[DataSource],
) -> tuple[dict[int, CollectionTask], dict[int, CollectionTask], dict[str, int], dict[str, str]]:
) -> tuple[dict[int, CollectionTask], dict[str, str]]:
datasource_ids = [datasource.id for datasource in datasources]
sources = [datasource.source for datasource in datasources]
@@ -185,10 +157,8 @@ async def _load_datasource_list_context(
if stale_datasource_ids:
running_tasks = await _load_latest_running_tasks(db, datasource_ids)
completed_tasks = await _load_latest_completed_tasks(db, datasource_ids)
data_counts = await _load_datasource_data_counts(db, sources)
endpoint_overrides = await _load_datasource_endpoint_overrides(db, sources)
return running_tasks, completed_tasks, data_counts, endpoint_overrides
return running_tasks, endpoint_overrides
async def get_datasource_record(db: AsyncSession, source_id: str) -> Optional[DataSource]:
@@ -401,26 +371,19 @@ async def list_datasources(
collector_list = []
config = get_data_sources_config()
running_tasks, completed_tasks, data_counts, endpoint_overrides = await _load_datasource_list_context(
db,
datasources,
)
running_tasks, endpoint_overrides = await _load_datasource_list_context(db, datasources)
for datasource in datasources:
running_task = running_tasks.get(datasource.id)
last_task = completed_tasks.get(datasource.id)
endpoint = endpoint_overrides.get(datasource.source) or config.get_yaml_url(
datasource.source,
)
data_count = data_counts.get(datasource.source, 0)
last_run_at = datasource.last_run_at or (last_task.completed_at if last_task else None)
last_run = to_iso8601_utc(last_run_at)
last_status = datasource.last_status or (last_task.status if last_task else None)
endpoint = endpoint_overrides.get(datasource.source) or config.get_yaml_url(datasource.source)
last_run_at = datasource.last_run_at
last_status = datasource.last_status
collector_list.append(
{
"id": datasource.id,
"source": datasource.source,
"name": datasource.name,
**datasource_metadata(datasource.source),
"module": datasource.module,
"priority": datasource.priority,
"frequency": format_frequency_label(datasource.frequency_minutes),
@@ -428,11 +391,9 @@ async def list_datasources(
"is_active": datasource.is_active,
"collector_class": datasource.collector_class,
"endpoint": endpoint,
"last_run": last_run,
"last_run": to_iso8601_utc(last_run_at),
"last_run_at": to_iso8601_utc(last_run_at),
"last_status": last_status,
"last_records_processed": last_task.records_processed if last_task else None,
"data_count": data_count,
"is_running": running_task is not None,
"task_id": running_task.id if running_task else None,
"progress": running_task.progress if running_task else None,
@@ -575,6 +536,7 @@ async def get_datasource(
return {
"id": datasource.id,
"name": datasource.name,
**datasource_metadata(datasource.source),
"module": datasource.module,
"priority": datasource.priority,
"frequency": format_frequency_label(datasource.frequency_minutes),

View File

@@ -9,10 +9,19 @@ from sqlalchemy.ext.asyncio import AsyncSession
from app.core.security import get_current_user
from app.core.time import to_iso8601_utc
from app.core.config import settings as app_settings
from app.core.data_sources import get_data_sources_config
from app.core.datasource_defaults import DEFAULT_DATASOURCES
from app.db.session import get_db
from app.models.datasource import DataSource
from app.models.datasource_config import DataSourceConfig
from app.models.system_setting import SystemSetting
from app.models.user import User
from app.services.llm_provider_catalog import (
get_fallback_llm_provider_preset,
list_fallback_llm_provider_presets,
refresh_llm_provider_preset,
)
from app.services.scheduler import sync_datasource_job
from app.services.tv_streams import DEFAULT_TV_SETTINGS, get_tv_settings_payload, normalize_tv_settings
@@ -39,6 +48,21 @@ DEFAULT_SETTINGS = {
"password_policy": "medium",
},
"tv": DEFAULT_TV_SETTINGS,
"external_integrations": {
"ai_provider": {
"service_url": "",
"service_token": "",
"provider": "minimax",
"provider_api": "anthropic-messages",
"base_url": "https://api.minimaxi.com/anthropic",
"model": "MiniMax-M2.7",
"api_key": "",
"max_tokens": 1200,
"anthropic_version": "2023-06-01",
"timeout_seconds": 60,
"retry_attempts": 2,
}
},
}
@@ -96,6 +120,34 @@ class TVSettingsUpdate(BaseModel):
sources: list[TVStreamSourceUpdate] = Field(default_factory=list)
class AIProviderIntegrationUpdate(BaseModel):
service_url: str = ""
service_token: Optional[str] = None
provider: str = Field(default="minimax", max_length=80)
provider_api: str = Field(default="anthropic-messages", max_length=80)
base_url: str = Field(default="", max_length=500)
model: str = Field(default="", max_length=200)
api_key: Optional[str] = None
max_tokens: int = Field(default=1200, ge=1, le=200000)
anthropic_version: str = Field(default="2023-06-01", max_length=40)
timeout_seconds: int = Field(default=60, ge=5, le=600)
retry_attempts: int = Field(default=2, ge=1, le=10)
clear_service_token: bool = False
clear_api_key: bool = False
class BarentsWatchIntegrationUpdate(BaseModel):
endpoint: str = ""
client_id: str = ""
client_secret: Optional[str] = None
clear_client_secret: bool = False
class ExternalIntegrationsUpdate(BaseModel):
ai_provider: AIProviderIntegrationUpdate
barentswatch: BarentsWatchIntegrationUpdate
def merge_with_defaults(category: str, payload: Optional[dict]) -> dict:
merged = deepcopy(DEFAULT_SETTINGS[category])
if payload:
@@ -146,6 +198,158 @@ async def save_setting_payload(db: AsyncSession, category: str, payload: dict) -
return merge_with_defaults(category, record.payload)
def _mask_secret(value: Optional[str]) -> dict:
if not value:
return {"configured": False, "preview": ""}
text = str(value)
if "-" in text:
prefix = text.split("-", 1)[0] + "-"
preview = prefix + ("*" * max(len(text) - len(prefix), 1))
else:
prefix_len = min(4, len(text))
preview = text[:prefix_len] + ("*" * max(len(text) - prefix_len, 1))
return {"configured": True, "preview": preview}
async def get_runtime_ai_provider_config(db: AsyncSession) -> dict:
runtime_record = await get_setting_record(db, "external_integrations")
payload = merge_with_defaults(
"external_integrations",
runtime_record.payload if runtime_record else None,
)
ai_payload = payload.get("ai_provider") or {}
has_runtime_llm_config = bool(
runtime_record
and isinstance(runtime_record.payload, dict)
and isinstance(runtime_record.payload.get("ai_provider"), dict)
)
return {
"service_url": ai_payload.get("service_url") or app_settings.AI_PROVIDER_SERVICE_URL,
"service_token": ai_payload.get("service_token") or app_settings.AI_PROVIDER_SERVICE_TOKEN,
"timeout_seconds": int(
ai_payload.get("timeout_seconds") or app_settings.AI_PROVIDER_TIMEOUT_SECONDS
),
"retry_attempts": int(
ai_payload.get("retry_attempts") or app_settings.AI_PROVIDER_RETRY_ATTEMPTS
),
"llm_config": {
"provider": ai_payload.get("provider") or "minimax",
"provider_api": ai_payload.get("provider_api") or "anthropic-messages",
"base_url": ai_payload.get("base_url") or "https://api.minimaxi.com/anthropic",
"model": ai_payload.get("model") or "MiniMax-M2.7",
"api_key": ai_payload.get("api_key") or "",
"max_tokens": int(ai_payload.get("max_tokens") or 1200),
"anthropic_version": ai_payload.get("anthropic_version") or "2023-06-01",
} if has_runtime_llm_config else {},
}
async def get_barentswatch_config_record(db: AsyncSession) -> Optional[DataSourceConfig]:
result = await db.execute(
select(DataSourceConfig)
.where(DataSourceConfig.name == "barentswatch_vessels")
.where(DataSourceConfig.is_active.is_(True))
)
return result.scalar_one_or_none()
async def serialize_external_integrations(db: AsyncSession) -> dict:
ai_config = await get_runtime_ai_provider_config(db)
runtime_setting = await get_setting_record(db, "external_integrations")
display_llm_config = ai_config["llm_config"] or DEFAULT_SETTINGS["external_integrations"]["ai_provider"]
barentswatch_record = await get_barentswatch_config_record(db)
yaml_config = get_data_sources_config()
barentswatch_auth = barentswatch_record.auth_config if barentswatch_record else {}
barentswatch_auth = barentswatch_auth or {}
return {
"ai_provider": {
"service_url": ai_config["service_url"],
"service_token": _mask_secret(ai_config["service_token"]),
"provider": display_llm_config.get("provider") or "minimax",
"provider_api": display_llm_config.get("provider_api") or "anthropic-messages",
"base_url": display_llm_config.get("base_url") or "https://api.minimaxi.com/anthropic",
"model": display_llm_config.get("model") or "MiniMax-M2.7",
"api_key": _mask_secret(display_llm_config.get("api_key")),
"max_tokens": int(display_llm_config.get("max_tokens") or 1200),
"anthropic_version": display_llm_config.get("anthropic_version") or "2023-06-01",
"timeout_seconds": ai_config["timeout_seconds"],
"retry_attempts": ai_config["retry_attempts"],
"source": "runtime" if runtime_setting else "env",
},
"barentswatch": {
"endpoint": (
barentswatch_record.endpoint
if barentswatch_record and barentswatch_record.endpoint
else yaml_config.get_yaml_url("barentswatch_vessels")
),
"client_id": barentswatch_auth.get("client_id") or "",
"client_secret": _mask_secret(barentswatch_auth.get("client_secret")),
"source": "datasource_config" if barentswatch_record else "default",
},
}
async def save_external_integrations_payload(
db: AsyncSession,
update: ExternalIntegrationsUpdate,
) -> dict:
current_payload = await get_setting_payload(db, "external_integrations")
current_ai = current_payload.get("ai_provider") or {}
ai_payload = {
"service_url": update.ai_provider.service_url.strip()
or app_settings.AI_PROVIDER_SERVICE_URL,
"service_token": current_ai.get("service_token") or "",
"provider": update.ai_provider.provider.strip() or "minimax",
"provider_api": update.ai_provider.provider_api.strip() or "anthropic-messages",
"base_url": update.ai_provider.base_url.strip(),
"model": update.ai_provider.model.strip(),
"api_key": current_ai.get("api_key") or "",
"max_tokens": update.ai_provider.max_tokens,
"anthropic_version": update.ai_provider.anthropic_version.strip() or "2023-06-01",
"timeout_seconds": update.ai_provider.timeout_seconds,
"retry_attempts": update.ai_provider.retry_attempts,
}
if update.ai_provider.clear_service_token:
ai_payload["service_token"] = ""
elif update.ai_provider.service_token not in (None, ""):
ai_payload["service_token"] = update.ai_provider.service_token
if update.ai_provider.clear_api_key:
ai_payload["api_key"] = ""
elif update.ai_provider.api_key not in (None, ""):
ai_payload["api_key"] = update.ai_provider.api_key
await save_setting_payload(db, "external_integrations", {"ai_provider": ai_payload})
default_endpoint = get_data_sources_config().get_yaml_url("barentswatch_vessels")
barentswatch_record = await get_barentswatch_config_record(db)
if barentswatch_record is None:
barentswatch_record = DataSourceConfig(
name="barentswatch_vessels",
description="BarentsWatch Live AIS credentials",
source_type="api",
endpoint=update.barentswatch.endpoint.strip() or default_endpoint,
auth_type="oauth_client_credentials",
auth_config={},
headers={},
config={},
is_active=True,
)
db.add(barentswatch_record)
current_auth = dict(barentswatch_record.auth_config or {})
if update.barentswatch.clear_client_secret:
current_auth.pop("client_secret", None)
elif update.barentswatch.client_secret not in (None, ""):
current_auth["client_secret"] = update.barentswatch.client_secret
current_auth["client_id"] = update.barentswatch.client_id.strip()
barentswatch_record.endpoint = update.barentswatch.endpoint.strip() or default_endpoint
barentswatch_record.auth_type = "oauth_client_credentials"
barentswatch_record.auth_config = current_auth
await db.commit()
return await serialize_external_integrations(db)
def format_frequency_label(minutes: int) -> str:
if minutes % 1440 == 0:
return f"{minutes // 1440}d"
@@ -155,9 +359,11 @@ def format_frequency_label(minutes: int) -> str:
def serialize_collector(datasource: DataSource) -> dict:
defaults = DEFAULT_DATASOURCES.get(datasource.source, {})
return {
"id": datasource.id,
"name": datasource.name,
"display_name": defaults.get("display_name") or datasource.name,
"source": datasource.source,
"module": datasource.module,
"priority": datasource.priority,
@@ -167,6 +373,10 @@ def serialize_collector(datasource: DataSource) -> dict:
"last_run_at": to_iso8601_utc(datasource.last_run_at),
"last_status": datasource.last_status,
"next_run_at": to_iso8601_utc(datasource.next_run_at),
"is_free": bool(defaults.get("is_free", True)),
"requires_credentials": bool(defaults.get("requires_credentials", False)),
"credential_provider": defaults.get("credential_provider"),
"credential_status": defaults.get("credential_status", "none"),
}
@@ -243,6 +453,46 @@ async def update_tv_settings(
return {"status": "updated", "tv": normalize_tv_settings(saved)}
@router.get("/integrations")
async def get_external_integrations(
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
return {"integrations": await serialize_external_integrations(db)}
@router.get("/integrations/ai-provider/presets")
async def get_ai_provider_presets(
current_user: User = Depends(get_current_user),
):
return {"data": list_fallback_llm_provider_presets()}
@router.post("/integrations/ai-provider/presets/{provider}/refresh")
async def refresh_ai_provider_preset(
provider: str,
current_user: User = Depends(get_current_user),
):
try:
return {"data": await refresh_llm_provider_preset(provider)}
except ValueError as exc:
raise HTTPException(status_code=404, detail=str(exc)) from exc
except Exception as exc:
fallback = get_fallback_llm_provider_preset(provider)
fallback["refresh_error"] = str(exc)
return {"data": fallback}
@router.put("/integrations")
async def update_external_integrations(
payload: ExternalIntegrationsUpdate,
current_user: User = Depends(get_current_user),
db: AsyncSession = Depends(get_db),
):
saved = await save_external_integrations_payload(db, payload)
return {"status": "updated", "integrations": saved}
@router.get("/collectors")
async def get_collector_settings(
current_user: User = Depends(get_current_user),
@@ -289,6 +539,7 @@ async def get_all_settings(
"notifications": setting_payloads["notifications"],
"security": setting_payloads["security"],
"tv": await get_tv_settings_payload(db),
"integrations": await serialize_external_integrations(db),
"collectors": [serialize_collector(datasource) for datasource in datasources],
"generated_at": to_iso8601_utc(datetime.now(UTC)),
}

View File

@@ -4,12 +4,15 @@ import os
import subprocess
import sys
from fastapi import APIRouter, Depends, HTTPException, status
from datetime import datetime
from fastapi import APIRouter, Depends, HTTPException, Query, Request, status
from pydantic import BaseModel
from app.core.config import ROOT_DIR
from app.core.security import get_current_user
from app.models.user import User
from app.services.persistent_logs import record_audit_log, record_system_log
from app.services.system_control import (
build_task_id,
clear_active_task_id,
@@ -23,6 +26,15 @@ from app.services.system_control import (
set_active_task_id,
upsert_task_state,
)
from app.services.system_logs import (
DEFAULT_LOG_LINE_LIMIT,
MAX_LOG_LINE_LIMIT,
SUPPORTED_LOG_LEVELS,
append_buffer_log,
list_log_sources,
normalize_log_level,
read_log_snapshot,
)
router = APIRouter()
@@ -47,6 +59,59 @@ class RestartTaskLogsResponse(BaseModel):
lines: list[str]
class SystemLogSourceSummary(BaseModel):
source_id: str
name: str
kind: str
location: str
description: str
category: str
status: str
class SystemLogSourcesResponse(BaseModel):
items: list[SystemLogSourceSummary]
class SystemLogDailyMarker(BaseModel):
date_token: str
total: int
dominant_level: str
class SystemLogSnapshotResponse(BaseModel):
source_id: str
name: str
kind: str
location: str
description: str
category: str
status: str
level: str
selected_levels: list[str] = []
search_query: str = ""
available_levels: list[str]
daily_markers: list[SystemLogDailyMarker] = []
line_limit: int
line_count: int
lines: list[str]
class EarthClientLogEventCreate(BaseModel):
level: str = "error"
message: str
category: str | None = None
url: str | None = None
module: str | None = None
detail: str | None = None
class EarthClientLogEventResponse(BaseModel):
accepted: bool
source_id: str
level: str
def ensure_super_admin(current_user: User) -> None:
if not require_super_admin(current_user.role):
raise HTTPException(
@@ -55,9 +120,22 @@ def ensure_super_admin(current_user: User) -> None:
)
def validate_log_date(raw_value: str | None, field_name: str) -> str | None:
if raw_value in {None, ""}:
return None
try:
return datetime.strptime(raw_value, "%Y-%m-%d").date().isoformat()
except ValueError as exc:
raise HTTPException(
status_code=status.HTTP_400_BAD_REQUEST,
detail=f"{field_name} must be in YYYY-MM-DD format",
) from exc
@router.post("/restart-tasks", response_model=RestartTaskResponse)
async def create_restart_task(
payload: RestartTaskCreate,
request: Request,
current_user: User = Depends(get_current_user),
):
ensure_super_admin(current_user)
@@ -133,11 +211,31 @@ async def create_restart_task(
requested_by=requested_by,
)
clear_active_task_id(task_id)
await record_audit_log(
action="system.restart_task.requested",
actor_id=current_user.id,
actor_name=current_user.username,
target_type="restart_task",
target_id=task_id,
result="failed",
ip=request.client.host if request.client else None,
details={"action": payload.action, "message": task_state["message"]},
)
raise HTTPException(
status_code=status.HTTP_500_INTERNAL_SERVER_ERROR,
detail=task_state["message"],
) from exc
await record_audit_log(
action="system.restart_task.requested",
actor_id=current_user.id,
actor_name=current_user.username,
target_type="restart_task",
target_id=task_id,
result="accepted",
ip=request.client.host if request.client else None,
details={"action": payload.action},
)
return task_state
@@ -165,3 +263,92 @@ async def get_restart_task_logs(
if task is None:
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="Restart task not found")
return {"task_id": task_id, "lines": get_task_logs(task_id)}
@router.get("/logs/sources", response_model=SystemLogSourcesResponse)
async def get_system_log_sources(
current_user: User = Depends(get_current_user),
):
ensure_super_admin(current_user)
return {"items": list_log_sources()}
@router.get("/logs/{source_id}", response_model=SystemLogSnapshotResponse)
async def get_system_log_snapshot(
source_id: str,
limit: int = DEFAULT_LOG_LINE_LIMIT,
level: str = "all",
levels: str | None = Query(None, description="Comma-separated log levels"),
start_date: str | None = Query(None, description="Filter logs from this date (YYYY-MM-DD)"),
end_date: str | None = Query(None, description="Filter logs until this date (YYYY-MM-DD)"),
search: str | None = Query(None, description="Case-insensitive substring search"),
current_user: User = Depends(get_current_user),
):
ensure_super_admin(current_user)
if limit < 1 or limit > MAX_LOG_LINE_LIMIT:
raise HTTPException(
status_code=status.HTTP_400_BAD_REQUEST,
detail=f"limit must be between 1 and {MAX_LOG_LINE_LIMIT}",
)
if str(level).strip().lower() not in SUPPORTED_LOG_LEVELS and normalize_log_level(level) == "all" and str(level).strip().lower() not in {"", "all"}:
raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail="Unsupported log level")
if levels:
for raw_level in str(levels).split(","):
normalized_level = str(raw_level).strip().lower()
if not normalized_level:
continue
if normalized_level not in SUPPORTED_LOG_LEVELS and normalize_log_level(normalized_level) == "all":
raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail="Unsupported log level")
normalized_start_date = validate_log_date(start_date, "start_date")
normalized_end_date = validate_log_date(end_date, "end_date")
if normalized_start_date and normalized_end_date and normalized_start_date > normalized_end_date:
raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail="start_date must be earlier than or equal to end_date")
snapshot = read_log_snapshot(
source_id,
limit,
level=level,
levels=levels,
start_date=normalized_start_date,
end_date=normalized_end_date,
search=search,
)
if snapshot is None:
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="Log source not found")
return snapshot
@router.post("/logs/earth-client", response_model=EarthClientLogEventResponse)
async def ingest_earth_client_log(
payload: EarthClientLogEventCreate,
request: Request,
):
normalized_level = normalize_log_level(payload.level)
append_buffer_log(
"earth-client",
level=normalized_level,
message=payload.message,
context={
"category": payload.category or "",
"url": payload.url or "",
"module": payload.module or "",
"detail": payload.detail or "",
},
)
await record_system_log(
source="earth-client",
service="earth",
module=payload.module or "earth-client",
event="earth.client.runtime_log",
level=normalized_level,
message=payload.message,
category=payload.category or "client-runtime",
context={
"url": payload.url or "",
"detail": payload.detail or "",
"module": payload.module or "",
"client_ip": request.client.host if request.client else "",
},
)
return {"accepted": True, "source_id": "earth-client", "level": normalized_level}

View File

@@ -4,25 +4,34 @@ Unified API for all visualization data sources.
Returns GeoJSON format compatible with Three.js, CesiumJS, and Unreal Cesium.
"""
from datetime import UTC, datetime
from datetime import UTC, datetime, timedelta
import math
from fastapi import APIRouter, HTTPException, Depends, Query
import httpx
from fastapi import APIRouter, HTTPException, Depends, Query, Response
from sqlalchemy.ext.asyncio import AsyncSession
from sqlalchemy import select, func
from typing import List, Dict, Any, Optional
from app.core.collected_data_fields import get_record_field
from app.core.countries import get_country_centroid
from app.core.satellite_tle import build_tle_lines_from_elements
from app.core.time import to_iso8601_utc
from app.db.session import get_db
from app.models.bgp_anomaly import BGPAnomaly
from app.models.bgp_incident import BGPIncident
from app.models.collected_data import CollectedData
from app.models.vessel import VesselPosition, VesselStatic
from app.services.bgp_collectors import build_bgp_collector_coverage
from app.services.cable_graph import build_graph_from_data, CableGraph, haversine_distance
from app.services.collectors.bgp_common import RIPE_RIS_COLLECTOR_COORDS
from app.services.persistent_logs import record_system_log
from app.core.logging import get_logger
router = APIRouter()
logger = get_logger(__name__, service="api")
TERRAIN_TILE_URL_TEMPLATE = (
"https://s3.amazonaws.com/elevation-tiles-prod/terrarium/{z}/{x}/{y}.png"
)
# ============== Converter Functions ==============
@@ -176,6 +185,12 @@ def convert_satellite_to_geojson(records: List[CollectedData]) -> Dict[str, Any]
mean_motion=metadata.get("mean_motion"),
)
constellation_group = _normalize_satellite_constellation_group(
metadata.get("constellation_group"),
record.name,
)
footprint_policy = _get_satellite_footprint_policy(constellation_group)
features.append(
{
"type": "Feature",
@@ -185,6 +200,8 @@ def convert_satellite_to_geojson(records: List[CollectedData]) -> Dict[str, Any]
"id": record.id,
"norad_cat_id": norad_id,
"name": record.name,
"constellation_group": constellation_group,
"footprint_policy": footprint_policy,
"international_designator": metadata.get("international_designator"),
"epoch": metadata.get("epoch"),
"inclination": metadata.get("inclination"),
@@ -205,6 +222,31 @@ def convert_satellite_to_geojson(records: List[CollectedData]) -> Dict[str, Any]
return {"type": "FeatureCollection", "features": features}
def _normalize_satellite_constellation_group(
raw_group: Any,
name: Optional[str],
) -> Optional[str]:
normalized_group = str(raw_group or "").strip().lower()
if normalized_group:
return normalized_group
normalized_name = str(name or "").strip().upper()
if normalized_name.startswith("STARLINK"):
return "starlink"
if normalized_name.startswith("IRIDIUM"):
return "iridium-next"
return None
def _get_satellite_footprint_policy(constellation_group: Optional[str]) -> str:
if constellation_group == "starlink":
return "starlink_ground_footprint"
if constellation_group == "iridium-next":
return "iridium_coverage_ring"
return "none"
def _current_collected_data_stmt(source: str):
return (
select(CollectedData)
@@ -359,6 +401,317 @@ def convert_gpu_cluster_to_geojson(records: List[CollectedData]) -> Dict[str, An
return {"type": "FeatureCollection", "features": features}
def _parse_float(value: Any) -> Optional[float]:
try:
if value in (None, ""):
return None
return float(value)
except (TypeError, ValueError):
return None
COMPUTE_CENTER_COORDINATE_HINTS = (
("el capitan", 37.6819, -121.7681),
("livermore", 37.6819, -121.7681),
("llnl", 37.6819, -121.7681),
("lawrence livermore", 37.6819, -121.7681),
("frontier", 35.9319, -84.3107),
("oak ridge", 35.9319, -84.3107),
("ornl", 35.9319, -84.3107),
("aurora", 41.7130, -87.9820),
("argonne", 41.7130, -87.9820),
("anl", 41.7130, -87.9820),
("fugaku", 34.6953, 135.1974),
("kobe", 34.6953, 135.1974),
("riken", 34.6953, 135.1974),
("summit", 35.9319, -84.3107),
("leonardo", 44.4949, 11.3426),
("bologna", 44.4949, 11.3426),
("alps", 46.0037, 8.9511),
("lugano", 46.0037, 8.9511),
("sunway taihulight", 31.4912, 120.3119),
("wuxi", 31.4912, 120.3119),
("tianhe-2", 23.1291, 113.2644),
("tianhe-2a", 23.1291, 113.2644),
("guangzhou", 23.1291, 113.2644),
("colossus", 35.1495, -90.0490),
("memphis", 35.1495, -90.0490),
("xai", 35.1495, -90.0490),
)
def _normalize_hint_text(*parts: Any) -> str:
return " ".join(
str(part).strip().lower()
for part in parts
if part not in (None, "")
)
def _resolve_compute_center_coordinates(
record: CollectedData,
metadata: Dict[str, Any],
) -> Dict[str, Any]:
latitude = _parse_float(get_record_field(record, "latitude"))
longitude = _parse_float(get_record_field(record, "longitude"))
if latitude not in (None, 0.0) and longitude not in (None, 0.0):
return {
"latitude": latitude,
"longitude": longitude,
"location_precision": "precise",
"geography_mode": "source_coordinates",
"is_estimated": False,
"estimated_reason": None,
}
hint_text = _normalize_hint_text(
record.name,
get_record_field(record, "city"),
get_record_field(record, "country"),
metadata.get("site"),
metadata.get("organization"),
metadata.get("operator"),
)
for needle, resolved_latitude, resolved_longitude in COMPUTE_CENTER_COORDINATE_HINTS:
if needle in hint_text:
return {
"latitude": resolved_latitude,
"longitude": resolved_longitude,
"location_precision": "estimated_site",
"geography_mode": "site_hint",
"is_estimated": True,
"estimated_reason": f"Matched known site hint: {needle}",
}
centroid = get_country_centroid(get_record_field(record, "country"))
if centroid:
return {
"latitude": centroid.get("latitude"),
"longitude": centroid.get("longitude"),
"location_precision": "estimated_country",
"geography_mode": "country_centroid",
"is_estimated": True,
"estimated_reason": "Estimated from country centroid",
}
return {
"latitude": latitude,
"longitude": longitude,
"location_precision": "unknown",
"geography_mode": "unknown",
"is_estimated": True,
"estimated_reason": "No resolvable location hints",
}
def _normalize_capacity_band(capacity_value: Optional[float], capacity_unit: str) -> str:
if capacity_value is None:
return "unknown"
unit = str(capacity_unit or "").strip().lower()
if unit in {"pflop/s", "pflops", "pflop"}:
normalized_tflops = capacity_value * 1000
elif unit in {"gflop/s", "gflops", "gflop"}:
normalized_tflops = capacity_value
else:
normalized_tflops = capacity_value
if normalized_tflops >= 1_000_000:
return "exascale"
if normalized_tflops >= 100_000:
return "ultra"
if normalized_tflops >= 10_000:
return "large"
if normalized_tflops > 0:
return "regional"
return "unknown"
def convert_compute_centers_to_geojson(records: List[CollectedData]) -> Dict[str, Any]:
"""Convert compute infrastructure records into a unified GeoJSON layer."""
features = []
for record in records:
metadata = record.extra_data or {}
coordinate_info = _resolve_compute_center_coordinates(record, metadata)
latitude = coordinate_info.get("latitude")
longitude = coordinate_info.get("longitude")
site_type = (
"supercomputer"
if record.source == "top500" or record.data_type == "supercomputer"
else "gpu_cluster"
)
if latitude in (None, 0.0) or longitude in (None, 0.0):
continue
if site_type == "supercomputer":
capacity_value = _parse_float(get_record_field(record, "rmax"))
capacity_unit = "GFlops"
else:
capacity_value = _parse_float(get_record_field(record, "value"))
capacity_unit = str(get_record_field(record, "unit") or "TFlop/s")
vendor = (
metadata.get("manufacturer")
or metadata.get("vendor")
or metadata.get("gpu_type")
)
operator = (
metadata.get("organization")
or metadata.get("operator")
or metadata.get("owner")
)
rank = metadata.get("rank")
if rank in (None, "") and site_type == "supercomputer":
rank = get_record_field(record, "rank")
updated_at = to_iso8601_utc(record.reference_date or record.collected_at)
features.append(
{
"type": "Feature",
"id": record.id,
"geometry": {
"type": "Point",
"coordinates": [longitude or 0, latitude or 0],
},
"properties": {
"id": record.id,
"source_id": record.source_id,
"name": record.name,
"site_type": site_type,
"country": get_record_field(record, "country"),
"city": get_record_field(record, "city"),
"latitude": latitude,
"longitude": longitude,
"operator": operator,
"vendor": vendor,
"capacity_value": capacity_value,
"capacity_unit": capacity_unit,
"capacity_band": _normalize_capacity_band(capacity_value, capacity_unit),
"rank": rank,
"gpu_count": metadata.get("gpu_count"),
"gpu_type": metadata.get("gpu_type"),
"cores": get_record_field(record, "cores"),
"power": get_record_field(record, "power"),
"source": record.source,
"updated_at": updated_at,
"status": "observed",
"location_precision": coordinate_info.get("location_precision"),
"geography_mode": coordinate_info.get("geography_mode"),
"is_estimated": coordinate_info.get("is_estimated", False),
"estimated_reason": coordinate_info.get("estimated_reason"),
"data_type": "compute_center",
"metadata": metadata,
},
}
)
return {"type": "FeatureCollection", "features": features}
VESSEL_TYPE_FILTERS = {
"cargo": lambda props: str(props.get("vessel_type_name", "")).lower() == "cargo"
or 70 <= int(props.get("vessel_type") or -1) <= 79,
"tanker": lambda props: str(props.get("vessel_type_name", "")).lower() == "tanker"
or 80 <= int(props.get("vessel_type") or -1) <= 89,
"passenger": lambda props: str(props.get("vessel_type_name", "")).lower() == "passenger"
or 60 <= int(props.get("vessel_type") or -1) <= 69,
"fishing": lambda props: str(props.get("vessel_type_name", "")).lower() == "fishing"
or int(props.get("vessel_type") or -1) == 30,
"military": lambda props: str(props.get("vessel_type_name", "")).lower() == "military"
or int(props.get("vessel_type") or -1) == 35,
"other": lambda props: str(props.get("vessel_type_name", "")).lower()
not in {"cargo", "tanker", "passenger", "fishing", "military"},
}
def convert_vessels_to_geojson(rows: List[Any]) -> Dict[str, Any]:
features = []
for position, static in rows:
if position.lat is None or position.lon is None:
continue
props = {
"mmsi": position.mmsi,
"name": getattr(static, "name", None) or f"MMSI {position.mmsi}",
"callsign": getattr(static, "callsign", None),
"imo": getattr(static, "imo", None),
"vessel_type": getattr(static, "vessel_type", None),
"vessel_type_name": getattr(static, "vessel_type_name", None) or "Other",
"flag": getattr(static, "flag", None),
"length": getattr(static, "length", None),
"width": getattr(static, "width", None),
"draught": getattr(static, "draught", None),
"sog": position.sog,
"cog": position.cog,
"heading": position.heading,
"nav_status": position.nav_status,
"received_at": to_iso8601_utc(position.received_at),
"data_type": "vessel",
}
features.append(
{
"type": "Feature",
"id": position.mmsi,
"geometry": {
"type": "Point",
"coordinates": [position.lon, position.lat],
},
"properties": props,
}
)
return {"type": "FeatureCollection", "features": features}
def _parse_bbox(value: Optional[str]) -> tuple[float, float, float, float] | None:
if not value:
return None
parts = [part.strip() for part in value.split(",")]
if len(parts) != 4:
raise HTTPException(status_code=400, detail="bbox must be lon_min,lat_min,lon_max,lat_max")
try:
lon_min, lat_min, lon_max, lat_max = [float(part) for part in parts]
except ValueError as exc:
raise HTTPException(status_code=400, detail="bbox values must be numbers") from exc
if lat_min > lat_max:
lat_min, lat_max = lat_max, lat_min
if lon_min > lon_max:
lon_min, lon_max = lon_max, lon_min
return lon_min, lat_min, lon_max, lat_max
def _matches_vessel_type(props: dict[str, Any], requested_types: set[str]) -> bool:
if not requested_types:
return True
for requested_type in requested_types:
predicate = VESSEL_TYPE_FILTERS.get(requested_type)
if predicate and predicate(props):
return True
return False
def _build_vessel_stats(features: List[dict[str, Any]]) -> dict[str, Any]:
by_type: dict[str, int] = {}
underway = 0
anchored_or_moored = 0
for feature in features:
props = feature.get("properties", {})
vessel_type = str(props.get("vessel_type_name") or "Other")
by_type[vessel_type] = by_type.get(vessel_type, 0) + 1
nav_status = props.get("nav_status")
if nav_status in (1, 5):
anchored_or_moored += 1
else:
underway += 1
return {
"total": len(features),
"by_type": by_type,
"underway": underway,
"anchored_or_moored": anchored_or_moored,
}
def convert_bgp_anomalies_to_geojson(
records: List[BGPAnomaly],
geography_hints: Optional[Dict[str, Dict[str, Any]]] = None,
@@ -776,15 +1129,41 @@ async def get_cables_geojson(db: AsyncSession = Depends(get_db)):
except HTTPException:
raise
except Exception as e:
logger.exception_event(
"Failed to build cables GeoJSON response",
event="visualization.cables.load_failed",
context={"error": str(e)},
)
await record_system_log(
source="backend",
service="api",
module=__name__,
event="visualization.cables.load_failed",
level="error",
message="Failed to build cables GeoJSON response",
category="visualization",
context={"error": str(e)},
)
raise HTTPException(status_code=500, detail=f"Internal error: {str(e)}")
@router.get("/geo/landing-points")
async def get_landing_points_geojson(db: AsyncSession = Depends(get_db)):
try:
records = await _load_current_collected_data(db, "arcgis_landing_points")
relation_records = await _load_current_collected_data(db, "arcgis_cable_landing_relation")
cable_records = await _load_current_collected_data(db, "arcgis_cables")
records_by_source = await _load_current_collected_data_by_sources(
db,
[
"arcgis_landing_points",
"arcgis_cable_landing_relation",
"arcgis_cables",
],
)
records = records_by_source.get("arcgis_landing_points", [])
relation_records = records_by_source.get(
"arcgis_cable_landing_relation",
[],
)
cable_records = records_by_source.get("arcgis_cables", [])
city_to_cable_ids_map, cable_id_to_name_map = _build_landing_point_cable_maps(
relation_records,
@@ -801,9 +1180,68 @@ async def get_landing_points_geojson(db: AsyncSession = Depends(get_db)):
except HTTPException:
raise
except Exception as e:
logger.exception_event(
"Failed to build landing points GeoJSON response",
event="visualization.landing_points.load_failed",
context={"error": str(e)},
)
await record_system_log(
source="backend",
service="api",
module=__name__,
event="visualization.landing_points.load_failed",
level="error",
message="Failed to build landing points GeoJSON response",
category="visualization",
context={"error": str(e)},
)
raise HTTPException(status_code=500, detail=f"Internal error: {str(e)}")
@router.get("/terrain/terrarium/{z}/{x}/{y}.png")
async def get_terrarium_tile(z: int, x: int, y: int):
"""Proxy Terrarium elevation tiles through the backend to avoid browser CORS issues."""
if z < 0 or x < 0 or y < 0:
raise HTTPException(status_code=400, detail="Invalid terrain tile coordinates")
url = TERRAIN_TILE_URL_TEMPLATE.format(z=z, x=x, y=y)
try:
async with httpx.AsyncClient(
timeout=20.0,
follow_redirects=True,
) as client:
upstream = await client.get(url)
upstream.raise_for_status()
except httpx.HTTPStatusError as exc:
raise HTTPException(
status_code=exc.response.status_code,
detail=f"Terrain tile upstream error: {exc.response.status_code}",
) from exc
except httpx.HTTPError as exc:
raise HTTPException(
status_code=502,
detail=f"Terrain tile fetch failed: {exc}",
) from exc
cache_control = upstream.headers.get("cache-control") or "public, max-age=86400"
etag = upstream.headers.get("etag")
last_modified = upstream.headers.get("last-modified")
headers = {
"Cache-Control": cache_control,
}
if etag:
headers["ETag"] = etag
if last_modified:
headers["Last-Modified"] = last_modified
return Response(
content=upstream.content,
media_type=upstream.headers.get("content-type", "image/png"),
headers=headers,
)
@router.get("/geo/all")
async def get_all_geojson(db: AsyncSession = Depends(get_db)):
records_by_source = await _load_current_collected_data_by_sources(
@@ -916,6 +1354,184 @@ async def get_gpu_clusters_geojson(
}
@router.get("/geo/compute-centers")
async def get_compute_centers_geojson(
limit: int = Query(200, ge=1, le=1000),
db: AsyncSession = Depends(get_db),
):
"""获取统一算力中心 GeoJSON 数据"""
records_by_source = await _load_current_collected_data_by_sources(
db,
["top500", "epoch_ai_gpu"],
)
records = _filter_known_records(
records_by_source.get("top500", []) + records_by_source.get("epoch_ai_gpu", []),
)
if limit is not None:
records = records[:limit]
if not records:
return {
"type": "FeatureCollection",
"features": [],
"count": 0,
"stats": {
"total": 0,
"supercomputers": 0,
"gpu_clusters": 0,
},
}
geojson = convert_compute_centers_to_geojson(records)
features = geojson.get("features", [])
return {
**geojson,
"count": len(features),
"stats": {
"total": len(features),
"supercomputers": sum(
1 for feature in features
if feature.get("properties", {}).get("site_type") == "supercomputer"
),
"gpu_clusters": sum(
1 for feature in features
if feature.get("properties", {}).get("site_type") == "gpu_cluster"
),
},
}
@router.get("/geo/vessels")
async def get_vessels_geojson(
bbox: Optional[str] = Query(
None,
description="Viewport bbox as lon_min,lat_min,lon_max,lat_max",
),
type: Optional[str] = Query(
None,
description="Comma-separated vessel types: cargo,tanker,passenger,fishing,military,other",
),
limit: int = Query(5000, ge=1, le=50000),
db: AsyncSession = Depends(get_db),
):
"""Return latest vessel positions as GeoJSON points."""
latest_times = (
select(
VesselPosition.mmsi.label("mmsi"),
func.max(VesselPosition.received_at).label("received_at"),
)
.group_by(VesselPosition.mmsi)
.subquery()
)
stmt = (
select(VesselPosition, VesselStatic)
.join(
latest_times,
(VesselPosition.mmsi == latest_times.c.mmsi)
& (VesselPosition.received_at == latest_times.c.received_at),
)
.outerjoin(VesselStatic, VesselStatic.mmsi == VesselPosition.mmsi)
.order_by(VesselPosition.received_at.desc())
.limit(limit)
)
parsed_bbox = _parse_bbox(bbox)
if parsed_bbox is not None:
lon_min, lat_min, lon_max, lat_max = parsed_bbox
stmt = stmt.where(
VesselPosition.lon >= lon_min,
VesselPosition.lon <= lon_max,
VesselPosition.lat >= lat_min,
VesselPosition.lat <= lat_max,
)
result = await db.execute(stmt)
rows = list(result.all())
geojson = convert_vessels_to_geojson(rows)
requested_types = {
item.strip().lower()
for item in (type or "").split(",")
if item.strip()
}
if requested_types:
geojson["features"] = [
feature
for feature in geojson.get("features", [])
if _matches_vessel_type(feature.get("properties", {}), requested_types)
]
features = geojson.get("features", [])
return {
**geojson,
"count": len(features),
"stats": _build_vessel_stats(features),
}
@router.get("/vessels/{mmsi}")
async def get_vessel_detail(mmsi: int, db: AsyncSession = Depends(get_db)):
latest_position_stmt = (
select(VesselPosition)
.where(VesselPosition.mmsi == mmsi)
.order_by(VesselPosition.received_at.desc())
.limit(1)
)
static = await db.get(VesselStatic, mmsi)
result = await db.execute(latest_position_stmt)
position = result.scalar_one_or_none()
if position is None:
raise HTTPException(status_code=404, detail="Vessel not found")
geojson = convert_vessels_to_geojson([(position, static)])
return {
**(geojson["features"][0]["properties"]),
"latitude": position.lat,
"longitude": position.lon,
}
@router.get("/vessels/{mmsi}/track")
async def get_vessel_track(
mmsi: int,
hours: int = Query(6, ge=1, le=24),
db: AsyncSession = Depends(get_db),
):
cutoff = datetime.now(UTC) - timedelta(hours=hours)
result = await db.execute(
select(VesselPosition)
.where(VesselPosition.mmsi == mmsi)
.where(VesselPosition.received_at >= cutoff)
.order_by(VesselPosition.received_at.asc())
)
positions = list(result.scalars().all())
if not positions:
return {
"type": "FeatureCollection",
"features": [],
"count": 0,
}
return {
"type": "FeatureCollection",
"features": [
{
"type": "Feature",
"geometry": {
"type": "LineString",
"coordinates": [[position.lon, position.lat] for position in positions],
},
"properties": {
"mmsi": mmsi,
"hours": hours,
"point_count": len(positions),
"start_at": to_iso8601_utc(positions[0].received_at),
"end_at": to_iso8601_utc(positions[-1].received_at),
},
}
],
"count": 1,
}
@router.get("/geo/bgp-anomalies")
async def get_bgp_anomalies_geojson(
severity: Optional[str] = Query(None),
@@ -971,6 +1587,76 @@ async def get_bgp_collectors_geojson(db: AsyncSession = Depends(get_db)):
return {**geojson, "count": len(geojson.get("features", []))}
@router.get("/geo/summary")
async def get_visualization_geo_summary(db: AsyncSession = Depends(get_db)):
"""Return lightweight Earth HUD counts without loading layer GeoJSON payloads."""
records_by_source = await _load_current_collected_data_by_sources(
db,
[
"arcgis_cables",
"arcgis_landing_points",
"celestrak_tle",
"top500",
"epoch_ai_gpu",
],
)
cables = convert_cable_to_geojson(records_by_source.get("arcgis_cables", []))
landing_points = convert_landing_point_to_geojson(
records_by_source.get("arcgis_landing_points", []),
)
satellites = convert_satellite_to_geojson(
_filter_known_records(records_by_source.get("celestrak_tle", [])),
)
compute_centers = convert_compute_centers_to_geojson(
_filter_known_records(
records_by_source.get("top500", [])
+ records_by_source.get("epoch_ai_gpu", []),
),
)
compute_features = compute_centers.get("features", [])
active_incident_result = await db.execute(
select(func.count(BGPIncident.id)).where(BGPIncident.status == "active"),
)
active_anomaly_result = await db.execute(
select(func.count(BGPAnomaly.id)).where(BGPAnomaly.status == "active"),
)
active_incident_count = int(active_incident_result.scalar() or 0)
active_anomaly_count = int(active_anomaly_result.scalar() or 0)
bgp_collectors = await build_bgp_collector_coverage(
db,
source_filter=("ris_live_bgp", "bgpstream_bgp"),
)
vessel_count_result = await db.execute(
select(func.count(func.distinct(VesselPosition.mmsi))),
)
vessel_count = int(vessel_count_result.scalar() or 0)
return {
"generated_at": to_iso8601_utc(datetime.now(UTC)),
"stats": {
"cable_count": len(cables.get("features", [])),
"landing_point_count": len(landing_points.get("features", [])),
"satellite_count": len(satellites.get("features", [])),
"compute_center_count": len(compute_features),
"vessel_count": vessel_count,
"supercomputer_count": sum(
1 for feature in compute_features
if feature.get("properties", {}).get("site_type") == "supercomputer"
),
"gpu_cluster_count": sum(
1 for feature in compute_features
if feature.get("properties", {}).get("site_type") == "gpu_cluster"
),
"bgp_event_count": active_incident_count or active_anomaly_count,
"bgp_incident_count": active_incident_count,
"bgp_anomaly_count": active_anomaly_count,
"bgp_collector_count": len([item for item in bgp_collectors if item.get("collector")]),
},
}
@router.get("/all")
async def get_all_visualization_data(db: AsyncSession = Depends(get_db)):
"""获取所有可视化数据的统一端点

View File

@@ -2,7 +2,6 @@
import asyncio
import json
import logging
from datetime import UTC, datetime
from typing import Optional
@@ -10,10 +9,11 @@ from fastapi import APIRouter, WebSocket, WebSocketDisconnect, Query
from jose import jwt, JWTError
from app.core.config import settings
from app.core.logging import get_logger
from app.core.time import to_iso8601_utc
from app.core.websocket.manager import manager
logger = logging.getLogger(__name__)
logger = get_logger(__name__, service="api")
router = APIRouter()
@@ -22,11 +22,18 @@ async def authenticate_token(token: str) -> Optional[dict]:
try:
payload = jwt.decode(token, settings.SECRET_KEY, algorithms=[settings.ALGORITHM])
if payload.get("type") != "access":
logger.warning(f"WebSocket auth failed: wrong token type")
logger.warning_event(
"WebSocket auth failed: wrong token type",
event="auth.websocket.invalid_token_type",
)
return None
return payload
except JWTError as e:
logger.warning(f"WebSocket auth failed: {e}")
logger.warning_event(
"WebSocket auth failed",
event="auth.websocket.decode_failed",
context={"error": str(e)},
)
return None
@@ -36,10 +43,17 @@ async def websocket_endpoint(
token: str = Query(...),
):
"""WebSocket endpoint for real-time data"""
logger.info(f"WebSocket connection attempt with token: {token[:20]}...")
logger.info_event(
"WebSocket connection attempt",
event="auth.websocket.connection_attempt",
context={"token_preview": f"{token[:8]}..."},
)
payload = await authenticate_token(token)
if payload is None:
logger.warning("WebSocket authentication failed, closing connection")
logger.warning_event(
"WebSocket authentication failed, closing connection",
event="auth.websocket.connection_rejected",
)
await websocket.close(code=4001)
return

View File

@@ -1,15 +1,15 @@
"""Redis caching service"""
import json
import logging
from datetime import timedelta
from typing import Optional, Any
import redis
from app.core.config import settings
from app.core.logging import get_logger
logger = logging.getLogger(__name__)
logger = get_logger(__name__)
# Lazy Redis client initialization
@@ -47,7 +47,7 @@ class CacheService:
return json.loads(value)
return None
except Exception as e:
logger.warning(f"Cache get error: {e}")
logger.warning_event("Cache get error", event="cache.get.failed", context={"error": str(e)})
return None
def set(
@@ -61,7 +61,7 @@ class CacheService:
serialized = json.dumps(value, default=str)
return self.client.setex(key, expire_seconds, serialized)
except Exception as e:
logger.warning(f"Cache set error: {e}")
logger.warning_event("Cache set error", event="cache.set.failed", context={"error": str(e)})
return False
def delete(self, key: str) -> bool:
@@ -69,7 +69,7 @@ class CacheService:
try:
return self.client.delete(key) > 0
except Exception as e:
logger.warning(f"Cache delete error: {e}")
logger.warning_event("Cache delete error", event="cache.delete.failed", context={"error": str(e)})
return False
def delete_pattern(self, pattern: str) -> int:
@@ -80,7 +80,7 @@ class CacheService:
return self.client.delete(*keys)
return 0
except Exception as e:
logger.warning(f"Cache delete_pattern error: {e}")
logger.warning_event("Cache delete_pattern error", event="cache.delete_pattern.failed", context={"error": str(e)})
return 0
def get_or_set(

View File

@@ -30,6 +30,8 @@ COLLECTOR_URL_KEYS = {
"iptoasn_prefix_geo": "iptoasn.combined_url",
"opengeofeed_prefix_geo": "opengeofeed.public_csv_url",
"nro_delegated_prefix_geo": "nro.delegated_stats_url",
"news_live_streams": "news_live_streams.channels_url",
"barentswatch_vessels": "barentswatch_vessels.url",
}

View File

@@ -86,3 +86,15 @@ opengeofeed:
nro:
# NRO delegated stats 下载地址
delegated_stats_url: "https://ftp.ripe.net/pub/stats/ripencc/nro-stats/latest/nro-delegated-stats"
news_live_streams:
# IPTV-org 频道元数据 JSON
channels_url: "https://iptv-org.github.io/api/channels.json"
# IPTV-org 频道播放流 JSON
streams_url: "https://iptv-org.github.io/api/streams.json"
# IPTV-org 台标 JSON
logos_url: "https://iptv-org.github.io/api/logos.json"
barentswatch_vessels:
# BarentsWatch Live AIS latest combined endpoint. Requires an AIS bearer token.
url: "https://live.ais.barentswatch.no/v1/latest/combined"

View File

@@ -4,163 +4,246 @@ DEFAULT_DATASOURCES = {
"top500": {
"id": 1,
"name": "TOP500 Supercomputers",
"display_name": "TOP500 超算榜单",
"module": "L1",
"priority": "P0",
"frequency_minutes": 240,
"is_free": True,
"requires_credentials": False,
},
"epoch_ai_gpu": {
"id": 2,
"name": "Epoch AI GPU Clusters",
"display_name": "Epoch AI GPU 集群",
"module": "L1",
"priority": "P0",
"frequency_minutes": 360,
"is_free": True,
"requires_credentials": False,
},
"huggingface_models": {
"id": 3,
"name": "HuggingFace Models",
"display_name": "Hugging Face 模型",
"module": "L2",
"priority": "P1",
"frequency_minutes": 720,
"is_free": True,
"requires_credentials": False,
},
"huggingface_datasets": {
"id": 4,
"name": "HuggingFace Datasets",
"display_name": "Hugging Face 数据集",
"module": "L2",
"priority": "P1",
"frequency_minutes": 720,
"is_free": True,
"requires_credentials": False,
},
"huggingface_spaces": {
"id": 5,
"name": "HuggingFace Spaces",
"display_name": "Hugging Face Spaces",
"module": "L2",
"priority": "P2",
"frequency_minutes": 1440,
"is_free": True,
"requires_credentials": False,
},
"peeringdb_ixp": {
"id": 6,
"name": "PeeringDB IXP",
"display_name": "PeeringDB 交换中心",
"module": "L2",
"priority": "P1",
"frequency_minutes": 1440,
"is_free": True,
"requires_credentials": False,
},
"peeringdb_network": {
"id": 7,
"name": "PeeringDB Networks",
"display_name": "PeeringDB 网络",
"module": "L2",
"priority": "P2",
"frequency_minutes": 2880,
"is_free": True,
"requires_credentials": False,
},
"peeringdb_facility": {
"id": 8,
"name": "PeeringDB Facilities",
"display_name": "PeeringDB 设施",
"module": "L2",
"priority": "P2",
"frequency_minutes": 2880,
"is_free": True,
"requires_credentials": False,
},
"telegeography_cables": {
"id": 9,
"name": "Submarine Cables",
"display_name": "海底光缆",
"module": "L2",
"priority": "P1",
"frequency_minutes": 10080,
"is_free": True,
"requires_credentials": False,
},
"telegeography_landing": {
"id": 10,
"name": "Cable Landing Points",
"display_name": "光缆登陆点",
"module": "L2",
"priority": "P2",
"frequency_minutes": 10080,
"is_free": True,
"requires_credentials": False,
},
"telegeography_systems": {
"id": 11,
"name": "Cable Systems",
"display_name": "光缆系统",
"module": "L2",
"priority": "P2",
"frequency_minutes": 10080,
"is_free": True,
"requires_credentials": False,
},
"arcgis_cables": {
"id": 15,
"name": "ArcGIS Submarine Cables",
"display_name": "ArcGIS 海底光缆",
"module": "L2",
"priority": "P1",
"frequency_minutes": 10080,
"is_free": True,
"requires_credentials": False,
},
"arcgis_landing_points": {
"id": 16,
"name": "ArcGIS Landing Points",
"display_name": "ArcGIS 登陆点",
"module": "L2",
"priority": "P1",
"frequency_minutes": 10080,
"is_free": True,
"requires_credentials": False,
},
"arcgis_cable_landing_relation": {
"id": 17,
"name": "ArcGIS Cable-Landing Relations",
"display_name": "ArcGIS 光缆登陆关系",
"module": "L2",
"priority": "P1",
"frequency_minutes": 10080,
"is_free": True,
"requires_credentials": False,
},
"fao_landing_points": {
"id": 18,
"name": "FAO Landing Points",
"display_name": "FAO 登陆点",
"module": "L2",
"priority": "P1",
"frequency_minutes": 10080,
"is_free": True,
"requires_credentials": False,
},
"spacetrack_tle": {
"id": 19,
"name": "Space-Track TLE",
"display_name": "Space-Track 轨道根数",
"module": "L3",
"priority": "P2",
"frequency_minutes": 1440,
"is_free": True,
"requires_credentials": True,
"credential_provider": "spacetrack",
"credential_status": "planned",
},
"celestrak_tle": {
"id": 20,
"name": "CelesTrak TLE",
"display_name": "CelesTrak 轨道根数",
"module": "L3",
"priority": "P2",
"frequency_minutes": 1440,
"is_free": True,
"requires_credentials": False,
},
"ris_live_bgp": {
"id": 21,
"name": "RIPE RIS Live BGP",
"display_name": "RIPE RIS 实时 BGP",
"module": "L3",
"priority": "P1",
"frequency_minutes": 15,
"is_free": True,
"requires_credentials": False,
},
"bgpstream_bgp": {
"id": 22,
"name": "CAIDA BGPStream Backfill",
"display_name": "CAIDA BGPStream 回填",
"module": "L3",
"priority": "P1",
"frequency_minutes": 360,
"is_free": True,
"requires_credentials": False,
},
"iptoasn_prefix_geo": {
"id": 23,
"name": "IPtoASN Prefix Geography",
"display_name": "IPtoASN 前缀地理",
"module": "L3",
"priority": "P1",
"frequency_minutes": 1440,
"is_free": True,
"requires_credentials": False,
},
"opengeofeed_prefix_geo": {
"id": 24,
"name": "OpenGeoFeed Prefix Geography",
"display_name": "OpenGeoFeed 前缀地理",
"module": "L3",
"priority": "P1",
"frequency_minutes": 1440,
"is_free": True,
"requires_credentials": False,
},
"nro_delegated_prefix_geo": {
"id": 25,
"name": "NRO Delegated Prefix Geography",
"display_name": "NRO 分配前缀地理",
"module": "L3",
"priority": "P1",
"frequency_minutes": 1440,
"is_free": True,
"requires_credentials": False,
},
"news_live_streams": {
"id": 26,
"name": "News Live Streams",
"display_name": "新闻直播源",
"module": "L4",
"priority": "P2",
"frequency_minutes": 720,
"is_free": True,
"requires_credentials": False,
},
"barentswatch_vessels": {
"id": 27,
"name": "BarentsWatch AIS Vessels",
"display_name": "BarentsWatch AIS 船舶",
"module": "L4",
"priority": "P1",
"frequency_minutes": 1,
"is_free": True,
"requires_credentials": True,
"credential_provider": "barentswatch",
"credential_status": "supported",
},
}

161
backend/app/core/logging.py Normal file
View File

@@ -0,0 +1,161 @@
from __future__ import annotations
import json
import logging
import os
import re
from collections.abc import Mapping, Sequence
from typing import Any
from app.core.request_context import get_request_id
DEFAULT_SERVICE = "backend"
DEFAULT_EVENT = "app.log"
DEFAULT_LOG_LEVEL = os.getenv("PLANET_LOG_LEVEL", "INFO").upper()
REDACTED = "[REDACTED]"
SENSITIVE_FIELD_NAMES = {
"access_token",
"api_key",
"authorization",
"cookie",
"password",
"refresh_token",
"secret",
"token",
}
SENSITIVE_TEXT_PATTERNS = (
re.compile(r"(?i)(authorization\s*[:=]\s*)(.+)"),
re.compile(r"(?i)(bearer\s+)([A-Za-z0-9._\-]+)"),
re.compile(r"(?i)(token\s*[:=]\s*)(.+)"),
re.compile(r"(?i)(password\s*[:=]\s*)(.+)"),
re.compile(r"(?i)(cookie\s*[:=]\s*)(.+)"),
)
def sanitize_log_value(value: Any) -> Any:
if isinstance(value, Mapping):
return {
str(key): (REDACTED if str(key).lower() in SENSITIVE_FIELD_NAMES else sanitize_log_value(item))
for key, item in value.items()
}
if isinstance(value, Sequence) and not isinstance(value, (str, bytes, bytearray)):
return [sanitize_log_value(item) for item in value]
if isinstance(value, str):
sanitized = value
for pattern in SENSITIVE_TEXT_PATTERNS:
sanitized = pattern.sub(lambda match: f"{match.group(1)}{REDACTED}", sanitized)
return sanitized
return value
def _normalize_context(context: Any) -> dict[str, Any]:
if context is None:
return {}
if isinstance(context, Mapping):
sanitized = sanitize_log_value(context)
return {str(key): value for key, value in sanitized.items()}
return {"value": sanitize_log_value(context)}
class PlanetContextFilter(logging.Filter):
def filter(self, record: logging.LogRecord) -> bool:
record.request_id = getattr(record, "request_id", None) or get_request_id() or "-"
record.service = getattr(record, "service", None) or DEFAULT_SERVICE
record.event = getattr(record, "event", None) or DEFAULT_EVENT
record.context = _normalize_context(getattr(record, "context", None))
record.message = sanitize_log_value(record.getMessage())
return True
class PlanetFormatter(logging.Formatter):
def format(self, record: logging.LogRecord) -> str:
timestamp = self.formatTime(record, self.datefmt)
level = record.levelname
service = getattr(record, "service", DEFAULT_SERVICE)
module_name = record.name
event = getattr(record, "event", DEFAULT_EVENT)
request_id = getattr(record, "request_id", "-")
message = sanitize_log_value(record.getMessage())
context = _normalize_context(getattr(record, "context", None))
context_suffix = ""
if context:
context_suffix = f" context={json.dumps(context, ensure_ascii=False, sort_keys=True)}"
rendered = (
f"{timestamp} {level} service={service} module={module_name} "
f"event={event} request_id={request_id} message={message}{context_suffix}"
)
if record.exc_info:
rendered = f"{rendered}\n{self.formatException(record.exc_info)}"
return rendered
class PlanetLoggerAdapter(logging.LoggerAdapter):
def process(self, msg: Any, kwargs: dict[str, Any]) -> tuple[Any, dict[str, Any]]:
extra = dict(self.extra)
extra.update(kwargs.get("extra", {}))
if "context" in extra:
extra["context"] = _normalize_context(extra.get("context"))
kwargs["extra"] = extra
return sanitize_log_value(msg), kwargs
def log_event(
self,
level: int,
message: str,
*,
event: str,
context: Mapping[str, Any] | None = None,
**extra: Any,
) -> None:
self.log(level, message, extra={"event": event, "context": context or {}, **extra})
def debug_event(self, message: str, *, event: str, context: Mapping[str, Any] | None = None, **extra: Any) -> None:
self.log_event(logging.DEBUG, message, event=event, context=context, **extra)
def info_event(self, message: str, *, event: str, context: Mapping[str, Any] | None = None, **extra: Any) -> None:
self.log_event(logging.INFO, message, event=event, context=context, **extra)
def warning_event(self, message: str, *, event: str, context: Mapping[str, Any] | None = None, **extra: Any) -> None:
self.log_event(logging.WARNING, message, event=event, context=context, **extra)
def error_event(self, message: str, *, event: str, context: Mapping[str, Any] | None = None, **extra: Any) -> None:
self.log_event(logging.ERROR, message, event=event, context=context, **extra)
def exception_event(
self,
message: str,
*,
event: str,
context: Mapping[str, Any] | None = None,
**extra: Any,
) -> None:
self.error(message, exc_info=True, extra={"event": event, "context": context or {}, **extra})
def get_logger(name: str, *, service: str = DEFAULT_SERVICE) -> PlanetLoggerAdapter:
return PlanetLoggerAdapter(logging.getLogger(name), {"service": service})
def configure_logging(level: str | None = None) -> None:
root_logger = logging.getLogger()
if getattr(configure_logging, "_configured", False):
if level:
root_logger.setLevel(level.upper())
return
handler = logging.StreamHandler()
handler.setFormatter(PlanetFormatter(datefmt="%Y-%m-%d %H:%M:%S"))
handler.addFilter(PlanetContextFilter())
root_logger.handlers.clear()
root_logger.addHandler(handler)
root_logger.setLevel((level or DEFAULT_LOG_LEVEL).upper())
for logger_name in ("uvicorn", "uvicorn.error", "uvicorn.access"):
target_logger = logging.getLogger(logger_name)
target_logger.handlers.clear()
target_logger.propagate = True
logging.captureWarnings(True)
configure_logging._configured = True

View File

@@ -0,0 +1,14 @@
from __future__ import annotations
from contextvars import ContextVar
request_id_context: ContextVar[str | None] = ContextVar("request_id", default=None)
def set_request_id(request_id: str | None) -> None:
request_id_context.set(request_id)
def get_request_id() -> str | None:
return request_id_context.get()

View File

@@ -0,0 +1,151 @@
"""Registry of target schemas supported by mapped custom data sources."""
from __future__ import annotations
from dataclasses import dataclass
from datetime import datetime
from typing import Any
from pydantic import BaseModel, Field, ValidationError, field_validator
class VesselAISRecord(BaseModel):
mmsi: int = Field(ge=100000000, le=999999999)
lat: float = Field(ge=-90, le=90)
lon: float = Field(ge=-180, le=180)
sog: float | None = None
cog: float | None = Field(default=None, ge=0, le=360)
heading: int | None = Field(default=None, ge=0, le=511)
name: str | None = None
vessel_type: str | int | None = None
received_at: datetime | None = None
class GeoPointRecord(BaseModel):
lat: float = Field(ge=-90, le=90)
lon: float = Field(ge=-180, le=180)
name: str | None = None
type: str | None = None
source_id: str | None = None
observed_at: datetime | None = None
metadata: dict[str, Any] = Field(default_factory=dict)
class GenericRecord(BaseModel):
data: dict[str, Any] = Field(default_factory=dict)
source_id: str | None = None
observed_at: datetime | None = None
@field_validator("data")
@classmethod
def require_payload(cls, value: dict[str, Any]) -> dict[str, Any]:
if not value:
raise ValueError("generic_records requires a non-empty data object")
return value
@dataclass(frozen=True)
class TargetField:
name: str
type: str
required: bool = False
description: str = ""
example: Any = None
def to_dict(self) -> dict[str, Any]:
return {
"name": self.name,
"type": self.type,
"required": self.required,
"description": self.description,
"example": self.example,
}
@dataclass(frozen=True)
class TargetSchema:
key: str
label: str
description: str
fields: tuple[TargetField, ...]
model: type[BaseModel]
destination: str
def to_dict(self) -> dict[str, Any]:
return {
"key": self.key,
"label": self.label,
"description": self.description,
"destination": self.destination,
"fields": [field.to_dict() for field in self.fields],
}
def validate_record(self, record: dict[str, Any]) -> tuple[dict[str, Any] | None, list[str]]:
try:
return self.model.model_validate(record).model_dump(mode="json"), []
except ValidationError as exc:
return None, [
".".join(str(part) for part in error["loc"]) + f": {error['msg']}"
for error in exc.errors()
]
TARGET_SCHEMAS: dict[str, TargetSchema] = {
"vessel_ais": TargetSchema(
key="vessel_ais",
label="船舶 AIS",
description="船只位置、航速、航向、MMSI 等 AIS 数据。",
destination="vessel_position",
model=VesselAISRecord,
fields=(
TargetField("mmsi", "integer", True, "MMSI 九位船舶标识", 257123000),
TargetField("lat", "float", True, "纬度", 59.91),
TargetField("lon", "float", True, "经度", 10.75),
TargetField("sog", "float", False, "对地航速,单位节", 12.4),
TargetField("cog", "float", False, "对地航向0-360 度", 184.5),
TargetField("heading", "integer", False, "船首向0-511", 186),
TargetField("name", "string", False, "船名", "OSLO EXPRESS"),
TargetField("vessel_type", "string", False, "船型", "cargo"),
TargetField("received_at", "datetime", False, "数据接收时间", "2026-04-28T00:00:00Z"),
),
),
"geo_points": TargetSchema(
key="geo_points",
label="通用地理点",
description="带经纬度的通用实体或事件点位。",
destination="generic_geo_points",
model=GeoPointRecord,
fields=(
TargetField("lat", "float", True, "纬度", 1.3),
TargetField("lon", "float", True, "经度", 103.8),
TargetField("name", "string", False, "点位名称", "Singapore"),
TargetField("type", "string", False, "点位类型", "datacenter"),
TargetField("source_id", "string", False, "来源侧 ID", "sg-1"),
TargetField("observed_at", "datetime", False, "观测时间", "2026-04-28T00:00:00Z"),
TargetField("metadata", "object", False, "扩展字段", {"provider": "example"}),
),
),
"generic_records": TargetSchema(
key="generic_records",
label="通用结构化记录",
description="未知结构数据沉淀,不直接进入 Earth 图层。",
destination="collected_data",
model=GenericRecord,
fields=(
TargetField("data", "object", True, "结构化记录主体", {"raw": "value"}),
TargetField("source_id", "string", False, "来源侧 ID", "record-1"),
TargetField("observed_at", "datetime", False, "观测时间", "2026-04-28T00:00:00Z"),
),
),
}
def list_target_schemas() -> list[dict[str, Any]]:
return [schema.to_dict() for schema in TARGET_SCHEMAS.values()]
def get_target_schema(key: str) -> TargetSchema:
try:
return TARGET_SCHEMAS[key]
except KeyError as exc:
raise ValueError(f"Unsupported target schema: {key}") from exc

View File

@@ -5,10 +5,22 @@ from sqlalchemy.ext.asyncio import AsyncSession, create_async_engine, async_sess
from sqlalchemy.orm import declarative_base
from app.core.config import settings
from app.core.logging import get_logger
logger = get_logger(__name__)
DB_POOL_CONFIG = {
"pool_pre_ping": True,
"pool_recycle": 1800,
"pool_size": 10,
"max_overflow": 20,
"pool_timeout": 30,
}
engine = create_async_engine(
settings.DATABASE_URL,
echo=settings.DEBUG if hasattr(settings, "DEBUG") else False,
**DB_POOL_CONFIG,
)
async_session_factory = async_sessionmaker(engine, class_=AsyncSession, expire_on_commit=False)
@@ -97,6 +109,21 @@ async def init_db():
import app.models.system_setting # noqa: F401
import app.models.playground_session # noqa: F401
import app.models.playground_message # noqa: F401
import app.models.system_log # noqa: F401
import app.models.vessel # noqa: F401
import app.models.datasource_mapping # noqa: F401
logger.warning_event(
"Database pool settings active",
event="database.pool.initialized",
context={
"pool_pre_ping": DB_POOL_CONFIG["pool_pre_ping"],
"pool_recycle": DB_POOL_CONFIG["pool_recycle"],
"pool_size": DB_POOL_CONFIG["pool_size"],
"max_overflow": DB_POOL_CONFIG["max_overflow"],
"pool_timeout": DB_POOL_CONFIG["pool_timeout"],
},
)
async with engine.begin() as conn:
await conn.run_sync(Base.metadata.create_all)

View File

@@ -1,4 +1,5 @@
from contextlib import asynccontextmanager
from uuid import uuid4
from fastapi import FastAPI
from fastapi.middleware.cors import CORSMiddleware
@@ -7,6 +8,8 @@ from starlette.middleware.base import BaseHTTPMiddleware
from app.api.main import api_router
from app.api.v1 import websocket
from app.core.config import settings
from app.core.logging import configure_logging
from app.core.request_context import set_request_id
from app.core.websocket.broadcaster import broadcaster
from app.db.session import init_db
from app.services.scheduler import (
@@ -17,6 +20,9 @@ from app.services.scheduler import (
)
configure_logging()
class WebSocketCORSMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request, call_next):
if request.url.path.startswith("/ws") and request.method == "GET":
@@ -28,6 +34,18 @@ class WebSocketCORSMiddleware(BaseHTTPMiddleware):
return await call_next(request)
class RequestContextMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request, call_next):
request_id = request.headers.get("X-Request-ID") or uuid4().hex
set_request_id(request_id)
try:
response = await call_next(request)
finally:
set_request_id(None)
response.headers["X-Request-ID"] = request_id
return response
@asynccontextmanager
async def lifespan(app: FastAPI):
await init_db()
@@ -58,6 +76,7 @@ app.add_middleware(
allow_headers=["*"],
)
app.add_middleware(RequestContextMiddleware)
app.add_middleware(WebSocketCORSMiddleware)
app.include_router(api_router, prefix="/api/v1")

View File

@@ -11,6 +11,9 @@ from app.models.bgp_observation import BGPObservation
from app.models.system_setting import SystemSetting
from app.models.playground_session import PlaygroundSession
from app.models.playground_message import PlaygroundMessage
from app.models.system_log import SystemLog, AuditLog
from app.models.vessel import VesselPosition, VesselStatic
from app.models.datasource_mapping import DataSourceMappingTemplate
__all__ = [
"User",
@@ -26,4 +29,9 @@ __all__ = [
"BGPAnomaly",
"BGPIncident",
"BGPObservation",
"SystemLog",
"AuditLog",
"VesselPosition",
"VesselStatic",
"DataSourceMappingTemplate",
]

View File

@@ -0,0 +1,32 @@
"""Mapping templates for user-defined data source payloads."""
from sqlalchemy import Boolean, Column, DateTime, ForeignKey, Integer, JSON, String
from sqlalchemy.sql import func
from app.db.session import Base
class DataSourceMappingTemplate(Base):
__tablename__ = "datasource_mapping_templates"
id = Column(Integer, primary_key=True, autoincrement=True)
datasource_config_id = Column(
Integer,
ForeignKey("datasource_configs.id"),
nullable=False,
index=True,
)
target_schema = Column(String(80), nullable=False, index=True)
mapping_json = Column(JSON, nullable=False, default={})
sample_payload_hash = Column(String(64), nullable=True)
validation_status = Column(String(30), nullable=False, default="draft")
version = Column(Integer, nullable=False, default=1)
is_active = Column(Boolean, nullable=False, default=False, index=True)
created_at = Column(DateTime(timezone=True), server_default=func.now())
updated_at = Column(DateTime(timezone=True), server_default=func.now(), onupdate=func.now())
def __repr__(self):
return (
f"<DataSourceMappingTemplate {self.id}: "
f"{self.datasource_config_id}/{self.target_schema}/v{self.version}>"
)

View File

@@ -0,0 +1,40 @@
from sqlalchemy import JSON, Column, DateTime, Integer, String, Text
from sqlalchemy.sql import func
from app.db.session import Base
class SystemLog(Base):
__tablename__ = "system_logs"
id = Column(Integer, primary_key=True, autoincrement=True)
occurred_at = Column(DateTime(timezone=True), server_default=func.now(), index=True)
source = Column(String(50), nullable=False, index=True)
service = Column(String(50), nullable=True)
module = Column(String(120), nullable=True)
event = Column(String(160), nullable=True, index=True)
level = Column(String(20), nullable=False, index=True)
message = Column(Text, nullable=False)
request_id = Column(String(64), nullable=True, index=True)
trace_id = Column(String(64), nullable=True)
user_id = Column(Integer, nullable=True, index=True)
category = Column(String(80), nullable=True, index=True)
context = Column(JSON, nullable=False, default=dict)
created_at = Column(DateTime(timezone=True), server_default=func.now())
class AuditLog(Base):
__tablename__ = "audit_logs"
id = Column(Integer, primary_key=True, autoincrement=True)
occurred_at = Column(DateTime(timezone=True), server_default=func.now(), index=True)
actor_id = Column(Integer, nullable=True, index=True)
actor_name = Column(String(255), nullable=True)
action = Column(String(120), nullable=False, index=True)
target_type = Column(String(80), nullable=True)
target_id = Column(String(120), nullable=True)
result = Column(String(40), nullable=True, index=True)
request_id = Column(String(64), nullable=True, index=True)
ip = Column(String(64), nullable=True)
details = Column(JSON, nullable=False, default=dict)
created_at = Column(DateTime(timezone=True), server_default=func.now())

View File

@@ -0,0 +1,75 @@
"""Vessel AIS models for live maritime tracking."""
from sqlalchemy import BigInteger, Column, DateTime, Float, Index, Integer, SmallInteger, String
from sqlalchemy.sql import func
from app.core.time import to_iso8601_utc
from app.db.session import Base
class VesselStatic(Base):
"""Slow-changing vessel identity and dimensions."""
__tablename__ = "vessel_static"
mmsi = Column(BigInteger, primary_key=True)
name = Column(String(128), nullable=True)
callsign = Column(String(16), nullable=True)
vessel_type = Column(SmallInteger, nullable=True, index=True)
vessel_type_name = Column(String(64), nullable=True, index=True)
flag = Column(String(4), nullable=True, index=True)
length = Column(Float, nullable=True)
width = Column(Float, nullable=True)
draught = Column(Float, nullable=True)
imo = Column(BigInteger, nullable=True)
updated_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now())
def to_dict(self) -> dict:
return {
"mmsi": self.mmsi,
"name": self.name,
"callsign": self.callsign,
"vessel_type": self.vessel_type,
"vessel_type_name": self.vessel_type_name,
"flag": self.flag,
"length": self.length,
"width": self.width,
"draught": self.draught,
"imo": self.imo,
"updated_at": to_iso8601_utc(self.updated_at),
}
class VesselPosition(Base):
"""Append-only AIS positions retained for short history windows."""
__tablename__ = "vessel_position"
id = Column(Integer, primary_key=True, autoincrement=True)
mmsi = Column(BigInteger, nullable=False, index=True)
lat = Column(Float, nullable=False)
lon = Column(Float, nullable=False)
sog = Column(Float, nullable=True)
cog = Column(Float, nullable=True)
heading = Column(SmallInteger, nullable=True)
nav_status = Column(SmallInteger, nullable=True, index=True)
received_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now(), index=True)
__table_args__ = (
Index("idx_vessel_pos_mmsi_time", "mmsi", "received_at"),
Index("idx_vessel_pos_time", "received_at"),
Index("idx_vessel_pos_lat_lon", "lat", "lon"),
)
def to_dict(self) -> dict:
return {
"id": self.id,
"mmsi": self.mmsi,
"lat": self.lat,
"lon": self.lon,
"sog": self.sog,
"cog": self.cog,
"heading": self.heading,
"nav_status": self.nav_status,
"received_at": to_iso8601_utc(self.received_at),
}

View File

@@ -3,9 +3,11 @@ from __future__ import annotations
import asyncio
import httpx
from fastapi import HTTPException, status
from fastapi import Depends, HTTPException, status
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.config import settings
from app.db.session import get_db
from app.schemas.ai import (
AIProviderStatusResponse,
SituationalAnalysisRequest,
@@ -14,11 +16,27 @@ from app.schemas.ai import (
class AIProviderClient:
def __init__(self) -> None:
self.service_url = settings.AI_PROVIDER_SERVICE_URL.rstrip("/")
self.service_token = settings.AI_PROVIDER_SERVICE_TOKEN
self.timeout = settings.AI_PROVIDER_TIMEOUT_SECONDS
self.retry_attempts = max(settings.AI_PROVIDER_RETRY_ATTEMPTS, 1)
def __init__(
self,
*,
service_url: str | None = None,
service_token: str | None = None,
timeout: int | None = None,
retry_attempts: int | None = None,
llm_config: dict | None = None,
) -> None:
self.service_url = (
service_url if service_url is not None else settings.AI_PROVIDER_SERVICE_URL
).rstrip("/")
self.service_token = (
service_token if service_token is not None else settings.AI_PROVIDER_SERVICE_TOKEN
)
self.timeout = timeout if timeout is not None else settings.AI_PROVIDER_TIMEOUT_SECONDS
self.retry_attempts = max(
retry_attempts if retry_attempts is not None else settings.AI_PROVIDER_RETRY_ATTEMPTS,
1,
)
self.llm_config = llm_config or {}
def _headers(self, request_id: str | None = None) -> dict[str, str]:
headers = {"Content-Type": "application/json"}
@@ -26,6 +44,19 @@ class AIProviderClient:
headers["X-Provider-Token"] = self.service_token
if request_id:
headers["X-Request-ID"] = request_id
llm_header_map = {
"provider": "X-AI-Provider",
"provider_api": "X-AI-Provider-API",
"base_url": "X-AI-Base-URL",
"api_key": "X-AI-API-Key",
"model": "X-AI-Model",
"max_tokens": "X-AI-Max-Tokens",
"anthropic_version": "X-AI-Anthropic-Version",
}
for key, header_name in llm_header_map.items():
value = self.llm_config.get(key)
if value not in (None, ""):
headers[header_name] = str(value)
return headers
async def get_status(self, request_id: str | None = None) -> AIProviderStatusResponse:
@@ -105,5 +136,14 @@ class AIProviderClient:
)
def get_ai_provider_client() -> AIProviderClient:
return AIProviderClient()
async def get_ai_provider_client(db: AsyncSession = Depends(get_db)) -> AIProviderClient:
from app.api.v1.settings import get_runtime_ai_provider_config
runtime_config = await get_runtime_ai_provider_config(db)
return AIProviderClient(
service_url=runtime_config["service_url"],
service_token=runtime_config["service_token"],
timeout=runtime_config["timeout_seconds"],
retry_attempts=runtime_config["retry_attempts"],
llm_config=runtime_config.get("llm_config") or {},
)

View File

@@ -36,6 +36,7 @@ from app.services.collectors.iptoasn import IPtoASNPrefixGeoCollector
from app.services.collectors.opengeofeed import OpenGeoFeedPrefixGeoCollector
from app.services.collectors.nro_delegated import NRODelegatedPrefixGeoCollector
from app.services.collectors.news_live_streams import NewsLiveStreamsCollector
from app.services.collectors.vessel_ais import VesselAISCollector
collector_registry.register(TOP500Collector())
collector_registry.register(EpochAIGPUCollector())
@@ -63,3 +64,4 @@ collector_registry.register(IPtoASNPrefixGeoCollector())
collector_registry.register(OpenGeoFeedPrefixGeoCollector())
collector_registry.register(NRODelegatedPrefixGeoCollector())
collector_registry.register(NewsLiveStreamsCollector())
collector_registry.register(VesselAISCollector())

View File

@@ -46,6 +46,9 @@ class CelesTrakTLECollector(BaseCollector):
if response.status_code == 200:
data = response.json()
if isinstance(data, list):
for item in data:
if isinstance(item, dict):
item["_celestrak_group"] = group
all_satellites.extend(data)
print(f"CelesTrak: Fetched {len(data)} satellites from group '{group}'")
except Exception as e:
@@ -78,6 +81,7 @@ class CelesTrakTLECollector(BaseCollector):
"name": item.get("OBJECT_NAME", "Unknown"),
"reference_date": item.get("EPOCH", ""),
"metadata": {
"constellation_group": item.get("_celestrak_group"),
"norad_cat_id": item.get("NORAD_CAT_ID"),
"international_designator": item.get("OBJECT_ID"),
"epoch": item.get("EPOCH"),

View File

@@ -1,10 +1,16 @@
from __future__ import annotations
import asyncio
import base64
from datetime import UTC, datetime
from typing import Any
from urllib.parse import urlparse
import httpx
from sqlalchemy import select
from app.core.data_sources import get_data_sources_config
from app.models.datasource_config import DataSourceConfig
from app.services.collectors.base import BaseCollector
@@ -18,52 +24,537 @@ class NewsLiveStreamsCollector(BaseCollector):
data_type = "news_live_stream"
fail_on_empty = False
DEFAULT_TIMEOUT = 45.0
DEFAULT_HEADERS = {
"User-Agent": "Planet-Intelligence-System/1.0 (Python/collector)",
"Accept": "application/json",
}
RESPONSE_CANDIDATE_KEYS = ("sources", "streams", "channels", "items", "results", "data")
DEFAULT_ADAPTER = "iptv_org"
DEFAULT_IPTV_ORG_STREAMS_URL = "https://iptv-org.github.io/api/streams.json"
DEFAULT_IPTV_ORG_LOGOS_URL = "https://iptv-org.github.io/api/logos.json"
DEFAULT_IPTV_ORG_NEWS_CATEGORIES = ("news", "business", "weather")
DEFAULT_IPTV_ORG_EXCLUDE_CATEGORIES = ("music", "sports", "kids", "entertainment")
DEFAULT_IPTV_ORG_MAX_SOURCES = 120
async def fetch(self) -> list[dict[str, Any]]:
request_url = (self._resolved_url or "").strip()
if not request_url:
return []
async with httpx.AsyncClient(timeout=45.0, follow_redirects=True) as client:
response = await client.get(
datasource_config = await self._load_datasource_config()
effective_config = self._get_effective_config(datasource_config)
adapter = str(effective_config.get("adapter") or "").strip().lower()
if adapter == "iptv_org":
return await self._fetch_iptv_org(request_url, effective_config)
request_headers = self._build_request_headers(datasource_config)
request_config = self._get_request_config(datasource_config)
request_params = self._build_request_params(datasource_config)
request_json = self._build_request_json_body(datasource_config)
request_data = self._build_request_form_body(datasource_config)
timeout = self._get_timeout(datasource_config)
async with httpx.AsyncClient(timeout=timeout, follow_redirects=True) as client:
response = await client.request(
request_config["method"],
request_url,
headers={
"User-Agent": "Planet-Intelligence-System/1.0 (Python/collector)",
"Accept": "application/json",
},
headers=request_headers,
params=request_params or None,
json=request_json,
data=request_data,
)
response.raise_for_status()
return self.parse_response(response.json())
return self.parse_response(
response.json(),
response_path=request_config["response_path"],
)
def parse_response(self, response: Any) -> list[dict[str, Any]]:
if isinstance(response, dict):
candidates = response.get("sources") or response.get("streams") or response.get("data") or []
elif isinstance(response, list):
candidates = response
async def _load_datasource_config(self) -> DataSourceConfig | None:
if not self._db_session:
return None
result = await self._db_session.execute(
select(DataSourceConfig)
.where(DataSourceConfig.name == self.name)
.where(DataSourceConfig.is_active.is_(True))
.limit(1)
)
return result.scalar_one_or_none()
def _get_effective_config(self, datasource_config: DataSourceConfig | None) -> dict[str, Any]:
payload = dict(datasource_config.config or {}) if datasource_config else {}
if payload:
return payload
yaml_config = get_data_sources_config()
return {
"adapter": self.DEFAULT_ADAPTER,
"streams_url": yaml_config.get_yaml_value("news_live_streams.streams_url")
or self.DEFAULT_IPTV_ORG_STREAMS_URL,
"logos_url": yaml_config.get_yaml_value("news_live_streams.logos_url")
or self.DEFAULT_IPTV_ORG_LOGOS_URL,
"news_categories": list(self.DEFAULT_IPTV_ORG_NEWS_CATEGORIES),
"exclude_categories": list(self.DEFAULT_IPTV_ORG_EXCLUDE_CATEGORIES),
"max_sources": self.DEFAULT_IPTV_ORG_MAX_SOURCES,
}
def _get_request_config(self, datasource_config: DataSourceConfig | None) -> dict[str, Any]:
payload = self._get_effective_config(datasource_config)
raw_method = payload.get("method") or payload.get("request_method") or "GET"
method = str(raw_method).strip().upper() or "GET"
if method not in {"GET", "POST"}:
method = "GET"
response_path = payload.get("response_path") or payload.get("payload_path") or payload.get("items_path")
if isinstance(response_path, str):
response_path = response_path.strip()
else:
candidates = []
response_path = None
return {
"method": method,
"response_path": response_path or None,
}
def _get_timeout(self, datasource_config: DataSourceConfig | None) -> float:
payload = self._get_effective_config(datasource_config)
try:
return float(payload.get("timeout", self.DEFAULT_TIMEOUT))
except (TypeError, ValueError):
return self.DEFAULT_TIMEOUT
def _build_request_headers(self, datasource_config: DataSourceConfig | None) -> dict[str, str]:
headers = dict(self.DEFAULT_HEADERS)
if datasource_config:
headers.update(self._normalize_headers(datasource_config.headers))
headers.update(self._build_auth_headers(datasource_config))
return headers
def _build_request_params(self, datasource_config: DataSourceConfig | None) -> dict[str, Any]:
params: dict[str, Any] = {}
if not datasource_config:
return params
payload = datasource_config.config or {}
candidate = payload.get("params") or payload.get("query_params")
if isinstance(candidate, dict):
params.update(candidate)
if datasource_config.auth_type == "api_key":
auth_config = datasource_config.auth_config or {}
if str(auth_config.get("in") or auth_config.get("location") or "header").lower() == "query":
api_key = auth_config.get("api_key")
key_name = auth_config.get("key_name") or auth_config.get("param_name") or "api_key"
if api_key and key_name:
params[str(key_name)] = api_key
return params
def _build_request_json_body(self, datasource_config: DataSourceConfig | None) -> Any:
if not datasource_config:
return None
payload = datasource_config.config or {}
body = payload.get("json_body")
if body is None and str(payload.get("body_type") or "").lower() in {"json", ""}:
candidate = payload.get("body")
if isinstance(candidate, (dict, list)):
body = candidate
return body
def _build_request_form_body(self, datasource_config: DataSourceConfig | None) -> Any:
if not datasource_config:
return None
payload = datasource_config.config or {}
form_body = payload.get("form_body")
if form_body is not None:
return form_body
if str(payload.get("body_type") or "").lower() == "form":
candidate = payload.get("body")
if isinstance(candidate, dict):
return candidate
return None
def _normalize_headers(self, headers: Any) -> dict[str, str]:
if not isinstance(headers, dict):
return {}
normalized: dict[str, str] = {}
for key, value in headers.items():
header_name = str(key).strip()
if not header_name or value is None:
continue
normalized[header_name] = str(value)
return normalized
def _build_auth_headers(self, datasource_config: DataSourceConfig | None) -> dict[str, str]:
if not datasource_config:
return {}
auth_type = str(datasource_config.auth_type or "none").lower()
auth_config = datasource_config.auth_config or {}
if auth_type == "bearer" and auth_config.get("token"):
return {"Authorization": f"Bearer {auth_config['token']}"}
if auth_type == "api_key" and auth_config.get("api_key"):
location = str(auth_config.get("in") or auth_config.get("location") or "header").lower()
if location == "query":
return {}
key_name = auth_config.get("key_name") or "X-API-Key"
return {str(key_name): str(auth_config["api_key"])}
if auth_type == "basic":
username = str(auth_config.get("username") or "")
password = str(auth_config.get("password") or "")
encoded = base64.b64encode(f"{username}:{password}".encode()).decode()
return {"Authorization": f"Basic {encoded}"}
return {}
def _extract_candidates(self, response: Any, response_path: str | None) -> list[Any]:
if response_path:
extracted = self._extract_from_path(response, response_path)
if isinstance(extracted, list):
return extracted
if isinstance(extracted, dict):
for key in self.RESPONSE_CANDIDATE_KEYS:
nested = extracted.get(key)
if isinstance(nested, list):
return nested
return [extracted]
if isinstance(response, dict):
for key in self.RESPONSE_CANDIDATE_KEYS:
nested = response.get(key)
if isinstance(nested, list):
return nested
return []
if isinstance(response, list):
return response
return []
def _extract_from_path(self, payload: Any, path: str) -> Any:
current = payload
for segment in (part.strip() for part in path.split(".") if part.strip()):
if isinstance(current, dict):
current = current.get(segment)
continue
if isinstance(current, list):
try:
current = current[int(segment)]
except (TypeError, ValueError, IndexError):
return None
continue
return None
return current
def _infer_source_type(self, item: dict[str, Any]) -> str:
explicit = str(item.get("source_type") or item.get("type") or "").strip().lower()
if explicit in {"iframe", "hls", "video", "external", "youtube"}:
return explicit
youtube_video_id = self._clean_text(
item.get("youtube_video_id")
or item.get("video_id")
or item.get("youtubeVideoId")
)
youtube_channel = self._clean_text(item.get("youtube_channel") or item.get("channel_handle"))
embed_url = self._clean_url(item.get("embed_url") or item.get("embed") or item.get("page_url"))
stream_url = self._clean_url(item.get("stream_url") or item.get("stream") or item.get("playback_url") or item.get("hls_url"))
homepage_url = self._clean_url(item.get("homepage_url") or item.get("source_url") or item.get("website"))
if youtube_video_id or youtube_channel:
return "youtube"
if stream_url.endswith(".m3u8"):
return "hls"
if stream_url:
return "video"
if embed_url:
parsed = urlparse(embed_url)
if "youtube.com" in (parsed.netloc or "") or "youtu.be" in (parsed.netloc or ""):
return "youtube"
return "iframe"
if homepage_url:
return "external"
return "iframe"
def _parse_enabled(self, item: dict[str, Any]) -> bool:
if "is_enabled" in item:
return self._to_bool(item.get("is_enabled"), default=True)
if "enabled" in item:
return self._to_bool(item.get("enabled"), default=True)
if "active" in item:
return self._to_bool(item.get("active"), default=True)
if "status" in item:
status = str(item.get("status") or "").strip().lower()
if status in {"disabled", "inactive", "offline"}:
return False
if status in {"enabled", "active", "online", "live"}:
return True
return True
def _to_bool(self, value: Any, *, default: bool) -> bool:
if isinstance(value, bool):
return value
if value in (None, ""):
return default
if isinstance(value, str):
lowered = value.strip().lower()
if lowered in {"1", "true", "yes", "on", "enabled", "active", "online", "live"}:
return True
if lowered in {"0", "false", "no", "off", "disabled", "inactive", "offline"}:
return False
return bool(value)
def _clean_text(self, value: Any) -> str:
if value is None:
return ""
return str(value).strip()
def _clean_url(self, value: Any) -> str:
text = self._clean_text(value)
if not text:
return ""
parsed = urlparse(text)
if parsed.scheme and parsed.scheme not in {"http", "https"}:
return ""
if parsed.scheme and not parsed.netloc:
return ""
return text
async def _fetch_iptv_org(self, channels_url: str, collector_config: dict[str, Any]) -> list[dict[str, Any]]:
streams_url = self._clean_url(collector_config.get("streams_url")) or self.DEFAULT_IPTV_ORG_STREAMS_URL
logos_url = self._clean_url(collector_config.get("logos_url")) or self.DEFAULT_IPTV_ORG_LOGOS_URL
news_categories = {
self._clean_text(value).lower()
for value in (collector_config.get("news_categories") or self.DEFAULT_IPTV_ORG_NEWS_CATEGORIES)
if self._clean_text(value)
}
exclude_categories = {
self._clean_text(value).lower()
for value in (collector_config.get("exclude_categories") or self.DEFAULT_IPTV_ORG_EXCLUDE_CATEGORIES)
if self._clean_text(value)
}
try:
max_sources = int(collector_config.get("max_sources", self.DEFAULT_IPTV_ORG_MAX_SOURCES))
except (TypeError, ValueError):
max_sources = self.DEFAULT_IPTV_ORG_MAX_SOURCES
timeout = self.DEFAULT_TIMEOUT
try:
timeout = float(collector_config.get("timeout", self.DEFAULT_TIMEOUT))
except (TypeError, ValueError):
timeout = self.DEFAULT_TIMEOUT
async with httpx.AsyncClient(timeout=timeout, follow_redirects=True) as client:
channels_payload, streams_payload, logos_payload = await self._gather_iptv_org_payloads(
client,
channels_url,
streams_url,
logos_url,
)
channels = channels_payload if isinstance(channels_payload, list) else []
streams = streams_payload if isinstance(streams_payload, list) else []
logos = logos_payload if isinstance(logos_payload, list) else []
logo_by_channel = {
self._clean_text(item.get("channel")): self._clean_url(item.get("url"))
for item in logos
if isinstance(item, dict) and self._clean_text(item.get("channel")) and self._clean_url(item.get("url"))
}
streams_by_channel: dict[str, list[dict[str, Any]]] = {}
for stream in streams:
if not isinstance(stream, dict):
continue
channel_id = self._clean_text(stream.get("channel"))
if not channel_id:
continue
streams_by_channel.setdefault(channel_id, []).append(stream)
normalized: list[dict[str, Any]] = []
for channel in channels:
if not isinstance(channel, dict):
continue
categories = [
self._clean_text(value).lower()
for value in (channel.get("categories") or [])
if self._clean_text(value)
]
if news_categories and not any(category in news_categories for category in categories):
continue
if exclude_categories and any(category in exclude_categories for category in categories):
continue
if channel.get("is_nsfw") is True:
continue
if channel.get("closed"):
continue
channel_id = self._clean_text(channel.get("id"))
if not channel_id:
continue
stream = self._pick_iptv_org_stream(streams_by_channel.get(channel_id) or [])
if not stream:
continue
stream_url = self._clean_url(stream.get("url"))
if not stream_url:
continue
name = self._clean_text(channel.get("name")) or channel_id
notes_parts = [
f"Imported from IPTV-org catalog ({channel_id})",
f"Categories: {', '.join(categories)}" if categories else "",
f"Quality: {self._clean_text(stream.get('quality'))}" if self._clean_text(stream.get("quality")) else "",
]
metadata = {
"provider": self._clean_text(channel.get("network")) or "IPTV-org",
"region": self._clean_text(channel.get("country")) or "Global",
"language": "und",
"source_type": "hls" if stream_url.endswith(".m3u8") else "video",
"embed_url": "",
"stream_url": stream_url,
"homepage_url": self._clean_url(channel.get("website")),
"poster_url": logo_by_channel.get(channel_id, ""),
"youtube_video_id": "",
"youtube_channel": "",
"sort_order": 400 + len(normalized),
"notes": "; ".join(part for part in notes_parts if part),
"is_enabled": True,
"collector_adapter": "iptv_org",
"channel_id": channel_id,
"categories": categories,
"quality": self._clean_text(stream.get("quality")),
"stream_label": self._clean_text(stream.get("label") or stream.get("title")),
"stream_referrer": self._clean_text(stream.get("referrer")),
"stream_user_agent": self._clean_text(stream.get("user_agent")),
}
normalized.append(
{
"source_id": channel_id,
"name": name,
"description": metadata["notes"],
"metadata": metadata,
"reference_date": datetime.now(UTC).isoformat(),
}
)
if len(normalized) >= max_sources:
break
return normalized
async def _gather_iptv_org_payloads(
self,
client: httpx.AsyncClient,
channels_url: str,
streams_url: str,
logos_url: str,
) -> tuple[Any, Any, Any]:
headers = dict(self.DEFAULT_HEADERS)
channels_payload, streams_payload, logos_payload = await asyncio.gather(
client.get(channels_url, headers=headers),
client.get(streams_url, headers=headers),
client.get(logos_url, headers=headers),
)
channels_payload.raise_for_status()
streams_payload.raise_for_status()
logos_payload.raise_for_status()
return channels_payload.json(), streams_payload.json(), logos_payload.json()
def _pick_iptv_org_stream(self, streams: list[dict[str, Any]]) -> dict[str, Any] | None:
if not streams:
return None
def score(stream: dict[str, Any]) -> tuple[int, int]:
url = self._clean_url(stream.get("url"))
quality = self._clean_text(stream.get("quality")).lower()
quality_score = 0
if quality.endswith("p"):
try:
quality_score = int(quality[:-1])
except ValueError:
quality_score = 0
stream_score = 1000 if url.endswith(".m3u8") else 0
return stream_score, quality_score
sorted_streams = sorted(streams, key=score, reverse=True)
return sorted_streams[0]
def parse_response(self, response: Any, *, response_path: str | None = None) -> list[dict[str, Any]]:
candidates = self._extract_candidates(response, response_path)
normalized: list[dict[str, Any]] = []
for index, item in enumerate(candidates):
if not isinstance(item, dict):
continue
stream_id = item.get("id") or item.get("source_id") or item.get("slug") or f"news-live-{index + 1}"
name = str(item.get("name") or item.get("title") or f"News Live {index + 1}").strip()
stream_id = (
item.get("id")
or item.get("source_id")
or item.get("slug")
or item.get("channel_id")
or item.get("code")
or f"news-live-{index + 1}"
)
name = self._clean_text(
item.get("name")
or item.get("title")
or item.get("channel")
or item.get("display_name")
or f"News Live {index + 1}"
)
if not name:
continue
source_type = self._infer_source_type(item)
stream_url = self._clean_url(
item.get("stream_url")
or item.get("stream")
or item.get("playback_url")
or item.get("hls_url")
or item.get("m3u8_url")
)
embed_url = self._clean_url(
item.get("embed_url")
or item.get("embed")
or item.get("page_url")
or (item.get("url") if source_type == "iframe" else "")
)
homepage_url = self._clean_url(
item.get("homepage_url")
or item.get("source_url")
or item.get("website")
or item.get("url")
)
metadata = {
"provider": item.get("provider") or item.get("publisher") or "Collector",
"region": item.get("region") or item.get("country") or "Global",
"language": item.get("language") or "und",
"source_type": item.get("source_type") or "iframe",
"embed_url": item.get("embed_url") or item.get("url") or "",
"stream_url": item.get("stream_url") or "",
"homepage_url": item.get("homepage_url") or item.get("source_url") or "",
"poster_url": item.get("poster_url") or "",
"provider": self._clean_text(item.get("provider") or item.get("publisher") or item.get("network")) or "Collector",
"region": self._clean_text(item.get("region") or item.get("country") or item.get("market")) or "Global",
"language": self._clean_text(item.get("language") or item.get("lang") or item.get("locale")) or "und",
"source_type": source_type,
"embed_url": embed_url,
"stream_url": stream_url,
"homepage_url": homepage_url,
"poster_url": self._clean_url(item.get("poster_url") or item.get("thumbnail_url") or item.get("logo_url")),
"youtube_video_id": self._clean_text(
item.get("youtube_video_id")
or item.get("video_id")
or item.get("youtubeVideoId")
),
"youtube_channel": self._clean_text(
item.get("youtube_channel")
or item.get("channel_handle")
or item.get("youtubeChannel")
),
"sort_order": item.get("sort_order", 200 + index),
"notes": item.get("notes") or item.get("description") or "",
"is_enabled": item.get("is_enabled", True),
"notes": self._clean_text(item.get("notes") or item.get("description") or item.get("summary")),
"is_enabled": self._parse_enabled(item),
}
normalized.append(
@@ -72,7 +563,7 @@ class NewsLiveStreamsCollector(BaseCollector):
"name": name,
"description": metadata["notes"],
"metadata": metadata,
"reference_date": item.get("reference_date", datetime.now(UTC).isoformat()),
"reference_date": item.get("reference_date") or datetime.now(UTC).isoformat(),
}
)

View File

@@ -0,0 +1,326 @@
"""BarentsWatch AIS collector for vessel tracking."""
from datetime import UTC, datetime, timedelta
import os
from typing import Any
import httpx
from sqlalchemy import delete, select
from sqlalchemy.ext.asyncio import AsyncSession
from app.models.vessel import VesselPosition, VesselStatic
from app.services.collectors.base import BaseCollector
BARENTSWATCH_LATEST_URL = "https://live.ais.barentswatch.no/v1/latest/combined"
BARENTSWATCH_TOKEN_URL = "https://id.barentswatch.no/connect/token"
VESSEL_TYPE_NAMES = {
30: "Fishing",
35: "Military",
60: "Passenger",
70: "Cargo",
80: "Tanker",
}
class VesselAISCollector(BaseCollector):
"""Collect latest AIS positions and append them to vessel tables."""
name = "barentswatch_vessels"
priority = "P1"
module = "L4"
frequency_hours = 1
data_type = "vessel_ais"
@property
def base_url(self) -> str:
return self._resolved_url or BARENTSWATCH_LATEST_URL
async def _load_datasource_config(self) -> dict[str, Any]:
if not self._db_session:
return {}
try:
from sqlalchemy import select
from app.models.datasource_config import DataSourceConfig
result = await self._db_session.execute(
select(DataSourceConfig)
.where(DataSourceConfig.name == self.name)
.where(DataSourceConfig.is_active.is_(True))
)
datasource_config = result.scalar_one_or_none()
except Exception:
return {}
if not datasource_config:
return {}
return {
"auth_config": datasource_config.auth_config or {},
"config": datasource_config.config or {},
}
async def _get_access_token(self, client: httpx.AsyncClient) -> str | None:
datasource_config = await self._load_datasource_config()
auth_config = datasource_config.get("auth_config") or {}
config = datasource_config.get("config") or {}
client_id = (
auth_config.get("client_id")
or config.get("client_id")
or os.getenv("BARENTSWATCH_CLIENT_ID")
or os.getenv("BARRENTSWATCH_CLIENT_ID")
)
client_secret = (
auth_config.get("client_secret")
or config.get("client_secret")
or os.getenv("BARENTSWATCH_CLIENT_SECRET")
or os.getenv("BARRENTSWATCH_CLIENT_SECRET")
)
if not client_id or not client_secret:
return None
response = await client.post(
BARENTSWATCH_TOKEN_URL,
data={
"client_id": client_id,
"client_secret": client_secret,
"scope": "ais",
"grant_type": "client_credentials",
},
headers={"Content-Type": "application/x-www-form-urlencoded"},
)
response.raise_for_status()
payload = response.json()
token = payload.get("access_token")
return str(token) if token else None
async def fetch(self) -> list[dict[str, Any]]:
async with httpx.AsyncClient(timeout=60.0) as client:
headers: dict[str, str] = {}
token = await self._get_access_token(client)
if token:
headers["Authorization"] = f"Bearer {token}"
response = await client.get(self.base_url, headers=headers)
if response.status_code == 401 and not token:
return self._get_sample_data()
response.raise_for_status()
payload = response.json()
if isinstance(payload, list):
return [item for item in payload if isinstance(item, dict)]
if isinstance(payload, dict):
for key in ("features", "data", "items", "vessels"):
value = payload.get(key)
if isinstance(value, list):
if key == "features":
return [
{
**(item.get("properties") or {}),
"geometry": item.get("geometry"),
}
for item in value
if isinstance(item, dict)
]
return [item for item in value if isinstance(item, dict)]
return self._get_sample_data()
def transform(self, raw_data: list[dict[str, Any]]) -> list[dict[str, Any]]:
transformed = []
for item in raw_data:
record = self._normalize_record(item)
if record:
transformed.append(record)
return transformed
async def _save_data(
self,
db: AsyncSession,
data: list[dict[str, Any]],
task_id: int | None = None,
snapshot_id: int | None = None,
) -> int:
now = datetime.now(UTC)
records_added = 0
for index, item in enumerate(data):
static = await db.get(VesselStatic, item["mmsi"])
if static is None:
static = VesselStatic(mmsi=item["mmsi"])
db.add(static)
for field in (
"name",
"callsign",
"vessel_type",
"vessel_type_name",
"flag",
"length",
"width",
"draught",
"imo",
):
value = item.get(field)
if value not in (None, ""):
setattr(static, field, value)
static.updated_at = now
db.add(
VesselPosition(
mmsi=item["mmsi"],
lat=item["lat"],
lon=item["lon"],
sog=item.get("sog"),
cog=item.get("cog"),
heading=item.get("heading"),
nav_status=item.get("nav_status"),
received_at=item.get("received_at") or now,
)
)
records_added += 1
if (index + 1) % 1000 == 0:
await self.update_progress(index + 1, commit=True)
await db.execute(
delete(VesselPosition).where(VesselPosition.received_at < now - timedelta(hours=24))
)
await db.commit()
await self.update_progress(records_added, force=True)
return records_added
def _normalize_record(self, item: dict[str, Any]) -> dict[str, Any] | None:
mmsi = _as_int(_pick(item, "mmsi", "MMSI", "Mmsi"))
lat = _as_float(_pick(item, "lat", "latitude", "Latitude"))
lon = _as_float(_pick(item, "lon", "lng", "longitude", "Longitude"))
geometry = item.get("geometry")
coordinates = geometry.get("coordinates") if isinstance(geometry, dict) else None
if (lat is None or lon is None) and isinstance(coordinates, list) and len(coordinates) >= 2:
lon = _as_float(coordinates[0])
lat = _as_float(coordinates[1])
if mmsi is None or lat is None or lon is None:
return None
if not (-90 <= lat <= 90 and -180 <= lon <= 180):
return None
vessel_type = _as_int(_pick(item, "vessel_type", "shipType", "ship_type", "ShipType"))
vessel_type_name = (
_pick(item, "vessel_type_name", "shipTypeName", "ship_type_name", "VesselTypeName")
or _vessel_type_name(vessel_type)
)
received_at = _parse_datetime(_pick(item, "received_at", "timestamp", "time", "msgtime"))
return {
"mmsi": mmsi,
"name": _pick(item, "name", "shipName", "ship_name", "Name"),
"callsign": _pick(item, "callsign", "callSign", "CallSign"),
"vessel_type": vessel_type,
"vessel_type_name": vessel_type_name,
"flag": _pick(item, "flag", "country", "Flag"),
"length": _as_float(_pick(item, "length", "shipLength", "Length")),
"width": _as_float(_pick(item, "width", "shipWidth", "Width")),
"draught": _as_float(_pick(item, "draught", "draft", "Draught")),
"imo": _as_int(_pick(item, "imo", "IMO", "imoNumber")),
"lat": lat,
"lon": lon,
"sog": _as_float(_pick(item, "sog", "speedOverGround", "SOG")),
"cog": _as_float(_pick(item, "cog", "courseOverGround", "COG")),
"heading": _as_int(_pick(item, "heading", "trueHeading", "Heading")),
"nav_status": _as_int(_pick(item, "nav_status", "navStatus", "NavigationalStatus")),
"received_at": received_at,
}
def _get_sample_data(self) -> list[dict[str, Any]]:
return [
{
"mmsi": 257123000,
"name": "OSLO TRADER",
"lat": 59.91,
"lon": 10.73,
"sog": 12.4,
"cog": 214,
"heading": 215,
"nav_status": 0,
"vessel_type": 70,
"vessel_type_name": "Cargo",
"flag": "NO",
"length": 185,
},
{
"mmsi": 257456000,
"name": "NORDIC FJORD",
"lat": 60.39,
"lon": 5.32,
"sog": 0.2,
"cog": 82,
"heading": 80,
"nav_status": 1,
"vessel_type": 60,
"vessel_type_name": "Passenger",
"flag": "NO",
"length": 126,
},
]
def _pick(item: dict[str, Any], *keys: str) -> Any:
for key in keys:
if key in item and item[key] not in (None, ""):
return item[key]
return None
def _as_float(value: Any) -> float | None:
try:
if value in (None, ""):
return None
return float(value)
except (TypeError, ValueError):
return None
def _as_int(value: Any) -> int | None:
try:
if value in (None, ""):
return None
return int(float(value))
except (TypeError, ValueError):
return None
def _parse_datetime(value: Any) -> datetime | None:
if isinstance(value, datetime):
return value if value.tzinfo else value.replace(tzinfo=UTC)
if not value:
return None
if isinstance(value, (int, float)):
timestamp = float(value)
if timestamp > 10_000_000_000:
timestamp /= 1000
return datetime.fromtimestamp(timestamp, UTC)
if isinstance(value, str):
try:
parsed = datetime.fromisoformat(value.replace("Z", "+00:00"))
return parsed if parsed.tzinfo else parsed.replace(tzinfo=UTC)
except ValueError:
return None
return None
def _vessel_type_name(vessel_type: int | None) -> str:
if vessel_type is None:
return "Other"
if 70 <= vessel_type <= 79:
return "Cargo"
if 80 <= vessel_type <= 89:
return "Tanker"
if 60 <= vessel_type <= 69:
return "Passenger"
if vessel_type == 30:
return "Fishing"
if vessel_type == 35:
return "Military"
return VESSEL_TYPE_NAMES.get(vessel_type, "Other")

View File

@@ -0,0 +1,358 @@
"""Deterministic mapping support for custom data sources."""
from __future__ import annotations
import hashlib
import json
import re
from datetime import UTC, datetime
from typing import Any
from sqlalchemy.ext.asyncio import AsyncSession
from app.core.target_schema_registry import TargetSchema, get_target_schema
SECRET_KEY_PATTERN = re.compile(
r"(token|secret|password|passwd|authorization|api[_-]?key|client[_-]?secret)",
re.IGNORECASE,
)
class MappingError(ValueError):
"""Raised when a mapping definition cannot be executed."""
def stable_payload_hash(payload: Any) -> str:
encoded = json.dumps(payload, ensure_ascii=False, sort_keys=True, default=str).encode()
return hashlib.sha256(encoded).hexdigest()
def redact_for_llm(value: Any) -> Any:
if isinstance(value, dict):
redacted = {}
for key, item in value.items():
if SECRET_KEY_PATTERN.search(str(key)):
redacted[key] = "[REDACTED]"
else:
redacted[key] = redact_for_llm(item)
return redacted
if isinstance(value, list):
return [redact_for_llm(item) for item in value[:20]]
return value
def extract_path(payload: Any, path: str | None) -> Any:
if not path or path == "$":
return payload
normalized = path.strip()
if normalized.startswith("$."):
normalized = normalized[2:]
elif normalized.startswith("$"):
normalized = normalized[1:]
normalized = normalized.strip(".")
if not normalized:
return payload
current = payload
for raw_segment in normalized.split("."):
segment = raw_segment.strip()
if not segment:
continue
list_all = segment.endswith("[*]")
if list_all:
segment = segment[:-3]
index = None
match = re.fullmatch(r"(.+)\[(\d+)\]", segment)
if match:
segment = match.group(1)
index = int(match.group(2))
if segment:
if isinstance(current, dict):
current = current.get(segment)
else:
return None
if list_all:
return current if isinstance(current, list) else []
if index is not None:
if not isinstance(current, list) or index >= len(current):
return None
current = current[index]
return current
def _convert_value(value: Any, target_type: str | None) -> Any:
if value is None or target_type in (None, "", "any"):
return value
if target_type == "string":
return str(value)
if target_type == "integer":
return int(value)
if target_type == "float":
return float(value)
if target_type == "boolean":
if isinstance(value, bool):
return value
if isinstance(value, str):
return value.strip().lower() in {"1", "true", "yes", "y", "on"}
return bool(value)
if target_type == "datetime":
if isinstance(value, datetime):
return value
if isinstance(value, (int, float)):
return datetime.fromtimestamp(value)
if isinstance(value, str):
return datetime.fromisoformat(value.replace("Z", "+00:00"))
return value
if target_type == "object":
if isinstance(value, dict):
return value
raise ValueError("expected object")
if target_type == "array":
if isinstance(value, list):
return value
raise ValueError("expected array")
return value
def _apply_enum(value: Any, enum_map: Any) -> Any:
if not isinstance(enum_map, dict):
return value
key = str(value)
return enum_map.get(key, enum_map.get(value, value))
def _map_one(item: Any, field_mapping: dict[str, Any]) -> tuple[dict[str, Any], list[str]]:
output: dict[str, Any] = {}
errors: list[str] = []
for field_name, rule in field_mapping.items():
if isinstance(rule, str):
rule = {"path": rule}
if not isinstance(rule, dict):
errors.append(f"{field_name}: mapping rule must be an object or path string")
continue
value = extract_path(item, rule.get("path"))
if value is None and "default" in rule:
value = rule.get("default")
value = _apply_enum(value, rule.get("enum"))
try:
value = _convert_value(value, rule.get("type"))
except (TypeError, ValueError) as exc:
errors.append(f"{field_name}: failed to convert value {value!r}: {exc}")
continue
if value is not None or rule.get("include_null", False):
output[field_name] = value
return output, errors
def execute_mapping(
payload: Any,
mapping_json: dict[str, Any],
target_schema: str | TargetSchema,
*,
limit: int | None = None,
) -> dict[str, Any]:
schema = get_target_schema(target_schema) if isinstance(target_schema, str) else target_schema
source = mapping_json.get("source") or {}
fields = mapping_json.get("fields")
if not isinstance(fields, dict) or not fields:
raise MappingError("mapping_json.fields must be a non-empty object")
items_path = source.get("items_path") or mapping_json.get("items_path") or "$"
items = extract_path(payload, items_path)
if isinstance(items, dict):
items = [items]
elif not isinstance(items, list):
items = []
if limit is not None:
items = items[:limit]
mapped_records: list[dict[str, Any]] = []
errors: list[dict[str, Any]] = []
for index, item in enumerate(items):
mapped, mapping_errors = _map_one(item, fields)
validated, validation_errors = schema.validate_record(mapped)
all_errors = mapping_errors + validation_errors
if all_errors:
errors.append({"index": index, "errors": all_errors, "record": mapped})
continue
if validated is not None:
mapped_records.append(validated)
return {
"target_schema": schema.key,
"total_items": len(items),
"mapped_count": len(mapped_records),
"failed_count": len(errors),
"records": mapped_records,
"errors": errors,
}
def build_heuristic_mapping(sample_payload: Any, target_schema_key: str) -> dict[str, Any]:
schema = get_target_schema(target_schema_key)
items_path = "$"
sample_item = sample_payload
if isinstance(sample_payload, dict):
for key in ("data", "items", "results", "features", "vessels"):
candidate = sample_payload.get(key)
if isinstance(candidate, list) and candidate:
items_path = f"$.{key}[*]"
sample_item = candidate[0]
break
elif isinstance(sample_payload, list) and sample_payload:
items_path = "$"
sample_item = sample_payload[0]
available = _flatten_keys(sample_item if isinstance(sample_item, dict) else {})
fields: dict[str, Any] = {}
for field in schema.fields:
candidate = _best_field_match(field.name, available)
if candidate:
fields[field.name] = {"path": f"$.{candidate}", "type": field.type}
elif field.name == "data" and target_schema_key == "generic_records":
fields[field.name] = {"path": "$", "type": "object"}
elif not field.required:
fields[field.name] = {"path": f"$.{field.name}", "type": field.type, "default": None}
return {
"source": {"items_path": items_path},
"fields": fields,
"meta": {
"generated_by": "heuristic",
"requires_review": True,
},
}
def _flatten_keys(payload: dict[str, Any], prefix: str = "") -> list[str]:
keys: list[str] = []
for key, value in payload.items():
dotted = f"{prefix}.{key}" if prefix else str(key)
keys.append(dotted)
if isinstance(value, dict):
keys.extend(_flatten_keys(value, dotted))
return keys
def _best_field_match(field_name: str, candidates: list[str]) -> str | None:
aliases = {
"lat": ("lat", "latitude", "y"),
"lon": ("lon", "lng", "longitude", "x"),
"mmsi": ("mmsi",),
"sog": ("sog", "speed", "speedOverGround"),
"cog": ("cog", "course", "courseOverGround"),
"received_at": ("received_at", "timestamp", "time", "updated_at"),
"observed_at": ("observed_at", "timestamp", "time", "updated_at"),
"source_id": ("id", "source_id", "uuid"),
}.get(field_name, (field_name,))
lowered = {candidate.lower(): candidate for candidate in candidates}
for alias in aliases:
if alias.lower() in lowered:
return lowered[alias.lower()]
for candidate in candidates:
tail = candidate.split(".")[-1].lower()
if tail in {alias.lower() for alias in aliases}:
return candidate
return None
def _parse_datetime(value: Any) -> datetime | None:
if value is None:
return None
if isinstance(value, datetime):
return value
if isinstance(value, str):
return datetime.fromisoformat(value.replace("Z", "+00:00"))
return None
async def persist_mapped_records(
db: AsyncSession,
*,
datasource_name: str,
datasource_config_id: int,
target_schema: str,
records: list[dict[str, Any]],
mapping_version: int,
) -> int:
"""Persist validated mapped records to the destination for a target schema."""
if target_schema == "vessel_ais":
from app.models.vessel import VesselPosition
for record in records:
db.add(
VesselPosition(
mmsi=record["mmsi"],
lat=record["lat"],
lon=record["lon"],
sog=record.get("sog"),
cog=record.get("cog"),
heading=record.get("heading"),
received_at=_parse_datetime(record.get("received_at")) or datetime.now(UTC),
)
)
await db.commit()
return len(records)
from app.models.collected_data import CollectedData
collected_at = datetime.now(UTC)
for index, record in enumerate(records):
if target_schema == "geo_points":
source_id = record.get("source_id") or f"{datasource_config_id}:{index}"
name = record.get("name")
metadata = {
"latitude": record.get("lat"),
"longitude": record.get("lon"),
"type": record.get("type"),
"mapping_version": mapping_version,
"target_schema": target_schema,
**(record.get("metadata") or {}),
}
reference_date = _parse_datetime(record.get("observed_at"))
else:
source_id = record.get("source_id") or f"{datasource_config_id}:{index}"
name = None
metadata = {
"data": record.get("data") or {},
"mapping_version": mapping_version,
"target_schema": target_schema,
}
reference_date = _parse_datetime(record.get("observed_at"))
db.add(
CollectedData(
source=datasource_name,
source_id=str(source_id),
entity_key=f"{datasource_name}:{source_id}",
data_type=target_schema,
name=name,
title=name,
extra_data=metadata,
collected_at=collected_at,
reference_date=reference_date,
is_valid=1,
is_current=True,
change_type="created",
change_summary={},
)
)
await db.commit()
return len(records)

View File

@@ -30,6 +30,14 @@ class RegionProfile:
accent: str
@dataclass(frozen=True)
class RegionAnchor:
region: str
label: str
latitude: float
longitude: float
@dataclass(frozen=True)
class NewsFeedSource:
id: str
@@ -95,6 +103,39 @@ REGION_PROFILES: dict[str, RegionProfile] = {
),
}
REGION_ANCHORS: dict[str, RegionAnchor] = {
"americas": RegionAnchor(
region="americas",
label="美洲",
latitude=37.0902,
longitude=-95.7129,
),
"europe": RegionAnchor(
region="europe",
label="欧洲",
latitude=50.1109,
longitude=8.6821,
),
"middle-east-africa": RegionAnchor(
region="middle-east-africa",
label="中东与非洲",
latitude=25.2048,
longitude=55.2708,
),
"asia-pacific": RegionAnchor(
region="asia-pacific",
label="亚太",
latitude=1.3521,
longitude=103.8198,
),
"global": RegionAnchor(
region="global",
label="全球",
latitude=20.0,
longitude=0.0,
),
}
def _google_news_feed(query: str, *, hl: str, gl: str, ceid: str) -> str:
return (
@@ -213,6 +254,10 @@ def get_region_profile(region: str) -> RegionProfile:
return REGION_PROFILES.get(region, REGION_PROFILES["global"])
def get_region_anchor(region: str) -> RegionAnchor:
return REGION_ANCHORS.get(region, REGION_ANCHORS["global"])
def get_sources_for_region(region: str) -> list[NewsFeedSource]:
return sorted(
[source for source in NEWS_FEED_SOURCES if source.region in {"global", region}],
@@ -342,6 +387,7 @@ def _serialize_sources(sources: list[NewsFeedSource]) -> list[dict[str, Any]]:
def _serialize_item(item: ParsedNewsItem, *, active_region: str) -> dict[str, Any]:
published_at = item.published_at
anchor = get_region_anchor(item.feed_region)
return {
"id": item.id,
"title": item.title,
@@ -352,6 +398,10 @@ def _serialize_item(item: ParsedNewsItem, *, active_region: str) -> dict[str, An
"region": item.feed_region,
"homepage_url": item.homepage_url,
"published_at": published_at.isoformat().replace("+00:00", "Z") if published_at else None,
"latitude": anchor.latitude,
"longitude": anchor.longitude,
"location_label": anchor.label,
"location_inferred": True,
"is_focus_match": item.feed_region == active_region,
}

View File

@@ -0,0 +1,150 @@
"""LLM provider presets used by Settings and the runtime AI provider bridge."""
from __future__ import annotations
from typing import Any
import httpx
MODELS_DEV_URL = "https://models.dev/api.json"
FALLBACK_LLM_PROVIDER_PRESETS: dict[str, dict[str, Any]] = {
"minimax": {
"provider": "minimax",
"label": "MiniMax",
"provider_api": "anthropic-messages",
"base_url": "https://api.minimaxi.com/anthropic",
"model": "MiniMax-M2.7",
"models": ["MiniMax-M2.7", "MiniMax-M2.7-highspeed", "MiniMax-M2.5", "MiniMax-M2"],
"api_key_env": "MINIMAX_API_KEY",
"source": "fallback",
},
"openai": {
"provider": "openai",
"label": "OpenAI",
"provider_api": "openai-completions",
"base_url": "https://api.openai.com/v1",
"model": "gpt-5.1",
"models": ["gpt-5.1", "gpt-5.1-codex", "gpt-4.1", "gpt-4o"],
"api_key_env": "OPENAI_API_KEY",
"source": "fallback",
},
"anthropic": {
"provider": "anthropic",
"label": "Anthropic",
"provider_api": "anthropic-messages",
"base_url": "https://api.anthropic.com/v1",
"model": "claude-sonnet-4-6",
"models": ["claude-sonnet-4-6", "claude-opus-4-5", "claude-3-5-haiku-20241022"],
"api_key_env": "ANTHROPIC_API_KEY",
"source": "fallback",
},
"deepseek": {
"provider": "deepseek",
"label": "DeepSeek",
"provider_api": "openai-completions",
"base_url": "https://api.deepseek.com/v1",
"model": "deepseek-chat",
"models": ["deepseek-chat", "deepseek-reasoner"],
"api_key_env": "DEEPSEEK_API_KEY",
"source": "fallback",
},
"alibaba": {
"provider": "alibaba",
"label": "Alibaba Qwen / DashScope",
"provider_api": "openai-completions",
"base_url": "https://dashscope.aliyuncs.com/compatible-mode/v1",
"model": "qwen3-max",
"models": ["qwen3-max", "qwen3.5-plus", "qwen-max", "qwen-plus"],
"api_key_env": "DASHSCOPE_API_KEY",
"source": "fallback",
},
"moonshotai": {
"provider": "moonshotai",
"label": "Moonshot AI / Kimi",
"provider_api": "openai-completions",
"base_url": "https://api.moonshot.ai/v1",
"model": "kimi-k2.5",
"models": ["kimi-k2.5", "kimi-k2-thinking", "kimi-k2-turbo-preview"],
"api_key_env": "MOONSHOT_API_KEY",
"source": "fallback",
},
"openrouter": {
"provider": "openrouter",
"label": "OpenRouter",
"provider_api": "openai-completions",
"base_url": "https://openrouter.ai/api/v1",
"model": "openai/gpt-5.1",
"models": ["openai/gpt-5.1", "anthropic/claude-sonnet-4.5", "qwen/qwen3-max"],
"api_key_env": "OPENROUTER_API_KEY",
"source": "fallback",
},
"ollama": {
"provider": "ollama",
"label": "Ollama Local",
"provider_api": "ollama-generate",
"base_url": "http://127.0.0.1:11434",
"model": "qwen2.5:7b",
"models": ["qwen2.5:7b", "llama3.1:8b", "mistral:7b"],
"api_key_env": "",
"source": "fallback",
},
}
MODELS_DEV_PROVIDER_KEYS = {
"minimax": "minimax",
"openai": "openai",
"anthropic": "anthropic",
"deepseek": "deepseek",
"alibaba": "alibaba",
"moonshotai": "moonshotai",
"openrouter": "openrouter",
}
def list_fallback_llm_provider_presets() -> list[dict[str, Any]]:
return [dict(value) for value in FALLBACK_LLM_PROVIDER_PRESETS.values()]
def get_fallback_llm_provider_preset(provider: str) -> dict[str, Any]:
key = provider.strip().lower()
if key not in FALLBACK_LLM_PROVIDER_PRESETS:
raise ValueError(f"Unsupported LLM provider preset: {provider}")
return dict(FALLBACK_LLM_PROVIDER_PRESETS[key])
async def refresh_llm_provider_preset(provider: str) -> dict[str, Any]:
fallback = get_fallback_llm_provider_preset(provider)
models_dev_key = MODELS_DEV_PROVIDER_KEYS.get(fallback["provider"])
if not models_dev_key:
return fallback
async with httpx.AsyncClient(timeout=15.0, follow_redirects=True) as client:
response = await client.get(
MODELS_DEV_URL,
headers={"User-Agent": "Planet/1.0"},
)
response.raise_for_status()
catalog = response.json()
upstream = catalog.get(models_dev_key)
if not isinstance(upstream, dict):
return fallback
upstream_models = upstream.get("models") if isinstance(upstream.get("models"), dict) else {}
model_ids = list(upstream_models.keys())[:80]
base_url = upstream.get("api") or fallback["base_url"]
if fallback["provider"] == "deepseek" and base_url == "https://api.deepseek.com":
base_url = "https://api.deepseek.com/v1"
refreshed = {
**fallback,
"label": upstream.get("name") or fallback["label"],
"base_url": base_url,
"model": model_ids[0] if model_ids else fallback["model"],
"models": model_ids or fallback["models"],
"api_key_env": (upstream.get("env") or [fallback["api_key_env"]])[0],
"source": MODELS_DEV_URL,
}
return refreshed

View File

@@ -0,0 +1,86 @@
from __future__ import annotations
from typing import Any
from app.core.logging import get_logger, sanitize_log_value
from app.core.request_context import get_request_id
from app.db.session import async_session_factory
from app.models.system_log import AuditLog, SystemLog
logger = get_logger(__name__)
async def record_system_log(
*,
source: str,
level: str,
message: str,
service: str | None = None,
module: str | None = None,
event: str | None = None,
request_id: str | None = None,
trace_id: str | None = None,
user_id: int | None = None,
category: str | None = None,
context: dict[str, Any] | None = None,
) -> None:
try:
async with async_session_factory() as session:
session.add(
SystemLog(
source=source,
service=service,
module=module,
event=event,
level=level.lower(),
message=str(sanitize_log_value(message)),
request_id=request_id or get_request_id(),
trace_id=trace_id,
user_id=user_id,
category=category,
context=sanitize_log_value(context or {}),
)
)
await session.commit()
except Exception:
logger.exception_event(
"Failed to persist system log",
event="system_log.persist.failed",
context={"event_name": event, "source": source},
)
async def record_audit_log(
*,
action: str,
actor_id: int | None = None,
actor_name: str | None = None,
target_type: str | None = None,
target_id: str | None = None,
result: str | None = None,
request_id: str | None = None,
ip: str | None = None,
details: dict[str, Any] | None = None,
) -> None:
try:
async with async_session_factory() as session:
session.add(
AuditLog(
actor_id=actor_id,
actor_name=actor_name,
action=action,
target_type=target_type,
target_id=target_id,
result=result,
request_id=request_id or get_request_id(),
ip=ip,
details=sanitize_log_value(details or {}),
)
)
await session.commit()
except Exception:
logger.exception_event(
"Failed to persist audit log",
event="audit_log.persist.failed",
context={"action": action},
)

View File

@@ -1,7 +1,6 @@
"""Task Scheduler for running collection jobs."""
import asyncio
import logging
from datetime import UTC, datetime, timedelta
from typing import Any, Dict, Optional
@@ -9,13 +8,14 @@ from apscheduler.schedulers.asyncio import AsyncIOScheduler
from apscheduler.triggers.interval import IntervalTrigger
from sqlalchemy import select
from app.core.logging import get_logger
from app.db.session import async_session_factory
from app.core.time import to_iso8601_utc
from app.models.datasource import DataSource
from app.models.task import CollectionTask
from app.services.collectors.registry import collector_registry
logger = logging.getLogger(__name__)
logger = get_logger(__name__)
scheduler = AsyncIOScheduler()
RUNNING_TASK_GUARD_TIMEOUT_MINUTES = 90
@@ -54,7 +54,11 @@ async def _update_next_run_at(datasource: DataSource, session) -> None:
async def _apply_datasource_schedule(datasource: DataSource, session) -> None:
collector = collector_registry.get(datasource.source)
if not collector:
logger.warning("Collector not found for datasource %s", datasource.source)
logger.warning_event(
"Collector not found for datasource",
event="collector.schedule.collector_missing",
context={"collector_name": datasource.source},
)
return
collector_registry.set_active(datasource.source, datasource.is_active)
@@ -72,13 +76,17 @@ async def _apply_datasource_schedule(datasource: DataSource, session) -> None:
replace_existing=True,
kwargs={"collector_name": datasource.source},
)
logger.info(
"Scheduled collector: %s (every %sm)",
datasource.source,
datasource.frequency_minutes,
logger.info_event(
"Scheduled collector",
event="collector.schedule.updated",
context={"collector_name": datasource.source, "frequency_minutes": datasource.frequency_minutes},
)
else:
logger.info("Collector disabled: %s", datasource.source)
logger.info_event(
"Collector disabled",
event="collector.schedule.disabled",
context={"collector_name": datasource.source},
)
await _update_next_run_at(datasource, session)
@@ -87,18 +95,30 @@ async def run_collector_task(collector_name: str):
"""Run a single collector task."""
collector = collector_registry.get(collector_name)
if not collector:
logger.error("Collector not found: %s", collector_name)
logger.error_event(
"Collector not found",
event="collector.run.collector_missing",
context={"collector_name": collector_name},
)
return
async with async_session_factory() as db:
result = await db.execute(select(DataSource).where(DataSource.source == collector_name))
datasource = result.scalar_one_or_none()
if not datasource:
logger.error("Datasource not found for collector: %s", collector_name)
logger.error_event(
"Datasource not found for collector",
event="collector.run.datasource_missing",
context={"collector_name": collector_name},
)
return
if not datasource.is_active:
logger.info("Skipping disabled collector: %s", collector_name)
logger.info_event(
"Skipping disabled collector",
event="collector.run.skipped_disabled",
context={"collector_name": collector_name},
)
return
running_result = await db.execute(
@@ -122,10 +142,10 @@ async def run_collector_task(collector_name: str):
and (now - started_at) > timedelta(minutes=RUNNING_TASK_GUARD_TIMEOUT_MINUTES)
)
if not is_stale:
logger.warning(
"Skipping collector %s trigger because task %s is already running",
collector_name,
existing_running.id,
logger.warning_event(
"Skipping collector trigger because task is already running",
event="collector.run.skipped_already_running",
context={"collector_name": collector_name, "task_id": existing_running.id},
)
return
@@ -143,31 +163,47 @@ async def run_collector_task(collector_name: str):
else stale_reason
)
await db.commit()
logger.warning(
"Marked stale running task %s as failed before rerun of %s",
existing_running.id,
collector_name,
logger.warning_event(
"Marked stale running task as failed before rerun",
event="collector.run.stale_task_failed",
context={"collector_name": collector_name, "task_id": existing_running.id},
)
try:
collector._datasource_id = datasource.id
logger.info("Running collector: %s (datasource_id=%s)", collector_name, datasource.id)
logger.info_event(
"Running collector",
event="collector.run.started",
context={"collector_name": collector_name, "datasource_id": datasource.id},
)
task_result = await collector.run(db)
datasource.last_run_at = datetime.now(UTC)
datasource.last_status = task_result.get("status")
await _update_next_run_at(datasource, db)
logger.info("Collector %s completed: %s", collector_name, task_result)
logger.info_event(
"Collector completed",
event="collector.run.completed",
context={"collector_name": collector_name, "datasource_id": datasource.id, "result": task_result},
)
except asyncio.CancelledError:
datasource.last_run_at = datetime.now(UTC)
datasource.last_status = "cancelled"
await db.commit()
logger.warning("Collector %s cancelled by operator", collector_name)
logger.warning_event(
"Collector cancelled by operator",
event="collector.run.cancelled",
context={"collector_name": collector_name, "datasource_id": datasource.id},
)
raise
except Exception as exc:
datasource.last_run_at = datetime.now(UTC)
datasource.last_status = "failed"
await db.commit()
logger.exception("Collector %s failed: %s", collector_name, exc)
logger.exception_event(
"Collector failed",
event="collector.run.failed",
context={"collector_name": collector_name, "datasource_id": datasource.id, "error": str(exc)},
)
async def cleanup_stale_running_tasks(max_age_hours: int = 2) -> int:
@@ -194,7 +230,11 @@ async def cleanup_stale_running_tasks(max_age_hours: int = 2) -> int:
if stale_tasks:
await db.commit()
logger.warning("Cleaned up %s stale running collection task(s)", len(stale_tasks))
logger.warning_event(
"Cleaned up stale running collection tasks",
event="collector.cleanup.stale_tasks_cleaned",
context={"count": len(stale_tasks)},
)
return len(stale_tasks)
@@ -203,14 +243,14 @@ def start_scheduler() -> None:
"""Start the scheduler."""
if not scheduler.running:
scheduler.start()
logger.info("Scheduler started")
logger.info_event("Scheduler started", event="scheduler.started")
def stop_scheduler() -> None:
"""Stop the scheduler."""
if scheduler.running:
scheduler.shutdown(wait=False)
logger.info("Scheduler stopped")
logger.info_event("Scheduler stopped", event="scheduler.stopped")
async def sync_scheduler_with_datasources() -> None:
@@ -271,12 +311,20 @@ def run_collector_now(collector_name: str) -> bool:
"""Run a collector immediately (not scheduled)."""
collector = collector_registry.get(collector_name)
if not collector:
logger.error("Collector not found: %s", collector_name)
logger.error_event(
"Collector not found",
event="collector.trigger.collector_missing",
context={"collector_name": collector_name},
)
return False
existing_task = get_running_collector_task(collector_name)
if existing_task is not None and not existing_task.done():
logger.warning("Collector %s is already running in-memory; skipping duplicate trigger", collector_name)
logger.warning_event(
"Collector is already running in-memory; skipping duplicate trigger",
event="collector.trigger.skipped_already_running",
context={"collector_name": collector_name},
)
return False
try:
@@ -289,10 +337,18 @@ def run_collector_now(collector_name: str) -> bool:
RUNNING_COLLECTOR_TASKS.pop(collector_name, None)
task.add_done_callback(_cleanup_task)
logger.info("Triggered collector: %s", collector_name)
logger.info_event(
"Triggered collector",
event="collector.trigger.started",
context={"collector_name": collector_name},
)
return True
except Exception as exc:
logger.error("Failed to trigger collector %s: %s", collector_name, exc)
logger.error_event(
"Failed to trigger collector",
event="collector.trigger.failed",
context={"collector_name": collector_name, "error": str(exc)},
)
return False

View File

@@ -0,0 +1,532 @@
from __future__ import annotations
import json
import re
import shutil
import subprocess
from collections import Counter, deque
from dataclasses import dataclass
from datetime import UTC, datetime
from pathlib import Path
from typing import Any
from app.core.security import redis_client
DEFAULT_LOG_LINE_LIMIT = 200
MAX_LOG_LINE_LIMIT = 1000
BUFFER_LOG_LIMIT = 1000
BUFFER_LOG_TTL_SECONDS = 7 * 24 * 60 * 60
LOG_BUFFER_KEY_PREFIX = "planet:system_logs"
LOG_LEVEL_ERROR = "error"
LOG_LEVEL_WARNING = "warning"
LOG_LEVEL_INFO = "info"
LOG_LEVEL_DEBUG = "debug"
LOG_LEVEL_ALL = "all"
SUPPORTED_LOG_LEVELS = {
LOG_LEVEL_ALL,
LOG_LEVEL_ERROR,
LOG_LEVEL_WARNING,
LOG_LEVEL_INFO,
LOG_LEVEL_DEBUG,
}
LOG_LEVEL_ALIASES = {
"warn": LOG_LEVEL_WARNING,
"warning": LOG_LEVEL_WARNING,
"err": LOG_LEVEL_ERROR,
"error": LOG_LEVEL_ERROR,
"info": LOG_LEVEL_INFO,
"information": LOG_LEVEL_INFO,
"debug": LOG_LEVEL_DEBUG,
"trace": LOG_LEVEL_DEBUG,
"critical": LOG_LEVEL_ERROR,
"fatal": LOG_LEVEL_ERROR,
}
TIMESTAMP_FORMATS = (
"%Y-%m-%d %H:%M:%S.%f",
"%Y-%m-%d %H:%M:%S",
"%Y-%m-%dT%H:%M:%S.%f",
"%Y-%m-%dT%H:%M:%S",
)
LEVEL_PATTERNS = (
("CRITICAL", LOG_LEVEL_ERROR),
("FATAL", LOG_LEVEL_ERROR),
("ERROR", LOG_LEVEL_ERROR),
("WARNING", LOG_LEVEL_WARNING),
("WARN", LOG_LEVEL_WARNING),
("INFO", LOG_LEVEL_INFO),
("DEBUG", LOG_LEVEL_DEBUG),
("TRACE", LOG_LEVEL_DEBUG),
)
LEADING_LEVEL_PATTERN = re.compile(
r"^\s*(?:\[[^\]]+\]\s*)?(CRITICAL|FATAL|ERROR|WARNING|WARN|INFO|DEBUG|TRACE)\b[:\s-]*",
re.IGNORECASE,
)
EMBEDDED_LEVEL_PATTERN = re.compile(
r"\b(CRITICAL|FATAL|ERROR|WARNING|WARN|INFO|DEBUG|TRACE)\b",
re.IGNORECASE,
)
CONTROL_CHAR_PATTERN = re.compile(r"[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]")
@dataclass(frozen=True)
class LogSource:
source_id: str
name: str
kind: str
location: str
description: str
category: str
status: str = "ok"
buffer_key: str | None = None
container_name: str | None = None
@dataclass
class StructuredLogEntry:
timestamp: datetime | None
level: str | None
display_line: str
raw_line: str
search_text: str
@dataclass
class DailyLogMarker:
date_token: str
total: int
dominant_level: str
LOG_SOURCES: dict[str, LogSource] = {
"backend": LogSource(
source_id="backend",
name="后端服务",
kind="file",
location="/tmp/planet_backend.log",
description="FastAPI 后端、调度器和采集任务共享日志。",
category="service",
),
"frontend": LogSource(
source_id="frontend",
name="前端开发服务",
kind="file",
location="/tmp/planet_frontend.log",
description="控制台与 Earth 前端开发服务输出。",
category="service",
),
"ai-provider": LogSource(
source_id="ai-provider",
name="AI Provider",
kind="docker",
location="docker://planet_aiprovider",
description="AI Provider 容器实时输出日志。",
category="service",
container_name="planet_aiprovider",
),
"earth-client": LogSource(
source_id="earth-client",
name="Earth 浏览器端",
kind="buffer",
location="redis://planet:system_logs:earth-client",
description="Earth 浏览器端上报的运行时错误与关键业务日志。",
category="client",
buffer_key=f"{LOG_BUFFER_KEY_PREFIX}:earth-client",
),
}
def normalize_log_level(level: str | None) -> str:
if level is None:
return LOG_LEVEL_ALL
normalized = str(level).strip().lower()
if normalized in {"", LOG_LEVEL_ALL}:
return LOG_LEVEL_ALL
return LOG_LEVEL_ALIASES.get(normalized, LOG_LEVEL_ALL)
def normalize_log_levels(level: str | None = None, levels: str | None = None) -> tuple[str, ...]:
normalized_levels: list[str] = []
if levels:
for item in str(levels).split(","):
normalized = normalize_log_level(item)
if normalized != LOG_LEVEL_ALL and normalized not in normalized_levels:
normalized_levels.append(normalized)
normalized_level = normalize_log_level(level)
if normalized_level != LOG_LEVEL_ALL and normalized_level not in normalized_levels:
normalized_levels.append(normalized_level)
return tuple(normalized_levels)
def get_source_status(source: LogSource) -> str:
if source.kind == "file":
path = Path(source.location)
if not path.exists():
return "missing"
return "ok" if path.stat().st_size > 0 else "empty"
if source.kind == "docker":
return "ok" if shutil.which("docker") else "docker_unavailable"
if source.kind == "buffer":
if not source.buffer_key:
return "source_unavailable"
try:
return "ok" if redis_client.llen(source.buffer_key) > 0 else "empty"
except Exception:
return "source_unavailable"
return "source_unavailable"
def list_log_sources() -> list[dict[str, str]]:
items: list[dict[str, str]] = []
for source in LOG_SOURCES.values():
items.append(
{
"source_id": source.source_id,
"name": source.name,
"kind": source.kind,
"location": source.location,
"description": source.description,
"category": source.category,
"status": get_source_status(source),
}
)
return items
def get_buffer_log_key(source_id: str) -> str:
return f"{LOG_BUFFER_KEY_PREFIX}:{source_id}"
def append_buffer_log(
source_id: str,
*,
level: str,
message: str,
context: dict[str, Any] | None = None,
) -> None:
payload = {
"timestamp": datetime.now(tz=UTC).isoformat(),
"level": normalize_log_level(level),
"message": message,
"context": context or {},
}
buffer_key = get_buffer_log_key(source_id)
redis_client.rpush(buffer_key, json.dumps(payload, ensure_ascii=False))
redis_client.ltrim(buffer_key, -BUFFER_LOG_LIMIT, -1)
redis_client.expire(buffer_key, BUFFER_LOG_TTL_SECONDS)
def parse_timestamp(raw_value: str | None) -> datetime | None:
if not raw_value:
return None
candidate = str(raw_value).strip()
if not candidate:
return None
candidate = candidate.replace("Z", "+00:00")
try:
parsed = datetime.fromisoformat(candidate)
return parsed if parsed.tzinfo else parsed.replace(tzinfo=UTC)
except ValueError:
pass
for fmt in TIMESTAMP_FORMATS:
try:
return datetime.strptime(candidate, fmt).replace(tzinfo=UTC)
except ValueError:
continue
return None
def parse_prefixed_timestamp(line: str) -> tuple[datetime | None, str]:
stripped = line.strip()
if not stripped:
return None, ""
for prefix_length in (35, 32, 29, 26, 23, 19):
if len(stripped) < prefix_length:
continue
prefix = stripped[:prefix_length]
timestamp = parse_timestamp(prefix)
if timestamp is not None:
return timestamp, stripped[prefix_length:].lstrip()
first_token = stripped.split(maxsplit=1)[0]
timestamp = parse_timestamp(first_token)
if timestamp is not None:
remainder = stripped[len(first_token):].lstrip()
return timestamp, remainder
return None, stripped
def infer_log_level_from_text(text: str, *, allow_embedded: bool = True) -> str | None:
leading_match = LEADING_LEVEL_PATTERN.match(text)
if leading_match:
return normalize_log_level(leading_match.group(1))
if allow_embedded:
embedded_match = EMBEDDED_LEVEL_PATTERN.search(text)
if embedded_match:
return normalize_log_level(embedded_match.group(1))
upper_text = text.upper()
for pattern, normalized in LEVEL_PATTERNS:
if f"{pattern}:" in upper_text or f"{pattern} " in upper_text:
return normalized
return None
def build_display_line(timestamp: datetime | None, level: str | None, message: str) -> str:
message_part = message.strip() if message else ""
parts = []
if timestamp is not None:
parts.append(timestamp.astimezone(UTC).strftime("%Y-%m-%d %H:%M:%S"))
if level:
parts.append(level.upper())
if message_part:
parts.append(message_part)
return " ".join(parts).strip()
def sanitize_text_log_line(line: str) -> str:
return CONTROL_CHAR_PATTERN.sub("", line)
def parse_text_log_entry(line: str) -> StructuredLogEntry:
sanitized_line = sanitize_text_log_line(line).rstrip("\n")
timestamp, remainder = parse_prefixed_timestamp(sanitized_line)
level = infer_log_level_from_text(remainder or sanitized_line, allow_embedded=False)
display_line = sanitized_line
return StructuredLogEntry(
timestamp=timestamp,
level=level,
display_line=display_line,
raw_line=display_line,
search_text=display_line.lower(),
)
def build_buffer_entry(payload: dict[str, Any]) -> StructuredLogEntry:
timestamp = parse_timestamp(str(payload.get("timestamp", "")).strip())
level = normalize_log_level(payload.get("level"))
if level == LOG_LEVEL_ALL:
level = None
message = str(payload.get("message", "")).strip()
context = payload.get("context")
context_map = context if isinstance(context, dict) else {}
context_fragments = []
for key in ("category", "module", "url", "detail"):
value = str(context_map.get(key, "")).strip()
if value:
context_fragments.append(f"{key}={value}")
message_with_context = " | ".join([message, *context_fragments]) if context_fragments else message
display_line = build_display_line(timestamp, level, message_with_context)
search_text = " ".join(
[
message,
json.dumps(context_map, ensure_ascii=False, sort_keys=True),
display_line,
]
).lower()
return StructuredLogEntry(
timestamp=timestamp,
level=level,
display_line=display_line,
raw_line=json.dumps(payload, ensure_ascii=False, sort_keys=True),
search_text=search_text,
)
def read_file_entries(source: LogSource, scan_limit: int) -> list[StructuredLogEntry]:
path = Path(source.location)
if not path.exists():
return []
with path.open("r", encoding="utf-8", errors="replace") as handle:
recent_lines = deque(handle, maxlen=scan_limit)
return [
parse_text_log_entry(line)
for line in recent_lines
if sanitize_text_log_line(line).strip()
]
def read_docker_entries(source: LogSource, scan_limit: int) -> list[StructuredLogEntry]:
if not shutil.which("docker") or not source.container_name:
return []
try:
completed = subprocess.run(
[
"docker",
"logs",
"--timestamps",
"--tail",
str(scan_limit),
source.container_name,
],
capture_output=True,
text=True,
check=False,
)
except OSError:
return []
if completed.returncode != 0:
return []
return [
parse_text_log_entry(line)
for line in completed.stdout.splitlines()
if line.strip()
]
def read_buffer_entries(source: LogSource, scan_limit: int) -> list[StructuredLogEntry]:
if not source.buffer_key:
return []
try:
raw_items = redis_client.lrange(source.buffer_key, -scan_limit, -1)
except Exception:
return []
entries: list[StructuredLogEntry] = []
for raw_item in raw_items:
try:
payload = json.loads(raw_item)
except json.JSONDecodeError:
entries.append(parse_text_log_entry(str(raw_item)))
continue
if isinstance(payload, dict):
entries.append(build_buffer_entry(payload))
else:
entries.append(parse_text_log_entry(str(raw_item)))
return entries
def read_source_entries(source: LogSource, scan_limit: int) -> list[StructuredLogEntry]:
if source.kind == "file":
return read_file_entries(source, scan_limit)
if source.kind == "docker":
return read_docker_entries(source, scan_limit)
if source.kind == "buffer":
return read_buffer_entries(source, scan_limit)
return []
def matches_levels(entry: StructuredLogEntry, selected_levels: tuple[str, ...]) -> bool:
if not selected_levels:
return True
return entry.level in selected_levels
def matches_date_range(
entry: StructuredLogEntry,
start_date: str | None,
end_date: str | None,
) -> bool:
if not start_date and not end_date:
return True
if entry.timestamp is None:
return False
date_token = entry.timestamp.astimezone(UTC).date().isoformat()
if start_date and date_token < start_date:
return False
if end_date and date_token > end_date:
return False
return True
def matches_search(entry: StructuredLogEntry, search: str | None) -> bool:
if search is None:
return True
query = search.strip().lower()
if not query:
return True
return query in entry.search_text
def build_daily_log_markers(entries: list[StructuredLogEntry]) -> list[dict[str, Any]]:
grouped: dict[str, list[StructuredLogEntry]] = {}
for entry in entries:
if entry.timestamp is None:
continue
date_token = entry.timestamp.astimezone(UTC).date().isoformat()
grouped.setdefault(date_token, []).append(entry)
markers: list[DailyLogMarker] = []
for date_token, group in sorted(grouped.items()):
level_counts = Counter(
entry.level
for entry in group
if entry.level in SUPPORTED_LOG_LEVELS and entry.level != LOG_LEVEL_ALL
)
dominant_level = LOG_LEVEL_INFO
if level_counts:
dominant_level = sorted(
level_counts.items(),
key=lambda item: (
-item[1],
("error", "warning", "info", "debug").index(item[0]),
),
)[0][0]
markers.append(
DailyLogMarker(
date_token=date_token,
total=len(group),
dominant_level=dominant_level,
)
)
return [marker.__dict__ for marker in markers]
def read_log_snapshot(
source_id: str,
limit: int,
*,
level: str = LOG_LEVEL_ALL,
levels: str | None = None,
start_date: str | None = None,
end_date: str | None = None,
search: str | None = None,
) -> dict[str, Any] | None:
source = LOG_SOURCES.get(source_id)
if source is None:
return None
selected_levels = normalize_log_levels(level, levels)
search_query = (search or "").strip()
scan_limit = max(min(MAX_LOG_LINE_LIMIT * 5, 5000), limit * 5, BUFFER_LOG_LIMIT if source.kind == "buffer" else 1000)
all_entries = read_source_entries(source, scan_limit)
marker_entries = [
entry
for entry in all_entries
if matches_levels(entry, selected_levels) and matches_search(entry, search_query)
]
filtered_entries = [
entry
for entry in marker_entries
if matches_date_range(entry, start_date, end_date)
]
visible_entries = filtered_entries[-limit:]
compatibility_level = selected_levels[0] if len(selected_levels) == 1 else LOG_LEVEL_ALL
return {
"source_id": source.source_id,
"name": source.name,
"kind": source.kind,
"location": source.location,
"description": source.description,
"category": source.category,
"status": get_source_status(source),
"level": compatibility_level,
"selected_levels": list(selected_levels),
"search_query": search_query,
"available_levels": [
LOG_LEVEL_ALL,
LOG_LEVEL_ERROR,
LOG_LEVEL_WARNING,
LOG_LEVEL_INFO,
LOG_LEVEL_DEBUG,
],
"daily_markers": build_daily_log_markers(marker_entries),
"line_limit": limit,
"line_count": len(visible_entries),
"lines": [entry.display_line for entry in visible_entries],
}

View File

@@ -17,7 +17,7 @@ TV_LIVE_SOURCE_COLLECTOR = "news_live_streams"
TV_LIVE_SOURCE_DATA_TYPE = "news_live_stream"
DEFAULT_TV_SETTINGS = {
"default_source_id": DEFAULT_TV_SOURCE_ID,
"default_source_id": DEFAULT_TV_SOURCE_ID,
"auto_fallback": True,
"sources": [
{
@@ -362,7 +362,7 @@ def _build_collected_tv_source(record: CollectedData, index: int) -> dict[str, A
"sort_order": metadata.get("sort_order", 200 + index),
"collector_source": record.source,
"notes": record.description or metadata.get("notes") or "",
"updated_at": to_iso8601_utc(record.updated_at or record.reference_date or datetime.now(UTC)),
"updated_at": to_iso8601_utc(record.collected_at or record.reference_date or datetime.now(UTC)),
},
index=index,
)

View File

@@ -35,6 +35,7 @@ async def test_health_check():
data = response.json()
assert data["status"] == "healthy"
assert "version" in data
assert response.headers["x-request-id"]
@pytest.mark.asyncio
@@ -161,6 +162,345 @@ async def test_alerts_endpoint_with_auth(auth_headers):
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_system_log_sources_requires_super_admin(auth_headers):
def override_get_current_user():
return User(
id=1,
username="testuser",
email="test@example.com",
password_hash="hashed",
role="admin",
is_active=True,
)
app.dependency_overrides = {
__import__("app.core.security", fromlist=["get_current_user"]).get_current_user: override_get_current_user,
}
transport = ASGITransport(app=app)
try:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get("/api/v1/system/logs/sources", headers=auth_headers)
assert response.status_code == 403
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_system_log_sources_with_super_admin(auth_headers):
def override_get_current_user():
return User(
id=1,
username="root",
email="root@example.com",
password_hash="hashed",
role="super_admin",
is_active=True,
)
app.dependency_overrides = {
__import__("app.core.security", fromlist=["get_current_user"]).get_current_user: override_get_current_user,
}
transport = ASGITransport(app=app)
try:
with patch(
"app.api.v1.system_control.list_log_sources",
return_value=[
{
"source_id": "backend",
"name": "后端服务",
"kind": "file",
"location": "/tmp/planet_backend.log",
"description": "FastAPI 后端、调度器和采集任务共享日志。",
"category": "service",
"status": "ok",
}
],
):
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get("/api/v1/system/logs/sources", headers=auth_headers)
assert response.status_code == 200
data = response.json()
assert data["items"][0]["source_id"] == "backend"
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_system_log_snapshot_with_super_admin(auth_headers):
def override_get_current_user():
return User(
id=1,
username="root",
email="root@example.com",
password_hash="hashed",
role="super_admin",
is_active=True,
)
app.dependency_overrides = {
__import__("app.core.security", fromlist=["get_current_user"]).get_current_user: override_get_current_user,
}
transport = ASGITransport(app=app)
try:
with patch(
"app.api.v1.system_control.read_log_snapshot",
return_value={
"source_id": "backend",
"name": "后端服务",
"kind": "file",
"location": "/tmp/planet_backend.log",
"description": "FastAPI 后端、调度器和采集任务共享日志。",
"category": "service",
"status": "ok",
"level": "all",
"selected_levels": [],
"search_query": "",
"available_levels": ["all", "error", "warning", "info", "debug"],
"daily_markers": [],
"line_limit": 50,
"line_count": 2,
"lines": ["line 1", "line 2"],
},
):
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get("/api/v1/system/logs/backend?limit=50", headers=auth_headers)
assert response.status_code == 200
data = response.json()
assert data["source_id"] == "backend"
assert data["line_count"] == 2
assert data["lines"] == ["line 1", "line 2"]
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_system_log_snapshot_supports_level_filter(auth_headers):
def override_get_current_user():
return User(
id=1,
username="root",
email="root@example.com",
password_hash="hashed",
role="super_admin",
is_active=True,
)
app.dependency_overrides = {
__import__("app.core.security", fromlist=["get_current_user"]).get_current_user: override_get_current_user,
}
transport = ASGITransport(app=app)
try:
with patch(
"app.api.v1.system_control.read_log_snapshot",
return_value={
"source_id": "backend",
"name": "后端服务",
"kind": "file",
"location": "/tmp/planet_backend.log",
"description": "FastAPI 后端、调度器和采集任务共享日志。",
"category": "service",
"status": "ok",
"level": "error",
"selected_levels": ["error"],
"search_query": "",
"available_levels": ["all", "error", "warning", "info", "debug"],
"daily_markers": [],
"line_limit": 50,
"line_count": 1,
"lines": ["ERROR: failed"],
},
):
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get("/api/v1/system/logs/backend?limit=50&level=error", headers=auth_headers)
assert response.status_code == 200
data = response.json()
assert data["level"] == "error"
assert data["lines"] == ["ERROR: failed"]
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_system_log_snapshot_supports_date_range_filter(auth_headers):
def override_get_current_user():
return User(
id=1,
username="root",
email="root@example.com",
password_hash="hashed",
role="super_admin",
is_active=True,
)
app.dependency_overrides = {
__import__("app.core.security", fromlist=["get_current_user"]).get_current_user: override_get_current_user,
}
transport = ASGITransport(app=app)
try:
with patch(
"app.api.v1.system_control.read_log_snapshot",
return_value={
"source_id": "backend",
"name": "后端服务",
"kind": "file",
"location": "/tmp/planet_backend.log",
"description": "FastAPI 后端、调度器和采集任务共享日志。",
"category": "service",
"status": "ok",
"level": "all",
"selected_levels": [],
"search_query": "",
"available_levels": ["all", "error", "warning", "info", "debug"],
"daily_markers": [],
"line_limit": 50,
"line_count": 1,
"lines": ["2026-04-23 INFO: service started"],
},
) as mock_read_log_snapshot:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get(
"/api/v1/system/logs/backend?limit=50&start_date=2026-04-20&end_date=2026-04-23",
headers=auth_headers,
)
assert response.status_code == 200
mock_read_log_snapshot.assert_called_once_with(
"backend",
50,
level="all",
levels=None,
start_date="2026-04-20",
end_date="2026-04-23",
search=None,
)
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_system_log_snapshot_supports_levels_and_search_filter(auth_headers):
def override_get_current_user():
return User(
id=1,
username="root",
email="root@example.com",
password_hash="hashed",
role="super_admin",
is_active=True,
)
app.dependency_overrides = {
__import__("app.core.security", fromlist=["get_current_user"]).get_current_user: override_get_current_user,
}
transport = ASGITransport(app=app)
try:
with patch(
"app.api.v1.system_control.read_log_snapshot",
return_value={
"source_id": "backend",
"name": "后端服务",
"kind": "file",
"location": "/tmp/planet_backend.log",
"description": "FastAPI 后端、调度器和采集任务共享日志。",
"category": "service",
"status": "ok",
"level": "all",
"selected_levels": ["error", "warning"],
"search_query": "timeout",
"available_levels": ["all", "error", "warning", "info", "debug"],
"daily_markers": [],
"line_limit": 50,
"line_count": 1,
"lines": ["2026-04-23 10:00:00 ERROR timeout"],
},
) as mock_read_log_snapshot:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get(
"/api/v1/system/logs/backend?limit=50&levels=error,warning&search=timeout",
headers=auth_headers,
)
assert response.status_code == 200
mock_read_log_snapshot.assert_called_once_with(
"backend",
50,
level="all",
levels="error,warning",
start_date=None,
end_date=None,
search="timeout",
)
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_system_log_snapshot_rejects_invalid_date_range(auth_headers):
def override_get_current_user():
return User(
id=1,
username="root",
email="root@example.com",
password_hash="hashed",
role="super_admin",
is_active=True,
)
app.dependency_overrides = {
__import__("app.core.security", fromlist=["get_current_user"]).get_current_user: override_get_current_user,
}
transport = ASGITransport(app=app)
try:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get(
"/api/v1/system/logs/backend?start_date=2026-04-31",
headers=auth_headers,
)
assert response.status_code == 400
assert "start_date must be in YYYY-MM-DD format" in response.json()["detail"]
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_ingest_earth_client_log_accepts_public_events():
transport = ASGITransport(app=app)
try:
with patch("app.api.v1.system_control.append_buffer_log") as mock_append_buffer_log:
with patch("app.api.v1.system_control.record_system_log", new_callable=AsyncMock) as mock_record_system_log:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.post(
"/api/v1/system/logs/earth-client",
json={
"level": "error",
"message": "登陆点加载失败: 登陆点接口返回 HTTP 500",
"category": "startup-load",
"module": "layer-startup",
},
)
assert response.status_code == 200
data = response.json()
assert data["accepted"] is True
assert data["source_id"] == "earth-client"
mock_append_buffer_log.assert_called_once()
mock_record_system_log.assert_awaited_once()
persisted_kwargs = mock_record_system_log.await_args.kwargs
assert persisted_kwargs["source"] == "earth-client"
assert persisted_kwargs["event"] == "earth.client.runtime_log"
assert persisted_kwargs["category"] == "startup-load"
assert persisted_kwargs["level"] == "error"
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_request_id_header_is_echoed_when_provided():
transport = ASGITransport(app=app)
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get("/health", headers={"X-Request-ID": "planet-test-request"})
assert response.status_code == 200
assert response.headers["x-request-id"] == "planet-test-request"
@pytest.mark.asyncio
async def test_invalid_token():
"""Test that invalid token is rejected"""
@@ -263,6 +603,8 @@ async def test_ai_situational_analysis_returns_503_when_disabled(auth_headers):
assert "content_blocks" in data
assert "text_blocks" in data
assert "thinking_blocks" in data
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
@@ -382,8 +724,6 @@ async def test_save_playground_session_with_auth(auth_headers):
assert data["state"]["objective"] == "测试目标"
finally:
app.dependency_overrides.clear()
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio

View File

@@ -0,0 +1,199 @@
from types import SimpleNamespace
import pytest
from httpx import ASGITransport, AsyncClient
from app.api.v1.datasource_config import get_ai_provider_client
from app.core.security import get_current_user
from app.core.target_schema_registry import get_target_schema, list_target_schemas
from app.main import app
from app.models.user import User
from app.services.datasource_mapping import execute_mapping, persist_mapped_records, redact_for_llm
SAMPLE_AIS = {
"data": [
{
"mmsi": "257123000",
"latitude": "59.91",
"longitude": "10.75",
"speedOverGround": "12.4",
"timestamp": "2026-04-28T00:00:00Z",
"api_token": "secret-value",
}
]
}
def test_registry_exposes_v1_target_schemas():
keys = {schema["key"] for schema in list_target_schemas()}
assert {"vessel_ais", "geo_points", "generic_records"}.issubset(keys)
assert get_target_schema("vessel_ais").destination == "vessel_position"
def test_mapping_engine_maps_and_validates_vessel_ais():
mapping = {
"source": {"items_path": "$.data[*]"},
"fields": {
"mmsi": {"path": "$.mmsi", "type": "integer"},
"lat": {"path": "$.latitude", "type": "float"},
"lon": {"path": "$.longitude", "type": "float"},
"sog": {"path": "$.speedOverGround", "type": "float"},
"received_at": {"path": "$.timestamp", "type": "datetime"},
},
}
result = execute_mapping(SAMPLE_AIS, mapping, "vessel_ais")
assert result["mapped_count"] == 1
assert result["failed_count"] == 0
assert result["records"][0]["mmsi"] == 257123000
assert result["records"][0]["lat"] == 59.91
def test_mapping_engine_reports_schema_errors():
mapping = {
"source": {"items_path": "$.data[*]"},
"fields": {
"mmsi": {"path": "$.mmsi", "type": "integer"},
"lat": {"path": "$.missing_lat", "type": "float"},
"lon": {"path": "$.longitude", "type": "float"},
},
}
result = execute_mapping(SAMPLE_AIS, mapping, "vessel_ais")
assert result["mapped_count"] == 0
assert result["failed_count"] == 1
assert any("lat" in error for error in result["errors"][0]["errors"])
def test_redact_for_llm_masks_secret_like_fields():
redacted = redact_for_llm(SAMPLE_AIS)
assert redacted["data"][0]["api_token"] == "[REDACTED]"
@pytest.mark.asyncio
async def test_persist_mapped_records_writes_generic_records():
class FakeDB:
def __init__(self):
self.added = []
self.committed = False
def add(self, value):
self.added.append(value)
async def commit(self):
self.committed = True
db = FakeDB()
count = await persist_mapped_records(
db,
datasource_name="custom_weather",
datasource_config_id=42,
target_schema="generic_records",
records=[{"source_id": "row-1", "data": {"temp": 25}}],
mapping_version=3,
)
assert count == 1
assert db.committed is True
assert db.added[0].source == "custom_weather"
assert db.added[0].data_type == "generic_records"
assert db.added[0].extra_data["mapping_version"] == 3
@pytest.mark.asyncio
async def test_mapping_preview_api_uses_deterministic_engine():
def override_get_current_user():
return User(
id=1,
username="testuser",
email="test@example.com",
password_hash="hashed",
role="admin",
is_active=True,
)
app.dependency_overrides = {get_current_user: override_get_current_user}
transport = ASGITransport(app=app)
try:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.post(
"/api/v1/datasources/mappings/preview",
json={
"sample_payload": SAMPLE_AIS,
"target_schema": "vessel_ais",
"mapping_json": {
"source": {"items_path": "$.data[*]"},
"fields": {
"mmsi": {"path": "$.mmsi", "type": "integer"},
"lat": {"path": "$.latitude", "type": "float"},
"lon": {"path": "$.longitude", "type": "float"},
},
},
},
)
finally:
app.dependency_overrides.clear()
assert response.status_code == 200
payload = response.json()
assert payload["success"] is True
assert payload["preview"]["records"][0]["mmsi"] == 257123000
@pytest.mark.asyncio
async def test_mapping_propose_api_redacts_sample_before_ai():
seen_context = {}
class FakeAIClient:
async def analyze(self, request, request_id=None):
seen_context.update(request.context)
return SimpleNamespace(
content=(
'{"source":{"items_path":"$.data[*]"},"fields":{'
'"mmsi":{"path":"$.mmsi","type":"integer"},'
'"lat":{"path":"$.latitude","type":"float"},'
'"lon":{"path":"$.longitude","type":"float"}}}'
)
)
def override_get_current_user():
return User(
id=1,
username="testuser",
email="test@example.com",
password_hash="hashed",
role="admin",
is_active=True,
)
def override_ai_client():
return FakeAIClient()
app.dependency_overrides = {
get_current_user: override_get_current_user,
get_ai_provider_client: override_ai_client,
}
transport = ASGITransport(app=app)
try:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.post(
"/api/v1/datasources/mappings/propose",
json={
"sample_payload": SAMPLE_AIS,
"target_schema": "vessel_ais",
"use_ai": True,
},
)
finally:
app.dependency_overrides.clear()
assert response.status_code == 200
payload = response.json()
assert payload["mapping_json"]["meta"]["generated_by"] == "ai_provider"
assert seen_context["sample_payload"]["data"][0]["api_token"] == "[REDACTED]"

View File

@@ -0,0 +1,49 @@
from datetime import UTC, datetime
from app.services.earth_news import ParsedNewsItem, _serialize_item
def test_serialize_item_includes_region_anchor_for_cruise():
item = ParsedNewsItem(
id="google-apac:test",
title="Example APAC story",
summary="Example summary",
url="https://example.com/story",
source="Example Source",
feed_name="Global Monitor / APAC",
feed_region="asia-pacific",
homepage_url="https://example.com",
published_at=datetime(2026, 4, 23, 2, 30, tzinfo=UTC),
)
payload = _serialize_item(item, active_region="asia-pacific")
assert payload["latitude"] == 1.3521
assert payload["longitude"] == 103.8198
assert payload["location_label"] == "亚太"
assert payload["location_inferred"] is True
assert payload["is_focus_match"] is True
assert payload["published_at"] == "2026-04-23T02:30:00Z"
def test_serialize_item_falls_back_to_global_anchor():
item = ParsedNewsItem(
id="custom:test",
title="Fallback story",
summary="Fallback summary",
url="https://example.com/fallback",
source="Fallback Source",
feed_name="Fallback Feed",
feed_region="unknown-region",
homepage_url="https://example.com",
published_at=None,
)
payload = _serialize_item(item, active_region="americas")
assert payload["latitude"] == 20.0
assert payload["longitude"] == 0.0
assert payload["location_label"] == "全球"
assert payload["location_inferred"] is True
assert payload["is_focus_match"] is False
assert payload["published_at"] is None

View File

@@ -0,0 +1,78 @@
from __future__ import annotations
import logging
from io import StringIO
from app.core.logging import PlanetContextFilter, PlanetFormatter, get_logger
from app.core.request_context import set_request_id
def _capture_output(callback):
stream = StringIO()
handler = logging.StreamHandler(stream)
handler.setFormatter(PlanetFormatter(datefmt="%Y-%m-%d %H:%M:%S"))
handler.addFilter(PlanetContextFilter())
adapter = get_logger("tests.logging")
target_logger = adapter.logger
original_handlers = list(target_logger.handlers)
original_level = target_logger.level
original_propagate = target_logger.propagate
target_logger.handlers = [handler]
target_logger.setLevel(logging.INFO)
target_logger.propagate = False
try:
callback(adapter)
finally:
handler.flush()
target_logger.handlers = original_handlers
target_logger.setLevel(original_level)
target_logger.propagate = original_propagate
return stream.getvalue()
def test_structured_logger_injects_request_id_and_event():
set_request_id("req-test-123")
try:
output = _capture_output(
lambda logger: logger.info_event(
"collector started",
event="collector.run.started",
context={"collector_name": "bgp_news"},
)
)
finally:
set_request_id(None)
assert "request_id=req-test-123" in output
assert "event=collector.run.started" in output
assert "service=backend" in output
assert '"collector_name": "bgp_news"' in output
def test_structured_logger_redacts_sensitive_text_and_context():
set_request_id("req-test-redact")
try:
output = _capture_output(
lambda logger: logger.error_event(
"Authorization: Bearer super-secret-token",
event="auth.token.failed",
context={
"token": "plain-secret",
"nested": {"password": "hunter2"},
"safe": "visible",
},
)
)
finally:
set_request_id(None)
assert "super-secret-token" not in output
assert "plain-secret" not in output
assert "hunter2" not in output
assert "[REDACTED]" in output
assert '"safe": "visible"' in output

View File

@@ -0,0 +1,217 @@
from __future__ import annotations
import json
from pathlib import Path
from app.services import system_logs
class FakeRedis:
def __init__(self) -> None:
self.store: dict[str, list[str]] = {}
def rpush(self, key: str, value: str) -> None:
self.store.setdefault(key, []).append(value)
def ltrim(self, key: str, start: int, end: int) -> None:
items = self.store.get(key, [])
normalized_end = None if end == -1 else end + 1
self.store[key] = items[start:normalized_end]
def expire(self, key: str, seconds: int) -> None:
return None
def lrange(self, key: str, start: int, end: int) -> list[str]:
items = self.store.get(key, [])
normalized_end = None if end == -1 else end + 1
return items[start:normalized_end]
def llen(self, key: str) -> int:
return len(self.store.get(key, []))
def test_read_log_snapshot_uses_structured_buffer_timestamp_level_and_search(monkeypatch):
fake_redis = FakeRedis()
monkeypatch.setattr(system_logs, "redis_client", fake_redis)
monkeypatch.setattr(
system_logs,
"LOG_SOURCES",
{
"earth-client": system_logs.LogSource(
source_id="earth-client",
name="Earth 浏览器端",
kind="buffer",
location="redis://planet:system_logs:earth-client",
description="Earth 浏览器端上报日志",
category="client",
buffer_key=system_logs.get_buffer_log_key("earth-client"),
)
},
)
fake_redis.rpush(
system_logs.get_buffer_log_key("earth-client"),
json.dumps(
{
"timestamp": "2026-04-22T10:15:30Z",
"level": "warning",
"message": "news feed degraded",
"context": {"module": "news", "detail": "timeout"},
},
ensure_ascii=False,
),
)
fake_redis.rpush(
system_logs.get_buffer_log_key("earth-client"),
json.dumps(
{
"timestamp": "2026-04-23T06:01:00Z",
"level": "error",
"message": "landing points failed",
"context": {"module": "layer-startup", "detail": "http 500"},
},
ensure_ascii=False,
),
)
snapshot = system_logs.read_log_snapshot(
"earth-client",
50,
levels="error,warning",
start_date="2026-04-23",
end_date="2026-04-23",
search="landing",
)
assert snapshot is not None
assert snapshot["selected_levels"] == ["error", "warning"]
assert snapshot["search_query"] == "landing"
assert snapshot["line_count"] == 1
assert snapshot["lines"][0].startswith("2026-04-23 06:01:00 ERROR landing points failed")
assert snapshot["daily_markers"] == [
{"date_token": "2026-04-23", "total": 1, "dominant_level": "error"}
]
def test_read_log_snapshot_parses_file_timestamp_and_builds_markers(tmp_path: Path, monkeypatch):
log_path = tmp_path / "backend.log"
log_path.write_text(
"\n".join(
[
"2026-04-22 08:00:00 INFO service booted",
"2026-04-23 09:15:00 WARNING disk pressure detected",
"2026-04-23 09:16:00 ERROR sync failed",
"2026-04-24 10:00:00 DEBUG collector trace",
]
),
encoding="utf-8",
)
monkeypatch.setattr(
system_logs,
"LOG_SOURCES",
{
"backend": system_logs.LogSource(
source_id="backend",
name="后端服务",
kind="file",
location=str(log_path),
description="测试文件日志",
category="service",
)
},
)
snapshot = system_logs.read_log_snapshot(
"backend",
50,
levels="warning,error",
search="failed",
)
assert snapshot is not None
assert snapshot["line_count"] == 1
assert snapshot["lines"] == ["2026-04-23 09:16:00 ERROR sync failed"]
assert snapshot["daily_markers"] == [
{"date_token": "2026-04-23", "total": 1, "dominant_level": "error"}
]
assert snapshot["status"] == "ok"
def test_append_buffer_log_persists_normalized_level(monkeypatch):
fake_redis = FakeRedis()
monkeypatch.setattr(system_logs, "redis_client", fake_redis)
system_logs.append_buffer_log(
"earth-client",
level="warn",
message="feed delayed",
context={"module": "news"},
)
stored_items = fake_redis.lrange(system_logs.get_buffer_log_key("earth-client"), 0, -1)
payload = json.loads(stored_items[0])
assert payload["level"] == "warning"
assert payload["message"] == "feed delayed"
def test_infer_log_level_prefers_leading_prefix_over_query_string():
line = 'INFO: 127.0.0.1 - "GET /api/v1/system/logs/backend?limit=200&level=error&levels=error HTTP/1.1" 200 OK'
entry = system_logs.parse_text_log_entry(line)
assert entry.level == "info"
def test_parse_text_log_entry_does_not_promote_exception_context_to_error():
line = "websockets.exceptions.ConnectionClosedError: sent 1011 (internal error) keepalive ping timeout"
entry = system_logs.parse_text_log_entry(line)
assert entry.level is None
def test_parse_text_log_entry_still_detects_explicit_error_prefix():
line = "ERROR: [Errno 98] Address already in use"
entry = system_logs.parse_text_log_entry(line)
assert entry.level == "error"
def test_read_log_snapshot_strips_nul_bytes_from_file_lines(tmp_path: Path, monkeypatch):
log_path = tmp_path / "backend.log"
log_path.write_bytes(
(
b"INFO: service booted\n"
b"ERROR: bind failed\n"
+ b"\x00" * 32
+ b"2026-04-23 23:41:32 INFO service=backend message=request served\n"
)
)
monkeypatch.setattr(
system_logs,
"LOG_SOURCES",
{
"backend": system_logs.LogSource(
source_id="backend",
name="后端服务",
kind="file",
location=str(log_path),
description="测试文件日志",
category="service",
)
},
)
snapshot = system_logs.read_log_snapshot("backend", 50)
assert snapshot is not None
assert snapshot["line_count"] == 3
assert snapshot["lines"] == [
"INFO: service booted",
"ERROR: bind failed",
"2026-04-23 23:41:32 INFO service=backend message=request served",
]

View File

@@ -0,0 +1,105 @@
from datetime import datetime, timedelta, timezone
import pytest
from httpx import ASGITransport, AsyncClient
from app.api.v1.visualization import convert_vessels_to_geojson
from app.db.session import get_db
from app.main import app
from app.models.vessel import VesselPosition, VesselStatic
from app.services.collectors.vessel_ais import VesselAISCollector
def test_vessel_collector_transforms_barentswatch_like_records():
collector = VesselAISCollector()
records = collector.transform(
[
{
"mmsi": "257123000",
"lat": "59.91",
"lon": "10.73",
"sog": 12.4,
"cog": 214,
"nav_status": 0,
"shipType": 70,
"name": "OSLO TRADER",
},
{"mmsi": "bad", "lat": 120, "lon": 10},
]
)
assert len(records) == 1
assert records[0]["mmsi"] == 257123000
assert records[0]["vessel_type_name"] == "Cargo"
assert records[0]["lat"] == pytest.approx(59.91)
def test_convert_vessels_to_geojson():
position = VesselPosition(
mmsi=257123000,
lat=59.91,
lon=10.73,
sog=12.4,
cog=214,
heading=215,
nav_status=0,
received_at=datetime(2026, 4, 28, 1, 0, tzinfo=timezone.utc),
)
static = VesselStatic(
mmsi=257123000,
name="OSLO TRADER",
vessel_type=70,
vessel_type_name="Cargo",
flag="NO",
length=185,
)
payload = convert_vessels_to_geojson([(position, static)])
assert payload["type"] == "FeatureCollection"
assert payload["features"][0]["geometry"]["coordinates"] == [10.73, 59.91]
assert payload["features"][0]["properties"]["mmsi"] == 257123000
assert payload["features"][0]["properties"]["vessel_type_name"] == "Cargo"
@pytest.mark.asyncio
async def test_vessels_geojson_endpoint_filters_type_and_bbox():
now = datetime(2026, 4, 28, 1, 0, tzinfo=timezone.utc)
rows = [
(
VesselPosition(mmsi=1, lat=59.9, lon=10.7, received_at=now),
VesselStatic(mmsi=1, name="Cargo Ship", vessel_type=70, vessel_type_name="Cargo"),
),
(
VesselPosition(mmsi=2, lat=60.3, lon=5.3, received_at=now - timedelta(minutes=1)),
VesselStatic(mmsi=2, name="Passenger Ship", vessel_type=60, vessel_type_name="Passenger"),
),
]
class _Result:
def all(self):
return rows
class _FakeSession:
async def execute(self, _query):
return _Result()
async def override_get_db():
yield _FakeSession()
app.dependency_overrides[get_db] = override_get_db
transport = ASGITransport(app=app)
try:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get(
"/api/v1/visualization/geo/vessels",
params={"bbox": "0,50,20,70", "type": "cargo"},
)
assert response.status_code == 200
data = response.json()
assert data["count"] == 1
assert data["features"][0]["properties"]["name"] == "Cargo Ship"
assert data["stats"]["by_type"]["Cargo"] == 1
finally:
app.dependency_overrides.clear()

View File

@@ -0,0 +1,347 @@
from datetime import datetime, timezone
import pytest
from httpx import ASGITransport, AsyncClient
from app.api.v1.visualization import convert_compute_centers_to_geojson
from app.db.session import get_db
from app.main import app
from app.models.collected_data import CollectedData
def _build_record(
*,
record_id: int,
source: str,
data_type: str,
name: str,
country: str,
city: str,
latitude: float,
longitude: float,
metadata: dict,
):
return CollectedData(
id=record_id,
source=source,
data_type=data_type,
source_id=f"{source}-{record_id}",
name=name,
extra_data={
"country": country,
"city": city,
"latitude": latitude,
"longitude": longitude,
**metadata,
},
collected_at=datetime(2026, 4, 22, tzinfo=timezone.utc),
reference_date=datetime(2026, 4, 21, tzinfo=timezone.utc),
is_current=True,
)
def test_convert_compute_centers_to_geojson_unifies_sources():
top500_record = _build_record(
record_id=1,
source="top500",
data_type="supercomputer",
name="Frontier",
country="United States",
city="Oak Ridge",
latitude=35.93,
longitude=-84.31,
metadata={
"rank": 1,
"manufacturer": "HPE",
"organization": "ORNL",
"rmax": 1102000.0,
"cores": 8730112,
"power": 21510.0,
},
)
gpu_record = _build_record(
record_id=2,
source="epoch_ai_gpu",
data_type="gpu_cluster",
name="Colossus",
country="United States",
city="Memphis",
latitude=35.15,
longitude=-90.05,
metadata={
"organization": "xAI",
"gpu_type": "H100",
"gpu_count": 100000,
"value": "20000",
"unit": "TFlop/s",
},
)
payload = convert_compute_centers_to_geojson([top500_record, gpu_record])
assert payload["type"] == "FeatureCollection"
assert len(payload["features"]) == 2
supercomputer_feature = payload["features"][0]
assert supercomputer_feature["properties"]["site_type"] == "supercomputer"
assert supercomputer_feature["properties"]["capacity_unit"] == "GFlops"
assert supercomputer_feature["properties"]["capacity_band"] == "exascale"
assert supercomputer_feature["properties"]["operator"] == "ORNL"
assert supercomputer_feature["properties"]["location_precision"] == "precise"
assert supercomputer_feature["properties"]["is_estimated"] is False
gpu_feature = payload["features"][1]
assert gpu_feature["properties"]["site_type"] == "gpu_cluster"
assert gpu_feature["properties"]["vendor"] == "H100"
assert gpu_feature["properties"]["gpu_count"] == 100000
assert gpu_feature["properties"]["capacity_band"] == "large"
assert gpu_feature["properties"]["location_precision"] == "precise"
def test_convert_compute_centers_to_geojson_uses_coordinate_hints():
hinted_record = _build_record(
record_id=3,
source="top500",
data_type="supercomputer",
name="Frontier",
country="United States",
city="",
latitude=0.0,
longitude=0.0,
metadata={
"organization": "Oak Ridge National Laboratory",
"rmax": 1102000.0,
},
)
payload = convert_compute_centers_to_geojson([hinted_record])
assert len(payload["features"]) == 1
coords = payload["features"][0]["geometry"]["coordinates"]
assert coords[0] == pytest.approx(-84.3107)
assert coords[1] == pytest.approx(35.9319)
assert payload["features"][0]["properties"]["is_estimated"] is True
assert payload["features"][0]["properties"]["location_precision"] == "estimated_site"
def test_convert_compute_centers_to_geojson_falls_back_to_country_centroid():
centroid_record = _build_record(
record_id=4,
source="epoch_ai_gpu",
data_type="gpu_cluster",
name="Unknown Cluster",
country="United States",
city="",
latitude=0.0,
longitude=0.0,
metadata={
"organization": "Unknown Operator",
"value": "10000",
"unit": "TFlop/s",
},
)
payload = convert_compute_centers_to_geojson([centroid_record])
assert len(payload["features"]) == 1
props = payload["features"][0]["properties"]
coords = payload["features"][0]["geometry"]["coordinates"]
assert coords[0] == pytest.approx(-98.5795)
assert coords[1] == pytest.approx(39.8283)
assert props["is_estimated"] is True
assert props["location_precision"] == "estimated_country"
assert props["geography_mode"] == "country_centroid"
@pytest.mark.asyncio
async def test_compute_centers_geojson_endpoint_returns_stats():
records = [
_build_record(
record_id=1,
source="top500",
data_type="supercomputer",
name="Frontier",
country="United States",
city="Oak Ridge",
latitude=35.93,
longitude=-84.31,
metadata={"rank": 1, "rmax": 1102000.0},
),
_build_record(
record_id=2,
source="epoch_ai_gpu",
data_type="gpu_cluster",
name="Colossus",
country="United States",
city="Memphis",
latitude=35.15,
longitude=-90.05,
metadata={"value": "20000", "unit": "TFlop/s"},
),
]
class _ScalarResult:
def __init__(self, rows):
self._rows = rows
def scalars(self):
class _Scalars:
def __init__(self, rows):
self._rows = rows
def all(self):
return self._rows
return _Scalars(self._rows)
class _FakeSession:
async def execute(self, _query):
return _ScalarResult(records)
async def override_get_db():
yield _FakeSession()
app.dependency_overrides[get_db] = override_get_db
transport = ASGITransport(app=app)
try:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get("/api/v1/visualization/geo/compute-centers")
assert response.status_code == 200
data = response.json()
assert data["count"] == 2
assert data["stats"]["supercomputers"] == 1
assert data["stats"]["gpu_clusters"] == 1
assert data["features"][0]["properties"]["data_type"] == "compute_center"
finally:
app.dependency_overrides.clear()
@pytest.mark.asyncio
async def test_visualization_geo_summary_returns_counts(monkeypatch):
records = [
_build_record(
record_id=1,
source="arcgis_cables",
data_type="submarine_cable",
name="Test Cable",
country="",
city="",
latitude=0,
longitude=0,
metadata={
"route_coordinates": [[[0, 0], [1, 1]]],
"status": "active",
},
),
_build_record(
record_id=2,
source="arcgis_landing_points",
data_type="landing_point",
name="Test Landing",
country="United States",
city="New York",
latitude=40.7,
longitude=-74.0,
metadata={"city_id": 10},
),
_build_record(
record_id=3,
source="celestrak_tle",
data_type="satellite_tle",
name="TESTSAT",
country="",
city="",
latitude=0,
longitude=0,
metadata={
"norad_cat_id": 12345,
"tle_line1": "1 12345U 98067A 24001.00000000 .00000000 00000-0 00000-0 0 9991",
"tle_line2": "2 12345 51.6000 100.0000 0001000 10.0000 20.0000 15.50000000 01",
},
),
_build_record(
record_id=4,
source="top500",
data_type="supercomputer",
name="Frontier",
country="United States",
city="Oak Ridge",
latitude=35.93,
longitude=-84.31,
metadata={"rank": 1, "rmax": 1102000.0},
),
_build_record(
record_id=5,
source="epoch_ai_gpu",
data_type="gpu_cluster",
name="Colossus",
country="United States",
city="Memphis",
latitude=35.15,
longitude=-90.05,
metadata={"value": "20000", "unit": "TFlop/s"},
),
]
class _ScalarResult:
def __init__(self, rows=None, scalar_value=None):
self._rows = rows or []
self._scalar_value = scalar_value
def scalar(self):
return self._scalar_value
def scalars(self):
class _Scalars:
def __init__(self, rows):
self._rows = rows
def all(self):
return self._rows
return _Scalars(self._rows)
class _FakeSession:
async def execute(self, query):
query_text = str(query)
if "bgp_incidents" in query_text:
return _ScalarResult(scalar_value=2)
if "bgp_anomalies" in query_text:
return _ScalarResult(scalar_value=3)
return _ScalarResult(rows=records)
async def override_get_db():
yield _FakeSession()
async def _fake_build_bgp_collector_coverage(*_args, **_kwargs):
return [
{"collector": "rrc00"},
{"collector": "rrc01"},
]
monkeypatch.setattr(
"app.api.v1.visualization.build_bgp_collector_coverage",
_fake_build_bgp_collector_coverage,
)
app.dependency_overrides[get_db] = override_get_db
transport = ASGITransport(app=app)
try:
async with AsyncClient(transport=transport, base_url="http://test") as client:
response = await client.get("/api/v1/visualization/geo/summary")
assert response.status_code == 200
stats = response.json()["stats"]
assert stats["cable_count"] == 1
assert stats["landing_point_count"] == 1
assert stats["satellite_count"] == 1
assert stats["compute_center_count"] == 2
assert stats["supercomputer_count"] == 1
assert stats["gpu_cluster_count"] == 1
assert stats["bgp_event_count"] == 2
assert stats["bgp_incident_count"] == 2
assert stats["bgp_anomaly_count"] == 3
assert stats["bgp_collector_count"] == 2
finally:
app.dependency_overrides.clear()

View File

@@ -18,6 +18,9 @@ services:
build:
context: .
dockerfile: aiprovider/Dockerfile
args:
PYTHON_IMAGE: ${PYTHON_IMAGE:-python:3.14-slim}
UV_IMAGE: ${UV_IMAGE:-ghcr.io/astral-sh/uv:latest}
container_name: planet_aiprovider
ports:
- "8010:8010"

View File

@@ -5,6 +5,9 @@ services:
build:
context: .
dockerfile: aiprovider/Dockerfile
args:
PYTHON_IMAGE: ${PYTHON_IMAGE:-python:3.14-slim}
UV_IMAGE: ${UV_IMAGE:-ghcr.io/astral-sh/uv:latest}
container_name: planet_aiprovider
ports:
- "8010:8010"

View File

@@ -5,6 +5,9 @@ services:
build:
context: .
dockerfile: aiprovider/Dockerfile
args:
PYTHON_IMAGE: ${PYTHON_IMAGE:-python:3.14-slim}
UV_IMAGE: ${UV_IMAGE:-ghcr.io/astral-sh/uv:latest}
env_file:
- ./aiprovider/.env
container_name: planet_aiprovider

View File

@@ -8,6 +8,486 @@ This project follows the repository versioning rule:
- `improvement` -> `+0.0.1`bugfix + 小功能混合)
- `bugfix` -> `+0.0.1`
## [0.43.1] — 2026-04-28
### 🐛 Fixes
- 修正 `planet.sh` 在全量 restart 后启动 AI Provider 时的提示语义,避免把预期内未就绪描述成异常不健康
---
## [0.43.0] — 2026-04-28
### ✨ Highlights
- 新增 Earth 船舶追踪链路,接入 BarentsWatch AIS 凭证配置、采集器、后端 vessel 模型/API 与前端 Earth 船舶图层
- 新增自定义数据源映射流程,支持样本抓取、目标 schema、AI 辅助生成映射、预览校验和映射执行
- Settings 拆分 AI Provider 与采集器凭证配置DataSources 只保留采集状态、运行参数和必要引导
### 🔧 Improvements
- AI Provider 支持运行时 LLM 配置、provider preset 下拉与刷新,并在 Playground 中引导到 AI 配置页
- Markdown 渲染器补齐代码块复制按钮、语言标签、任务列表、图片、自动链接、删除线和文档主题样式
- Docs 公开导航改为显式元数据白名单,避免开发任务文档自动出现在“其他”分组
- 将 Codex/Claude cleanup、docs、goal-driven、release 流程补充 CLI-first 约束,并把 `rules.md` 整理成可按模块加载的工程规则
---
## [0.42.2] — 2026-04-28
### 🐛 Fixes
- Docs 中文模式下补齐左侧分组、文档标题、页头分类与搜索结果分类翻译,并更新文档站品牌标题/副标题文案
---
## [0.42.1] — 2026-04-28
### 🐛 Fixes
- 修正 release skill 的 feature 版本计算规则minor 进位时 patch 必须重置为 `0`,例如 `0.41.2` 应发布为 `0.42.0`
---
## [0.42.0] — 2026-04-28
### ✨ Highlights
- 新增公开 `/docs` 文档站支持中英文技术文档、使用手册、Quickstart、搜索、目录锚点与浅色/深色/跟随系统主题
- Earth 在无高清材质时新增轻量 Fresnel 边缘提示,并调整卫星覆盖默认显示与地表材质可读性
### 🔧 Improvements
- 将技术文档整理为 `docs/technical/zh``docs/technical/en`,并补充控制台、`planet.sh`、Earth 与公共组件使用说明
- 新增 `SegmentedControl` 公共滑块组件,支持缩放参数,复用到 docs 语言与主题切换
- Markdown 渲染器接入自定义滚动条,表格与代码块在深色模式和 overflow 场景下保持可读
- Docs 搜索结果支持内部滚动、点击外部关闭、重新聚焦恢复上次搜索结果
- Earth 工具栏展开状态与设置持久化版本迁移继续收口,改善默认面板和快捷关闭行为
---
## [0.41.2] — 2026-04-27
### 🔧 Improvements
- `planet.sh` 启动链路新增 verbose 滚动输出窗口,并在后端端口占用时打印目标地址和监听进程诊断
- Docker 构建支持通过 build args 覆盖 Python 与 uv 镜像,方便 Docker Hub 不稳定时切换镜像源
### 🐛 Fixes
- Earth 海缆登陆点改为基于相机射线与地球遮挡判断可见性,修复旋转后 pin 可见性滞后一帧的问题
### 🔧 Improvements
- `docker-compose*.yml` 为 AI Provider 构建传入 `PYTHON_IMAGE` / `UV_IMAGE` 参数,默认仍使用官方镜像
- 后端启动失败遇到 `Address already in use` 时输出 `lsof``ss` 与 PID 命令行信息
- verbose 模式下 AI Provider build、后端与前端启动日志会在 spinner 下方保留最新 5 行滚动展示
---
## [0.41.1] — 2026-04-27
### 🐛 Fixes
- 修复新闻直播面板设置项持久化失效:`closeTransientMobileOverlays` 通过旁路路径隐藏面板导致下次 persist 快照到错误状态,改为不重新从 DOM 读取面板可见性
- 修复登陆点 pin 在地球侧面被半截遮挡改为在接近地平线前dot < 0.05)主动隐藏,避免深度测试切片
### 🔧 Improvements
- 将所有画布绘制的图标抽取为 SVG存入 `frontend/public/earth/assets/icons/`,新增图标规范到 `rules.md`
---
## [0.41.0] — 2026-04-27
### ✨ Highlights
- Earth 图层系统完成地表到天空的注册顺序与关注优先的面板顺序拆分支持基座海陆色块、国界、高清材质、云图、地形、算力、BGP、卫星、轨迹与海缆的稳定层级
- 国界层新增真实行政区轮廓交互与中国/台湾联动高亮修复高清材质、地形、footprint、卫星与经纬线之间的遮挡和 hover 竞争
### 🔧 Improvements
- 新增无轮廓基座地图,所有图层关闭时仍保留 `#010609` 海洋与 `#080f1b` 陆地色块
- 将大气云图抽象为独立图层并接入桌面/移动端图层开关、持久化状态与启动同步
- 高清材质改为独立纹理覆盖层,地形显示在高清材质上方,并在高清材质关闭/恢复时保持原地形开关意图
- 补充 Earth 渲染层级与图层样式文档,记录正式图层名、变量名、材质颜色、线宽与 renderOrder
---
## [0.40.5] — 2026-04-26
### 🔧 Improvements
- 卫星拖尾改用 Instanced screen-space ribbon单 draw call 渲染所有轨迹段,支持像素级宽度控制
- Iridium 地面覆盖重写为球面投影径向网格,修复填充光晕不可见问题;新增外圈 LineLoop
- 搜索面板打开时改用双 rAF 延迟聚焦输入框,确保 CSS 过渡完成后焦点可靠触发
- 代码清理:提取 `IRIDIUM_OVERLAY_COLOR``IRIDIUM_REFERENCE_ALTITUDE_KM` 常量,消除重复三角函数调用
---
## [0.39.0] — 2026-04-24
## [0.40.4] — 2026-04-26
### 🔧 Improvements
- 新增页面可见性恢复处理,页面从后台切回前台时主动刷新卫星位置,避免累积后台时间在下一帧一次性回放
- 抽出卫星轨迹状态与轨迹几何清理 helper统一后台恢复与清空数据时的轨迹重置路径
### 🐛 Fixes
- 修复页面在后台停留较久后恢复前台时,卫星轨迹因超大 `deltaTime` 突然跳变、拖尾异常拉长的问题
- 修复后台恢复后首帧仍沿用旧轨迹缓存,导致轨迹与当前卫星位置短时错位的问题
---
## [0.40.3] — 2026-04-25
### 🔧 Improvements
- 卫星点云升级为自定义 ShaderMaterial支持 per-point alpha 控制,锁定/悬停卫星从点云中精确隐藏
- 修复锁定环与自发光选中标记的 depthTest 错误false → true消除远端渲染穿透 artifact
- 新增锁定环悬停态缩放与线宽LOCKED_RING_HOVER_SCALE / LOCKED_RING_HOVER_LINE_WIDTH
- 修复 updateLockedDotWorldTransform / updateLockedHaloWorldTransform 未强制刷新 matrixWorld 导致的位置漂移
---
## [0.40.2] — 2026-04-24
### 🔧 Improvements
- 卫星点大小随镜头缩放动态调整,拉近变大、拉远变小,响应与相机距离线性对应
- 调小卫星点默认基础尺寸dotSize 2.8),缩放范围更合理
---
## [0.40.1] — 2026-04-24
### 🔧 Improvements
- 卫星选中标记lockedring / lockeddot / 光晕)颜色统一跟随图例轨道倾角分类配色
- 修复 Starlink footprint 在特定视角下遮蔽卫星点的渲染顺序问题Group renderOrder 影响子 Mesh 排序)
- footprint 材质改为 `depthTest: false` + 相机朝向 limbFade替代 polygonOffset 深度竞争方案
- 修复选中海缆时误触发附近卫星高亮(该行为属于 BGP 事件点逻辑,不应用于海缆)
---
## [0.40.0] — 2026-04-24
### ✨ Highlights
- Earth 卫星 footprint 正式按星座能力分层Starlink 保留专用地表覆盖Iridium 改为独立外圈覆盖表达,其它非 Starlink 星座不再误用同一套 footprint
- Earth 卫星详情卡补齐覆盖能力与当前显示说明,用户现在可以直接看见每颗卫星为什么显示 footprint、为何回退为自身发光
### 🔧 Improvements
- 后端可视化接口新增并透传 `constellation_group``footprint_policy`,前端据此执行 capability-gated footprint renderer
- 新增 Iridium 独立 coverage ring adapter并继续保留 Starlink 专用 footprint 调校与昼夜可读性增强
- 新增 Earth 卫星 footprint 策略技术文档,明确 GNSS、generic LEO、GEO 与 Iridium 的显示边界
### 🐛 Fixes
- 修复前后端对 Iridium footprint policy 命名不一致,导致策略分发语义含混的问题
- 清理 Starlink footprint 渲染中的未使用常量与过时命名,减少后续继续调校时的歧义
---
## [0.39.0] — 2026-04-24
### ✨ Highlights
- 后端正式落下统一结构化日志地基:请求上下文、事件名、脱敏与持久化链路开始收口为可扩展的企业级日志体系
- 系统日志页重构为真正的日志工作台:顶部筛选更紧凑,终端日志区成为主视觉,移动端 Earth 新闻/态势细节交互继续补稳
### 🔧 Improvements
- 新增 `backend/app/core/logging.py`,统一 `request_id``service``event` 注入与敏感字段脱敏,并接入后端主入口、调度器、缓存、数据库和可视化链路
- 系统日志页筛选区重排为更紧凑的两层结构,信息摘要并入终端工具栏 tooltip日志终端区留出更稳定的按钮避让空间
- Earth 移动端态势抽屉补齐宽度约束与图例换行规则,新闻详情抽屉在巡航切换时可同步更新标题和摘要
### 🐛 Fixes
- 修复 `/tmp/planet_backend.log` 中混入空字节时,日志摘要条行数与实际可见日志不一致的问题
- 修复移动端“态势”tab 在内容渲染后被图例文本撑宽、超出一屏的问题
- 修复移动端新闻详情抽屉在巡航切换下一条新闻时标题更新但 summary 不同步的问题
---
## [0.38.0] — 2026-04-23
### ✨ Highlights
- Earth 新闻正式接入通用巡航层:新闻和 BGP 统一进入可配置巡航模块,桌面端与移动端都能在巡航聚焦时展示对应新闻卡片
- 系统日志页升级为结构化过滤链路:按真实时间戳、结构化级别和字符串检索统一筛选,不再依赖前端或后端从日志文本里猜结果
### 🔧 Improvements
- 新闻巡航补齐业务适配层:按发生地与时间生成巡航目标,桌面端与移动端统一标题 + summary 卡片风格,并增加连线与打字机摘要展示
- 日志页筛选体验重排,统一服务源、级别、行数、时间和检索布局,日历标记改为由后端返回的结构化每日聚合结果驱动
- 后端补充 `system_logs` 结构化解析与多级别精确过滤能力Earth 浏览器端日志缓冲与系统日志 API 现在走同一套筛选语义
### 🐛 Fixes
- 修复新闻巡航模块开启后难以关闭、桌面/移动端设置状态互相污染的问题
- 修复新闻巡航卡片缺少摘要、移动端详情样式不统一、新闻巡航缺少连线的问题
- 修复日志级别筛选会被访问日志 query string 中的 `level=error` 等参数污染,从而把 `INFO` 行误判为 `ERROR` 的问题
---
## [0.37.2] — 2026-04-23
### ✨ Highlights
- Earth 图层系统新增经纬线开关,桌面图层面板与移动端抽屉都可直接控制
### 🔧 Improvements
- 经纬线正式接入 Earth layer registry复用现有图层切换、移动端图层卡片与设置持久化流
### 🐛 Fixes
- 修复经纬线只能默认常驻、无法作为独立图层开关控制的问题
---
## [0.37.1] — 2026-04-23
### ✨ Highlights
- `planet.sh` 后端重启链路修复 `uvicorn --reload` 残留 worker 场景,`restart` 现在能真正替换旧实例
### 🔧 Improvements
- 收口后端清理逻辑,统一按 `uvicorn` 进程、端口占用进程和进程组执行清理,减少 reload 场景漏杀分支
### 🐛 Fixes
- 修复部分机器执行 `./planet.sh restart --allow-lan` 后后端仍停留旧实例,导致 `/api/v1/visualization/geo/compute-centers` 返回 `404` 的问题
---
## [0.35.1] — 2026-04-22
## [0.37.0] — 2026-04-23
### ✨ Highlights
- Earth 连线系统正式从巡航里解耦成通用 callout connector桌面端和移动端统一支持对象级锚点、四边切换与临界区边缘滑动
- BGP 巡航展示继续收口为稳定的“先定位卡片、再连真实锚点、再展示卡片”链路,移动端 popup 与桌面 info panel 的路线规则统一
### 🔧 Improvements
- connector 配置从 `CRUISE_CONFIG` 拆到独立 `CONNECTOR_CONFIG`,默认类名、动画名和实例命名也全部去 cruise 语义
- 移动端 popup 增加更稳定的 dock/obstacle 处理,拖动卡片时连线起终点会持续按几何关系自适应刷新
- Earth 多个图层与控制逻辑继续收口,补充算力中心/BGP 风格对齐、layer panel 与相关交互细节调整
### 🐛 Fixes
- 修复巡航模式下终点只像“视觉锚点”而不是真实绑定对象的问题,卡片拖动后终点现在会跟随
- 修复移动端与桌面端多类连线路线异常:压线、反向、临界区折返、起点遮挡事件点等问题
- 修复对象矩形临界区内连线仍强制中点到中点导致路线像“先钻进 source 内部”再出去的问题
---
## [0.35.1] — 2026-04-22
## [0.36.0] — 2026-04-22
### ✨ Highlights
- Earth 新增统一“算力中心”图层:接入超算与 GPU 集群,支持搜索、统计、图例、详情卡与独立图层开关
- 算力中心支持精确位置与估算位置两种状态,估算点会以问号角标区分,避免数据不全时整批节点在地图上消失
### 🔧 Improvements
- Earth 详情卡拖拽与地球拖拽交互继续收口,减少拖动卡片和旋转地球时的选中文本与 pointer 竞争
- `planet.sh` 改为通过独立脚本计算 AI Provider 依赖指纹,降低与根仓库依赖版本文件的无关耦合
- README 补充 WSL / Windows 局域网访问排查与转发配置说明,便于开发环境联调
### 🐛 Fixes
- 修复 Earth 算力中心图层在无原始坐标时无法显示的问题,支持站点提示和国家级估算回退
- 修复信息卡拖拽事件可能被卡片级 stopPropagation 吞掉,导致拖拽流中断的问题
---
## [0.35.1] — 2026-04-22
### ✨ Highlights
- Earth 统计展示改为统一 `data-earth-stat` 绑定机制,桌面 HUD 和移动端抽屉复用同一套状态更新入口
### 🔧 Improvements
- 收口海缆、登陆点、卫星、BGP 事件与 BGP 状态的统计写入逻辑,减少后续继续补桌面/移动双写分支的成本
### 🐛 Fixes
- 修复移动端态势抽屉中的海缆、登陆点与 BGP 统计在图层切换后可能停留旧值的问题
---
## [0.35.0] — 2026-04-22
### ✨ Highlights
- Earth 移动端底部抽屉系统全面上线响应式布局自动切换、Tab 导航、手势上拉/下滑开合、惯性速度判定
- 移动端点击可交互物件海缆、登陆点、卫星、BGP后弹出智能定位悬浮卡片可拖动点击跳转详情
### 🔧 Improvements
- 抽屉把手区域缩小至 36pxcollapsed 时仅露出把手,不遮挡地球操作区)
- 抽屉定期弹跳动画提示用户可上拉5 秒间隔,打开后自动停止
- 通知胶囊位置调整,不再覆盖品牌 logo
- 移动端单指旋转、双指捏合缩放地球触控事件冲突修复pointer-events 级联)
### 🐛 Fixes
- 修复移动端抽屉 shell 因 layout 高度240px+遮挡地球触控区域pointer-events 改为按层级精确控制
- 修复悬浮卡片因 setPointerCapture 在 iOS Safari 抑制合成 click 事件导致无法点击的问题
---
## [0.34.0] — 2026-04-22
### ✨ Highlights
- Earth 搜索面板正式接入支持搜索海缆、登陆点、卫星、BGP 事件与观测站,并可直接聚焦到对应对象
- `planet.sh --allow-lan` 打通 Bun + Vite 的局域网开放链路,启动成功后自动打印推荐访问地址与后端健康检查地址
### 🔧 Improvements
- 前端开发启动链统一改成 Bun 直接执行 Vite 入口,不再依赖 shell 中额外暴露的 Node 路径
- Earth 搜索结果接入登陆点详情卡片与对象聚焦,搜索后可直接进入对应详情流
- `planet.sh` 补充局域网 IPv4 自动识别与推荐地址输出,减少 WSL 局域网调试成本
### 🐛 Fixes
- 修复 `./planet.sh restart --allow-lan` 全量重启时未把 `--allow-lan` 继续传给 `start()`,导致前端退回本机监听的问题
- 修复 WSL + Bun 环境下前端偶发因 Vite 启动链不稳定而无法正确监听 `0.0.0.0:3000` 的问题
---
## [0.33.0] — 2026-04-22
### ✨ Highlights
- `news_live_streams` 采集器默认接入 `iptv-org` 频道目录,并将采集结果稳定并入 Earth TV 直播源列表
- 数据源页支持直接编辑内置数据源 override并为内置源提供一键恢复默认配置入口
### 🔧 Improvements
- `News Live Streams` 现在作为可直接触发的内置默认数据源提供,无需先手工补 override 才能采集
- TV 播放源菜单会直接区分 `[内置]``[采集]` 来源,频道来源信息也会同步展示
- 新增 [earth-news-source-configuration-and-collector-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-news-source-configuration-and-collector-plan.md),正式规划 Earth 态势新闻源配置化与后续采集器化路线
### 🐛 Fixes
- 修复 `news_live_streams` 采集完成后 `/api/v1/tv/streams` 因读取不存在的 `updated_at` 字段而导致默认频道全部消失的问题
- 修复内置数据源操作列按钮显示不全,以及编辑抽屉中多个 `Collapse` 紧贴的问题
---
## [0.32.0] — 2026-04-22
### ✨ Highlights
- Earth 设置新增“地球默认大小”持久化项,重置视角、缩放百分比重置和 BGP 巡航视图现在统一复用这一份默认 zoom
- 卫星焦点层次继续收口:巡航进入 presentation 前不再过早 dim非焦点卫星改成“降亮度/尾迹/背板”而不是去饱和度
### 🔧 Improvements
- Earth 设置面板区块和左右留白进一步收紧,整体更贴近 HUD 面板的密度
- toolbar 展开边界缓存改为按需刷新,减少 document 级 mousemove 期间的重复布局读取
- Scrollbar 和 ScrollbarOverlay 收窄 observer 范围,减少大表格和动态菜单下的额外刷新成本
- 更新 [earth-frontend-context.md](/home/ray/dev/linkong/planet/docs/technical/earth-frontend-context.md),补充默认视图大小已进入 Earth 设置持久化真源
### 🐛 Fixes
- 修复开启巡航后,尚未进入连线/presentation 时卫星已经整体变暗的问题
- 修复默认大小重置链路分散在多个入口、实际 reset/cruise/缩放提示不一致的问题
- 修复开启地形后卫星反馈层与地球背面可见性之间的一组表现问题,保留正面反馈同时恢复背面轨道遮挡
---
## [0.31.3] — 2026-04-22
### ✨ Highlights
- Earth 图层注册表和启动任务框架继续收口,启动顺序、启动模式、启动提示和任务注册现在都能从统一入口扩展
- 修复 Earth 普通旋转模式与巡航模式切换时的一组交互回归,同时让卫星/地形/昼夜模式的表现更稳定
### 🔧 Improvements
- 新增 [layer-startup-tasks.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/layer-startup-tasks.js) 启动任务注册表,支持 `registerLayerStartupTask(id, taskFactory)`,并拆成海缆 / 卫星 / BGP 独立注册函数
- Earth 图层控制改成注册表驱动,统一承载 `startupPriority``startupMode``startupLabel``startupMessage` 与图层持久化元信息
- Earth 设置支持持久化图层开关、旋转模式、HUD 面板显示状态、地形透明度与日夜模式,并提供一键重置
- 更新 [earth-frontend-context.md](/home/ray/dev/linkong/planet/docs/technical/earth-frontend-context.md) 记录图层注册表、启动任务、设置持久化与巡航适配边界
### 🐛 Fixes
- 修复普通旋转模式下点击海缆 / 卫星 / BGP 后卡片和选中表现会被异常清空的问题
- 修复巡航模式切回旋转再切回巡航后无法继续自动巡航的问题
- 修复开启地形后卫星选中反馈层被高海拔区域吞掉的问题,并恢复轨道只在地球前半侧可见
- 修复关闭日夜模式后地球照明仍沿真实昼夜切换、亮部过曝和偏色的问题,改成更中性的 inspection lighting
- 修复 toolbar 收起态仍挡住地球交互,以及首帧短暂展开闪现的问题
---
## [0.31.2] — 2026-04-21
### ✨ Highlights
- Earth 巡航模式重构为“通用巡航队列 + 通用连线动画 + BGP 业务适配”三层结构,后续扩到海缆、卫星或新闻巡航时不必再复制一套 `main.js` 状态机
- 修复巡航重构后的交互回归:空白点击重新稳定切到下一项,连线按“起点 → 引导线 → 终点”顺序入场
### 🔧 Improvements
- 新增 [cruise-sequencer.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/cruise-sequencer.js) 统一管理队列推进、停留时长、打断与恢复
- 新增 [callout-connector.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/callout-connector.js) 统一管理 SVG 连线、折线路径与描边动画
- 新增 [bgp-cruise-adapter.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/bgp-cruise-adapter.js) 收口 BGP 巡航目标排序、卡片落点、轮询去重与连线适配
- 更新 [earth-frontend-context.md](/home/ray/dev/linkong/planet/docs/technical/earth-frontend-context.md) 说明新的巡航分层与复用边界
### 🐛 Fixes
- 修复巡航模式下点击空白处无法稳定跳转到下一项、切回旋转再切回巡航后直接卡住的问题
- 修复巡航连线被实时重定位覆盖导致“直接出现”而非绘制动画的问题
- 修复连线动画节点入场节奏不对的问题,改为先出现起点,再绘制连线,最后出现终点
---
## [0.31.1] — 2026-04-21
### ✨ Highlights
- Earth 图层开关状态统一成可复用的 `active / loading` 状态机,首次启用地形和卫星时不再像按钮失效
- 文档目录重构为 `docs/technical``docs/plans``docs/deprecated`,并吸收 `.sisyphus/plans` 中有价值的 Earth / 卫星 / UE5 草案
### 🔧 Improvements
- 新增 [layer-button-state.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/layer-button-state.js),统一按钮 tooltip、`aria-busy`、禁用态和状态文本同步
- 地形图层支持 hover/focus 预热与空闲预热,首次点击等待前移,加载中状态持续可见
- 卫星图层启用前会立即切换为 `loading` 中间态,请求完成后再切回正常开关表现
### 🐛 Fixes
- 修复地形首次加载时通知过早消失、开关仍像关闭状态导致用户误判按钮损坏的问题
- 修复卫星接口较慢时按钮没有任何中间态反馈的问题
---
## [0.31.0] — 2026-04-21
### ✨ Features
- Earth 新增"巡航展示"模式:自动轮播 BGP 异常事件逐帧追踪连接线位置支持外部交互立即中断序列cancel notifier 模式)
- 巡航目标事件点高亮显示hover 外观 + 锁定脉冲动画,并与点击行为统一展示周边受影响卫星与海缆
- BGP 事件图标新增填充 W 形波动符号flap 类型),替换原有难以辨认的贝塞尔细线
- 巡航/点击激活时其余卫星自动降饱和度 + 增加透明度以突出焦点;海缆未受影响时同步变暗
### 🔧 Improvements
- 修复巡航轮播期间 BGP 事件 polling 刷新导致标记闪烁消失的问题clearBGPData 延迟到请求完成后执行)
- 点击与巡航锁定颜色统一为 hover 色0.92, 0.98, 1.0 全透明),移除锁定态脉冲动画
- 巡航连接折线转折点从尖角调整为钝角linkElbowDropPx提升连线可读性
---
## [0.30.0] — 2026-04-21
### ✨ Features
- Earth 新增真实地形图层:后端代理 Terrarium DEM 瓦片(`/api/v1/visualization/terrain/terrarium/{z}/{x}/{y}.png`),前端新增 `terrain.js` 负责瓦片拉取、顶点位移与按海拔着色
- 设置弹窗新增"地形"分组,支持通过滑块实时调整地形图层透明度
### 🔧 Improvements
- 地形按钮改为异步加载,首次点击显示进度提示并在失败时自动回退
- 启动阶段改用 `applyImmediateView` 直接应用初始视角,`showStatusMessage` / `queueStatusMessage` 区分即时与队列态状态消息,加载中不再被临时状态打断
- 控制面板抽取 `applyTerrainUiState` / `getViewRotation` 收敛地形切换与视角旋转的重复 UI 同步逻辑
---
## [0.29.2] — 2026-04-21
### ✨ Highlights
- Earth 继续收口 HUD 交互与设置面板表现,设置弹窗改成更接近从按钮展开的窗口感,同时加入系统级 admin 入口
- 修正天球太阳方向与地球受光解耦后的日照逻辑,地表昼夜判断改为按太阳直射点经纬度落到地球贴图坐标
### 🔧 Improvements
- toolbar 进一步收成更贴近 hub 的浅弓形排列,并统一成与 HUD panel 一致的液态玻璃配色与透明度
- 设置弹窗与各 HUD panel 继续统一样式、等比缩放和头部基线,设置列表补充系统分组与 admin 跳转
- 所有 HUD panel 增加更统一的液态玻璃高光与 hover / press 反馈
### 🐛 Fixes
- 修复设置弹窗仍像旧圆角矩形、标题文案重复和从底边直直飞出的动画问题
- 修复天球与太阳方向混用显示校准导致中国白天仍落在夜面的日照错误
---
## [0.29.1] — 2026-04-20
### ✨ Highlights
- Earth 加载状态条改成单一队列式通知面板,加载阶段不再因为步骤文案变化而回缩,也不会被其他通知打断
- 调整 brand panel 的呈现方式与昼夜/选中态可读性,让品牌区更自然、交互高亮在白天和黑夜里都更稳定
### 🔧 Improvements
- 移除旧的地球加载浮层结构,统一由 HUD 状态消息承载三点脉冲加载过程
- brand panel 改为无边框品牌层,仅保留轻微氛围光,不再因为非常规尺寸显得像第五块功能面板
- 温和收敛地球昼夜材质与主背光强度,保留昼夜辨识度的同时提升白天地表纹理和夜面交互可见性
### 🐛 Fixes
- 修复加载地球时通知条在步骤切换中反复缩短、其他状态消息抢占加载流程的问题
- 修复海缆、登陆点和 BGP 选中高亮在黑夜中过暗、在高光中过亮导致难以辨识的问题
---
## [0.29.0] — 2026-04-20
### ✨ Highlights
- Earth 新增天球层第一版:引入真实全天星图、亮星层与太阳/月亮位置计算,地球场景首次具备可校准的天文背景
- 地球昼夜分隔升级为更明显的日夜增强效果,夜面、晨昏带和太阳方向联动更容易直接读出来
### 🔧 Improvements
- 新增 `celestial.js` 模块和 `assets/celestial/` 资源目录,统一管理星图、亮星数据以及太阳/月亮与光照同步
- 卫星图例改为按倾角分组,严格固定为“赤道轨道 → 低倾角轨道 → 中倾角轨道 → 高倾角轨道 → 逆行轨道”顺序,并全部中文化
- 图层面板补齐关闭按钮,拖拽脱离左列后不再被流布局 margin 影响,能够真正贴到品牌面板下沿
### 🐛 Fixes
- 修复天球球壳放大后被相机 far plane 裁剪导致的外层黑环问题
- 修复图层面板在左侧上移时始终与 brand panel 保持额外间距的问题
---
## [0.28.2] — 2026-04-20
### ✨ Highlights
@@ -229,7 +709,7 @@ Released: 2026-04-12
- Added [backend/app/api/v1/tv.py](/home/ray/dev/linkong/planet/backend/app/api/v1/tv.py), [backend/app/services/tv_streams.py](/home/ray/dev/linkong/planet/backend/app/services/tv_streams.py), and [backend/app/services/collectors/news_live_streams.py](/home/ray/dev/linkong/planet/backend/app/services/collectors/news_live_streams.py) to provide TV source configuration, public stream payloads, a guarded HLS proxy path, and a collector entry point for future world-news live-source ingestion.
- Added the Earth TV HUD workspace through [frontend/public/earth/index.html](/home/ray/dev/linkong/planet/frontend/public/earth/index.html), [frontend/public/earth/js/tv.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/tv.js), and [frontend/public/earth/css/tv-panel.css](/home/ray/dev/linkong/planet/frontend/public/earth/css/tv-panel.css), including toolbar access, draggable/closable behavior, resize support, direct video/HLS playback, iframe fallback, and per-channel external-open handling.
- Added [docs/deprecated/earth-tv-live-module-plan.md](/home/ray/dev/linkong/planet/docs/deprecated/earth-tv-live-module-plan.md) and [docs/earth/news-live-streams-collector-format.md](/home/ray/dev/linkong/planet/docs/earth/news-live-streams-collector-format.md) to document the TV module rollout plan and the expected collector payload format for future curated live-channel ingestion.
- Added [docs/deprecated/earth-tv-live-module-plan.md](/home/ray/dev/linkong/planet/docs/deprecated/earth-tv-live-module-plan.md) and [docs/earth/technical/news-live-streams-collector-format.md](/home/ray/dev/linkong/planet/docs/technical/earth-news-live-streams-collector-format.md) to document the TV module rollout plan and the expected collector payload format for future curated live-channel ingestion.
### Improved
@@ -315,7 +795,7 @@ Released: 2026-04-10
- Improved [frontend/src/pages/Playground/Playground.tsx](/home/ray/dev/linkong/planet/frontend/src/pages/Playground/Playground.tsx) and [frontend/src/index.css](/home/ray/dev/linkong/planet/frontend/src/index.css) by rebuilding Playground into a true chatbox workflow with persistent history, edit-and-resend behavior, grounded message actions, responsive composer behavior, bottom-stick scrolling, and tighter mobile layout handling.
- Improved [frontend/src/components/AppLayout/AppLayout.tsx](/home/ray/dev/linkong/planet/frontend/src/components/AppLayout/AppLayout.tsx), [frontend/src/App.tsx](/home/ray/dev/linkong/planet/frontend/src/App.tsx), and [frontend/src/pages/Alerts/Alerts.tsx](/home/ray/dev/linkong/planet/frontend/src/pages/Alerts/Alerts.tsx) by reorganizing navigation around `采集与数据`, `专题观测`, and split alert entries so the app can scale to more observability and situational modules without turning the top-level UI into a single overloaded page.
- Improved [README.md](/home/ray/dev/linkong/planet/README.md) and [docs/agents/situational-awareness-foundation-plan.md](/home/ray/dev/linkong/planet/docs/agents/situational-awareness-foundation-plan.md) by documenting the current AI/alerts base, planned situational-awareness direction, and the new persistent Playground foundation.
- Improved [README.md](/home/ray/dev/linkong/planet/README.md) and [docs/agents/situational-awareness-foundation-plan.md](/home/ray/dev/linkong/planet/docs/plans/agents-situational-awareness-foundation-plan.md) by documenting the current AI/alerts base, planned situational-awareness direction, and the new persistent Playground foundation.
### Fixed
@@ -353,7 +833,7 @@ Released: 2026-04-10
### Improved
- Improved [rules.md](/home/ray/dev/linkong/planet/rules.md) by adding mandatory release-workflow requirements and a new frontend layout constraint section covering single-screen workspaces, overflow ownership, tab-pane behavior, compact-mode expectations, and readable-card fallbacks.
- Improved [frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/frontend/frontend-layout-guidelines.md) by summarizing the recurring Earth, Playground, BGP, and admin-layout regressions into concrete constraints for future frontend work, including “prefer scrollbars over unreadable compression” and “do not treat every tab as a table pane.”
- Improved [frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/technical/frontend-layout-guidelines.md) by summarizing the recurring Earth, Playground, BGP, and admin-layout regressions into concrete constraints for future frontend work, including “prefer scrollbars over unreadable compression” and “do not treat every tab as a table pane.”
## 0.24.6
@@ -370,7 +850,7 @@ Released: 2026-04-10
- Improved [backend/app/services/bgp_incidents.py](/home/ray/dev/linkong/planet/backend/app/services/bgp_incidents.py) and [backend/app/services/bgp_enrichment.py](/home/ray/dev/linkong/planet/backend/app/services/bgp_enrichment.py) by avoiding historical full-table infrastructure scans, narrowing observation baseline payloads to required columns, and pushing more ASN filtering into the database.
- Improved [backend/app/api/v1/alerts.py](/home/ray/dev/linkong/planet/backend/app/api/v1/alerts.py), [backend/app/api/v1/dashboard.py](/home/ray/dev/linkong/planet/backend/app/api/v1/dashboard.py), and [backend/app/api/v1/settings.py](/home/ray/dev/linkong/planet/backend/app/api/v1/settings.py) by collapsing several repeated count and settings queries into fewer aggregate or batched reads.
- Improved [frontend/src/pages/BGP/BGP.tsx](/home/ray/dev/linkong/planet/frontend/src/pages/BGP/BGP.tsx), [frontend/src/index.css](/home/ray/dev/linkong/planet/frontend/src/index.css), and [frontend/src/components/MarkdownRenderer/MarkdownRenderer.tsx](/home/ray/dev/linkong/planet/frontend/src/components/MarkdownRenderer/MarkdownRenderer.tsx) by rebuilding the `AI 简报` tab layout, fixing saved brief scrolling behavior, and extending the renderer to handle tables, separators, and stored metadata comments more gracefully.
- Improved [docs/frontend/ai-playground-development-plan.md](/home/ray/dev/linkong/planet/docs/frontend/ai-playground-development-plan.md) by explicitly recording that the current BGP brief is only the first-stage summary flow and that regional prefix-geography analysis remains a planned Phase B follow-up.
- Improved [docs/frontend/plans/ai-playground-development-plan.md](/home/ray/dev/linkong/planet/docs/plans/frontend-ai-playground-development-plan.md) by explicitly recording that the current BGP brief is only the first-stage summary flow and that regional prefix-geography analysis remains a planned Phase B follow-up.
### Fixed
@@ -476,8 +956,8 @@ Released: 2026-04-09
### Added
- Added [frontend/src/pages/Playground/Playground.tsx](/home/ray/dev/linkong/planet/frontend/src/pages/Playground/Playground.tsx), introducing the first dedicated AI testing workspace with provider status visibility, prompt/result tabs, and collapsible operator guidance.
- Added [docs/frontend/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/frontend/frontend-layout-guidelines.md), documenting the repository standard for one-screen admin workspaces and module-local overflow handling.
- Added [docs/frontend/ai-playground-development-plan.md](/home/ray/dev/linkong/planet/docs/frontend/ai-playground-development-plan.md), capturing the completed AI gateway/UI work and the next delivery phases for BGP briefs, evidence-first inputs, and future agent runtime expansion.
- Added [docs/frontend/technical/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/technical/frontend-layout-guidelines.md), documenting the repository standard for one-screen admin workspaces and module-local overflow handling.
- Added [docs/frontend/plans/ai-playground-development-plan.md](/home/ray/dev/linkong/planet/docs/plans/frontend-ai-playground-development-plan.md), capturing the completed AI gateway/UI work and the next delivery phases for BGP briefs, evidence-first inputs, and future agent runtime expansion.
### Improved
@@ -582,7 +1062,7 @@ Released: 2026-04-07
- Added [backend/app/services/ai_client.py](/home/ray/dev/linkong/planet/backend/app/services/ai_client.py), introducing an internal HTTP client for `backend -> aiprovider` calls with request-id propagation and lightweight retry.
- Added [aiprovider/main.py](/home/ray/dev/linkong/planet/aiprovider/main.py), [aiprovider/provider_service.py](/home/ray/dev/linkong/planet/aiprovider/provider_service.py), and related config/schema files to stand up the dedicated adapter service.
- Added [aiprovider/.env.example](/home/ray/dev/linkong/planet/aiprovider/.env.example) and [docker-compose.local-model.yml](/home/ray/dev/linkong/planet/docker-compose.local-model.yml) as ready-to-edit local-model templates.
- Added [docs/agents/aiprovider.md](/home/ray/dev/linkong/planet/docs/agents/aiprovider.md), documenting architecture, configuration, single-machine and multi-machine deployment, and cross-service calling patterns.
- Added [docs/agents/aiprovider.md](/home/ray/dev/linkong/planet/docs/technical/agents-aiprovider.md), documenting architecture, configuration, single-machine and multi-machine deployment, and cross-service calling patterns.
- Added a dedicated `重启 AI Provider` control path in [Dashboard.tsx](/home/ray/dev/linkong/planet/frontend/src/pages/Dashboard/Dashboard.tsx), [system_control.py](/home/ray/dev/linkong/planet/backend/app/services/system_control.py), and [system_restart_runner.py](/home/ray/dev/linkong/planet/backend/scripts/system_restart_runner.py).
### Improved
@@ -708,7 +1188,7 @@ Released: 2026-04-02
- Added a new `IPtoASN Prefix Geography` collector in [iptoasn.py](/home/ray/dev/linkong/planet/backend/app/services/collectors/iptoasn.py) and registered it through [data_sources.yaml](/home/ray/dev/linkong/planet/backend/app/core/data_sources.yaml), [data_sources.py](/home/ray/dev/linkong/planet/backend/app/core/data_sources.py), [datasource_defaults.py](/home/ray/dev/linkong/planet/backend/app/core/datasource_defaults.py), and [collectors/__init__.py](/home/ray/dev/linkong/planet/backend/app/services/collectors/__init__.py).
- Added country centroid helpers in [countries.py](/home/ray/dev/linkong/planet/backend/app/core/countries.py) so country-level prefix geography can produce map coordinates instead of only labels.
- Added a dedicated prefix-geography implementation note in [prefix-geography-plan.md](/home/ray/dev/linkong/planet/docs/earth/prefix-geography-plan.md).
- Added a dedicated prefix-geography implementation note in [prefix-geography-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-prefix-geography-plan.md).
- Added recent `15m` collector activity dimensions to BGP coverage output in [bgp_collectors.py](/home/ray/dev/linkong/planet/backend/app/services/bgp_collectors.py) and [visualization.py](/home/ray/dev/linkong/planet/backend/app/api/v1/visualization.py).
- Added additional BGP detector coverage for `route_leak_candidate` and `path_flap` flows in [test_bgp.py](/home/ray/dev/linkong/planet/backend/tests/test_bgp.py).
- Added a local Earth cloud texture at [earth_clouds_1024.png](/home/ray/dev/linkong/planet/frontend/public/earth/assets/earth_clouds_1024.png) to avoid remote cloud-map dependency failures.
@@ -723,7 +1203,7 @@ Released: 2026-04-02
- Improved Earth event animation semantics in [bgp.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/bgp.js) by separating icon pulse from ring expansion so the center marker can breathe while the ring expands independently.
- Improved Earth texture reliability in [earth.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/earth.js) by switching clouds back to a local static asset under the restored `public/earth` runtime.
- Improved frontend boot noise in [frontend/index.html](/home/ray/dev/linkong/planet/frontend/index.html) by removing the default Vite favicon request that was generating irrelevant `vite.svg` timeouts during Earth debugging.
- Improved project planning docs in [bgp-context.md](/home/ray/dev/linkong/planet/docs/earth/bgp-context.md) and [TODO.md](/home/ray/dev/linkong/planet/TODO.md) so the roadmap now explicitly prioritizes `activity layer`, `prefix-centric geography`, and follow-up geofeed/whois work.
- Improved project planning docs in [bgp-context.md](/home/ray/dev/linkong/planet/docs/technical/earth-bgp-context.md) and [TODO.md](/home/ray/dev/linkong/planet/TODO.md) so the roadmap now explicitly prioritizes `activity layer`, `prefix-centric geography`, and follow-up geofeed/whois work.
### Fixed
@@ -881,7 +1361,7 @@ Released: 2026-03-31
- Added restart-task Redis helpers and whitelist command mapping in [system_control.py](/home/ray/dev/linkong/planet/backend/app/services/system_control.py).
- Added detached restart runner orchestration in [system_restart_runner.py](/home/ray/dev/linkong/planet/backend/scripts/system_restart_runner.py).
- Added `-d` / `--database` support to [planet.sh](/home/ray/dev/linkong/planet/planet.sh) for database-only restarts.
- Added restart control documentation in [system-service-control.md](/home/ray/dev/linkong/planet/docs/backend/system-service-control.md).
- Added restart control documentation in [system-service-control.md](/home/ray/dev/linkong/planet/docs/technical/backend-system-service-control.md).
### Improved

View File

@@ -15,3 +15,8 @@
- 明确写明“已完成”的计划,优先归档
- 已被正式实现替代、继续放在 `docs/` 根目录会误导后续开发的计划,归档
- 仍然指导未来开发、尚未完成或仍有明确执行价值的文档,继续保留在 `docs/`
补充说明:
- 一部分归档文档来自外部或临时工作流草案,例如 sisyphus 生成的初稿
- 这类文档如果有可用内容,应先吸收到 `docs/plans/``docs/technical/`,再归档保留来源记录

View File

@@ -1,3 +1,5 @@
> Archived note: this document was originally created by sisyphus and later reviewed against the main docs set. Useful content has been absorbed into `docs/plans/` where appropriate.
# 地球3D可视化架构重构计划
## 背景

View File

@@ -1,3 +1,5 @@
> Archived note: this document was originally created by sisyphus and later reviewed against the main docs set. Useful content has been absorbed into `docs/plans/` where appropriate.
# 卫星预测轨道显示功能
## TL;DR

View File

@@ -1,3 +1,5 @@
> Archived note: this document was originally created by sisyphus and later reviewed against the main docs set. Useful content has been absorbed into `docs/plans/` where appropriate.
# UE5 3D 大屏客户端开发计划
## 项目概述

View File

@@ -1,3 +1,5 @@
> Archived note: this document was originally created by sisyphus and later reviewed against the main docs set. Useful content has been absorbed into `docs/plans/` where appropriate.
# WebGL Instancing 卫星渲染优化计划
## 背景

39
docs/plans/README.md Normal file
View File

@@ -0,0 +1,39 @@
# Plans Docs
这里放“未来实施方案和未完成计划”的文档,重点回答:
- 我们准备做什么
- 为什么要做
- 分几期做
- 当前差距和下一步是什么
适合放入这里的内容:
- Earth / BGP / 地形 / 天球实施方案
- AI Playground 发展计划
- backend / datasource / agent roadmap
- UE5 MVP 方案
当前重点入口:
- [earth-mobile-drawer-ui-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-mobile-drawer-ui-plan.md)
- [earth-compute-center-bgp-style-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-compute-center-bgp-style-plan.md)
- [earth-renderer-architecture-separation-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-renderer-architecture-separation-plan.md)
- [earth-country-boundary-overlay-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-country-boundary-overlay-plan.md)
- [earth-predicted-orbit-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-predicted-orbit-plan.md)
- [earth-webgl-instancing-satellites-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-webgl-instancing-satellites-plan.md)
- [earth-real-terrain-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-real-terrain-plan.md)
- [earth-news-source-configuration-and-collector-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-news-source-configuration-and-collector-plan.md)
- [frontend-public-docs-site-plan.md](/home/ray/dev/linkong/planet/docs/plans/frontend-public-docs-site-plan.md)
- [frontend-ai-playground-development-plan.md](/home/ray/dev/linkong/planet/docs/plans/frontend-ai-playground-development-plan.md)
- [ue5-mvp-fused-plan.md](/home/ray/dev/linkong/planet/docs/plans/ue5-mvp-fused-plan.md)
不适合放入这里的内容:
- 当前代码结构说明
- 组件现状和实现入口
- 已经落地的技术上下文说明
这些应放入:
- [docs/technical/README.md](/home/ray/dev/linkong/planet/docs/technical/README.md)

View File

@@ -10,9 +10,9 @@ This document connects three existing planning threads into one implementation r
Related documents:
- [aiprovider](/home/ray/dev/linkong/planet/docs/agents/aiprovider.md)
- [datasource-health-plan](/home/ray/dev/linkong/planet/docs/agents/datasource-health-plan.md)
- [agent-architecture-plan](/home/ray/dev/linkong/planet/docs/agents/agent-architecture-plan.md)
- [aiprovider](/home/ray/dev/linkong/planet/docs/technical/agents-aiprovider.md)
- [datasource-health-plan](/home/ray/dev/linkong/planet/docs/plans/agents-datasource-health-plan.md)
- [agent-architecture-plan](/home/ray/dev/linkong/planet/docs/plans/agents-agent-architecture-plan.md)
## Big Picture

View File

@@ -0,0 +1,424 @@
# 自定义 API 数据源与 LLM 映射系统 — 实施计划
**状态**:规划中
**创建日期**2026-04-28
**核心原则**LLM 辅助生成映射配置;生产采集使用确定性转换引擎
## 已确认决策
| 项目 | 决策 |
|-----|------|
| 自定义 API 的定位 | 作为内置数据源的补充入口,不直接等同于 Earth 新功能 |
| LLM 的职责 | 探索未知 API、分析样本 JSON、生成 mapping 草案 |
| 采集时是否调用 LLM | 不调用;采集链路必须确定性、可审计、可复现 |
| 自定义数据如何进入 Earth | 必须映射到已支持的目标 schema或先进入通用数据沉淀 |
| 外部凭证放置位置 | Settings / 外部集成统一管理 provider tokenDataSources 引用 provider profile |
| TimescaleDB | 放入 TODO高频时序数据稳定后再评估迁移 |
---
## 一、背景与问题
当前系统已经有 `datasource_configs`,可以配置自定义数据源的 endpoint、auth、headers、config也已经有部分 collector 会读取这些配置。但这只能解决“怎么请求数据”,还没有解决以下问题:
- API 返回 JSON 后,如何转换成系统已有领域模型。
- 自定义数据源是补充已有能力,还是全新数据沉淀。
- 转换规则由谁生成、谁校验、谁执行。
- 未知数据是否能自动在 Earth 上展示。
- 外部 token 是放在全局配置中心,还是放在每个 datasource 下。
专业做法是把“请求配置”“外部凭证”“目标 schema”“字段映射”“采集执行”拆开
- Settings 管外部集成凭证,例如 AI Provider、BarentsWatch、未来付费 AIS API。
- DataSources 管具体数据源实例,例如 endpoint、调度频率、目标 schema、mapping 版本。
- LLM 只在配置阶段辅助生成 mapping不进入生产采集链路。
- Earth 只消费明确 schema 的数据,不消费任意未知 JSON。
---
## 二、目标架构
### 2.1 自定义 API 数据源生命周期
```mermaid
flowchart LR
A[配置 endpoint/auth/request] --> B[抓取 sample JSON]
B --> C[选择目标 schema]
C --> D[LLM 生成 mapping 草案]
D --> E[确定性 mapping engine 预览]
E --> F[schema validation]
F --> G[保存 mapping version]
G --> H[scheduler 执行 mapped collector]
H --> I[写入目标表或 generic_records]
```
### 2.2 目标 schema 分层
| schema | 用途 | Earth 可视化 |
|-------|------|-------------|
| `vessel_ais` | 船只 AIS 位置、航速、航向、MMSI 等 | 进入船舶图层 |
| `geo_points` | 通用点位数据,包含经纬度、名称、类型、时间 | 进入通用 geo layerTODO |
| `news_events` | 新闻/事件类数据,带时间、地点、摘要、来源 | 复用新闻/事件链路 |
| `compute_centers` | 算力中心、机房、数据中心数据 | 复用算力中心图层 |
| `generic_records` | 未知结构化数据沉淀 | 不直接展示 |
v1 建议优先实现:
- `vessel_ais`
- `geo_points`
- `generic_records`
其他 schema 可先在 registry 中预留名称,但不承诺完整落库与可视化。
### 2.3 LLM 的边界
LLM 可以做:
- 根据 API 文档或 sample JSON 解释字段含义。
- 推荐目标 schema。
- 生成 mapping JSON 草案。
- 给出字段置信度和需要人工确认的字段。
- 帮用户发现分页、数组路径、时间字段、坐标字段。
LLM 不应该做:
- 在正式采集时参与每批数据转换。
- 生成并执行 Python/JavaScript 代码。
- 接触 API key、bearer token、basic auth password。
- 自动创建新的 Earth 图层或数据库表。
---
## 三、后端实施计划
### Phase 1 — Target Schema Registry
新增代码级 registry统一描述系统支持的目标数据类型。
每个 target schema 至少包含:
- `key`:例如 `vessel_ais`
- `label`:前端展示名称。
- `description`:适用场景。
- `fields`:字段名、类型、是否必填、说明、示例。
- `validator`Pydantic 或等价校验器。
- `destination`:写入目标,例如 vessel 表、generic_records、future geo layer。
示例概念:
```json
{
"key": "vessel_ais",
"fields": [
{"name": "mmsi", "type": "integer", "required": true},
{"name": "lat", "type": "float", "required": true},
{"name": "lon", "type": "float", "required": true},
{"name": "sog", "type": "float", "required": false},
{"name": "cog", "type": "float", "required": false},
{"name": "received_at", "type": "datetime", "required": false}
]
}
```
### Phase 2 — Mapping Template Model
新增 mapping 配置持久化表,建议命名为 `datasource_mapping_templates`
关键字段:
- `id`
- `datasource_config_id`
- `target_schema`
- `mapping_json`
- `sample_payload_hash`
- `validation_status`
- `version`
- `is_active`
- `created_at`
- `updated_at`
`mapping_json` 是声明式 DSL不允许任意代码执行。
示例:
```json
{
"source": {
"items_path": "$.data.vessels[*]"
},
"fields": {
"mmsi": {"path": "$.mmsi", "type": "integer"},
"lat": {"path": "$.latitude", "type": "float"},
"lon": {"path": "$.longitude", "type": "float"},
"sog": {"path": "$.speedOverGround", "type": "float", "default": null},
"received_at": {"path": "$.timestamp", "type": "datetime"}
}
}
```
### Phase 3 — Deterministic Mapping Engine
实现独立 mapping engine输入 sample/raw payload 和 mapping JSON输出目标 schema 记录。
v1 支持能力:
- JSONPath/JMESPath 风格路径提取。
- 数组展开。
- 默认值。
- 基础类型转换string、integer、float、boolean、datetime。
- 坐标范围校验。
- 简单枚举映射。
- 错误收集:缺字段、类型转换失败、路径不存在。
明确不支持:
- 任意表达式执行。
- 用户提交脚本。
- LLM runtime 修复。
### Phase 4 — LLM Mapping Assistant API
新增配置阶段 API
- `POST /api/v1/datasources/custom/sample`
- 按 datasource 请求配置抓取 sample JSON。
- `GET /api/v1/datasources/target-schemas`
- 返回可选目标 schema 和字段说明。
- `POST /api/v1/datasources/mappings/propose`
- 输入 sample JSON + target schema调用 AI provider 生成 mapping 草案。
- `POST /api/v1/datasources/mappings/preview`
- 使用确定性 mapping engine 预览转换结果。
- `POST /api/v1/datasources/mappings`
- 保存 mapping 版本。
- `PUT /api/v1/datasources/mappings/{id}`
- 更新 mapping生成新版本或覆盖草稿。
- `POST /api/v1/datasources/{id}/run-mapped`
- 手动触发一次 mapped collector。
安全要求:
- `propose` 请求发送给 LLM 前必须脱敏 sample。
- auth headers、token、password 不进入 prompt。
- LLM 返回结果必须再经过 mapping schema 校验。
### Phase 5 — Generic Mapped HTTP Collector
新增通用 collector
- 读取 `DataSourceConfig` 请求配置。
- 读取 active mapping template。
- 拉取 API 数据。
- 使用 mapping engine 转换。
- 使用 target schema validator 校验。
- 调用 destination handler 写入目标表或 generic storage。
- 将失败记录写入错误日志或 dead-letter 结构。
对于 `generic_records`
- 保存 datasource id。
- 保存 target schema。
- 保存 normalized JSON。
- 保存 raw payload 摘要或 raw reference。
- 保存采集时间、source timestamp、mapping version。
---
## 四、前端实施计划
### Phase 1 — Settings 外部集成
Settings 中保留统一外部集成配置:
- AI Providerbase URL、model、API key。
- BarentsWatchclient id/client secret 或 bearer token。
- 未来付费接口AISHub、MarineTraffic、VesselFinder 等 provider profile。
DataSources 不直接管理全局 secret只引用 provider profile。
### Phase 2 — DataSources 自定义源向导
自定义数据源配置改成向导或右侧 drawer
1. Request
- endpoint
- method
- auth profile
- headers
- query/body config
- schedule
2. Sample
- 点击抓取 sample
- 展示 JSON tree
- 支持选择数组根路径
3. Target Schema
- 选择 `vessel_ais``geo_points``generic_records`
- 展示该 schema 必填字段
4. Mapping Proposal
- 调用 LLM 生成 mapping 草案
- 显示字段匹配置信度
- 标出需要人工确认的字段
5. Preview
- 用确定性 engine 预览前 N 条转换结果
- 展示校验错误
6. Save & Enable
- 保存 mapping version
- 启用调度或仅保存草稿
### Phase 3 — 运维视图
为 mapped datasource 展示:
- 上次运行时间。
- 成功记录数。
- 失败记录数。
- 当前 mapping version。
- 目标 schema。
- 最近错误。
- 手动运行按钮。
---
## 五、数据库与存储策略
### v1继续使用 PostgreSQL
PostgreSQL 可以承载当前规模的采集、关系查询、JSONB 沉淀和基础时序查询。v1 不必因为“时序数据”立刻引入 TimescaleDB。
适合继续用 PostgreSQL 的场景:
- 数据量可控。
- 最近状态查询为主。
- 历史保留窗口较短。
- 查询模式还没稳定。
- 需要快速迭代 schema 与 mapping。
### TODOTimescaleDB
以下条件满足后,再评估 TimescaleDB
- AIS、遥测、轨迹类数据达到高频持续写入。
- 需要按时间窗口做聚合、降采样、retention policy。
- 单表时间序列查询明显成为瓶颈。
- 历史轨迹保留从 24h 扩展到数周或数月。
候选迁移对象:
- `vessel_position`
- future telemetry tables
- future generic time-series records
备选方案:
- PostgreSQL 原生按天/月分区。
- TimescaleDB hypertable。
- 热数据 PostgreSQL冷数据对象存储。
---
## 六、安全与治理
### Secret 管理
- Settings 中保存 provider credentials。
- API 返回配置时必须 mask secret。
- LLM prompt 只能包含脱敏 sample 和 schema 说明。
- 后续 TODO引入字段级加密或 KMS。
### Mapping 治理
- 每次 mapping 变更保留版本。
- active mapping 只能有一个。
- 允许保存 draft mapping。
- 运行记录关联 mapping version。
- 校验失败不能自动启用。
### 错误处理
常见错误类型:
- API 401/403凭证错误或过期。
- API 429限流需要调整 schedule。
- JSON path 不存在:上游结构变化。
- 类型转换失败mapping 规则错误。
- schema validation failed转换结果不满足目标模型。
每次运行需要记录:
- datasource id。
- mapping version。
- started_at / finished_at。
- fetched count。
- mapped count。
- written count。
- failed count。
- error summary。
---
## 七、测试计划
### Backend Unit Tests
- mapping engine
- path 提取。
- 数组展开。
- 默认值。
- 类型转换。
- datetime parse。
- 枚举映射。
- 缺字段错误。
- target schema registry
- `vessel_ais` 必填字段校验。
- `geo_points` 经纬度范围校验。
- `generic_records` 接受未知结构。
- LLM assistant
- mock provider 返回 mapping。
- 验证 secret 不进入 prompt。
- 验证非法 mapping 被拒绝。
### Backend Integration Tests
- sample JSON -> propose mapping -> preview -> save mapping。
- mapped collector 使用保存的 mapping 写入 `generic_records`
- `vessel_ais` sample 写入船舶相关目标结构。
- 上游 JSON 结构变化时,运行失败并记录错误。
### Frontend Tests
- 自定义数据源向导完整流程。
- 未配置 AI Provider 时,提示去 Settings 配置,但允许手写 mapping。
- LLM 返回不完整 mapping 时Preview 阶段显示校验错误。
- 保存 mapping 后展示 active version 和运行状态。
---
## 八、分期工作量
| 阶段 | 内容 | 估算 |
|-----|------|------|
| Phase 0 | 完成本规划、确认 schema registry 设计 | 0.5 天 |
| Phase 1 | target schema registry + mapping template model | 12 天 |
| Phase 2 | deterministic mapping engine | 23 天 |
| Phase 3 | sample/propose/preview/save API | 23 天 |
| Phase 4 | DataSources 自定义源向导 | 35 天 |
| Phase 5 | generic mapped collector + run history | 24 天 |
| Phase 6 | vessel_ais / geo_points destination handler | 24 天 |
---
## 九、当前差距与下一步
当前差距:
- `datasource_configs` 只描述请求配置,不描述目标 schema 和 mapping。
- 自定义源没有 sample -> schema -> mapping -> preview -> save 的闭环。
- 生产采集还没有通用 mapped collector。
- Settings 与 DataSources 的职责边界需要在 UI 上进一步明确。
- Earth 还没有通用 `geo_points` 图层。
下一步建议:
1. 先实现 target schema registry 和 mapping engine不急着接 LLM。
2. 用固定 sample JSON 做 `vessel_ais``generic_records` 的单元测试。
3. 再接 LLM propose API让 LLM 产出的只是 mapping 草案。
4. 最后做前端向导,把人工确认和 preview 放到启用之前。

View File

@@ -17,7 +17,7 @@ It is an aggregation/view-model layer:
## Why This Layer Exists
Current product gap from [bgp-context.md](/home/ray/dev/linkong/planet/docs/earth/bgp-context.md):
Current product gap from [bgp-context.md](/home/ray/dev/linkong/planet/docs/technical/earth-bgp-context.md):
- incident density is naturally low
- anomaly density is higher, but still not enough to keep the globe expressive all the time
@@ -290,7 +290,7 @@ Each feature should include:
## Earth Rendering Plan
Detailed visual layering guidance is expanded in [bgp-earth-rendering-plan.md](/home/ray/dev/linkong/planet/docs/earth/bgp-earth-rendering-plan.md).
Detailed visual layering guidance is expanded in [bgp-earth-rendering-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-bgp-earth-rendering-plan.md).
### Layer Relationship

View File

@@ -112,6 +112,285 @@
- [Three.js SpriteMaterial](https://threejs.org/docs/pages/SpriteMaterial.html)
## 天球背景资源与星体数据来源
为避免把“视觉背景”和“可计算天体位置”混为一谈,本方案明确分成两类资源:
### 1. 背景资源:全天星图贴图
用于 Phase 1 的“真实天空背景”。
推荐优先来源:
- NASA SVS 的 Tycho 全天星图
- [The Tycho Catalog Skymap - Version 2.0](https://svs.gsfc.nasa.gov/3572/)
- NASA Deep Star Maps 2020
- SatelliteMap.space 在 credits 中明确提到其使用了 `NASA Deep Star Maps 2020 - High-resolution star field (1.7 billion stars from Gaia DR2)` 作为星空视觉资源
- 这说明行业内成熟实现并不一定直接渲染全部星表点,而很可能先使用一张高质量官方深空星图作为背景层
- 如需后续替换,也可评估 ESA / Gaia 的全天 sky map 资源
- [Gaia DR3 stories](https://www.cosmos.esa.int/web/gaia/dr3-stories)
建议要求:
- 使用官方来源或官方衍生可复用资源
- 等距矩形投影equirectangular
- 坐标定义尽量明确为赤道坐标展开
- 分辨率建议至少 `4k`
- 颜色不要过亮,避免压过 Earth HUD 前景
- 尽量优先选择官方天文机构已经生产好的深空图,而不是自行拼接低质量星空纹理
建议本地资源目录:
- `frontend/public/earth/assets/celestial/starmap_equatorial_4k.jpg`
### 2. 位置数据:星表与天体计算
用于 Phase 2+ 的“位置正确的星体”。
推荐来源分两层:
- 太阳、月亮位置
- 使用 [Astronomy Engine](https://github.com/cosinekitty/astronomy)
- 恒星位置
- 第一优先Hipparcos / Tycho
- [Hipparcos overview](https://www.cosmos.esa.int/web/Hipparcos)
- [Hipparcos catalogues](https://www.cosmos.esa.int/web/hipparcos/catalogues)
- 第二优先Gaia
- [Gaia DR3 stories](https://www.cosmos.esa.int/web/gaia/dr3-stories)
建议策略:
- V1背景球壳只用全天星图不立即生成全量恒星点
- V2只挑选亮星例如星等 `< 5.5`)生成恒星点层
- V3如果确实需要更丰富的星场再逐步扩展到更深星等
这样做的原因:
- 背景球壳负责“天球真实感”
- 亮星点负责“位置正确、可后续标注和高亮”
- 不需要一开始就处理数十万甚至数百万颗星
### 3. 对外部成熟实现的参考结论
`SatelliteMap.space` 的公开 credits 提供了一个很有价值的参考样板:
- 图形渲染使用 `TWGL.js`
- 天文计算使用 `Skyfield``Astronomia`
- 星空/天球视觉资源使用 `NASA Deep Star Maps 2020`
这给本项目的启发是:
- “真实感强的天球背景”完全可以先依赖官方高质量深空图
- “位置正确的动态天体”则应依赖单独的天文计算链路
- 没有必要在第一版就直接渲染完整星表
因此本项目推荐继续坚持两层拆分:
- 背景层:官方深空图 / 全天星图
- 计算层:太阳、月亮与后续亮星点
## 如何保证星体位置正确
位置正确不是只看“图看起来像”,而是要统一参考系和转换链路。
### 1. 统一坐标基准
本方案推荐统一使用:
- `J2000` 赤道坐标系作为恒星位置基准
原因:
- Hipparcos / Tycho 资料和大量天文可视化都容易映射到该基准
- 太阳、月亮也可以通过 Astronomy Engine 转到同一坐标系
- 这样背景、恒星点、太阳、月亮就能共用一套 sky orientation
### 2. 背景贴图与点位必须使用同一展开逻辑
如果背景球壳使用赤道坐标全天图,那么:
- 亮星点也必须按赤道坐标贴到同一球面方向
- 太阳/月亮 sprite 也必须按赤道坐标转换后落到同一 world-space
否则会出现:
- 背景银河带是对的
- 但太阳/月亮或亮星点飘到不匹配的位置
### 3. RA / Dec 到 Three.js 坐标的落点方式
亮星点和日月方向最终都要转成单位球面向量。
概念步骤:
1. 读取赤经 `RA`
2. 读取赤纬 `Dec`
3. 转成弧度
4. 映射到单位球面向量
5. 再根据 Three.js 当前世界坐标定义做轴向映射
参考公式:
```text
x = cos(dec) * cos(ra)
y = sin(dec)
z = cos(dec) * sin(ra)
```
实际接入 Three.js 时,需要做一次项目内坐标轴校准:
- 验证 `RA = 0h`
- 验证 `RA = 6h`
- 验证北天极
- 验证银河带主方向
然后确定最终的:
- `x/y/z` 对应 Three.js 哪个轴
- 是否需要 `z` 取反
- 是否需要整体再做一个固定 `rotation`
建议把这层显式封装在:
```js
function equatorialToWorldVector(raRad, decRad)
```
不要把轴映射散落在不同模块里。
### 4. 背景球壳与恒星点的关系
推荐最终组合:
- 背景层:全天星图球壳
- 点位层:亮星点
- 动态层:太阳 / 月亮
这样有三个好处:
- 背景层提供密集真实的天空纹理
- 亮星点提供位置正确、可扩展的标注基础
- 太阳/月亮提供与时间相关的真实动态对象
## 数据与资源建议清单
### 推荐首批引入资源
1. 全天星图
- 来源NASA Tycho all-sky map
- 用途:背景球壳纹理
2. 月亮纹理
- 用途Phase 4 月相表现
- 路径建议:
- `frontend/public/earth/assets/celestial/moon_albedo_2k.jpg`
3. 太阳 glow 贴图
- 用途:太阳 sprite halo
- 路径建议:
- `frontend/public/earth/assets/celestial/sun_glow.png`
### 推荐首批数据文件
如果要上亮星层,建议新增一个预处理后的轻量数据文件:
- `frontend/public/earth/assets/celestial/bright-stars.json`
建议字段:
```json
[
{
"id": 32349,
"name": "Sirius",
"raDeg": 101.2875,
"decDeg": -16.7161,
"mag": -1.46,
"colorIndex": 0.00
}
]
```
建议不要在浏览器里直接吞原始 Gaia 大表,而是先离线裁剪成:
- 只保留亮星
- 只保留渲染必需字段
- JSON 或二进制轻量格式
## 资源与数据实施路线
### 路线 A先做可用版本推荐
1. 引入 NASA Tycho 全天图
- 或评估替换为更接近 SatelliteMap.space 路线的 `NASA Deep Star Maps 2020`
2. 实现背景球壳
3. 用 Astronomy Engine 计算太阳/月亮方向
4. 暂不做亮星点
优点:
- 最快见效
- 风险最低
- 就能明显提升天球真实感
### 路线 B在 A 基础上增强
1. 离线生成 `bright-stars.json`
2. 浏览器端渲染亮星点
3. 后续可加:
- 星座线
- 亮星名称
- 特定星体高亮
优点:
- 背景真实感和“位置正确的可交互星体”同时兼顾
## 代码模块建议细化
### 新增模块
- `frontend/public/earth/js/celestial.js`
- 管理天球背景
- 管理太阳/月亮
- 管理亮星层(后续)
- `frontend/public/earth/js/celestial-data.js`
- 资源路径
- 星图方向配置
- 亮星数据加载(后续)
### 建议函数设计
```js
export function initCelestialLayer(scene)
export function updateCelestialLayer(date)
export function setCelestialVisibility(visible)
export function disposeCelestialLayer()
function loadStarMapTexture()
function createSkySphere(texture)
function createSunSprite()
function createMoonSprite()
function getSunEquatorialPosition(date)
function getMoonEquatorialPosition(date)
function equatorialToWorldVector(raRad, decRad)
```
### 推荐后续预处理脚本
如要引入亮星层,建议单独做离线脚本:
- `scripts/build_bright_stars.py`
职责:
- 从 Hipparcos / Tycho 源数据读取
- 过滤亮星
- 生成 `bright-stars.json`
这样浏览器端只消费轻量结果,不承担大表解析成本。
## 分阶段实施
## Phase 1真实天球背景

View File

@@ -0,0 +1,372 @@
# Earth Compute Center BGP-Style Plan
## Goal
这份文档定义如何按照 BGP 模块的产品方式,把“算力中心”提升为 Earth 上的一级能力。
这里的“按 BGP 方式”指的是:
- 有独立的数据语义和接口入口
- 有独立的 Earth 图层与图例
- 有独立的 hover / click / 选中态 / 详情卡
- 有独立的统计口径与后续专题页扩展空间
这里的“按 BGP 方式”不指:
- 机械复制 BGP 的 anomaly / incident / collector 三层事件模型
- 为静态算力设施强行引入不必要的复杂告警语义
算力中心本质上更接近“长期基础设施分布层”,不是“高频动态异常层”。
因此应该复用 BGP 的模块化方法,而不是照搬 BGP 的事件结构。
## Why
当前仓库里已经有算力相关基础:
- 后端已有 `top500``epoch_ai_gpu` 数据采集
- 可视化接口已有 `/api/v1/visualization/geo/supercomputers``/api/v1/visualization/geo/gpu-clusters`
- Earth 信息卡已对 `supercomputer``gpu_cluster` 做了基础类型兼容
但当前能力还停留在“数据可取到”的阶段,没有形成像 BGP 那样完整的可视化模块:
- Earth 缺少独立的算力图层加载模块
- 缺少算力 marker 体系和视觉层级
- 缺少算力图例、统计、开关和搜索接入
- 缺少与海缆、BGP、卫星的关系表达
- 缺少算力专题页和后续告警/研判扩展入口
所以当前真正的缺口不是“有没有数据”,而是“有没有产品级模块”。
## Core Principle
算力中心应当采用和 BGP 一致的模块化分层:
1. 数据层:稳定的数据契约和 GeoJSON 输出
2. 渲染层:独立的 Earth 图层、marker 和视觉状态管理
3. 交互层hover、click、锁定态、详情卡、图例和统计
4. 扩展层:后续专题页、关系分析、告警和 AI 研判
但语义上必须保持算力中心自身的特点:
- `site / center` 是主对象,不是事件
- `capacity / rank / vendor / operator / status` 是主信息,不是异常严重度
- `distribution / concentration / dependency` 是后续分析方向,不是第一阶段必须项
## Recommended Scope
第一版“算力中心”建议统一承载两类对象:
- `supercomputer`
- `gpu_cluster`
并在 Earth 上收口为一个主题层:`compute_centers`
这样做有几个好处:
- 用户看到的是统一的“算力基础设施”语义,而不是零散数据源
- 后端仍可保留 `top500``epoch_ai_gpu` 的来源差异
- 前端可以在一个图层里再细分两种 marker 语言
## Current Gap
和 BGP 对比,当前差距主要在下面几层。
### 1. Data Contract Gap
现在的算力 GeoJSON 还是通用 `collected_data` 输出思路,字段较轻:
- `gpu_cluster` 只有基础名称和地点
- `supercomputer` 只暴露一部分性能字段
- 缺少统一的 `site_type / operator / capacity_band / source / updated_at / confidence`
- 缺少统一的算力层聚合出口
### 2. Earth Rendering Gap
当前 Earth 里没有类似 `bgp.js` 的算力模块:
- `constants.js` 没有算力 API 路径和视觉配置
- `main.js` 没有算力加载、拾取、状态同步和 HUD 更新
- `controls.js` 没有算力图层开关和启动加载优先级
- `layer-startup-tasks.js` 没有算力启动任务
- `legend.js` / `ui.js` 没有算力统计与图例模式
### 3. Interaction Gap
虽然 `info-card.js` 支持基础字段,但还没有形成 BGP 那种完整交互链路:
- 没有 hover / selected / dimmed 的视觉状态
- 没有算力对象专属 tooltip 与摘要文案
- 没有锁定后与其他基础设施的联动高亮
- 没有搜索、统计卡和详情组织方式
### 4. Product Expansion Gap
当前还没有“算力中心”专题页与分析语义:
- 没有全球分布/国家聚合/厂商聚合视图
- 没有算力与海缆/BGP/区域的关系表达
- 没有 AI brief / assessment 的后续落点
## Architecture Direction
推荐把算力中心做成“BGP 同级能力”,但采用更适合静态基础设施的结构。
### Backend
建议新增统一聚合接口,例如:
- `/api/v1/visualization/geo/compute-centers`
它的职责是把:
- `top500`
- `epoch_ai_gpu`
统一转换成一个主题层输出,同时保留对象细分类型:
- `site_type: supercomputer | gpu_cluster`
建议统一字段至少包括:
- `id`
- `name`
- `site_type`
- `country`
- `city`
- `latitude`
- `longitude`
- `operator`
- `vendor`
- `capacity_value`
- `capacity_unit`
- `capacity_band`
- `rank`
- `source`
- `updated_at`
- `location_precision`
- `geography_mode`
- `is_estimated`
- `estimated_reason`
- `metadata`
这里建议优先做“统一聚合出口”,而不是一开始就新增独立数据库表。
原因:
- 当前源数据更新频率低,先复用 `collected_data` 成本更低
- 可以先把 Earth 产品体验做完整
- 如果后续要做历史趋势、关系推断、告警,再评估是否拆成独立模型
### Frontend Earth
建议新增独立模块,例如:
- `frontend/public/earth/js/compute-centers.js`
职责参照 `bgp.js`
- 拉取算力中心 GeoJSON
- 创建 marker
- 管理 hover / selected / dimmed 状态
- 输出图例项
- 输出统计摘要
- 提供 overlay 和详情格式化辅助函数
推荐视觉分层:
1. `supercomputer` 用更稳定、更规整的设施型符号
2. `gpu_cluster` 用更活跃、更现代的密度型符号
3. 选中态通过 halo / ring / related infrastructure highlight 表达
视觉上应避免把算力中心做成“BGP 事件点”那种高频脉冲风格。
它应该更像长期存在的高价值设施。
## Phases
## Phase 1: Unified Earth Layer
目标:
- 先把算力中心做成 Earth 上可用、可点、可解释的一级图层
工作项:
- 新增统一算力 GeoJSON 接口
- 新增 `compute-centers.js`
-`constants.js` 增加 API 路径和视觉配置
-`controls.js` 增加算力图层开关与启动元数据
-`layer-startup-tasks.js` 增加算力启动加载任务
-`main.js` 接入算力拾取、hover、click、锁定态和 HUD 统计
-`ui.js` / `legend.js` / `index.html` 增加算力统计与图例入口
-`info-card.js` 提升算力详情字段组织
- 对无法精确定位、但可按国家或弱线索推测的大概位置,仍然生成地图点位
- 这类对象必须带显式“估算位置”状态,例如图标问号角标与详情说明
完成标准:
- Earth 上能独立显示/隐藏算力中心
- 两类对象有可区分的视觉表达
- hover / click / 详情卡 / 图例 / 统计全部打通
- 精确位置与估算位置在图标或文案上可区分,不会误导为同一精度
- 不干扰现有海缆、卫星、BGP 的交互链路
## Phase 2: Relationship Layer
目标:
- 让算力中心不只是“点”,而是和其他基础设施产生上下文关系
工作项:
- 建立算力中心与国家/区域聚合摘要
- 增加与附近海缆登陆点的关系提示
- 增加与 BGP 事件/观测范围的空间邻近提示
- 增加与卫星覆盖或区域连通性的实验性提示
完成标准:
- 点击算力中心时,用户能看到“它和哪些基础设施相关”
- 信息表达以辅助判断为主,不做夸张推断
## Phase 3: Compute Center Observatory
目标:
- 把算力中心从 Earth 图层扩展成独立专题观测能力
工作项:
- 新增算力中心专题页
- 提供国家/厂商/类型/容量分布统计
- 支持列表、筛选、详情和历史快照
- 预留 AI brief / assessment 入口
完成标准:
- 算力中心不再只是 Earth 上的视觉点位
- 能作为独立业务上下文进入日常观察与研判
## Phase 4: Alerts And Assessment
目标:
- 在不滥造“假动态告警”的前提下,引入真正有价值的变化感知
候选方向:
- 新增大规模算力中心
- 既有中心容量显著变化
- 国家/区域集中度显著变化
- 高价值中心与关键网络基础设施关系变化
完成标准:
- 告警来自可解释的结构变化
- 不把静态数据硬做成噪声式实时事件流
## Implementation Notes
建议按下面顺序推进:
1. 先统一 GeoJSON 契约
2. 再做 Earth 独立模块和图层开关
3. 再补详情卡、图例和统计
4. 最后才做关系层和专题页
这样可以避免一开始把范围摊得过大。
## Unknown Location Strategy
由于部分算力数据源不会直接提供经纬度,未知位置补全不能只依赖“继续找 API 字段”。
更稳妥的方式是做成一条分层富化链路,而不是单一猜测规则。
推荐按下面优先级推进:
1. 直接源信息
- 源记录显式给出 `latitude / longitude`
- 源记录给出 `city / region / facility / campus / operator`
- 源页面详情、内嵌 JSON、结构化元数据、新闻稿链接里能抽出地点线索
2. 名称与机构归一化
- 建立 `canonical_name / aliases / operator / facility` 归一化表
-`cluster name``operator``campus name` 归一到同一个实体
- 优先解决同一对象多写法导致的命中失败,而不是先扩大猜测范围
3. 本地位置注册表
- 用仓库内可维护的 registry 保存高价值对象的位置知识
- 每条记录至少包含:`canonical_name``aliases``operator``country``region``city``lat``lon``confidence``source_note`
- 转换层优先读取 registry避免地点知识长期散落在转换代码里
4. 分层回退定位
- `precise`
- `estimated_site`
- `estimated_city`
- `estimated_region`
- `estimated_national_hub`
- `estimated_country`
这里建议把“国家内主要算力城市”作为国家质心之前的一层。
例如没有美国精确位置时,优先考虑已知的主要算力/数据中心城市候选,而不是直接落在几何质心。
5. 候选证据富化
- 如果源 API 无地点信息,可以允许采集链路读取公开辅助证据
- 例如机构官网、数据中心介绍页、新闻稿、百科型页面、公开 PDF
- 但只提取“地点线索”,不把外部页面上的经纬度当真值直接写回
6. 人工校验闭环
- 对高价值且仍然未知的对象输出待核验清单
- 把人工确认结果回写到位置注册表
- 后续采集继续优先复用这层人工确认结果
### Additional Solution Paths
除了静态映射表,还可以考虑下面这些办法:
- 基于国家和运营方建立“主要园区候选集”,用稳定散列把同国未知节点分散到若干可信城市,而不是全部压到一个点
- 基于数据中心/云厂商公开 region 列表建立 `operator -> city set` 候选映射,用于云 GPU 集群类对象
- 把“估算依据”结构化,例如 `matched_alias``matched_operator``matched_city_text``fallback_country_hub`
- 给位置补全增加 `last_verified_at`,便于后续按时间重新校验老旧映射
- 单独维护“不可可靠定位”状态;这类对象仍可在国家级聚合统计中出现,但可以允许用户在地图上过滤掉
- 后续如果你们愿意投入更多,可把这条链路做成小型 enrichment pipeline而不是仅在 API 转换时临时判断
## Non-Goals
第一阶段不建议做这些内容:
- 不复制 BGP 巡航模式到算力中心
- 不先做复杂实时 websocket 推送
- 不先引入独立 `compute_center_incident` 一类模型
- 不先做全量 AI 分析面板
原因是算力中心的第一需求是“被看清楚”,不是“被实时播报”。
但“被看清楚”不等于“只显示精确坐标对象”。
对于没有精确经纬度、但能推测到国家或区域级位置的算力中心,应优先以上图并标注估算状态的方式处理,而不是直接在地图上消失。
## Acceptance Checklist
- 后端存在统一的算力中心 GeoJSON 出口
- Earth 有独立算力图层模块,而不是散落在 `main.js`
- 页面上有清晰的算力开关、图例和统计
- `supercomputer``gpu_cluster` 在视觉和详情上都可区分
- 估算位置对象在地图和详情中都有明确状态提示
- 现有 BGP / 海缆 / 卫星功能无回归
- 代码结构上为后续专题页和关系分析留出了明确扩展点
## Summary
这项工作的本质不是“再多画几个点”。
它应该把算力中心从已有数据源,升级成与 BGP 同级的 Earth 观测主题:
- 有独立语义
- 有独立图层
- 有独立交互
- 有后续分析扩展能力
推荐先完成 Phase 1把算力中心做成真正可用的 Earth 一级模块,再继续推进关系层和专题页。

View File

@@ -0,0 +1,400 @@
# Earth Mobile Drawer UI Plan
## 背景
当前 Earth 移动端已经补上了基础触控能力,例如:
- 单指拖拽旋转地球
- 双指缩放
- 点击阈值和基础事件隔离
但移动端 UI 仍然存在一个根本问题:
它还在沿用桌面 HUD 的内容切分方式,只是把原来的 panel、modal、toolbar 改位置、改层级、改容器。这样虽然能快速复用旧代码,但手机端体验仍然是生硬的,因为:
- 信息密度和结构是按桌面设计的
- 面板标题、关闭、折叠、开关项是桌面心智,不是手机心智
- 很多内容只是“被塞进抽屉”,而不是为抽屉重新设计
- 设置里仍然带有“显示/隐藏某些 panel”的思路但移动端本来就不应该存在那些独立 panel
因此本计划进一步收紧:
移动端不只是“底部抽屉化”,而是**重新设计一套 fit 抽屉体系的 mobile-first UI**。
## 新目标
1. 手机端不再使用现有 `toolbar` 作为主入口。
2. 手机端不再使用现有独立 `panel / modal / sheet` 作为直接 UI 单元。
3. 手机端统一采用“底部抽屉 + 顶部标题 + tab 切换 + 卡片内容”的单前景模式。
4. 抽屉内部每个 tab 页面都按移动端重新设计内容结构,而不是直接复用旧 panel 结构。
5. 设置页移除“显示/隐藏 panel”的桌面遗留配置。
6. 媒体页拆成两个移动端页面:`新闻``TV`,都归入抽屉体系。
7. 桌面端保持现有 HUD 体系,不回退。
## 核心原则
### 1. 只复用数据和状态,不复用桌面 UI 结构
可复用:
- 图层注册表
- 搜索结果数据
- BGP / 海缆 / 卫星详情数据
- 媒体数据
- 旋转、缩放、选择、高亮等运行时状态
不直接复用:
- 桌面 panel DOM 结构
- 桌面 panel header / close / collapse 交互
- 桌面 settings 项里的“显示某 panel”逻辑
- 桌面媒体面板布局
### 2. 抽屉是唯一主前景层
移动端同一时刻只有一个主前景层:底部抽屉。
抽屉内部切换内容页,而不是多个悬浮层互相覆盖。
### 3. 每个 tab 都是移动端页面,而不是 panel 容器
抽屉中的每一项都应视为一个移动端子页面:
- 有自己的标题
- 有自己的内容层次
- 有自己的滚动区域
- 有自己的主操作
而不是简单挂一个旧面板进去。
### 4. 移动端状态提示不占据屏幕正中
桌面端当前很多通知、状态提示、胶囊消息更适合在屏幕上方居中出现,但移动端不应继续沿用这套布局。
移动端统一改为:
- 通知栏放在右上角安全区
- 胶囊提示放在右上角堆叠
- 不遮挡地球中心视野
- 不与底部抽屉主交互区冲突
## 交互模型
### 默认态
移动端默认只显示:
- 地球主画布
- 底部半露出的抽屉头部
不再单独显示上箭头按钮。
### 展开态
用户从底边直接上拉抽屉,或点击抽屉头部展开。
展开后显示:
- 当前页面标题
- tab 导航
- 当前页面内容
### 收起态
用户下拉抽屉头部收起,或点击背景收起。
## 信息架构
移动端抽屉内的一级页面重定为:
1. 图层
2. 搜索
3. 态势
4. 新闻
5. TV
6. 设置
7. 详情(按需出现,不固定常驻 tab
其中 `新闻``TV` 不再共享同一个移动端媒体面板。
## 页面重设计要求
### 图层页
目标:
- 成为移动端最核心的控制页
- 强调快速开关,不强调桌面 panel 感
内容建议:
- 顶部摘要:当前已启用图层数量
- 图层列表卡片
- 每个图层项只保留:
- 图标
- 中文名
- 英文副标题
- 开关
- 去掉桌面式 header / collapse / close 结构
### 搜索页
目标:
- 成为抽屉中的完整搜索页
- 避免看起来像桌面 modal 被塞进抽屉
内容建议:
- 顶部搜索输入框
- 搜索提示文案
- 结果列表
- 结果项更适合手指点击
- 结果点击后:
- 聚焦地球对象
- 自动切换到详情页
### 态势页
目标:
- 合并原来的 `stats + legend` 思路
- 成为移动端全局态势页
内容建议:
- 顶部核心统计卡
- 海缆数量
- 登陆点数量
- 卫星数量
- BGP 事件数量
- 当前关注层图例
- BGP 状态摘要
- 不再出现独立 legend 面板和独立 stats 面板
### 新闻页
目标:
- 从原媒体面板中拆出单独的移动端新闻页
内容建议:
- 当前区域焦点
- 新闻源数量
- 新闻卡片列表
- 卡片内显示标题、来源、时间、区域
- 外链操作更清晰
### TV 页
目标:
- 从原媒体面板中拆出单独的移动端 TV 页
内容建议:
- 顶部频道选择
- 直播状态
- 当前频道说明
- 视频播放器区域
- 刷新和外链按钮
不再保留桌面式“新闻/TV tab 共处一个 panel”的结构。
### 设置页
目标:
- 只保留对移动端仍有意义的系统配置
必须移除:
- 图层控制 panel 显示/隐藏
- 图例 panel 显示/隐藏
- 全球态势 panel 显示/隐藏
- 媒体 panel 显示/隐藏
保留项建议:
- 旋转模式
- 日夜模式
- 地球默认大小
- 地形透明度
- 系统入口
原因:
移动端已经没有这些独立 panel 了,所以继续保留这些开关会制造错误心智。
### 详情页
目标:
- 成为海缆 / BGP / 卫星对象的统一移动端详情页
内容建议:
- 标题区
- 类型标签
- 关键属性列表
- 相关对象摘要
- 相关图层或态势提示
行为建议:
- 点击对象后自动切入详情页
- 搜索结果点击后也切入详情页
## 阶段重定义
### 阶段 2抽屉壳层
目标:
1. 实现底部抽屉基本壳层。
2. 支持上拉展开、下拉收起、背景点击关闭。
3. `mobile` 模式下隐藏旧 toolbar。
4. `mobile` 模式下不再直接显示旧 panel。
完成标准:
1. 手机端只有地球主视图和抽屉。
2. 抽屉开合稳定。
### 阶段 3基础页面重做
目标:
1. 重新设计并实现图层页。
2. 重新设计并实现搜索页。
3. 重新设计并实现设置页。
完成标准:
1. 这三个页面不再是旧 panel 原样移植。
2. 设置页已移除 panel 可见性开关。
### 阶段 4态势与详情重做
目标:
1. 将 stats 和 legend 合并为新的态势页。
2. 实现统一详情页。
3. 对象点击与搜索结果点击都可切入详情页。
完成标准:
1. 不再存在移动端独立 legend / stats 面板。
2. 详情页成为统一对象信息入口。
### 阶段 5媒体拆分重做
目标:
1. 将原媒体面板拆成两个移动端页面新闻页、TV 页。
2. 分别重做这两个页面的布局。
3. 保留各自必要操作,但不继续共享桌面 panel 结构。
完成标准:
1. 新闻与 TV 各自成为独立移动端页面。
2. 不再使用桌面媒体 panel 的 tab 结构作为移动端主体。
### 阶段 6手感与真机修正
目标:
1. 调整抽屉高度、节奏、手势阈值。
2. 调整 tab 密度与文字层级。
3. 优化 iPhone / Android 安全区。
4. 优化抽屉滚动与地球拖拽边界。
完成标准:
1. 抽屉和地球不会抢手势。
2. 手机端各页面信息层次清晰。
3. 真机下无遮挡、无死层、无错误交互心智。
## 技术落点调整
### [frontend/public/earth/index.html](/home/ray/dev/linkong/planet/frontend/public/earth/index.html)
职责:
- 只保留移动端抽屉壳层
- 为各页面提供新的页面容器
不再把旧 panel 作为最终结构直接塞进抽屉。
### [frontend/public/earth/js/controls.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/controls.js)
职责:
- 管理抽屉开合
- 管理 tab 切换
- 管理详情页切入
- 管理 mobile / desktop 分流
### [frontend/public/earth/js/search.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/search.js)
职责:
- 保留搜索能力和结果逻辑
- 输出给新的移动端搜索页
### [frontend/public/earth/js/info-card.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/info-card.js)
职责:
- 从桌面 info-card 逻辑中提取可复用的数据层
- 服务新的移动端详情页
### [frontend/public/earth/js/tv.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/tv.js)
职责:
- 为新的 TV 页面提供数据和状态
- 不再直接主导移动端媒体 panel 壳层
### [frontend/public/earth/js/news.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/news.js)
职责:
- 为新的新闻页面提供列表和区域焦点数据
### CSS
需要新增真正的移动端页面样式,而不是继续在旧 panel class 上堆条件分支:
- 图层页样式
- 搜索页样式
- 态势页样式
- 新闻页样式
- TV 页样式
- 设置页样式
- 详情页样式
- 移动端右上角通知 / 胶囊提示样式
## 验收标准
1. `mobile` 模式下不再显示旧 toolbar。
2. `mobile` 模式下不再把旧 panel 直接作为最终 UI。
3. 图层、搜索、态势、新闻、TV、设置都是重新设计的移动端页面。
4. 设置页不再包含移动端无意义的 panel 显示/隐藏项。
5. 新闻与 TV 已拆分为两个移动端页面。
6. legend / stats 已整合为态势页。
7. 详情页成为统一对象详情入口。
8. 移动端通知栏和胶囊提示已统一放到右上角安全区,而不是屏幕正中。
## 结论
本计划进一步明确:
移动端目标不是“把桌面 HUD 放进抽屉”,而是“以抽屉为载体,重做一套适合手机端的信息页面”。
后续开发必须以此为准:
- 复用数据
- 重做界面
- 清除桌面遗留心智

View File

@@ -0,0 +1,156 @@
# Earth News Source Configuration And Collector Plan
## Why
当前 Earth 的“态势新闻”由 [earth_news.py](/home/ray/dev/linkong/planet/backend/app/services/earth_news.py) 直接在请求时抓取 RSS / Google News feed再按当前地球视角中心区域聚合返回。
这条链已经可用,但存在两个明显限制:
- 新闻源写死在代码里,不能像 TV 直播源一样从后台维护
- 新闻并未进入统一采集体系,没有采集状态、失败监控、历史数据和后续 AI 复用能力
因此这块更合理的路线不是一步到位重写,而是分阶段推进:
1. 先做“新闻源配置化”
2. 再做“新闻采集器化”
## Current State
当前实现分布在:
- 新闻接口
- [news.py](/home/ray/dev/linkong/planet/backend/app/api/v1/news.py)
- 实时聚合逻辑
- [earth_news.py](/home/ray/dev/linkong/planet/backend/app/services/earth_news.py)
- 前端消费
- [news.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/news.js)
当前新闻源包含:
- `BBC World` RSS
- `DW Top Stories` RSS
- 按区域关键词拼出来的 `Google News RSS`
- `Global`
- `Americas`
- `Europe`
- `Middle East / Africa`
- `Asia Pacific`
当前不是采集器,也不落库,只做内存缓存。
## Phase 1: Source Configuration
### Goal
`NEWS_FEED_SOURCES` 从硬编码列表升级成可配置新闻源目录,但继续保留当前“实时聚合”的工作方式。
### Scope
- 为 Earth news 建立独立配置结构
- 支持后台维护 feed 源
- 支持启用/禁用、优先级、区域、源类型
- 保持现有 `/api/v1/news/earth-feed` 输出协议不变
### Proposed Shape
建议配置字段至少包括:
- `id`
- `name`
- `region`
- `feed_url`
- `homepage_url`
- `source_type`
- `priority`
- `is_enabled`
- 可选 `query_profile`
- 可选 `language`
- 可选 `notes`
### Suggested Storage
优先走系统设置或单独的 news source settings payload而不是先建复杂新表。
推荐原因:
- 改动小
- 易上线
- 和当前 TV settings 维护体验更接近
- 先解决“写死在代码里”的问题
### Non-goals
这一阶段不做:
- 新闻入库
- 新闻历史回看
- 新闻采集任务监控
- 新闻去重流水线
## Phase 2: News Collectorization
### Goal
把“态势新闻”升级为真正的采集器链路,使其进入采集系统和数据层。
### Scope
- 新增专用 news collector
- 按配置源定时采集 RSS / feed
- 做标题/链接级去重
- 建立统一新闻记录模型
- 为 Earth、控制台、AI 研判复用同一份新闻数据
### Benefits
- 有采集状态
- 有失败监控
- 有历史缓存
- 可以做时间轴 / 区域新闻基线
- 可以作为 AI 引用证据
### Required Design Work
需要提前明确:
- 新闻数据模型
- 去重策略
- 过期清理策略
- 区域映射策略
- 聚合排序策略
- 新闻与 Earth 当前视角/区域的关联方式
### Candidate Output Model
至少应包含:
- `source_id`
- `headline`
- `summary`
- `url`
- `publisher`
- `region`
- `published_at`
- `language`
- `tags`
- `raw_feed_source`
- `reference_date`
## Recommended Order
推荐执行顺序:
1. 先完成 Phase 1 配置化
2. 保持 Earth 继续实时聚合,但改为读取配置源
3. 等新闻源稳定后,再设计 Phase 2 的 collector / storage / dedupe
## Decision
当前结论:
- TV 直播源:优先采集器化
- 态势新闻:优先配置化,再采集器化
## Source Note
This plan is newly created for the Planet repo to separate the short-term "configurable source directory" work from the longer-term "collectorized news pipeline" work.

View File

@@ -0,0 +1,98 @@
# Earth Predicted Orbit Plan
> Source note: this plan absorbs useful ideas from a sisyphus-created draft formerly stored at `.sisyphus/plans/predicted-orbit.md`.
## Goal
在 Earth 中锁定卫星时,显示“预测轨道”而不是只有历史尾迹:
- 从当前时刻开始
- 绕地球一圈
- 当前点最亮
- 向后沿轨道逐步衰减
## Current State
当前已经有:
- 卫星历史轨迹
- 锁定卫星
- 轨道高亮与相关联动
但“预测轨道”仍然不是一套稳定、可验证的单独功能计划。
## Why It Is Valuable
预测轨道可以明显提升:
- 锁定卫星后的空间可读性
- 轨道类型辨识
- 演示解释力
相比短历史尾迹,预测轨道更符合用户对“这颗卫星接下来会怎么走”的预期。
## Scope
### Phase 1
- 锁定卫星时显示一整圈预测轨道
- 解锁时隐藏
- 不替代现有普通轨迹系统
### Phase 2
- 根据轨道类型调整采样率
- GEO / MEO / LEO 不同密度
- 进一步减少 fallback 轨迹的比例
## Implementation Direction
### 1. Orbit period
基于 `meanMotion` 估算轨道周期。
### 2. Predicted samples
以固定采样步长从 `now -> now + period` 推算轨迹点。
### 3. Render object lifecycle
预测轨道应是一个独立渲染对象:
- show
- update
- hide
- dispose
### 4. Visual semantics
预测轨道不应与普通尾迹混淆:
- 更稳定
- 更完整
- 透明度沿轨道衰减
- 当前点附近更亮
## Known Risks
### 1. TLE propagation gaps
部分卫星可能出现 SGP4 计算不足,需要 fallback。
### 2. Multiple orbit lines
必须确保:
- 锁定切换前先清旧轨道
- 页面隐藏/销毁时清理
### 3. Performance
GEO 轨道点数高,采样率需要按轨道类型分层。
## Acceptance
1. 锁定单颗卫星时只显示一条预测轨道
2. 解锁后轨道立即清除
3. 不同轨道类型下点数可控
4. 页面切换回来不会闪出旧轨道残留

View File

@@ -0,0 +1,472 @@
# Earth Real Terrain Plan
## Goal
将 Earth 页当前的“程序噪声假地形”替换成基于真实 DEM 的可用地形层,使 `地形 terrain` 开关真正显示全球海拔起伏,而不是占位效果。
当前占位实现位于:
- [earth.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/earth.js)
具体问题:
- `createTerrain()` 直接对球体顶点应用 `simplex noise`
- 没有真实海拔数据来源
- 没有分辨率分层
- 没有和当前相机/视角配套的性能控制
## Constraints
本计划必须贴合当前 Earth 架构,而不是引入一套全新的地形引擎:
- 地球主体仍然是一个 Three.js sphere
- 海缆、登陆点、卫星、BGP 都已经建立在当前球体坐标系之上
- 不能为了地形把整页改成 Cesium/MapLibre Globe 之类的全栈替换
- 第一阶段优先做“真实可用”,不是一步到位做摄影测量级地形
## Recommended Data Source
### Primary recommendation
使用公开的 Terrarium 编码高程瓦片作为浏览器端高度来源,第一阶段优先接入:
- Mapzen/AWS `Terrarium` elevation tiles
参考:[Mapzen terrain tile format / Terrarium](https://www.mapzen.com/blog/terrain-tile-service/)
原因:
- 已经是全球瓦片化高程
- 浏览器端按 tile 请求,最适合当前 Earth 这种在线 globe
- 编码简单稳定:
- `heightMeters = (R * 256 + G + B / 256) - 32768`
- 不需要我们先离线拼整球 DEM
### Data quality upgrade path
如果后面第一阶段效果确认可用,再逐步升级到底层源:
- Copernicus DEM GLO-30
参考:[Copernicus DEM docs](https://documentation.dataspace.copernicus.eu/APIs/SentinelHub/Data/DEM.html)
- 或用 Copernicus / SRTM / ASTER 等离线切成我们自己的 terrain tiles
这条升级路径适合第二阶段,不建议一开始就直接自建全球瓦片服务。
## Why Not Replace the Engine
不建议为了地形直接切到 Cesium terrain / quantized mesh 引擎,原因:
- 现有 Earth 业务对象都依附当前球面坐标
- 切引擎会同时波及:
- 海缆绘制
- 卫星/轨迹
- BGP 标记
- HUD 与交互
- 这是“重做一页”,不是“给地形层接真实数据”
所以推荐路线是:
- 保持当前 sphere globe
- 为 sphere 增加真实高度位移层
## Implementation Strategy
分三期推进。
### Phase 1 — Global Heightmap Terrain Overlay
目标:
- 地形层切换后显示真实海拔起伏
- 全球范围可用
- 性能可控
做法:
1. 新增 terrain 数据模块
建议文件:
- `frontend/public/earth/js/terrain.js`
职责:
- 选择 DEM zoom level
- 请求 Terrarium tiles
- 解码 tile 高程
- 将高程重采样到当前地形球体网格
2. 替换 `createTerrain()`
当前:
- 在 [earth.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/earth.js) 中同步生成噪声地形
调整后:
- `createTerrain()` 只负责创建 terrain mesh 骨架
- 真正的顶点位移由 terrain 模块异步注入
3. 第一阶段采用“整球低分辨率位移”
不要一上来做动态 patch stitching。第一阶段更稳的办法是
- 保留一张全球 terrain sphere
- 使用较低分辨率几何
- 例如 `SphereGeometry(radius, 192, 192)``256/256`
- 运行时按一个固定地形 zoom`z=4``z=5`)抓取覆盖全球的 Terrarium tiles
- 将 tile 解码后重投影到经纬度采样网格
- 将每个球面顶点按真实高度抬升
这样第一阶段就能做到:
- 有真实地形
- 不需要复杂的局部 LOD
- 不会让现有球体对象体系爆炸
### Phase 2 — View-Aware Refinement
目标:
- 正面可见区域更精细
- 背面与远处维持低成本
做法:
- 引入“基础全球地形 + 当前视角高分局部补丁”
- 正面区域额外抓更高 zoom 的高程 tile
- 只替换局部顶点位移或局部 overlay mesh
这一阶段适合在第一阶段稳定后做。
### Phase 3 — Normals / Shading / Terrain UX
目标:
- 地形不仅有起伏,还更好看、更可读
包括:
- 根据高度生成更合理的 normals
- 调整 terrain material使山脉/高原更易读
- 可选加入:
- hillshade
- contour lines
- snowline / bathymetry tint
## Calibration Overlay Before More Terrain Tuning
在当前项目里terrain 看起来“不像真地形”,不一定只是 DEM 或 exaggeration 不够,也可能是因为缺少稳定参照物。
没有清晰的海岸线、国界线和地表分层时,人眼很难判断:
- 山脉是不是在应该高的地方高
- terrain 是否真的贴在正确的大陆位置上
- 地球纹理、本初子午线、terrain 采样之间是否存在偏移
这里要明确区分两件事:
- 国界线不会修好错误的 terrain
- 但海岸线 / 国界线会让我们更容易判断 terrain 有没有贴准
所以在继续盲调 terrain 参数之前,建议先插入一个“校准参照层”阶段。
### Recommended order for the calibration layer
1. 海岸线
2. 国界线
3. 再继续调 terrain
原因:
- 海岸线比国界线更基础,也更接近真实地表边界
- 判断 terrain 是否贴准,最重要的是大陆边缘和山脉/海岸关系
- 国界线更多是政治边界,只能作为辅助参照
如果只加国界线,不加海岸线,效果仍然可能会怪,因为:
- 很多国界线本来就是人为直线
- 它们并不总是跟真实地形走
### Suggested layer order during debugging
建议调试期临时把地球层次明确成:
1. base earth texture
2. coastline / borders overlay
3. terrain relief
4. cables / landing points / bgp / satellites
这样会比现在更容易判断:
- 山脉是否位于正确区域
- terrain 是否和地表对齐
- 国界/海岸是否漂移
### Suggested data source for the calibration overlay
优先用 `Natural Earth` 的轻量全球矢量数据:
- 海岸线coastline
- Admin 0 国界线country borders
优点:
- 全球一致
- 轻量
- 很适合当前 Three.js globe 做 overlay
### Recommended execution path
#### Phase A — Add reference overlays
先加两层可开关的参考线:
- 海岸线
- 国界线
这两层的目标不是最终美术表现,而是调试 / 校准。
#### Phase B — Recalibrate terrain against coastline
有了海岸线以后,再重新看 terrain
- terrain 是否和大陆边缘错位
- 地球纹理、本初子午线、terrain 采样之间是否有固定偏移
#### Phase C — Decide whether to keep the current terrain path
这时再决定后面的路线:
- 如果发现真实高程整体是对的,只是缺少 shading / readability
继续保留当前 DEM + terrain overlay 路线
- 如果发现整球采样投影、本初子午线或 overlay 关系本身就很别扭
再考虑重做 terrain pipeline
### Practical recommendation
当前阶段不建议“从头开始重做 terrain”。
更稳的策略是:
- 暂停继续盲调 terrain 参数
- 先补海岸线 / 国界线作为校准参照层
- 再基于参照层判断 terrain 是“参数没调好”,还是“整条实现路径有偏移”
## Recommended Geometry Model
### First usable model
保留一层独立 terrain sphere
- base earth sphere贴纹理、昼夜、海洋
- terrain sphere略高于地球半径真实高程位移
建议:
- `terrainBaseRadius = CONFIG.earthRadius + 0.2`
- 高度缩放使用真实米制换算,再乘一个可调 exaggeration
示例关系:
- `heightWorld = (elevationMeters / 6371000) * CONFIG.earthRadius * exaggeration`
建议第一阶段 `exaggeration = 1.3 ~ 1.8`
因为完全真实比例在全球球体上会太平,看不出来。
## Tile Decoding Plan
### Terrarium decode
对于每个高程 tile 像素:
```text
heightMeters = (R * 256 + G + B / 256) - 32768
```
### Sampling path
对于 terrain mesh 上每个顶点:
1. 将顶点方向转成经纬度
2. 将经纬度映射到 Web Mercator tile 坐标
3. 找到对应的 tile 和像素
4. 解码高程
5. 将顶点沿法线方向抬升
### Needed helpers
建议新增:
- `latLonToTileXY(lat, lon, z)`
- `tilePixelFromLatLon(lat, lon, z, tileSize)`
- `decodeTerrariumHeight(r, g, b)`
## Caching Strategy
为了不让地形开关每次重开都重新抓全量 tile
- terrain tile 按 `z/x/y` 存到内存缓存
- terrain mesh 结果也缓存一份
- 当用户关闭/开启 terrain
- 直接复用已有位移结果
建议:
- `Map<string, Float32Array | ImageBitmap>`
## Material Strategy
第一阶段不要复杂化。
建议 terrain material
- 半透明低饱和地形色
- 比 base earth 稍亮或稍偏冷
- 保留当前 HUD 风格下的可读性
第一阶段不需要:
- 真实土地覆被纹理
- 独立卫星影像贴 terrain
因为那会和现有地球纹理、云层、昼夜 shader 打架。
## Integration Points
### Files to change
- [earth.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/earth.js)
- 重写 `createTerrain()`
- 删除 simplex noise 占位逻辑
- [main.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/main.js)
- 初始化 terrain 数据加载
- 控制 terrain readiness / loading message
- [controls.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/controls.js)
- `toggleTerrain` 逻辑保持,但应能区分:
- mesh 已就绪
- 正在加载
- [constants.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/constants.js)
- 新增 `TERRAIN_CONFIG`
- 新文件:
- `frontend/public/earth/js/terrain.js`
### Suggested new config
建议新增:
```js
export const TERRAIN_CONFIG = {
enabled: true,
tileSize: 256,
baseZoom: 4,
baseRadiusOffset: 0.2,
exaggeration: 1.5,
opacity: 0.55,
color: 0x6c876f,
maxConcurrentRequests: 8,
cacheEnabled: true,
};
```
## Loading UX
地形第一次开启时,不能像现在一样瞬时切换。
建议:
- 如果地形数据尚未准备:
- 顶部状态条显示:`正在加载真实地形数据...`
- 完成后:
- `真实地形已就绪`
如果加载失败:
- 保留 base earth
- 显示轻量错误提示
- 不要让 terrain 开关卡死在“开”状态
## Risks
### 1. Global tile count too high
即使 `z=5` 全球 tile 数也不少。
缓解:
- 第一阶段限定低 zoom
- 并发上限
- 缓存
### 2. Mesh resolution too low
如果球面分段太低,山脉会被抹平。
缓解:
- 第一阶段先选一个中等分辨率
- 用 exaggeration 保证可见性
### 3. Existing overlays may z-fight with terrain
海缆、登陆点、BGP、卫星相关对象都假设地球半径固定。
缓解:
- terrain sphere 单独作为 overlay
- overlay 保持略低或略高的固定 offset
- 必要时局部调整 landing point / cable altitude offset
### 4. Mercator sampling distortion near poles
Web Mercator 在高纬会有失真。
缓解:
- 第一阶段接受
- 后续若需要更严格极区质量,再上 geodetic reprojection pipeline
## Acceptance Criteria
第一阶段完成后,应满足:
1. `地形 terrain` 开关开启时,地表起伏明显不再是随机噪声
2. 喜马拉雅、安第斯、落基山、东非高原等全球大尺度地形可辨认
3. 关闭/重新开启 terrain 不重复全量请求
4. 不破坏:
- 海缆
- 卫星
- BGP
- 地球昼夜
- 天球层
## Suggested Execution Order
1. 引入 `TERRAIN_CONFIG`
2. 新建 `terrain.js`
3. 实现 Terrarium tile 请求与 decode
4. 用低 zoom 全球 tile 构建真实 terrain sphere
5. 接管 `toggleTerrain()`
6. 调整 terrain material 和高度 exaggeration
7. 做缓存
8. 再考虑第二阶段局部高分 refinement
## Source References
- Mapzen Terrarium / AWS terrain tiles
[Mapzen Terrain Tile Service](https://www.mapzen.com/blog/terrain-tile-service/)
- Terrarium tile experiments / format background
[mapzen/terrarium](https://github.com/mapzen/terrarium)
- Copernicus DEM overview
[Copernicus DEM docs](https://documentation.dataspace.copernicus.eu/APIs/SentinelHub/Data/DEM.html)
## Recommendation Summary
如果现在就要开始做,我建议直接按这条路线开工:
- 第一阶段接入 Terrarium 全球高程 tile
- 替换掉当前 simplex 假地形
- 先做一层真实可见的全球 terrain overlay
- 等第一阶段稳定,再做视角高分 refinement
这是对当前项目风险最低、最贴合现有 Earth 架构的一条路。

View File

@@ -0,0 +1,111 @@
# Earth Renderer / Logic Separation Plan
> Source note: this plan absorbs useful ideas from a sisyphus-created draft formerly stored at `.sisyphus/plans/earth-architecture-refactor.md`.
## Goal
将 Earth 前端继续往“逻辑层 / 状态层 / 渲染层”分离推进,降低后续这几类工作的耦合成本:
- Three.js 渲染重构
- 部分图层替换实现
- 未来 UE / Cesium 客户端迁移
- Earth 行为逻辑复用
## Why This Matters
当前 Earth 已经有一些良好分层,例如:
- 图层显隐入口
- Cable state 枚举与状态 map
- 交互逻辑与实际视觉效果的部分分离
但还没有形成一套更明确的统一规则。现在的风险是:
- 同一类对象的 hover / locked / hidden / loading 语义不一致
- 状态和渲染更新散落在多个模块
- 后续再加新图层时容易复制旧逻辑
## Target Architecture
Earth 对每类对象都尽量拆成三层:
1. `state layer`
- 保存对象状态
- 例如:`normal / hovered / locked / hidden / loading`
2. `logic layer`
- 处理点击、悬停、锁定、过滤、显隐切换
- 不直接关心 Three.js 具体材质怎么改
3. `renderer layer`
- 根据状态更新 Three.js / HUD 外观
- 是最容易针对不同渲染引擎替换的一层
## Current Good Signals
当前已经接近这条方向的地方:
- cable 状态管理
- 部分 landing point 状态同步
- layer button 的统一状态入口
- tooltip / legend / info-card 开始朝状态驱动靠拢
## Next Steps
### 1. Standardize object state enums
优先为这些对象建立更稳定的状态语义:
- cables
- satellites
- landing points
- BGP markers
- media / news 面板入口按钮
### 2. Unify state-to-visual adapters
为各模块建立更清晰的渲染适配函数,例如:
- `applyCableVisualState()`
- `applySatelliteVisualState()`
- `applyBGPVisualState()`
要求:
- 逻辑层只改状态
- 视觉层负责把状态映射到材质、透明度、发光、尺寸、文字
### 3. Separate Earth UI state from render state
HUD / 面板 / 图层按钮状态也需要和渲染状态分离:
- `loading`
- `active`
- `locked`
- `hidden`
- `error`
不要再让 UI 通过“猜渲染结果”推导业务状态。
### 4. Prepare migration-safe boundaries
后续如果做 UE / Cesium 客户端,尽量保留:
- 状态枚举
- 交互规则
- 数据层接口
只替换:
- Three.js 具体渲染实现
- HUD 展示实现
## Practical Rule
后续 Earth 新功能开发时,优先问三个问题:
1. 这个状态由谁持有?
2. 这个交互逻辑在哪一层处理?
3. 这个视觉变化是否能在不改逻辑的情况下单独替换?
如果答不上来,就说明还在把状态、逻辑、渲染揉在一起。

View File

@@ -0,0 +1,272 @@
# 实时船只监控系统 — 实施计划
**状态**:规划中
**创建日期**2026-04-27
**优先数据源**BarentsWatch免费→ AISHub / MarineTrafficTODO付费
## 已确认决策
| 项目 | 决策 |
|-----|------|
| 数据源 | BarentsWatch 先行AISHub / MarineTraffic TODO |
| 船只规模 | BarentsWatch 阶段全部显示;全球数据接入后按需加船型过滤(默认 Cargo + Tanker + Passenger |
| 更新频率 | 准实时:前端 5 分钟轮询,后端 Collector 每分钟拉取写库 |
| 历史轨迹 | 保留(`vessel_position` 表保留 24h后期按需扩展 |
| 推送方式 | HTTP 轮询(不用 WebSocket换实时数据源后再评估升级 |
---
## 一、技术背景
船只通过 AIS自动识别系统每 210 秒广播位置、航速、航向、目的地等信息。全球约 50 万艘持证船只在线,实时数据通过以下方式获取:
| 来源类型 | 典型服务 | 覆盖范围 | 成本 | 状态 |
|---------|---------|---------|------|------|
| **BarentsWatch Open API** | live.ais.barentswatch.no | 挪威海域实时 | 完全免费 | **当前使用** |
| **AISHub** | aishub.net | 全球实时 | 免费/小额 | TODO付费接入 |
| **MarineTraffic API** | marinetraffic.com | 全球实时 | $50$500/月 | TODO评估 tier |
| **VesselFinder API** | vesselfinder.com | 全球实时 | $50$300/月 | TODO备选 |
| **自建 SDR 接收** | RTL-SDR + AIS-catcher | 仅本地 3050km | 硬件 $30 | 不考虑 |
| **NOAA 历史数据** | Marine Cadastre | 美国近海历史 | 免费 | 可用于冷启动 |
### BarentsWatch API
- 端点:`https://live.ais.barentswatch.no/v1/latest/combined`
- 无需注册,直接 GET返回挪威近海 20005000 艘船只 JSON
- 字段mmsi, lat, lon, sog, cog, heading, nav_status, name, vessel_type, flag
- 刷新频率:数据约 3060s 更新一次,可随意轮询
### TODO付费数据源接入
- [ ] 评估 AISHub 订阅(全球覆盖,约 $30/月),接入全球实时流
- [ ] 评估 MarineTraffic API tier对比 AISHub 数据质量与成本
- [ ] 实现多数据源适配器,通过 `datasource_config` 切换
- [ ] 真实高频 AIS 稳定接入后,评估将 `vessel_position` 迁移为 TimescaleDB hypertable保留 Postgres 原生分区作为备选)
---
## 二、实施计划
### Phase 0 — 数据源验证与链路打通12 天)
- 接入 BarentsWatch Open API验证数据格式与字段
- 构建全球 mock 数据生成器(用于前端渲染压测,补充 BarentsWatch 的地域限制)
- 确认前端可渲染船只点,整条链路走通
### Phase 1 — 后端基础设施34 天)
#### 1.1 数据库 Schema
```sql
-- 船只静态信息(每 6h 刷新一次)
CREATE TABLE vessel_static (
mmsi BIGINT PRIMARY KEY,
name VARCHAR(128),
callsign VARCHAR(16),
vessel_type SMALLINT,
vessel_type_name VARCHAR(64),
flag VARCHAR(4), -- ISO 国家码
length FLOAT,
width FLOAT,
draught FLOAT,
imo BIGINT,
updated_at TIMESTAMPTZ
);
-- 船只实时位置(高频写入,保留 24h 轨迹)
CREATE TABLE vessel_position (
id BIGSERIAL PRIMARY KEY,
mmsi BIGINT NOT NULL,
lat FLOAT NOT NULL,
lon FLOAT NOT NULL,
sog FLOAT, -- Speed over ground
cog FLOAT, -- Course over ground
heading SMALLINT, -- 真北航向
nav_status SMALLINT, -- 0=航行 1=锚泊 5=停靠 ...
received_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
);
CREATE INDEX idx_vessel_pos_mmsi_time ON vessel_position(mmsi, received_at DESC);
CREATE INDEX idx_vessel_pos_time ON vessel_position(received_at DESC);
-- 最新位置物化视图(地图渲染主数据源,避免全表扫描)
CREATE MATERIALIZED VIEW vessel_latest AS
SELECT DISTINCT ON (mmsi)
vp.*, vs.name, vs.vessel_type_name, vs.flag, vs.length
FROM vessel_position vp
LEFT JOIN vessel_static vs USING (mmsi)
ORDER BY mmsi, received_at DESC;
CREATE UNIQUE INDEX ON vessel_latest(mmsi);
```
> 后期如需完整历史轨迹查询,迁移 `vessel_position` 到 TimescaleDB 或按天分区。
#### 1.2 CollectorVesselAISCollector
文件:`backend/app/services/collectors/vessel_ais.py`
- 继承 `BaseCollector`,注册到 `collector_registry`
- 轮询间隔3060s由数据源限速决定
- 支持多数据源切换,通过 `datasource_config` 配置 URL + API Key
- 写入逻辑upsert `vessel_latest`append `vessel_position`
- 接入现有调度系统(`scheduler.py`
#### 1.3 API 端点
```
GET /api/v1/visualization/geo/vessels
?bbox=lon_min,lat_min,lon_max,lat_max # 视口裁剪
?type=cargo,tanker,passenger # 船型过滤
?limit=5000
→ GeoJSON FeatureCollectionPoint
GET /api/v1/visualization/vessels/{mmsi} # 单船详情
GET /api/v1/visualization/vessels/{mmsi}/track # 历史轨迹(默认 6h
?hours=6
→ GeoJSON LineString
```
GeoJSON Feature 格式:
```json
{
"type": "Feature",
"geometry": { "type": "Point", "coordinates": [lon, lat] },
"properties": {
"mmsi": 123456789,
"name": "EVER GIVEN",
"vessel_type": 70,
"vessel_type_name": "Cargo",
"flag": "PA",
"sog": 12.4,
"cog": 247.0,
"heading": 245,
"nav_status": 0,
"length": 400,
"received_at": "2026-04-27T10:00:00Z"
}
}
```
#### 1.4 更新机制
**HTTP 轮询**(不使用 WebSocket
- 前端 `setInterval(fetchVessels, 5 * 60 * 1000)` 定期拉取最新快照
- 后端 Collector 每 60s 从 BarentsWatch 拉取并写库,`vessel_latest` 物化视图随时可查
- WebSocket 留给告警/事件驱动场景BGP、系统通知不混入周期性位置刷新
- 换用 AISHub / MarineTraffic 实时流后,届时再评估是否升级为 WebSocket delta push
---
### Phase 2 — 前端渲染34 天)
文件:`frontend/public/earth/js/vessels.js`
#### 2.1 渲染方案
参考现有卫星系统(`satellites.js`)的 InstancedMesh 模式:
- `THREE.InstancedMesh`:每个实例 = 一艘船,矩阵包含位置 + 旋转(朝向 COG
- 行进船:三角箭头图标,朝向 COG 方向
- 静止/锚泊船:圆点图标
- SVG 图标输出到 `frontend/public/earth/assets/icons/vessel-arrow.svg``vessel-dot.svg`
#### 2.2 船型颜色规范
| 船型 | 颜色 |
|-----|------|
| 货轮 Cargo | `#4A90D9` 蓝 |
| 油轮 Tanker | `#E85D04` 橙红 |
| 客船 Passenger | `#06D6A0` 绿 |
| 渔船 Fishing | `#FFD166` 黄 |
| 军舰 Military | `#73797E` 灰 |
| 其他 | `#9B9B9B` 浅灰 |
| 锚泊/停靠 | 降低饱和度 0.4x |
#### 2.3 LOD相机距离细节层次
| 相机距离 | 渲染策略 |
|---------|---------|
| > 400 | 仅渲染 top 1000 艘(按数据新鲜度 + 船型优先级) |
| 200400 | 渲染 top 5000 艘 |
| < 200 | 渲染当前视口 bbox 内全部船只 |
前端根据相机位置动态计算 bbox附加到 API 请求中。
#### 2.4 图层集成
接入现有图层系统,新增"船只"图层项,支持:
- 图层开/关,状态持久化
- 子过滤(按船型选择显示哪类,可在图例或设置面板中配置)
- 与海缆、BGP、卫星层级共存renderOrder 待定,参考现有层级文档)
#### 2.5 Info Card
复用 `showInfoCard` 机制,点击船只弹出:
```
EVER GIVEN 🚢
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
MMSI 123456789
IMO 9811000
旗帜 巴拿马 🇵🇦
船型 散货轮
当前航速 12.4 kn
航向 247°
状态 航行中
目的地 ROTTERDAM
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
[ 查看轨迹 ] [ MarineTraffic ↗ ]
```
#### 2.6 轨迹可视化
点击"查看轨迹" → 请求 `/vessels/{mmsi}/track` → 用 `THREE.CatmullRomCurve3` 渲染插值轨迹线,风格与海缆一致。
---
### Phase 3 — 功能完善23 天)
| 功能 | 说明 |
|-----|------|
| **船只搜索** | 接入现有搜索面板,按名称 / MMSI 搜索 |
| **统计 HUD** | 显示当前在线船只数、各类型分布 |
| **密度热图** | 超低 zoom 时切换为 hex-bin 热力图(避免点云爆炸) |
| **港口标注** | 加载 WorldPorts 数据集,显示主要港口标记 |
| **关键水道监控** | 马六甲、霍尔木兹、苏伊士等高亮 + 流量统计 |
---
### Phase 4 — 性能与生产化23 天)
- `vessel_position` 按天分区7 天自动清理
- TODO真实数据量达到百万级/日后,将 `vessel_position` 升级为 TimescaleDB hypertable配置 retention policy 与压缩策略
- GeoJSON endpoint 用 Redis 缓存 15s
- 若需 bbox 精确查询,引入 PostGIS `geography` + `ST_DWithin`
- InstancedMesh + frustum culling目标 5 万船只 60fps
---
## 三、工作量估算
| Phase | 内容 | 估计时间 |
|-------|-----|---------|
| Phase 0 | 数据源验证、mock | 12 天 |
| Phase 1 | 后端 Schema + Collector + API | 34 天 |
| Phase 2 | 前端渲染InstancedMesh + 图层 + Info Card | 34 天 |
| Phase 3 | 搜索 + 统计 + 轨迹 | 23 天 |
| Phase 4 | 性能优化 + 生产数据源接入 | 23 天 |
| **合计** | | **约 23 周** |
---
## 四、参考资料
- BarentsWatch AIS API 文档https://www.barentswatch.no/en/developer/ais-api/
- MarineTraffic APIhttps://www.marinetraffic.com/en/ais-api-services
- AISHubhttps://www.aishub.net/api
- AIS 导航状态码ITU-R M.1371-5
- 船型编码vessel_typeITU/IMO AIS Message 5 Type and Cargo
- WorldPorts 数据集https://msi.nga.mil/Publications/WPI

View File

@@ -0,0 +1,82 @@
# Earth WebGL Instancing Satellites Plan
> Source note: this plan absorbs useful ideas from a sisyphus-created draft formerly stored at `.sisyphus/plans/webgl-instancing-satellites.md`.
## Goal
把 Earth 卫星渲染从当前方案继续推进到更适合高数量卫星的 instancing 方向,目标是:
- 支持更多卫星
- 降低渲染压力
- 仍然保留当前数据层和交互层
## Why It Matters
当前卫星系统已经具备:
- 数据加载
- 轨迹
- 选择/锁定
- 图例
- 相关区域联动
但当卫星数量持续增加时,渲染层会越来越接近瓶颈。
## Recommended Direction
优先调研并原型验证:
- `InstancedBufferGeometry + custom shader`
而不是一开始就推倒重写成 raw WebGL。
原因:
- 仍能保留 Three.js 主架构
- 更容易渐进迁移
- 比继续堆普通点渲染更有上限
## What Should Stay
尽量保留这些层:
- 卫星数据获取
- 位置计算
- 锁定/悬停逻辑
- legend / info-card / 相关联动
主要替换的是:
- 卫星点渲染实现
- 颜色/大小等实例属性更新方式
## Phases
### Phase 1: Prototype
- 用 instancing 做最小原型
- 先只渲染卫星点
- 不碰轨迹系统
### Phase 2: Integrate
- 接入当前 `satellites.js` 数据层
- 保留当前选择和高亮语义
### Phase 3: Tune
- 调整可视大小
- 调整选中高亮方式
- 评估是否需要分层 LOD
## Risks
1. 透明度排序更复杂
2. Shader 调试成本更高
3. 选中态和 hover 态不能简单复用旧材质逻辑
## Acceptance
1. 在更高卫星数量下保持可接受帧率
2. 不破坏现有锁定/高亮语义
3. 图例、信息卡、相关卫星联动仍然成立

View File

@@ -0,0 +1,793 @@
# Planet 企业级日志系统实施计划
## Goal
把 Planet 当前“能看一点运行输出”的日志能力,升级为一套真正可用、可定位、可纠错、可追责、可演进的企业级日志系统。
这里的“企业级”不是指一上来就接入很重的外部平台,而是指这套系统需要同时满足下面五件事:
1. 排障可用
2. 历史可查
3. 业务可解释
4. 权限操作可追责
5. 出错后能够反向定位到请求、任务、模块和操作者
最终目标不是“把更多 stdout 放到日志页里”,而是建立一套统一的日志契约与落地链路:
- 统一日志字段
- 统一事件命名
- 统一采集入口
- 统一查询视图
- 清晰的实时日志、持久化事件、审计日志分层
## Why
当前仓库已经有一些日志基础,但离真正可用的日志系统还有明显距离。
已有基础:
- 后端运行日志可通过 `/tmp/planet_backend.log` 查看
- 前端开发服务日志可通过 `/tmp/planet_frontend.log` 查看
- AI Provider 可从 Docker 容器读取日志
- Earth 浏览器端关键日志可上报到后端并进入 Redis 缓冲
- 已有 `system_logs` / `audit_logs` 持久化能力
- 管理台已有“系统日志”页面,支持来源、级别、日期、搜索
当前缺口:
- 后端日志仍以 `uvicorn` / 文本输出为主,不是统一结构化事件流
- 不同模块的日志格式不一致,很多地方只有 message没有 event 语义
- 还没有统一的后端 logger 封装与字段注入机制
- 前端虽然能上报错误,但还没有统一 logger API 和统一事件词汇
- Earth 与管理台之间的错误事件还没有形成可串联的事件链路
- 历史持久化还偏点状,很多高价值失败并没有系统性落库
- 系统日志页当前更像“运行输出查看器”,不是“多层日志查询台”
- 审计日志与运行日志尚未形成明确的产品级联动
所以当前真正的问题不是“有没有日志页”,而是:
**当前系统能看见输出,但还不能稳定回答“发生了什么、影响了谁、在哪条链路上坏了、是否已修复、是谁触发的”。**
## Current State
截至 2026-04-23当前代码中的日志相关能力大致如下。
### 1. 日志来源
当前系统日志页主要读取以下来源:
- `backend`
读取 `/tmp/planet_backend.log`
- `frontend`
读取 `/tmp/planet_frontend.log`
- `ai-provider`
读取 Docker 容器日志
- `earth-client`
读取 Redis 缓冲的浏览器端日志
这些来源定义在:
- [backend/app/services/system_logs.py](/home/ray/dev/linkong/planet/backend/app/services/system_logs.py)
### 2. 当前日志读取模型
当前 `read_log_snapshot()` 的职责是:
- 读取某个来源的最近若干行
- 解析基础级别与时间
- 按级别、日期、搜索进行过滤
- 返回用于日志页展示的快照
这个模型适合“运维查看器”,但不适合企业级日志系统,原因是:
- 读取基于文本尾部扫描,不是基于事件模型
- 不同来源的结构粒度完全不同
- 过滤依赖文本解析,准确率有限
- 没有请求、任务、用户、资源、动作等核心关联字段
### 3. 已有持久化能力
当前已经存在两个持久化入口:
- `record_system_log(...)`
- `record_audit_log(...)`
位置:
- [backend/app/services/persistent_logs.py](/home/ray/dev/linkong/planet/backend/app/services/persistent_logs.py)
这说明系统并不是从 0 开始,但也说明当前最大的问题是:
**持久化能力存在,但没有成为统一默认路径。**
### 4. 已有 request_id 基础
当前系统已具备 `request_id` 相关基础,部分持久化能力也会尝试写入 `request_id`
这为后续做:
- 请求链路排障
- 前后端关联查询
- 任务执行追踪
提供了很好的基础。
### 5. 当前日志页定位
当前日志页已经具备:
- 来源切换
- 级别筛选
- 日期筛选
- 搜索
- 文本控制台视图
但它仍然是“单层视图”:
- 上面是筛选器
- 下面是一块文本控制台
它还不是:
- 运行日志 + 事件日志 + 审计日志 的统一入口
- 也没有事件详情、关联跳转、纠错建议、链路追踪能力
## Core Principles
这套日志系统后续必须遵循下面几个原则。
### 1. 分层,而不是混存
日志必须拆成三层:
1. 运行日志
2. 持久化事件日志
3. 审计日志
它们的用途不同,绝不能继续混成一个概念。
#### 运行日志
用于:
- 实时排障
- 观察服务运行状态
- 看 stdout / stderr / exception / collector 输出
特点:
- 数据量大
- 时效性强
- 保留周期短
- 不要求每条都落库
#### 持久化事件日志
用于:
- 记录高价值错误
- 记录关键业务失败
- 支撑历史追溯
- 支撑趋势分析
特点:
- 只持久化有价值事件
- 必须结构化
- 必须有统一 event 命名
#### 审计日志
用于:
- 留痕
- 追责
- 还原高权限操作
特点:
- 必须单独建模
- 不与普通运行日志混用
### 2. 结构化优先
正式日志必须可拆字段,不能长期依赖自由文本。
最低要求至少能拿到:
- `timestamp`
- `level`
- `service`
- `module`
- `event`
- `message`
- `request_id`
- `trace_id`
- `user_id` / `actor`
- `context`
### 3. 事件命名优先于 message 命名
人看的 message 可以变化,但机器查询和跨模块关联必须依赖稳定事件名。
例如:
- `collector.run.started`
- `collector.run.completed`
- `collector.run.failed`
- `earth.layer.load_failed`
- `earth.cruise.route_build_failed`
- `system.restart_task.failed`
- `auth.websocket.invalid_token`
### 4. 查询链路必须可串联
企业级日志系统的核心不是“有很多日志”,而是“能串起来”。
最终一条高价值事件,至少要能回链到下面任意几类对象:
- 某个请求
- 某个任务
- 某个用户
- 某个数据源
- 某个 Earth 模块
- 某个管理动作
### 5. 默认脱敏
日志体系必须明确禁止记录:
- token
- password
- Authorization header
- cookie
- session
- 明文敏感个人信息
并且需要有统一脱敏器,而不是靠调用者自觉。
### 6. “可纠错”不是一句口号
这里的“可纠错”至少包含三层:
1. 日志字段足够解释错误,方便人排查
2. 系统能识别常见错误模式并给出纠偏建议
3. 关键错误支持闭环动作,例如重试、重建索引、重新触发采集、跳转到对应对象
也就是说,这套日志系统最终不只是“告诉你出错了”,而要尽量接近“告诉你为什么出错、怎么修、去哪修”。
## Non-Goals
第一阶段不追求:
- 全量接入 ELK / Loki / Datadog / OpenTelemetry 全家桶
- 做分布式 trace 全链路可视化大屏
- 把所有历史日志都迁进数据库
- 先做特别复杂的规则引擎
第一阶段追求的是:
- 在当前仓库和当前部署方式下,先把基础日志体系做正确
- 再为后续平台化接入预留好接口
## Target Architecture
推荐目标架构如下。
### Layer 1: Runtime Logs
职责:
- 承载后端、前端开发服务、容器输出、浏览器端缓冲事件
- 提供最近窗口内的实时查看能力
来源:
- 文件
- Docker
- Redis 缓冲
- 后续可扩展到 stdout collector
接口:
- `GET /api/v1/system/logs/sources`
- `GET /api/v1/system/logs/{source_id}`
这层继续保留,但需要做结构化增强和来源补强。
### Layer 2: Persistent System Events
职责:
- 只存高价值事件
- 供历史追溯、事件列表、趋势和纠错使用
数据来源:
- 后端关键异常
- 浏览器端关键失败
- 采集器/调度器关键失败
- 业务关键告警与降级事件
接口建议:
- `GET /api/v1/system/events`
- `GET /api/v1/system/events/{id}`
- `POST /api/v1/system/events/{id}/actions/...`(后续)
### Layer 3: Audit Logs
职责:
- 留痕高权限操作
- 记录操作者、对象、结果、请求号
接口建议:
- `GET /api/v1/system/audit-logs`
### Layer 4: Error Intelligence / Triage
职责:
- 对高频错误做归类
- 对已知错误给出解释与建议动作
- 对相同错误进行 fingerprint 聚合
这是“可纠错”能力的关键层。
建议字段:
- `fingerprint`
- `root_cause_type`
- `known_fix_hint`
- `runbook_url`
- `related_resource_type`
- `related_resource_id`
## Canonical Event Model
推荐统一事件字段模型如下。
### Runtime Log Record
```json
{
"timestamp": "2026-04-23T10:15:30Z",
"level": "error",
"service": "backend",
"module": "app.services.scheduler",
"event": "collector.run.failed",
"message": "Collector bgp_news failed",
"request_id": "req_xxx",
"trace_id": "trace_xxx",
"user_id": null,
"actor": null,
"resource_type": "collector",
"resource_id": "bgp_news",
"context": {
"datasource_id": 12,
"exception_type": "TimeoutError"
}
}
```
### Persistent System Event
```json
{
"id": 1024,
"event": "earth.layer.load_failed",
"level": "error",
"source": "earth-client",
"service": "earth",
"module": "cables",
"message": "Failed to load cable layer",
"fingerprint": "earth.layer.load_failed:cables:network_timeout",
"request_id": "req_xxx",
"trace_id": null,
"user_id": 1,
"resource_type": "earth_layer",
"resource_id": "cables",
"category": "visualization",
"status": "open",
"context": {
"url": "/api/v1/visualization/geo/cables"
},
"created_at": "2026-04-23T10:15:30Z"
}
```
### Audit Log
```json
{
"id": 88,
"action": "system.restart_task.requested",
"actor_id": 1,
"actor_name": "root",
"target_type": "restart_task",
"target_id": "restart_20260423_xxx",
"result": "success",
"request_id": "req_xxx",
"ip": "127.0.0.1",
"details": {
"action": "restart_backend"
},
"created_at": "2026-04-23T10:15:30Z"
}
```
## Implementation Plan
## Phase 0: Logging Inventory And Naming Freeze
目标:
- 先统一“记录什么”和“怎么命名”,避免后面越做越乱
工作项:
- 盘点当前所有 `logging.getLogger` 使用点
- 盘点裸 `print`
- 盘点 `record_system_log` / `record_audit_log` 已落点位
- 建立统一事件命名表
- 定义 service / module / category / resource 字段枚举
- 输出日志字段白名单和脱敏规范
完成标准:
- 有一份稳定的事件命名清单
- 有一份字段规范清单
- 后续新增日志不再“临时起名”
## Phase 1: Backend Structured Logging Foundation
目标:
- 把后端从“散落 logging + 文本输出”升级成“统一结构化 logger”
工作项:
- 新增统一后端 logger helper例如 `app/core/logging.py`
- 自动注入:
- `service`
- `module`
- `request_id`
- `trace_id`
- 增加统一脱敏 filter
- 把关键模块先切到统一 logger
- API 层
- scheduler
- collectors
- websocket
- visualization
- system control
- 约束:
- 正式路径禁止裸 `print`
- 正式异常优先 `logger.exception(..., extra={...})`
完成标准:
- 后端关键模块都有稳定 `event`
- request 日志和异常日志能挂上 `request_id`
- 不再依赖只看 `uvicorn` 原生文本输出来定位问题
## Phase 2: Persistent Event Layer
目标:
- 把“值得长期保留的错误和关键事件”系统性落库
工作项:
- 重新定义 `record_system_log()` 的使用边界
- 明确哪些事件必须持久化:
- API 关键失败
- 调度器失败
- 采集器失败
- Earth 客户端关键错误
- 数据源不可用
- 业务降级与恢复
- 补齐字段:
- `event`
- `resource_type`
- `resource_id`
- `category`
- `fingerprint`
- `status`
- 增加高频错误去重/聚合策略
完成标准:
- 高价值错误不再只存在于运行日志里
- 能查询最近一周/一月的关键失败事件
- 相同错误具备聚合基础
## Phase 3: Frontend And Earth Unified Logger
目标:
- 把前端从“点状 error 上报”升级成统一前端事件流
工作项:
- 在前端新增统一 logger API
- 统一方法:
- `debug`
- `info`
- `warn`
- `error`
- 统一字段:
- `page`
- `module`
- `event`
- `message`
- `url`
- `user_agent`
- `context`
- Earth 模块优先接入:
- layer load failed
- cruise build failed
- popup render failed
- connector render failed
- websocket dropped
- 管理台优先接入:
- settings save failed
- datasource toggle failed
- restart task submit failed
完成标准:
- 前端日志事件名与后端可对齐
- Earth 和管理台关键失败不再只停留在 console
- 浏览器端关键问题能进入统一系统日志/事件层
## Phase 4: Audit Logging Completion
目标:
- 把管理员与高权限操作真正做成企业级审计
工作项:
- 扩大审计覆盖面:
- 系统重启
- 数据源启停
- 调度规则变更
- 配置变更
- 人工触发采集
- 删除/修改关键配置
- 增加字段:
- actor
- target
- before / after
- request_id
- IP
- 审计页支持:
- 动作筛选
- 操作者筛选
- 时间筛选
- 目标对象筛选
完成标准:
- 所有高权限操作都能追到人、时间、对象、结果
## Phase 5: Log Console To Enterprise Observability UI
目标:
- 把当前“系统日志”页升级为真正的多层日志工作台
工作项:
- 将页面拆为三个主视图:
1. 运行日志
2. 关键事件
3. 审计日志
- 运行日志视图:
- 保留大控制台
- 支持来源、级别、日期、搜索
- 关键事件视图:
- 列表化展示高价值事件
- 支持聚合、状态、指纹、对象筛选
- 审计视图:
- 列表化展示管理员动作
- 增加详情抽屉:
- 原始 message
- context
- request_id
- related resource
- recommended action
完成标准:
- 日志页不再只是“终端文本窗口”
- 运维排障、历史追溯、审计留痕三者分层清晰
## Phase 6: Corrective Intelligence
目标:
- 让系统从“能看日志”进化到“能辅助修错”
工作项:
- 引入错误 fingerprint
- 对已知错误配置:
- 根因类型
- 修复建议
- runbook 链接
- 推荐动作
- 支持常见纠错动作:
- 重试采集任务
- 重载配置
- 跳转到对应模块/资源
- 打开相关日志过滤视图
- 高频错误支持聚合与静默窗口
完成标准:
- 已知错误能给出明确建议
- 运维不需要每次都从零猜
## Recommended Module Changes
### Backend
建议新增/增强的模块:
- `backend/app/core/logging.py`
- 统一 logger 封装
- formatter
- filter
- request/trace 注入
- `backend/app/services/persistent_logs.py`
- 扩展字段
- 统一持久化策略
- `backend/app/services/system_logs.py`
- 逐步从“文本尾部查看器”升级为“运行日志聚合器”
- `backend/app/services/log_classification.py`
- 指纹
- 根因分类
- 纠错建议
- `backend/app/api/v1/system_control.py`
- 补充事件 / 审计 / 日志多视图接口
### Frontend
建议新增/增强:
- `frontend/src/lib/logger.ts`
- 统一前端 logger API
- `frontend/src/pages/Logs/Logs.tsx`
- 升级为多层工作台
- `frontend/public/earth/js/...`
- 各 Earth 模块接入统一事件 logger
## Event Naming Convention
建议采用:
`<domain>.<resource>.<action>.<result>`
示例:
- `collector.datasource.run.started`
- `collector.datasource.run.failed`
- `earth.layer.cables.load.failed`
- `earth.cruise.route.build.failed`
- `system.restart_task.requested`
- `system.restart_task.completed`
- `auth.websocket.connect.failed`
- `settings.datasource.priority.updated`
规则:
- 不用自然语言句子
- 不把 ID 塞进 event 名里
- 资源对象通过字段承载,不通过 event 名承载
## Query Model
最终推荐支持的查询维度:
- 时间范围
- level
- source
- service
- module
- event
- request_id
- trace_id
- user_id / actor
- resource_type / resource_id
- category
- fingerprint
- status
- full-text search
## Retention Strategy
推荐保留策略:
- 运行日志:
- 文件 / 容器 / Redis 缓冲保留短周期
- 持久化事件:
- 保留中长期
- 审计日志:
- 长期保留
初版可以先这样:
- 运行日志7 到 14 天
- 关键事件90 到 180 天
- 审计日志180 天以上
后续再根据存储与合规要求调整。
## Security And Compliance
必须落实:
- 敏感字段脱敏
- 前端上报白名单
- 防止日志注入
- 审计日志不可被普通管理员随意篡改
- 高敏感纠错动作必须再次鉴权
## Success Criteria
当下面这些条件成立时,才算这套日志系统真的“成了”:
1. 一个后端请求失败时,能通过 `request_id` 在运行日志、持久化事件、审计日志之间串联查询
2. 一个 Earth 前端错误能定位到页面、模块、事件名和上下文
3. 一个采集器失败能同时看到运行日志、持久化事件和可执行纠错动作
4. 一个管理员操作能查到操作者、目标对象、结果和 request_id
5. 日志页不再只是文本控制台,而是完整的“运行日志 / 关键事件 / 审计日志”工作台
6. 高频已知错误能聚合并给出修复建议
## Delivery Order
推荐严格按下面顺序做,不要乱跳:
1. Phase 0 命名与字段规范冻结
2. Phase 1 后端结构化 logging 基础
3. Phase 2 高价值事件持久化
4. Phase 3 前端 / Earth 统一 logger
5. Phase 4 审计覆盖补齐
6. Phase 5 日志工作台 UI 重构
7. Phase 6 指纹 / 纠错 / runbook
原因:
- 如果不先统一字段和命名,后面 UI 和持久化会越来越乱
- 如果不先做后端结构化基础,前端上报再多也串不起来
- 如果不先补持久化层,就只有“实时可看”,没有“历史可查”
## First Actionable Milestone
如果要从明天就开始做,最合理的第一个里程碑是:
### M1: 让后端关键路径全部拥有统一结构化事件
范围:
- API 请求入口/出口
- scheduler
- collectors
- websocket
- visualization
- system control
交付物:
- 统一 logger helper
- 统一 event naming 表
- 统一 request_id 注入
- 统一脱敏策略
- 关键模块替换完成
完成这个里程碑后Planet 才算真正拥有了“企业级日志系统的地基”。

View File

@@ -30,7 +30,7 @@
- [backend/app/services/ai_client.py](/home/ray/dev/linkong/planet/backend/app/services/ai_client.py)
- [aiprovider/main.py](/home/ray/dev/linkong/planet/aiprovider/main.py)
- [aiprovider/provider_service.py](/home/ray/dev/linkong/planet/aiprovider/provider_service.py)
- [docs/agents/aiprovider.md](/home/ray/dev/linkong/planet/docs/agents/aiprovider.md)
- [docs/technical/agents-aiprovider.md](/home/ray/dev/linkong/planet/docs/technical/agents-aiprovider.md)
### 2. 本地运行与配置打通
@@ -77,7 +77,7 @@
相关文件:
- [docs/frontend/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/frontend/frontend-layout-guidelines.md)
- [docs/technical/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/technical/frontend-layout-guidelines.md)
- [frontend/src/pages/BGP/BGP.tsx](/home/ray/dev/linkong/planet/frontend/src/pages/BGP/BGP.tsx)
## 当前限制

View File

@@ -0,0 +1,97 @@
# Markdown 渲染器完善计划
## 背景
Planet 控制台当前有三类主要 Markdown 使用场景:
- 文档中心:技术文档、计划文档、运行手册。
- AI Playground模型回复、分析结果、代码片段。
- BGP 简报:由系统生成并保存的态势报告。
这些场景都复用 `frontend/src/components/MarkdownRenderer/MarkdownRenderer.tsx`。因此 Markdown 能力应该集中在共享渲染器内完成,页面只负责传入内容、链接转换和布局约束,不能让每篇文档或每个页面手写复制按钮、表格样式、列表样式等交互细节。
## 目标
建设一个稳定、可复用、适合技术文档和 AI 输出的 Markdown 渲染器,优先覆盖常用语法、代码块操作和清晰的阅读样式,并为后续语法高亮、锚点导航、内容安全策略留出接口。
## 成功标准
- 代码块支持 fenced language、语言标签、复制按钮、复制成功状态和横向滚动。
- 常用块语法稳定渲染:标题 1-6、段落、引用、分割线、表格、无序列表、有序列表、任务列表。
- 常用行内语法稳定渲染:链接、自动链接、图片、行内代码、粗体、斜体、删除线。
- 文档中心、AI Playground、BGP 简报继续复用同一个组件,不出现页面级重复实现。
- 样式在普通业务面板和文档中心都有合理表现,文档中心可以通过 `.docs-markdown` 覆盖主题变量。
- 前端 TypeScript build 通过,`git diff --check` 无空白错误。
## 当前实施范围
### 第一阶段:共享渲染器补齐
-`MarkdownRenderer` 内解析 fenced code block 的语言信息。
- 引入 `MarkdownCodeBlock` 子组件,负责语言标签、复制按钮和复制状态。
- 保留现有 `Scrollbar` 横向滚动能力,避免长代码撑破页面。
- 扩展标题渲染到 h1-h6并保留 `getHeadingId` 对文档目录的支持。
- 扩展列表解析,支持 `-``*``+``1.``1)` 和 GitHub 风格任务列表。
- 扩展行内解析,支持图片、自动链接、删除线。
### 第二阶段:样式统一
- 全局 Markdown 样式覆盖业务场景,保持紧凑、清晰、可扫描。
- 文档中心用 `.docs-markdown` 适配主题变量,避免硬编码颜色破坏明暗主题。
- 代码块 toolbar 和 copy button 不依赖具体页面。
- 图片默认响应式展示,避免超出内容区域。
### 第三阶段:验证
- 使用前端 build 验证 TypeScript 和 Vite 构建。
- 使用 `git diff --check` 验证补丁格式。
- 手动检查至少一个文档页中代码块复制按钮、语言标签和表格滚动是否出现。
## 后续增强项
### 语法高亮
当前不新增高亮依赖,避免一次性引入过重运行时代码。后续可以在以下方案中二选一:
- `shiki`:适合文档中心,视觉质量高,但包体和初始化成本更高。
- `highlight.js`:接入简单,覆盖语言广,但样式控制需要额外约束。
建议当文档代码块数量稳定增加后再引入,并做按需加载或懒加载。
### 更完整 CommonMark 支持
当前渲染器覆盖 Planet 常见内容,不追求完整 CommonMark 兼容。后续如果需要完整规范,建议切换到成熟生态:
- `react-markdown`
- `remark-gfm`
- `rehype-sanitize`
- `rehype-slug`
切换前需要评估:链接转换、目录 ID、现有样式、AI 输出安全策略和包体影响。
### 安全策略
目前渲染器不解析原始 HTML这是正确默认值。后续如需支持 HTML必须先明确
- 是否允许用户输入 Markdown。
- 是否需要 HTML 白名单。
- 是否需要 `rehype-sanitize`
- 图片和链接是否需要域名策略。
### 文档页能力
可继续补齐:
- 标题锚点悬浮复制。
- Mermaid 图表。
- 代码块折叠。
- 文档内搜索结果定位到代码块。
- 复制按钮埋点,用于判断文档片段是否真正被使用。
## 维护约束
- Markdown 语法能力优先放在共享渲染器,不在具体文档页面散落实现。
- 文档内容只表达内容,不承载 UI 行为。
- 新增 Markdown 能力必须同时考虑文档中心、AI Playground、BGP 简报三个调用方。
- 不解析原始 HTML除非同步引入明确的 sanitize 策略。
- 与主题相关的样式优先走页面容器变量覆盖,不在组件内写死文档中心颜色。

View File

@@ -0,0 +1,486 @@
# Frontend Public Docs Site Plan
## 目标
新增一个公开访问的 `/docs` 页面,作为 Planet 的开发设计文档与使用手册入口。
这个页面应类似常见开源软件文档站:
- 不需要登录即可访问
-`/earth` 和 admin 后台平级,但视觉和信息架构独立
- 直接整理并展示仓库内 `docs/technical` 的 Markdown 文档
- 支持搜索、分类导航、文档目录和内部跳转
-`docs/technical` 继续作为文档真源,避免页面内容和仓库文档漂移
## 非目标
本阶段不做:
- 后端全文搜索服务
- 数据库驱动的 CMS
- 独立文档构建系统,例如 Docusaurus / VitePress
- 每篇文档单独手写 React 页面
- 用户权限、编辑器、在线保存或评论功能
-`docs/plans``docs/deprecated` 全量公开为正式手册
后续可以再决定是否把 plans / deprecated 做成独立的“路线图 / 历史归档”分区。
## 技术路线
### 推荐方案Markdown 直接渲染
使用 Vite 在前端构建阶段直接加载 `docs/technical/**/*.md`
```ts
const modules = import.meta.glob('../../../docs/technical/**/*.md', {
query: '?raw',
import: 'default',
})
```
这样每篇 Markdown 文件仍然留在仓库文档目录中,`/docs` 页面只是读取、索引和渲染这些文档。
当前项目已经满足主要前提:
- 前端使用 Vite + React
- `frontend/vite.config.ts` 已配置 `server.fs.allow: ['..']`
- 已有 `MarkdownRenderer` 可作为基础
- `docs/technical` 文档数量较少,前端本地搜索足够
### 不推荐方案:每篇文档单独写 React
不建议把每篇文档重写成 `.tsx` 页面,因为:
- 文档会出现两份真源
- 修改技术文档时还要同步 UI 页面
- 计划文档、技术上下文、变量表这类内容天然适合 Markdown
- 后续新增文档的成本会变高
只有当某篇文档需要强交互演示、实时图表或复杂 UI 时,才考虑给该文档补充一个 React 组件扩展。
## 信息架构
### 公开路由
新增:
- `/docs`
- `/docs/:slug`
路由行为:
- `/docs` 默认打开 `docs/technical/README.md`,或打开人工指定的首页文档
- `/docs/:slug` 打开对应技术文档
- 未找到文档时显示 docs 专属 404而不是跳回 admin
- `/docs` 加入 `App.tsx` 的公开路由白名单
### 文档分类
`docs/technical` 中的现有文档整理进以下分组:
#### Overview
- `README.md`
#### Earth
- `earth-frontend-context.md`
- `earth-layer-style-reference.md`
- `earth-render-layer-order.md`
- `earth-satellite-footprint-policy.md`
- `earth-bgp-context.md`
- `earth-news-live-streams-collector-format.md`
#### Frontend
- `frontend-admin-frontend-context.md`
- `frontend-layout-guidelines.md`
#### Backend
- `backend-collectors.md`
- `backend-system-service-control.md`
#### Agents
- `agents-aiprovider.md`
#### Ops
- `ops-docker-compose-buildx-upgrade.md`
### 页面布局
桌面端:
- 顶部:产品名、搜索框、当前文档标题
- 左侧:文档分组导航
- 中间Markdown 正文
- 右侧:当前文档目录,也就是 h2 / h3 anchors
移动端:
- 顶部固定搜索入口
- 导航折叠为抽屉或下拉
- 正文单列显示
- 当前文档目录折叠为“本文目录”
视觉风格:
- 像开源软件 docs 页面,清晰、安静、可长时间阅读
- 不复用 admin 后台的重操作感布局
- 不做 Earth 的沉浸式深色 HUD 风格
- 优先阅读性、扫描效率和代码/表格可读性
## 前端实现设计
### 文件结构
建议新增:
```text
frontend/src/pages/Docs/
Docs.tsx
docs-content.ts
docs-search.ts
docs-slugs.ts
Docs.css
```
可选拆分:
```text
frontend/src/pages/Docs/components/
DocsSidebar.tsx
DocsSearch.tsx
DocsToc.tsx
DocsMarkdown.tsx
```
如果初版代码量不大,可以先保持在 `Docs.tsx` + 少量 helper 文件中,避免过度拆分。
### 文档注册表
创建一个 registry负责将 Markdown 文件路径映射为文档元信息:
```ts
interface DocsEntry {
slug: string
path: string
title: string
group: string
order: number
loader: () => Promise<string>
}
```
slug 规则:
- `docs/technical/README.md` -> `overview`
- `docs/technical/earth-layer-style-reference.md` -> `earth-layer-style-reference`
- 只暴露稳定 slug不暴露本机绝对路径
标题规则:
- 优先读取 Markdown 第一个 `# heading`
- 没有 h1 时用人工 registry title
- 再 fallback 到文件名转换标题
### Markdown 渲染
初版可以复用现有:
- [frontend/src/components/MarkdownRenderer/MarkdownRenderer.tsx](/home/ray/dev/linkong/planet/frontend/src/components/MarkdownRenderer/MarkdownRenderer.tsx)
但建议增强或包装为 docs 专用渲染:
- heading 生成稳定 `id`
- 右侧 TOC 使用同一套 heading 解析结果
- 内部 Markdown 链接转换为 `/docs/:slug`
- 外部链接保留 `target="_blank" rel="noreferrer"`
- 表格横向滚动
- 代码块保留等宽字体和语言标记
- 支持 GitHub 风格的相对文档链接
内部链接转换示例:
- `earth-render-layer-order.md` -> `/docs/earth-render-layer-order`
- `./earth-layer-style-reference.md` -> `/docs/earth-layer-style-reference`
- `/home/ray/dev/linkong/planet/docs/technical/foo.md` -> `/docs/foo`
对非 `docs/technical` 的链接:
- 初版可保留原始链接文本
- 或显示为不可跳转的 repo path
- 后续再扩展为跨文档区导航
### 搜索
初版使用纯前端本地搜索。
索引字段:
- title
- slug
- group
- headings
- markdown 正文纯文本
搜索策略:
- 页面首次加载后异步加载所有 `docs/technical` Markdown
- 生成内存索引
- 用户输入时本地过滤
- 简单打分即可:
- 标题命中权重最高
- heading 命中其次
- 文件名 / slug 命中其次
- 正文命中最低
搜索结果展示:
- 文档标题
- 分组
- 命中的 heading 或正文摘要
- 点击跳转到文档
当前只有 13 篇文档,不需要 Lunr、Fuse 或后端搜索。后续文档数量显著增长时,再考虑引入轻量搜索库。
### 路由接入
修改:
- [frontend/src/App.tsx](/home/ray/dev/linkong/planet/frontend/src/App.tsx)
新增 lazy import
```ts
const Docs = lazy(() => import('./pages/Docs/Docs'))
```
公开路由:
```ts
const publicPaths = new Set(['/', '/earth', '/docs'])
```
注意:`/docs/:slug` 不能只用精确匹配 `Set`
建议改为:
```ts
const isPublicRoute =
window.location.pathname === '/' ||
window.location.pathname === '/earth' ||
window.location.pathname === '/docs' ||
window.location.pathname.startsWith('/docs/')
```
新增 routes
```tsx
<Route path="/docs" element={<Docs />} />
<Route path="/docs/:slug" element={<Docs />} />
```
### 样式
建议独立 `Docs.css`,不依赖 admin 页面布局。
核心样式要求:
- 文档正文最大宽度控制在适合阅读的范围
- 表格横向滚动,不撑破布局
- 代码块横向滚动
- 左侧导航固定或 sticky
- 右侧 TOC sticky
- 移动端隐藏右侧 TOC导航折叠
- 搜索结果浮层或独立面板不遮挡正文阅读
注意:
- 不做营销 hero
- 不做卡片堆叠式首页
- 首页第一屏应直接是文档入口和内容,而不是宣传页
## 实施阶段
### Phase 1基础文档站
目标:
- `/docs` 可公开访问
- 能看到 `docs/technical` 文档列表
- 能打开每篇 Markdown
- 能基本渲染标题、段落、列表、代码块、表格
任务:
- 新增 `Docs` 页面
- 新增 docs registry
- 接入 Vite raw Markdown loading
- 接入 `/docs``/docs/:slug`
- 加入公开路由白名单
- 初版 CSS 布局
验收:
- 未登录访问 `/docs` 不跳转登录
- `/docs/earth-layer-style-reference` 可打开样式参考文档
- `/docs/backend-collectors` 可打开后端采集器文档
- 构建通过:`source ~/.zshrc && bun run build`
### Phase 2搜索与 TOC
目标:
- 支持本地搜索所有 technical 文档
- 当前文档右侧显示目录
- 搜索结果可跳转
任务:
- 实现 heading parser
- 实现 TOC 组件
- 实现 search index
- 搜索结果显示文档标题、分组和摘要
- 当前文档标题与 active nav 高亮
验收:
- 搜索 `Fresnel` 能找到 Earth 图层样式文档
- 搜索 `collector` 能找到 backend collectors
- 点击搜索结果进入对应文档
- 右侧 TOC 点击后滚动到对应 heading
### Phase 3链接清理与文档体验
目标:
- Markdown 内部链接在 docs 站内自然跳转
- 长表格、代码块、绝对路径链接的显示更友好
任务:
- 转换 `docs/technical/*.md` 相对链接
- 转换 repo 内 technical 文档绝对路径
- 外链新窗口打开
- 文件路径链接以代码样式显示
- 增强空状态和 404
验收:
-`docs/technical/README.md` 点击 technical 文档链接进入 `/docs/:slug`
- 不支持的 repo 内路径不会导致前端崩溃
- 外部链接行为正常
### Phase 4文档内容整理
目标:
- `docs/technical` 的首页适合作为公开手册入口
- 每篇文档标题、摘要和分类清晰
任务:
- 检查每篇文档是否有唯一 h1
- 给 README 补公开手册导览
- 必要时补文档摘要
- 保持文档内容仍然服务开发维护,不改成营销语气
验收:
- `/docs` 首页能说明各技术文档用途
- 左侧分类和 README 内容一致
- 没有明显重复、过期或找不到的主入口
## 需要改动的文件
预计新增:
- `frontend/src/pages/Docs/Docs.tsx`
- `frontend/src/pages/Docs/Docs.css`
- `frontend/src/pages/Docs/docs-content.ts`
- `frontend/src/pages/Docs/docs-search.ts`
预计修改:
- `frontend/src/App.tsx`
- `frontend/src/components/MarkdownRenderer/MarkdownRenderer.tsx` 或新增 docs 专用 wrapper
- `docs/technical/README.md`
可选修改:
- `frontend/src/index.css`,只放全局极少量 docs shell reset 时才需要
- `docs/CHANGELOG.md`,实施完成后记录
- `docs/version-history.md`,若进入版本发布流程再更新
## 风险与注意事项
### 构建路径风险
Vite 从 `frontend/src` 读取 `../../../docs/technical/**/*.md` 时,需要确认开发和生产构建都可解析。
缓解:
- 使用相对路径 glob
- 构建验证必须跑 `source ~/.zshrc && bun run build`
- 不使用运行时 `fetch('/docs/...')` 读取仓库文件,避免生产环境缺文件
### Markdown 能力不足
现有 `MarkdownRenderer` 是轻量实现,可能不完整支持所有 GitHub Markdown。
缓解:
- 初版优先覆盖当前 `docs/technical` 实际用到的语法
- 若后续需要脚注、嵌套列表、复杂代码高亮,再考虑引入 `react-markdown` 等依赖
### Bundle 体积
把所有 Markdown 打进前端 bundle 会增加体积。
当前文档数量少,风险可接受。
缓解:
- 使用 lazy page chunk
- Markdown loader 保持异步
- 搜索索引在 `/docs` 页面内初始化,不影响 `/earth` 和 admin 首屏
### 公开内容边界
`docs/technical` 会被公开展示,需要避免包含密钥、内部机器地址、临时方案或不应公开的操作细节。
缓解:
- 实施前快速审阅 `docs/technical`
- 暂不公开 `docs/plans``docs/deprecated`
- 以后如需公开更多文档,先建立 allowlist
## 验收清单
- `/docs` 未登录可访问
- `/docs/:slug` 未登录可访问
- `/docs` 不影响 `/earth`
- 未登录访问 admin 仍然跳登录
- 左侧导航包含所有 `docs/technical` 文档
- 文档按 Overview / Earth / Frontend / Backend / Agents / Ops 分类
- Markdown 表格正常显示并可横向滚动
- 代码块正常显示并可横向滚动
- 搜索可搜索标题、heading 和正文
- 搜索结果点击可跳转
- 当前文档 TOC 可跳转
- 不存在的 slug 显示 docs 404
- `source ~/.zshrc && bun run build` 通过
## 后续增强
- 给文档页面增加复制 heading 链接按钮
- 给代码块增加复制按钮
- 增加“上一页 / 下一页”导航
- 增加最近更新信息
- 从 git metadata 读取文档更新时间
- 引入轻量全文搜索库
- 支持 plans / deprecated 独立分区
- 增加页面内反馈入口

View File

@@ -979,3 +979,37 @@ Content/
如果你按这份方案推进,一期最现实的目标不是“立刻做出完整 UE 大屏”,而是:
**在 14 天左右,做出一个能显示真实地球、能显示超算点、能点击看详情、能接后端的可用 UE 客户端 MVP。**
---
# 附录:来自 sisyphus 草案的补充
> 这部分吸收自一个 sisyphus-created draft原始草案已归档不再单独维护为主计划。
## 1. 项目骨架建议
原草案给过一个更偏“工程初始化”的目录示意,适合拿来做一期的命名参考:
- `Levels/`
- `Blueprints/`
- `Materials/`
- `Widgets/`
- `Source/PlanetAPI/`
- `Source/CesiumIntegration/`
- `Source/Visualization/`
这不是强制结构,但对 UE 初期整理目录很有帮助。
## 2. API 契约意识
原草案有一个很对的提醒:
- 一期虽然可以先走 HTTP
- 但数据模型命名不应只服务于一次性演示
- 后续 WebSocket 接入时,字段设计最好能沿用
所以当前主计划继续建议:
- 先做 HTTP 拉取
- 尽量把 UE 侧数据模型定义清楚
- 不要在蓝图各处散写临时 JSON 字段解析

View File

@@ -0,0 +1,35 @@
# Technical Docs
This directory holds "current implementation and current structure" documentation, focusing on:
- How the code is organized right now
- Where the current entry points are
- How state and components work
- Which implementation boundaries future changes should follow
What belongs here:
- Quickstart and user manual
- Frontend context
- Earth frontend structure
- Earth satellite footprint policy
- Earth render layer order
- Earth layer style property index
- Backend runtime control
- Collector status
- Collection format conventions
## Entry Points
- [quickstart.md](/home/ray/dev/linkong/planet/docs/technical/en/quickstart.md): The shortest path to getting Planet running from scratch
- [manual.md](/home/ray/dev/linkong/planet/docs/technical/en/manual.md): Complete usage guide for the console, `planet.sh`, Earth, and Docs
What does not belong here:
- Incomplete roadmaps
- Future iteration plans
- Large-scale refactor proposals
Those belong in:
- [docs/plans/README.md](/home/ray/dev/linkong/planet/docs/plans/README.md)

View File

@@ -0,0 +1,264 @@
# Data Collectors
## I. System Architecture
```
┌─────────────────────────────────────────────────────────────────┐
│ Data Collection Architecture │
├─────────────────────────────────────────────────────────────────┤
│ │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │
│ │ TOP500 │ │ Epoch AI │ │ HuggingFace │ │
│ │ Collector │ │ Collector │ │ Collector │ │
│ └──────┬──────┘ └──────┬──────┘ └──────┬──────┘ │
│ │ │ │ │
│ └───────────────────┼───────────────────┘ │
│ ▼ │
│ ┌─────────────────────┐ │
│ │ BaseCollector │◄── Base class (unified) │
│ │ run() method │ │
│ └─────────┬───────────┘ │
│ │ │
│ ┌─────────────────┼─────────────────┐ │
│ ▼ ▼ ▼ │
│ ┌───────────┐ ┌───────────┐ ┌───────────┐ │
│ │ fetch() │ │transform()│ │ _save_data│ │
│ │ raw data │ │ transform │ │ save to DB│ │
│ └───────────┘ └───────────┘ └───────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────────┐ │
│ │ CollectedData table│◄── Unified storage │
│ └─────────────────────┘ │
│ │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Scheduler (APScheduler) │ │
│ │ Scheduled tasks: every 4h/6h/12h/1d auto-execute │ │
│ └─────────────────────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────┘
```
## II. Pipeline
```python
# 1. Scheduler triggers (scheduled or manual)
# ↓
# 2. run() executes the full pipeline
async def run(self, db):
# 2.1 Check if collector is enabled
if not collector_registry.is_active(self.name):
return {"status": "skipped"}
# 2.2 Record task start
task = CollectionTask(status="running")
db.add(task)
await db.commit()
# 2.3 FETCH — get raw data (implemented by subclass)
raw_data = await self.fetch()
# 2.4 TRANSFORM — convert to unified format
data = self.transform(raw_data)
# 2.5 SAVE — persist to database
records_count = await self._save_data(db, data)
# 2.6 Record task completion
task.status = "success"
task.records_processed = records_count
await db.commit()
```
**Core file**: `backend/app/services/collectors/base.py`
## III. Collector List
| Collector | Data type | Content | Frequency |
|-----------|-----------|---------|-----------|
| TOP500 | supercomputer | Global supercomputer rankings (compute, performance) | 4 hours |
| Epoch AI | gpu_cluster | GPU compute cluster info | 6 hours |
| HuggingFace Models | model | AI model information | 12 hours |
| HuggingFace Datasets | dataset | Dataset information | 12 hours |
| HuggingFace Spaces | space | Demo applications | 1 day |
| PeeringDB | ixp/network/facility | Internet exchange points / networks / facilities | 1-2 days |
| TeleGeography | submarine_cable | Submarine cable information | 7 days |
## IV. Data Format (stored in CollectedData table)
```python
# Each collector's parse_response() return format
{
"source_id": "top500_1", # Original system ID (required)
"name": "El Capitan", # Name (required)
"description": "System desc...", # Description
"country": "United States", # Country
"city": "Livermore, CA", # City
"latitude": "37.6819", # Latitude (string)
"longitude": "-121.7681", # Longitude (string)
"value": "1742.00", # Performance value (e.g. compute)
"unit": "PFlop/s", # Unit
"metadata": { # Extra data (JSON)
"rank": 1,
"r_peak": 2746.38,
"cores": 11039616
},
"reference_date": "2025-11-01" # Data reference date
}
```
## V. Database Schema
**CollectedData table** (`collected_data`)
| Field | Type | Description |
|-------|------|-------------|
| id | SERIAL | Primary key |
| source | VARCHAR(100) | Data source name (top500, huggingface, etc.) |
| source_id | VARCHAR(100) | Original data ID |
| data_type | VARCHAR(50) | Data type (supercomputer, model, etc.) |
| name | VARCHAR(500) | Name |
| title | VARCHAR(500) | Title |
| description | TEXT | Description |
| country | VARCHAR(100) | Country |
| city | VARCHAR(100) | City |
| latitude | VARCHAR(50) | Latitude |
| longitude | VARCHAR(50) | Longitude |
| value | VARCHAR(100) | Performance value |
| unit | VARCHAR(20) | Unit |
| metadata | JSONB | Extra metadata |
| collected_at | TIMESTAMP | Collection time |
| reference_date | TIMESTAMP | Data reference date |
| is_valid | INTEGER | Whether valid |
**Core file**: `backend/app/models/collected_data.py`
## VI. TOP500 Collector Example (full pipeline)
```python
# 1. fetch() — get HTML from the web
async def fetch(self):
url = "https://top500.org/lists/top500/list/2025/11/"
response = await client.get(url)
return response.text # returns HTML
# 2. parse_response() — parse HTML into unified format
def parse_response(self, html):
soup = BeautifulSoup(html, "html.parser")
table = soup.find("table")
for row in table.find_all("tr")[1:]: # skip header
cells = row.find_all("td")
entry = {
"source_id": f"top500_{cells[0].text}",
"name": cells[1].text.strip(),
"country": cells[2].text.strip(),
"city": "",
"latitude": "",
"longitude": "",
"value": "1742.00",
"unit": "PFlop/s",
"metadata": {
"rank": 1,
"cores": "11340000"
},
"reference_date": "2025-11-01"
}
data.append(entry)
return data
# 3. run() automatically calls _save_data() to save to database
```
**Core file**: `backend/app/services/collectors/top500.py`
## VII. Scheduler
```python
# Register all collectors into scheduled tasks at startup
def start_scheduler():
for name, collector in collectors.items():
if collector_registry.is_active(name):
scheduler.add_job(
run_collector_task,
trigger=IntervalTrigger(hours=collector.frequency_hours),
id=name,
name=name
)
```
| Collector | Frequency |
|-----------|-----------|
| TOP500 | Every 4 hours |
| Epoch AI | Every 6 hours |
| HuggingFace | Every 12 hours |
| PeeringDB | Every 1-2 days |
| TeleGeography | Every 7 days |
**Core file**: `backend/app/services/scheduler.py`
## VIII. Code Files
```
backend/app/services/collectors/
├── base.py # Base class: run() pipeline, _save_data() persistence
├── registry.py # Collector registry
├── scheduler.py # Scheduled task dispatch (APScheduler)
├── top500.py # TOP500 collector
├── epoch_ai.py # Epoch AI collector
├── huggingface.py # HuggingFace collector
├── peeringdb.py # PeeringDB collector
└── telegeraphy.py # TeleGeography submarine cable collector
backend/app/models/
└── collected_data.py # Unified data model
```
## IX. Data Usage
Collected data ultimately:
1. **Visualization** — displays supercomputers, GPU clusters, and submarine cables' geographic positions
2. **Situational analysis** — global compute distribution statistics and growth trends
3. **Alert system** — detects changes to important nodes
## X. Collector Registration
Collectors are automatically registered at application startup:
```python
# backend/app/services/collectors/__init__.py
collector_registry.register(TOP500Collector())
collector_registry.register(EpochAIGPUCollector())
collector_registry.register(HuggingFaceModelCollector())
collector_registry.register(HuggingFaceDatasetCollector())
collector_registry.register(HuggingFaceSpacesCollector())
collector_registry.register(PeeringDBIXPCollector())
collector_registry.register(PeeringDBNetworkCollector())
collector_registry.register(PeeringDBFacilityCollector())
collector_registry.register(TeleGeographyCableCollector())
collector_registry.register(TeleGeographyLandingPointCollector())
collector_registry.register(TeleGeographyCableSystemCollector())
```
**Core file**: `backend/app/services/collectors/registry.py`
## XI. Triggering Collection
### Method 1: Scheduled
At startup, APScheduler automatically creates scheduled tasks based on each collector's `frequency_hours` setting.
### Method 2: Manual API trigger
```bash
# Trigger TOP500 collection
curl -X POST http://localhost:8000/api/v1/datasources/1/trigger \
-H "Authorization: Bearer <token>"
```
**Core file**: `backend/app/api/v1/datasources.py`

View File

@@ -187,7 +187,7 @@ Current reality:
- that is expected, because incidents are aggregated and de-noised
- but incident-first rendering makes the Earth view look too quiet unless there is another always-available activity layer
Implementation detail for the recommended `activity layer` is expanded in [bgp-region-aggregation-plan.md](/home/ray/dev/linkong/planet/docs/earth/bgp-region-aggregation-plan.md).
Implementation detail for the recommended `activity layer` is expanded in [bgp-region-aggregation-plan.md](/home/ray/dev/linkong/planet/docs/plans/earth-bgp-region-aggregation-plan.md).
So the immediate next milestone is:

View File

@@ -0,0 +1,252 @@
# Earth Frontend Context
This document describes the current real structure of the Earth display frontend. The focus is on helping future changes to the HUD, layers, media panel, real terrain, and BGP visualization avoid repeating past structural and state-sync pitfalls.
Related references:
- [rules.md](/home/ray/dev/linkong/planet/rules.md)
- [frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/technical/en/frontend-layout-guidelines.md)
## Current Goal
The Earth frontend is not an ordinary admin page — it is an independent large-screen display frontend. Current product goals:
- Maintain the spatial depth and readability of the globe view
- Keep HUD, layers, media panel, BGP, satellites, cables, and similar elements in a unified interaction model
- Clearly represent states like loading, enabled, hidden, and locked
## Current Entry Point
React route entry:
- [Earth.tsx](/home/ray/dev/linkong/planet/frontend/src/pages/Earth/Earth.tsx)
The current approach is simple:
- The React page only provides a full-screen `iframe`
- The actual Earth application runs at:
- [index.html](/home/ray/dev/linkong/planet/frontend/public/earth/index.html)
Earth frontend is essentially a standalone static application under `public/earth`.
## Current File Layers
### 1. Page Entry and Structure
- [index.html](/home/ray/dev/linkong/planet/frontend/public/earth/index.html)
Responsibilities:
- Base HUD DOM
- Layer panel
- Media panel
- Toolbar
- Settings dialog
- Legacy element ID compatibility
### 2. Main Runtime
- [main.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/main.js)
Responsibilities:
- Globe initialization
- Three.js scene assembly
- Data loading and refresh
- Layer module integration
- Earth-level state synchronization
### 3. Earth Control Layer
- [controls.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/controls.js)
Responsibilities:
- Toolbar interaction
- Layer panel interaction
- Rotation / zoom / layout
- HUD panel drag
- Layer toggle state machine
- Earth settings read, persist, and reset
This is currently the most critical UI control entry point for the Earth frontend.
### 4. UI and Status Messages
- [ui.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/ui.js)
Responsibilities:
- Loading panel
- Status message
- Tooltip / error / cleanup logic
### 5. Globe and Terrain
- [earth.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/earth.js)
- [terrain.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/terrain.js)
Responsibilities:
- Globe sphere, cloud layer, atmosphere
- Real terrain mesh
- Terrain tile fetch, decode, displacement, and shading
### 6. Layer Modules
- [satellites.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/satellites.js)
- [cables.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/cables.js)
- [bgp.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/bgp.js)
- [bgp-cruise-adapter.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/bgp-cruise-adapter.js)
- [compute-centers.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/compute-centers.js)
- [country-boundaries.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/country-boundaries.js)
Each module is responsible for its own:
- Data fetching
- Three.js mesh creation and update
- State tracking (loaded, visible, hover, locked)
- Self-cleanup (dispose on scene destroy)
### 7. HUD Panels and Search
- [hud-panels.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/hud-panels.js)
- [info-card.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/info-card.js)
- [search.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/search.js)
- [legend.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/legend.js)
### 8. Cruise Mode
- [cruise-sequencer.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/cruise-sequencer.js)
- [callout-connector.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/callout-connector.js)
The cruise sequencer handles generic logic: current target, queue order, camera focus, and dwell / hide / switch. Business modules supply target queues and content — they should not contain camera control logic.
### 9. Constants
- [constants.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/constants.js)
All material, layer, satellite, BGP, cable, terrain, celestial, and other style parameters are maintained here. Do not scatter magic numbers in module files.
## Current Style Layers
CSS files in `frontend/public/earth/css/` each correspond to a specific component scope. Do not write global Earth styles into `base.css` unless they genuinely apply to everything.
## Current Layer Toggle State Semantics
### `data-status-target`
Layer toggle buttons use `data-status-target` attributes to link button state to layer state. The state machine in `controls.js` handles:
- `loading`: showing the loading indicator
- `enabled`: layer is active
- `hidden`: layer is hidden
- `error`: layer failed to load
This is the canonical way to synchronize button visual state with actual layer state. Do not maintain separate boolean flags for button display.
## Current Settings Persistence
Earth settings are stored in `localStorage`. The key is typically a namespaced string defined in `constants.js`. `controls.js` handles read, write, and reset.
Settings that affect visual layers (terrain opacity, day/night mode, satellite display style, etc.) are read during initialization and applied immediately.
## Current Terrain Pipeline
1. `terrain.js` creates a sphere geometry with enough segments
2. On load, fetches Terrarium-format elevation tiles from the backend
3. Decodes R/G/B into elevation values
4. Displaces vertex positions radially based on elevation
5. Applies a vertex alpha that fades terrain edges at coastlines
6. Terrain writes to the scene as a mesh above the HD texture layer
When HD texture is off, terrain is temporarily hidden and its state is remembered. When HD texture comes back on, terrain restores its prior visibility.
## Current High-Frequency Risk Points
### 1. Visual State and Business State Out of Sync
The most common class of Earth bugs:
- Button shows "loaded," but layer has no objects rendered
- Button shows "hidden," but objects are still visible
- Loading ended, but button still looks like it hasn't
All future changes must prioritize checking state sync.
### 2. HUD Layout: Check Structure First, Not CSS Patches
Earth HUD has repeatedly experienced:
- Panel compressed to a sliver
- Markdown content clipped
- Tabs/iframe content consumed by `overflow: hidden`
Inspection order:
1. Who is responsible for height
2. Who is responsible for scrolling
3. Which layer is doing the clipping
Do not immediately add `overflow: hidden` or extra wrapper layers.
### 3. Transitional Paths Must Be Closed Off
Earth has gone through multiple rounds of HUD, toolbar, and media panel refactoring, making it easy to accumulate:
- Old helpers
- Old classes
- Old fallback logic
- Deprecated variants
After each major feature is complete, do a cleanup pass.
### 4. Cruise Mode and Business Events Must Not Be Deeply Coupled
The correct boundary:
- The generic cruise layer only knows:
- Current target
- Queue order
- Camera focus
- Dwell / hide / switch
- Business modules only supply:
- Target queues
- Focus coordinates
- Card content
- Highlight / layer side effects
If future cable, satellite, or news cruise is added, do not copy a new set of `main.js` state variables. Instead reuse:
- [cruise-sequencer.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/cruise-sequencer.js)
- [callout-connector.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/callout-connector.js)
- The business adapter pattern from [bgp-cruise-adapter.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/bgp-cruise-adapter.js)
## Recommended Change Approach
For future Earth changes:
1. First identify what you're changing:
- Three.js rendering layer
- HUD structure layer
- Layer state layer
- Panel content layer
2. If involving layer buttons, connect to the unified state machine
3. If involving visibility toggle, check whether tooltip / legend / info-card / lock all close together
4. If involving panel layout, check structure before touching CSS
## Current Boundary with the Console Frontend
The Earth frontend and the console frontend are not the same UI system:
- Console frontend: React + Ant Design workbench
- Earth frontend: native HUD + Three.js display under `public/earth`
Therefore:
- Earth should not directly reuse Ant Table / AppLayout semantics
- The console should not copy Earth HUD animations and glass-layer design language
For console structure, see:
- [frontend-admin-frontend-context.md](/home/ray/dev/linkong/planet/docs/technical/en/frontend-admin-frontend-context.md)

View File

@@ -0,0 +1,224 @@
# Earth Layer Style Property Index
This document records the material, color, opacity, line width, radius offset, and `renderOrder` style properties of all Earth frontend layers. For layer ordering relationships, see [earth-render-layer-order.md](/home/ray/dev/linkong/planet/docs/technical/en/earth-render-layer-order.md).
## Naming Conventions
| Category | Convention | Example |
| --- | --- | --- |
| Global config objects | `*_CONFIG` | `COUNTRY_BOUNDARY_CONFIG` |
| Layer radius offsets | `*AltitudeOffset` / `radiusOffset` | `lineAltitudeOffset`, `GRID_CONFIG.radiusOffset` |
| Opacity | `*Opacity` | `hoverLineOpacity` |
| Render order | `*RenderOrder` | `textureOverlayRenderOrder` |
| Color | `*Color`, hex number or CSS color value | `lineColor`, `colors.supercomputer` |
| Line width | `lineWidth` / `*LineWidth` | `GRID_CONFIG.lineWidth` |
## Earth Base and HD Texture
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Earth base radius | `CONFIG.earthRadius` | `100` | `earth.js:createEarth()` |
| Earth base color | `EARTH_MATERIAL_CONFIG.color` | `0x010609` | `MeshPhongMaterial.color` |
| Earth base emissive | `EARTH_MATERIAL_CONFIG.emissive` | `0x010609` | `MeshPhongMaterial.emissive` |
| Earth base specular | `EARTH_MATERIAL_CONFIG.specular` | `0x1a2d45` | `MeshPhongMaterial.specular` |
| Earth base shininess | `EARTH_MATERIAL_CONFIG.shininess` | `12` | `MeshPhongMaterial.shininess` |
| Earth base opacity | `EARTH_MATERIAL_CONFIG.opacity` | `1` | `MeshPhongMaterial.opacity` |
| HD texture radius offset | `EARTH_MATERIAL_CONFIG.textureOverlayAltitudeOffset` | `0.1` | Standalone HD texture sphere radius |
| HD texture opacity | `EARTH_MATERIAL_CONFIG.textureOverlayOpacity` | `0.88` | HD texture `MeshPhongMaterial.opacity` |
| HD texture renderOrder | `EARTH_MATERIAL_CONFIG.textureOverlayRenderOrder` | `0.96` | `_earthTextureOverlay.renderOrder` |
| HD texture specular | `EARTH_MATERIAL_CONFIG.textureOverlaySpecular` | `0x05080d` | Reduces specular highlight in direct-light areas to avoid blown-out texture |
| HD texture shininess | `EARTH_MATERIAL_CONFIG.textureOverlayShininess` | `4` | Reduces specular concentration |
| HD texture color multiplier | inline | `0xffffff` | `_earthTextureOverlayMaterial.color` |
## Earth Occluder and Day/Night
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Occluder radius factor | `EARTH_MATERIAL_CONFIG.occluderRadiusFactor` | `0.999` | Depth occluder sphere radius |
| Occluder segments | `EARTH_MATERIAL_CONFIG.occluderSegments` | `48` | Occluder geometry segments |
| Occluder renderOrder | inline | `-1` | `occluder.renderOrder` |
| Day/night sun direction | `EARTH_MATERIAL_CONFIG.dayNight.sunDirection` | `{ x: 1, y: 0.2, z: 0.4 }` | Custom day/night shader |
| Night-side minimum brightness | `EARTH_MATERIAL_CONFIG.dayNight.nightFloor` | `0.24` | Shader uniform |
| Day-side boost | `EARTH_MATERIAL_CONFIG.dayNight.dayBoost` | `1.12` | Shader uniform |
| Twilight width | `EARTH_MATERIAL_CONFIG.dayNight.twilightWidth` | `0.2` | Shader uniform |
| Twilight intensity | `EARTH_MATERIAL_CONFIG.dayNight.twilightIntensity` | `0.14` | Shader uniform |
| Twilight color | `EARTH_MATERIAL_CONFIG.dayNight.twilightColor` | `0x4ea0ff` | Shader uniform |
| Night tint color | `EARTH_MATERIAL_CONFIG.dayNight.nightTintColor` | `0x0b1830` | Shader uniform |
| Night tint intensity | `EARTH_MATERIAL_CONFIG.dayNight.nightTintIntensity` | `0.08` | Shader uniform |
## Atmospheric Glow and Clouds
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Inner atmosphere radius factor | `EARTH_MATERIAL_CONFIG.atmosInnerRadiusFactor` | `1.01` | `atmosInnerGeo` |
| Inner atmosphere segments | `EARTH_MATERIAL_CONFIG.atmosInnerSegments` | `64` | `atmosInnerGeo` |
| Inner atmosphere color | `EARTH_MATERIAL_CONFIG.atmosInnerColor` | `[0.25, 0.62, 1.0]` | Shader RGB |
| Inner atmosphere rim power | `EARTH_MATERIAL_CONFIG.atmosInnerRimPower` | `3.2` | Shader rim falloff |
| Inner atmosphere intensity | `EARTH_MATERIAL_CONFIG.atmosInnerIntensity` | `0.18` | Shader alpha multiplier |
| Outer atmosphere radius factor | `EARTH_MATERIAL_CONFIG.atmosOuterRadiusFactor` | `1.016` | `atmosOuterGeo` |
| Outer atmosphere segments | `EARTH_MATERIAL_CONFIG.atmosOuterSegments` | `48` | `atmosOuterGeo` |
| Outer atmosphere color | `EARTH_MATERIAL_CONFIG.atmosOuterColor` | `[0.18, 0.45, 0.9]` | Shader RGB |
| Outer atmosphere rim power | `EARTH_MATERIAL_CONFIG.atmosOuterRimPower` | `5.0` | Shader rim falloff |
| Outer atmosphere intensity | `EARTH_MATERIAL_CONFIG.atmosOuterIntensity` | `0.02` | Shader alpha multiplier |
| Atmosphere blending | inline | `THREE.AdditiveBlending` | `ShaderMaterial.blending` |
| Atmosphere renderOrder | inline | `1` | `atmosInner/Outer.renderOrder` |
| No-HD-texture rim glow color | `EARTH_MATERIAL_CONFIG.rimGlowColor` | `[0.35, 0.65, 1.0]` | Fresnel shell RGB when HD texture is hidden or unavailable |
| No-HD-texture rim glow power | `EARTH_MATERIAL_CONFIG.rimGlowPower` | `3.8` | Shader rim falloff; higher = narrower edge |
| No-HD-texture rim glow intensity | `EARTH_MATERIAL_CONFIG.rimGlowIntensity` | `0.28` | Shader alpha multiplier |
| No-HD-texture rim glow segments | `EARTH_MATERIAL_CONFIG.rimGlowSegments` | `64` | `earth-rim-glow` geometry segments |
| Cloud layer radius offset | `CLOUD_LAYER_CONFIG.radiusOffset` | `3` | Cloud sphere radius |
| Cloud layer segments | `CLOUD_LAYER_CONFIG.widthSegments / heightSegments` | `64 / 64` | Cloud sphere geometry |
| Cloud layer opacity | `CLOUD_LAYER_CONFIG.opacity` | `0.15` | `MeshPhongMaterial.opacity` |
| Cloud texture | `CLOUD_LAYER_CONFIG.textureUrl` | `"./assets/earth_clouds_1024.png"` | Cloud texture map |
| Cloud blending | inline | `THREE.AdditiveBlending` | `MeshPhongMaterial.blending` |
## Land/Ocean Base and Country Borders
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Country border data path | `COUNTRY_BOUNDARY_CONFIG.dataPath` | `"/earth/data/countries-admin0.min.geojson"` | GeoJSON input |
| Ocean fill color | local `OCEAN_HEX` | `0x010609` | Land/ocean base canvas background |
| Land fill color | `COUNTRY_BOUNDARY_CONFIG.landColor` | `0x080f1b` | Land/ocean base canvas land |
| Land/ocean base opacity | `COUNTRY_BOUNDARY_CONFIG.landOpacity` | `1.0` | `MeshBasicMaterial.opacity` |
| Land/ocean base radius offset | `COUNTRY_BOUNDARY_CONFIG.landAltitudeOffset` | `0.08` | `country-land-ocean` radius |
| Land/ocean base renderOrder | `COUNTRY_BOUNDARY_CONFIG.landRenderOrder` | `0.86` | `country-land-ocean.renderOrder` |
| Land/ocean mask size | `landMaskWidth / landMaskHeight` | `2048 / 1024` | Canvas / DataTexture size |
| Country tint color | `COUNTRY_BOUNDARY_CONFIG.tintColor` | `0x0b1830` | Tint when HD texture is off |
| Country tint radius offset | `COUNTRY_BOUNDARY_CONFIG.tintAltitudeOffset` | `0.04` | `country-tint` radius |
| Country tint renderOrder | `COUNTRY_BOUNDARY_CONFIG.tintRenderOrder` | `0.2` | `country-tint.renderOrder` |
| Border line color | `COUNTRY_BOUNDARY_CONFIG.lineColor` | `0x7fc7ff` | Normal border line |
| Border line opacity | `COUNTRY_BOUNDARY_CONFIG.lineOpacity` | `0.58` | Normal border line opacity |
| Border dimmed opacity on hover | `COUNTRY_BOUNDARY_CONFIG.dimmedLineOpacity` | `0.18` | Normal border opacity during hover |
| Border line radius offset | `COUNTRY_BOUNDARY_CONFIG.lineAltitudeOffset` | `0.24` | Normal border line radius |
| Border line renderOrder | `COUNTRY_BOUNDARY_CONFIG.lineRenderOrder` | `2.2` | Normal border line level |
| Border hover color | `COUNTRY_BOUNDARY_CONFIG.hoverLineColor` | `0xff3b1f` | Neon red-orange |
| Border hover opacity | `COUNTRY_BOUNDARY_CONFIG.hoverLineOpacity` | `1.0` | Hover line opacity |
| Border hover radius offset | `COUNTRY_BOUNDARY_CONFIG.hoverAltitudeOffset` | `0.32` | Hover line radius |
| Border hover renderOrder | `COUNTRY_BOUNDARY_CONFIG.hoverLineRenderOrder` | `2.3` | Hover line level |
| Border hover glow opacity | `COUNTRY_BOUNDARY_CONFIG.hoverGlowOpacity` | `0.38` | Glow line opacity |
| Border hover glow line width | `COUNTRY_BOUNDARY_CONFIG.hoverGlowLineWidth` | `3` | Glow `LineBasicMaterial.linewidth` |
| Border hover glow level offset | `COUNTRY_BOUNDARY_CONFIG.hoverGlowRenderOrderOffset` | `0.01` | Glow renderOrder = `2.29` |
| Border hover glow radius offset | `COUNTRY_BOUNDARY_CONFIG.hoverGlowRadiusOffset` | `0.04` | Glow radius = hover radius + 0.04 |
## Real Terrain
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Terrain tile size | `TERRAIN_CONFIG.tileSize` | `256` | Terrarium tile read size |
| Terrain base zoom | `TERRAIN_CONFIG.baseZoom` | `4` | Terrain sampling zoom |
| Terrain geometry segments | `geometryWidthSegments / geometryHeightSegments` | `320 / 320` | Terrain sphere geometry |
| Terrain base radius offset | `TERRAIN_CONFIG.baseRadiusOffset` | `0.16` | Terrain overlays HD texture |
| Terrain exaggeration | `TERRAIN_CONFIG.exaggeration` | `34` | Elevation to world units |
| Terrain land fade height | `TERRAIN_CONFIG.landRevealFadeMeters` | `220` | Vertex alpha for coastline fade |
| Terrain opacity | `TERRAIN_CONFIG.opacity` | `0.68` | `MeshPhongMaterial.opacity` |
| Terrain color | `TERRAIN_CONFIG.color` | `0x8aa884` | `MeshPhongMaterial.color` |
| Terrain emissive | `TERRAIN_CONFIG.emissive` | `0x030704` | Reduces self-emission to preserve terrain shading |
| Terrain specular | `TERRAIN_CONFIG.specular` | `0x344438` | Gives terrain local sheen without boosting HD texture brightness |
| Terrain shininess | `TERRAIN_CONFIG.shininess` | `16` | Tightens terrain highlight |
| Terrain renderOrder | inline | `1.2` | `terrain.renderOrder` |
| Terrain polygonOffset | inline | `factor -1`, `units -1` | Reduces z-fighting near sphere surface |
## Grid Lines
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Grid radius offset | `GRID_CONFIG.radiusOffset` | `0.14` | Grid sphere radius |
| Grid color | `GRID_CONFIG.color` | `0xc0e0ff` | `LineBasicMaterial.color` |
| Grid opacity | `GRID_CONFIG.opacity` | `0.08` | `LineBasicMaterial.opacity` |
| Grid line width | `GRID_CONFIG.lineWidth` | `1` | `LineBasicMaterial.linewidth` |
| Grid renderOrder | `GRID_CONFIG.renderOrder` | `2.05` | Grid level |
| Latitude step | `GRID_CONFIG.latitudeStep` | `15` | Latitude line generation step |
| Longitude step | `GRID_CONFIG.longitudeStep` | `30` | Longitude line generation step |
| Segment sample step | `GRID_CONFIG.segmentStep` | `5` | Grid line sample step |
## Submarine Cables and Landing Points
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Default cable color | `CABLE_COLORS.default` | `0xffff44` | Used when no data color available |
| Cable radius offset | `CABLE_CONFIG.line.altitudeOffset` | `0.2` | Cable line radius |
| Cable line width | `CABLE_CONFIG.line.lineWidth` | `1` | `LineBasicMaterial.linewidth` |
| Cable opacity | `CABLE_CONFIG.line.opacity` | `1.0` | Cable line opacity |
| Cable renderOrder | `CABLE_CONFIG.line.renderOrder` | `1` | Cable line level |
| Landing point radius offset | `CABLE_CONFIG.landingPoint.altitudeOffset` | `0.48` | Aligns with compute center marker height |
| Landing point icon texture size | `CABLE_CONFIG.landingPoint.textureSize` | `256` | Canvas size for solid map-pin icon |
| Landing point icon aspect ratio | `CABLE_CONFIG.landingPoint.iconAspectRatio` | `0.82` | `Sprite.scale.x = height * aspect` |
| Landing point icon anchor | `CABLE_CONFIG.landingPoint.anchorX / anchorY` | `0.52 / 0.276` | `Sprite.center`, aligns pin tip to landing point lat/lon |
| Landing point base scale | `CABLE_CONFIG.landingPoint.baseScale` | `12` | Matches compute center sprite height |
| Landing point color | `CABLE_CONFIG.landingPoint.color` | `0xffaa00` | `SpriteMaterial.color` |
| Landing point opacity | `CABLE_CONFIG.landingPoint.opacity` | `1.0` | `SpriteMaterial.opacity` |
| Landing point renderOrder | `CABLE_CONFIG.landingPoint.renderOrder` | `4.5` | Aligns with compute center surface level |
| Landing point dim brightness | `landingPointVisual.dimBrightness` | `0.62` | Dim state color multiplier |
| Dimmed landing point color | `landingPointVisual.dimmed.colorRGB` | `{ r: 180, g: 116, b: 28 }` | Dim state color; avoids dark base showing through as a dark hole |
| Dimmed landing point emissive | `landingPointVisual.dimmed.emissive` | `0x3a2200` | Dim state weak amber self-emission |
| Dimmed landing point opacity | `landingPointVisual.dimmed.opacity` | `0.78` | Dim state opacity; no longer uses low alpha blending with dark base |
## Satellites, Trails, and Footprints
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Satellite display radius offset | `SATELLITE_CONFIG.displayAltitudeOffset` | `8` | Satellite point position |
| Satellite dot base pixel size | `SATELLITE_CONFIG.dotBaseSize` | `2.8` | Point shader size |
| Satellite backdrop dot scale | `SATELLITE_CONFIG.dotBackdropScale` | `1.28` | Backdrop dot size |
| Satellite dot opacity range | `dotOpacityMin / dotOpacityMax` | `0.7 / 1.0` | Breathing animation |
| Satellite dot breathing speed | `SATELLITE_CONFIG.dotBreathingSpeed` | `0.12` | Dot opacity animation |
| Satellite backdrop renderOrder | inline | `5` | `satelliteBackdropPoints.renderOrder` |
| Satellite dot renderOrder | inline | `6` | `satellitePoints.renderOrder` |
| Satellite trail length | `SATELLITE_CONFIG.trailLength` | `10` | Trail buffer |
| Satellite trail line width | `SATELLITE_CONFIG.trailLineWidth` | `3` | Ribbon shader uniform |
| Selected ring size | `SATELLITE_CONFIG.ringSize` | `0.07` | Hover / locked ring sprite |
| Satellite overlay renderOrder | `SATELLITE_CONFIG.overlayRenderOrder` | `12` | Locked ring / halo / orbit |
| Footprint renderOrder | local `GROUND_FOOTPRINT_RENDER_ORDER` | `3` | Footprint fill |
## Compute Centers
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Compute center radius offset | `COMPUTE_CENTER_CONFIG.altitudeOffset` | `0.48` | Marker position |
| Compute center base opacity | `COMPUTE_CENTER_CONFIG.marker.baseOpacity` | `0.88` | `SpriteMaterial.opacity` |
| Supercomputer marker scale | `COMPUTE_CENTER_CONFIG.marker.supercomputerScale` | `12` | Supercomputer marker |
| GPU cluster marker scale | `COMPUTE_CENTER_CONFIG.marker.gpuClusterScale` | `12` | GPU marker |
| Hover scale | `COMPUTE_CENTER_CONFIG.marker.hoverScale` | `1.16` | Hover state |
| Locked scale | `COMPUTE_CENTER_CONFIG.marker.lockedScale` | `1.22` | Locked state |
| Dimmed scale / opacity | `dimmedScale / dimmedOpacity` | `0.82 / 0.34` | Dim state |
| Supercomputer color | `COMPUTE_CENTER_CONFIG.colors.supercomputer` | `"#38bdf8"` | Marker texture |
| GPU cluster color | `COMPUTE_CENTER_CONFIG.colors.gpu_cluster` | `"#2dd4bf"` | Marker texture |
| Linked color | `COMPUTE_CENTER_CONFIG.colors.linked` | `"#f8fafc"` | Linked state |
| Compute center renderOrder | local `COMPUTE_CENTER_RENDER_ORDER` | `4.5` | Surface facility below satellites |
## BGP Observation
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| BGP event radius offset | `BGP_CONFIG.altitudeOffset` | `2.1` | Anomaly marker |
| BGP collector radius offset | `BGP_CONFIG.collectorAltitudeOffset` | `1.6` | Collector marker |
| Event base scale | `BGP_CONFIG.marker.eventBaseScale` | `6.2` | Anomaly sprite |
| Collector base scale | `BGP_CONFIG.marker.collectorBaseScale` | `7.4` | Collector plane |
| Hover / dim scale | `hoverScale / dimmedScale` | `1.16 / 0.92` | Interaction states |
| Normal event opacity | `BGP_CONFIG.opacity.normal` | `0.78` | Anomaly sprite |
| Hover opacity | `BGP_CONFIG.opacity.hover` | `1.0` | Hover state |
| Dimmed opacity | `BGP_CONFIG.opacity.dimmed` | `0.24` | Dim state |
| Collector opacity | `BGP_CONFIG.opacity.collector` | `0.62` | Collector state |
| Critical color | `BGP_CONFIG.severityColors.critical` | `0xff4d4f` | Critical event |
| High color | `BGP_CONFIG.severityColors.high` | `0xff9f43` | High-severity event |
| Medium color | `BGP_CONFIG.severityColors.medium` | `0xffd166` | Medium-severity event |
| Low color | `BGP_CONFIG.severityColors.low` | `0x4dabf7` | Low-severity event |
| Collector base color | `BGP_CONFIG.collectorColor` | `0x6db7ff` | Default collector color |
| Region color | `BGP_CONFIG.regionColor` | `0x2dd4bf` | Region overlay |
## Celestial and Starfield
| Name | Variable | Current Value | Location / Notes |
| --- | --- | --- | --- |
| Sky sphere radius | `CELESTIAL_CONFIG.skyRadius` | `2600` | Celestial background |
| Sky opacity | `CELESTIAL_CONFIG.skyOpacity` | `1` | Background material |
| Sun distance / scale | `sunDistance / sunScale` | `2150 / 78` | Sun sprite |
| Moon distance / scale | `moonDistance / moonScale` | `2050 / 38` | Moon sprite |
| Sun halo scale | `CELESTIAL_CONFIG.sunHaloScale` | `136` | Sun halo |
| Moon halo scale | `CELESTIAL_CONFIG.moonHaloScale` | `62` | Moon halo |
| Sun light color / intensity | `sunLightColor / sunLightIntensity` | `0xfff4df / 1.02` | Scene light |
| Back light color / intensity | `backLightColor / backLightIntensity` | `0x2b4c78 / 0.3` | Scene light |
| Star count | `STARFIELD_CONFIG.count` | `8000` | `createStars()` |
| Star radius range | `minRadius + radiusJitter` | `800 + 200` | Random distribution |
| Star color | `STARFIELD_CONFIG.color` | `0xffffff` | `PointsMaterial.color` |
| Star size | `STARFIELD_CONFIG.size` | `0.5` | `PointsMaterial.size` |

View File

@@ -0,0 +1,175 @@
# News Live Streams Collector Format
The `news_live_streams` collector accepts a "channel directory JSON" as input rather than scraping web pages directly.
Goals:
- Allow the backend to stably ingest live news streams from around the world
- Ensure the Earth page TV module always consumes a consistent structure
- Make it easy to integrate channel directories like `worldmonitor` that mix YouTube / HLS / iframe sources
## Recommended JSON Structure
```json
{
"sources": [
{
"id": "bbc-world-news",
"name": "BBC World News",
"provider": "BBC",
"region": "UK",
"language": "en",
"source_type": "youtube",
"youtube_video_id": "dQw4w9WgXcQ",
"youtube_channel": "https://www.youtube.com/@BBCNews",
"embed_url": "",
"stream_url": "",
"homepage_url": "https://www.youtube.com/@BBCNews/live",
"poster_url": "",
"sort_order": 220,
"is_enabled": true,
"notes": "Primary English global news channel"
},
{
"id": "france24-en",
"name": "France 24 English",
"provider": "France 24",
"region": "France",
"language": "en",
"source_type": "hls",
"stream_url": "https://example.com/live.m3u8",
"homepage_url": "https://www.france24.com/en/live",
"sort_order": 230,
"is_enabled": true
},
{
"id": "cctv4-page",
"name": "CCTV-4 Chinese International",
"provider": "CCTV",
"region": "China",
"language": "zh-CN",
"source_type": "iframe",
"embed_url": "https://tv.cctv.com/live/cctv4/",
"homepage_url": "https://tv.cctv.com/live/cctv4/",
"sort_order": 10,
"is_enabled": true
}
]
}
```
## Field Conventions
- `id`: unique identifier, should be stable
- `name`: channel display name
- `provider`: provider name
- `region`: country or region
- `language`: language code
- `source_type`: `iframe` / `hls` / `video` / `external` / `youtube`
- `embed_url`: page suitable for iframe embedding
- `stream_url`: direct video stream URL
- `homepage_url`: official website or channel page
- `youtube_video_id`: YouTube live video ID
- `youtube_channel`: YouTube channel handle or channel URL
- `poster_url`: cover image, optional
- `sort_order`: sort value, smaller = higher in the list
- `is_enabled`: whether enabled
- `notes`: brief notes
## Panel Behavior Conventions
- `youtube`
- Prefers `youtube_video_id`
- When embedding is not possible, at least keep `youtube_channel` or `homepage_url` for external opening
- `hls` / `video`
- Prefers `stream_url`
- `iframe`
- Prefers `embed_url`
- `external`
- No embedding attempt; only keeps external open link
## Current Implementation Status
- The backend settings page supports manually maintaining channel directories
- The Earth TV module merges:
- Manually configured sources
- Sources collected by the `news_live_streams` collector
- The current default fallback source is CCTV-4 Chinese International
- When no override is configured, `news_live_streams` defaults to `iptv-org`:
- `channels.json`
- `streams.json`
- `logos.json`
and automatically filters for news-category channel directories
## Collector Configuration
`news_live_streams` does not need a separate new page; it reuses the existing data source configuration:
- `endpoint`
- Channel directory JSON API URL
- `auth_type`
- `none` / `bearer` / `api_key` / `basic`
- `headers`
- Additional request headers
- `config`
- Collector request and parsing behavior
### Supported `config` Fields
```json
{
"timeout": 30,
"method": "GET",
"params": {
"region": "global"
},
"body_type": "json",
"body": {
"include_disabled": false
},
"response_path": "payload.channels"
}
```
- `timeout`: request timeout in seconds
- `method`: `GET` or `POST`
- `params`: query parameter object
- `body_type`: `json` or `form`
- `body`: request body for `POST`
- `json_body`: explicit JSON request body, takes priority over `body`
- `form_body`: explicit form request body, takes priority over `body`
- `response_path`: path to the channel array in the response JSON, supports dot notation, e.g.:
- `payload.channels`
- `data.items`
- `result.streams`
### Authentication Details
- `bearer`: uses `Authorization: Bearer <token>`
- `api_key`: sent as request header by default; if `auth_config.in = "query"`, sent as query param
- `basic`: uses HTTP Basic Authorization
## Compatible Response Structures
The collector first tries to read:
- Top-level array
- Or an array under these common fields:
- `sources`
- `streams`
- `channels`
- `items`
- `results`
- `data`
It also accepts these field aliases:
- `id` / `source_id` / `slug` / `channel_id` / `code`
- `name` / `title` / `channel` / `display_name`
- `provider` / `publisher` / `network`
- `stream_url` / `stream` / `playback_url` / `hls_url` / `m3u8_url`
- `embed_url` / `embed` / `page_url`
- `homepage_url` / `source_url` / `website`
- `language` / `lang` / `locale`
- `youtube_video_id` / `video_id`
- `youtube_channel` / `channel_handle`

Some files were not shown because too many files have changed in this diff Show More