Compare commits
39 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
eb4c4b7904 | ||
|
|
887fec972e | ||
|
|
5bf5c73ca0 | ||
| e65267fe21 | |||
| ae982e51cd | |||
|
|
65e6a96c0d | ||
|
|
37e92e7572 | ||
|
|
a37d4b6289 | ||
|
|
69789d7505 | ||
|
|
4f124121e7 | ||
|
|
085bdf9a80 | ||
|
|
fbca381512 | ||
|
|
5c65ee24d6 | ||
|
|
81970a1d05 | ||
|
|
9b913a3b83 | ||
|
|
93eb41a9f7 | ||
|
|
dd176a6ae6 | ||
|
|
f14ff6ec0f | ||
|
|
39854b9983 | ||
|
|
3b4347c87d | ||
|
|
d9efd98d26 | ||
|
|
b87cb310fd | ||
|
|
b15d097b9c | ||
|
|
8955c58d19 | ||
|
|
1cb51b1172 | ||
|
|
455b8360d0 | ||
|
|
e1984c7a35 | ||
|
|
bb9183b8a4 | ||
|
|
421234301a | ||
|
|
f22079d33a | ||
|
|
9f737fdb89 | ||
|
|
7418ce2fc1 | ||
|
|
b1a5934b80 | ||
|
|
ba54545ac7 | ||
|
|
9dafbf4f6e | ||
|
|
a87537e903 | ||
|
|
87594a95ff | ||
|
|
2da25376bd | ||
|
|
ac69d5d354 |
@@ -12,6 +12,19 @@ allowed-tools: ["Read", "Edit", "Bash", "Grep", "Glob"]
|
|||||||
|
|
||||||
若 `$ARGUMENTS` 非空,则只检查指定文件/目录;否则检查所有未提交修改(`git diff HEAD`)。
|
若 `$ARGUMENTS` 非空,则只检查指定文件/目录;否则检查所有未提交修改(`git diff HEAD`)。
|
||||||
|
|
||||||
|
## 节省上下文规则
|
||||||
|
|
||||||
|
优先用确定性的 CLI 检查缩小范围,不要一上来把完整文件或大 diff 读入上下文:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git diff --name-only HEAD
|
||||||
|
git diff --unified=0 HEAD -- <path>
|
||||||
|
git diff --check
|
||||||
|
rg -n "TODO|FIXME|console\.log|debugger|print\(" <changed-paths>
|
||||||
|
```
|
||||||
|
|
||||||
|
只有 focused diff 不足以安全判断或修改时,才读取完整文件。
|
||||||
|
|
||||||
## 审查清单
|
## 审查清单
|
||||||
|
|
||||||
按优先级检查以下问题(只报告在本次 diff 中**新增或修改**的代码里存在的问题):
|
按优先级检查以下问题(只报告在本次 diff 中**新增或修改**的代码里存在的问题):
|
||||||
@@ -59,8 +72,13 @@ git diff HEAD --name-only
|
|||||||
|
|
||||||
### Step 2 — 逐文件阅读并分析
|
### Step 2 — 逐文件阅读并分析
|
||||||
|
|
||||||
- 用 Read 工具读取完整文件(不只读 diff)
|
先从 focused diff 开始:
|
||||||
- 对照审查清单,记录每个问题:文件名、行号、问题类型、建议修复方式
|
|
||||||
|
```bash
|
||||||
|
git diff --unified=0 HEAD -- <file>
|
||||||
|
```
|
||||||
|
|
||||||
|
用 `rg`、`git diff --check`、编译器或 linter 输出确认确定性问题。只有需要上下文时才用 Read 读取完整文件。对照审查清单,记录每个问题:文件名、行号、问题类型、建议修复方式。
|
||||||
|
|
||||||
### Step 3 — 报告问题清单
|
### Step 3 — 报告问题清单
|
||||||
|
|
||||||
@@ -95,6 +113,7 @@ git diff HEAD --name-only
|
|||||||
- 只改在审查清单中发现的问题,不做额外优化
|
- 只改在审查清单中发现的问题,不做额外优化
|
||||||
- 每次 Edit 只修改确实有问题的行,保持 diff 最小
|
- 每次 Edit 只修改确实有问题的行,保持 diff 最小
|
||||||
- 改完后用 `grep` 验证旧的坏代码已消失
|
- 改完后用 `grep` 验证旧的坏代码已消失
|
||||||
|
- 优先做精确补丁;只有仓库已有对应格式化流程时,才运行格式化工具
|
||||||
|
|
||||||
### Step 5 — 输出总结
|
### Step 5 — 输出总结
|
||||||
|
|
||||||
|
|||||||
104
.claude/commands/docs.md
Normal file
104
.claude/commands/docs.md
Normal file
@@ -0,0 +1,104 @@
|
|||||||
|
---
|
||||||
|
description: Create or update repository documentation from current code changes
|
||||||
|
argument-hint: Optional: topic to document, or leave empty to infer from git diff
|
||||||
|
allowed-tools: ["Read", "Edit", "Write", "Bash", "Glob", "Grep"]
|
||||||
|
---
|
||||||
|
|
||||||
|
# /docs — Documentation Workflow
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Create or update documentation that explains why a change exists, how it behaves, and what maintainers need to know. Keep this command generic. Repository-specific coverage rules live in the repository and must be loaded separately.
|
||||||
|
|
||||||
|
## Repository Rules
|
||||||
|
|
||||||
|
Before deciding scope, check whether the repository has a documentation rules file:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
test -f docs/documentation-coverage-rules.md && sed -n '1,240p' docs/documentation-coverage-rules.md
|
||||||
|
```
|
||||||
|
|
||||||
|
If it exists, apply it as the project-specific coverage checklist. If it does not exist, continue with the generic workflow below.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
### Step 1 — Understand The Change
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git diff HEAD --stat
|
||||||
|
git diff HEAD --name-only
|
||||||
|
git log --oneline -10
|
||||||
|
rg --files docs
|
||||||
|
```
|
||||||
|
|
||||||
|
If `$ARGUMENTS` specifies a topic, focus on that topic. Otherwise infer the documentation topic from the changed files. Do not read the full repository diff by default; inspect focused files only:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git diff HEAD -- <path>
|
||||||
|
rg -n "class |def |function |export |router|@router|interface |type " <path>
|
||||||
|
```
|
||||||
|
|
||||||
|
### Step 2 — Decide Scope
|
||||||
|
|
||||||
|
- Prefer updating an existing relevant document over creating a duplicate.
|
||||||
|
- Use one document for one coherent topic.
|
||||||
|
- Split documents only when the change crosses meaningful domains.
|
||||||
|
- Keep filenames lowercase and hyphenated.
|
||||||
|
- Apply the repository-specific rules file before writing.
|
||||||
|
|
||||||
|
#### Document Audience Routing (Planet)
|
||||||
|
|
||||||
|
In this repository, classify the action's performer before picking a target file:
|
||||||
|
|
||||||
|
- Browser/UI end user → `docs/technical/{zh,en}/manual.md` or `quickstart.md`.
|
||||||
|
- Shell / Docker / log paths / `planet.sh` / SMTP fallbacks / port forwarding → `docs/technical/{zh,en}/ops-runbook.md` (or an existing `ops-*.md`).
|
||||||
|
- Second-party developers → existing `*-context.md` / `backend-*.md` / `earth-*.md` files.
|
||||||
|
|
||||||
|
Never put shell commands, log paths, or Docker operations into `manual.md` / `quickstart.md`. Never put UI button labels or screenshots into `ops-*.md`. When the same action has both a UI and a CLI path, write each in its own home and cross-link them with one sentence.
|
||||||
|
|
||||||
|
For ambiguous or large documentation changes, briefly state the intended doc plan before editing. For clear small changes, proceed directly.
|
||||||
|
|
||||||
|
### Step 3 — Write
|
||||||
|
|
||||||
|
Explain:
|
||||||
|
|
||||||
|
- Background/problem: what was wrong or missing before.
|
||||||
|
- Core design decisions and rationale.
|
||||||
|
- Operational or user-facing impact.
|
||||||
|
- Relevant code paths, only when useful for future maintainers.
|
||||||
|
|
||||||
|
Style:
|
||||||
|
|
||||||
|
- Follow the repository’s existing language and heading conventions.
|
||||||
|
- Use fenced code blocks with language tags.
|
||||||
|
- Prefer tables for comparisons or parameter lists.
|
||||||
|
- Keep snippets concise and relevant.
|
||||||
|
- For UI labels, chart labels, feature names, datasource names, and other terms that may become mixed Chinese/English copy, check `docs/technical/{zh,en}/naming-glossary.md` and use the documented display name. If a confusing term is missing, update the glossary in both languages as part of the docs change.
|
||||||
|
|
||||||
|
### Step 4 — Verify
|
||||||
|
|
||||||
|
- Read the completed docs once for clarity and stale statements.
|
||||||
|
- Verify referenced paths exist with `test -e` or `rg --files`.
|
||||||
|
- Run applicable checks from `docs/documentation-coverage-rules.md`.
|
||||||
|
- Check Markdown links use readable user-facing titles unless repository rules allow otherwise.
|
||||||
|
|
||||||
|
### Step 5 — Report
|
||||||
|
|
||||||
|
Summarize changed docs and verification:
|
||||||
|
|
||||||
|
```md
|
||||||
|
Updated:
|
||||||
|
- path/to/doc.md — what changed
|
||||||
|
|
||||||
|
Verified:
|
||||||
|
- checks that passed
|
||||||
|
- checks that could not be run, if any
|
||||||
|
```
|
||||||
|
|
||||||
|
## Hard Constraints
|
||||||
|
|
||||||
|
- Do not leave placeholder docs.
|
||||||
|
- Do not duplicate bilingual files byte-for-byte.
|
||||||
|
- Do not reference PR numbers, issue numbers, or the current conversation unless explicitly requested.
|
||||||
|
- Do not write changelog-style lists without the reasoning and tradeoffs behind the change.
|
||||||
|
- Keep docs maintainable and concise.
|
||||||
@@ -72,6 +72,8 @@ Verification
|
|||||||
## 执行风格
|
## 执行风格
|
||||||
|
|
||||||
- 重证据,轻口头判断
|
- 重证据,轻口头判断
|
||||||
|
- 优先使用确定性工具证据:`rg`、`git diff --stat`、`git diff -- <path>`、测试、构建、lint、`curl`、数据库查询等能直接证明成功标准的方式
|
||||||
|
- 不把大段命令输出粘进回复;保留在工具调用里,回复只总结关键证据
|
||||||
- 重验收,轻自我感觉
|
- 重验收,轻自我感觉
|
||||||
- 优先用测试、日志、产物、对比结果来证明完成
|
- 优先用测试、日志、产物、对比结果来证明完成
|
||||||
- 对长期任务保持“未达标就继续”的节奏
|
- 对长期任务保持“未达标就继续”的节奏
|
||||||
|
|||||||
@@ -28,6 +28,19 @@ allowed-tools: ["Read", "Edit", "Bash", "Glob", "Grep"]
|
|||||||
- `docs/CHANGELOG.md`
|
- `docs/CHANGELOG.md`
|
||||||
- `docs/version-history.md`
|
- `docs/version-history.md`
|
||||||
|
|
||||||
|
## 节省上下文规则
|
||||||
|
|
||||||
|
发版判断应以确定性 CLI 证据为主,优先使用紧凑命令和定点读取:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git status --short
|
||||||
|
git diff --stat HEAD
|
||||||
|
git diff --name-only HEAD
|
||||||
|
rg -n "version|^## |^Released:|当前开发版本|current" VERSION frontend/package.json pyproject.toml docs/CHANGELOG.md docs/version-history.md
|
||||||
|
```
|
||||||
|
|
||||||
|
除非需要判断某个代码变更是否属于本次发版,否则不要读取完整 diff。
|
||||||
|
|
||||||
## 执行步骤
|
## 执行步骤
|
||||||
|
|
||||||
### Step 1 — 环境检查
|
### Step 1 — 环境检查
|
||||||
@@ -45,7 +58,7 @@ cat VERSION # 读取当前版本
|
|||||||
### Step 2 — 确定发版类型与新版本号
|
### Step 2 — 确定发版类型与新版本号
|
||||||
|
|
||||||
- 若 `$ARGUMENTS` 提供了明确类型(`feature` / `bugfix`),直接使用
|
- 若 `$ARGUMENTS` 提供了明确类型(`feature` / `bugfix`),直接使用
|
||||||
- 否则根据当前 `git diff HEAD` 和 `git log` 推断
|
- 否则根据 `git diff --stat HEAD`、`git diff --name-only HEAD`、必要的 focused diff 和 `git log` 推断
|
||||||
- 计算新版本号(例:`0.26.2` → bugfix → `0.26.3`)
|
- 计算新版本号(例:`0.26.2` → bugfix → `0.26.3`)
|
||||||
- **先输出发版计划供用户确认**:
|
- **先输出发版计划供用户确认**:
|
||||||
|
|
||||||
@@ -91,12 +104,13 @@ cat VERSION # 读取当前版本
|
|||||||
|
|
||||||
针对本次变更范围做最小验证:
|
针对本次变更范围做最小验证:
|
||||||
|
|
||||||
- Python 文件有修改:`python3 -m py_compile <changed_files>`
|
- Python 文件有修改:先用 `git diff --name-only HEAD -- '*.py'` 列出,再运行 `python3 -m py_compile <changed_files>`
|
||||||
- Frontend 文件有修改:运行项目标准检查(若无则跳过并说明)
|
- Frontend 文件有修改:先用 `git diff --name-only HEAD -- frontend` 判断范围,再运行项目标准检查(若无则跳过并说明)
|
||||||
- 版本号一致性检查:用 grep 确认 VERSION、package.json、pyproject.toml 中的版本号完全一致
|
- 版本号一致性检查:用 grep 确认 VERSION、package.json、pyproject.toml 中的版本号完全一致
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
grep -h "version" VERSION frontend/package.json pyproject.toml
|
cat VERSION
|
||||||
|
rg -n "\"version\":|^version =|version = " frontend/package.json pyproject.toml uv.lock
|
||||||
```
|
```
|
||||||
|
|
||||||
### Step 7 — 提交前预览
|
### Step 7 — 提交前预览
|
||||||
|
|||||||
@@ -21,6 +21,19 @@ If the user specifies a file or directory, check only that. Otherwise check all
|
|||||||
|
|
||||||
Only report issues present in **newly added or modified** lines of this diff — do not audit unchanged code.
|
Only report issues present in **newly added or modified** lines of this diff — do not audit unchanged code.
|
||||||
|
|
||||||
|
## Token-Saving Rule
|
||||||
|
|
||||||
|
Prefer deterministic CLI checks before reading files into model context:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git diff --name-only HEAD
|
||||||
|
git diff --unified=0 HEAD -- <path>
|
||||||
|
git diff --check
|
||||||
|
rg -n "TODO|FIXME|console\.log|debugger|print\(" <changed-paths>
|
||||||
|
```
|
||||||
|
|
||||||
|
Read full files only when the focused diff does not provide enough surrounding context to make a safe edit.
|
||||||
|
|
||||||
## Checklist
|
## Checklist
|
||||||
|
|
||||||
### 1. Duplicate Logic
|
### 1. Duplicate Logic
|
||||||
@@ -64,7 +77,13 @@ Filter to the user-specified path if one was provided.
|
|||||||
|
|
||||||
### Step 2 — Read and analyze each file
|
### Step 2 — Read and analyze each file
|
||||||
|
|
||||||
Read the full file (not just the diff) with the Read tool. For each file, record every issue found: filename, line number, category, and suggested fix.
|
Start with focused diffs:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git diff --unified=0 HEAD -- <file>
|
||||||
|
```
|
||||||
|
|
||||||
|
Use `rg`, `git diff --check`, and compiler/linter output for deterministic findings. Read the full file only for files that need surrounding context. For each issue found, record filename, line number, category, and suggested fix.
|
||||||
|
|
||||||
### Step 3 — Report findings before touching anything
|
### Step 3 — Report findings before touching anything
|
||||||
|
|
||||||
@@ -99,6 +118,7 @@ Principles:
|
|||||||
- Only fix issues identified in the checklist — no extra improvements
|
- Only fix issues identified in the checklist — no extra improvements
|
||||||
- Keep each Edit as small as possible
|
- Keep each Edit as small as possible
|
||||||
- After fixing, verify the old bad pattern is gone with grep
|
- After fixing, verify the old bad pattern is gone with grep
|
||||||
|
- Prefer `apply_patch` for targeted edits; use formatters only when the repository already uses them for the touched file type
|
||||||
|
|
||||||
### Step 5 — Summary
|
### Step 5 — Summary
|
||||||
|
|
||||||
|
|||||||
83
.codex/skills/docs/SKILL.md
Normal file
83
.codex/skills/docs/SKILL.md
Normal file
@@ -0,0 +1,83 @@
|
|||||||
|
---
|
||||||
|
name: docs
|
||||||
|
description: Create or update repository documentation from current code changes. Use when the user asks to write docs, update docs, summarize implementation changes into docs, or check documentation coverage. Load repository-specific coverage rules from docs/documentation-coverage-rules.md when present.
|
||||||
|
---
|
||||||
|
|
||||||
|
# Docs
|
||||||
|
|
||||||
|
Use this skill when the task is documentation work: creating, updating, checking, or summarizing docs for code or behavior changes.
|
||||||
|
|
||||||
|
## Goal
|
||||||
|
|
||||||
|
Write documentation that explains why a change exists, how it behaves, and what maintainers need to know. Keep the skill generic; repository-specific rules belong in the repository, not in this skill.
|
||||||
|
|
||||||
|
## Repository Rules
|
||||||
|
|
||||||
|
Before deciding scope, check whether the repository has a documentation rules file:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
test -f docs/documentation-coverage-rules.md && sed -n '1,240p' docs/documentation-coverage-rules.md
|
||||||
|
```
|
||||||
|
|
||||||
|
If it exists, apply it as the project-specific coverage checklist. If it does not exist, continue with the generic workflow below.
|
||||||
|
|
||||||
|
## Workflow
|
||||||
|
|
||||||
|
1. Gather focused context:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git diff HEAD --stat
|
||||||
|
git diff HEAD --name-only
|
||||||
|
git log --oneline -10
|
||||||
|
rg --files docs
|
||||||
|
```
|
||||||
|
|
||||||
|
If the user gives a topic, focus on that topic. Otherwise infer the doc topic from changed files. Avoid reading large full diffs by default; inspect focused files and symbols:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git diff HEAD -- <path>
|
||||||
|
rg -n "class |def |function |export |router|@router|interface |type " <path>
|
||||||
|
```
|
||||||
|
|
||||||
|
2. Decide scope:
|
||||||
|
|
||||||
|
- Prefer updating an existing relevant doc over creating a duplicate.
|
||||||
|
- Use one document for one coherent topic.
|
||||||
|
- Split documents only when changes cross meaningful domains.
|
||||||
|
- Keep filenames lowercase and hyphenated.
|
||||||
|
|
||||||
|
3. Write the doc:
|
||||||
|
|
||||||
|
- Explain background/problem, design decisions, constraints, and operational impact.
|
||||||
|
- Keep code snippets short and directly relevant.
|
||||||
|
- List related files only when they help future maintainers navigate.
|
||||||
|
- Use the repository’s existing language, heading style, and naming conventions.
|
||||||
|
- For UI labels, chart labels, feature names, datasource names, and other terms that may become mixed Chinese/English copy, check `docs/technical/{zh,en}/naming-glossary.md` and use the documented display name. If a confusing term is missing, update the glossary in both languages as part of the docs change.
|
||||||
|
|
||||||
|
4. Verify:
|
||||||
|
|
||||||
|
- Read the completed doc once for clarity and stale statements.
|
||||||
|
- Verify important referenced paths exist with `test -e` or `rg --files`.
|
||||||
|
- Run repository-specific doc checks from `docs/documentation-coverage-rules.md` when present.
|
||||||
|
- For Markdown links, check that user-facing titles are readable and not raw filenames unless the repository rules allow it.
|
||||||
|
|
||||||
|
## Hard Constraints
|
||||||
|
|
||||||
|
- Do not leave placeholder docs or copied source text pretending to be documentation.
|
||||||
|
- Do not duplicate bilingual files byte-for-byte.
|
||||||
|
- Do not reference PR numbers, issue numbers, or the current conversation unless explicitly requested.
|
||||||
|
- Do not write changelog-style lists without the reasoning, constraints, and tradeoffs behind the change.
|
||||||
|
- Keep docs concise enough to maintain.
|
||||||
|
|
||||||
|
## Recommended Output
|
||||||
|
|
||||||
|
After editing, summarize:
|
||||||
|
|
||||||
|
```md
|
||||||
|
Updated:
|
||||||
|
- path/to/doc.md — what changed
|
||||||
|
|
||||||
|
Verified:
|
||||||
|
- checks that passed
|
||||||
|
- checks that could not be run, if any
|
||||||
|
```
|
||||||
@@ -72,6 +72,8 @@ In Codex, only use actual subagents when the user explicitly asks for delegation
|
|||||||
## Operating Rules
|
## Operating Rules
|
||||||
|
|
||||||
- Prefer objective checks over self-reported completion.
|
- Prefer objective checks over self-reported completion.
|
||||||
|
- Prefer deterministic tool evidence over long model summaries: use `rg`, `git diff --stat`, targeted `git diff -- <path>`, tests, builds, linters, `curl`, or database queries when they can prove a criterion.
|
||||||
|
- Do not paste large command output into the conversation; summarize the evidence and keep raw output in tool calls.
|
||||||
- Do not confuse progress with completion.
|
- Do not confuse progress with completion.
|
||||||
- If the worker says "done", verify it.
|
- If the worker says "done", verify it.
|
||||||
- If verification fails, continue from the gap instead of restarting blindly.
|
- If verification fails, continue from the gap instead of restarting blindly.
|
||||||
|
|||||||
@@ -35,6 +35,19 @@ Use `git rev-parse --show-toplevel` to get the repo root. All paths are relative
|
|||||||
- `docs/CHANGELOG.md`
|
- `docs/CHANGELOG.md`
|
||||||
- `docs/version-history.md`
|
- `docs/version-history.md`
|
||||||
|
|
||||||
|
## Token-Saving Rule
|
||||||
|
|
||||||
|
Release work should be driven by deterministic CLI evidence. Prefer compact commands and targeted file reads:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git status --short
|
||||||
|
git diff --stat HEAD
|
||||||
|
git diff --name-only HEAD
|
||||||
|
rg -n "version|^## |^Released:|current" VERSION frontend/package.json pyproject.toml docs/CHANGELOG.md docs/version-history.md
|
||||||
|
```
|
||||||
|
|
||||||
|
Do not inspect full diffs unless deciding whether changed code belongs in the release.
|
||||||
|
|
||||||
## Workflow
|
## Workflow
|
||||||
|
|
||||||
### Step 1 — Environment check
|
### Step 1 — Environment check
|
||||||
@@ -52,7 +65,7 @@ If unrelated uncommitted changes exist, list them and ask the user whether to in
|
|||||||
### Step 2 — Determine release type and next version
|
### Step 2 — Determine release type and next version
|
||||||
|
|
||||||
- If the user provided an explicit type (`feature` / `bugfix`), use it
|
- If the user provided an explicit type (`feature` / `bugfix`), use it
|
||||||
- Otherwise infer from `git diff HEAD` and recent `git log`
|
- Otherwise infer from `git diff --stat HEAD`, `git diff --name-only HEAD`, focused diffs for changed code, and recent `git log`
|
||||||
- Compute the next version:
|
- Compute the next version:
|
||||||
- `feature`: increment minor and reset patch to `0` (e.g. `0.41.2` → `0.42.0`)
|
- `feature`: increment minor and reset patch to `0` (e.g. `0.41.2` → `0.42.0`)
|
||||||
- `bugfix`: increment patch only (e.g. `0.26.2` → `0.26.3`)
|
- `bugfix`: increment patch only (e.g. `0.26.2` → `0.26.3`)
|
||||||
@@ -106,12 +119,13 @@ Get today's date with `date +%Y-%m-%d`.
|
|||||||
|
|
||||||
Run the smallest relevant validation for the changes in scope:
|
Run the smallest relevant validation for the changes in scope:
|
||||||
|
|
||||||
- Python files changed: `python3 -m py_compile <changed_files>`
|
- Python files changed: list changed Python files with `git diff --name-only HEAD -- '*.py'`, then run `python3 -m py_compile <changed_files>`
|
||||||
- Frontend files changed: run the project-standard check if available; otherwise skip and say so
|
- Frontend files changed: list changed frontend files with `git diff --name-only HEAD -- frontend`, then run the project-standard check if available; otherwise skip and say so
|
||||||
- Version consistency: confirm VERSION, package.json, pyproject.toml, and uv.lock all show the same version
|
- Version consistency: confirm VERSION, package.json, pyproject.toml, and uv.lock all show the same version
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
grep -h "version" VERSION frontend/package.json pyproject.toml
|
cat VERSION
|
||||||
|
rg -n "\"version\":|^version =|version = " frontend/package.json pyproject.toml uv.lock
|
||||||
```
|
```
|
||||||
|
|
||||||
### Step 7 — Pre-commit preview
|
### Step 7 — Pre-commit preview
|
||||||
|
|||||||
18
.dockerignore
Normal file
18
.dockerignore
Normal file
@@ -0,0 +1,18 @@
|
|||||||
|
**
|
||||||
|
|
||||||
|
!pyproject.toml
|
||||||
|
!uv.lock
|
||||||
|
!VERSION
|
||||||
|
!backend/
|
||||||
|
!backend/**
|
||||||
|
!aiprovider/
|
||||||
|
!aiprovider/**
|
||||||
|
|
||||||
|
backend/.env
|
||||||
|
backend/.env.*
|
||||||
|
aiprovider/.env
|
||||||
|
aiprovider/.env.*
|
||||||
|
!aiprovider/.env.example
|
||||||
|
**/__pycache__/
|
||||||
|
**/*.pyc
|
||||||
|
**/*.pyo
|
||||||
64
.gitea/workflows/ci.yaml
Normal file
64
.gitea/workflows/ci.yaml
Normal file
@@ -0,0 +1,64 @@
|
|||||||
|
name: ci
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches:
|
||||||
|
- dev
|
||||||
|
- main
|
||||||
|
pull_request:
|
||||||
|
|
||||||
|
env:
|
||||||
|
REGISTRY: gitea.rclaw.top
|
||||||
|
IMAGE_NAMESPACE: linkong/planet
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
backend:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@v4
|
||||||
|
- name: Install uv
|
||||||
|
run: curl -LsSf https://astral.sh/uv/install.sh | sh
|
||||||
|
- name: Sync Python dependencies
|
||||||
|
run: ~/.local/bin/uv sync --group dev
|
||||||
|
- name: Run backend smoke tests
|
||||||
|
working-directory: backend
|
||||||
|
run: PYTHONPATH=. "$GITHUB_WORKSPACE/.venv/bin/python" -m pytest -s tests/test_api.py tests/test_realtime_sources.py -q
|
||||||
|
|
||||||
|
frontend:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@v4
|
||||||
|
- name: Install Bun
|
||||||
|
run: curl -fsSL https://bun.sh/install | bash
|
||||||
|
- name: Build frontend
|
||||||
|
working-directory: frontend
|
||||||
|
run: |
|
||||||
|
~/.bun/bin/bun install --frozen-lockfile
|
||||||
|
~/.bun/bin/bun run build
|
||||||
|
|
||||||
|
delivery:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
needs:
|
||||||
|
- backend
|
||||||
|
- frontend
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@v4
|
||||||
|
- name: Install Helm
|
||||||
|
run: |
|
||||||
|
mkdir -p "$HOME/.local/bin"
|
||||||
|
curl -fsSL https://get.helm.sh/helm-v3.15.4-linux-amd64.tar.gz -o /tmp/helm.tar.gz
|
||||||
|
tar -xzf /tmp/helm.tar.gz -C /tmp
|
||||||
|
mv /tmp/linux-amd64/helm "$HOME/.local/bin/helm"
|
||||||
|
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
|
||||||
|
- name: Docker build smoke
|
||||||
|
run: |
|
||||||
|
docker build -t "$REGISTRY/$IMAGE_NAMESPACE/frontend:${GITHUB_SHA}" ./frontend
|
||||||
|
docker build -t "$REGISTRY/$IMAGE_NAMESPACE/backend:${GITHUB_SHA}" -f backend/Dockerfile .
|
||||||
|
docker build -t "$REGISTRY/$IMAGE_NAMESPACE/aiprovider:${GITHUB_SHA}" -f aiprovider/Dockerfile .
|
||||||
|
- name: Helm template smoke
|
||||||
|
run: |
|
||||||
|
helm lint deploy/helm/planet
|
||||||
|
helm template planet-staging deploy/helm/planet \
|
||||||
|
--namespace planet-staging \
|
||||||
|
-f deploy/helm/planet/values.single-node.yaml \
|
||||||
|
--set image.tag="${GITHUB_SHA}" >/tmp/planet-rendered.yaml
|
||||||
67
.gitea/workflows/deploy-staging.yaml
Normal file
67
.gitea/workflows/deploy-staging.yaml
Normal file
@@ -0,0 +1,67 @@
|
|||||||
|
name: deploy-staging
|
||||||
|
|
||||||
|
on:
|
||||||
|
workflow_dispatch:
|
||||||
|
push:
|
||||||
|
branches:
|
||||||
|
- main
|
||||||
|
|
||||||
|
env:
|
||||||
|
REGISTRY: gitea.rclaw.top
|
||||||
|
IMAGE_NAMESPACE: linkong/planet
|
||||||
|
RELEASE_NAME: planet-staging
|
||||||
|
NAMESPACE: planet-staging
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
deploy:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@v4
|
||||||
|
- name: Install deploy tools
|
||||||
|
run: |
|
||||||
|
mkdir -p "$HOME/.local/bin"
|
||||||
|
curl -fsSL https://get.helm.sh/helm-v3.15.4-linux-amd64.tar.gz -o /tmp/helm.tar.gz
|
||||||
|
tar -xzf /tmp/helm.tar.gz -C /tmp
|
||||||
|
mv /tmp/linux-amd64/helm "$HOME/.local/bin/helm"
|
||||||
|
curl -fsSL https://dl.k8s.io/release/v1.30.5/bin/linux/amd64/kubectl -o /tmp/kubectl
|
||||||
|
install -m 0755 /tmp/kubectl "$HOME/.local/bin/kubectl"
|
||||||
|
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
|
||||||
|
- name: Configure kubeconfig
|
||||||
|
run: |
|
||||||
|
mkdir -p "$HOME/.kube"
|
||||||
|
printf "%s" "${{ secrets.KUBE_CONFIG_STAGING }}" | base64 -d > "$HOME/.kube/config"
|
||||||
|
chmod 600 "$HOME/.kube/config"
|
||||||
|
- name: Deploy Helm release
|
||||||
|
run: |
|
||||||
|
kubectl create namespace "$NAMESPACE" --dry-run=client -o yaml | kubectl apply -f -
|
||||||
|
helm upgrade --install "$RELEASE_NAME" deploy/helm/planet \
|
||||||
|
--namespace "$NAMESPACE" \
|
||||||
|
-f deploy/helm/planet/values.single-node.yaml \
|
||||||
|
--set global.imageRegistry="$REGISTRY" \
|
||||||
|
--set global.imageNamespace="$IMAGE_NAMESPACE" \
|
||||||
|
--set image.tag="${GITHUB_SHA}"
|
||||||
|
- name: Wait for rollout
|
||||||
|
run: |
|
||||||
|
kubectl rollout status deployment/planet-frontend -n "$NAMESPACE" --timeout=180s
|
||||||
|
kubectl rollout status deployment/planet-backend -n "$NAMESPACE" --timeout=180s
|
||||||
|
kubectl rollout status deployment/planet-aiprovider -n "$NAMESPACE" --timeout=180s
|
||||||
|
- name: Smoke test services
|
||||||
|
run: |
|
||||||
|
kubectl run planet-smoke-${GITHUB_RUN_NUMBER} \
|
||||||
|
--rm -i --restart=Never \
|
||||||
|
--namespace "$NAMESPACE" \
|
||||||
|
--image=curlimages/curl:8.11.1 \
|
||||||
|
--command -- sh -c '
|
||||||
|
set -eu
|
||||||
|
curl -fsS http://planet-frontend:3000/ >/dev/null
|
||||||
|
curl -fsS http://planet-frontend:3000/health >/dev/null
|
||||||
|
curl -fsS http://planet-frontend:3000/api/health >/dev/null
|
||||||
|
curl -fsS http://planet-backend:8000/health >/dev/null
|
||||||
|
curl -fsS http://planet-aiprovider:8010/health >/dev/null
|
||||||
|
'
|
||||||
|
- name: Collect diagnostics on failure
|
||||||
|
if: failure()
|
||||||
|
run: |
|
||||||
|
kubectl get all -n "$NAMESPACE" -o wide || true
|
||||||
|
kubectl describe pods -n "$NAMESPACE" || true
|
||||||
|
kubectl logs -n "$NAMESPACE" -l app.kubernetes.io/instance="$RELEASE_NAME" --all-containers --tail=200 || true
|
||||||
58
.gitea/workflows/release.yaml
Normal file
58
.gitea/workflows/release.yaml
Normal file
@@ -0,0 +1,58 @@
|
|||||||
|
name: release
|
||||||
|
|
||||||
|
on:
|
||||||
|
push:
|
||||||
|
branches:
|
||||||
|
- main
|
||||||
|
tags:
|
||||||
|
- "v*"
|
||||||
|
workflow_dispatch:
|
||||||
|
|
||||||
|
env:
|
||||||
|
REGISTRY: gitea.rclaw.top
|
||||||
|
IMAGE_NAMESPACE: linkong/planet
|
||||||
|
|
||||||
|
jobs:
|
||||||
|
images:
|
||||||
|
runs-on: ubuntu-latest
|
||||||
|
steps:
|
||||||
|
- uses: actions/checkout@v4
|
||||||
|
- name: Resolve image tags
|
||||||
|
id: meta
|
||||||
|
run: |
|
||||||
|
echo "sha_tag=${GITHUB_SHA}" >> "$GITHUB_OUTPUT"
|
||||||
|
if printf "%s" "${GITHUB_REF}" | grep -q '^refs/tags/v'; then
|
||||||
|
echo "release_tag=${GITHUB_REF_NAME}" >> "$GITHUB_OUTPUT"
|
||||||
|
else
|
||||||
|
echo "release_tag=" >> "$GITHUB_OUTPUT"
|
||||||
|
fi
|
||||||
|
- name: Login to registry
|
||||||
|
run: echo "${{ secrets.REGISTRY_PASSWORD }}" | docker login "$REGISTRY" -u "${{ secrets.REGISTRY_USER }}" --password-stdin
|
||||||
|
- name: Build and push images
|
||||||
|
run: |
|
||||||
|
for service in frontend backend aiprovider; do
|
||||||
|
case "$service" in
|
||||||
|
frontend)
|
||||||
|
dockerfile="./frontend/Dockerfile"
|
||||||
|
context="./frontend"
|
||||||
|
;;
|
||||||
|
backend)
|
||||||
|
dockerfile="backend/Dockerfile"
|
||||||
|
context="."
|
||||||
|
;;
|
||||||
|
aiprovider)
|
||||||
|
dockerfile="aiprovider/Dockerfile"
|
||||||
|
context="."
|
||||||
|
;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
image="$REGISTRY/$IMAGE_NAMESPACE/$service:${{ steps.meta.outputs.sha_tag }}"
|
||||||
|
docker build -t "$image" -f "$dockerfile" "$context"
|
||||||
|
docker push "$image"
|
||||||
|
|
||||||
|
if [ -n "${{ steps.meta.outputs.release_tag }}" ]; then
|
||||||
|
release_image="$REGISTRY/$IMAGE_NAMESPACE/$service:${{ steps.meta.outputs.release_tag }}"
|
||||||
|
docker tag "$image" "$release_image"
|
||||||
|
docker push "$release_image"
|
||||||
|
fi
|
||||||
|
done
|
||||||
7
.gitignore
vendored
7
.gitignore
vendored
@@ -8,6 +8,7 @@
|
|||||||
.env
|
.env
|
||||||
.env.local
|
.env.local
|
||||||
.env.*.local
|
.env.*.local
|
||||||
|
config/earth-boundary-sources.local.json
|
||||||
*.pem
|
*.pem
|
||||||
*.key
|
*.key
|
||||||
*.crt
|
*.crt
|
||||||
@@ -150,3 +151,9 @@ temp/
|
|||||||
# Runtime Data
|
# Runtime Data
|
||||||
# ----------------------
|
# ----------------------
|
||||||
data/ai/bgp-briefs/
|
data/ai/bgp-briefs/
|
||||||
|
data/earth-boundary-sources/
|
||||||
|
|
||||||
|
# Generated Earth boundary tile artifacts. Keep source configs and builders in
|
||||||
|
# Git; publish PMTiles/MVT artifacts through release/deploy storage instead of
|
||||||
|
# committing thousands of generated tile files.
|
||||||
|
frontend/public/earth/data/boundaries/
|
||||||
|
|||||||
264
README.md
264
README.md
@@ -8,68 +8,54 @@
|
|||||||
|
|
||||||
## 系统架构
|
## 系统架构
|
||||||
|
|
||||||
|
当前仓库的核心形态是“Web Earth 可视化 + React 运维台 + FastAPI 数据与 AI 编排后端 + 独立模型适配层”。物理大屏与 UE 客户端仍是长期方向,但不再作为本地开发和当前发布的必需运行单元。
|
||||||
|
|
||||||
```
|
```
|
||||||
┌─────────────────────────────────────────────────────────────────────────┐
|
┌─────────────────────────────────────────────────────────────────────┐
|
||||||
│ 物理大屏展示层 │
|
│ 浏览器展示与运维层 │
|
||||||
│ ┌─────────────────────────────────────────────────────────────────┐ │
|
│ ┌──────────────────────────────┐ ┌──────────────────────────────┐ │
|
||||||
│ │ 偏振片3D大屏 (2m×3m, 4K, 120Hz, 眼镜式) │ │
|
│ │ Web Earth │ │ React 运维台 │ │
|
||||||
│ │ ┌─────────────────────────────────────────────────────────┐ │ │
|
│ │ frontend/public/earth │ │ frontend/src │ │
|
||||||
│ │ │ 虚幻引擎 UE5 客户端 │ │ │
|
│ │ Three.js 地球 / HUD / 新闻 │ │ 数据源 / 告警 / AI 设置 │ │
|
||||||
│ │ │ ├── 3D地球渲染 (Cesium for UE) │ │ │
|
│ │ 国界精度 / 品牌内容配置 │ │ 提示词配置 / 用户与系统配置 │ │
|
||||||
│ │ │ ├── 算力点可视化 (GPU集群、智算中心) │ │ │
|
│ └──────────────────────────────┘ └──────────────────────────────┘ │
|
||||||
│ │ │ ├── 连接弧线 (光缆、路由、数据流向) │ │ │
|
└─────────────────────────────────────────────────────────────────────┘
|
||||||
│ │ │ ├── 粒子效果 (数据流动、告警提示) │ │ │
|
│ REST / WebSocket
|
||||||
│ │ │ └── 自动巡航相机 + 交互控制 │ │ │
|
▼
|
||||||
│ │ └─────────────────────────────────────────────────────────┘ │ │
|
┌─────────────────────────────────────────────────────────────────────┐
|
||||||
│ └─────────────────────────────────────────────────────────────────┘ │
|
│ FastAPI 业务与编排后端 │
|
||||||
└─────────────────────────────────────────────────────────────────────────┘
|
│ ┌────────────────────┐ ┌────────────────────┐ ┌─────────────────┐ │
|
||||||
▲
|
│ │ 数据 API 与认证 │ │ Earth 新闻增强 │ │ 告警与态势简报 │ │
|
||||||
│ WebSocket (实时推送)
|
│ │ JWT / 权限 / 审计 │ │ 位置推断 / 本地化 │ │ BGP / 告警研判 │ │
|
||||||
│ 120Hz 心跳 / 数据帧同步
|
│ └────────────────────┘ └────────────────────┘ └─────────────────┘ │
|
||||||
▼
|
│ ┌────────────────────┐ ┌────────────────────┐ ┌─────────────────┐ │
|
||||||
┌─────────────────────────────────────────────────────────────────────────┐
|
│ │ 系统运行配置 │ │ 默认提示词注册表 │ │ 未来 Agent Runtime│ │
|
||||||
│ 数据中台服务层 (FastAPI) │
|
│ │ system_settings │ │ 代码发布 + DB 覆盖 │ │ 工具/证据/工作流 │ │
|
||||||
│ ┌─────────────────────────────────────────────────────────────────┐ │
|
│ └────────────────────┘ └────────────────────┘ └─────────────────┘ │
|
||||||
│ │ API Gateway (Redis 限流) │ │
|
└─────────────────────────────────────────────────────────────────────┘
|
||||||
│ └─────────────────────────────────────────────────────────────────┘ │
|
│ SQLAlchemy / Redis Stream │ 纯净 LLM 调用
|
||||||
│ │ │
|
▼ ▼
|
||||||
│ ┌───────────────────┬──────────────────────────┬──────────────────┐ │
|
┌──────────────────────────────┐ ┌──────────────────────────────┐
|
||||||
│ │ 数据采集服务 │ 核心业务服务 │ 运维管理服务 │ │
|
│ PostgreSQL / Redis │ │ aiprovider │
|
||||||
│ │ ┌─────────────┐ │ ┌─────────────────┐ │ ┌─────────────┐ │ │
|
│ 用户、配置、采集结果、新闻 │ │ provider + protocol adapter │
|
||||||
│ │ │ 调度中心 │ │ │ WebSocket 服务 │ │ │ 用户管理 │ │ │
|
│ Stream、缓存、运行状态 │ │ OpenAI / MiniMax / Ollama 等 │
|
||||||
│ │ │ (Celery) │ │ │ (FastAPI) │ │ │ (JWT Auth) │ │ │
|
└──────────────────────────────┘ └──────────────────────────────┘
|
||||||
│ │ └─────────────┘ │ └─────────────────┘ │ └─────────────┘ │ │
|
▲
|
||||||
│ │ ┌─────────────┐ │ ┌─────────────────┐ │ ┌─────────────┐ │ │
|
│ 采集器 / 外部数据源
|
||||||
│ │ │ 采集器池 │ │ │ 数据查询 API │ │ │ 数据源配置 │ │ │
|
▼
|
||||||
│ │ │ (10+源) │ │ │ (REST) │ │ │ 监控告警 │ │ │
|
┌─────────────────────────────────────────────────────────────────────┐
|
||||||
│ │ └─────────────┘ │ └─────────────────┘ │ └─────────────┘ │ │
|
│ RSS 新闻、BGP 观测、公开数据源、后续 WebSearch/OCR/语音识别等工具 │
|
||||||
│ │ ┌─────────────┐ │ ┌─────────────────┐ │ ┌─────────────┐ │ │
|
└─────────────────────────────────────────────────────────────────────┘
|
||||||
│ │ │ 消息队列 │ │ │ 态势分析引擎 │ │ │ 系统配置 │ │ │
|
|
||||||
│ │ │ (Kafka) │ │ │ (计算/聚合) │ │ │ 日志审计 │ │ │
|
|
||||||
│ │ └─────────────┘ │ └─────────────────┘ │ └─────────────┘ │ │
|
|
||||||
│ └───────────────────┴──────────────────────────┴──────────────────┘ │
|
|
||||||
└─────────────────────────────────────────────────────────────────────────┘
|
|
||||||
▲
|
|
||||||
│ 内部 API 调用
|
|
||||||
▼
|
|
||||||
┌─────────────────────────────────────────────────────────────────────────┐
|
|
||||||
│ Web管理端 (React Admin) │
|
|
||||||
│ ┌─────────────────────────────────────────────────────────────────┐ │
|
|
||||||
│ │ 登录页 │ 仪表盘 │ 用户管理 │ 数据源配置 │ 任务监控 │ 系统配置 │ │
|
|
||||||
│ └─────────────────────────────────────────────────────────────────┘ │
|
|
||||||
└─────────────────────────────────────────────────────────────────────────┘
|
|
||||||
▲
|
|
||||||
│ PostgreSQL / Redis
|
|
||||||
▼
|
|
||||||
┌─────────────────────────────────────────────────────────────────────────┐
|
|
||||||
│ 数据存储层 │
|
|
||||||
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │
|
|
||||||
│ │ PostgreSQL │ │ TimescaleDB │ │ Redis │ │ MinIO │ │
|
|
||||||
│ │ (用户/配置) │ │ (时序数据) │ │ (缓存/会话) │ │ (文件存储) │ │
|
|
||||||
│ └─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘ │
|
|
||||||
└─────────────────────────────────────────────────────────────────────────┘
|
|
||||||
```
|
```
|
||||||
|
|
||||||
|
架构边界:
|
||||||
|
|
||||||
|
- `backend` 负责业务语义、证据收集、提示词选择、AI 任务编排、权限和数据落库。
|
||||||
|
- `aiprovider` 只负责把纯净模型请求适配到不同供应商或协议,不内置具体业务提示词。
|
||||||
|
- 默认提示词随代码发布并保存在 `backend/app/ai_tasks/default_prompts.json`,运维台可在数据库中保存覆盖值,重置时回到当前代码版本的默认提示词。
|
||||||
|
- Earth 新闻保留英文原文,中文展示结果存入 `localizations`,前端默认展示 `zh-CN` 的 `display_title`、`display_summary` 和中文地域/状态文案。
|
||||||
|
- Earth LLM 指令、语音识别、多角色态势研判属于后续 Agent Runtime 方向,计划见 [docs/plans/agents-earth-command-runtime-plan.md](/home/ray/dev/linkong/planet/docs/plans/agents-earth-command-runtime-plan.md)。
|
||||||
|
|
||||||
## 四大核心要素
|
## 四大核心要素
|
||||||
|
|
||||||
| 层级 | 要素 | 描述 |
|
| 层级 | 要素 | 描述 |
|
||||||
@@ -87,11 +73,10 @@
|
|||||||
|------|------|------|
|
|------|------|------|
|
||||||
| FastAPI | 0.109+ | Web 框架 |
|
| FastAPI | 0.109+ | Web 框架 |
|
||||||
| SQLAlchemy | 2.0+ | ORM |
|
| SQLAlchemy | 2.0+ | ORM |
|
||||||
| Alembic | - | 数据库迁移 |
|
| uv | - | Python 依赖与命令运行 |
|
||||||
| Celery | 5.3+ | 任务队列 |
|
| Redis | 7.0+ | 缓存、Stream 与运行协调 |
|
||||||
| Redis | 7.0+ | 缓存/消息 |
|
|
||||||
| Kafka | 3.0+ | 事件流 |
|
|
||||||
| PyJWT | - | 认证 |
|
| PyJWT | - | 认证 |
|
||||||
|
| APScheduler / 后台任务 | - | 采集、增强与运行时任务 |
|
||||||
|
|
||||||
### 前端 (React Admin)
|
### 前端 (React Admin)
|
||||||
|
|
||||||
@@ -102,6 +87,7 @@
|
|||||||
| Axios | HTTP 客户端 |
|
| Axios | HTTP 客户端 |
|
||||||
| Socket.io-client | WebSocket 客户端 |
|
| Socket.io-client | WebSocket 客户端 |
|
||||||
| ECharts | 统计图表 |
|
| ECharts | 统计图表 |
|
||||||
|
| Three.js | Earth 3D 地球渲染 |
|
||||||
| Bun | 前端包管理与脚本运行 |
|
| Bun | 前端包管理与脚本运行 |
|
||||||
|
|
||||||
前端工程统一使用 Bun:
|
前端工程统一使用 Bun:
|
||||||
@@ -110,22 +96,16 @@
|
|||||||
- 运行脚本使用 `bun run <script>`
|
- 运行脚本使用 `bun run <script>`
|
||||||
- 不使用 `npm`、`pnpm`、`yarn`
|
- 不使用 `npm`、`pnpm`、`yarn`
|
||||||
|
|
||||||
### 虚幻引擎客户端
|
### 大屏与 3D 展示方向
|
||||||
|
|
||||||
| 组件 | 版本 | 用途 |
|
当前发布优先使用浏览器 Web Earth。UE5 / Cesium for Unreal / Niagara 可作为后续物理大屏方向接入,但不是本地开发闭环的必需组件。
|
||||||
|------|------|------|
|
|
||||||
| Unreal Engine 5 | 5.3+ | 3D 渲染引擎 |
|
|
||||||
| Cesium for Unreal | 1.5+ | 地理可视化 |
|
|
||||||
| Niagara | - | 粒子系统 |
|
|
||||||
|
|
||||||
### 数据库
|
### 数据库
|
||||||
|
|
||||||
| 组件 | 用途 |
|
| 组件 | 用途 |
|
||||||
|------|------|
|
|------|------|
|
||||||
| PostgreSQL 15+ | 关系数据 |
|
| PostgreSQL 15+ | 关系数据 |
|
||||||
| TimescaleDB | 时序数据扩展 |
|
| Redis 7+ | 缓存、Stream、运行状态 |
|
||||||
| Redis 7+ | 缓存/会话 |
|
|
||||||
| MinIO | S3 兼容存储 |
|
|
||||||
|
|
||||||
### 部署
|
### 部署
|
||||||
|
|
||||||
@@ -152,9 +132,9 @@
|
|||||||
| P0 | Epoch AI | 每小时 |
|
| P0 | Epoch AI | 每小时 |
|
||||||
| P0 | Hugging Face | 每 2 小时 |
|
| P0 | Hugging Face | 每 2 小时 |
|
||||||
| P0 | GitHub | 每 4 小时 |
|
| P0 | GitHub | 每 4 小时 |
|
||||||
| P0 每日 |
|
| P0 | 海底光缆 / IXP / 卫星等基础设施数据 | 每日或按源刷新 |
|
||||||
| P0 | PeeringDB | 每 2 小时 |
|
| P0 | PeeringDB | 每 2 小时 |
|
||||||
| P1 | Cloudflare Radar | | TeleGeography | 每小时 |
|
| P1 | Cloudflare Radar / TeleGeography | 每小时 |
|
||||||
| P1 | CAIDA BGPStream | 每 15 分钟 |
|
| P1 | CAIDA BGPStream | 每 15 分钟 |
|
||||||
|
|
||||||
## 项目结构
|
## 项目结构
|
||||||
@@ -166,20 +146,18 @@
|
|||||||
│ │ ├── core/ # 核心配置
|
│ │ ├── core/ # 核心配置
|
||||||
│ │ ├── models/ # 数据模型
|
│ │ ├── models/ # 数据模型
|
||||||
│ │ ├── schemas/ # Pydantic 模型
|
│ │ ├── schemas/ # Pydantic 模型
|
||||||
│ │ ├── services/ # 业务逻辑
|
│ │ ├── services/ # 业务逻辑与 AI 任务编排
|
||||||
│ │ └── tasks/ # Celery 任务
|
│ │ └── ai_tasks/ # 默认提示词与 AI 任务定义
|
||||||
│ └── tests/
|
│ └── tests/
|
||||||
|
├── aiprovider/ # 独立模型供应商适配层
|
||||||
├── frontend/ # React 管理后台
|
├── frontend/ # React 管理后台
|
||||||
│ ├── src/
|
│ ├── src/
|
||||||
│ │ ├── components/ # 组件
|
│ │ ├── components/ # 组件
|
||||||
│ │ ├── pages/ # 页面
|
│ │ ├── pages/ # 页面
|
||||||
│ │ ├── services/ # API 服务
|
│ │ ├── services/ # API 服务
|
||||||
│ │ └── store/ # 状态管理
|
│ │ └── store/ # 状态管理
|
||||||
│ └── tests/
|
│ ├── public/earth/ # Web Earth 静态应用
|
||||||
├── unreal/ # UE5 大屏客户端
|
│ └── tests/ # 前端测试
|
||||||
│ ├── Content/
|
|
||||||
│ ├── Source/
|
|
||||||
│ └── Plugins/
|
|
||||||
├── data/ # 数据文件
|
├── data/ # 数据文件
|
||||||
├── docs/ # 文档
|
├── docs/ # 文档
|
||||||
├── scripts/ # 脚本
|
├── scripts/ # 脚本
|
||||||
@@ -191,10 +169,11 @@
|
|||||||
## 快速启动
|
## 快速启动
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# 新机器首次初始化
|
# 新机器或空项目首次初始化
|
||||||
./scripts/bootstrap-dev.sh
|
./planet.sh init
|
||||||
# 会自动安装/检查 uv、bun,并同步 Python/前端依赖
|
# 会自动安装/检查 uv、bun,同步 Python/前端依赖
|
||||||
# 会在缺少时生成 backend/.env、aiprovider/.env、frontend/.env.local
|
# 会在缺少时生成 backend/.env、aiprovider/.env、frontend/.env.local
|
||||||
|
# 会启动 PostgreSQL/Redis,并创建表、默认数据源和本地默认用户
|
||||||
|
|
||||||
# 启动前后端服务
|
# 启动前后端服务
|
||||||
./planet.sh start
|
./planet.sh start
|
||||||
@@ -210,6 +189,9 @@
|
|||||||
|
|
||||||
# 查看服务状态
|
# 查看服务状态
|
||||||
./planet.sh health
|
./planet.sh health
|
||||||
|
|
||||||
|
# 删除容器、卷、镜像和本地编译状态,执行前需要输入 Y 确认
|
||||||
|
./planet.sh destroy
|
||||||
```
|
```
|
||||||
|
|
||||||
前端命令约定:
|
前端命令约定:
|
||||||
@@ -236,13 +218,15 @@ bun run build
|
|||||||
|
|
||||||
推荐按下面顺序排查和配置。
|
推荐按下面顺序排查和配置。
|
||||||
|
|
||||||
|
端口占用、`iphlpsvc` / portproxy、摄像头和依赖问题的集中排障入口见 [常见问题](/home/ray/dev/linkong/planet/docs/technical/zh/faq.md)。
|
||||||
|
|
||||||
### 1. 在 WSL 中启动服务
|
### 1. 在 WSL 中启动服务
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
./planet.sh start --allow-lan
|
./planet.sh start --allow-lan
|
||||||
```
|
```
|
||||||
|
|
||||||
这会让前端监听 `0.0.0.0:3000`,后端监听 `0.0.0.0:8000`。
|
这会让前端监听 `0.0.0.0:3000`,后端监听 `0.0.0.0:8000`,AI Provider 通过 Docker 发布到 `0.0.0.0:8010`。启动前脚本会检查这三个端口;如果 WSL/Linux 侧无法释放端口,并检测到 Windows 侧 listener 或旧 `portproxy`,会请求管理员 PowerShell 清理。
|
||||||
|
|
||||||
### 2. 先确认 WSL 内部服务正常
|
### 2. 先确认 WSL 内部服务正常
|
||||||
|
|
||||||
@@ -251,14 +235,16 @@ bun run build
|
|||||||
```bash
|
```bash
|
||||||
curl http://localhost:3000
|
curl http://localhost:3000
|
||||||
curl http://localhost:8000/health
|
curl http://localhost:8000/health
|
||||||
ss -ltnp | grep -E ':3000|:8000'
|
curl http://localhost:8010/health
|
||||||
|
ss -ltnp | grep -E ':3000|:8000|:8010'
|
||||||
```
|
```
|
||||||
|
|
||||||
预期:
|
预期:
|
||||||
|
|
||||||
- `3000` 返回前端 HTML
|
- `3000` 返回前端 HTML
|
||||||
- `8000/health` 返回健康检查 JSON
|
- `8000/health` 返回健康检查 JSON
|
||||||
- `ss` 中能看到 `0.0.0.0:3000` 和 `0.0.0.0:8000`
|
- `8010/health` 返回 AI Provider 健康检查 JSON
|
||||||
|
- `ss` 中能看到 `0.0.0.0:3000`、`0.0.0.0:8000` 和 `0.0.0.0:8010`,或 Docker 已发布 `8010`
|
||||||
|
|
||||||
如果这一步不通,先不要继续做 Windows 转发。
|
如果这一步不通,先不要继续做 Windows 转发。
|
||||||
|
|
||||||
@@ -269,42 +255,31 @@ ss -ltnp | grep -E ':3000|:8000'
|
|||||||
```powershell
|
```powershell
|
||||||
curl http://localhost:3000
|
curl http://localhost:3000
|
||||||
curl http://localhost:8000/health
|
curl http://localhost:8000/health
|
||||||
|
curl http://localhost:8010/health
|
||||||
```
|
```
|
||||||
|
|
||||||
在常见的 WSL2 开发环境下,Windows 通常可以直接通过 `localhost` 访问 WSL 中的服务。
|
在常见的 WSL2 开发环境下,Windows 通常可以直接通过 `localhost` 访问 WSL 中的服务。
|
||||||
|
|
||||||
### 4. 如果需要让局域网设备访问,再做 Windows 端口转发
|
### 4. 如果需要让局域网设备访问,清理端口和防火墙
|
||||||
|
|
||||||
注意:下面的命令必须在“以管理员身份运行”的 PowerShell 中执行。
|
`./planet.sh start --allow-lan` 不再启动额外的 Windows 端口转发进程。它直接让开发服务对 `3000` / `8000` / `8010` 开放,并在启动前尝试释放这些端口。端口被 Windows 侧 listener 或旧 `portproxy` 占用时,脚本会请求一次管理员 PowerShell 清理。
|
||||||
|
|
||||||
先把 Windows 对外网卡上的 `3000` / `8000` 转发到 Windows 本机 `127.0.0.1`:
|
如果以前手动配置过持久 `portproxy`,若自动请求被取消,可以手动清理,避免 `iphlpsvc` 继续占用端口:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=3000
|
netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=3000
|
||||||
netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=8000
|
netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=8000
|
||||||
|
netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=8010
|
||||||
netsh interface portproxy add v4tov4 listenaddress=0.0.0.0 listenport=3000 connectaddress=127.0.0.1 connectport=3000
|
|
||||||
netsh interface portproxy add v4tov4 listenaddress=0.0.0.0 listenport=8000 connectaddress=127.0.0.1 connectport=8000
|
|
||||||
```
|
```
|
||||||
|
|
||||||
再放行 Windows 防火墙:
|
脚本会检测 Windows 防火墙是否已放行 `3000` / `8000` / `8010`。如果缺少规则,会触发一次 Windows UAC 管理员 PowerShell 请求来自动创建。若自动请求被取消,也可以手动执行:
|
||||||
|
|
||||||
```powershell
|
```powershell
|
||||||
New-NetFirewallRule -DisplayName "WSL Planet 3000" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 3000
|
New-NetFirewallRule -DisplayName "WSL Planet 3000" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 3000
|
||||||
New-NetFirewallRule -DisplayName "WSL Planet 8000" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 8000
|
New-NetFirewallRule -DisplayName "WSL Planet 8000" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 8000
|
||||||
|
New-NetFirewallRule -DisplayName "WSL Planet 8010" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 8010
|
||||||
```
|
```
|
||||||
|
|
||||||
检查转发规则是否生效:
|
|
||||||
|
|
||||||
```powershell
|
|
||||||
netsh interface portproxy show all
|
|
||||||
```
|
|
||||||
|
|
||||||
预期能看到:
|
|
||||||
|
|
||||||
- `0.0.0.0:3000 -> 127.0.0.1:3000`
|
|
||||||
- `0.0.0.0:8000 -> 127.0.0.1:8000`
|
|
||||||
|
|
||||||
### 5. 查 Windows 局域网 IP,并让其他设备访问
|
### 5. 查 Windows 局域网 IP,并让其他设备访问
|
||||||
|
|
||||||
在 Windows PowerShell 中执行:
|
在 Windows PowerShell 中执行:
|
||||||
@@ -319,6 +294,8 @@ ipconfig
|
|||||||
|
|
||||||
- `http://<Windows局域网IP>:3000/earth`
|
- `http://<Windows局域网IP>:3000/earth`
|
||||||
- `http://<Windows局域网IP>:3000/admin`
|
- `http://<Windows局域网IP>:3000/admin`
|
||||||
|
- `http://<Windows局域网IP>:8000/health`
|
||||||
|
- `http://<Windows局域网IP>:8010/health`
|
||||||
|
|
||||||
例如:
|
例如:
|
||||||
|
|
||||||
@@ -327,7 +304,7 @@ ipconfig
|
|||||||
### 6. 常见现象与判断
|
### 6. 常见现象与判断
|
||||||
|
|
||||||
- WSL 中 `curl localhost:3000` 能通,但 Windows 访问 `WSL 的局域网 IP:3000` 不通:这是正常现象之一,优先验证 Windows 的 `localhost:3000`
|
- WSL 中 `curl localhost:3000` 能通,但 Windows 访问 `WSL 的局域网 IP:3000` 不通:这是正常现象之一,优先验证 Windows 的 `localhost:3000`
|
||||||
- Windows `localhost:3000` 能通,但局域网设备访问 `Windows 局域网 IP:3000` 不通:通常缺少 `portproxy` 或防火墙放行
|
- Windows `localhost:3000` 能通,但局域网设备访问 `Windows 局域网 IP:3000` 不通:通常是 Windows 防火墙、网络配置或旧 `portproxy` 残留
|
||||||
- `whoami /groups` 中 `S-1-5-32-544` 显示 `deny only`:说明当前 PowerShell 不是提权管理员窗口
|
- `whoami /groups` 中 `S-1-5-32-544` 显示 `deny only`:说明当前 PowerShell 不是提权管理员窗口
|
||||||
|
|
||||||
### 7. 本项目一次性验证顺序
|
### 7. 本项目一次性验证顺序
|
||||||
@@ -336,10 +313,12 @@ ipconfig
|
|||||||
|
|
||||||
1. WSL 中执行 `curl http://localhost:3000`
|
1. WSL 中执行 `curl http://localhost:3000`
|
||||||
2. WSL 中执行 `curl http://localhost:8000/health`
|
2. WSL 中执行 `curl http://localhost:8000/health`
|
||||||
3. Windows 中执行 `curl http://localhost:3000`
|
3. WSL 中执行 `curl http://localhost:8010/health`
|
||||||
4. Windows 中执行 `curl http://localhost:8000/health`
|
4. Windows 中执行 `curl http://localhost:3000`
|
||||||
5. 管理员 PowerShell 配置 `portproxy` 和防火墙
|
5. Windows 中执行 `curl http://localhost:8000/health`
|
||||||
6. 用手机或其他电脑访问 `http://<Windows局域网IP>:3000/earth`
|
6. Windows 中执行 `curl http://localhost:8010/health`
|
||||||
|
7. 按脚本提示完成 Windows 防火墙或端口清理 UAC 请求
|
||||||
|
8. 用手机或其他电脑访问 Windows 对外端口,例如 `http://<Windows局域网IP>:3000/earth`
|
||||||
|
|
||||||
## 启动容错参数
|
## 启动容错参数
|
||||||
|
|
||||||
@@ -363,18 +342,27 @@ DATABASE_RETRY_INTERVAL=10 \
|
|||||||
- `AI_PROVIDER_START_MAX_RETRIES` / `AI_PROVIDER_RETRY_INTERVAL`: 控制 `aiprovider` 的构建/启动与容器重启自愈,默认 `3` 次、`5` 秒
|
- `AI_PROVIDER_START_MAX_RETRIES` / `AI_PROVIDER_RETRY_INTERVAL`: 控制 `aiprovider` 的构建/启动与容器重启自愈,默认 `3` 次、`5` 秒
|
||||||
- `BACKEND_MAX_RETRIES`: 控制后端进程启动重试次数,默认 `3`
|
- `BACKEND_MAX_RETRIES`: 控制后端进程启动重试次数,默认 `3`
|
||||||
- `FRONTEND_MAX_RETRIES`: 控制前端 dev server 启动重试次数,默认 `3`
|
- `FRONTEND_MAX_RETRIES`: 控制前端 dev server 启动重试次数,默认 `3`
|
||||||
- `BACKEND_HEALTH_CHECK_ATTEMPTS` / `BACKEND_HEALTH_CHECK_INTERVAL`: 控制后端 HTTP 健康检查等待次数与间隔,默认 `10` 次、`2` 秒
|
- `BACKEND_HEALTH_CHECK_ATTEMPTS` / `BACKEND_HEALTH_CHECK_INTERVAL`: 控制后端 HTTP 健康检查等待次数与间隔,默认 `60` 次、`2` 秒
|
||||||
- `FRONTEND_HEALTH_CHECK_ATTEMPTS` / `FRONTEND_HEALTH_CHECK_INTERVAL`: 控制前端 HTTP 可访问检查等待次数与间隔,默认 `10` 次、`2` 秒
|
- `FRONTEND_HEALTH_CHECK_ATTEMPTS` / `FRONTEND_HEALTH_CHECK_INTERVAL`: 控制前端 HTTP 可访问检查等待次数与间隔,默认 `10` 次、`2` 秒
|
||||||
- `AI_PROVIDER_HEALTH_CHECK_ATTEMPTS` / `AI_PROVIDER_HEALTH_CHECK_INTERVAL`: 控制 `aiprovider` HTTP 健康检查等待次数与间隔,默认 `10` 次、`2` 秒
|
- `AI_PROVIDER_HEALTH_CHECK_ATTEMPTS` / `AI_PROVIDER_HEALTH_CHECK_INTERVAL`: 控制 `aiprovider` HTTP 健康检查等待次数与间隔,默认 `10` 次、`2` 秒
|
||||||
|
|
||||||
## AI 接口预留
|
## AI 与智能体接口
|
||||||
|
|
||||||
项目现在采用“两层”设计:
|
项目现在采用“三段式”边界:
|
||||||
|
|
||||||
- 主后端暴露稳定业务接口: `GET /api/v1/ai/provider/status`、`POST /api/v1/ai/situational-awareness/analyze`
|
- `backend`: 暴露业务接口,负责选择任务提示词、组织证据、调用工具、保存 AI 设置和结果。
|
||||||
- 独立 `aiprovider` 服务负责适配具体模型供应商
|
- `aiprovider`: 暴露模型网关接口,只负责 provider / protocol 适配,不写入 BGP、新闻、告警等业务提示词。
|
||||||
|
- 模型供应商: OpenAI 兼容、MiniMax、Anthropic、Ollama 或其他兼容网关。
|
||||||
|
|
||||||
这样前端和业务代码不直接依赖 OpenAI、本地模型网关或其他订阅服务,后续切换部署方式只需要调整环境变量。
|
这样前端和业务代码不直接依赖某个模型供应商,后续增加 Agent Runtime、Earth 一键 LLM 指令、语音识别或多角色态势研判时,也可以把业务工作流放在后端,而不是污染模型适配层。
|
||||||
|
|
||||||
|
当前已落地的 AI 配置能力:
|
||||||
|
|
||||||
|
- 运维台 AI 设置可维护 provider、模型、协议、超时、token 等运行配置。
|
||||||
|
- 运维台 AI 设置中的“提示词”页可选择不同功能入口,手动覆盖提示词,并一键重置到默认值。
|
||||||
|
- 默认提示词随代码发布,位于 [backend/app/ai_tasks/default_prompts.json](/home/ray/dev/linkong/planet/backend/app/ai_tasks/default_prompts.json)。
|
||||||
|
- 覆盖值保存在数据库运行配置中,升级代码后可继续保留现场配置,也可重置到新版本默认提示词。
|
||||||
|
- 态势摘要、告警研判、新闻本地化等入口应使用各自任务提示词;调用 `aiprovider` 时只传递当前任务所需的 `prompt` / `system_prompt`。
|
||||||
|
|
||||||
主后端建议配置:
|
主后端建议配置:
|
||||||
|
|
||||||
@@ -387,39 +375,33 @@ AI_PROVIDER_TIMEOUT_SECONDS=60
|
|||||||
`aiprovider` 服务建议配置:
|
`aiprovider` 服务建议配置:
|
||||||
|
|
||||||
```env
|
```env
|
||||||
AI_PROVIDER=openai_compatible
|
AI_PROVIDER=minimax
|
||||||
AI_BASE_URL=https://api.openai.com/v1
|
AI_PROVIDER_API=anthropic-messages
|
||||||
|
AI_BASE_URL=https://api.minimaxi.com/anthropic
|
||||||
AI_API_KEY=your_api_key
|
AI_API_KEY=your_api_key
|
||||||
AI_MODEL=gpt-4o-mini
|
AI_MODEL=MiniMax-M2.7
|
||||||
AI_TIMEOUT_SECONDS=60
|
AI_TIMEOUT_SECONDS=60
|
||||||
AI_PROVIDER_SERVICE_TOKEN=change_me
|
AI_PROVIDER_SERVICE_TOKEN=change_me
|
||||||
```
|
```
|
||||||
|
|
||||||
OpenAI 兼容场景推荐使用:
|
推荐映射关系:
|
||||||
|
|
||||||
- `AI_PROVIDER=openai_compatible`
|
- `vLLM` / `LM Studio` / `One API`: `AI_PROVIDER=openai` + `AI_PROVIDER_API=openai-completions`
|
||||||
|
- `MiniMax`: `AI_PROVIDER=minimax` + `AI_PROVIDER_API=anthropic-messages`
|
||||||
|
- Claude 兼容网关: `AI_PROVIDER=anthropic` + `AI_PROVIDER_API=anthropic-messages`
|
||||||
|
- `Ollama`: `AI_PROVIDER=ollama` + `AI_PROVIDER_API=ollama-generate`
|
||||||
|
|
||||||
Claude 兼容场景推荐使用:
|
比如 MiniMax 可以这样配置:
|
||||||
|
|
||||||
- `AI_PROVIDER=anthropic`
|
|
||||||
- `AI_PROVIDER=anthropic_compatible`
|
|
||||||
- `AI_PROVIDER=claude_compatible`
|
|
||||||
|
|
||||||
Ollama 原生场景推荐使用:
|
|
||||||
|
|
||||||
- `AI_PROVIDER=ollama`
|
|
||||||
|
|
||||||
比如 MiniMax 或其他 Claude 兼容网关,可以这样配置:
|
|
||||||
|
|
||||||
```env
|
```env
|
||||||
AI_PROVIDER=claude_compatible
|
AI_PROVIDER=minimax
|
||||||
AI_BASE_URL=https://your-claude-compatible-endpoint.example.com
|
AI_PROVIDER_API=anthropic-messages
|
||||||
|
AI_BASE_URL=https://api.minimaxi.com/anthropic
|
||||||
AI_API_KEY=your_api_key
|
AI_API_KEY=your_api_key
|
||||||
AI_MODEL=your-claude-compatible-model
|
AI_MODEL=MiniMax-M2.7
|
||||||
AI_TIMEOUT_SECONDS=60
|
AI_TIMEOUT_SECONDS=60
|
||||||
AI_MAX_TOKENS=1200
|
AI_MAX_TOKENS=1200
|
||||||
AI_ANTHROPIC_VERSION=2023-06-01
|
AI_ANTHROPIC_VERSION=2023-06-01
|
||||||
AI_PROVIDER_SERVICE_TOKEN=change_me
|
|
||||||
```
|
```
|
||||||
|
|
||||||
如果你要本地直接起模型适配层,项目里已经补了模板:
|
如果你要本地直接起模型适配层,项目里已经补了模板:
|
||||||
@@ -427,12 +409,6 @@ AI_PROVIDER_SERVICE_TOKEN=change_me
|
|||||||
- [aiprovider/.env.example](/home/ray/dev/linkong/planet/aiprovider/.env.example)
|
- [aiprovider/.env.example](/home/ray/dev/linkong/planet/aiprovider/.env.example)
|
||||||
- [docker-compose.local-model.yml](/home/ray/dev/linkong/planet/docker-compose.local-model.yml)
|
- [docker-compose.local-model.yml](/home/ray/dev/linkong/planet/docker-compose.local-model.yml)
|
||||||
|
|
||||||
推荐映射关系:
|
|
||||||
|
|
||||||
- `vLLM` / `LM Studio` / `One API`: `AI_PROVIDER=openai_compatible`
|
|
||||||
- `MiniMax` / Claude 兼容网关: `AI_PROVIDER=claude_compatible`
|
|
||||||
- `Ollama`: `AI_PROVIDER=ollama`
|
|
||||||
|
|
||||||
运行与调用补充:
|
运行与调用补充:
|
||||||
|
|
||||||
- `./planet.sh start` 默认会启动 `aiprovider`
|
- `./planet.sh start` 默认会启动 `aiprovider`
|
||||||
@@ -442,11 +418,13 @@ AI_PROVIDER_SERVICE_TOKEN=change_me
|
|||||||
|
|
||||||
详细文档:
|
详细文档:
|
||||||
|
|
||||||
- [docs/technical/agents-aiprovider.md](/home/ray/dev/linkong/planet/docs/technical/agents-aiprovider.md)
|
- [docs/technical/zh/agents-aiprovider.md](/home/ray/dev/linkong/planet/docs/technical/zh/agents-aiprovider.md)
|
||||||
|
- [docs/technical/en/agents-aiprovider.md](/home/ray/dev/linkong/planet/docs/technical/en/agents-aiprovider.md)
|
||||||
- [aiprovider/README.md](/home/ray/dev/linkong/planet/aiprovider/README.md)
|
- [aiprovider/README.md](/home/ray/dev/linkong/planet/aiprovider/README.md)
|
||||||
- [docs/technical/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/technical/frontend-layout-guidelines.md)
|
- [docs/technical/frontend-layout-guidelines.md](/home/ray/dev/linkong/planet/docs/technical/frontend-layout-guidelines.md)
|
||||||
- [docs/plans/frontend-ai-playground-development-plan.md](/home/ray/dev/linkong/planet/docs/plans/frontend-ai-playground-development-plan.md)
|
- [docs/plans/frontend-ai-playground-development-plan.md](/home/ray/dev/linkong/planet/docs/plans/frontend-ai-playground-development-plan.md)
|
||||||
- [docs/plans/agents-situational-awareness-foundation-plan.md](/home/ray/dev/linkong/planet/docs/plans/agents-situational-awareness-foundation-plan.md)
|
- [docs/plans/agents-situational-awareness-foundation-plan.md](/home/ray/dev/linkong/planet/docs/plans/agents-situational-awareness-foundation-plan.md)
|
||||||
|
- [docs/plans/agents-earth-command-runtime-plan.md](/home/ray/dev/linkong/planet/docs/plans/agents-earth-command-runtime-plan.md)
|
||||||
|
|
||||||
## 前端页面布局规范
|
## 前端页面布局规范
|
||||||
|
|
||||||
|
|||||||
120
TODO.md
120
TODO.md
@@ -1,37 +1,87 @@
|
|||||||
# TODO
|
# TODO
|
||||||
|
|
||||||
- [x] 把 BGP 观测站和异常点的 `hover/click` 手感再磨细一点
|
This file is the active backlog only. Completed history belongs in `docs/CHANGELOG.md`; detailed designs belong in `docs/plans/`.
|
||||||
- [x] 开始做 BGP 异常和海缆/区域的关联展示
|
|
||||||
- [x] 做 Earth 侧的 `BGP activity layer`,让低 incident 密度时地图仍然有持续可感知的观测存在感
|
## Earth
|
||||||
- [x] 给 Earth BGP 补三层状态表达:`平稳观测态 / 局部波动态 / 事件活跃态`
|
|
||||||
- [x] 把“当前无活跃事件”改造成“观测网络仍在运行、当前未发现聚合级事件”的状态表达
|
- [ ] Earth AI command entry: merge natural-language and speech-triggered LLM commands into the existing Earth search panel as described in [Agent Runtime, Earth LLM Command, And Speech Entry Plan](/home/ray/dev/linkong/planet/docs/plans/agents-earth-command-runtime-plan.md).
|
||||||
- [x] 做 collector / region 近 15 分钟 activity score 聚合接口或动态聚合逻辑
|
- [ ] Earth action executor: implement safe visualization actions for layer toggles, batch highlights, filters, focus, result panels, and clear-highlight behavior.
|
||||||
- [x] 把 Earth 的 BGP incident 改成 `紧凑事件核 + 向外扩张环形 pulse`,替换当前大面积 glow
|
- [ ] Earth entity matching: support stable entity ids and batch matching for Beidou satellites, mainland China compute centers, BGP, news, vessels, and cables.
|
||||||
- [x] 为 BGP incident 建立符号系统:按事件类型用不同 marker,而不是都用同一种亮点
|
- [ ] Replace debug GeoJSON boundary tiles with the real `earth-boundaries-china-pov-v1.pmtiles` production artifact after audited admin-0 / coastline / claim-line sources and the PMTiles toolchain are available.
|
||||||
- [x] 把 incident 地理定位从 `collector-centric` 改成 `prefix-centric`,优先使用 `prefix_geography`,其次 `prefix_scope`,再次 ASN 区域,最后才回退到观测区域质心
|
- [ ] Import authoritative China POV / coastline / claim-line source packages through the three standard Earth boundary source collectors, then rebuild a versioned PMTiles artifact so highest zoom `8-10` preserves trusted source geometry instead of seed data.
|
||||||
- [x] 新增 `prefix_geography` 数据层,不再把 `prefix_scope` 当成 prefix 地理归属本身
|
- [ ] Earth boundary data: acquire or generate auditable China POV geometry for Zangnan, Aksai Chin, Taiwan/Penghu, Diaoyu Dao and affiliated islands, Chiwei Yu, South China Sea islands, Kosovo, Gaza, and the official dashed maritime claim line before implementing final visual changes.
|
||||||
- [x] 接入 `IPtoASN / IPtoCountry` 作为 prefix-centric geography 的主数据源
|
- [ ] Earth high-resolution basemap tiles: implement the viewport-loaded imagery layer described in [Earth High Resolution Basemap Tiles Plan](/home/ray/dev/linkong/planet/docs/plans/earth-high-resolution-basemap-tiles-plan.md), using high-precision coastline as the alignment reference instead of replacing the globe with one huge texture.
|
||||||
- [x] 接入 `OpenGeoFeed` 作为 prefix geography 的高质量覆盖/override 数据源
|
- [ ] Presentation controller ownership: replace the singleton card fallback in [presentation-controller.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/presentation-controller.js) with a presentation/card token check before BGP/News migrate onto the shared controller, so connectors only attach to their owning card.
|
||||||
- [x] 把 RIR delegated 设计成 prefix geography 的 fallback,而不是主来源
|
- [ ] BGP frontend maintainability: split [bgp.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/bgp.js) by responsibility into data loading, marker rendering, overlays, and animation once the current interaction behavior is stable.
|
||||||
- [ ] 为 `aiprovider` 建立 `provider -> api adapter -> compat policy` 的配置中心,优先落成 `json` 或 `yaml` 文件,运行时按 `provider/model` 读取兼容设置,而不是把专项兼容继续散落在 Python 分支里
|
- [ ] Optional BGP marker experiment: evaluate HTML markers for BGP incident/collector points if WebGL marker density or fixed screen-size clickability becomes a real blocker.
|
||||||
- [ ] 为市面上主流 AI 服务补专项兼容配置并固化到配置文件中,至少覆盖 `OpenAI / Anthropic / MiniMax / Ollama / Moonshot / DeepSeek / Qwen / GLM / Gemini / OpenRouter / vLLM / LM Studio / One API`
|
- [ ] Earth news cruise: connect Earth news to the generic cruise queue via a news adapter rather than coupling news-specific sequencing into `main.js`.
|
||||||
- [ ] 在兼容配置中补齐可声明项:`api adapter`、`base_url pattern`、`auth header`、`thinking default`、`reasoning block mapping`、`stream path`、`tool-call capability`、`multimodal capability`、`provider-specific request patch`
|
|
||||||
- [ ] 接入 `inetnum` / `inet6num` whois 作为比 RIR 更细粒度的后备层
|
## Compute Centers And Location
|
||||||
- [x] 在 activity layer 之后继续补 `route leak` 和 `path instability / flap` detector
|
|
||||||
- [ ] 对 [frontend/public/earth/js/bgp.js](/home/ray/dev/linkong/planet/frontend/public/earth/js/bgp.js) 做按职责拆分的小重构,拆成 data / markers / overlays / animation,降低后续维护复杂度
|
- [ ] Unknown compute-center locations: continue reducing unresolved records through the shared location pipeline, with confidence, precision, reason, and verification date preserved in GeoJSON/details.
|
||||||
- [ ] 可选优化(非必做):将 BGP incident/collector 标点改为 HTML marker(参考 worldmonitor 的 `htmlElementsData` 思路),实现近乎固定屏幕尺寸与更高密度可点击性
|
- [ ] Compute-center registry: keep expanding the local canonical location registry with `canonical_name`, aliases, operator, country/region/city, coordinates, confidence, and source notes.
|
||||||
- [ ] 保持 Earth 当前这批纯个人偏好设置继续走本地持久化:`旋转模式`、HUD 面板显示/隐藏、`地形透明度` 暂不升级到后端系统设置,避免把设备级偏好过早做成全局配置
|
- [ ] Compute-center enrichment: improve source-page parsing for Epoch AI and related collectors by extracting location clues from detail pages, embedded JSON, schema.org, OpenGraph, script variables, PDFs, and press releases.
|
||||||
- [ ] 如果后续明确需要“账号级同步 Earth 偏好”,再单独设计 `Earth user preferences`:优先按用户维度而不是全局系统设置保存,并规划 `localStorage -> backend` 的平滑迁移策略
|
- [ ] Compute-center identity normalization: normalize operator / cluster / facility aliases such as `xAI / Colossus / Memphis`, `OpenAI / Stargate`, `CoreWeave`, `Lambda`, and `Crusoe`.
|
||||||
- [ ] 为 Planet / Earth 补一个可用的日志查看系统:先明确前后端/AI Provider/采集任务的日志入口、最近日志聚合、筛选与 tail 能力,再决定是先做脚本级统一入口还是控制台内置日志面板
|
- [ ] Compute-center manual review: add an export/review/import workflow for unresolved or estimated locations and feed confirmed results back into the registry.
|
||||||
- [ ] 把 Earth 态势新闻源从 [earth_news.py](/home/ray/dev/linkong/planet/backend/app/services/earth_news.py) 的硬编码列表抽成可配置目录,优先保持当前“实时聚合”链路不变,只先解决新闻源不可配置的问题
|
|
||||||
- [ ] 为 Earth 态势新闻设计后续采集器化方案:明确新闻数据模型、去重策略、区域映射、过期清理和 Earth/AI 复用方式,再决定何时把新闻从实时抓取升级成正式 collector
|
## AIS / Vessels
|
||||||
- [ ] 为 Earth 地球表面增加一层与基础纹理对齐的材质/纹理 overlay,并在同层叠加国界轮廓参考线;要求国界线与底图稳定对齐,且 hover 到国家轮廓时能高亮当前国家,便于校准地表和增强交互
|
|
||||||
- [ ] 把 Earth 新闻接入通用巡航队列:按新闻发生地和时间排序生成巡航目标,巡航聚焦到新闻事件时显示对应新闻卡片,并保持实现边界为“通用巡航层 + 新闻业务适配层”,不要再把新闻逻辑直接耦合回 `main.js` 状态机
|
- [ ] AIS aggregation strategy v4: expose source priority, field-level merge rules, freshness windows, and protected dynamic-field rules in configuration, with validation and strategy version returned by vessel APIs.
|
||||||
- [ ] 为未知位置的算力中心建立分层坐标补全链路:优先 `精确坐标 > 站点/园区命中 > 城市 > 州/省 > 国家内主要算力城市 > 国家质心`,并把每次回退的 `confidence / reason / precision` 明确写进统一 GeoJSON
|
- [ ] AIS vessel enrichment v5: add asynchronous vessel profile enrichment for ship type detail, AIS class, flag, dimensions, build year, operator, and cached media. Do not fetch third-party pages in the realtime AIS request path.
|
||||||
- [ ] 为算力中心补一份可维护的本地位置注册表,例如 `canonical_name / aliases / operator / country / region / city / lat / lon / confidence / source_note`,避免把地点知识长期硬编码在 `visualization.py`
|
- [ ] AIS identity cleanup: continue identifying vessels whose display name is only `MMSI <number>` and backfill names from AISStream static messages, BarentsWatch static fields, or enrichment cache.
|
||||||
- [ ] 增强 `epoch_ai_gpu` 和相关算力采集器的源页面解析:即使公开 API 不给坐标,也继续尝试从详情页、HTML、内嵌 JSON、schema.org、OpenGraph、脚本变量和 PDF/新闻稿链接里抽地点线索
|
|
||||||
- [ ] 为未知位置算力中心增加外部富化策略评估:可选接入公开知识源或搜索兜底,只抓“站点名/园区名/城市名”级别线索,不直接抓经纬度结论,并把结果作为候选证据而不是真值
|
## AI Provider And Agents
|
||||||
- [ ] 为算力中心建立 `operator / cluster name / facility alias` 归一化层,先解决 `xAI / Colossus / Memphis`、`OpenAI / Stargate`、`CoreWeave`、`Lambda`、`Crusoe` 这类同一对象多种写法导致的地点匹配失败
|
|
||||||
- [ ] 为估算位置增加更细的视觉和产品表达:除了问号角标,还要支持 tooltip/详情中的“估算依据”“精度级别”“最后核验时间”,并允许在设置中单独开关“仅看精确位置”
|
- [ ] Unified integration config schema: implement the shared low-code schema engine for datasource, AI Provider, Web Search, and OCR configuration described in [Integration Config Schema System Plan](/home/ray/dev/linkong/planet/docs/plans/integration-config-schema-system-plan.md).
|
||||||
- [ ] 为国家级估算点设计更合理的落点策略:优先落在“该国主要算力/数据中心城市候选集”而不是几何质心,必要时同国多节点做稳定散列分配,避免大量节点堆在荒漠或海上
|
- [ ] AI provider routing: finish the OpenClaw-style provider/model routing refactor described in [AI Provider OpenClaw-Style Routing Plan](/home/ray/dev/linkong/planet/docs/plans/ai-provider-openclaw-style-routing-plan.md), so model-specific transport rules live in provider metadata rather than runtime hardcoding.
|
||||||
- [ ] 为未知位置算力中心建立人工校验工作流:支持导出待核验清单、记录人工确认结果,并把人工确认反哺到位置注册表,逐步减少问号点比例
|
- [ ] AI provider catalog: replace the temporary `model_provider_apis` bridge with structured `models_metadata`, discovery descriptors, and incremental model sync with stale marking.
|
||||||
|
- [ ] AI provider connectivity: keep the plug action as lightweight network/auth/model-directory validation only, and keep real generation tests inside Playground or explicit “trial run” actions.
|
||||||
|
- [ ] Agent runtime foundation: add auditable agent runs, steps, evidence, proposals, and the Agent operations UI described in [Agent Runtime, Earth LLM Command, And Speech Entry Plan](/home/ray/dev/linkong/planet/docs/plans/agents-earth-command-runtime-plan.md).
|
||||||
|
- [ ] Agent tool protocol: add backend JSON tool-call fallback, optional provider-native tool compatibility, tool whitelist validation, and policy-gated proposal application.
|
||||||
|
- [ ] Speech/ASR integration for agents: add provider-neutral transcription settings and API, defaulting to Whisper-compatible API providers while keeping text commands usable when ASR is unavailable.
|
||||||
|
- [ ] Earth voice wake: add device-local configurable wake-word preferences, microphone fallback states, and post-wake instruction upload for Earth commands.
|
||||||
|
- [ ] AI provider compatibility center: move provider/model compatibility rules into a JSON/YAML config read by runtime, instead of continuing to scatter provider-specific branches through Python code.
|
||||||
|
- [ ] Provider compatibility coverage: add explicit config for OpenAI, Anthropic, MiniMax, Ollama, Moonshot, DeepSeek, Qwen, GLM, Gemini, OpenRouter, vLLM, LM Studio, and One API.
|
||||||
|
- [ ] Compatibility schema: cover adapter type, base URL pattern, auth header, thinking/reasoning defaults, stream path, tool-call capability, multimodal capability, and provider-specific request patches.
|
||||||
|
- [ ] BGP geography fallback: evaluate `inetnum` / `inet6num` whois as a finer fallback layer after `prefix_geography`, `OpenGeoFeed`, and RIR delegated data.
|
||||||
|
|
||||||
|
## Platform
|
||||||
|
|
||||||
|
- [ ] Earth preferences scope: keep current device-local Earth preferences in `localStorage`; only design backend user preferences if account-level synchronization becomes a real product requirement.
|
||||||
|
- [ ] System logs: finish a usable Planet log viewing flow that covers backend, frontend, AI Provider, and collector/task logs, with filtering and tailing.
|
||||||
|
- [ ] Console UI modernization: gradually replace Ant Design with Planet-owned components and a consistent Tabler Icons based icon system.
|
||||||
|
- [ ] Earth live sync: design a unified realtime invalidation path for summary/BGP/satellite updates if polling and current WebSocket channels become insufficient.
|
||||||
|
|
||||||
|
## Archive
|
||||||
|
|
||||||
|
Archived items stay here so old context is not lost. Completed items remain checked; obsolete, invalid, or superseded items stay unchecked and include the reason.
|
||||||
|
|
||||||
|
### Completed
|
||||||
|
|
||||||
|
- [x] Implemented the high-precision country boundary tile framework from [Earth High Precision Boundary Tiles Plan](/home/ray/dev/linkong/planet/docs/plans/earth-high-precision-boundary-tiles-plan.md): static vector tile builder, versioned seed output, frontend bbox tile loader, debounce, in-flight dedupe, and LRU cache.
|
||||||
|
- [x] Added the `pmtiles-mvt` frontend tile provider contract, MVT decoder dependencies, static PMTiles Nginx handling, collector artifact registration, production readiness check, and user operation docs for Earth boundaries.
|
||||||
|
- [x] Split Earth boundary ingestion into standard source collectors (`earth_admin0_boundaries`, `earth_coastline`, `earth_claim_lines`) plus the downstream `earth_boundary_tiles` PMTiles builder.
|
||||||
|
- [x] Refined BGP observer and anomaly `hover/click` feel.
|
||||||
|
- [x] Added BGP anomaly relationship display with cables / regions.
|
||||||
|
- [x] Added the Earth BGP activity layer so the map still feels alive when incident density is low.
|
||||||
|
- [x] Added BGP state expression for stable observation, local fluctuation, and active incident states.
|
||||||
|
- [x] Reframed "no active incident" as "observation network is running; no aggregate incident detected".
|
||||||
|
- [x] Added collector / region recent activity scoring.
|
||||||
|
- [x] Replaced oversized BGP incident glow with compact incident core plus outward pulse rings.
|
||||||
|
- [x] Added BGP incident symbol types instead of using one generic bright marker.
|
||||||
|
- [x] Switched BGP incident geography from collector-centric to prefix-centric priority.
|
||||||
|
- [x] Added `prefix_geography` as a separate data layer instead of treating `prefix_scope` as prefix geography.
|
||||||
|
- [x] Added IPtoASN / IPtoCountry as the main prefix-centric geography source.
|
||||||
|
- [x] Added OpenGeoFeed as a high-quality prefix geography override source.
|
||||||
|
- [x] Made RIR delegated data a prefix geography fallback rather than the primary source.
|
||||||
|
- [x] Added route leak and path instability / flap detectors after the activity layer work.
|
||||||
|
|
||||||
|
### Obsolete Or Superseded
|
||||||
|
|
||||||
|
- [ ] AIS v3.1 old `/geo/vessels` full-merge requirement. Superseded by `/api/v1/vessels/snapshot`, controlled legacy fallback, and diagnostics in the AIS aggregation plan.
|
||||||
|
- [ ] AIS v3.2 old framing of AISStream as a batch collector that needed conversion. Superseded by the implemented long-lived AISStream collector and realtime stream UI.
|
||||||
|
- [ ] AIS v3.3 old one-shot REST progress semantics for AISStream. Superseded by realtime stream status handling.
|
||||||
|
- [ ] AIS v3.4 broad identity cleanup wording. Folded into the active AIS identity cleanup and v5 enrichment tasks.
|
||||||
|
- [ ] Earth surface material overlay for boundary calibration. Superseded by the high-precision boundary tile plan; future work must use source-faithful boundary/coastline data rather than overlay calibration against the coarse base map.
|
||||||
|
- [ ] Hardcoded Earth news source extraction as a standalone task. Superseded by the broader Earth news source configuration and collector plans.
|
||||||
|
- [ ] Country-level compute-center fallback placement as a standalone task. Superseded by the shared location pipeline and registry/manual-review backlog.
|
||||||
|
|||||||
@@ -32,6 +32,15 @@ AI_API_KEY=sk-cp-change-me
|
|||||||
AI_MAX_TOKENS=1200
|
AI_MAX_TOKENS=1200
|
||||||
AI_ANTHROPIC_VERSION=2023-06-01
|
AI_ANTHROPIC_VERSION=2023-06-01
|
||||||
|
|
||||||
|
# Optional provider-specific keys used by Settings fallback before AI_API_KEY
|
||||||
|
# MINIMAX_API_KEY=sk-cp-change-me
|
||||||
|
# OPENAI_API_KEY=sk-change-me
|
||||||
|
# ANTHROPIC_API_KEY=sk-ant-change-me
|
||||||
|
# DEEPSEEK_API_KEY=sk-change-me
|
||||||
|
# DASHSCOPE_API_KEY=sk-change-me
|
||||||
|
# MOONSHOT_API_KEY=sk-change-me
|
||||||
|
# OPENROUTER_API_KEY=sk-or-change-me
|
||||||
|
|
||||||
# OpenAI-compatible example (vLLM / LM Studio / One API / local gateway)
|
# OpenAI-compatible example (vLLM / LM Studio / One API / local gateway)
|
||||||
# AI_PROVIDER=openai
|
# AI_PROVIDER=openai
|
||||||
# AI_PROVIDER_API=openai-completions
|
# AI_PROVIDER_API=openai-completions
|
||||||
|
|||||||
@@ -1,9 +1,15 @@
|
|||||||
|
# syntax=docker/dockerfile:1.7
|
||||||
|
|
||||||
ARG PYTHON_IMAGE=python:3.14-slim
|
ARG PYTHON_IMAGE=python:3.14-slim
|
||||||
ARG UV_IMAGE=ghcr.io/astral-sh/uv:latest
|
ARG UV_IMAGE=ghcr.io/astral-sh/uv:latest
|
||||||
|
ARG AI_PROVIDER_BUILD_FINGERPRINT=unknown
|
||||||
|
|
||||||
FROM ${UV_IMAGE} AS uv
|
FROM ${UV_IMAGE} AS uv
|
||||||
FROM ${PYTHON_IMAGE}
|
FROM ${PYTHON_IMAGE}
|
||||||
|
|
||||||
|
ARG AI_PROVIDER_BUILD_FINGERPRINT
|
||||||
|
LABEL planet.aiprovider.build-fingerprint="${AI_PROVIDER_BUILD_FINGERPRINT}"
|
||||||
|
|
||||||
COPY --from=uv /uv /uvx /bin/
|
COPY --from=uv /uv /uvx /bin/
|
||||||
|
|
||||||
WORKDIR /app
|
WORKDIR /app
|
||||||
@@ -13,15 +19,22 @@ ENV PYTHONUNBUFFERED=1
|
|||||||
ENV UV_COMPILE_BYTECODE=1
|
ENV UV_COMPILE_BYTECODE=1
|
||||||
ENV UV_LINK_MODE=copy
|
ENV UV_LINK_MODE=copy
|
||||||
|
|
||||||
|
RUN mkdir -p /root/.config/uv
|
||||||
|
|
||||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||||
curl \
|
curl \
|
||||||
&& rm -rf /var/lib/apt/lists/*
|
&& rm -rf /var/lib/apt/lists/*
|
||||||
|
|
||||||
COPY pyproject.toml uv.lock /app/
|
COPY pyproject.toml uv.lock /app/
|
||||||
RUN uv sync --frozen --no-dev
|
RUN --mount=type=cache,target=/root/.cache/uv \
|
||||||
|
--mount=type=secret,id=planet_uv_config,target=/root/.config/uv/uv.toml,required=false \
|
||||||
|
uv sync --frozen --no-dev
|
||||||
|
|
||||||
COPY . /app
|
COPY aiprovider /app/aiprovider
|
||||||
|
|
||||||
EXPOSE 8010
|
EXPOSE 8010
|
||||||
|
|
||||||
CMD ["uv", "run", "--frozen", "--no-dev", "--project", "/app", "python", "-m", "uvicorn", "aiprovider.main:app", "--host", "0.0.0.0", "--port", "8010", "--reload"]
|
HEALTHCHECK --interval=30s --timeout=5s --retries=3 \
|
||||||
|
CMD curl -fsS http://127.0.0.1:8010/health >/dev/null || exit 1
|
||||||
|
|
||||||
|
CMD ["uv", "run", "--frozen", "--no-dev", "--project", "/app", "python", "-m", "uvicorn", "aiprovider.main:app", "--host", "0.0.0.0", "--port", "8010"]
|
||||||
|
|||||||
@@ -17,9 +17,6 @@ class Settings(BaseSettings):
|
|||||||
AI_HTTP_RETRY_ATTEMPTS: int = 2
|
AI_HTTP_RETRY_ATTEMPTS: int = 2
|
||||||
AI_MAX_TOKENS: int = 1200
|
AI_MAX_TOKENS: int = 1200
|
||||||
AI_ANTHROPIC_VERSION: str = "2023-06-01"
|
AI_ANTHROPIC_VERSION: str = "2023-06-01"
|
||||||
AI_ANALYSIS_SYSTEM_PROMPT: str = (
|
|
||||||
"你是态势感知分析助手。请基于输入的上下文、观测与约束,输出结构化、克制、可执行的分析。"
|
|
||||||
)
|
|
||||||
|
|
||||||
AI_PROVIDER_SERVICE_TOKEN: str = ""
|
AI_PROVIDER_SERVICE_TOKEN: str = ""
|
||||||
|
|
||||||
|
|||||||
@@ -37,8 +37,28 @@ def verify_service_token(x_provider_token: str | None = Header(default=None)) ->
|
|||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
def get_provider_service() -> ProviderService:
|
def get_provider_service(
|
||||||
return ProviderService()
|
x_ai_provider: str | None = Header(default=None),
|
||||||
|
x_ai_provider_api: str | None = Header(default=None),
|
||||||
|
x_ai_base_url: str | None = Header(default=None),
|
||||||
|
x_ai_api_key: str | None = Header(default=None),
|
||||||
|
x_ai_model: str | None = Header(default=None),
|
||||||
|
x_ai_max_tokens: str | None = Header(default=None),
|
||||||
|
x_ai_anthropic_version: str | None = Header(default=None),
|
||||||
|
x_ai_model_provider_apis: str | None = Header(default=None),
|
||||||
|
) -> ProviderService:
|
||||||
|
overrides = {
|
||||||
|
"provider": x_ai_provider,
|
||||||
|
"provider_api": x_ai_provider_api,
|
||||||
|
"base_url": x_ai_base_url,
|
||||||
|
"api_key": x_ai_api_key,
|
||||||
|
"model": x_ai_model,
|
||||||
|
"anthropic_version": x_ai_anthropic_version,
|
||||||
|
"model_provider_apis": x_ai_model_provider_apis,
|
||||||
|
}
|
||||||
|
if x_ai_max_tokens:
|
||||||
|
overrides["max_tokens"] = x_ai_max_tokens
|
||||||
|
return ProviderService({key: value for key, value in overrides.items() if value not in (None, "")})
|
||||||
|
|
||||||
|
|
||||||
@app.get("/health")
|
@app.get("/health")
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import asyncio
|
import asyncio
|
||||||
|
import json
|
||||||
from typing import Any
|
from typing import Any
|
||||||
|
|
||||||
import httpx
|
import httpx
|
||||||
@@ -14,7 +15,6 @@ from aiprovider.schemas import (
|
|||||||
SituationalAnalysisResponse,
|
SituationalAnalysisResponse,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
def _normalize_provider(value: str) -> str:
|
def _normalize_provider(value: str) -> str:
|
||||||
return (value or "disabled").strip().lower()
|
return (value or "disabled").strip().lower()
|
||||||
|
|
||||||
@@ -46,20 +46,25 @@ def _resolve_provider_api(provider: str, configured_api: str) -> str:
|
|||||||
|
|
||||||
|
|
||||||
class ProviderService:
|
class ProviderService:
|
||||||
def __init__(self) -> None:
|
def __init__(self, overrides: dict[str, Any] | None = None) -> None:
|
||||||
self.provider = _normalize_provider(settings.AI_PROVIDER)
|
overrides = overrides or {}
|
||||||
|
self.provider = _normalize_provider(overrides.get("provider") or settings.AI_PROVIDER)
|
||||||
self.provider_api = _resolve_provider_api(
|
self.provider_api = _resolve_provider_api(
|
||||||
self.provider,
|
self.provider,
|
||||||
_normalize_provider_api(settings.AI_PROVIDER_API),
|
_normalize_provider_api(overrides.get("provider_api") or settings.AI_PROVIDER_API),
|
||||||
)
|
)
|
||||||
self.base_url = settings.AI_BASE_URL.rstrip("/")
|
self.base_url = str(overrides.get("base_url") or settings.AI_BASE_URL).rstrip("/")
|
||||||
self.api_key = settings.AI_API_KEY
|
self.api_key = str(overrides.get("api_key") or settings.AI_API_KEY)
|
||||||
self.default_model = settings.AI_MODEL
|
self.default_model = str(overrides.get("model") or settings.AI_MODEL)
|
||||||
self.timeout = settings.AI_TIMEOUT_SECONDS
|
self.timeout = settings.AI_TIMEOUT_SECONDS
|
||||||
self.http_retry_attempts = max(settings.AI_HTTP_RETRY_ATTEMPTS, 1)
|
self.http_retry_attempts = max(settings.AI_HTTP_RETRY_ATTEMPTS, 1)
|
||||||
self.max_tokens = settings.AI_MAX_TOKENS
|
self.max_tokens = int(overrides.get("max_tokens") or settings.AI_MAX_TOKENS)
|
||||||
self.anthropic_version = settings.AI_ANTHROPIC_VERSION
|
self.anthropic_version = str(
|
||||||
self.system_prompt = settings.AI_ANALYSIS_SYSTEM_PROMPT
|
overrides.get("anthropic_version") or settings.AI_ANTHROPIC_VERSION
|
||||||
|
)
|
||||||
|
self.model_provider_apis = self._parse_model_provider_apis(
|
||||||
|
overrides.get("model_provider_apis")
|
||||||
|
)
|
||||||
|
|
||||||
def get_status(self) -> AIProviderStatusResponse:
|
def get_status(self) -> AIProviderStatusResponse:
|
||||||
enabled = self.provider != "disabled"
|
enabled = self.provider != "disabled"
|
||||||
@@ -91,16 +96,23 @@ class ProviderService:
|
|||||||
|
|
||||||
prompt = self._build_prompt(payload)
|
prompt = self._build_prompt(payload)
|
||||||
|
|
||||||
if self.provider_api == "openai-completions":
|
provider_api = self._resolve_model_provider_api(model)
|
||||||
data = await self._request_openai_compatible(model, prompt)
|
|
||||||
|
if provider_api == "openai-completions":
|
||||||
|
data = await self._request_openai_compatible(model, prompt, payload.system_prompt)
|
||||||
content = self._extract_openai_content(data)
|
content = self._extract_openai_content(data)
|
||||||
content_blocks = self._extract_openai_blocks(data)
|
content_blocks = self._extract_openai_blocks(data)
|
||||||
elif self.provider_api == "anthropic-messages":
|
elif provider_api == "anthropic-messages":
|
||||||
data = await self._request_anthropic_messages(model, prompt, payload.thinking)
|
data = await self._request_anthropic_messages(
|
||||||
|
model,
|
||||||
|
prompt,
|
||||||
|
payload.thinking,
|
||||||
|
payload.system_prompt,
|
||||||
|
)
|
||||||
content = self._extract_anthropic_content(data)
|
content = self._extract_anthropic_content(data)
|
||||||
content_blocks = self._extract_anthropic_blocks(data)
|
content_blocks = self._extract_anthropic_blocks(data)
|
||||||
elif self.provider_api == "ollama-generate":
|
elif provider_api == "ollama-generate":
|
||||||
data = await self._request_ollama(model, prompt)
|
data = await self._request_ollama(model, prompt, payload.system_prompt)
|
||||||
content = self._extract_ollama_content(data)
|
content = self._extract_ollama_content(data)
|
||||||
content_blocks = self._extract_ollama_blocks(data)
|
content_blocks = self._extract_ollama_blocks(data)
|
||||||
else:
|
else:
|
||||||
@@ -125,6 +137,26 @@ class ProviderService:
|
|||||||
def _requires_api_key(self) -> bool:
|
def _requires_api_key(self) -> bool:
|
||||||
return self.provider_api != "ollama-generate"
|
return self.provider_api != "ollama-generate"
|
||||||
|
|
||||||
|
def _resolve_model_provider_api(self, model: str) -> str:
|
||||||
|
return self.model_provider_apis.get(model) or self.provider_api
|
||||||
|
|
||||||
|
def _parse_model_provider_apis(self, value: Any) -> dict[str, str]:
|
||||||
|
if isinstance(value, dict):
|
||||||
|
raw = value
|
||||||
|
elif isinstance(value, str) and value.strip():
|
||||||
|
try:
|
||||||
|
parsed = json.loads(value)
|
||||||
|
except json.JSONDecodeError:
|
||||||
|
return {}
|
||||||
|
raw = parsed if isinstance(parsed, dict) else {}
|
||||||
|
else:
|
||||||
|
raw = {}
|
||||||
|
return {
|
||||||
|
str(model): _normalize_provider_api(str(provider_api))
|
||||||
|
for model, provider_api in raw.items()
|
||||||
|
if model and provider_api
|
||||||
|
}
|
||||||
|
|
||||||
def _build_prompt(self, payload: SituationalAnalysisRequest) -> str:
|
def _build_prompt(self, payload: SituationalAnalysisRequest) -> str:
|
||||||
sections = [
|
sections = [
|
||||||
f"任务标题:\n{payload.title}",
|
f"任务标题:\n{payload.title}",
|
||||||
@@ -136,19 +168,28 @@ class ProviderService:
|
|||||||
sections.append("约束条件:\n" + "\n".join(f"- {item}" for item in payload.constraints))
|
sections.append("约束条件:\n" + "\n".join(f"- {item}" for item in payload.constraints))
|
||||||
if payload.context:
|
if payload.context:
|
||||||
sections.append(f"附加上下文:\n{payload.context}")
|
sections.append(f"附加上下文:\n{payload.context}")
|
||||||
sections.append(
|
|
||||||
"请输出: 1) 态势摘要 2) 关键风险 3) 研判依据 4) 建议动作 5) 还缺少的数据。"
|
|
||||||
)
|
|
||||||
return "\n\n".join(sections)
|
return "\n\n".join(sections)
|
||||||
|
|
||||||
async def _request_openai_compatible(self, model: str, prompt: str) -> dict[str, Any]:
|
def _resolve_system_prompt(self, system_prompt: str | None) -> str | None:
|
||||||
|
resolved = str(system_prompt or "").strip()
|
||||||
|
return resolved or None
|
||||||
|
|
||||||
|
async def _request_openai_compatible(
|
||||||
|
self,
|
||||||
|
model: str,
|
||||||
|
prompt: str,
|
||||||
|
system_prompt: str | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
messages = []
|
||||||
|
resolved_system_prompt = self._resolve_system_prompt(system_prompt)
|
||||||
|
if resolved_system_prompt:
|
||||||
|
messages.append({"role": "system", "content": resolved_system_prompt})
|
||||||
|
messages.append({"role": "user", "content": prompt})
|
||||||
request_body = {
|
request_body = {
|
||||||
"model": model,
|
"model": model,
|
||||||
"messages": [
|
"messages": messages,
|
||||||
{"role": "system", "content": self.system_prompt},
|
|
||||||
{"role": "user", "content": prompt},
|
|
||||||
],
|
|
||||||
"temperature": 0.2,
|
"temperature": 0.2,
|
||||||
|
"max_tokens": self.max_tokens,
|
||||||
}
|
}
|
||||||
return await self._post(
|
return await self._post(
|
||||||
path="/chat/completions",
|
path="/chat/completions",
|
||||||
@@ -164,10 +205,10 @@ class ProviderService:
|
|||||||
model: str,
|
model: str,
|
||||||
prompt: str,
|
prompt: str,
|
||||||
thinking: dict[str, Any] | None = None,
|
thinking: dict[str, Any] | None = None,
|
||||||
|
system_prompt: str | None = None,
|
||||||
) -> dict[str, Any]:
|
) -> dict[str, Any]:
|
||||||
request_body = {
|
request_body = {
|
||||||
"model": model,
|
"model": model,
|
||||||
"system": self.system_prompt,
|
|
||||||
"messages": [
|
"messages": [
|
||||||
{
|
{
|
||||||
"role": "user",
|
"role": "user",
|
||||||
@@ -182,6 +223,9 @@ class ProviderService:
|
|||||||
"max_tokens": self.max_tokens,
|
"max_tokens": self.max_tokens,
|
||||||
"temperature": 0.2,
|
"temperature": 0.2,
|
||||||
}
|
}
|
||||||
|
resolved_system_prompt = self._resolve_system_prompt(system_prompt)
|
||||||
|
if resolved_system_prompt:
|
||||||
|
request_body["system"] = resolved_system_prompt
|
||||||
resolved_thinking = self._resolve_anthropic_thinking(thinking)
|
resolved_thinking = self._resolve_anthropic_thinking(thinking)
|
||||||
if resolved_thinking:
|
if resolved_thinking:
|
||||||
request_body["thinking"] = resolved_thinking
|
request_body["thinking"] = resolved_thinking
|
||||||
@@ -215,19 +259,28 @@ class ProviderService:
|
|||||||
model: str,
|
model: str,
|
||||||
prompt: str,
|
prompt: str,
|
||||||
thinking: dict[str, Any] | None = None,
|
thinking: dict[str, Any] | None = None,
|
||||||
|
system_prompt: str | None = None,
|
||||||
) -> dict[str, Any]:
|
) -> dict[str, Any]:
|
||||||
return await self._request_anthropic_messages(model, prompt, thinking)
|
return await self._request_anthropic_messages(model, prompt, thinking, system_prompt)
|
||||||
|
|
||||||
async def _request_ollama(self, model: str, prompt: str) -> dict[str, Any]:
|
async def _request_ollama(
|
||||||
|
self,
|
||||||
|
model: str,
|
||||||
|
prompt: str,
|
||||||
|
system_prompt: str | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
request_body = {
|
request_body = {
|
||||||
"model": model,
|
"model": model,
|
||||||
"stream": False,
|
"stream": False,
|
||||||
"system": self.system_prompt,
|
|
||||||
"prompt": prompt,
|
"prompt": prompt,
|
||||||
"options": {
|
"options": {
|
||||||
"temperature": 0.2,
|
"temperature": 0.2,
|
||||||
|
"num_predict": self.max_tokens,
|
||||||
},
|
},
|
||||||
}
|
}
|
||||||
|
resolved_system_prompt = self._resolve_system_prompt(system_prompt)
|
||||||
|
if resolved_system_prompt:
|
||||||
|
request_body["system"] = resolved_system_prompt
|
||||||
return await self._post(
|
return await self._post(
|
||||||
path="/api/generate",
|
path="/api/generate",
|
||||||
headers={
|
headers={
|
||||||
@@ -286,13 +339,19 @@ class ProviderService:
|
|||||||
message = choices[0].get("message") or {}
|
message = choices[0].get("message") or {}
|
||||||
content = message.get("content")
|
content = message.get("content")
|
||||||
if isinstance(content, str):
|
if isinstance(content, str):
|
||||||
return content
|
if content:
|
||||||
|
return content
|
||||||
|
reasoning_content = message.get("reasoning_content")
|
||||||
|
return reasoning_content if isinstance(reasoning_content, str) else ""
|
||||||
if isinstance(content, list):
|
if isinstance(content, list):
|
||||||
return "".join(
|
return "".join(
|
||||||
item.get("text", "")
|
item.get("text", "")
|
||||||
for item in content
|
for item in content
|
||||||
if isinstance(item, dict)
|
if isinstance(item, dict)
|
||||||
)
|
)
|
||||||
|
reasoning_content = message.get("reasoning_content")
|
||||||
|
if isinstance(reasoning_content, str):
|
||||||
|
return reasoning_content
|
||||||
return ""
|
return ""
|
||||||
|
|
||||||
def _extract_openai_blocks(self, payload: dict[str, Any]) -> list[AIContentBlock]:
|
def _extract_openai_blocks(self, payload: dict[str, Any]) -> list[AIContentBlock]:
|
||||||
@@ -303,9 +362,14 @@ class ProviderService:
|
|||||||
message = choices[0].get("message") or {}
|
message = choices[0].get("message") or {}
|
||||||
content = message.get("content")
|
content = message.get("content")
|
||||||
if isinstance(content, str):
|
if isinstance(content, str):
|
||||||
return [AIContentBlock(type="text", text=content)]
|
blocks = [AIContentBlock(type="text", text=content)] if content else []
|
||||||
|
reasoning_content = message.get("reasoning_content")
|
||||||
|
if isinstance(reasoning_content, str) and reasoning_content:
|
||||||
|
blocks.append(AIContentBlock(type="thinking", thinking=reasoning_content))
|
||||||
|
return blocks
|
||||||
if not isinstance(content, list):
|
if not isinstance(content, list):
|
||||||
return []
|
reasoning_content = message.get("reasoning_content")
|
||||||
|
return [AIContentBlock(type="thinking", thinking=reasoning_content)] if isinstance(reasoning_content, str) and reasoning_content else []
|
||||||
|
|
||||||
blocks: list[AIContentBlock] = []
|
blocks: list[AIContentBlock] = []
|
||||||
for item in content:
|
for item in content:
|
||||||
@@ -318,7 +382,11 @@ class ProviderService:
|
|||||||
metadata={k: v for k, v in item.items() if k not in {"type", "text"}},
|
metadata={k: v for k, v in item.items() if k not in {"type", "text"}},
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
reasoning_content = message.get("reasoning_content")
|
||||||
|
if isinstance(reasoning_content, str) and reasoning_content:
|
||||||
|
blocks.append(AIContentBlock(type="thinking", thinking=reasoning_content))
|
||||||
return blocks
|
return blocks
|
||||||
|
|
||||||
def _extract_anthropic_content(self, payload: dict[str, Any]) -> str:
|
def _extract_anthropic_content(self, payload: dict[str, Any]) -> str:
|
||||||
content = payload.get("content")
|
content = payload.get("content")
|
||||||
if isinstance(content, str):
|
if isinstance(content, str):
|
||||||
|
|||||||
@@ -13,10 +13,11 @@ class AIContentBlock(BaseModel):
|
|||||||
|
|
||||||
class SituationalAnalysisRequest(BaseModel):
|
class SituationalAnalysisRequest(BaseModel):
|
||||||
title: str = Field(..., min_length=1, max_length=200)
|
title: str = Field(..., min_length=1, max_length=200)
|
||||||
objective: str = Field(..., min_length=1, max_length=1000)
|
objective: str = Field(..., min_length=1, max_length=20000)
|
||||||
context: dict[str, Any] = Field(default_factory=dict)
|
context: dict[str, Any] = Field(default_factory=dict)
|
||||||
observations: list[str] = Field(default_factory=list)
|
observations: list[str] = Field(default_factory=list)
|
||||||
constraints: list[str] = Field(default_factory=list)
|
constraints: list[str] = Field(default_factory=list)
|
||||||
|
system_prompt: str | None = Field(default=None, max_length=8000)
|
||||||
preferred_model: str | None = Field(default=None, max_length=200)
|
preferred_model: str | None = Field(default=None, max_length=200)
|
||||||
thinking: dict[str, Any] | None = None
|
thinking: dict[str, Any] | None = None
|
||||||
|
|
||||||
|
|||||||
@@ -1,3 +1,5 @@
|
|||||||
|
# syntax=docker/dockerfile:1.7
|
||||||
|
|
||||||
ARG PYTHON_IMAGE=python:3.14-slim
|
ARG PYTHON_IMAGE=python:3.14-slim
|
||||||
ARG UV_IMAGE=ghcr.io/astral-sh/uv:latest
|
ARG UV_IMAGE=ghcr.io/astral-sh/uv:latest
|
||||||
|
|
||||||
@@ -12,17 +14,25 @@ ENV PYTHONDONTWRITEBYTECODE=1
|
|||||||
ENV PYTHONUNBUFFERED=1
|
ENV PYTHONUNBUFFERED=1
|
||||||
ENV UV_COMPILE_BYTECODE=1
|
ENV UV_COMPILE_BYTECODE=1
|
||||||
ENV UV_LINK_MODE=copy
|
ENV UV_LINK_MODE=copy
|
||||||
|
ENV PYTHONPATH=/app/backend
|
||||||
|
|
||||||
|
RUN mkdir -p /root/.config/uv
|
||||||
|
|
||||||
RUN apt-get update && apt-get install -y --no-install-recommends \
|
RUN apt-get update && apt-get install -y --no-install-recommends \
|
||||||
curl \
|
curl \
|
||||||
&& rm -rf /var/lib/apt/lists/*
|
&& rm -rf /var/lib/apt/lists/*
|
||||||
|
|
||||||
COPY pyproject.toml uv.lock /app/
|
COPY pyproject.toml uv.lock /app/
|
||||||
RUN uv sync --frozen --no-dev
|
RUN --mount=type=cache,target=/root/.cache/uv \
|
||||||
|
--mount=type=secret,id=planet_uv_config,target=/root/.config/uv/uv.toml,required=false \
|
||||||
|
uv sync --frozen --no-dev
|
||||||
|
|
||||||
COPY backend /app/backend
|
COPY backend /app/backend
|
||||||
COPY VERSION /app/VERSION
|
COPY VERSION /app/VERSION
|
||||||
|
|
||||||
EXPOSE 8000
|
EXPOSE 8000
|
||||||
|
|
||||||
CMD ["uv", "run", "--frozen", "--no-dev", "--project", "/app", "python", "-m", "uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000", "--reload"]
|
HEALTHCHECK --interval=30s --timeout=5s --retries=3 \
|
||||||
|
CMD curl -fsS http://127.0.0.1:8000/health >/dev/null || exit 1
|
||||||
|
|
||||||
|
CMD ["uv", "run", "--frozen", "--no-dev", "--project", "/app", "python", "-m", "uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000"]
|
||||||
|
|||||||
2
backend/app/ai_tasks/__init__.py
Normal file
2
backend/app/ai_tasks/__init__.py
Normal file
@@ -0,0 +1,2 @@
|
|||||||
|
"""AI task prompt registry and runtime helpers."""
|
||||||
|
|
||||||
74
backend/app/ai_tasks/default_prompts.json
Normal file
74
backend/app/ai_tasks/default_prompts.json
Normal file
@@ -0,0 +1,74 @@
|
|||||||
|
[
|
||||||
|
{
|
||||||
|
"key": "earth.news.enrich",
|
||||||
|
"label": "Earth 新闻汉化与定位",
|
||||||
|
"group": "Earth 新闻",
|
||||||
|
"version": "2026-05-16.2",
|
||||||
|
"system_prompt": "",
|
||||||
|
"prompt": "Return exactly one strict JSON object with a location object and a localizations object. Infer the most likely physical event location and produce a faithful Simplified Chinese title plus a one-sentence newswire-style Chinese summary based only on the supplied RSS headline, description, source, and date. The summary should read like a concise breaking-news lead, not a label, slogan, or keyword headline."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "alerts.brief",
|
||||||
|
"label": "系统告警研判",
|
||||||
|
"group": "告警研判",
|
||||||
|
"version": "2026-05-16.1",
|
||||||
|
"system_prompt": "你是告警研判助手。请基于输入的告警事实、上下文与约束,输出结构化、克制、可执行的值班研判;明确区分事实、推断与建议,不要夸大证据不足的风险。",
|
||||||
|
"prompt": "基于当前告警总量、严重度、状态、数据源分布与最近告警摘录,生成一份面向值班人员的简明告警态势简报,突出待处理风险、告警集中点和优先动作。"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "alerts.situational.brief",
|
||||||
|
"label": "跨模块态势告警研判",
|
||||||
|
"group": "告警研判",
|
||||||
|
"version": "2026-05-16.1",
|
||||||
|
"system_prompt": "你是告警研判助手。请基于输入的告警事实、上下文与约束,输出结构化、克制、可执行的值班研判;明确区分事实、推断与建议,不要夸大证据不足的风险。",
|
||||||
|
"prompt": "综合系统告警、BGP incidents、BGP anomalies 与近期 BGP AI 简报,生成一份面向值班人员的态势告警简报,指出当前最需要关注的风险域、跨模块联动迹象和优先动作。"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "bgp.brief",
|
||||||
|
"label": "BGP 态势简报",
|
||||||
|
"group": "BGP",
|
||||||
|
"version": "2026-05-16.2",
|
||||||
|
"system_prompt": "你是 BGP 值班分析师。请直接输出面向值班人员的中文 Markdown 简报,只写最终研判内容;不要复述用户需求、提示词、写作计划、字段清单或“我将如何回答”。",
|
||||||
|
"prompt": "基于当前 BGP incidents、anomalies、原始观测事件、观测站覆盖与 prefix geography 证据,生成一份面向操作员的简明态势简报,突出区域热点、观测偏差、当前风险、证据和优先动作。"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "location.factcheck.normalize",
|
||||||
|
"label": "位置事实核查结构化",
|
||||||
|
"group": "位置解析",
|
||||||
|
"version": "2026-05-16.1",
|
||||||
|
"system_prompt": "",
|
||||||
|
"prompt": "Convert the supplied location factcheck text into exactly one strict JSON object. Extract only facts present in the text or original query."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "location.factcheck.resolve",
|
||||||
|
"label": "位置事实核查兜底",
|
||||||
|
"group": "位置解析",
|
||||||
|
"version": "2026-05-16.1",
|
||||||
|
"system_prompt": "",
|
||||||
|
"prompt": "Return exactly one JSON object for the most likely physical location. Use only fact-checkable public knowledge; return null fields rather than guessing when evidence is weak."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "datasource.mapping",
|
||||||
|
"label": "数据源映射生成",
|
||||||
|
"group": "采集配置",
|
||||||
|
"version": "2026-05-16.1",
|
||||||
|
"system_prompt": "",
|
||||||
|
"prompt": "Return only JSON for a deterministic mapping DSL. The JSON must contain source.items_path and fields. Do not include prose or code."
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "credential.guide",
|
||||||
|
"label": "采集器凭据教程",
|
||||||
|
"group": "采集配置",
|
||||||
|
"version": "2026-05-16.1",
|
||||||
|
"system_prompt": "",
|
||||||
|
"prompt": "生成一份中文采集器凭据配置教程。只能根据 context.search_evidence 中的来源生成教程;如果证据不足,明确说明需要以官方页面为准。"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"key": "ai.connection_test",
|
||||||
|
"label": "AI Provider 连接测试",
|
||||||
|
"group": "运维测试",
|
||||||
|
"version": "2026-05-16.1",
|
||||||
|
"system_prompt": "",
|
||||||
|
"prompt": "Reply OK."
|
||||||
|
}
|
||||||
|
]
|
||||||
182
backend/app/ai_tasks/prompts.py
Normal file
182
backend/app/ai_tasks/prompts.py
Normal file
@@ -0,0 +1,182 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
import json
|
||||||
|
from functools import lru_cache
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from sqlalchemy import select
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.models.system_setting import SystemSetting
|
||||||
|
|
||||||
|
AI_PROMPTS_CATEGORY = "ai_prompts"
|
||||||
|
DEFAULT_PROMPTS_PATH = Path(__file__).with_name("default_prompts.json")
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class AIPromptDefinition:
|
||||||
|
key: str
|
||||||
|
label: str
|
||||||
|
group: str
|
||||||
|
version: str
|
||||||
|
system_prompt: str
|
||||||
|
prompt: str
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class EffectiveAIPrompt:
|
||||||
|
key: str
|
||||||
|
label: str
|
||||||
|
group: str
|
||||||
|
version: str
|
||||||
|
default_system_prompt: str
|
||||||
|
default_prompt: str
|
||||||
|
system_prompt: str
|
||||||
|
prompt: str
|
||||||
|
is_custom: bool
|
||||||
|
updated_at: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
@lru_cache(maxsize=1)
|
||||||
|
def list_prompt_definitions() -> tuple[AIPromptDefinition, ...]:
|
||||||
|
raw_items = json.loads(DEFAULT_PROMPTS_PATH.read_text(encoding="utf-8"))
|
||||||
|
return tuple(
|
||||||
|
AIPromptDefinition(
|
||||||
|
key=str(item["key"]),
|
||||||
|
label=str(item["label"]),
|
||||||
|
group=str(item["group"]),
|
||||||
|
version=str(item["version"]),
|
||||||
|
system_prompt=str(item.get("system_prompt") or ""),
|
||||||
|
prompt=str(item.get("prompt") or ""),
|
||||||
|
)
|
||||||
|
for item in raw_items
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def get_prompt_definition(task_key: str) -> AIPromptDefinition:
|
||||||
|
for definition in list_prompt_definitions():
|
||||||
|
if definition.key == task_key:
|
||||||
|
return definition
|
||||||
|
raise KeyError(task_key)
|
||||||
|
|
||||||
|
|
||||||
|
async def _get_prompt_setting(db: AsyncSession) -> SystemSetting | None:
|
||||||
|
result = await db.execute(
|
||||||
|
select(SystemSetting).where(SystemSetting.category == AI_PROMPTS_CATEGORY)
|
||||||
|
)
|
||||||
|
return result.scalar_one_or_none()
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_overrides(payload: dict[str, Any] | None) -> dict[str, dict[str, Any]]:
|
||||||
|
raw = (payload or {}).get("overrides")
|
||||||
|
if not isinstance(raw, dict):
|
||||||
|
return {}
|
||||||
|
return {
|
||||||
|
str(key): dict(value)
|
||||||
|
for key, value in raw.items()
|
||||||
|
if isinstance(value, dict)
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def get_prompt_overrides(db: AsyncSession) -> dict[str, dict[str, Any]]:
|
||||||
|
if not hasattr(db, "execute"):
|
||||||
|
return {}
|
||||||
|
setting = await _get_prompt_setting(db)
|
||||||
|
return _normalize_overrides(setting.payload if setting else None)
|
||||||
|
|
||||||
|
|
||||||
|
def _effective_prompt(
|
||||||
|
definition: AIPromptDefinition,
|
||||||
|
override: dict[str, Any] | None,
|
||||||
|
) -> EffectiveAIPrompt:
|
||||||
|
override = override or {}
|
||||||
|
custom_system = override.get("system_prompt")
|
||||||
|
custom_prompt = override.get("prompt")
|
||||||
|
has_custom_system = isinstance(custom_system, str)
|
||||||
|
has_custom_prompt = isinstance(custom_prompt, str)
|
||||||
|
return EffectiveAIPrompt(
|
||||||
|
key=definition.key,
|
||||||
|
label=definition.label,
|
||||||
|
group=definition.group,
|
||||||
|
version=definition.version,
|
||||||
|
default_system_prompt=definition.system_prompt,
|
||||||
|
default_prompt=definition.prompt,
|
||||||
|
system_prompt=custom_system if has_custom_system else definition.system_prompt,
|
||||||
|
prompt=custom_prompt if has_custom_prompt else definition.prompt,
|
||||||
|
is_custom=has_custom_system or has_custom_prompt,
|
||||||
|
updated_at=str(override.get("updated_at") or "") or None,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def list_effective_prompts(db: AsyncSession) -> list[EffectiveAIPrompt]:
|
||||||
|
overrides = await get_prompt_overrides(db)
|
||||||
|
return [
|
||||||
|
_effective_prompt(definition, overrides.get(definition.key))
|
||||||
|
for definition in list_prompt_definitions()
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
async def get_effective_prompt(db: AsyncSession | None, task_key: str) -> EffectiveAIPrompt:
|
||||||
|
definition = get_prompt_definition(task_key)
|
||||||
|
if db is None:
|
||||||
|
return _effective_prompt(definition, None)
|
||||||
|
overrides = await get_prompt_overrides(db)
|
||||||
|
return _effective_prompt(definition, overrides.get(task_key))
|
||||||
|
|
||||||
|
|
||||||
|
async def save_prompt_override(
|
||||||
|
db: AsyncSession,
|
||||||
|
task_key: str,
|
||||||
|
*,
|
||||||
|
system_prompt: str,
|
||||||
|
prompt: str,
|
||||||
|
) -> EffectiveAIPrompt:
|
||||||
|
definition = get_prompt_definition(task_key)
|
||||||
|
setting = await _get_prompt_setting(db)
|
||||||
|
payload = dict(setting.payload or {}) if setting else {}
|
||||||
|
overrides = _normalize_overrides(payload)
|
||||||
|
overrides[definition.key] = {
|
||||||
|
"system_prompt": system_prompt,
|
||||||
|
"prompt": prompt,
|
||||||
|
"updated_at": datetime.now(UTC).isoformat().replace("+00:00", "Z"),
|
||||||
|
}
|
||||||
|
payload["overrides"] = overrides
|
||||||
|
if setting is None:
|
||||||
|
setting = SystemSetting(category=AI_PROMPTS_CATEGORY, payload=payload)
|
||||||
|
db.add(setting)
|
||||||
|
else:
|
||||||
|
setting.payload = payload
|
||||||
|
await db.commit()
|
||||||
|
return _effective_prompt(definition, overrides[definition.key])
|
||||||
|
|
||||||
|
|
||||||
|
async def reset_prompt_override(db: AsyncSession, task_key: str) -> EffectiveAIPrompt:
|
||||||
|
definition = get_prompt_definition(task_key)
|
||||||
|
setting = await _get_prompt_setting(db)
|
||||||
|
if setting is None:
|
||||||
|
return _effective_prompt(definition, None)
|
||||||
|
payload = dict(setting.payload or {})
|
||||||
|
overrides = _normalize_overrides(payload)
|
||||||
|
overrides.pop(definition.key, None)
|
||||||
|
payload["overrides"] = overrides
|
||||||
|
setting.payload = payload
|
||||||
|
await db.commit()
|
||||||
|
return _effective_prompt(definition, None)
|
||||||
|
|
||||||
|
|
||||||
|
def serialize_effective_prompt(prompt: EffectiveAIPrompt) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"key": prompt.key,
|
||||||
|
"label": prompt.label,
|
||||||
|
"group": prompt.group,
|
||||||
|
"version": prompt.version,
|
||||||
|
"default_system_prompt": prompt.default_system_prompt,
|
||||||
|
"default_prompt": prompt.default_prompt,
|
||||||
|
"system_prompt": prompt.system_prompt,
|
||||||
|
"prompt": prompt.prompt,
|
||||||
|
"is_custom": prompt.is_custom,
|
||||||
|
"updated_at": prompt.updated_at,
|
||||||
|
}
|
||||||
@@ -5,15 +5,22 @@ from app.api.v1 import (
|
|||||||
users,
|
users,
|
||||||
datasource_config,
|
datasource_config,
|
||||||
datasources,
|
datasources,
|
||||||
|
docs,
|
||||||
|
earth,
|
||||||
tasks,
|
tasks,
|
||||||
dashboard,
|
dashboard,
|
||||||
websocket,
|
|
||||||
alerts,
|
alerts,
|
||||||
settings,
|
settings,
|
||||||
collected_data,
|
collected_data,
|
||||||
|
data_products,
|
||||||
|
layers,
|
||||||
visualization,
|
visualization,
|
||||||
|
vessel_aggregation,
|
||||||
|
vessels,
|
||||||
bgp,
|
bgp,
|
||||||
news,
|
news,
|
||||||
|
interactables,
|
||||||
|
realtime_sources,
|
||||||
system_control,
|
system_control,
|
||||||
tv,
|
tv,
|
||||||
)
|
)
|
||||||
@@ -28,12 +35,24 @@ api_router.include_router(
|
|||||||
)
|
)
|
||||||
api_router.include_router(datasources.router, prefix="/datasources", tags=["datasources"])
|
api_router.include_router(datasources.router, prefix="/datasources", tags=["datasources"])
|
||||||
api_router.include_router(collected_data.router, prefix="/collected", tags=["collected-data"])
|
api_router.include_router(collected_data.router, prefix="/collected", tags=["collected-data"])
|
||||||
|
api_router.include_router(docs.router, prefix="/docs", tags=["docs"])
|
||||||
|
api_router.include_router(earth.router, prefix="/earth", tags=["earth"])
|
||||||
api_router.include_router(tasks.router, prefix="/tasks", tags=["tasks"])
|
api_router.include_router(tasks.router, prefix="/tasks", tags=["tasks"])
|
||||||
api_router.include_router(dashboard.router, prefix="/dashboard", tags=["dashboard"])
|
api_router.include_router(dashboard.router, prefix="/dashboard", tags=["dashboard"])
|
||||||
api_router.include_router(alerts.router, prefix="/alerts", tags=["alerts"])
|
api_router.include_router(alerts.router, prefix="/alerts", tags=["alerts"])
|
||||||
api_router.include_router(settings.router, prefix="/settings", tags=["settings"])
|
api_router.include_router(settings.router, prefix="/settings", tags=["settings"])
|
||||||
api_router.include_router(system_control.router, prefix="/system", tags=["system"])
|
api_router.include_router(system_control.router, prefix="/system", tags=["system"])
|
||||||
|
api_router.include_router(data_products.router, prefix="/data-products", tags=["data-products"])
|
||||||
|
api_router.include_router(layers.router, prefix="/layers", tags=["layers"])
|
||||||
api_router.include_router(visualization.router, prefix="/visualization", tags=["visualization"])
|
api_router.include_router(visualization.router, prefix="/visualization", tags=["visualization"])
|
||||||
|
api_router.include_router(
|
||||||
|
vessel_aggregation.router,
|
||||||
|
prefix="/vessel-aggregation",
|
||||||
|
tags=["vessel-aggregation"],
|
||||||
|
)
|
||||||
|
api_router.include_router(vessels.router, prefix="/vessels", tags=["vessels"])
|
||||||
api_router.include_router(bgp.router, prefix="/bgp", tags=["bgp"])
|
api_router.include_router(bgp.router, prefix="/bgp", tags=["bgp"])
|
||||||
api_router.include_router(tv.router, prefix="/tv", tags=["tv"])
|
api_router.include_router(tv.router, prefix="/tv", tags=["tv"])
|
||||||
api_router.include_router(news.router, prefix="/news", tags=["news"])
|
api_router.include_router(news.router, prefix="/news", tags=["news"])
|
||||||
|
api_router.include_router(interactables.router, prefix="/interactables", tags=["interactables"])
|
||||||
|
api_router.include_router(realtime_sources.router, prefix="/realtime-sources", tags=["realtime-sources"])
|
||||||
|
|||||||
@@ -3,6 +3,7 @@ from uuid import uuid4
|
|||||||
from fastapi import APIRouter, Depends, HTTPException, Request, Response
|
from fastapi import APIRouter, Depends, HTTPException, Request, Response
|
||||||
from sqlalchemy.ext.asyncio import AsyncSession
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.logging import get_logger
|
||||||
from app.core.security import get_current_user
|
from app.core.security import get_current_user
|
||||||
from app.db.session import get_db
|
from app.db.session import get_db
|
||||||
from app.models.user import User
|
from app.models.user import User
|
||||||
@@ -47,8 +48,10 @@ from app.services.playground_chat_service import (
|
|||||||
stop_message,
|
stop_message,
|
||||||
)
|
)
|
||||||
from app.services.situational_alert_ai_brief import build_situational_alert_brief_request
|
from app.services.situational_alert_ai_brief import build_situational_alert_brief_request
|
||||||
|
from app.services.business_logs import emit_business_log, exception_context
|
||||||
|
|
||||||
router = APIRouter()
|
router = APIRouter()
|
||||||
|
logger = get_logger(__name__, service="api")
|
||||||
|
|
||||||
|
|
||||||
@router.get("/provider/status", response_model=AIProviderStatusResponse)
|
@router.get("/provider/status", response_model=AIProviderStatusResponse)
|
||||||
@@ -122,6 +125,16 @@ async def create_playground_message(
|
|||||||
provider_client: AIProviderClient = Depends(get_ai_provider_client),
|
provider_client: AIProviderClient = Depends(get_ai_provider_client),
|
||||||
db: AsyncSession = Depends(get_db),
|
db: AsyncSession = Depends(get_db),
|
||||||
):
|
):
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.playground.message.create",
|
||||||
|
message="Playground message creation requested",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"session_key": payload.session_key, "preset": payload.selected_preset_key},
|
||||||
|
)
|
||||||
return await create_turn(
|
return await create_turn(
|
||||||
db,
|
db,
|
||||||
user_id=current_user.id,
|
user_id=current_user.id,
|
||||||
@@ -136,6 +149,16 @@ async def stop_playground_message(
|
|||||||
current_user: User = Depends(get_current_user),
|
current_user: User = Depends(get_current_user),
|
||||||
db: AsyncSession = Depends(get_db),
|
db: AsyncSession = Depends(get_db),
|
||||||
):
|
):
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.playground.message.stop",
|
||||||
|
message="Playground message stop requested",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"session_key": payload.session_key, "message_id": payload.message_id},
|
||||||
|
)
|
||||||
return await stop_message(
|
return await stop_message(
|
||||||
db,
|
db,
|
||||||
user_id=current_user.id,
|
user_id=current_user.id,
|
||||||
@@ -150,6 +173,16 @@ async def resend_playground_message(
|
|||||||
provider_client: AIProviderClient = Depends(get_ai_provider_client),
|
provider_client: AIProviderClient = Depends(get_ai_provider_client),
|
||||||
db: AsyncSession = Depends(get_db),
|
db: AsyncSession = Depends(get_db),
|
||||||
):
|
):
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.playground.message.resend",
|
||||||
|
message="Playground message resend requested",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"session_key": payload.session_key, "user_message_id": payload.user_message_id},
|
||||||
|
)
|
||||||
return await resend_turn(
|
return await resend_turn(
|
||||||
db,
|
db,
|
||||||
user_id=current_user.id,
|
user_id=current_user.id,
|
||||||
@@ -214,16 +247,70 @@ async def analyze_bgp_brief(
|
|||||||
anomaly_limit=payload.anomaly_limit,
|
anomaly_limit=payload.anomaly_limit,
|
||||||
collector_limit=payload.collector_limit,
|
collector_limit=payload.collector_limit,
|
||||||
)
|
)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.brief.bgp.facts_collected",
|
||||||
|
message="BGP brief facts collected",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={
|
||||||
|
"incident_limit": payload.incident_limit,
|
||||||
|
"anomaly_limit": payload.anomaly_limit,
|
||||||
|
"collector_limit": payload.collector_limit,
|
||||||
|
"fact_count": len(facts or []),
|
||||||
|
},
|
||||||
|
)
|
||||||
brief_request.preferred_model = payload.preferred_model
|
brief_request.preferred_model = payload.preferred_model
|
||||||
brief_request.thinking = payload.thinking
|
brief_request.thinking = payload.thinking
|
||||||
|
|
||||||
analysis = await provider_client.analyze(brief_request, request_id=request_id)
|
await emit_business_log(
|
||||||
return save_bgp_brief_record(
|
logger,
|
||||||
analysis,
|
event="ai.brief.bgp.start",
|
||||||
|
message="BGP brief AI analysis started",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
request_id=request_id,
|
request_id=request_id,
|
||||||
facts=facts,
|
user_id=current_user.id,
|
||||||
context=context,
|
context={"preferred_model": payload.preferred_model},
|
||||||
)
|
)
|
||||||
|
try:
|
||||||
|
analysis = await provider_client.analyze(brief_request, request_id=request_id)
|
||||||
|
record = save_bgp_brief_record(
|
||||||
|
analysis,
|
||||||
|
request_id=request_id,
|
||||||
|
facts=facts,
|
||||||
|
context=context,
|
||||||
|
)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.brief.bgp.completed",
|
||||||
|
message="BGP brief AI analysis saved",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"provider": analysis.provider, "model": analysis.model, "brief_id": record.id},
|
||||||
|
)
|
||||||
|
return record
|
||||||
|
except Exception as exc:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.brief.bgp.failed",
|
||||||
|
message="BGP brief AI analysis failed",
|
||||||
|
category="ai",
|
||||||
|
level="error",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context=exception_context(exc, {"preferred_model": payload.preferred_model}),
|
||||||
|
)
|
||||||
|
raise
|
||||||
|
|
||||||
|
|
||||||
@router.post("/alerts/brief", response_model=AlertBriefResponse)
|
@router.post("/alerts/brief", response_model=AlertBriefResponse)
|
||||||
@@ -242,17 +329,65 @@ async def analyze_alert_brief(
|
|||||||
db,
|
db,
|
||||||
alert_limit=payload.alert_limit,
|
alert_limit=payload.alert_limit,
|
||||||
)
|
)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.brief.alerts.facts_collected",
|
||||||
|
message="Alert brief facts collected",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"alert_limit": payload.alert_limit, "fact_count": len(facts or [])},
|
||||||
|
)
|
||||||
brief_request.preferred_model = payload.preferred_model
|
brief_request.preferred_model = payload.preferred_model
|
||||||
brief_request.thinking = payload.thinking
|
brief_request.thinking = payload.thinking
|
||||||
|
|
||||||
analysis = await provider_client.analyze(brief_request, request_id=request_id)
|
await emit_business_log(
|
||||||
return AlertBriefResponse(
|
logger,
|
||||||
**analysis.model_dump(),
|
event="ai.brief.alerts.start",
|
||||||
title=brief_request.title,
|
message="Alert brief AI analysis started",
|
||||||
objective=brief_request.objective,
|
category="ai",
|
||||||
facts=facts,
|
service="api",
|
||||||
context=context,
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"preferred_model": payload.preferred_model},
|
||||||
)
|
)
|
||||||
|
try:
|
||||||
|
analysis = await provider_client.analyze(brief_request, request_id=request_id)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.brief.alerts.completed",
|
||||||
|
message="Alert brief AI analysis completed",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"provider": analysis.provider, "model": analysis.model},
|
||||||
|
)
|
||||||
|
return AlertBriefResponse(
|
||||||
|
**analysis.model_dump(),
|
||||||
|
title=brief_request.title,
|
||||||
|
objective=brief_request.objective,
|
||||||
|
facts=facts,
|
||||||
|
context=context,
|
||||||
|
)
|
||||||
|
except Exception as exc:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.brief.alerts.failed",
|
||||||
|
message="Alert brief AI analysis failed",
|
||||||
|
category="ai",
|
||||||
|
level="error",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context=exception_context(exc, {"preferred_model": payload.preferred_model}),
|
||||||
|
)
|
||||||
|
raise
|
||||||
|
|
||||||
|
|
||||||
@router.post("/situational-alerts/brief", response_model=SituationalAlertBriefResponse)
|
@router.post("/situational-alerts/brief", response_model=SituationalAlertBriefResponse)
|
||||||
@@ -268,14 +403,62 @@ async def analyze_situational_alert_brief(
|
|||||||
response.headers["X-Request-ID"] = request_id
|
response.headers["X-Request-ID"] = request_id
|
||||||
|
|
||||||
brief_request, facts, context = await build_situational_alert_brief_request(db)
|
brief_request, facts, context = await build_situational_alert_brief_request(db)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.brief.situational_alerts.facts_collected",
|
||||||
|
message="Situational alert brief facts collected",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"fact_count": len(facts or [])},
|
||||||
|
)
|
||||||
brief_request.preferred_model = payload.preferred_model
|
brief_request.preferred_model = payload.preferred_model
|
||||||
brief_request.thinking = payload.thinking
|
brief_request.thinking = payload.thinking
|
||||||
|
|
||||||
analysis = await provider_client.analyze(brief_request, request_id=request_id)
|
await emit_business_log(
|
||||||
return SituationalAlertBriefResponse(
|
logger,
|
||||||
**analysis.model_dump(),
|
event="ai.brief.situational_alerts.start",
|
||||||
title=brief_request.title,
|
message="Situational alert brief AI analysis started",
|
||||||
objective=brief_request.objective,
|
category="ai",
|
||||||
facts=facts,
|
service="api",
|
||||||
context=context,
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"preferred_model": payload.preferred_model},
|
||||||
)
|
)
|
||||||
|
try:
|
||||||
|
analysis = await provider_client.analyze(brief_request, request_id=request_id)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.brief.situational_alerts.completed",
|
||||||
|
message="Situational alert brief AI analysis completed",
|
||||||
|
category="ai",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context={"provider": analysis.provider, "model": analysis.model},
|
||||||
|
)
|
||||||
|
return SituationalAlertBriefResponse(
|
||||||
|
**analysis.model_dump(),
|
||||||
|
title=brief_request.title,
|
||||||
|
objective=brief_request.objective,
|
||||||
|
facts=facts,
|
||||||
|
context=context,
|
||||||
|
)
|
||||||
|
except Exception as exc:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.brief.situational_alerts.failed",
|
||||||
|
message="Situational alert brief AI analysis failed",
|
||||||
|
category="ai",
|
||||||
|
level="error",
|
||||||
|
service="api",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
user_id=current_user.id,
|
||||||
|
context=exception_context(exc, {"preferred_model": payload.preferred_model}),
|
||||||
|
)
|
||||||
|
raise
|
||||||
|
|||||||
@@ -1,26 +1,85 @@
|
|||||||
from datetime import timedelta
|
|
||||||
|
|
||||||
from fastapi import APIRouter, Depends, HTTPException, status
|
from fastapi import APIRouter, Depends, HTTPException, status
|
||||||
from fastapi.security import OAuth2PasswordRequestForm
|
from fastapi.security import OAuth2PasswordRequestForm
|
||||||
from sqlalchemy.ext.asyncio import AsyncSession
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
from sqlalchemy import text
|
from sqlalchemy import text
|
||||||
|
|
||||||
from app.core.config import settings
|
from app.core.config import settings
|
||||||
|
from app.core.logging import get_logger
|
||||||
from app.core.security import (
|
from app.core.security import (
|
||||||
create_access_token,
|
create_access_token,
|
||||||
create_refresh_token,
|
create_refresh_token,
|
||||||
blacklist_token,
|
|
||||||
get_current_user,
|
get_current_user,
|
||||||
|
get_password_hash,
|
||||||
verify_password,
|
verify_password,
|
||||||
)
|
)
|
||||||
from app.db.session import get_db
|
from app.db.session import get_db
|
||||||
from app.models.user import User
|
from app.models.user import User
|
||||||
from app.schemas.token import Token
|
from app.schemas.token import Token
|
||||||
from app.schemas.user import UserCreate, UserResponse
|
from app.schemas.user import (
|
||||||
|
ForgotPasswordRequest,
|
||||||
|
ResendCodeRequest,
|
||||||
|
ResetPasswordRequest,
|
||||||
|
UserRegister,
|
||||||
|
UserResponse,
|
||||||
|
VerifyEmailRequest,
|
||||||
|
)
|
||||||
|
from app.services import otp
|
||||||
|
from app.services.email import (
|
||||||
|
EmailError,
|
||||||
|
EmailNotConfiguredError,
|
||||||
|
send_verification_email,
|
||||||
|
)
|
||||||
|
|
||||||
|
logger = get_logger(__name__)
|
||||||
|
|
||||||
router = APIRouter()
|
router = APIRouter()
|
||||||
|
|
||||||
|
|
||||||
|
def _token_response(user: User) -> dict:
|
||||||
|
access_token = create_access_token(data={"sub": user.id})
|
||||||
|
refresh = create_refresh_token(data={"sub": user.id})
|
||||||
|
expires_in = (
|
||||||
|
settings.ACCESS_TOKEN_EXPIRE_MINUTES * 60
|
||||||
|
if settings.ACCESS_TOKEN_EXPIRE_MINUTES > 0
|
||||||
|
else None
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"access_token": access_token,
|
||||||
|
"token_type": "bearer",
|
||||||
|
"expires_in": expires_in,
|
||||||
|
"refresh_token": refresh,
|
||||||
|
"user": {
|
||||||
|
"id": user.id,
|
||||||
|
"username": user.username,
|
||||||
|
"role": user.role,
|
||||||
|
"gatekeeper_groups": user.gatekeeper_groups or [],
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def _load_user_by_email(db: AsyncSession, email: str) -> User | None:
|
||||||
|
result = await db.execute(
|
||||||
|
text(
|
||||||
|
"SELECT id, username, email, password_hash, role, is_active, gatekeeper_groups, email_verified "
|
||||||
|
"FROM users WHERE email = :email"
|
||||||
|
),
|
||||||
|
{"email": email},
|
||||||
|
)
|
||||||
|
row = result.fetchone()
|
||||||
|
if row is None:
|
||||||
|
return None
|
||||||
|
user = User()
|
||||||
|
user.id = row[0]
|
||||||
|
user.username = row[1]
|
||||||
|
user.email = row[2]
|
||||||
|
user.password_hash = row[3]
|
||||||
|
user.role = row[4]
|
||||||
|
user.is_active = row[5]
|
||||||
|
user.gatekeeper_groups = row[6] or []
|
||||||
|
user.email_verified = bool(row[7])
|
||||||
|
return user
|
||||||
|
|
||||||
|
|
||||||
@router.post("/login", response_model=Token)
|
@router.post("/login", response_model=Token)
|
||||||
async def login(
|
async def login(
|
||||||
form_data: OAuth2PasswordRequestForm = Depends(),
|
form_data: OAuth2PasswordRequestForm = Depends(),
|
||||||
@@ -28,7 +87,8 @@ async def login(
|
|||||||
):
|
):
|
||||||
result = await db.execute(
|
result = await db.execute(
|
||||||
text(
|
text(
|
||||||
"SELECT id, username, email, password_hash, role, is_active FROM users WHERE username = :username"
|
"SELECT id, username, email, password_hash, role, is_active, gatekeeper_groups, email_verified "
|
||||||
|
"FROM users WHERE username = :username"
|
||||||
),
|
),
|
||||||
{"username": form_data.username},
|
{"username": form_data.username},
|
||||||
)
|
)
|
||||||
@@ -46,6 +106,8 @@ async def login(
|
|||||||
user.password_hash = row[3]
|
user.password_hash = row[3]
|
||||||
user.role = row[4]
|
user.role = row[4]
|
||||||
user.is_active = row[5]
|
user.is_active = row[5]
|
||||||
|
user.gatekeeper_groups = row[6] or []
|
||||||
|
user.email_verified = bool(row[7])
|
||||||
|
|
||||||
if not verify_password(form_data.password, user.password_hash):
|
if not verify_password(form_data.password, user.password_hash):
|
||||||
raise HTTPException(
|
raise HTTPException(
|
||||||
@@ -57,24 +119,13 @@ async def login(
|
|||||||
status_code=status.HTTP_401_UNAUTHORIZED,
|
status_code=status.HTTP_401_UNAUTHORIZED,
|
||||||
detail="User is inactive",
|
detail="User is inactive",
|
||||||
)
|
)
|
||||||
|
if not user.email_verified:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_403_FORBIDDEN,
|
||||||
|
detail={"code": "EMAIL_NOT_VERIFIED", "email": user.email},
|
||||||
|
)
|
||||||
|
|
||||||
access_token = create_access_token(data={"sub": user.id})
|
return _token_response(user)
|
||||||
refresh_token = create_refresh_token(data={"sub": user.id})
|
|
||||||
|
|
||||||
expires_in = None
|
|
||||||
if settings.ACCESS_TOKEN_EXPIRE_MINUTES > 0:
|
|
||||||
expires_in = settings.ACCESS_TOKEN_EXPIRE_MINUTES * 60
|
|
||||||
|
|
||||||
return {
|
|
||||||
"access_token": access_token,
|
|
||||||
"token_type": "bearer",
|
|
||||||
"expires_in": expires_in,
|
|
||||||
"user": {
|
|
||||||
"id": user.id,
|
|
||||||
"username": user.username,
|
|
||||||
"role": user.role,
|
|
||||||
},
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
@router.post("/refresh", response_model=Token)
|
@router.post("/refresh", response_model=Token)
|
||||||
@@ -95,6 +146,7 @@ async def refresh_token(
|
|||||||
"id": current_user.id,
|
"id": current_user.id,
|
||||||
"username": current_user.username,
|
"username": current_user.username,
|
||||||
"role": current_user.role,
|
"role": current_user.role,
|
||||||
|
"gatekeeper_groups": current_user.gatekeeper_groups or [],
|
||||||
},
|
},
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -111,6 +163,181 @@ async def get_me(current_user: User = Depends(get_current_user)):
|
|||||||
"username": current_user.username,
|
"username": current_user.username,
|
||||||
"email": current_user.email,
|
"email": current_user.email,
|
||||||
"role": current_user.role,
|
"role": current_user.role,
|
||||||
|
"gatekeeper_groups": current_user.gatekeeper_groups or [],
|
||||||
"is_active": current_user.is_active,
|
"is_active": current_user.is_active,
|
||||||
|
"email_verified": getattr(current_user, "email_verified", True),
|
||||||
"created_at": current_user.created_at,
|
"created_at": current_user.created_at,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def _send_code_or_raise(db: AsyncSession, email: str, code: str, purpose: str) -> None:
|
||||||
|
try:
|
||||||
|
await send_verification_email(db, to=email, code=code, purpose=purpose)
|
||||||
|
except EmailNotConfiguredError as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
|
||||||
|
detail={"code": exc.code, "message": str(exc)},
|
||||||
|
) from exc
|
||||||
|
except EmailError as exc:
|
||||||
|
logger.warning_event(
|
||||||
|
"SMTP send failed",
|
||||||
|
event="auth.email.send_failed",
|
||||||
|
context={"email": email, "purpose": purpose, "error": str(exc)},
|
||||||
|
)
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_502_BAD_GATEWAY,
|
||||||
|
detail={"code": exc.code, "message": str(exc)},
|
||||||
|
) from exc
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/register", status_code=status.HTTP_201_CREATED)
|
||||||
|
async def register(payload: UserRegister, db: AsyncSession = Depends(get_db)):
|
||||||
|
existing = await db.execute(
|
||||||
|
text("SELECT id, email_verified FROM users WHERE username = :u OR email = :e"),
|
||||||
|
{"u": payload.username, "e": payload.email},
|
||||||
|
)
|
||||||
|
row = existing.fetchone()
|
||||||
|
if row is not None:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_409_CONFLICT,
|
||||||
|
detail={"code": "USER_ALREADY_EXISTS", "message": "Username or email already in use"},
|
||||||
|
)
|
||||||
|
|
||||||
|
user = User(
|
||||||
|
username=payload.username,
|
||||||
|
email=payload.email,
|
||||||
|
password_hash=get_password_hash(payload.password),
|
||||||
|
role="viewer",
|
||||||
|
is_active=True,
|
||||||
|
email_verified=False,
|
||||||
|
)
|
||||||
|
db.add(user)
|
||||||
|
await db.commit()
|
||||||
|
|
||||||
|
try:
|
||||||
|
code = otp.issue_code(payload.email, "register")
|
||||||
|
except otp.OtpResendRateLimited as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_429_TOO_MANY_REQUESTS,
|
||||||
|
detail={"code": exc.code, "retry_after_seconds": exc.retry_after_seconds},
|
||||||
|
) from exc
|
||||||
|
await _send_code_or_raise(db, payload.email, code, "register")
|
||||||
|
return {"status": "pending_verification", "email": payload.email}
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/verify-email", response_model=Token)
|
||||||
|
async def verify_email(payload: VerifyEmailRequest, db: AsyncSession = Depends(get_db)):
|
||||||
|
user = await _load_user_by_email(db, payload.email)
|
||||||
|
if user is None:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_404_NOT_FOUND,
|
||||||
|
detail={"code": "USER_NOT_FOUND"},
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
otp.verify_code(payload.email, "register", payload.code)
|
||||||
|
except otp.OtpExpired as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_410_GONE,
|
||||||
|
detail={"code": exc.code, "message": str(exc)},
|
||||||
|
) from exc
|
||||||
|
except otp.OtpAttemptsExceeded as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_429_TOO_MANY_REQUESTS,
|
||||||
|
detail={"code": exc.code, "message": str(exc)},
|
||||||
|
) from exc
|
||||||
|
except otp.OtpInvalid as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_400_BAD_REQUEST,
|
||||||
|
detail={"code": exc.code, "message": str(exc)},
|
||||||
|
) from exc
|
||||||
|
|
||||||
|
await db.execute(
|
||||||
|
text("UPDATE users SET email_verified = TRUE WHERE id = :id"),
|
||||||
|
{"id": user.id},
|
||||||
|
)
|
||||||
|
await db.commit()
|
||||||
|
user.email_verified = True
|
||||||
|
return _token_response(user)
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/resend-code")
|
||||||
|
async def resend_code(payload: ResendCodeRequest, db: AsyncSession = Depends(get_db)):
|
||||||
|
user = await _load_user_by_email(db, payload.email)
|
||||||
|
if user is None:
|
||||||
|
# Avoid email enumeration; pretend success.
|
||||||
|
return {"status": "ok"}
|
||||||
|
if payload.purpose == "register" and user.email_verified:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_409_CONFLICT,
|
||||||
|
detail={"code": "ALREADY_VERIFIED"},
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
code = otp.issue_code(payload.email, payload.purpose)
|
||||||
|
except otp.OtpResendRateLimited as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_429_TOO_MANY_REQUESTS,
|
||||||
|
detail={"code": exc.code, "retry_after_seconds": exc.retry_after_seconds},
|
||||||
|
) from exc
|
||||||
|
await _send_code_or_raise(db, payload.email, code, payload.purpose)
|
||||||
|
return {"status": "ok"}
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/forgot-password")
|
||||||
|
async def forgot_password(payload: ForgotPasswordRequest, db: AsyncSession = Depends(get_db)):
|
||||||
|
user = await _load_user_by_email(db, payload.email)
|
||||||
|
if user is None:
|
||||||
|
# Don't leak whether an email is registered.
|
||||||
|
return {"status": "ok"}
|
||||||
|
try:
|
||||||
|
code = otp.issue_code(payload.email, "reset_password")
|
||||||
|
except otp.OtpResendRateLimited:
|
||||||
|
# Silently accept; the user can retry after the cooldown.
|
||||||
|
return {"status": "ok"}
|
||||||
|
try:
|
||||||
|
await send_verification_email(db, to=payload.email, code=code, purpose="reset_password")
|
||||||
|
except EmailNotConfiguredError as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
|
||||||
|
detail={"code": exc.code, "message": str(exc)},
|
||||||
|
) from exc
|
||||||
|
except EmailError as exc:
|
||||||
|
logger.warning_event(
|
||||||
|
"SMTP send failed",
|
||||||
|
event="auth.email.send_failed",
|
||||||
|
context={"email": payload.email, "purpose": "reset_password", "error": str(exc)},
|
||||||
|
)
|
||||||
|
return {"status": "ok"}
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/reset-password")
|
||||||
|
async def reset_password(payload: ResetPasswordRequest, db: AsyncSession = Depends(get_db)):
|
||||||
|
user = await _load_user_by_email(db, payload.email)
|
||||||
|
if user is None:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_400_BAD_REQUEST,
|
||||||
|
detail={"code": "OTP_INVALID"},
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
otp.verify_code(payload.email, "reset_password", payload.code)
|
||||||
|
except otp.OtpExpired as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_410_GONE,
|
||||||
|
detail={"code": exc.code, "message": str(exc)},
|
||||||
|
) from exc
|
||||||
|
except otp.OtpAttemptsExceeded as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_429_TOO_MANY_REQUESTS,
|
||||||
|
detail={"code": exc.code, "message": str(exc)},
|
||||||
|
) from exc
|
||||||
|
except otp.OtpInvalid as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_400_BAD_REQUEST,
|
||||||
|
detail={"code": exc.code, "message": str(exc)},
|
||||||
|
) from exc
|
||||||
|
|
||||||
|
await db.execute(
|
||||||
|
text("UPDATE users SET password_hash = :p, email_verified = TRUE WHERE id = :id"),
|
||||||
|
{"p": get_password_hash(payload.new_password), "id": user.id},
|
||||||
|
)
|
||||||
|
await db.commit()
|
||||||
|
return {"status": "ok"}
|
||||||
|
|||||||
@@ -5,13 +5,26 @@ from fastapi import APIRouter, Depends, HTTPException, Query
|
|||||||
from sqlalchemy import func, select
|
from sqlalchemy import func, select
|
||||||
from sqlalchemy.ext.asyncio import AsyncSession
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from pydantic import BaseModel
|
||||||
|
|
||||||
from app.core.security import get_current_user
|
from app.core.security import get_current_user
|
||||||
from app.db.session import get_db
|
from app.db.session import get_db
|
||||||
from app.models.bgp_anomaly import BGPAnomaly
|
from app.models.bgp_anomaly import BGPAnomaly
|
||||||
from app.models.bgp_incident import BGPIncident
|
from app.models.bgp_incident import BGPIncident
|
||||||
from app.models.bgp_observation import BGPObservation
|
from app.models.bgp_observation import BGPObservation
|
||||||
from app.models.user import User
|
from app.models.user import User
|
||||||
|
from app.services.bgp_collector_locations import (
|
||||||
|
build_bgp_collector_location_query,
|
||||||
|
collect_bgp_collector_location_candidates,
|
||||||
|
get_bgp_collector_location_dict,
|
||||||
|
)
|
||||||
from app.services.bgp_collectors import build_bgp_collector_coverage
|
from app.services.bgp_collectors import build_bgp_collector_coverage
|
||||||
|
from app.services.ai_client import get_ai_provider_client
|
||||||
|
from app.api.v1.settings import get_web_search_client
|
||||||
|
from app.services.location.llm_fallback import (
|
||||||
|
collect_llm_location_fallback_candidate,
|
||||||
|
collect_location_search_evidence,
|
||||||
|
)
|
||||||
|
|
||||||
router = APIRouter()
|
router = APIRouter()
|
||||||
|
|
||||||
@@ -264,6 +277,120 @@ async def get_bgp_collector_summary(
|
|||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class CollectBGPCollectorLocationRequest(BaseModel):
|
||||||
|
city: Optional[str] = None
|
||||||
|
country: Optional[str] = None
|
||||||
|
site: Optional[str] = None
|
||||||
|
operator: Optional[str] = None
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/collectors/{collector_id}/collect-location")
|
||||||
|
async def collect_bgp_collector_location(
|
||||||
|
collector_id: str,
|
||||||
|
payload: CollectBGPCollectorLocationRequest,
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Run the shared location pipeline for a BGP route collector.
|
||||||
|
|
||||||
|
Mirrors ``POST /api/v1/visualization/compute-centers/{source_id}/collect-location``.
|
||||||
|
Returns ranked candidates from source coordinates and Nominatim queries
|
||||||
|
built around the collector's stored context (IXP / city / country). Stored
|
||||||
|
collector locations provide context only; they are not emitted as
|
||||||
|
candidates.
|
||||||
|
"""
|
||||||
|
if not collector_id or not collector_id.strip():
|
||||||
|
raise HTTPException(status_code=400, detail="collector_id is required")
|
||||||
|
|
||||||
|
legacy = get_bgp_collector_location_dict(collector_id) or {}
|
||||||
|
site = payload.site or legacy.get("matched_location_name")
|
||||||
|
city = payload.city or legacy.get("city")
|
||||||
|
country = payload.country or legacy.get("country")
|
||||||
|
operator = payload.operator or "RIPE NCC"
|
||||||
|
|
||||||
|
candidates, attempted_queries = collect_bgp_collector_location_candidates(
|
||||||
|
collector=collector_id,
|
||||||
|
site=site,
|
||||||
|
city=city,
|
||||||
|
country=country,
|
||||||
|
operator=operator,
|
||||||
|
)
|
||||||
|
llm_failure_reason = None
|
||||||
|
if not candidates:
|
||||||
|
query = build_bgp_collector_location_query(
|
||||||
|
collector=collector_id,
|
||||||
|
site=site,
|
||||||
|
city=city,
|
||||||
|
country=country,
|
||||||
|
operator=operator,
|
||||||
|
)
|
||||||
|
llm_result = None
|
||||||
|
try:
|
||||||
|
web_search_client = await get_web_search_client(db)
|
||||||
|
search_result = await collect_location_search_evidence(
|
||||||
|
web_search_client=web_search_client,
|
||||||
|
query=query,
|
||||||
|
entity_type="bgp_collector",
|
||||||
|
)
|
||||||
|
attempted_queries = [*attempted_queries, *search_result.attempted_queries]
|
||||||
|
if not search_result.evidence:
|
||||||
|
llm_failure_reason = search_result.failure_reason
|
||||||
|
raise RuntimeError(search_result.failure_reason or "no WebSearch evidence")
|
||||||
|
provider_client = await get_ai_provider_client(db)
|
||||||
|
llm_result = await collect_llm_location_fallback_candidate(
|
||||||
|
provider_client=provider_client,
|
||||||
|
query=query,
|
||||||
|
entity_type="bgp_collector",
|
||||||
|
db=db,
|
||||||
|
attempted_queries=attempted_queries,
|
||||||
|
search_evidence=search_result.evidence,
|
||||||
|
)
|
||||||
|
except Exception as exc:
|
||||||
|
if llm_failure_reason is None:
|
||||||
|
llm_failure_reason = f"LLM location factcheck unavailable: {exc}"
|
||||||
|
attempted_queries = [
|
||||||
|
*attempted_queries,
|
||||||
|
f"llm_factcheck:bgp_collector:{collector_id or 'unknown'}",
|
||||||
|
]
|
||||||
|
if llm_result is not None:
|
||||||
|
attempted_queries = [*attempted_queries, *llm_result.attempted_queries]
|
||||||
|
candidates = llm_result.candidates
|
||||||
|
llm_failure_reason = llm_result.failure_reason
|
||||||
|
|
||||||
|
context = {
|
||||||
|
"collector": collector_id,
|
||||||
|
"site": site,
|
||||||
|
"city": city,
|
||||||
|
"country": country,
|
||||||
|
"operator": operator,
|
||||||
|
}
|
||||||
|
|
||||||
|
if not candidates:
|
||||||
|
return {
|
||||||
|
"collector_id": collector_id,
|
||||||
|
"name": collector_id,
|
||||||
|
"success": False,
|
||||||
|
"failure_reason": (
|
||||||
|
"No source coordinates or online geocoding result reached"
|
||||||
|
" city-level precision for this collector."
|
||||||
|
),
|
||||||
|
"candidates": [],
|
||||||
|
"attempted_queries": list(attempted_queries),
|
||||||
|
"llm_failure_reason": llm_failure_reason,
|
||||||
|
"context": context,
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
"collector_id": collector_id,
|
||||||
|
"name": collector_id,
|
||||||
|
"success": True,
|
||||||
|
"candidates": [candidate.to_dict() for candidate in candidates],
|
||||||
|
"best_candidate": candidates[0].to_dict(),
|
||||||
|
"attempted_queries": list(attempted_queries),
|
||||||
|
"context": context,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
@router.get("/overview/summary")
|
@router.get("/overview/summary")
|
||||||
async def get_bgp_overview_summary(
|
async def get_bgp_overview_summary(
|
||||||
current_user: User = Depends(get_current_user),
|
current_user: User = Depends(get_current_user),
|
||||||
|
|||||||
@@ -102,7 +102,12 @@ def build_search_rank_sql(search: Optional[str]) -> str:
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
|
|
||||||
def serialize_collected_row(row, source_name_map: dict[str, str] | None = None) -> dict:
|
def serialize_collected_row(
|
||||||
|
row,
|
||||||
|
source_name_map: dict[str, str] | None = None,
|
||||||
|
*,
|
||||||
|
include_metadata: bool = True,
|
||||||
|
) -> dict:
|
||||||
metadata = row[7]
|
metadata = row[7]
|
||||||
source = row[1]
|
source = row[1]
|
||||||
return {
|
return {
|
||||||
@@ -120,7 +125,7 @@ def serialize_collected_row(row, source_name_map: dict[str, str] | None = None)
|
|||||||
"longitude": get_metadata_field(metadata, "longitude"),
|
"longitude": get_metadata_field(metadata, "longitude"),
|
||||||
"value": get_metadata_field(metadata, "value"),
|
"value": get_metadata_field(metadata, "value"),
|
||||||
"unit": get_metadata_field(metadata, "unit"),
|
"unit": get_metadata_field(metadata, "unit"),
|
||||||
"metadata": metadata,
|
"metadata": metadata if include_metadata else None,
|
||||||
"cores": get_metadata_field(metadata, "cores"),
|
"cores": get_metadata_field(metadata, "cores"),
|
||||||
"rmax": get_metadata_field(metadata, "rmax"),
|
"rmax": get_metadata_field(metadata, "rmax"),
|
||||||
"rpeak": get_metadata_field(metadata, "rpeak"),
|
"rpeak": get_metadata_field(metadata, "rpeak"),
|
||||||
@@ -145,6 +150,7 @@ async def list_collected_data(
|
|||||||
search: Optional[str] = Query(None, description="搜索名称"),
|
search: Optional[str] = Query(None, description="搜索名称"),
|
||||||
page: int = Query(1, ge=1, description="页码"),
|
page: int = Query(1, ge=1, description="页码"),
|
||||||
page_size: int = Query(20, ge=1, le=100, description="每页数量"),
|
page_size: int = Query(20, ge=1, le=100, description="每页数量"),
|
||||||
|
include_metadata: bool = Query(True, description="是否返回完整 metadata 字段"),
|
||||||
current_user: User = Depends(get_current_user),
|
current_user: User = Depends(get_current_user),
|
||||||
db: AsyncSession = Depends(get_db),
|
db: AsyncSession = Depends(get_db),
|
||||||
):
|
):
|
||||||
@@ -201,7 +207,7 @@ async def list_collected_data(
|
|||||||
|
|
||||||
data = []
|
data = []
|
||||||
for row in rows:
|
for row in rows:
|
||||||
data.append(serialize_collected_row(row[:11], source_name_map))
|
data.append(serialize_collected_row(row[:11], source_name_map, include_metadata=include_metadata))
|
||||||
|
|
||||||
return {
|
return {
|
||||||
"total": total,
|
"total": total,
|
||||||
|
|||||||
98
backend/app/api/v1/data_products.py
Normal file
98
backend/app/api/v1/data_products.py
Normal file
@@ -0,0 +1,98 @@
|
|||||||
|
from datetime import UTC, datetime
|
||||||
|
|
||||||
|
from fastapi import APIRouter, Depends, HTTPException
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.api.v1.visualization import get_visualization_geo_summary
|
||||||
|
from app.core.time import to_iso8601_utc
|
||||||
|
from app.db.session import get_db
|
||||||
|
|
||||||
|
router = APIRouter()
|
||||||
|
|
||||||
|
PRODUCT_DEFINITIONS: dict[str, dict] = {
|
||||||
|
"vessels": {
|
||||||
|
"name": "船只",
|
||||||
|
"sources": ["aisstream_vessels", "barentswatch_vessels"],
|
||||||
|
"primary_stat_key": "vessel_count",
|
||||||
|
"stat_keys": ["vessel_count", "vessel_raw_unique_mmsi", "vessel_legacy_unique_mmsi"],
|
||||||
|
},
|
||||||
|
"cables": {
|
||||||
|
"name": "海底光缆",
|
||||||
|
"sources": [
|
||||||
|
"arcgis_cables",
|
||||||
|
"arcgis_landing_points",
|
||||||
|
"arcgis_cable_landing_relation",
|
||||||
|
"telegeography_cables",
|
||||||
|
"telegeography_landing",
|
||||||
|
"telegeography_systems",
|
||||||
|
"fao_landing_points",
|
||||||
|
],
|
||||||
|
"primary_stat_key": "cable_count",
|
||||||
|
"stat_keys": ["cable_count", "landing_point_count"],
|
||||||
|
},
|
||||||
|
"satellites": {
|
||||||
|
"name": "卫星",
|
||||||
|
"sources": ["celestrak_tle", "spacetrack_tle"],
|
||||||
|
"primary_stat_key": "satellite_count",
|
||||||
|
"stat_keys": ["satellite_count"],
|
||||||
|
},
|
||||||
|
"bgp": {
|
||||||
|
"name": "BGP",
|
||||||
|
"sources": [
|
||||||
|
"ris_live_bgp",
|
||||||
|
"bgpstream_bgp",
|
||||||
|
"iptoasn_prefix_geo",
|
||||||
|
"opengeofeed_prefix_geo",
|
||||||
|
"nro_delegated_prefix_geo",
|
||||||
|
],
|
||||||
|
"primary_stat_key": "bgp_event_count",
|
||||||
|
"stat_keys": ["bgp_event_count", "bgp_incident_count", "bgp_anomaly_count", "bgp_collector_count"],
|
||||||
|
},
|
||||||
|
"compute": {
|
||||||
|
"name": "算力",
|
||||||
|
"sources": ["top500", "epoch_ai_gpu"],
|
||||||
|
"primary_stat_key": "compute_center_count",
|
||||||
|
"stat_keys": ["compute_center_count", "supercomputer_count", "gpu_cluster_count"],
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _build_product_status(product_id: str, summary: dict) -> dict:
|
||||||
|
definition = PRODUCT_DEFINITIONS[product_id]
|
||||||
|
stats = summary.get("stats", {})
|
||||||
|
product_stats = {key: stats.get(key, 0) for key in definition["stat_keys"]}
|
||||||
|
total_count = int(product_stats.get(definition["primary_stat_key"]) or 0)
|
||||||
|
return {
|
||||||
|
"product_id": product_id,
|
||||||
|
"name": definition["name"],
|
||||||
|
"sources": definition["sources"],
|
||||||
|
"generated_at": summary.get("generated_at") or to_iso8601_utc(datetime.now(UTC)),
|
||||||
|
"total_count": total_count,
|
||||||
|
"stats": product_stats,
|
||||||
|
"build_state": "ready",
|
||||||
|
"stats_scope": "global",
|
||||||
|
"stats_freshness": "cached_or_indexed",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("")
|
||||||
|
async def list_data_products(db: AsyncSession = Depends(get_db)):
|
||||||
|
summary = await get_visualization_geo_summary(db)
|
||||||
|
return {
|
||||||
|
"generated_at": summary.get("generated_at"),
|
||||||
|
"data": [
|
||||||
|
_build_product_status(product_id, summary)
|
||||||
|
for product_id in PRODUCT_DEFINITIONS
|
||||||
|
],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/{product_id}/status")
|
||||||
|
async def get_data_product_status(
|
||||||
|
product_id: str,
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
if product_id not in PRODUCT_DEFINITIONS:
|
||||||
|
raise HTTPException(status_code=404, detail="Unknown data product")
|
||||||
|
summary = await get_visualization_geo_summary(db)
|
||||||
|
return _build_product_status(product_id, summary)
|
||||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
102
backend/app/api/v1/docs.py
Normal file
102
backend/app/api/v1/docs.py
Normal file
@@ -0,0 +1,102 @@
|
|||||||
|
"""Authenticated documentation APIs."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from fastapi import APIRouter, Depends, HTTPException, status
|
||||||
|
from fastapi.security import HTTPAuthorizationCredentials, HTTPBearer
|
||||||
|
from sqlalchemy import text
|
||||||
|
|
||||||
|
from app.core.security import decode_token
|
||||||
|
from app.db.session import async_session_factory
|
||||||
|
from app.models.user import User
|
||||||
|
from app.services.docs_gatekeeper import (
|
||||||
|
DOCS_BY_SLUG,
|
||||||
|
VALID_DOCS_LANGS,
|
||||||
|
can_read_doc,
|
||||||
|
catalog_for_user,
|
||||||
|
doc_path_for,
|
||||||
|
title_for,
|
||||||
|
)
|
||||||
|
|
||||||
|
router = APIRouter()
|
||||||
|
optional_bearer = HTTPBearer(auto_error=False)
|
||||||
|
|
||||||
|
|
||||||
|
async def get_optional_current_user(
|
||||||
|
credentials: HTTPAuthorizationCredentials | None = Depends(optional_bearer),
|
||||||
|
) -> User | None:
|
||||||
|
if credentials is None:
|
||||||
|
return None
|
||||||
|
|
||||||
|
payload = decode_token(credentials.credentials)
|
||||||
|
if payload is None or payload.get("type") != "access" or payload.get("sub") is None:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_401_UNAUTHORIZED,
|
||||||
|
detail="Invalid token",
|
||||||
|
)
|
||||||
|
|
||||||
|
async with async_session_factory() as db:
|
||||||
|
result = await db.execute(
|
||||||
|
text(
|
||||||
|
"SELECT id, username, email, password_hash, role, is_active, gatekeeper_groups FROM users WHERE id = :id"
|
||||||
|
),
|
||||||
|
{"id": int(payload["sub"])},
|
||||||
|
)
|
||||||
|
row = result.fetchone()
|
||||||
|
if row is None or not row[5]:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_401_UNAUTHORIZED,
|
||||||
|
detail="User not found or inactive",
|
||||||
|
)
|
||||||
|
|
||||||
|
user = User()
|
||||||
|
user.id = row[0]
|
||||||
|
user.username = row[1]
|
||||||
|
user.email = row[2]
|
||||||
|
user.password_hash = row[3]
|
||||||
|
user.role = row[4]
|
||||||
|
user.is_active = row[5]
|
||||||
|
user.gatekeeper_groups = row[6] or []
|
||||||
|
return user
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/catalog")
|
||||||
|
async def get_docs_catalog(current_user: User | None = Depends(get_optional_current_user)):
|
||||||
|
return {
|
||||||
|
"items": catalog_for_user(current_user),
|
||||||
|
"authenticated": current_user is not None,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/{lang}/{slug}")
|
||||||
|
async def get_doc_content(
|
||||||
|
lang: str,
|
||||||
|
slug: str,
|
||||||
|
current_user: User | None = Depends(get_optional_current_user),
|
||||||
|
):
|
||||||
|
if lang not in VALID_DOCS_LANGS:
|
||||||
|
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="Document not found")
|
||||||
|
|
||||||
|
entry = DOCS_BY_SLUG.get(slug)
|
||||||
|
if entry is None:
|
||||||
|
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="Document not found")
|
||||||
|
|
||||||
|
path = doc_path_for(entry, lang)
|
||||||
|
if not path.exists():
|
||||||
|
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="Document not found")
|
||||||
|
|
||||||
|
if not can_read_doc(entry, current_user):
|
||||||
|
if current_user is None:
|
||||||
|
raise HTTPException(status_code=status.HTTP_401_UNAUTHORIZED, detail="Authentication required")
|
||||||
|
raise HTTPException(status_code=status.HTTP_403_FORBIDDEN, detail="Insufficient Docs permissions")
|
||||||
|
|
||||||
|
return {
|
||||||
|
"slug": entry.slug,
|
||||||
|
"filename": entry.filename,
|
||||||
|
"lang": lang,
|
||||||
|
"title": title_for(entry, lang),
|
||||||
|
"group": entry.group,
|
||||||
|
"order": entry.order,
|
||||||
|
"access": entry.access,
|
||||||
|
"markdown": path.read_text(encoding="utf-8"),
|
||||||
|
}
|
||||||
441
backend/app/api/v1/earth.py
Normal file
441
backend/app/api/v1/earth.py
Normal file
@@ -0,0 +1,441 @@
|
|||||||
|
"""Earth asset management APIs."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
from uuid import uuid4
|
||||||
|
|
||||||
|
from fastapi import APIRouter, Depends, File, HTTPException, Request, UploadFile, status
|
||||||
|
from fastapi.security import HTTPAuthorizationCredentials, HTTPBearer
|
||||||
|
from pydantic import BaseModel, Field
|
||||||
|
from sqlalchemy import delete, func, select, text
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.config import settings as app_settings
|
||||||
|
from app.core.security import decode_token, get_current_user, redis_client
|
||||||
|
from app.db.session import get_db
|
||||||
|
from app.models.collected_data import CollectedData
|
||||||
|
from app.models.datasource import DataSource
|
||||||
|
from app.models.datasource_config import DataSourceConfig
|
||||||
|
from app.models.system_setting import SystemSetting
|
||||||
|
from app.models.user import User
|
||||||
|
from app.services.tv_streams import get_tv_settings_payload
|
||||||
|
from app.services.earth_boundaries import (
|
||||||
|
EarthBoundaryBuildError,
|
||||||
|
get_boundary_build_status,
|
||||||
|
get_boundary_status,
|
||||||
|
save_boundary_config,
|
||||||
|
start_boundary_build_job,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
router = APIRouter()
|
||||||
|
optional_bearer = HTTPBearer(auto_error=False)
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[4]
|
||||||
|
EARTH_BRAND_ASSET_DIR = REPO_ROOT / "data" / "earth-brand"
|
||||||
|
EARTH_BRAND_ASSET_URL_PREFIX = "/earth-brand-assets"
|
||||||
|
EARTH_BRAND_CATEGORY = "earth_brand"
|
||||||
|
EARTH_ABOUT_CATEGORY = "earth_about"
|
||||||
|
SYSTEM_SETTINGS_CATEGORY = "system"
|
||||||
|
MAX_EARTH_BRAND_ASSET_BYTES = 3 * 1024 * 1024
|
||||||
|
ALLOWED_EARTH_BRAND_EXTENSIONS = {".png", ".jpg", ".jpeg", ".webp", ".svg"}
|
||||||
|
|
||||||
|
|
||||||
|
def _app_version_label() -> str:
|
||||||
|
version = str(app_settings.VERSION or "").strip() or "0.0.0"
|
||||||
|
return version if version.startswith("v") else f"v{version}"
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_EARTH_BRAND = {
|
||||||
|
"logo_src": "/earth/assets/brand/earth-logo.png",
|
||||||
|
"title_src": "/earth/assets/brand/title-zh.png",
|
||||||
|
"title_text": "智能星球计划",
|
||||||
|
"subtitle": "现实层宇宙全息感知系统",
|
||||||
|
"description": "卫星 · 海底光缆 · 算力基础设施",
|
||||||
|
"aria_label": "智能星球计划品牌标识",
|
||||||
|
"title_alt": "智能星球计划",
|
||||||
|
}
|
||||||
|
|
||||||
|
DEFAULT_EARTH_ABOUT = {
|
||||||
|
"logo_src": "/earth/assets/brand/lim-logo.png",
|
||||||
|
"kicker": "About",
|
||||||
|
"title": "智能星球计划",
|
||||||
|
"version": _app_version_label(),
|
||||||
|
"description": "面向临空场景下的智能媒体研究、全球态势感知与多源开放数据巡航,提供可视化观测、事件聚合与交互式探索能力。",
|
||||||
|
"meta": [
|
||||||
|
{"label": "出品方", "value": "浙江大学临空智能媒体研究院"},
|
||||||
|
{"label": "策划人", "value": "方兴东、黄柳青"},
|
||||||
|
{"label": "产品兼开发者", "value": "钱坤、张鸽、齐鹏"},
|
||||||
|
],
|
||||||
|
}
|
||||||
|
EARTH_ABOUT_LEGACY_PLANNER_VALUE = "黄柳青"
|
||||||
|
|
||||||
|
|
||||||
|
class EarthBoundaryConfigPayload(BaseModel):
|
||||||
|
config: dict[str, Any] = Field(default_factory=dict)
|
||||||
|
|
||||||
|
|
||||||
|
class EarthBrandPayload(BaseModel):
|
||||||
|
logo_src: str = Field(default=DEFAULT_EARTH_BRAND["logo_src"], max_length=1000)
|
||||||
|
title_src: str = Field(default=DEFAULT_EARTH_BRAND["title_src"], max_length=1000)
|
||||||
|
title_text: str = Field(default=DEFAULT_EARTH_BRAND["title_text"], max_length=120)
|
||||||
|
subtitle: str = Field(default=DEFAULT_EARTH_BRAND["subtitle"], max_length=160)
|
||||||
|
description: str = Field(default=DEFAULT_EARTH_BRAND["description"], max_length=200)
|
||||||
|
aria_label: str = Field(default=DEFAULT_EARTH_BRAND["aria_label"], max_length=200)
|
||||||
|
title_alt: str = Field(default=DEFAULT_EARTH_BRAND["title_alt"], max_length=200)
|
||||||
|
|
||||||
|
|
||||||
|
class EarthAboutMetaItem(BaseModel):
|
||||||
|
label: str = Field(default="", max_length=80)
|
||||||
|
value: str = Field(default="", max_length=240)
|
||||||
|
|
||||||
|
|
||||||
|
class EarthAboutPayload(BaseModel):
|
||||||
|
logo_src: str = Field(default=DEFAULT_EARTH_ABOUT["logo_src"], max_length=1000)
|
||||||
|
kicker: str = Field(default=DEFAULT_EARTH_ABOUT["kicker"], max_length=80)
|
||||||
|
title: str = Field(default=DEFAULT_EARTH_ABOUT["title"], max_length=160)
|
||||||
|
version: str = Field(default=DEFAULT_EARTH_ABOUT["version"], max_length=80)
|
||||||
|
description: str = Field(default=DEFAULT_EARTH_ABOUT["description"], max_length=800)
|
||||||
|
meta: list[EarthAboutMetaItem] = Field(default_factory=list)
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_earth_brand_payload(payload: dict[str, Any] | None) -> dict[str, str]:
|
||||||
|
merged = DEFAULT_EARTH_BRAND.copy()
|
||||||
|
if payload:
|
||||||
|
for key in DEFAULT_EARTH_BRAND:
|
||||||
|
value = payload.get(key)
|
||||||
|
if value is not None:
|
||||||
|
merged[key] = str(value).strip()
|
||||||
|
|
||||||
|
if not merged["title_text"]:
|
||||||
|
merged["title_text"] = DEFAULT_EARTH_BRAND["title_text"]
|
||||||
|
if not merged["aria_label"]:
|
||||||
|
merged["aria_label"] = merged["title_text"]
|
||||||
|
if not merged["title_alt"]:
|
||||||
|
merged["title_alt"] = merged["title_text"]
|
||||||
|
return merged
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_earth_about_payload(payload: dict[str, Any] | None) -> dict[str, Any]:
|
||||||
|
merged: dict[str, Any] = {
|
||||||
|
key: value
|
||||||
|
for key, value in DEFAULT_EARTH_ABOUT.items()
|
||||||
|
if key != "meta"
|
||||||
|
}
|
||||||
|
raw_meta = DEFAULT_EARTH_ABOUT["meta"]
|
||||||
|
if payload:
|
||||||
|
for key in ("logo_src", "kicker", "title", "description"):
|
||||||
|
value = payload.get(key)
|
||||||
|
if value is not None:
|
||||||
|
merged[key] = str(value).strip()
|
||||||
|
raw_meta = payload.get("meta") if isinstance(payload.get("meta"), list) else raw_meta
|
||||||
|
merged["version"] = _app_version_label()
|
||||||
|
|
||||||
|
for key, default_value in DEFAULT_EARTH_ABOUT.items():
|
||||||
|
if key == "meta":
|
||||||
|
continue
|
||||||
|
if not merged.get(key):
|
||||||
|
merged[key] = default_value
|
||||||
|
|
||||||
|
normalized_meta: list[dict[str, str]] = []
|
||||||
|
for item in raw_meta:
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
continue
|
||||||
|
label = str(item.get("label") or "").strip()
|
||||||
|
value = str(item.get("value") or "").strip()
|
||||||
|
if label == "策划人" and value == EARTH_ABOUT_LEGACY_PLANNER_VALUE:
|
||||||
|
value = "方兴东、黄柳青"
|
||||||
|
if label or value:
|
||||||
|
normalized_meta.append({"label": label, "value": value})
|
||||||
|
if not normalized_meta:
|
||||||
|
normalized_meta = [dict(item) for item in DEFAULT_EARTH_ABOUT["meta"]]
|
||||||
|
merged["meta"] = normalized_meta
|
||||||
|
return merged
|
||||||
|
|
||||||
|
|
||||||
|
def _is_demo_mode_enabled(payload: Any) -> bool:
|
||||||
|
return bool(payload.get("demo_mode")) if isinstance(payload, dict) else False
|
||||||
|
|
||||||
|
|
||||||
|
async def _get_earth_brand_record(db: AsyncSession) -> SystemSetting | None:
|
||||||
|
result = await db.execute(
|
||||||
|
select(SystemSetting).where(SystemSetting.category == EARTH_BRAND_CATEGORY)
|
||||||
|
)
|
||||||
|
return result.scalar_one_or_none()
|
||||||
|
|
||||||
|
|
||||||
|
async def _get_earth_brand_payload(db: AsyncSession) -> dict[str, Any]:
|
||||||
|
record = await _get_earth_brand_record(db)
|
||||||
|
return {
|
||||||
|
"brand": _normalize_earth_brand_payload(record.payload if record else None),
|
||||||
|
"is_default": record is None,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def _get_earth_about_record(db: AsyncSession) -> SystemSetting | None:
|
||||||
|
result = await db.execute(
|
||||||
|
select(SystemSetting).where(SystemSetting.category == EARTH_ABOUT_CATEGORY)
|
||||||
|
)
|
||||||
|
return result.scalar_one_or_none()
|
||||||
|
|
||||||
|
|
||||||
|
async def _get_earth_about_payload(db: AsyncSession) -> dict[str, Any]:
|
||||||
|
record = await _get_earth_about_record(db)
|
||||||
|
return {
|
||||||
|
"about": _normalize_earth_about_payload(record.payload if record else None),
|
||||||
|
"is_default": record is None,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def _get_optional_current_user(
|
||||||
|
credentials: HTTPAuthorizationCredentials | None = Depends(optional_bearer),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
) -> User | None:
|
||||||
|
if credentials is None:
|
||||||
|
return None
|
||||||
|
token = credentials.credentials
|
||||||
|
if redis_client.sismember("blacklisted_tokens", token):
|
||||||
|
return None
|
||||||
|
payload = decode_token(token)
|
||||||
|
if payload is None or payload.get("type") != "access":
|
||||||
|
return None
|
||||||
|
user_id = payload.get("sub")
|
||||||
|
if user_id is None:
|
||||||
|
return None
|
||||||
|
result = await db.execute(
|
||||||
|
text(
|
||||||
|
"SELECT id, username, email, password_hash, role, is_active, gatekeeper_groups FROM users WHERE id = :id"
|
||||||
|
),
|
||||||
|
{"id": int(user_id)},
|
||||||
|
)
|
||||||
|
row = result.fetchone()
|
||||||
|
if row is None or not row[5]:
|
||||||
|
return None
|
||||||
|
user = User()
|
||||||
|
user.id = row[0]
|
||||||
|
user.username = row[1]
|
||||||
|
user.email = row[2]
|
||||||
|
user.password_hash = row[3]
|
||||||
|
user.role = row[4]
|
||||||
|
user.is_active = row[5]
|
||||||
|
user.gatekeeper_groups = row[6] or []
|
||||||
|
return user
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/brand")
|
||||||
|
async def get_earth_brand(db: AsyncSession = Depends(get_db)):
|
||||||
|
return await _get_earth_brand_payload(db)
|
||||||
|
|
||||||
|
|
||||||
|
@router.put("/brand")
|
||||||
|
async def update_earth_brand(
|
||||||
|
payload: EarthBrandPayload,
|
||||||
|
_current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
normalized = _normalize_earth_brand_payload(payload.model_dump())
|
||||||
|
record = await _get_earth_brand_record(db)
|
||||||
|
if record is None:
|
||||||
|
record = SystemSetting(category=EARTH_BRAND_CATEGORY, payload=normalized)
|
||||||
|
db.add(record)
|
||||||
|
else:
|
||||||
|
record.payload = normalized
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(record)
|
||||||
|
return {"status": "updated", "brand": _normalize_earth_brand_payload(record.payload), "is_default": False}
|
||||||
|
|
||||||
|
|
||||||
|
@router.delete("/brand")
|
||||||
|
@router.post("/brand/reset")
|
||||||
|
async def reset_earth_brand(
|
||||||
|
_current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
await db.execute(delete(SystemSetting).where(SystemSetting.category == EARTH_BRAND_CATEGORY))
|
||||||
|
await db.commit()
|
||||||
|
return {"status": "reset", "brand": DEFAULT_EARTH_BRAND.copy(), "is_default": True}
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/brand/assets")
|
||||||
|
async def upload_earth_brand_asset(
|
||||||
|
file: UploadFile = File(...),
|
||||||
|
_current_user: User = Depends(get_current_user),
|
||||||
|
):
|
||||||
|
original_name = file.filename or ""
|
||||||
|
extension = Path(original_name).suffix.lower()
|
||||||
|
if extension not in ALLOWED_EARTH_BRAND_EXTENSIONS:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=400,
|
||||||
|
detail={
|
||||||
|
"code": "unsupported_file_type",
|
||||||
|
"message": "Only png, jpg, jpeg, webp, and svg brand assets are supported.",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
content = await file.read(MAX_EARTH_BRAND_ASSET_BYTES + 1)
|
||||||
|
if len(content) > MAX_EARTH_BRAND_ASSET_BYTES:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=400,
|
||||||
|
detail={
|
||||||
|
"code": "file_too_large",
|
||||||
|
"message": "Brand asset must be 3 MB or smaller.",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
EARTH_BRAND_ASSET_DIR.mkdir(parents=True, exist_ok=True)
|
||||||
|
safe_name = f"{uuid4().hex}{extension}"
|
||||||
|
destination = EARTH_BRAND_ASSET_DIR / safe_name
|
||||||
|
destination.write_bytes(content)
|
||||||
|
asset_url = f"{EARTH_BRAND_ASSET_URL_PREFIX}/{safe_name}"
|
||||||
|
return {"url": asset_url, "filename": safe_name, "content_type": file.content_type}
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/about")
|
||||||
|
async def get_earth_about(db: AsyncSession = Depends(get_db)):
|
||||||
|
return await _get_earth_about_payload(db)
|
||||||
|
|
||||||
|
|
||||||
|
@router.put("/about")
|
||||||
|
async def update_earth_about(
|
||||||
|
payload: EarthAboutPayload,
|
||||||
|
_current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
normalized = _normalize_earth_about_payload(payload.model_dump())
|
||||||
|
record = await _get_earth_about_record(db)
|
||||||
|
if record is None:
|
||||||
|
record = SystemSetting(category=EARTH_ABOUT_CATEGORY, payload=normalized)
|
||||||
|
db.add(record)
|
||||||
|
else:
|
||||||
|
record.payload = normalized
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(record)
|
||||||
|
return {"status": "updated", "about": _normalize_earth_about_payload(record.payload), "is_default": False}
|
||||||
|
|
||||||
|
|
||||||
|
@router.delete("/about")
|
||||||
|
async def reset_earth_about(
|
||||||
|
_current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
await db.execute(delete(SystemSetting).where(SystemSetting.category == EARTH_ABOUT_CATEGORY))
|
||||||
|
await db.commit()
|
||||||
|
return {"status": "reset", "about": _normalize_earth_about_payload(None), "is_default": True}
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/oobe-status")
|
||||||
|
async def get_earth_oobe_status(
|
||||||
|
current_user: User | None = Depends(_get_optional_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
current_count_result = await db.execute(
|
||||||
|
select(func.count(CollectedData.id)).where(CollectedData.is_current.is_(True))
|
||||||
|
)
|
||||||
|
current_record_count = int(current_count_result.scalar() or 0)
|
||||||
|
system_result = await db.execute(
|
||||||
|
select(SystemSetting).where(SystemSetting.category == SYSTEM_SETTINGS_CATEGORY)
|
||||||
|
)
|
||||||
|
system_record = system_result.scalar_one_or_none()
|
||||||
|
demo_mode = _is_demo_mode_enabled(system_record.payload if system_record else None)
|
||||||
|
|
||||||
|
datasource_count_result = await db.execute(select(func.count(DataSource.id)))
|
||||||
|
datasource_count = int(datasource_count_result.scalar() or 0)
|
||||||
|
active_datasource_count_result = await db.execute(
|
||||||
|
select(func.count(DataSource.id)).where(DataSource.is_active.is_(True))
|
||||||
|
)
|
||||||
|
active_datasource_count = int(active_datasource_count_result.scalar() or 0)
|
||||||
|
config_result = await db.execute(select(func.count(DataSourceConfig.id)))
|
||||||
|
custom_config_count = int(config_result.scalar() or 0)
|
||||||
|
|
||||||
|
tv_payload = await get_tv_settings_payload(db)
|
||||||
|
tv_sources = tv_payload.get("sources") if isinstance(tv_payload, dict) else []
|
||||||
|
tv_source_count = len(tv_sources) if isinstance(tv_sources, list) else 0
|
||||||
|
|
||||||
|
boundary_status = get_boundary_status()
|
||||||
|
has_core_layers = bool(boundary_status.get("ready") or boundary_status.get("available") or boundary_status.get("status") in {"ready", "built", "ok"})
|
||||||
|
has_collected_data = current_record_count > 0
|
||||||
|
ready = has_collected_data
|
||||||
|
|
||||||
|
suggestions: list[str] = []
|
||||||
|
if demo_mode:
|
||||||
|
suggestions.append("演示模式已开启")
|
||||||
|
if not current_user:
|
||||||
|
suggestions.append("登录控制台")
|
||||||
|
if not has_collected_data:
|
||||||
|
suggestions.append("触发数据源采集")
|
||||||
|
if not custom_config_count:
|
||||||
|
suggestions.append("确认采集器配置")
|
||||||
|
if not has_core_layers:
|
||||||
|
suggestions.append("构建或启用 Earth 图层")
|
||||||
|
|
||||||
|
return {
|
||||||
|
"ready": ready,
|
||||||
|
"demo_mode": demo_mode,
|
||||||
|
"authenticated": current_user is not None,
|
||||||
|
"needs_login": current_user is None and not ready and not demo_mode,
|
||||||
|
"has_collected_data": has_collected_data,
|
||||||
|
"has_tv_sources": tv_source_count > 0,
|
||||||
|
"has_core_layers": has_core_layers,
|
||||||
|
"current_record_count": current_record_count,
|
||||||
|
"datasource_count": datasource_count,
|
||||||
|
"active_datasource_count": active_datasource_count,
|
||||||
|
"custom_config_count": custom_config_count,
|
||||||
|
"tv_source_count": tv_source_count,
|
||||||
|
"suggestions": suggestions,
|
||||||
|
"login_url": "/login?next=/datasources",
|
||||||
|
"datasources_url": "/datasources",
|
||||||
|
"collection_url": "/collection-management",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/boundaries/status")
|
||||||
|
async def get_earth_boundary_status():
|
||||||
|
return get_boundary_status()
|
||||||
|
|
||||||
|
def _is_loopback_request(request: Request) -> bool:
|
||||||
|
host = request.client.host if request.client else ""
|
||||||
|
return host in {"127.0.0.1", "::1", "localhost"} or host.startswith("127.")
|
||||||
|
|
||||||
|
|
||||||
|
def _require_local_or_user(request: Request, user: User | None) -> None:
|
||||||
|
if user is not None or _is_loopback_request(request):
|
||||||
|
return
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_401_UNAUTHORIZED,
|
||||||
|
detail="Authentication required outside localhost",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@router.put("/boundaries/config")
|
||||||
|
async def update_earth_boundary_config(
|
||||||
|
payload: EarthBoundaryConfigPayload,
|
||||||
|
_current_user: User = Depends(get_current_user),
|
||||||
|
):
|
||||||
|
try:
|
||||||
|
return save_boundary_config(payload.config)
|
||||||
|
except EarthBoundaryBuildError as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=400,
|
||||||
|
detail={"code": exc.code, "message": str(exc), "details": exc.details},
|
||||||
|
) from exc
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/boundaries/build")
|
||||||
|
async def build_earth_boundary_assets(
|
||||||
|
request: Request,
|
||||||
|
current_user: User | None = Depends(_get_optional_current_user),
|
||||||
|
):
|
||||||
|
_require_local_or_user(request, current_user)
|
||||||
|
try:
|
||||||
|
return await start_boundary_build_job()
|
||||||
|
except EarthBoundaryBuildError as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=400,
|
||||||
|
detail={"code": exc.code, "message": str(exc), "details": exc.details},
|
||||||
|
) from exc
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/boundaries/build/status")
|
||||||
|
async def get_earth_boundary_build_status():
|
||||||
|
return get_boundary_build_status()
|
||||||
190
backend/app/api/v1/interactables.py
Normal file
190
backend/app/api/v1/interactables.py
Normal file
@@ -0,0 +1,190 @@
|
|||||||
|
"""CRUD APIs for persistent Earth interactables."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from fastapi import APIRouter, Depends, HTTPException, Query, Response, status
|
||||||
|
from pydantic import BaseModel, Field, field_validator
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.security import get_current_user
|
||||||
|
from app.db.session import get_db
|
||||||
|
from app.models.earth_interactable import EarthInteractable
|
||||||
|
from app.models.user import User
|
||||||
|
from app.services.earth_interactables import (
|
||||||
|
build_interactable_event,
|
||||||
|
interactables_to_geojson,
|
||||||
|
invalidate_interactable_cache,
|
||||||
|
list_interactables,
|
||||||
|
normalize_interactable_id,
|
||||||
|
publish_interactable_event,
|
||||||
|
serialize_interactable,
|
||||||
|
)
|
||||||
|
from app.services.earth_layer_cache import (
|
||||||
|
EarthLayerCachePolicy,
|
||||||
|
earth_layer_cache,
|
||||||
|
get_or_build_layer_payload,
|
||||||
|
)
|
||||||
|
|
||||||
|
router = APIRouter()
|
||||||
|
INTERACTABLE_CACHE_POLICY = EarthLayerCachePolicy(
|
||||||
|
fresh_ttl_seconds=60,
|
||||||
|
stale_ttl_seconds=10 * 60,
|
||||||
|
max_features=5000,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class InteractableCreate(BaseModel):
|
||||||
|
id: str | None = Field(default=None, max_length=160)
|
||||||
|
layer: str = Field(default="default", min_length=1, max_length=80)
|
||||||
|
kind: str = Field(default="default", min_length=1, max_length=80)
|
||||||
|
label: str = Field(default="", max_length=255)
|
||||||
|
description: str = Field(default="", max_length=4000)
|
||||||
|
latitude: float = Field(ge=-90, le=90)
|
||||||
|
longitude: float = Field(ge=-180, le=180)
|
||||||
|
altitude: float | None = None
|
||||||
|
properties: dict[str, Any] = Field(default_factory=dict)
|
||||||
|
|
||||||
|
@field_validator("layer", "kind")
|
||||||
|
@classmethod
|
||||||
|
def normalize_key(cls, value: str) -> str:
|
||||||
|
normalized = str(value or "").strip()
|
||||||
|
if not normalized:
|
||||||
|
raise ValueError("must not be empty")
|
||||||
|
return normalized
|
||||||
|
|
||||||
|
|
||||||
|
class InteractableUpdate(BaseModel):
|
||||||
|
layer: str | None = Field(default=None, min_length=1, max_length=80)
|
||||||
|
kind: str | None = Field(default=None, min_length=1, max_length=80)
|
||||||
|
label: str | None = Field(default=None, max_length=255)
|
||||||
|
description: str | None = Field(default=None, max_length=4000)
|
||||||
|
latitude: float | None = Field(default=None, ge=-90, le=90)
|
||||||
|
longitude: float | None = Field(default=None, ge=-180, le=180)
|
||||||
|
altitude: float | None = None
|
||||||
|
properties: dict[str, Any] | None = None
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("")
|
||||||
|
async def get_interactables(
|
||||||
|
response: Response,
|
||||||
|
layer: str | None = Query(default=None),
|
||||||
|
include_deleted: bool = Query(default=False),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
items = await list_interactables(db, layer=layer, include_deleted=include_deleted)
|
||||||
|
response.headers["X-Planet-Interactables-Count"] = str(len(items))
|
||||||
|
return {"items": [serialize_interactable(item) for item in items]}
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/geojson")
|
||||||
|
async def get_interactables_geojson(
|
||||||
|
response: Response,
|
||||||
|
layer: str | None = Query(default=None),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
async def build_payload() -> dict[str, Any]:
|
||||||
|
items = await list_interactables(db, layer=layer)
|
||||||
|
return interactables_to_geojson(items)
|
||||||
|
|
||||||
|
payload = await get_or_build_layer_payload(
|
||||||
|
key=earth_layer_cache.key("interactables", layer=layer or "all"),
|
||||||
|
policy=INTERACTABLE_CACHE_POLICY,
|
||||||
|
builder=build_payload,
|
||||||
|
response=response,
|
||||||
|
)
|
||||||
|
response.headers["X-Planet-Interactables-Count"] = str(len(payload.get("features") or []))
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("", status_code=status.HTTP_201_CREATED)
|
||||||
|
async def create_interactable(
|
||||||
|
payload: InteractableCreate,
|
||||||
|
_current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
record_id = normalize_interactable_id(payload.id)
|
||||||
|
existing = await db.get(EarthInteractable, record_id)
|
||||||
|
if existing and not existing.is_deleted:
|
||||||
|
raise HTTPException(status_code=409, detail="Interactable already exists")
|
||||||
|
|
||||||
|
if existing is None:
|
||||||
|
record = EarthInteractable(id=record_id)
|
||||||
|
db.add(record)
|
||||||
|
else:
|
||||||
|
record = existing
|
||||||
|
record.is_deleted = False
|
||||||
|
record.deleted_at = None
|
||||||
|
record.revision += 1
|
||||||
|
|
||||||
|
record.layer = payload.layer
|
||||||
|
record.kind = payload.kind
|
||||||
|
record.label = payload.label
|
||||||
|
record.description = payload.description
|
||||||
|
record.latitude = payload.latitude
|
||||||
|
record.longitude = payload.longitude
|
||||||
|
record.altitude = payload.altitude
|
||||||
|
record.properties = payload.properties
|
||||||
|
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(record)
|
||||||
|
invalidate_interactable_cache(record.layer)
|
||||||
|
await publish_interactable_event("created", record)
|
||||||
|
return {"item": serialize_interactable(record)}
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/{interactable_id}")
|
||||||
|
async def get_interactable(interactable_id: str, db: AsyncSession = Depends(get_db)):
|
||||||
|
record = await db.get(EarthInteractable, interactable_id)
|
||||||
|
if record is None or record.is_deleted:
|
||||||
|
raise HTTPException(status_code=404, detail="Interactable not found")
|
||||||
|
return {"item": serialize_interactable(record)}
|
||||||
|
|
||||||
|
|
||||||
|
@router.patch("/{interactable_id}")
|
||||||
|
async def update_interactable(
|
||||||
|
interactable_id: str,
|
||||||
|
payload: InteractableUpdate,
|
||||||
|
_current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
record = await db.get(EarthInteractable, interactable_id)
|
||||||
|
if record is None or record.is_deleted:
|
||||||
|
raise HTTPException(status_code=404, detail="Interactable not found")
|
||||||
|
|
||||||
|
previous_layer = record.layer
|
||||||
|
patch = payload.model_dump(exclude_unset=True)
|
||||||
|
for key, value in patch.items():
|
||||||
|
setattr(record, key, value)
|
||||||
|
record.revision += 1
|
||||||
|
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(record)
|
||||||
|
invalidate_interactable_cache(previous_layer)
|
||||||
|
if record.layer != previous_layer:
|
||||||
|
invalidate_interactable_cache(record.layer)
|
||||||
|
await publish_interactable_event("updated", record)
|
||||||
|
return {"item": serialize_interactable(record)}
|
||||||
|
|
||||||
|
|
||||||
|
@router.delete("/{interactable_id}")
|
||||||
|
async def delete_interactable(
|
||||||
|
interactable_id: str,
|
||||||
|
_current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
record = await db.get(EarthInteractable, interactable_id)
|
||||||
|
if record is None or record.is_deleted:
|
||||||
|
raise HTTPException(status_code=404, detail="Interactable not found")
|
||||||
|
|
||||||
|
record.is_deleted = True
|
||||||
|
record.deleted_at = datetime.now(UTC)
|
||||||
|
record.revision += 1
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(record)
|
||||||
|
invalidate_interactable_cache(record.layer)
|
||||||
|
await publish_interactable_event("deleted", record)
|
||||||
|
event = build_interactable_event(action="deleted", record=record, include_item=True)
|
||||||
|
return {"deleted": True, "event": event}
|
||||||
231
backend/app/api/v1/layers.py
Normal file
231
backend/app/api/v1/layers.py
Normal file
@@ -0,0 +1,231 @@
|
|||||||
|
from typing import Any, Optional
|
||||||
|
|
||||||
|
from fastapi import APIRouter, Depends, HTTPException, Query, Response
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.api.v1.visualization import (
|
||||||
|
_parse_bbox,
|
||||||
|
get_bgp_anomalies_geojson,
|
||||||
|
get_bgp_collectors_geojson,
|
||||||
|
get_bgp_incidents_geojson,
|
||||||
|
get_cables_geojson,
|
||||||
|
get_landing_points_geojson,
|
||||||
|
get_satellites_geojson,
|
||||||
|
)
|
||||||
|
from app.api.v1.vessels import build_vessel_snapshot_response
|
||||||
|
from app.db.session import get_db
|
||||||
|
|
||||||
|
router = APIRouter()
|
||||||
|
|
||||||
|
DEFAULT_LAYER_LIMIT = 1000
|
||||||
|
MAX_LAYER_LIMIT = 5000
|
||||||
|
LOW_ZOOM_FEATURE_LIMIT = 500
|
||||||
|
|
||||||
|
|
||||||
|
def _clamp_limit(limit: int, zoom: int) -> tuple[int, bool]:
|
||||||
|
clamped = min(max(limit, 1), MAX_LAYER_LIMIT)
|
||||||
|
if zoom <= 3:
|
||||||
|
return min(clamped, LOW_ZOOM_FEATURE_LIMIT), clamped != limit or clamped > LOW_ZOOM_FEATURE_LIMIT
|
||||||
|
return clamped, clamped != limit
|
||||||
|
|
||||||
|
|
||||||
|
def _coordinate_in_bbox(coord: Any, bbox: tuple[float, float, float, float]) -> bool:
|
||||||
|
if not isinstance(coord, (list, tuple)) or len(coord) < 2:
|
||||||
|
return False
|
||||||
|
try:
|
||||||
|
lon = float(coord[0])
|
||||||
|
lat = float(coord[1])
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return False
|
||||||
|
lon_min, lat_min, lon_max, lat_max = bbox
|
||||||
|
return lon_min <= lon <= lon_max and lat_min <= lat <= lat_max
|
||||||
|
|
||||||
|
|
||||||
|
def _geometry_intersects_bbox(geometry: dict, bbox: tuple[float, float, float, float]) -> bool:
|
||||||
|
coordinates = geometry.get("coordinates")
|
||||||
|
geometry_type = geometry.get("type")
|
||||||
|
if geometry_type == "Point":
|
||||||
|
return _coordinate_in_bbox(coordinates, bbox)
|
||||||
|
if geometry_type in {"LineString", "MultiPoint"}:
|
||||||
|
return any(_coordinate_in_bbox(coord, bbox) for coord in coordinates or [])
|
||||||
|
if geometry_type in {"Polygon", "MultiLineString"}:
|
||||||
|
return any(
|
||||||
|
_coordinate_in_bbox(coord, bbox)
|
||||||
|
for line in coordinates or []
|
||||||
|
for coord in line
|
||||||
|
)
|
||||||
|
if geometry_type == "MultiPolygon":
|
||||||
|
return any(
|
||||||
|
_coordinate_in_bbox(coord, bbox)
|
||||||
|
for polygon in coordinates or []
|
||||||
|
for line in polygon
|
||||||
|
for coord in line
|
||||||
|
)
|
||||||
|
return False
|
||||||
|
|
||||||
|
|
||||||
|
def _guard_geojson_layer(
|
||||||
|
geojson: dict,
|
||||||
|
*,
|
||||||
|
bbox: tuple[float, float, float, float],
|
||||||
|
zoom: int,
|
||||||
|
limit: int,
|
||||||
|
) -> dict:
|
||||||
|
bounded_limit, limit_clamped = _clamp_limit(limit, zoom)
|
||||||
|
features = [
|
||||||
|
feature
|
||||||
|
for feature in geojson.get("features", [])
|
||||||
|
if _geometry_intersects_bbox(feature.get("geometry") or {}, bbox)
|
||||||
|
]
|
||||||
|
visible_count = len(features)
|
||||||
|
returned_features = features[:bounded_limit]
|
||||||
|
return {
|
||||||
|
**geojson,
|
||||||
|
"features": returned_features,
|
||||||
|
"visible_count": visible_count,
|
||||||
|
"returned_count": len(returned_features),
|
||||||
|
"diagnostics": {
|
||||||
|
"bbox_limited": True,
|
||||||
|
"limit": bounded_limit,
|
||||||
|
"limit_clamped": limit_clamped,
|
||||||
|
"truncated": visible_count > len(returned_features),
|
||||||
|
"degraded": zoom <= 3 or visible_count > len(returned_features),
|
||||||
|
"stats_scope": "viewport",
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_layer_bbox(bbox: str) -> tuple[float, float, float, float]:
|
||||||
|
parsed = _parse_bbox(bbox)
|
||||||
|
if parsed is None:
|
||||||
|
raise HTTPException(status_code=400, detail="bbox is required")
|
||||||
|
return parsed
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/vessels/snapshot")
|
||||||
|
async def get_vessel_layer_snapshot(
|
||||||
|
bbox: str = Query(..., description="lon_min,lat_min,lon_max,lat_max"),
|
||||||
|
zoom: int = Query(..., ge=1, le=20),
|
||||||
|
limit: int = Query(DEFAULT_LAYER_LIMIT, ge=1),
|
||||||
|
vessel_type: Optional[str] = Query(None, alias="type"),
|
||||||
|
since_minutes: int = Query(60, ge=1, le=1440),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
response: Response = None,
|
||||||
|
):
|
||||||
|
parsed_bbox = _parse_layer_bbox(bbox)
|
||||||
|
return await build_vessel_snapshot_response(
|
||||||
|
db,
|
||||||
|
bbox=parsed_bbox,
|
||||||
|
zoom=zoom,
|
||||||
|
limit=limit,
|
||||||
|
type_filter=vessel_type,
|
||||||
|
since_minutes=since_minutes,
|
||||||
|
response=response,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/cables")
|
||||||
|
async def get_cable_layer(
|
||||||
|
bbox: str = Query(..., description="lon_min,lat_min,lon_max,lat_max"),
|
||||||
|
zoom: int = Query(..., ge=1, le=20),
|
||||||
|
limit: int = Query(DEFAULT_LAYER_LIMIT, ge=1),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
return _guard_geojson_layer(
|
||||||
|
await get_cables_geojson(db),
|
||||||
|
bbox=_parse_layer_bbox(bbox),
|
||||||
|
zoom=zoom,
|
||||||
|
limit=limit,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/landing-points")
|
||||||
|
async def get_landing_point_layer(
|
||||||
|
bbox: str = Query(..., description="lon_min,lat_min,lon_max,lat_max"),
|
||||||
|
zoom: int = Query(..., ge=1, le=20),
|
||||||
|
limit: int = Query(DEFAULT_LAYER_LIMIT, ge=1),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
return _guard_geojson_layer(
|
||||||
|
await get_landing_points_geojson(db),
|
||||||
|
bbox=_parse_layer_bbox(bbox),
|
||||||
|
zoom=zoom,
|
||||||
|
limit=limit,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/satellites")
|
||||||
|
async def get_satellite_layer(
|
||||||
|
bbox: str = Query(..., description="lon_min,lat_min,lon_max,lat_max"),
|
||||||
|
zoom: int = Query(..., ge=1, le=20),
|
||||||
|
limit: int = Query(DEFAULT_LAYER_LIMIT, ge=1),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
bounded_limit, _ = _clamp_limit(limit, zoom)
|
||||||
|
return _guard_geojson_layer(
|
||||||
|
await get_satellites_geojson(limit=bounded_limit, db=db),
|
||||||
|
bbox=_parse_layer_bbox(bbox),
|
||||||
|
zoom=zoom,
|
||||||
|
limit=limit,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/bgp/anomalies")
|
||||||
|
async def get_bgp_anomaly_layer(
|
||||||
|
bbox: str = Query(..., description="lon_min,lat_min,lon_max,lat_max"),
|
||||||
|
zoom: int = Query(..., ge=1, le=20),
|
||||||
|
limit: int = Query(DEFAULT_LAYER_LIMIT, ge=1),
|
||||||
|
severity: Optional[str] = Query(None),
|
||||||
|
status: Optional[str] = Query("active"),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
bounded_limit, _ = _clamp_limit(limit, zoom)
|
||||||
|
return _guard_geojson_layer(
|
||||||
|
await get_bgp_anomalies_geojson(
|
||||||
|
severity=severity,
|
||||||
|
status=status,
|
||||||
|
limit=bounded_limit,
|
||||||
|
db=db,
|
||||||
|
),
|
||||||
|
bbox=_parse_layer_bbox(bbox),
|
||||||
|
zoom=zoom,
|
||||||
|
limit=limit,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/bgp/incidents")
|
||||||
|
async def get_bgp_incident_layer(
|
||||||
|
bbox: str = Query(..., description="lon_min,lat_min,lon_max,lat_max"),
|
||||||
|
zoom: int = Query(..., ge=1, le=20),
|
||||||
|
limit: int = Query(DEFAULT_LAYER_LIMIT, ge=1),
|
||||||
|
severity: Optional[str] = Query(None),
|
||||||
|
status: Optional[str] = Query("active"),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
bounded_limit, _ = _clamp_limit(limit, zoom)
|
||||||
|
return _guard_geojson_layer(
|
||||||
|
await get_bgp_incidents_geojson(
|
||||||
|
severity=severity,
|
||||||
|
status=status,
|
||||||
|
limit=min(bounded_limit, 500),
|
||||||
|
db=db,
|
||||||
|
),
|
||||||
|
bbox=_parse_layer_bbox(bbox),
|
||||||
|
zoom=zoom,
|
||||||
|
limit=limit,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/bgp/collectors")
|
||||||
|
async def get_bgp_collector_layer(
|
||||||
|
bbox: str = Query(..., description="lon_min,lat_min,lon_max,lat_max"),
|
||||||
|
zoom: int = Query(..., ge=1, le=20),
|
||||||
|
limit: int = Query(DEFAULT_LAYER_LIMIT, ge=1),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
return _guard_geojson_layer(
|
||||||
|
await get_bgp_collectors_geojson(db),
|
||||||
|
bbox=_parse_layer_bbox(bbox),
|
||||||
|
zoom=zoom,
|
||||||
|
limit=limit,
|
||||||
|
)
|
||||||
@@ -1,5 +1,7 @@
|
|||||||
from fastapi import APIRouter, Query
|
from fastapi import APIRouter, Depends, Query
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.db.session import get_db
|
||||||
from app.services.earth_news import get_earth_news_payload
|
from app.services.earth_news import get_earth_news_payload
|
||||||
|
|
||||||
router = APIRouter()
|
router = APIRouter()
|
||||||
@@ -9,5 +11,6 @@ router = APIRouter()
|
|||||||
async def get_earth_feed(
|
async def get_earth_feed(
|
||||||
lat: float | None = Query(None, description="Current Earth view center latitude"),
|
lat: float | None = Query(None, description="Current Earth view center latitude"),
|
||||||
lon: float | None = Query(None, description="Current Earth view center longitude"),
|
lon: float | None = Query(None, description="Current Earth view center longitude"),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
):
|
):
|
||||||
return await get_earth_news_payload(lat=lat, lon=lon)
|
return await get_earth_news_payload(lat=lat, lon=lon, db=db)
|
||||||
|
|||||||
280
backend/app/api/v1/realtime_sources.py
Normal file
280
backend/app/api/v1/realtime_sources.py
Normal file
@@ -0,0 +1,280 @@
|
|||||||
|
"""Realtime datasource operations and runtime statistics."""
|
||||||
|
|
||||||
|
from datetime import UTC, datetime, timedelta
|
||||||
|
import os
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from fastapi import APIRouter, Depends, HTTPException
|
||||||
|
from sqlalchemy import distinct, func, select
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.data_sources import get_data_sources_config
|
||||||
|
from app.core.security import get_current_user
|
||||||
|
from app.core.time import to_iso8601_utc
|
||||||
|
from app.db.session import get_db
|
||||||
|
from app.models.datasource import DataSource
|
||||||
|
from app.models.datasource_config import DataSourceConfig
|
||||||
|
from app.models.user import User
|
||||||
|
from app.models.vessel import AISRawObservation, AISSourceHealth
|
||||||
|
from app.services.custom_datasource_runtime import (
|
||||||
|
get_custom_stream_status,
|
||||||
|
start_custom_stream,
|
||||||
|
stop_custom_stream,
|
||||||
|
)
|
||||||
|
from app.services.scheduler import (
|
||||||
|
cancel_running_collector_now,
|
||||||
|
is_collector_running,
|
||||||
|
run_collector_now,
|
||||||
|
)
|
||||||
|
from app.services.vessel_ais_aggregation import VESSEL_AIS_SCHEMA, update_ais_source_health
|
||||||
|
|
||||||
|
router = APIRouter()
|
||||||
|
|
||||||
|
BUILTIN_REALTIME_SOURCES = {"aisstream_vessels"}
|
||||||
|
REALTIME_SOURCE_TYPES = {"websocket", "ws"}
|
||||||
|
|
||||||
|
|
||||||
|
def _is_realtime_config(config: DataSourceConfig) -> bool:
|
||||||
|
return str(config.source_type or "").lower() in REALTIME_SOURCE_TYPES
|
||||||
|
|
||||||
|
|
||||||
|
def _safe_config_dict(value: Any) -> dict[str, Any]:
|
||||||
|
return value if isinstance(value, dict) else {}
|
||||||
|
|
||||||
|
|
||||||
|
def _credential_configured(config: DataSourceConfig | None) -> bool:
|
||||||
|
if config is not None:
|
||||||
|
auth_config = _safe_config_dict(config.auth_config)
|
||||||
|
config_payload = _safe_config_dict(config.config)
|
||||||
|
if auth_config.get("api_key") or config_payload.get("api_key"):
|
||||||
|
return True
|
||||||
|
return bool(os.getenv("AISSTREAM_API_KEY"))
|
||||||
|
|
||||||
|
|
||||||
|
async def _load_realtime_stats(db: AsyncSession, source: str) -> dict[str, Any]:
|
||||||
|
now = datetime.now(UTC)
|
||||||
|
observed_24h = now - timedelta(hours=24)
|
||||||
|
observed_1h = now - timedelta(hours=1)
|
||||||
|
payload_mmsi = AISRawObservation.entity_key
|
||||||
|
|
||||||
|
result = await db.execute(
|
||||||
|
select(
|
||||||
|
func.count(AISRawObservation.id).label("total_observations"),
|
||||||
|
func.count(AISRawObservation.id)
|
||||||
|
.filter(AISRawObservation.observed_at >= observed_24h)
|
||||||
|
.label("observations_24h"),
|
||||||
|
func.count(AISRawObservation.id)
|
||||||
|
.filter(AISRawObservation.observed_at >= observed_1h)
|
||||||
|
.label("observations_1h"),
|
||||||
|
func.count(distinct(payload_mmsi)).label("unique_mmsi_total"),
|
||||||
|
func.count(distinct(payload_mmsi))
|
||||||
|
.filter(AISRawObservation.observed_at >= observed_24h)
|
||||||
|
.label("unique_mmsi_24h"),
|
||||||
|
func.max(AISRawObservation.observed_at).label("latest_observed_at"),
|
||||||
|
func.max(AISRawObservation.collected_at).label("latest_collected_at"),
|
||||||
|
)
|
||||||
|
.where(AISRawObservation.target_schema == VESSEL_AIS_SCHEMA)
|
||||||
|
.where(AISRawObservation.source == source)
|
||||||
|
)
|
||||||
|
row = result.mappings().one()
|
||||||
|
return {
|
||||||
|
"total_observations": int(row["total_observations"] or 0),
|
||||||
|
"observations_24h": int(row["observations_24h"] or 0),
|
||||||
|
"observations_1h": int(row["observations_1h"] or 0),
|
||||||
|
"unique_mmsi_total": int(row["unique_mmsi_total"] or 0),
|
||||||
|
"unique_mmsi_24h": int(row["unique_mmsi_24h"] or 0),
|
||||||
|
"latest_observed_at": to_iso8601_utc(row["latest_observed_at"]),
|
||||||
|
"latest_collected_at": to_iso8601_utc(row["latest_collected_at"]),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _runtime_status_for_builtin(source: str) -> dict[str, Any]:
|
||||||
|
running = is_collector_running(source)
|
||||||
|
return {
|
||||||
|
"running": running,
|
||||||
|
"done": False,
|
||||||
|
"runtime": "collector",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _runtime_status_for_custom(config_id: int) -> dict[str, Any]:
|
||||||
|
status = get_custom_stream_status(config_id)
|
||||||
|
return {
|
||||||
|
"running": bool(status.get("running")),
|
||||||
|
"done": bool(status.get("done")),
|
||||||
|
"runtime": "custom_stream",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def _serialize_builtin_aisstream(
|
||||||
|
db: AsyncSession,
|
||||||
|
datasource: DataSource,
|
||||||
|
config: DataSourceConfig | None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
health = await db.get(AISSourceHealth, datasource.source)
|
||||||
|
config_payload = _safe_config_dict(config.config if config else {})
|
||||||
|
endpoint = (
|
||||||
|
(config.endpoint if config else None)
|
||||||
|
or get_data_sources_config().get_yaml_url(datasource.source)
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"source": datasource.source,
|
||||||
|
"name": datasource.name,
|
||||||
|
"display_name": "AISStream 实时船舶",
|
||||||
|
"kind": "builtin",
|
||||||
|
"source_type": "websocket",
|
||||||
|
"endpoint": endpoint,
|
||||||
|
"is_active": bool(datasource.is_active),
|
||||||
|
"credential_configured": _credential_configured(config),
|
||||||
|
"message_types": config_payload.get("message_types") or ["PositionReport", "ShipStaticData"],
|
||||||
|
"bounding_boxes": config_payload.get("bounding_boxes") or [[[-90, -180], [90, 180]]],
|
||||||
|
"config": config_payload,
|
||||||
|
"runtime": _runtime_status_for_builtin(datasource.source),
|
||||||
|
"health": health.to_dict() if health else None,
|
||||||
|
"stats": await _load_realtime_stats(db, datasource.source),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def _serialize_custom_stream(
|
||||||
|
db: AsyncSession,
|
||||||
|
config: DataSourceConfig,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
health = await db.get(AISSourceHealth, config.name)
|
||||||
|
config_payload = _safe_config_dict(config.config)
|
||||||
|
return {
|
||||||
|
"source": config.name,
|
||||||
|
"name": config.name,
|
||||||
|
"display_name": config.description or config.name,
|
||||||
|
"kind": "custom",
|
||||||
|
"config_id": config.id,
|
||||||
|
"source_type": config.source_type,
|
||||||
|
"endpoint": config.endpoint,
|
||||||
|
"is_active": bool(config.is_active),
|
||||||
|
"credential_configured": config.auth_type == "none" or bool(_safe_config_dict(config.auth_config)),
|
||||||
|
"message_types": config_payload.get("message_types") or [],
|
||||||
|
"bounding_boxes": config_payload.get("bounding_boxes") or [],
|
||||||
|
"config": config_payload,
|
||||||
|
"runtime": _runtime_status_for_custom(config.id),
|
||||||
|
"health": health.to_dict() if health else None,
|
||||||
|
"stats": await _load_realtime_stats(db, config.name),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def _load_builtin_aisstream(db: AsyncSession) -> tuple[DataSource | None, DataSourceConfig | None]:
|
||||||
|
result = await db.execute(select(DataSource).where(DataSource.source == "aisstream_vessels"))
|
||||||
|
datasource = result.scalar_one_or_none()
|
||||||
|
config_result = await db.execute(
|
||||||
|
select(DataSourceConfig)
|
||||||
|
.where(DataSourceConfig.name == "aisstream_vessels")
|
||||||
|
.where(DataSourceConfig.is_active.is_(True))
|
||||||
|
.order_by(DataSourceConfig.id.desc())
|
||||||
|
.limit(1)
|
||||||
|
)
|
||||||
|
return datasource, config_result.scalar_one_or_none()
|
||||||
|
|
||||||
|
|
||||||
|
async def _load_custom_realtime_config(db: AsyncSession, source: str) -> DataSourceConfig | None:
|
||||||
|
result = await db.execute(
|
||||||
|
select(DataSourceConfig)
|
||||||
|
.where(DataSourceConfig.name == source)
|
||||||
|
.order_by(DataSourceConfig.id.desc())
|
||||||
|
.limit(1)
|
||||||
|
)
|
||||||
|
config = result.scalar_one_or_none()
|
||||||
|
return config if config is not None and _is_realtime_config(config) else None
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("")
|
||||||
|
async def list_realtime_sources(
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
sources: list[dict[str, Any]] = []
|
||||||
|
datasource, builtin_config = await _load_builtin_aisstream(db)
|
||||||
|
if datasource is not None:
|
||||||
|
sources.append(await _serialize_builtin_aisstream(db, datasource, builtin_config))
|
||||||
|
|
||||||
|
custom_result = await db.execute(
|
||||||
|
select(DataSourceConfig)
|
||||||
|
.where(func.lower(DataSourceConfig.source_type).in_(REALTIME_SOURCE_TYPES))
|
||||||
|
.order_by(DataSourceConfig.name)
|
||||||
|
)
|
||||||
|
for config in custom_result.scalars().all():
|
||||||
|
if config.name in BUILTIN_REALTIME_SOURCES:
|
||||||
|
continue
|
||||||
|
sources.append(await _serialize_custom_stream(db, config))
|
||||||
|
|
||||||
|
return {"total": len(sources), "data": sources}
|
||||||
|
|
||||||
|
|
||||||
|
async def _ensure_builtin_startable(db: AsyncSession) -> DataSourceConfig | None:
|
||||||
|
datasource, config = await _load_builtin_aisstream(db)
|
||||||
|
if datasource is None:
|
||||||
|
raise HTTPException(status_code=404, detail="Realtime source not found")
|
||||||
|
if not datasource.is_active:
|
||||||
|
raise HTTPException(status_code=400, detail="Realtime source is disabled")
|
||||||
|
if not _credential_configured(config):
|
||||||
|
raise HTTPException(status_code=400, detail="AISStream API key is not configured")
|
||||||
|
return config
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/{source}/start")
|
||||||
|
async def start_realtime_source(
|
||||||
|
source: str,
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
if source == "aisstream_vessels":
|
||||||
|
await _ensure_builtin_startable(db)
|
||||||
|
if is_collector_running(source):
|
||||||
|
return {"status": "already_running", "source": source, "runtime": _runtime_status_for_builtin(source)}
|
||||||
|
if not run_collector_now(source):
|
||||||
|
raise HTTPException(status_code=409, detail="Realtime source could not be started")
|
||||||
|
return {"status": "started", "source": source, "runtime": _runtime_status_for_builtin(source)}
|
||||||
|
|
||||||
|
config = await _load_custom_realtime_config(db, source)
|
||||||
|
if config is None:
|
||||||
|
raise HTTPException(status_code=404, detail="Realtime source not found")
|
||||||
|
if not config.is_active:
|
||||||
|
raise HTTPException(status_code=400, detail="Realtime source is disabled")
|
||||||
|
started = start_custom_stream(config.id)
|
||||||
|
return {
|
||||||
|
"status": "started" if started else "already_running",
|
||||||
|
"source": source,
|
||||||
|
"runtime": _runtime_status_for_custom(config.id),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/{source}/stop")
|
||||||
|
async def stop_realtime_source(
|
||||||
|
source: str,
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
if source == "aisstream_vessels":
|
||||||
|
stopped = await cancel_running_collector_now(source)
|
||||||
|
await update_ais_source_health(db, source=source, connection_state="disconnected", last_error=None)
|
||||||
|
await db.commit()
|
||||||
|
return {"status": "stopped" if stopped else "not_running", "source": source, "runtime": _runtime_status_for_builtin(source)}
|
||||||
|
|
||||||
|
config = await _load_custom_realtime_config(db, source)
|
||||||
|
if config is None:
|
||||||
|
raise HTTPException(status_code=404, detail="Realtime source not found")
|
||||||
|
stopped = await stop_custom_stream(config.id)
|
||||||
|
await update_ais_source_health(db, source=source, connection_state="disconnected", last_error=None)
|
||||||
|
await db.commit()
|
||||||
|
return {
|
||||||
|
"status": "stopped" if stopped else "not_running",
|
||||||
|
"source": source,
|
||||||
|
"runtime": _runtime_status_for_custom(config.id),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/{source}/restart")
|
||||||
|
async def restart_realtime_source(
|
||||||
|
source: str,
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
await stop_realtime_source(source, current_user=current_user, db=db)
|
||||||
|
return await start_realtime_source(source, current_user=current_user, db=db)
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -1,6 +1,7 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import os
|
import os
|
||||||
|
import json
|
||||||
import subprocess
|
import subprocess
|
||||||
import sys
|
import sys
|
||||||
|
|
||||||
@@ -8,9 +9,13 @@ from datetime import datetime
|
|||||||
|
|
||||||
from fastapi import APIRouter, Depends, HTTPException, Query, Request, status
|
from fastapi import APIRouter, Depends, HTTPException, Query, Request, status
|
||||||
from pydantic import BaseModel
|
from pydantic import BaseModel
|
||||||
|
from sqlalchemy import select
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
from app.core.config import ROOT_DIR
|
from app.core.config import ROOT_DIR
|
||||||
from app.core.security import get_current_user
|
from app.core.security import get_current_user
|
||||||
|
from app.db.session import get_db
|
||||||
|
from app.models.system_log import AuditLog, SystemLog
|
||||||
from app.models.user import User
|
from app.models.user import User
|
||||||
from app.services.persistent_logs import record_audit_log, record_system_log
|
from app.services.persistent_logs import record_audit_log, record_system_log
|
||||||
from app.services.system_control import (
|
from app.services.system_control import (
|
||||||
@@ -35,10 +40,47 @@ from app.services.system_logs import (
|
|||||||
normalize_log_level,
|
normalize_log_level,
|
||||||
read_log_snapshot,
|
read_log_snapshot,
|
||||||
)
|
)
|
||||||
|
from app.services.earth_layer_cache import earth_layer_cache
|
||||||
|
|
||||||
router = APIRouter()
|
router = APIRouter()
|
||||||
|
|
||||||
|
|
||||||
|
def _compact_log_context(context: dict | None) -> str:
|
||||||
|
if not context:
|
||||||
|
return ""
|
||||||
|
allowed = {
|
||||||
|
key: value
|
||||||
|
for key, value in (context or {}).items()
|
||||||
|
if key
|
||||||
|
in {
|
||||||
|
"status",
|
||||||
|
"duration_ms",
|
||||||
|
"provider",
|
||||||
|
"model",
|
||||||
|
"result_provider",
|
||||||
|
"result_model",
|
||||||
|
"collector_name",
|
||||||
|
"datasource_id",
|
||||||
|
"task_id",
|
||||||
|
"snapshot_id",
|
||||||
|
"raw_count",
|
||||||
|
"transformed_count",
|
||||||
|
"saved_count",
|
||||||
|
"created",
|
||||||
|
"updated",
|
||||||
|
"unchanged",
|
||||||
|
"deleted",
|
||||||
|
"result_count",
|
||||||
|
"status_code",
|
||||||
|
"error_type",
|
||||||
|
"error",
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if not allowed:
|
||||||
|
return ""
|
||||||
|
return json.dumps(allowed, ensure_ascii=False, sort_keys=True)
|
||||||
|
|
||||||
|
|
||||||
class RestartTaskCreate(BaseModel):
|
class RestartTaskCreate(BaseModel):
|
||||||
action: str
|
action: str
|
||||||
|
|
||||||
@@ -112,6 +154,17 @@ class EarthClientLogEventResponse(BaseModel):
|
|||||||
level: str
|
level: str
|
||||||
|
|
||||||
|
|
||||||
|
class EarthLayerCacheStatusResponse(BaseModel):
|
||||||
|
prefix: str
|
||||||
|
key_count: int
|
||||||
|
memory_bytes: int
|
||||||
|
layers: dict[str, dict[str, int]]
|
||||||
|
|
||||||
|
|
||||||
|
class EarthLayerCacheClearResponse(BaseModel):
|
||||||
|
deleted: int
|
||||||
|
|
||||||
|
|
||||||
def ensure_super_admin(current_user: User) -> None:
|
def ensure_super_admin(current_user: User) -> None:
|
||||||
if not require_super_admin(current_user.role):
|
if not require_super_admin(current_user.role):
|
||||||
raise HTTPException(
|
raise HTTPException(
|
||||||
@@ -132,6 +185,34 @@ def validate_log_date(raw_value: str | None, field_name: str) -> str | None:
|
|||||||
) from exc
|
) from exc
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/cache/earth-layers", response_model=EarthLayerCacheStatusResponse)
|
||||||
|
async def get_earth_layer_cache_status(
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
):
|
||||||
|
ensure_super_admin(current_user)
|
||||||
|
try:
|
||||||
|
return earth_layer_cache.status()
|
||||||
|
except Exception as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
|
||||||
|
detail=f"Unable to read Earth layer cache status: {exc}",
|
||||||
|
) from exc
|
||||||
|
|
||||||
|
|
||||||
|
@router.delete("/cache/earth-layers", response_model=EarthLayerCacheClearResponse)
|
||||||
|
async def clear_earth_layer_cache(
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
):
|
||||||
|
ensure_super_admin(current_user)
|
||||||
|
try:
|
||||||
|
return {"deleted": earth_layer_cache.delete_pattern()}
|
||||||
|
except Exception as exc:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
|
||||||
|
detail=f"Unable to clear Earth layer cache: {exc}",
|
||||||
|
) from exc
|
||||||
|
|
||||||
|
|
||||||
@router.post("/restart-tasks", response_model=RestartTaskResponse)
|
@router.post("/restart-tasks", response_model=RestartTaskResponse)
|
||||||
async def create_restart_task(
|
async def create_restart_task(
|
||||||
payload: RestartTaskCreate,
|
payload: RestartTaskCreate,
|
||||||
@@ -270,7 +351,123 @@ async def get_system_log_sources(
|
|||||||
current_user: User = Depends(get_current_user),
|
current_user: User = Depends(get_current_user),
|
||||||
):
|
):
|
||||||
ensure_super_admin(current_user)
|
ensure_super_admin(current_user)
|
||||||
return {"items": list_log_sources()}
|
return {
|
||||||
|
"items": [
|
||||||
|
*list_log_sources(),
|
||||||
|
{
|
||||||
|
"source_id": "system-db",
|
||||||
|
"name": "系统事件",
|
||||||
|
"kind": "database",
|
||||||
|
"location": "table://system_logs",
|
||||||
|
"description": "后端持久化系统事件、AI 和采集器操作日志。",
|
||||||
|
"category": "database",
|
||||||
|
"status": "ok",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"source_id": "audit-db",
|
||||||
|
"name": "审计事件",
|
||||||
|
"kind": "database",
|
||||||
|
"location": "table://audit_logs",
|
||||||
|
"description": "管理员敏感操作和密钥 reveal 审计记录。",
|
||||||
|
"category": "audit",
|
||||||
|
"status": "ok",
|
||||||
|
},
|
||||||
|
]
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def read_database_log_snapshot(
|
||||||
|
source_id: str,
|
||||||
|
*,
|
||||||
|
limit: int,
|
||||||
|
level: str,
|
||||||
|
levels: str | None,
|
||||||
|
start_date: str | None,
|
||||||
|
end_date: str | None,
|
||||||
|
search: str | None,
|
||||||
|
db: AsyncSession,
|
||||||
|
) -> dict | None:
|
||||||
|
selected_levels = set(normalize_log_level(item) for item in (levels or level).split(",") if item.strip())
|
||||||
|
selected_levels.discard("all")
|
||||||
|
search_query = (search or "").strip().lower()
|
||||||
|
lines: list[str] = []
|
||||||
|
|
||||||
|
if source_id == "system-db":
|
||||||
|
query = select(SystemLog).order_by(SystemLog.occurred_at.desc().nullslast(), SystemLog.id.desc()).limit(limit * 5)
|
||||||
|
result = await db.execute(query)
|
||||||
|
records = result.scalars().all()
|
||||||
|
for record in records:
|
||||||
|
record_level = normalize_log_level(record.level)
|
||||||
|
if selected_levels and record_level not in selected_levels:
|
||||||
|
continue
|
||||||
|
occurred_at = record.occurred_at.date().isoformat() if record.occurred_at else ""
|
||||||
|
if start_date and occurred_at and occurred_at < start_date:
|
||||||
|
continue
|
||||||
|
if end_date and occurred_at and occurred_at > end_date:
|
||||||
|
continue
|
||||||
|
line = " ".join(
|
||||||
|
part
|
||||||
|
for part in [
|
||||||
|
record.occurred_at.isoformat() if record.occurred_at else "",
|
||||||
|
record_level.upper(),
|
||||||
|
record.source,
|
||||||
|
record.category or "",
|
||||||
|
record.event or "",
|
||||||
|
f"request_id={record.request_id}" if record.request_id else "",
|
||||||
|
record.message,
|
||||||
|
_compact_log_context(record.context),
|
||||||
|
]
|
||||||
|
if part
|
||||||
|
)
|
||||||
|
if search_query and search_query not in line.lower():
|
||||||
|
continue
|
||||||
|
lines.append(line)
|
||||||
|
elif source_id == "audit-db":
|
||||||
|
query = select(AuditLog).order_by(AuditLog.occurred_at.desc().nullslast(), AuditLog.id.desc()).limit(limit * 5)
|
||||||
|
result = await db.execute(query)
|
||||||
|
records = result.scalars().all()
|
||||||
|
for record in records:
|
||||||
|
occurred_at = record.occurred_at.date().isoformat() if record.occurred_at else ""
|
||||||
|
if start_date and occurred_at and occurred_at < start_date:
|
||||||
|
continue
|
||||||
|
if end_date and occurred_at and occurred_at > end_date:
|
||||||
|
continue
|
||||||
|
line = " ".join(
|
||||||
|
part
|
||||||
|
for part in [
|
||||||
|
record.occurred_at.isoformat() if record.occurred_at else "",
|
||||||
|
"INFO",
|
||||||
|
record.action,
|
||||||
|
record.target_type or "",
|
||||||
|
record.target_id or "",
|
||||||
|
record.result or "",
|
||||||
|
]
|
||||||
|
if part
|
||||||
|
)
|
||||||
|
if search_query and search_query not in line.lower():
|
||||||
|
continue
|
||||||
|
lines.append(line)
|
||||||
|
else:
|
||||||
|
return None
|
||||||
|
|
||||||
|
lines = list(reversed(lines[:limit]))
|
||||||
|
return {
|
||||||
|
"source_id": source_id,
|
||||||
|
"name": "系统事件" if source_id == "system-db" else "审计事件",
|
||||||
|
"kind": "database",
|
||||||
|
"location": "table://system_logs" if source_id == "system-db" else "table://audit_logs",
|
||||||
|
"description": "数据库持久化日志",
|
||||||
|
"category": "database" if source_id == "system-db" else "audit",
|
||||||
|
"status": "ok" if lines else "empty",
|
||||||
|
"level": level,
|
||||||
|
"selected_levels": sorted(selected_levels),
|
||||||
|
"search_query": search or "",
|
||||||
|
"available_levels": ["all", "error", "warning", "info", "debug"],
|
||||||
|
"daily_markers": [],
|
||||||
|
"line_limit": limit,
|
||||||
|
"line_count": len(lines),
|
||||||
|
"lines": lines,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
@router.get("/logs/{source_id}", response_model=SystemLogSnapshotResponse)
|
@router.get("/logs/{source_id}", response_model=SystemLogSnapshotResponse)
|
||||||
@@ -283,6 +480,7 @@ async def get_system_log_snapshot(
|
|||||||
end_date: str | None = Query(None, description="Filter logs until this date (YYYY-MM-DD)"),
|
end_date: str | None = Query(None, description="Filter logs until this date (YYYY-MM-DD)"),
|
||||||
search: str | None = Query(None, description="Case-insensitive substring search"),
|
search: str | None = Query(None, description="Case-insensitive substring search"),
|
||||||
current_user: User = Depends(get_current_user),
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
):
|
):
|
||||||
ensure_super_admin(current_user)
|
ensure_super_admin(current_user)
|
||||||
|
|
||||||
@@ -305,15 +503,26 @@ async def get_system_log_snapshot(
|
|||||||
if normalized_start_date and normalized_end_date and normalized_start_date > normalized_end_date:
|
if normalized_start_date and normalized_end_date and normalized_start_date > normalized_end_date:
|
||||||
raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail="start_date must be earlier than or equal to end_date")
|
raise HTTPException(status_code=status.HTTP_400_BAD_REQUEST, detail="start_date must be earlier than or equal to end_date")
|
||||||
|
|
||||||
snapshot = read_log_snapshot(
|
snapshot = await read_database_log_snapshot(
|
||||||
source_id,
|
source_id,
|
||||||
limit,
|
limit=limit,
|
||||||
level=level,
|
level=level,
|
||||||
levels=levels,
|
levels=levels,
|
||||||
start_date=normalized_start_date,
|
start_date=normalized_start_date,
|
||||||
end_date=normalized_end_date,
|
end_date=normalized_end_date,
|
||||||
search=search,
|
search=search,
|
||||||
|
db=db,
|
||||||
)
|
)
|
||||||
|
if snapshot is None:
|
||||||
|
snapshot = read_log_snapshot(
|
||||||
|
source_id,
|
||||||
|
limit,
|
||||||
|
level=level,
|
||||||
|
levels=levels,
|
||||||
|
start_date=normalized_start_date,
|
||||||
|
end_date=normalized_end_date,
|
||||||
|
search=search,
|
||||||
|
)
|
||||||
if snapshot is None:
|
if snapshot is None:
|
||||||
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="Log source not found")
|
raise HTTPException(status_code=status.HTTP_404_NOT_FOUND, detail="Log source not found")
|
||||||
return snapshot
|
return snapshot
|
||||||
|
|||||||
@@ -27,7 +27,9 @@ async def list_tasks(
|
|||||||
offset = (page - 1) * page_size
|
offset = (page - 1) * page_size
|
||||||
query = """
|
query = """
|
||||||
SELECT ct.id, ct.datasource_id, ds.name as datasource_name, ct.status,
|
SELECT ct.id, ct.datasource_id, ds.name as datasource_name, ct.status,
|
||||||
ct.started_at, ct.completed_at, ct.records_processed, ct.error_message
|
ct.started_at, ct.completed_at, ct.records_processed, ct.error_message,
|
||||||
|
ct.phase, ct.phase_progress, ct.phase_message, ct.phase_current,
|
||||||
|
ct.phase_total, ct.phase_unit, ct.total_records, ct.progress
|
||||||
FROM collection_tasks ct
|
FROM collection_tasks ct
|
||||||
JOIN data_sources ds ON ct.datasource_id = ds.id
|
JOIN data_sources ds ON ct.datasource_id = ds.id
|
||||||
WHERE 1=1
|
WHERE 1=1
|
||||||
@@ -66,6 +68,14 @@ async def list_tasks(
|
|||||||
"completed_at": to_iso8601_utc(t[5]),
|
"completed_at": to_iso8601_utc(t[5]),
|
||||||
"records_processed": t[6],
|
"records_processed": t[6],
|
||||||
"error_message": t[7],
|
"error_message": t[7],
|
||||||
|
"phase": t[8],
|
||||||
|
"phase_progress": t[9],
|
||||||
|
"phase_message": t[10],
|
||||||
|
"phase_current": t[11],
|
||||||
|
"phase_total": t[12],
|
||||||
|
"phase_unit": t[13],
|
||||||
|
"total_records": t[14],
|
||||||
|
"progress": t[15],
|
||||||
}
|
}
|
||||||
for t in tasks
|
for t in tasks
|
||||||
],
|
],
|
||||||
@@ -81,7 +91,9 @@ async def get_task(
|
|||||||
result = await db.execute(
|
result = await db.execute(
|
||||||
text("""
|
text("""
|
||||||
SELECT ct.id, ct.datasource_id, ds.name as datasource_name, ct.status,
|
SELECT ct.id, ct.datasource_id, ds.name as datasource_name, ct.status,
|
||||||
ct.started_at, ct.completed_at, ct.records_processed, ct.error_message
|
ct.started_at, ct.completed_at, ct.records_processed, ct.error_message,
|
||||||
|
ct.phase, ct.phase_progress, ct.phase_message, ct.phase_current,
|
||||||
|
ct.phase_total, ct.phase_unit, ct.total_records, ct.progress
|
||||||
FROM collection_tasks ct
|
FROM collection_tasks ct
|
||||||
JOIN data_sources ds ON ct.datasource_id = ds.id
|
JOIN data_sources ds ON ct.datasource_id = ds.id
|
||||||
WHERE ct.id = :id
|
WHERE ct.id = :id
|
||||||
@@ -105,6 +117,14 @@ async def get_task(
|
|||||||
"completed_at": to_iso8601_utc(task[5]),
|
"completed_at": to_iso8601_utc(task[5]),
|
||||||
"records_processed": task[6],
|
"records_processed": task[6],
|
||||||
"error_message": task[7],
|
"error_message": task[7],
|
||||||
|
"phase": task[8],
|
||||||
|
"phase_progress": task[9],
|
||||||
|
"phase_message": task[10],
|
||||||
|
"phase_current": task[11],
|
||||||
|
"phase_total": task[12],
|
||||||
|
"phase_unit": task[13],
|
||||||
|
"total_records": task[14],
|
||||||
|
"progress": task[15],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -1,3 +1,4 @@
|
|||||||
|
import json
|
||||||
from typing import List
|
from typing import List
|
||||||
|
|
||||||
from fastapi import APIRouter, Depends, HTTPException, status
|
from fastapi import APIRouter, Depends, HTTPException, status
|
||||||
@@ -7,10 +8,12 @@ from sqlalchemy import text
|
|||||||
from app.core.security import get_current_user, get_password_hash
|
from app.core.security import get_current_user, get_password_hash
|
||||||
from app.db.session import get_db
|
from app.db.session import get_db
|
||||||
from app.models.user import User
|
from app.models.user import User
|
||||||
from app.schemas.user import UserCreate, UserResponse, UserUpdate
|
from app.schemas.user import UserCreate, UserUpdate
|
||||||
|
|
||||||
router = APIRouter()
|
router = APIRouter()
|
||||||
|
|
||||||
|
VALID_GATEKEEPER_GROUPS = {"docs_user", "docs_developer", "docs_admin"}
|
||||||
|
|
||||||
|
|
||||||
def check_permission(current_user: User, required_roles: List[str]) -> bool:
|
def check_permission(current_user: User, required_roles: List[str]) -> bool:
|
||||||
user_role_value = (
|
user_role_value = (
|
||||||
@@ -52,7 +55,7 @@ async def list_users(
|
|||||||
|
|
||||||
offset = (page - 1) * page_size
|
offset = (page - 1) * page_size
|
||||||
query = text(
|
query = text(
|
||||||
f"SELECT id, username, email, role, is_active, last_login_at, created_at FROM users WHERE {where_sql} ORDER BY created_at DESC LIMIT {page_size} OFFSET {offset}"
|
f"SELECT id, username, email, role, is_active, last_login_at, created_at, gatekeeper_groups FROM users WHERE {where_sql} ORDER BY created_at DESC LIMIT {page_size} OFFSET {offset}"
|
||||||
)
|
)
|
||||||
count_query = text(f"SELECT COUNT(*) FROM users WHERE {where_sql}")
|
count_query = text(f"SELECT COUNT(*) FROM users WHERE {where_sql}")
|
||||||
|
|
||||||
@@ -75,6 +78,7 @@ async def list_users(
|
|||||||
"is_active": u[4],
|
"is_active": u[4],
|
||||||
"last_login_at": u[5],
|
"last_login_at": u[5],
|
||||||
"created_at": u[6],
|
"created_at": u[6],
|
||||||
|
"gatekeeper_groups": u[7] or [],
|
||||||
}
|
}
|
||||||
for u in users
|
for u in users
|
||||||
],
|
],
|
||||||
@@ -95,7 +99,7 @@ async def get_user(
|
|||||||
|
|
||||||
result = await db.execute(
|
result = await db.execute(
|
||||||
text(
|
text(
|
||||||
"SELECT id, username, email, role, is_active, last_login_at, created_at FROM users WHERE id = :id"
|
"SELECT id, username, email, role, is_active, last_login_at, created_at, gatekeeper_groups FROM users WHERE id = :id"
|
||||||
),
|
),
|
||||||
{"id": user_id},
|
{"id": user_id},
|
||||||
)
|
)
|
||||||
@@ -114,6 +118,7 @@ async def get_user(
|
|||||||
"is_active": user[4],
|
"is_active": user[4],
|
||||||
"last_login_at": user[5],
|
"last_login_at": user[5],
|
||||||
"created_at": user[6],
|
"created_at": user[6],
|
||||||
|
"gatekeeper_groups": user[7] or [],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
@@ -128,6 +133,12 @@ async def create_user(
|
|||||||
status_code=status.HTTP_403_FORBIDDEN,
|
status_code=status.HTTP_403_FORBIDDEN,
|
||||||
detail="Only super_admin can create users",
|
detail="Only super_admin can create users",
|
||||||
)
|
)
|
||||||
|
invalid_groups = sorted(set(user_data.gatekeeper_groups) - VALID_GATEKEEPER_GROUPS)
|
||||||
|
if invalid_groups:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_400_BAD_REQUEST,
|
||||||
|
detail=f"Unsupported Gatekeeper groups: {', '.join(invalid_groups)}",
|
||||||
|
)
|
||||||
|
|
||||||
result = await db.execute(
|
result = await db.execute(
|
||||||
text("SELECT id FROM users WHERE username = :username OR email = :email"),
|
text("SELECT id FROM users WHERE username = :username OR email = :email"),
|
||||||
@@ -142,13 +153,14 @@ async def create_user(
|
|||||||
hashed_password = get_password_hash(user_data.password)
|
hashed_password = get_password_hash(user_data.password)
|
||||||
|
|
||||||
await db.execute(
|
await db.execute(
|
||||||
text("""INSERT INTO users (username, email, password_hash, role, is_active, created_at, updated_at)
|
text("""INSERT INTO users (username, email, password_hash, role, gatekeeper_groups, is_active, created_at, updated_at)
|
||||||
VALUES (:username, :email, :password_hash, :role, :is_active, NOW(), NOW())"""),
|
VALUES (:username, :email, :password_hash, :role, CAST(:gatekeeper_groups AS jsonb), :is_active, NOW(), NOW())"""),
|
||||||
{
|
{
|
||||||
"username": user_data.username,
|
"username": user_data.username,
|
||||||
"email": user_data.email,
|
"email": user_data.email,
|
||||||
"password_hash": hashed_password,
|
"password_hash": hashed_password,
|
||||||
"role": user_data.role,
|
"role": user_data.role,
|
||||||
|
"gatekeeper_groups": json.dumps(user_data.gatekeeper_groups),
|
||||||
"is_active": True,
|
"is_active": True,
|
||||||
},
|
},
|
||||||
)
|
)
|
||||||
@@ -172,6 +184,7 @@ async def create_user(
|
|||||||
"username": user_data.username,
|
"username": user_data.username,
|
||||||
"email": user_data.email,
|
"email": user_data.email,
|
||||||
"role": user_data.role,
|
"role": user_data.role,
|
||||||
|
"gatekeeper_groups": user_data.gatekeeper_groups,
|
||||||
"is_active": True,
|
"is_active": True,
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -194,6 +207,18 @@ async def update_user(
|
|||||||
status_code=status.HTTP_403_FORBIDDEN,
|
status_code=status.HTTP_403_FORBIDDEN,
|
||||||
detail="Only super_admin can change user role",
|
detail="Only super_admin can change user role",
|
||||||
)
|
)
|
||||||
|
if not check_permission(current_user, ["super_admin"]) and user_data.gatekeeper_groups is not None:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_403_FORBIDDEN,
|
||||||
|
detail="Only super_admin can change Gatekeeper groups",
|
||||||
|
)
|
||||||
|
if user_data.gatekeeper_groups is not None:
|
||||||
|
invalid_groups = sorted(set(user_data.gatekeeper_groups) - VALID_GATEKEEPER_GROUPS)
|
||||||
|
if invalid_groups:
|
||||||
|
raise HTTPException(
|
||||||
|
status_code=status.HTTP_400_BAD_REQUEST,
|
||||||
|
detail=f"Unsupported Gatekeeper groups: {', '.join(invalid_groups)}",
|
||||||
|
)
|
||||||
|
|
||||||
result = await db.execute(
|
result = await db.execute(
|
||||||
text("SELECT id FROM users WHERE id = :id"),
|
text("SELECT id FROM users WHERE id = :id"),
|
||||||
@@ -213,6 +238,9 @@ async def update_user(
|
|||||||
if user_data.role is not None:
|
if user_data.role is not None:
|
||||||
update_fields.append("role = :role")
|
update_fields.append("role = :role")
|
||||||
params["role"] = user_data.role
|
params["role"] = user_data.role
|
||||||
|
if user_data.gatekeeper_groups is not None:
|
||||||
|
update_fields.append("gatekeeper_groups = CAST(:gatekeeper_groups AS jsonb)")
|
||||||
|
params["gatekeeper_groups"] = json.dumps(user_data.gatekeeper_groups)
|
||||||
if user_data.is_active is not None:
|
if user_data.is_active is not None:
|
||||||
update_fields.append("is_active = :is_active")
|
update_fields.append("is_active = :is_active")
|
||||||
params["is_active"] = user_data.is_active
|
params["is_active"] = user_data.is_active
|
||||||
|
|||||||
132
backend/app/api/v1/vessel_aggregation.py
Normal file
132
backend/app/api/v1/vessel_aggregation.py
Normal file
@@ -0,0 +1,132 @@
|
|||||||
|
"""v4 strategy + v5 conflict-promotion + enrichment APIs for vessel_ais."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from fastapi import APIRouter, Depends, HTTPException
|
||||||
|
from sqlalchemy import select
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.security import get_current_user
|
||||||
|
from app.db.session import get_db
|
||||||
|
from app.models.user import User
|
||||||
|
from app.models.vessel import AISConflictRecord
|
||||||
|
from app.services.vessel_aggregation_strategy import (
|
||||||
|
StrategyValidationError,
|
||||||
|
load_strategy,
|
||||||
|
reset_strategy,
|
||||||
|
save_strategy,
|
||||||
|
)
|
||||||
|
from app.services.vessel_enrichment import (
|
||||||
|
get_vessel_enrichment_bundle,
|
||||||
|
upsert_vessel_media_enrichment,
|
||||||
|
upsert_vessel_profile_enrichment,
|
||||||
|
)
|
||||||
|
|
||||||
|
router = APIRouter()
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/strategy")
|
||||||
|
async def get_aggregation_strategy(db: AsyncSession = Depends(get_db)):
|
||||||
|
return await load_strategy(db)
|
||||||
|
|
||||||
|
|
||||||
|
@router.put("/strategy")
|
||||||
|
async def put_aggregation_strategy(
|
||||||
|
payload: dict[str, Any],
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
try:
|
||||||
|
return await save_strategy(db, payload)
|
||||||
|
except StrategyValidationError as exc:
|
||||||
|
raise HTTPException(status_code=400, detail=str(exc)) from exc
|
||||||
|
|
||||||
|
|
||||||
|
@router.delete("/strategy")
|
||||||
|
async def reset_aggregation_strategy(
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
return await reset_strategy(db)
|
||||||
|
|
||||||
|
|
||||||
|
@router.post("/conflicts/{mmsi}/{field}/promote-to-rule")
|
||||||
|
async def promote_conflict_to_rule(
|
||||||
|
mmsi: int,
|
||||||
|
field: str,
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
"""Lift the current conflict resolution into a persistent strategy rule."""
|
||||||
|
|
||||||
|
result = await db.execute(
|
||||||
|
select(AISConflictRecord)
|
||||||
|
.where(AISConflictRecord.target_schema == "vessel_ais")
|
||||||
|
.where(AISConflictRecord.entity_key == str(mmsi))
|
||||||
|
.where(AISConflictRecord.field == field)
|
||||||
|
.order_by(AISConflictRecord.updated_at.desc(), AISConflictRecord.id.desc())
|
||||||
|
.limit(1)
|
||||||
|
)
|
||||||
|
record = result.scalar_one_or_none()
|
||||||
|
if record is None or not record.selected_source:
|
||||||
|
raise HTTPException(status_code=404, detail="Conflict record with selected_source not found")
|
||||||
|
|
||||||
|
strategy = await load_strategy(db)
|
||||||
|
vessel_ais = dict(strategy.get("vessel_ais") or {})
|
||||||
|
field_rules = dict(vessel_ais.get("field_rules") or {})
|
||||||
|
field_rules[field] = {"mode": "source_priority", "source_priority": [record.selected_source]}
|
||||||
|
vessel_ais["field_rules"] = field_rules
|
||||||
|
|
||||||
|
incoming = {"version": int(strategy.get("version") or 0), "vessel_ais": vessel_ais}
|
||||||
|
try:
|
||||||
|
return await save_strategy(db, incoming)
|
||||||
|
except StrategyValidationError as exc:
|
||||||
|
raise HTTPException(status_code=400, detail=str(exc)) from exc
|
||||||
|
|
||||||
|
|
||||||
|
@router.delete("/conflicts/{mmsi}/{field}/promote-to-rule")
|
||||||
|
async def revert_conflict_rule(
|
||||||
|
mmsi: int,
|
||||||
|
field: str,
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
strategy = await load_strategy(db)
|
||||||
|
vessel_ais = dict(strategy.get("vessel_ais") or {})
|
||||||
|
field_rules = dict(vessel_ais.get("field_rules") or {})
|
||||||
|
if field in field_rules:
|
||||||
|
del field_rules[field]
|
||||||
|
vessel_ais["field_rules"] = field_rules
|
||||||
|
|
||||||
|
incoming = {"version": int(strategy.get("version") or 0), "vessel_ais": vessel_ais}
|
||||||
|
try:
|
||||||
|
return await save_strategy(db, incoming)
|
||||||
|
except StrategyValidationError as exc:
|
||||||
|
raise HTTPException(status_code=400, detail=str(exc)) from exc
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/enrichment/{mmsi}")
|
||||||
|
async def get_vessel_enrichment(mmsi: int, db: AsyncSession = Depends(get_db)):
|
||||||
|
return await get_vessel_enrichment_bundle(db, mmsi)
|
||||||
|
|
||||||
|
|
||||||
|
@router.put("/enrichment/{mmsi}/profile")
|
||||||
|
async def put_vessel_profile_enrichment(
|
||||||
|
mmsi: int,
|
||||||
|
payload: dict[str, Any],
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
return await upsert_vessel_profile_enrichment(db, mmsi=mmsi, payload=payload)
|
||||||
|
|
||||||
|
|
||||||
|
@router.put("/enrichment/{mmsi}/media")
|
||||||
|
async def put_vessel_media_enrichment(
|
||||||
|
mmsi: int,
|
||||||
|
payload: dict[str, Any],
|
||||||
|
current_user: User = Depends(get_current_user),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
):
|
||||||
|
return await upsert_vessel_media_enrichment(db, mmsi=mmsi, payload=payload)
|
||||||
41
backend/app/api/v1/vessels.py
Normal file
41
backend/app/api/v1/vessels.py
Normal file
@@ -0,0 +1,41 @@
|
|||||||
|
"""Bounded vessel snapshot APIs for viewport-first consumers."""
|
||||||
|
|
||||||
|
from typing import Optional
|
||||||
|
|
||||||
|
from fastapi import APIRouter, Depends, HTTPException, Query, Response
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.api.v1.visualization import _parse_bbox, build_vessel_snapshot_response
|
||||||
|
from app.db.session import get_db
|
||||||
|
from app.services.vessel_ais_aggregation import MAX_SNAPSHOT_LIMIT
|
||||||
|
|
||||||
|
router = APIRouter()
|
||||||
|
|
||||||
|
|
||||||
|
@router.get("/snapshot")
|
||||||
|
async def get_vessel_snapshot(
|
||||||
|
bbox: Optional[str] = Query(None, description="Viewport bbox as lon_min,lat_min,lon_max,lat_max"),
|
||||||
|
zoom: int = Query(..., ge=1, le=20, description="Current map zoom level"),
|
||||||
|
type: Optional[str] = Query(
|
||||||
|
None,
|
||||||
|
description="Comma-separated vessel types: cargo,tanker,passenger,fishing,military,other",
|
||||||
|
),
|
||||||
|
limit: int = Query(1000, ge=1, le=MAX_SNAPSHOT_LIMIT),
|
||||||
|
since_minutes: int = Query(60, ge=1, le=1440),
|
||||||
|
db: AsyncSession = Depends(get_db),
|
||||||
|
response: Response = None,
|
||||||
|
):
|
||||||
|
if not bbox:
|
||||||
|
raise HTTPException(status_code=400, detail="bbox is required")
|
||||||
|
parsed_bbox = _parse_bbox(bbox)
|
||||||
|
if parsed_bbox is None:
|
||||||
|
raise HTTPException(status_code=400, detail="bbox is required")
|
||||||
|
return await build_vessel_snapshot_response(
|
||||||
|
db,
|
||||||
|
bbox=parsed_bbox,
|
||||||
|
zoom=zoom,
|
||||||
|
type_filter=type,
|
||||||
|
limit=limit,
|
||||||
|
since_minutes=since_minutes,
|
||||||
|
response=response,
|
||||||
|
)
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -1,7 +1,6 @@
|
|||||||
"""WebSocket API endpoints"""
|
"""WebSocket API endpoints"""
|
||||||
|
|
||||||
import asyncio
|
import asyncio
|
||||||
import json
|
|
||||||
from datetime import UTC, datetime
|
from datetime import UTC, datetime
|
||||||
from typing import Optional
|
from typing import Optional
|
||||||
|
|
||||||
@@ -15,6 +14,7 @@ from app.core.websocket.manager import manager
|
|||||||
|
|
||||||
logger = get_logger(__name__, service="api")
|
logger = get_logger(__name__, service="api")
|
||||||
router = APIRouter()
|
router = APIRouter()
|
||||||
|
EARTH_UPDATES_CHANNEL = "earth_updates"
|
||||||
|
|
||||||
|
|
||||||
async def authenticate_token(token: str) -> Optional[dict]:
|
async def authenticate_token(token: str) -> Optional[dict]:
|
||||||
@@ -40,16 +40,16 @@ async def authenticate_token(token: str) -> Optional[dict]:
|
|||||||
@router.websocket("/ws")
|
@router.websocket("/ws")
|
||||||
async def websocket_endpoint(
|
async def websocket_endpoint(
|
||||||
websocket: WebSocket,
|
websocket: WebSocket,
|
||||||
token: str = Query(...),
|
token: str | None = Query(None),
|
||||||
):
|
):
|
||||||
"""WebSocket endpoint for real-time data"""
|
"""WebSocket endpoint for real-time data"""
|
||||||
logger.info_event(
|
logger.info_event(
|
||||||
"WebSocket connection attempt",
|
"WebSocket connection attempt",
|
||||||
event="auth.websocket.connection_attempt",
|
event="auth.websocket.connection_attempt",
|
||||||
context={"token_preview": f"{token[:8]}..."},
|
context={"token_preview": f"{token[:8]}..." if token else "anonymous"},
|
||||||
)
|
)
|
||||||
payload = await authenticate_token(token)
|
payload = await authenticate_token(token) if token else None
|
||||||
if payload is None:
|
if token and payload is None:
|
||||||
logger.warning_event(
|
logger.warning_event(
|
||||||
"WebSocket authentication failed, closing connection",
|
"WebSocket authentication failed, closing connection",
|
||||||
event="auth.websocket.connection_rejected",
|
event="auth.websocket.connection_rejected",
|
||||||
@@ -57,7 +57,19 @@ async def websocket_endpoint(
|
|||||||
await websocket.close(code=4001)
|
await websocket.close(code=4001)
|
||||||
return
|
return
|
||||||
|
|
||||||
user_id = str(payload.get("sub"))
|
is_anonymous = payload is None
|
||||||
|
user_id = str(payload.get("sub")) if payload else f"anonymous:{id(websocket)}"
|
||||||
|
supported_channels = ["vessels", "earth_news", EARTH_UPDATES_CHANNEL] if is_anonymous else [
|
||||||
|
"gpu_clusters",
|
||||||
|
"submarine_cables",
|
||||||
|
"ixp_nodes",
|
||||||
|
"alerts",
|
||||||
|
"dashboard",
|
||||||
|
"datasource_tasks",
|
||||||
|
"vessels",
|
||||||
|
"earth_news",
|
||||||
|
EARTH_UPDATES_CHANNEL,
|
||||||
|
]
|
||||||
await manager.connect(websocket, user_id)
|
await manager.connect(websocket, user_id)
|
||||||
|
|
||||||
try:
|
try:
|
||||||
@@ -68,14 +80,7 @@ async def websocket_endpoint(
|
|||||||
"connection_id": f"conn_{user_id}",
|
"connection_id": f"conn_{user_id}",
|
||||||
"server_version": settings.VERSION,
|
"server_version": settings.VERSION,
|
||||||
"heartbeat_interval": 30,
|
"heartbeat_interval": 30,
|
||||||
"supported_channels": [
|
"supported_channels": supported_channels,
|
||||||
"gpu_clusters",
|
|
||||||
"submarine_cables",
|
|
||||||
"ixp_nodes",
|
|
||||||
"alerts",
|
|
||||||
"dashboard",
|
|
||||||
"datasource_tasks",
|
|
||||||
],
|
|
||||||
},
|
},
|
||||||
}
|
}
|
||||||
)
|
)
|
||||||
@@ -92,11 +97,53 @@ async def websocket_endpoint(
|
|||||||
}
|
}
|
||||||
)
|
)
|
||||||
elif data.get("type") == "subscribe":
|
elif data.get("type") == "subscribe":
|
||||||
channels = data.get("data", {}).get("channels", [])
|
payload_data = data.get("data", {})
|
||||||
|
if not isinstance(payload_data, dict):
|
||||||
|
payload_data = {}
|
||||||
|
channels = payload_data.get("channels", [])
|
||||||
|
if isinstance(channels, str):
|
||||||
|
channels = [channels]
|
||||||
|
elif not isinstance(channels, list):
|
||||||
|
channels = []
|
||||||
|
channel = payload_data.get("channel")
|
||||||
|
if channel and channel not in channels:
|
||||||
|
channels = [*channels, channel]
|
||||||
|
if is_anonymous:
|
||||||
|
channels = [channel for channel in channels if channel in supported_channels]
|
||||||
|
vessel_subscription = None
|
||||||
|
if "vessels" in channels and "bbox" in payload_data:
|
||||||
|
try:
|
||||||
|
vessel_subscription = manager.subscribe_vessels(websocket, payload_data)
|
||||||
|
except ValueError as exc:
|
||||||
|
await websocket.send_json(
|
||||||
|
{
|
||||||
|
"type": "subscription_error",
|
||||||
|
"data": {"channel": "vessels", "detail": str(exc)},
|
||||||
|
}
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
channels = [channel for channel in channels if channel != "vessels"]
|
||||||
|
manager.subscribe(websocket, channels)
|
||||||
await websocket.send_json(
|
await websocket.send_json(
|
||||||
{
|
{
|
||||||
"type": "subscription_confirmed",
|
"type": "subscription_confirmed",
|
||||||
"data": {"action": "subscribe", "channels": channels},
|
"data": {
|
||||||
|
"action": "subscribe",
|
||||||
|
"channels": [
|
||||||
|
*channels,
|
||||||
|
*(["vessels"] if vessel_subscription else []),
|
||||||
|
],
|
||||||
|
"vessels": vessel_subscription,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
)
|
||||||
|
elif data.get("type") == "unsubscribe":
|
||||||
|
channels = data.get("data", {}).get("channels", [])
|
||||||
|
manager.unsubscribe(websocket, channels)
|
||||||
|
await websocket.send_json(
|
||||||
|
{
|
||||||
|
"type": "subscription_confirmed",
|
||||||
|
"data": {"action": "unsubscribe", "channels": channels},
|
||||||
}
|
}
|
||||||
)
|
)
|
||||||
elif data.get("type") == "control_frame":
|
elif data.get("type") == "control_frame":
|
||||||
|
|||||||
@@ -4,8 +4,8 @@ from typing import Any, Dict, Optional
|
|||||||
FIELD_ALIASES = {
|
FIELD_ALIASES = {
|
||||||
"country": ("country",),
|
"country": ("country",),
|
||||||
"city": ("city",),
|
"city": ("city",),
|
||||||
"latitude": ("latitude",),
|
"latitude": ("latitude", "lat"),
|
||||||
"longitude": ("longitude",),
|
"longitude": ("longitude", "lon", "lng"),
|
||||||
"value": ("value",),
|
"value": ("value",),
|
||||||
"unit": ("unit",),
|
"unit": ("unit",),
|
||||||
"cores": ("cores",),
|
"cores": ("cores",),
|
||||||
@@ -14,6 +14,28 @@ FIELD_ALIASES = {
|
|||||||
"power": ("power",),
|
"power": ("power",),
|
||||||
}
|
}
|
||||||
|
|
||||||
|
NESTED_FIELD_ALIASES = {
|
||||||
|
"latitude": (
|
||||||
|
("location", "latitude"),
|
||||||
|
("location", "lat"),
|
||||||
|
("geo", "latitude"),
|
||||||
|
("geo", "lat"),
|
||||||
|
("coordinates", "latitude"),
|
||||||
|
("coordinates", "lat"),
|
||||||
|
),
|
||||||
|
"longitude": (
|
||||||
|
("location", "longitude"),
|
||||||
|
("location", "lon"),
|
||||||
|
("location", "lng"),
|
||||||
|
("geo", "longitude"),
|
||||||
|
("geo", "lon"),
|
||||||
|
("geo", "lng"),
|
||||||
|
("coordinates", "longitude"),
|
||||||
|
("coordinates", "lon"),
|
||||||
|
("coordinates", "lng"),
|
||||||
|
),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
def get_metadata_field(metadata: Optional[Dict[str, Any]], field: str, fallback: Any = None) -> Any:
|
def get_metadata_field(metadata: Optional[Dict[str, Any]], field: str, fallback: Any = None) -> Any:
|
||||||
if isinstance(metadata, dict):
|
if isinstance(metadata, dict):
|
||||||
@@ -21,9 +43,34 @@ def get_metadata_field(metadata: Optional[Dict[str, Any]], field: str, fallback:
|
|||||||
value = metadata.get(key)
|
value = metadata.get(key)
|
||||||
if value not in (None, ""):
|
if value not in (None, ""):
|
||||||
return value
|
return value
|
||||||
|
for path in NESTED_FIELD_ALIASES.get(field, ()):
|
||||||
|
current: Any = metadata
|
||||||
|
for key in path:
|
||||||
|
if not isinstance(current, dict):
|
||||||
|
current = None
|
||||||
|
break
|
||||||
|
current = current.get(key)
|
||||||
|
if current not in (None, ""):
|
||||||
|
return current
|
||||||
|
if field in {"latitude", "longitude"}:
|
||||||
|
value = _get_coordinate_sequence_value(metadata, field)
|
||||||
|
if value not in (None, ""):
|
||||||
|
return value
|
||||||
return fallback
|
return fallback
|
||||||
|
|
||||||
|
|
||||||
|
def _get_coordinate_sequence_value(metadata: Dict[str, Any], field: str) -> Any:
|
||||||
|
for key in ("coordinates", "coord", "coords"):
|
||||||
|
value = metadata.get(key)
|
||||||
|
if not isinstance(value, (list, tuple)) or len(value) < 2:
|
||||||
|
continue
|
||||||
|
# GeoJSON uses [longitude, latitude]. Most raw collector tuples in this
|
||||||
|
# codebase use explicit field names, so only sequence aliases are treated
|
||||||
|
# as GeoJSON-shaped to avoid guessing.
|
||||||
|
return value[1] if field == "latitude" else value[0]
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
def build_dynamic_metadata(
|
def build_dynamic_metadata(
|
||||||
metadata: Optional[Dict[str, Any]],
|
metadata: Optional[Dict[str, Any]],
|
||||||
*,
|
*,
|
||||||
|
|||||||
@@ -1,7 +1,6 @@
|
|||||||
import os
|
import os
|
||||||
import yaml
|
import yaml
|
||||||
from functools import lru_cache
|
from functools import lru_cache
|
||||||
from typing import Optional
|
|
||||||
|
|
||||||
|
|
||||||
COLLECTOR_URL_KEYS = {
|
COLLECTOR_URL_KEYS = {
|
||||||
@@ -31,6 +30,8 @@ COLLECTOR_URL_KEYS = {
|
|||||||
"opengeofeed_prefix_geo": "opengeofeed.public_csv_url",
|
"opengeofeed_prefix_geo": "opengeofeed.public_csv_url",
|
||||||
"nro_delegated_prefix_geo": "nro.delegated_stats_url",
|
"nro_delegated_prefix_geo": "nro.delegated_stats_url",
|
||||||
"news_live_streams": "news_live_streams.channels_url",
|
"news_live_streams": "news_live_streams.channels_url",
|
||||||
|
"barentswatch_vessels": "barentswatch_vessels.url",
|
||||||
|
"aisstream_vessels": "aisstream_vessels.url",
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
@@ -73,7 +74,7 @@ class DataSourcesConfig:
|
|||||||
from app.models.datasource_config import DataSourceConfig
|
from app.models.datasource_config import DataSourceConfig
|
||||||
|
|
||||||
query = select(DataSourceConfig).where(
|
query = select(DataSourceConfig).where(
|
||||||
DataSourceConfig.name == collector_name, DataSourceConfig.is_active == True
|
DataSourceConfig.name == collector_name, DataSourceConfig.is_active
|
||||||
)
|
)
|
||||||
result = await db.execute(query)
|
result = await db.execute(query)
|
||||||
db_config = result.scalar_one_or_none()
|
db_config = result.scalar_one_or_none()
|
||||||
|
|||||||
@@ -94,3 +94,11 @@ news_live_streams:
|
|||||||
streams_url: "https://iptv-org.github.io/api/streams.json"
|
streams_url: "https://iptv-org.github.io/api/streams.json"
|
||||||
# IPTV-org 台标 JSON
|
# IPTV-org 台标 JSON
|
||||||
logos_url: "https://iptv-org.github.io/api/logos.json"
|
logos_url: "https://iptv-org.github.io/api/logos.json"
|
||||||
|
|
||||||
|
barentswatch_vessels:
|
||||||
|
# BarentsWatch Live AIS latest combined endpoint. Requires an AIS bearer token.
|
||||||
|
url: "https://live.ais.barentswatch.no/v1/latest/combined"
|
||||||
|
|
||||||
|
aisstream_vessels:
|
||||||
|
# AISStream realtime WebSocket endpoint. Requires an AISStream API key.
|
||||||
|
url: "wss://stream.aisstream.io/v0/stream"
|
||||||
|
|||||||
@@ -4,163 +4,268 @@ DEFAULT_DATASOURCES = {
|
|||||||
"top500": {
|
"top500": {
|
||||||
"id": 1,
|
"id": 1,
|
||||||
"name": "TOP500 Supercomputers",
|
"name": "TOP500 Supercomputers",
|
||||||
|
"display_name": "TOP500 超算榜单",
|
||||||
"module": "L1",
|
"module": "L1",
|
||||||
"priority": "P0",
|
"priority": "P0",
|
||||||
"frequency_minutes": 240,
|
"frequency_minutes": 240,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"epoch_ai_gpu": {
|
"epoch_ai_gpu": {
|
||||||
"id": 2,
|
"id": 2,
|
||||||
"name": "Epoch AI GPU Clusters",
|
"name": "Epoch AI GPU Clusters",
|
||||||
|
"display_name": "Epoch AI GPU 集群",
|
||||||
"module": "L1",
|
"module": "L1",
|
||||||
"priority": "P0",
|
"priority": "P0",
|
||||||
"frequency_minutes": 360,
|
"frequency_minutes": 360,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"huggingface_models": {
|
"huggingface_models": {
|
||||||
"id": 3,
|
"id": 3,
|
||||||
"name": "HuggingFace Models",
|
"name": "HuggingFace Models",
|
||||||
|
"display_name": "Hugging Face 模型",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 720,
|
"frequency_minutes": 720,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"huggingface_datasets": {
|
"huggingface_datasets": {
|
||||||
"id": 4,
|
"id": 4,
|
||||||
"name": "HuggingFace Datasets",
|
"name": "HuggingFace Datasets",
|
||||||
|
"display_name": "Hugging Face 数据集",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 720,
|
"frequency_minutes": 720,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"huggingface_spaces": {
|
"huggingface_spaces": {
|
||||||
"id": 5,
|
"id": 5,
|
||||||
"name": "HuggingFace Spaces",
|
"name": "HuggingFace Spaces",
|
||||||
|
"display_name": "Hugging Face Spaces",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P2",
|
"priority": "P2",
|
||||||
"frequency_minutes": 1440,
|
"frequency_minutes": 1440,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"peeringdb_ixp": {
|
"peeringdb_ixp": {
|
||||||
"id": 6,
|
"id": 6,
|
||||||
"name": "PeeringDB IXP",
|
"name": "PeeringDB IXP",
|
||||||
|
"display_name": "PeeringDB 交换中心",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 1440,
|
"frequency_minutes": 1440,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"peeringdb_network": {
|
"peeringdb_network": {
|
||||||
"id": 7,
|
"id": 7,
|
||||||
"name": "PeeringDB Networks",
|
"name": "PeeringDB Networks",
|
||||||
|
"display_name": "PeeringDB 网络",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P2",
|
"priority": "P2",
|
||||||
"frequency_minutes": 2880,
|
"frequency_minutes": 2880,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"peeringdb_facility": {
|
"peeringdb_facility": {
|
||||||
"id": 8,
|
"id": 8,
|
||||||
"name": "PeeringDB Facilities",
|
"name": "PeeringDB Facilities",
|
||||||
|
"display_name": "PeeringDB 设施",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P2",
|
"priority": "P2",
|
||||||
"frequency_minutes": 2880,
|
"frequency_minutes": 2880,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"telegeography_cables": {
|
"telegeography_cables": {
|
||||||
"id": 9,
|
"id": 9,
|
||||||
"name": "Submarine Cables",
|
"name": "Submarine Cables",
|
||||||
|
"display_name": "海底光缆",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 10080,
|
"frequency_minutes": 10080,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"telegeography_landing": {
|
"telegeography_landing": {
|
||||||
"id": 10,
|
"id": 10,
|
||||||
"name": "Cable Landing Points",
|
"name": "Cable Landing Points",
|
||||||
|
"display_name": "光缆登陆点",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P2",
|
"priority": "P2",
|
||||||
"frequency_minutes": 10080,
|
"frequency_minutes": 10080,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"telegeography_systems": {
|
"telegeography_systems": {
|
||||||
"id": 11,
|
"id": 11,
|
||||||
"name": "Cable Systems",
|
"name": "Cable Systems",
|
||||||
|
"display_name": "光缆系统",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P2",
|
"priority": "P2",
|
||||||
"frequency_minutes": 10080,
|
"frequency_minutes": 10080,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"arcgis_cables": {
|
"arcgis_cables": {
|
||||||
"id": 15,
|
"id": 15,
|
||||||
"name": "ArcGIS Submarine Cables",
|
"name": "ArcGIS Submarine Cables",
|
||||||
|
"display_name": "ArcGIS 海底光缆",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 10080,
|
"frequency_minutes": 10080,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"arcgis_landing_points": {
|
"arcgis_landing_points": {
|
||||||
"id": 16,
|
"id": 16,
|
||||||
"name": "ArcGIS Landing Points",
|
"name": "ArcGIS Landing Points",
|
||||||
|
"display_name": "ArcGIS 登陆点",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 10080,
|
"frequency_minutes": 10080,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"arcgis_cable_landing_relation": {
|
"arcgis_cable_landing_relation": {
|
||||||
"id": 17,
|
"id": 17,
|
||||||
"name": "ArcGIS Cable-Landing Relations",
|
"name": "ArcGIS Cable-Landing Relations",
|
||||||
|
"display_name": "ArcGIS 光缆登陆关系",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 10080,
|
"frequency_minutes": 10080,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"fao_landing_points": {
|
"fao_landing_points": {
|
||||||
"id": 18,
|
"id": 18,
|
||||||
"name": "FAO Landing Points",
|
"name": "FAO Landing Points",
|
||||||
|
"display_name": "FAO 登陆点",
|
||||||
"module": "L2",
|
"module": "L2",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 10080,
|
"frequency_minutes": 10080,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"spacetrack_tle": {
|
"spacetrack_tle": {
|
||||||
"id": 19,
|
"id": 19,
|
||||||
"name": "Space-Track TLE",
|
"name": "Space-Track TLE",
|
||||||
|
"display_name": "Space-Track 轨道根数",
|
||||||
"module": "L3",
|
"module": "L3",
|
||||||
"priority": "P2",
|
"priority": "P2",
|
||||||
"frequency_minutes": 1440,
|
"frequency_minutes": 1440,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": True,
|
||||||
|
"credential_provider": "spacetrack",
|
||||||
|
"credential_status": "planned",
|
||||||
},
|
},
|
||||||
"celestrak_tle": {
|
"celestrak_tle": {
|
||||||
"id": 20,
|
"id": 20,
|
||||||
"name": "CelesTrak TLE",
|
"name": "CelesTrak TLE",
|
||||||
|
"display_name": "CelesTrak 轨道根数",
|
||||||
"module": "L3",
|
"module": "L3",
|
||||||
"priority": "P2",
|
"priority": "P2",
|
||||||
"frequency_minutes": 1440,
|
"frequency_minutes": 1440,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"ris_live_bgp": {
|
"ris_live_bgp": {
|
||||||
"id": 21,
|
"id": 21,
|
||||||
"name": "RIPE RIS Live BGP",
|
"name": "RIPE RIS Live BGP",
|
||||||
|
"display_name": "RIPE RIS 实时 BGP",
|
||||||
"module": "L3",
|
"module": "L3",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 15,
|
"frequency_minutes": 15,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"bgpstream_bgp": {
|
"bgpstream_bgp": {
|
||||||
"id": 22,
|
"id": 22,
|
||||||
"name": "CAIDA BGPStream Backfill",
|
"name": "CAIDA BGPStream Backfill",
|
||||||
|
"display_name": "CAIDA BGPStream 回填",
|
||||||
"module": "L3",
|
"module": "L3",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 360,
|
"frequency_minutes": 360,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"iptoasn_prefix_geo": {
|
"iptoasn_prefix_geo": {
|
||||||
"id": 23,
|
"id": 23,
|
||||||
"name": "IPtoASN Prefix Geography",
|
"name": "IPtoASN Prefix Geography",
|
||||||
|
"display_name": "IPtoASN 前缀地理",
|
||||||
"module": "L3",
|
"module": "L3",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 1440,
|
"frequency_minutes": 1440,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"opengeofeed_prefix_geo": {
|
"opengeofeed_prefix_geo": {
|
||||||
"id": 24,
|
"id": 24,
|
||||||
"name": "OpenGeoFeed Prefix Geography",
|
"name": "OpenGeoFeed Prefix Geography",
|
||||||
|
"display_name": "OpenGeoFeed 前缀地理",
|
||||||
"module": "L3",
|
"module": "L3",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 1440,
|
"frequency_minutes": 1440,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"nro_delegated_prefix_geo": {
|
"nro_delegated_prefix_geo": {
|
||||||
"id": 25,
|
"id": 25,
|
||||||
"name": "NRO Delegated Prefix Geography",
|
"name": "NRO Delegated Prefix Geography",
|
||||||
|
"display_name": "NRO 分配前缀地理",
|
||||||
"module": "L3",
|
"module": "L3",
|
||||||
"priority": "P1",
|
"priority": "P1",
|
||||||
"frequency_minutes": 1440,
|
"frequency_minutes": 1440,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
"news_live_streams": {
|
"news_live_streams": {
|
||||||
"id": 26,
|
"id": 26,
|
||||||
"name": "News Live Streams",
|
"name": "News Live Streams",
|
||||||
|
"display_name": "新闻直播源",
|
||||||
"module": "L4",
|
"module": "L4",
|
||||||
"priority": "P2",
|
"priority": "P2",
|
||||||
"frequency_minutes": 720,
|
"frequency_minutes": 720,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
|
},
|
||||||
|
"barentswatch_vessels": {
|
||||||
|
"id": 27,
|
||||||
|
"name": "BarentsWatch AIS Vessels",
|
||||||
|
"display_name": "BarentsWatch AIS 船舶",
|
||||||
|
"module": "L4",
|
||||||
|
"priority": "P1",
|
||||||
|
"frequency_minutes": 1,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": True,
|
||||||
|
"credential_provider": "barentswatch",
|
||||||
|
"credential_status": "supported",
|
||||||
|
},
|
||||||
|
"aisstream_vessels": {
|
||||||
|
"id": 28,
|
||||||
|
"name": "AISStream Vessels",
|
||||||
|
"display_name": "AISStream 实时船舶",
|
||||||
|
"module": "L4",
|
||||||
|
"priority": "P1",
|
||||||
|
"frequency_minutes": 1,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": True,
|
||||||
|
"credential_provider": "aisstream",
|
||||||
|
"credential_status": "supported",
|
||||||
|
},
|
||||||
|
"media_news_archive": {
|
||||||
|
"id": 33,
|
||||||
|
"name": "Media News Archive",
|
||||||
|
"display_name": "媒体新闻归档",
|
||||||
|
"module": "L4",
|
||||||
|
"priority": "P2",
|
||||||
|
"frequency_minutes": 720,
|
||||||
|
"is_free": True,
|
||||||
|
"requires_credentials": False,
|
||||||
},
|
},
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -105,7 +105,7 @@ async def get_current_user(
|
|||||||
)
|
)
|
||||||
result = await db.execute(
|
result = await db.execute(
|
||||||
text(
|
text(
|
||||||
"SELECT id, username, email, password_hash, role, is_active FROM users WHERE id = :id"
|
"SELECT id, username, email, password_hash, role, is_active, gatekeeper_groups FROM users WHERE id = :id"
|
||||||
),
|
),
|
||||||
{"id": int(user_id)},
|
{"id": int(user_id)},
|
||||||
)
|
)
|
||||||
@@ -122,6 +122,7 @@ async def get_current_user(
|
|||||||
user.password_hash = row[3]
|
user.password_hash = row[3]
|
||||||
user.role = row[4]
|
user.role = row[4]
|
||||||
user.is_active = row[5]
|
user.is_active = row[5]
|
||||||
|
user.gatekeeper_groups = row[6] or []
|
||||||
return user
|
return user
|
||||||
|
|
||||||
|
|
||||||
@@ -144,7 +145,7 @@ async def get_current_user_refresh(
|
|||||||
)
|
)
|
||||||
result = await db.execute(
|
result = await db.execute(
|
||||||
text(
|
text(
|
||||||
"SELECT id, username, email, password_hash, role, is_active FROM users WHERE id = :id"
|
"SELECT id, username, email, password_hash, role, is_active, gatekeeper_groups FROM users WHERE id = :id"
|
||||||
),
|
),
|
||||||
{"id": int(user_id)},
|
{"id": int(user_id)},
|
||||||
)
|
)
|
||||||
@@ -161,6 +162,7 @@ async def get_current_user_refresh(
|
|||||||
user.password_hash = row[3]
|
user.password_hash = row[3]
|
||||||
user.role = row[4]
|
user.role = row[4]
|
||||||
user.is_active = row[5]
|
user.is_active = row[5]
|
||||||
|
user.gatekeeper_groups = row[6] or []
|
||||||
return user
|
return user
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
157
backend/app/core/target_schema_registry.py
Normal file
157
backend/app/core/target_schema_registry.py
Normal file
@@ -0,0 +1,157 @@
|
|||||||
|
"""Registry of target schemas supported by mapped custom data sources."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from datetime import datetime
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from pydantic import BaseModel, Field, ValidationError, field_validator
|
||||||
|
|
||||||
|
|
||||||
|
class VesselAISRecord(BaseModel):
|
||||||
|
mmsi: int = Field(ge=100000000, le=999999999)
|
||||||
|
lat: float = Field(ge=-90, le=90)
|
||||||
|
lon: float = Field(ge=-180, le=180)
|
||||||
|
sog: float | None = None
|
||||||
|
cog: float | None = Field(default=None, ge=0, le=360)
|
||||||
|
heading: int | None = Field(default=None, ge=0, le=511)
|
||||||
|
nav_status: int | None = None
|
||||||
|
name: str | None = None
|
||||||
|
callsign: str | None = None
|
||||||
|
vessel_type: str | int | None = None
|
||||||
|
vessel_type_name: str | None = None
|
||||||
|
received_at: datetime | None = None
|
||||||
|
|
||||||
|
|
||||||
|
class GeoPointRecord(BaseModel):
|
||||||
|
lat: float = Field(ge=-90, le=90)
|
||||||
|
lon: float = Field(ge=-180, le=180)
|
||||||
|
name: str | None = None
|
||||||
|
type: str | None = None
|
||||||
|
source_id: str | None = None
|
||||||
|
observed_at: datetime | None = None
|
||||||
|
metadata: dict[str, Any] = Field(default_factory=dict)
|
||||||
|
|
||||||
|
|
||||||
|
class GenericRecord(BaseModel):
|
||||||
|
data: dict[str, Any] = Field(default_factory=dict)
|
||||||
|
source_id: str | None = None
|
||||||
|
observed_at: datetime | None = None
|
||||||
|
|
||||||
|
@field_validator("data")
|
||||||
|
@classmethod
|
||||||
|
def require_payload(cls, value: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
if not value:
|
||||||
|
raise ValueError("generic_records requires a non-empty data object")
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class TargetField:
|
||||||
|
name: str
|
||||||
|
type: str
|
||||||
|
required: bool = False
|
||||||
|
description: str = ""
|
||||||
|
example: Any = None
|
||||||
|
|
||||||
|
def to_dict(self) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"name": self.name,
|
||||||
|
"type": self.type,
|
||||||
|
"required": self.required,
|
||||||
|
"description": self.description,
|
||||||
|
"example": self.example,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class TargetSchema:
|
||||||
|
key: str
|
||||||
|
label: str
|
||||||
|
description: str
|
||||||
|
fields: tuple[TargetField, ...]
|
||||||
|
model: type[BaseModel]
|
||||||
|
destination: str
|
||||||
|
|
||||||
|
def to_dict(self) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"key": self.key,
|
||||||
|
"label": self.label,
|
||||||
|
"description": self.description,
|
||||||
|
"destination": self.destination,
|
||||||
|
"fields": [field.to_dict() for field in self.fields],
|
||||||
|
}
|
||||||
|
|
||||||
|
def validate_record(self, record: dict[str, Any]) -> tuple[dict[str, Any] | None, list[str]]:
|
||||||
|
try:
|
||||||
|
return self.model.model_validate(record).model_dump(mode="json"), []
|
||||||
|
except ValidationError as exc:
|
||||||
|
return None, [
|
||||||
|
".".join(str(part) for part in error["loc"]) + f": {error['msg']}"
|
||||||
|
for error in exc.errors()
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
TARGET_SCHEMAS: dict[str, TargetSchema] = {
|
||||||
|
"vessel_ais": TargetSchema(
|
||||||
|
key="vessel_ais",
|
||||||
|
label="船舶 AIS",
|
||||||
|
description="船只位置、航速、航向、MMSI 等 AIS 数据。",
|
||||||
|
destination="vessel_position",
|
||||||
|
model=VesselAISRecord,
|
||||||
|
fields=(
|
||||||
|
TargetField("mmsi", "integer", True, "MMSI 九位船舶标识", 257123000),
|
||||||
|
TargetField("lat", "float", True, "纬度", 59.91),
|
||||||
|
TargetField("lon", "float", True, "经度", 10.75),
|
||||||
|
TargetField("sog", "float", False, "对地航速,单位节", 12.4),
|
||||||
|
TargetField("cog", "float", False, "对地航向,0-360 度", 184.5),
|
||||||
|
TargetField("heading", "integer", False, "船首向,0-511", 186),
|
||||||
|
TargetField("nav_status", "integer", False, "导航状态码", 0),
|
||||||
|
TargetField("name", "string", False, "船名", "OSLO EXPRESS"),
|
||||||
|
TargetField("callsign", "string", False, "呼号", "LAAB"),
|
||||||
|
TargetField("vessel_type", "string", False, "船型代码", 70),
|
||||||
|
TargetField("vessel_type_name", "string", False, "船型名称", "Cargo"),
|
||||||
|
TargetField("received_at", "datetime", False, "数据接收时间", "2026-04-28T00:00:00Z"),
|
||||||
|
),
|
||||||
|
),
|
||||||
|
"geo_points": TargetSchema(
|
||||||
|
key="geo_points",
|
||||||
|
label="通用地理点",
|
||||||
|
description="带经纬度的通用实体或事件点位。",
|
||||||
|
destination="generic_geo_points",
|
||||||
|
model=GeoPointRecord,
|
||||||
|
fields=(
|
||||||
|
TargetField("lat", "float", True, "纬度", 1.3),
|
||||||
|
TargetField("lon", "float", True, "经度", 103.8),
|
||||||
|
TargetField("name", "string", False, "点位名称", "Singapore"),
|
||||||
|
TargetField("type", "string", False, "点位类型", "datacenter"),
|
||||||
|
TargetField("source_id", "string", False, "来源侧 ID", "sg-1"),
|
||||||
|
TargetField("observed_at", "datetime", False, "观测时间", "2026-04-28T00:00:00Z"),
|
||||||
|
TargetField("metadata", "object", False, "扩展字段", {"provider": "example"}),
|
||||||
|
),
|
||||||
|
),
|
||||||
|
"generic_records": TargetSchema(
|
||||||
|
key="generic_records",
|
||||||
|
label="通用结构化记录",
|
||||||
|
description="未知结构数据沉淀,不直接进入 Earth 图层。",
|
||||||
|
destination="collected_data",
|
||||||
|
model=GenericRecord,
|
||||||
|
fields=(
|
||||||
|
TargetField("data", "object", True, "结构化记录主体", {"raw": "value"}),
|
||||||
|
TargetField("source_id", "string", False, "来源侧 ID", "record-1"),
|
||||||
|
TargetField("observed_at", "datetime", False, "观测时间", "2026-04-28T00:00:00Z"),
|
||||||
|
),
|
||||||
|
),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def list_target_schemas() -> list[dict[str, Any]]:
|
||||||
|
return [schema.to_dict() for schema in TARGET_SCHEMAS.values()]
|
||||||
|
|
||||||
|
|
||||||
|
def get_target_schema(key: str) -> TargetSchema:
|
||||||
|
try:
|
||||||
|
return TARGET_SCHEMAS[key]
|
||||||
|
except KeyError as exc:
|
||||||
|
raise ValueError(f"Unsupported target schema: {key}") from exc
|
||||||
@@ -2,12 +2,14 @@
|
|||||||
|
|
||||||
import asyncio
|
import asyncio
|
||||||
from datetime import UTC, datetime
|
from datetime import UTC, datetime
|
||||||
from typing import Dict, Any, Optional
|
from typing import Dict, Any
|
||||||
|
|
||||||
from app.core.time import to_iso8601_utc
|
from app.core.time import to_iso8601_utc
|
||||||
from app.core.websocket.manager import manager
|
from app.core.websocket.manager import manager
|
||||||
|
|
||||||
|
|
||||||
|
EARTH_UPDATES_CHANNEL = "earth_updates"
|
||||||
|
|
||||||
|
|
||||||
class DataBroadcaster:
|
class DataBroadcaster:
|
||||||
"""Periodically broadcasts data to connected WebSocket clients"""
|
"""Periodically broadcasts data to connected WebSocket clients"""
|
||||||
@@ -15,6 +17,8 @@ class DataBroadcaster:
|
|||||||
def __init__(self):
|
def __init__(self):
|
||||||
self.running = False
|
self.running = False
|
||||||
self.tasks: Dict[str, asyncio.Task] = {}
|
self.tasks: Dict[str, asyncio.Task] = {}
|
||||||
|
self._pending_vessel_updates: Dict[str, Dict[str, Any]] = {}
|
||||||
|
self._vessel_flush_interval = 1.0
|
||||||
|
|
||||||
async def get_dashboard_stats(self) -> Dict[str, Any]:
|
async def get_dashboard_stats(self) -> Dict[str, Any]:
|
||||||
"""Get dashboard statistics"""
|
"""Get dashboard statistics"""
|
||||||
@@ -68,6 +72,9 @@ class DataBroadcaster:
|
|||||||
|
|
||||||
async def broadcast_custom(self, channel: str, data: Dict[str, Any]):
|
async def broadcast_custom(self, channel: str, data: Dict[str, Any]):
|
||||||
"""Broadcast custom data to a specific channel"""
|
"""Broadcast custom data to a specific channel"""
|
||||||
|
if channel == "vessels":
|
||||||
|
self.enqueue_vessel_update(data)
|
||||||
|
return
|
||||||
await manager.broadcast(
|
await manager.broadcast(
|
||||||
{
|
{
|
||||||
"type": "data_frame",
|
"type": "data_frame",
|
||||||
@@ -75,9 +82,65 @@ class DataBroadcaster:
|
|||||||
"timestamp": to_iso8601_utc(datetime.now(UTC)),
|
"timestamp": to_iso8601_utc(datetime.now(UTC)),
|
||||||
"payload": data,
|
"payload": data,
|
||||||
},
|
},
|
||||||
channel=channel if channel in manager.active_connections else "all",
|
channel=channel,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
async def broadcast_earth_update(self, data: Dict[str, Any]):
|
||||||
|
"""Broadcast Earth visualization refresh hints to connected clients."""
|
||||||
|
await self.broadcast_custom(EARTH_UPDATES_CHANNEL, data)
|
||||||
|
|
||||||
|
def enqueue_vessel_update(self, data: Dict[str, Any]):
|
||||||
|
vessels = data.get("vessels") if isinstance(data, dict) else None
|
||||||
|
if not isinstance(vessels, list):
|
||||||
|
return
|
||||||
|
source = data.get("source")
|
||||||
|
action = data.get("action") or "upsert"
|
||||||
|
created = data.get("created")
|
||||||
|
for vessel in vessels:
|
||||||
|
if not isinstance(vessel, dict):
|
||||||
|
continue
|
||||||
|
mmsi = vessel.get("mmsi")
|
||||||
|
if mmsi in (None, ""):
|
||||||
|
continue
|
||||||
|
self._pending_vessel_updates[str(mmsi)] = {
|
||||||
|
**vessel,
|
||||||
|
"_source": source,
|
||||||
|
"_action": action,
|
||||||
|
"_created": created,
|
||||||
|
}
|
||||||
|
|
||||||
|
async def flush_vessel_updates(self):
|
||||||
|
if not self._pending_vessel_updates:
|
||||||
|
return
|
||||||
|
pending = self._pending_vessel_updates
|
||||||
|
self._pending_vessel_updates = {}
|
||||||
|
vessels = []
|
||||||
|
for item in pending.values():
|
||||||
|
vessel = dict(item)
|
||||||
|
source = vessel.pop("_source", None)
|
||||||
|
action = vessel.pop("_action", "upsert")
|
||||||
|
created = vessel.pop("_created", None)
|
||||||
|
vessel["source"] = source
|
||||||
|
vessel["action"] = action
|
||||||
|
vessel["created"] = created
|
||||||
|
vessels.append(vessel)
|
||||||
|
await manager.broadcast_vessels(
|
||||||
|
{
|
||||||
|
"action": "upsert",
|
||||||
|
"source": "mixed",
|
||||||
|
"created": None,
|
||||||
|
"vessels": vessels,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
async def broadcast_vessels_periodically(self):
|
||||||
|
while self.running:
|
||||||
|
try:
|
||||||
|
await self.flush_vessel_updates()
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
await asyncio.sleep(self._vessel_flush_interval)
|
||||||
|
|
||||||
async def broadcast_datasource_task_update(self, data: Dict[str, Any]):
|
async def broadcast_datasource_task_update(self, data: Dict[str, Any]):
|
||||||
"""Broadcast datasource task progress updates to connected clients."""
|
"""Broadcast datasource task progress updates to connected clients."""
|
||||||
await manager.broadcast(
|
await manager.broadcast(
|
||||||
@@ -87,7 +150,7 @@ class DataBroadcaster:
|
|||||||
"timestamp": to_iso8601_utc(datetime.now(UTC)),
|
"timestamp": to_iso8601_utc(datetime.now(UTC)),
|
||||||
"payload": data,
|
"payload": data,
|
||||||
},
|
},
|
||||||
channel="all",
|
channel="datasource_tasks",
|
||||||
)
|
)
|
||||||
|
|
||||||
def start(self):
|
def start(self):
|
||||||
@@ -95,6 +158,7 @@ class DataBroadcaster:
|
|||||||
if not self.running:
|
if not self.running:
|
||||||
self.running = True
|
self.running = True
|
||||||
self.tasks["dashboard"] = asyncio.create_task(self.broadcast_stats(5))
|
self.tasks["dashboard"] = asyncio.create_task(self.broadcast_stats(5))
|
||||||
|
self.tasks["vessels"] = asyncio.create_task(self.broadcast_vessels_periodically())
|
||||||
|
|
||||||
def stop(self):
|
def stop(self):
|
||||||
"""Stop all broadcasters"""
|
"""Stop all broadcasters"""
|
||||||
@@ -102,6 +166,7 @@ class DataBroadcaster:
|
|||||||
for task in self.tasks.values():
|
for task in self.tasks.values():
|
||||||
task.cancel()
|
task.cancel()
|
||||||
self.tasks.clear()
|
self.tasks.clear()
|
||||||
|
self._pending_vessel_updates.clear()
|
||||||
|
|
||||||
|
|
||||||
broadcaster = DataBroadcaster()
|
broadcaster = DataBroadcaster()
|
||||||
|
|||||||
@@ -1,20 +1,25 @@
|
|||||||
"""WebSocket Connection Manager"""
|
"""WebSocket Connection Manager"""
|
||||||
|
|
||||||
import json
|
from datetime import UTC, datetime
|
||||||
import asyncio
|
from typing import Any, Dict, Set, Optional
|
||||||
from typing import Dict, Set, Optional
|
|
||||||
from datetime import datetime
|
|
||||||
from fastapi import WebSocket
|
from fastapi import WebSocket
|
||||||
import redis.asyncio as redis
|
import redis.asyncio as redis
|
||||||
|
|
||||||
from app.core.config import settings
|
from app.core.config import settings
|
||||||
|
|
||||||
|
MAX_VESSEL_SUBSCRIPTION_LIMIT = 5000
|
||||||
|
MAX_VESSEL_WS_MESSAGE_ITEMS = 1000
|
||||||
|
MAX_VESSEL_BBOX_AREA = 2500.0
|
||||||
|
|
||||||
|
|
||||||
class ConnectionManager:
|
class ConnectionManager:
|
||||||
"""Manages WebSocket connections"""
|
"""Manages WebSocket connections"""
|
||||||
|
|
||||||
def __init__(self):
|
def __init__(self):
|
||||||
self.active_connections: Dict[str, Set[WebSocket]] = {} # user_id -> connections
|
self.active_connections: Dict[str, Set[WebSocket]] = {} # user_id -> connections
|
||||||
|
self.channel_subscriptions: Dict[str, Set[WebSocket]] = {}
|
||||||
|
self.websocket_channels: Dict[WebSocket, Set[str]] = {}
|
||||||
|
self.vessel_subscriptions: Dict[WebSocket, dict[str, Any]] = {}
|
||||||
self.redis_client: Optional[redis.Redis] = None
|
self.redis_client: Optional[redis.Redis] = None
|
||||||
|
|
||||||
async def connect(self, websocket: WebSocket, user_id: str):
|
async def connect(self, websocket: WebSocket, user_id: str):
|
||||||
@@ -40,6 +45,83 @@ class ConnectionManager:
|
|||||||
self.active_connections[user_id].discard(websocket)
|
self.active_connections[user_id].discard(websocket)
|
||||||
if not self.active_connections[user_id]:
|
if not self.active_connections[user_id]:
|
||||||
del self.active_connections[user_id]
|
del self.active_connections[user_id]
|
||||||
|
self.unsubscribe_all(websocket)
|
||||||
|
|
||||||
|
def subscribe(self, websocket: WebSocket, channels: list[str]):
|
||||||
|
normalized_channels = {
|
||||||
|
str(channel).strip()
|
||||||
|
for channel in channels
|
||||||
|
if str(channel).strip()
|
||||||
|
}
|
||||||
|
if not normalized_channels:
|
||||||
|
return
|
||||||
|
|
||||||
|
socket_channels = self.websocket_channels.setdefault(websocket, set())
|
||||||
|
for channel in normalized_channels:
|
||||||
|
self.channel_subscriptions.setdefault(channel, set()).add(websocket)
|
||||||
|
socket_channels.add(channel)
|
||||||
|
|
||||||
|
def unsubscribe(self, websocket: WebSocket, channels: list[str]):
|
||||||
|
for channel in {str(channel).strip() for channel in channels if str(channel).strip()}:
|
||||||
|
subscribers = self.channel_subscriptions.get(channel)
|
||||||
|
if subscribers is not None:
|
||||||
|
subscribers.discard(websocket)
|
||||||
|
if not subscribers:
|
||||||
|
del self.channel_subscriptions[channel]
|
||||||
|
socket_channels = self.websocket_channels.get(websocket)
|
||||||
|
if socket_channels is not None:
|
||||||
|
socket_channels.discard(channel)
|
||||||
|
if not socket_channels:
|
||||||
|
del self.websocket_channels[websocket]
|
||||||
|
|
||||||
|
def unsubscribe_all(self, websocket: WebSocket):
|
||||||
|
channels = list(self.websocket_channels.get(websocket, set()))
|
||||||
|
if channels:
|
||||||
|
self.unsubscribe(websocket, channels)
|
||||||
|
self.vessel_subscriptions.pop(websocket, None)
|
||||||
|
|
||||||
|
def subscribe_vessels(self, websocket: WebSocket, config: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
subscription = self._normalize_vessel_subscription(config)
|
||||||
|
self.channel_subscriptions.setdefault("vessels", set()).add(websocket)
|
||||||
|
self.websocket_channels.setdefault(websocket, set()).add("vessels")
|
||||||
|
self.vessel_subscriptions[websocket] = subscription
|
||||||
|
return subscription
|
||||||
|
|
||||||
|
def _normalize_vessel_subscription(self, config: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
bbox = config.get("bbox")
|
||||||
|
if not isinstance(bbox, (list, tuple)) or len(bbox) != 4:
|
||||||
|
raise ValueError("vessels subscription requires bbox=[lon_min,lat_min,lon_max,lat_max]")
|
||||||
|
try:
|
||||||
|
lon_min, lat_min, lon_max, lat_max = [float(value) for value in bbox]
|
||||||
|
except (TypeError, ValueError) as exc:
|
||||||
|
raise ValueError("bbox values must be numbers") from exc
|
||||||
|
if lat_min > lat_max:
|
||||||
|
lat_min, lat_max = lat_max, lat_min
|
||||||
|
if lon_min > lon_max:
|
||||||
|
lon_min, lon_max = lon_max, lon_min
|
||||||
|
if not (-180 <= lon_min <= 180 and -180 <= lon_max <= 180):
|
||||||
|
raise ValueError("bbox longitude values must be between -180 and 180")
|
||||||
|
if not (-90 <= lat_min <= 90 and -90 <= lat_max <= 90):
|
||||||
|
raise ValueError("bbox latitude values must be between -90 and 90")
|
||||||
|
if (lon_max - lon_min) * (lat_max - lat_min) > MAX_VESSEL_BBOX_AREA:
|
||||||
|
raise ValueError("bbox is too large; zoom in or request a smaller viewport")
|
||||||
|
|
||||||
|
zoom = int(config.get("zoom") or 1)
|
||||||
|
if zoom < 1 or zoom > 20:
|
||||||
|
raise ValueError("zoom must be between 1 and 20")
|
||||||
|
limit = min(max(int(config.get("limit") or 1000), 1), MAX_VESSEL_SUBSCRIPTION_LIMIT)
|
||||||
|
vessel_types = {
|
||||||
|
str(item).strip().lower()
|
||||||
|
for item in str(config.get("type") or "").split(",")
|
||||||
|
if str(item).strip()
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
"bbox": (lon_min, lat_min, lon_max, lat_max),
|
||||||
|
"zoom": zoom,
|
||||||
|
"limit": limit,
|
||||||
|
"type": vessel_types,
|
||||||
|
"last_sent_at": None,
|
||||||
|
}
|
||||||
|
|
||||||
async def send_personal_message(self, message: dict, user_id: str):
|
async def send_personal_message(self, message: dict, user_id: str):
|
||||||
if user_id in self.active_connections:
|
if user_id in self.active_connections:
|
||||||
@@ -54,13 +136,72 @@ class ConnectionManager:
|
|||||||
for user_id in self.active_connections:
|
for user_id in self.active_connections:
|
||||||
await self.send_personal_message(message, user_id)
|
await self.send_personal_message(message, user_id)
|
||||||
else:
|
else:
|
||||||
await self.send_personal_message(message, channel)
|
for connection in list(self.channel_subscriptions.get(channel, set())):
|
||||||
|
try:
|
||||||
|
await connection.send_json(message)
|
||||||
|
except Exception:
|
||||||
|
self.unsubscribe_all(connection)
|
||||||
|
|
||||||
|
async def broadcast_vessels(self, data: dict[str, Any]):
|
||||||
|
vessels = data.get("vessels") if isinstance(data, dict) else None
|
||||||
|
if not isinstance(vessels, list) or not vessels:
|
||||||
|
return
|
||||||
|
|
||||||
|
for connection, subscription in list(self.vessel_subscriptions.items()):
|
||||||
|
matched = [
|
||||||
|
vessel
|
||||||
|
for vessel in vessels
|
||||||
|
if self._vessel_matches_subscription(vessel, subscription)
|
||||||
|
][: min(subscription["limit"], MAX_VESSEL_WS_MESSAGE_ITEMS)]
|
||||||
|
if not matched:
|
||||||
|
continue
|
||||||
|
subscription["last_sent_at"] = datetime.now(UTC)
|
||||||
|
message = {
|
||||||
|
"type": "data_frame",
|
||||||
|
"channel": "vessels",
|
||||||
|
"timestamp": subscription["last_sent_at"].isoformat(),
|
||||||
|
"payload": {
|
||||||
|
**data,
|
||||||
|
"vessels": matched,
|
||||||
|
"subscription": {
|
||||||
|
"bbox": list(subscription["bbox"]),
|
||||||
|
"zoom": subscription["zoom"],
|
||||||
|
"limit": subscription["limit"],
|
||||||
|
},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
try:
|
||||||
|
await connection.send_json(message)
|
||||||
|
except Exception:
|
||||||
|
self.unsubscribe_all(connection)
|
||||||
|
|
||||||
|
def _vessel_matches_subscription(
|
||||||
|
self,
|
||||||
|
vessel: dict[str, Any],
|
||||||
|
subscription: dict[str, Any],
|
||||||
|
) -> bool:
|
||||||
|
try:
|
||||||
|
lon = float(vessel.get("lon"))
|
||||||
|
lat = float(vessel.get("lat"))
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return False
|
||||||
|
lon_min, lat_min, lon_max, lat_max = subscription["bbox"]
|
||||||
|
if not (lon_min <= lon <= lon_max and lat_min <= lat <= lat_max):
|
||||||
|
return False
|
||||||
|
requested_types = subscription.get("type") or set()
|
||||||
|
if not requested_types:
|
||||||
|
return True
|
||||||
|
type_name = str(vessel.get("vessel_type_name") or "").lower()
|
||||||
|
return any(requested_type in type_name for requested_type in requested_types)
|
||||||
|
|
||||||
async def close_all(self):
|
async def close_all(self):
|
||||||
for user_id in self.active_connections:
|
for user_id in self.active_connections:
|
||||||
for connection in self.active_connections[user_id]:
|
for connection in self.active_connections[user_id]:
|
||||||
await connection.close()
|
await connection.close()
|
||||||
self.active_connections.clear()
|
self.active_connections.clear()
|
||||||
|
self.channel_subscriptions.clear()
|
||||||
|
self.websocket_channels.clear()
|
||||||
|
self.vessel_subscriptions.clear()
|
||||||
|
|
||||||
|
|
||||||
manager = ConnectionManager()
|
manager = ConnectionManager()
|
||||||
|
|||||||
328
backend/app/data/seeds/ripe_ris_collector_locations_seed.json
Normal file
328
backend/app/data/seeds/ripe_ris_collector_locations_seed.json
Normal file
@@ -0,0 +1,328 @@
|
|||||||
|
{
|
||||||
|
"_comment": "Seed payload for the bgp_collector_locations DB table. Coordinates were migrated from the legacy RIPE_RIS_COLLECTOR_COORDS table and default to city-center; seeded rows are unverified and should be upgraded in the database with source evidence when known.",
|
||||||
|
"locations": [
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc00",
|
||||||
|
"aliases": ["rrc00", "RIPE RIS rrc00", "AMS-IX"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "AMS-IX",
|
||||||
|
"city": "Amsterdam",
|
||||||
|
"country": "Netherlands",
|
||||||
|
"latitude": 52.3676,
|
||||||
|
"longitude": 4.9041,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc01",
|
||||||
|
"aliases": ["rrc01", "RIPE RIS rrc01", "LINX"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "LINX",
|
||||||
|
"city": "London",
|
||||||
|
"country": "United Kingdom",
|
||||||
|
"latitude": 51.5072,
|
||||||
|
"longitude": -0.1276,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc03",
|
||||||
|
"aliases": ["rrc03", "RIPE RIS rrc03", "AMS-IX"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "AMS-IX",
|
||||||
|
"city": "Amsterdam",
|
||||||
|
"country": "Netherlands",
|
||||||
|
"latitude": 52.3676,
|
||||||
|
"longitude": 4.9041,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc04",
|
||||||
|
"aliases": ["rrc04", "RIPE RIS rrc04", "CIXP", "CERN Internet Exchange Point"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "CIXP",
|
||||||
|
"city": "Geneva",
|
||||||
|
"country": "Switzerland",
|
||||||
|
"latitude": 46.2044,
|
||||||
|
"longitude": 6.1432,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc05",
|
||||||
|
"aliases": ["rrc05", "RIPE RIS rrc05", "VIX", "Vienna Internet Exchange"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "VIX",
|
||||||
|
"city": "Vienna",
|
||||||
|
"country": "Austria",
|
||||||
|
"latitude": 48.2082,
|
||||||
|
"longitude": 16.3738,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc06",
|
||||||
|
"aliases": ["rrc06", "RIPE RIS rrc06", "JPIX", "Otemachi"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "JPIX",
|
||||||
|
"city": "Otemachi",
|
||||||
|
"country": "Japan",
|
||||||
|
"latitude": 35.686,
|
||||||
|
"longitude": 139.7671,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc07",
|
||||||
|
"aliases": ["rrc07", "RIPE RIS rrc07", "Netnod", "Netnod Stockholm"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "Netnod Stockholm",
|
||||||
|
"city": "Stockholm",
|
||||||
|
"country": "Sweden",
|
||||||
|
"latitude": 59.3293,
|
||||||
|
"longitude": 18.0686,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc10",
|
||||||
|
"aliases": ["rrc10", "RIPE RIS rrc10", "MIX", "Milan Internet Exchange"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "MIX",
|
||||||
|
"city": "Milan",
|
||||||
|
"country": "Italy",
|
||||||
|
"latitude": 45.4642,
|
||||||
|
"longitude": 9.19,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc11",
|
||||||
|
"aliases": ["rrc11", "RIPE RIS rrc11", "NYIIX", "New York International Internet Exchange"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "NYIIX",
|
||||||
|
"city": "New York",
|
||||||
|
"country": "United States",
|
||||||
|
"latitude": 40.7128,
|
||||||
|
"longitude": -74.006,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc12",
|
||||||
|
"aliases": ["rrc12", "RIPE RIS rrc12", "DE-CIX", "DE-CIX Frankfurt"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "DE-CIX Frankfurt",
|
||||||
|
"city": "Frankfurt",
|
||||||
|
"country": "Germany",
|
||||||
|
"latitude": 50.1109,
|
||||||
|
"longitude": 8.6821,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc13",
|
||||||
|
"aliases": ["rrc13", "RIPE RIS rrc13", "MSK-IX"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "MSK-IX",
|
||||||
|
"city": "Moscow",
|
||||||
|
"country": "Russia",
|
||||||
|
"latitude": 55.7558,
|
||||||
|
"longitude": 37.6173,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc14",
|
||||||
|
"aliases": ["rrc14", "RIPE RIS rrc14", "PAIX", "Palo Alto Internet Exchange"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "PAIX",
|
||||||
|
"city": "Palo Alto",
|
||||||
|
"country": "United States",
|
||||||
|
"latitude": 37.4419,
|
||||||
|
"longitude": -122.143,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc15",
|
||||||
|
"aliases": ["rrc15", "RIPE RIS rrc15", "PTT.br Sao Paulo", "PTTMetro Sao Paulo"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "PTT.br",
|
||||||
|
"city": "Sao Paulo",
|
||||||
|
"country": "Brazil",
|
||||||
|
"latitude": -23.5558,
|
||||||
|
"longitude": -46.6396,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc16",
|
||||||
|
"aliases": ["rrc16", "RIPE RIS rrc16", "Equinix Miami", "NOTA Miami"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "Equinix Miami",
|
||||||
|
"city": "Miami",
|
||||||
|
"country": "United States",
|
||||||
|
"latitude": 25.7617,
|
||||||
|
"longitude": -80.1918,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc18",
|
||||||
|
"aliases": ["rrc18", "RIPE RIS rrc18", "CATNIX"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "CATNIX",
|
||||||
|
"city": "Barcelona",
|
||||||
|
"country": "Spain",
|
||||||
|
"latitude": 41.3874,
|
||||||
|
"longitude": 2.1686,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc19",
|
||||||
|
"aliases": ["rrc19", "RIPE RIS rrc19", "NAPAfrica", "JINX", "NAPAfrica Johannesburg"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "NAPAfrica Johannesburg",
|
||||||
|
"city": "Johannesburg",
|
||||||
|
"country": "South Africa",
|
||||||
|
"latitude": -26.2041,
|
||||||
|
"longitude": 28.0473,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc20",
|
||||||
|
"aliases": ["rrc20", "RIPE RIS rrc20", "SwissIX"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "SwissIX",
|
||||||
|
"city": "Zurich",
|
||||||
|
"country": "Switzerland",
|
||||||
|
"latitude": 47.3769,
|
||||||
|
"longitude": 8.5417,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc21",
|
||||||
|
"aliases": ["rrc21", "RIPE RIS rrc21", "France-IX Paris"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "France-IX Paris",
|
||||||
|
"city": "Paris",
|
||||||
|
"country": "France",
|
||||||
|
"latitude": 48.8566,
|
||||||
|
"longitude": 2.3522,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc22",
|
||||||
|
"aliases": ["rrc22", "RIPE RIS rrc22", "InterLAN Bucharest"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "InterLAN Bucharest",
|
||||||
|
"city": "Bucharest",
|
||||||
|
"country": "Romania",
|
||||||
|
"latitude": 44.4268,
|
||||||
|
"longitude": 26.1025,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc23",
|
||||||
|
"aliases": ["rrc23", "RIPE RIS rrc23", "Equinix Singapore"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "Equinix Singapore",
|
||||||
|
"city": "Singapore",
|
||||||
|
"country": "Singapore",
|
||||||
|
"latitude": 1.3521,
|
||||||
|
"longitude": 103.8198,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc24",
|
||||||
|
"aliases": ["rrc24", "RIPE RIS rrc24", "LACNIC Montevideo"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "LACNIC Montevideo",
|
||||||
|
"city": "Montevideo",
|
||||||
|
"country": "Uruguay",
|
||||||
|
"latitude": -34.9011,
|
||||||
|
"longitude": -56.1645,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc25",
|
||||||
|
"aliases": ["rrc25", "RIPE RIS rrc25", "AMS-IX"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "AMS-IX",
|
||||||
|
"city": "Amsterdam",
|
||||||
|
"country": "Netherlands",
|
||||||
|
"latitude": 52.3676,
|
||||||
|
"longitude": 4.9041,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_name": "RIPE RIS rrc26",
|
||||||
|
"aliases": ["rrc26", "RIPE RIS rrc26", "UAE-IX"],
|
||||||
|
"operator": "RIPE NCC",
|
||||||
|
"site": "UAE-IX",
|
||||||
|
"city": "Dubai",
|
||||||
|
"country": "United Arab Emirates",
|
||||||
|
"latitude": 25.2048,
|
||||||
|
"longitude": 55.2708,
|
||||||
|
"precision": "city",
|
||||||
|
"confidence": 0.85,
|
||||||
|
"source_note": "Migrated from RIPE_RIS_COLLECTOR_COORDS legacy table",
|
||||||
|
"verified_at": null
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"city_fallbacks": []
|
||||||
|
}
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
from typing import AsyncGenerator
|
from typing import AsyncGenerator
|
||||||
|
|
||||||
from sqlalchemy import text
|
from sqlalchemy import bindparam, text
|
||||||
from sqlalchemy.ext.asyncio import AsyncSession, create_async_engine, async_sessionmaker
|
from sqlalchemy.ext.asyncio import AsyncSession, create_async_engine, async_sessionmaker
|
||||||
from sqlalchemy.orm import declarative_base
|
from sqlalchemy.orm import declarative_base
|
||||||
|
|
||||||
@@ -72,25 +72,112 @@ async def seed_default_datasources(session: AsyncSession):
|
|||||||
await session.commit()
|
await session.commit()
|
||||||
|
|
||||||
|
|
||||||
|
LEGACY_EARTH_BOUNDARY_SOURCES = (
|
||||||
|
"earth_admin0_boundaries",
|
||||||
|
"earth_coastline",
|
||||||
|
"earth_claim_lines",
|
||||||
|
"earth_boundary_tiles",
|
||||||
|
)
|
||||||
|
LEGACY_EARTH_BOUNDARY_DATATYPES = (
|
||||||
|
"earth_boundary_source",
|
||||||
|
"earth_boundary_tiles",
|
||||||
|
)
|
||||||
|
LEGACY_EARTH_BOUNDARY_IDS = (29, 30, 31, 32)
|
||||||
|
|
||||||
|
|
||||||
|
async def purge_legacy_earth_boundary_datasources(session: AsyncSession) -> None:
|
||||||
|
source_names = tuple(LEGACY_EARTH_BOUNDARY_SOURCES)
|
||||||
|
source_ids = tuple(LEGACY_EARTH_BOUNDARY_IDS)
|
||||||
|
data_types = tuple(LEGACY_EARTH_BOUNDARY_DATATYPES)
|
||||||
|
await session.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
DELETE FROM datasource_mapping_templates
|
||||||
|
WHERE target_schema IN :data_types
|
||||||
|
OR datasource_config_id IN (
|
||||||
|
SELECT id FROM datasource_configs WHERE name IN :source_names
|
||||||
|
)
|
||||||
|
"""
|
||||||
|
).bindparams(bindparam("source_names", expanding=True), bindparam("data_types", expanding=True)),
|
||||||
|
{"source_names": list(source_names), "data_types": list(data_types)},
|
||||||
|
)
|
||||||
|
await session.execute(
|
||||||
|
text("DELETE FROM datasource_configs WHERE name IN :source_names").bindparams(
|
||||||
|
bindparam("source_names", expanding=True)
|
||||||
|
),
|
||||||
|
{"source_names": list(source_names)},
|
||||||
|
)
|
||||||
|
await session.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
DELETE FROM collected_data
|
||||||
|
WHERE source IN :source_names OR data_type IN :data_types
|
||||||
|
"""
|
||||||
|
).bindparams(bindparam("source_names", expanding=True), bindparam("data_types", expanding=True)),
|
||||||
|
{"source_names": list(source_names), "data_types": list(data_types)},
|
||||||
|
)
|
||||||
|
await session.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
DELETE FROM data_snapshots
|
||||||
|
WHERE source IN :source_names OR datasource_id IN :source_ids
|
||||||
|
"""
|
||||||
|
).bindparams(bindparam("source_names", expanding=True), bindparam("source_ids", expanding=True)),
|
||||||
|
{"source_names": list(source_names), "source_ids": list(source_ids)},
|
||||||
|
)
|
||||||
|
await session.execute(
|
||||||
|
text("DELETE FROM collection_tasks WHERE datasource_id IN :source_ids").bindparams(
|
||||||
|
bindparam("source_ids", expanding=True)
|
||||||
|
),
|
||||||
|
{"source_ids": list(source_ids)},
|
||||||
|
)
|
||||||
|
await session.execute(
|
||||||
|
text("DELETE FROM data_sources WHERE source IN :source_names OR id IN :source_ids").bindparams(
|
||||||
|
bindparam("source_names", expanding=True), bindparam("source_ids", expanding=True)
|
||||||
|
),
|
||||||
|
{"source_names": list(source_names), "source_ids": list(source_ids)},
|
||||||
|
)
|
||||||
|
await session.commit()
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_LOGIN_USERS = (
|
||||||
|
{
|
||||||
|
"username": "admin",
|
||||||
|
"email": "admin@planet.local",
|
||||||
|
"password": "admin123",
|
||||||
|
"role": "super_admin",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"username": "linkong",
|
||||||
|
"email": "linkong@planet.local",
|
||||||
|
"password": "LK12345678",
|
||||||
|
"role": "super_admin",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
async def ensure_default_admin_user(session: AsyncSession):
|
async def ensure_default_admin_user(session: AsyncSession):
|
||||||
from app.core.security import get_password_hash
|
from app.core.security import get_password_hash
|
||||||
from app.models.user import User
|
from app.models.user import User
|
||||||
|
|
||||||
result = await session.execute(
|
for default_user in DEFAULT_LOGIN_USERS:
|
||||||
text("SELECT id FROM users WHERE username = 'admin'")
|
result = await session.execute(
|
||||||
)
|
text("SELECT id FROM users WHERE username = :username"),
|
||||||
if result.fetchone():
|
{"username": default_user["username"]},
|
||||||
return
|
)
|
||||||
|
if result.fetchone():
|
||||||
session.add(
|
continue
|
||||||
User(
|
|
||||||
username="admin",
|
session.add(
|
||||||
email="admin@planet.local",
|
User(
|
||||||
password_hash=get_password_hash("admin123"),
|
username=default_user["username"],
|
||||||
role="super_admin",
|
email=default_user["email"],
|
||||||
is_active=True,
|
password_hash=get_password_hash(default_user["password"]),
|
||||||
|
role=default_user["role"],
|
||||||
|
is_active=True,
|
||||||
|
email_verified=True,
|
||||||
|
)
|
||||||
)
|
)
|
||||||
)
|
|
||||||
await session.commit()
|
await session.commit()
|
||||||
|
|
||||||
|
|
||||||
@@ -103,13 +190,20 @@ async def init_db():
|
|||||||
import app.models.datasource_config # noqa: F401
|
import app.models.datasource_config # noqa: F401
|
||||||
import app.models.alert # noqa: F401
|
import app.models.alert # noqa: F401
|
||||||
import app.models.bgp_anomaly # noqa: F401
|
import app.models.bgp_anomaly # noqa: F401
|
||||||
|
import app.models.bgp_collector_location # noqa: F401
|
||||||
import app.models.bgp_incident # noqa: F401
|
import app.models.bgp_incident # noqa: F401
|
||||||
import app.models.bgp_observation # noqa: F401
|
import app.models.bgp_observation # noqa: F401
|
||||||
import app.models.collected_data # noqa: F401
|
import app.models.collected_data # noqa: F401
|
||||||
|
import app.models.compute_center_location # noqa: F401
|
||||||
import app.models.system_setting # noqa: F401
|
import app.models.system_setting # noqa: F401
|
||||||
import app.models.playground_session # noqa: F401
|
import app.models.playground_session # noqa: F401
|
||||||
import app.models.playground_message # noqa: F401
|
import app.models.playground_message # noqa: F401
|
||||||
import app.models.system_log # noqa: F401
|
import app.models.system_log # noqa: F401
|
||||||
|
import app.models.vessel # noqa: F401
|
||||||
|
import app.models.vessel_enrichment # noqa: F401
|
||||||
|
import app.models.datasource_mapping # noqa: F401
|
||||||
|
import app.models.earth_news # noqa: F401
|
||||||
|
import app.models.earth_interactable # noqa: F401
|
||||||
|
|
||||||
logger.warning_event(
|
logger.warning_event(
|
||||||
"Database pool settings active",
|
"Database pool settings active",
|
||||||
@@ -125,6 +219,31 @@ async def init_db():
|
|||||||
|
|
||||||
async with engine.begin() as conn:
|
async with engine.begin() as conn:
|
||||||
await conn.run_sync(Base.metadata.create_all)
|
await conn.run_sync(Base.metadata.create_all)
|
||||||
|
users_email_verified_existed = (
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
SELECT 1
|
||||||
|
FROM information_schema.columns
|
||||||
|
WHERE table_name = 'users' AND column_name = 'email_verified'
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
).fetchone() is not None
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
ALTER TABLE users
|
||||||
|
ADD COLUMN IF NOT EXISTS gatekeeper_groups JSONB DEFAULT '[]'::jsonb,
|
||||||
|
ADD COLUMN IF NOT EXISTS email_verified BOOLEAN NOT NULL DEFAULT FALSE,
|
||||||
|
ADD COLUMN IF NOT EXISTS pending_email VARCHAR(255)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
if not users_email_verified_existed:
|
||||||
|
await conn.execute(
|
||||||
|
text("UPDATE users SET email_verified = TRUE WHERE email_verified = FALSE")
|
||||||
|
)
|
||||||
await conn.execute(
|
await conn.execute(
|
||||||
text(
|
text(
|
||||||
"""
|
"""
|
||||||
@@ -140,11 +259,448 @@ async def init_db():
|
|||||||
"""
|
"""
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE TABLE IF NOT EXISTS earth_data_change_events (
|
||||||
|
id BIGSERIAL PRIMARY KEY,
|
||||||
|
table_name VARCHAR(128) NOT NULL,
|
||||||
|
operation VARCHAR(16) NOT NULL,
|
||||||
|
source VARCHAR(128),
|
||||||
|
entity_key VARCHAR(255),
|
||||||
|
payload JSONB NOT NULL,
|
||||||
|
occurred_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
||||||
|
consumed_at TIMESTAMPTZ
|
||||||
|
)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_earth_data_change_events_unconsumed
|
||||||
|
ON earth_data_change_events (consumed_at, id)
|
||||||
|
WHERE consumed_at IS NULL
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE OR REPLACE FUNCTION planet_emit_earth_data_changed_statement(
|
||||||
|
change_table TEXT,
|
||||||
|
change_operation TEXT,
|
||||||
|
change_source TEXT,
|
||||||
|
source_record_count INTEGER,
|
||||||
|
source_entity_keys TEXT[]
|
||||||
|
)
|
||||||
|
RETURNS VOID AS $$
|
||||||
|
DECLARE
|
||||||
|
change_event_id BIGINT;
|
||||||
|
change_payload JSONB;
|
||||||
|
BEGIN
|
||||||
|
change_payload := jsonb_build_object(
|
||||||
|
'event', 'earth.layer.changed',
|
||||||
|
'table', change_table,
|
||||||
|
'operation', change_operation,
|
||||||
|
'source', change_source,
|
||||||
|
'entity_key', NULL,
|
||||||
|
'entity_keys', COALESCE(to_jsonb(source_entity_keys), '[]'::jsonb),
|
||||||
|
'records_processed', COALESCE(source_record_count, 0),
|
||||||
|
'occurred_at', NOW()
|
||||||
|
);
|
||||||
|
|
||||||
|
INSERT INTO earth_data_change_events (
|
||||||
|
table_name,
|
||||||
|
operation,
|
||||||
|
source,
|
||||||
|
entity_key,
|
||||||
|
payload,
|
||||||
|
occurred_at
|
||||||
|
) VALUES (
|
||||||
|
change_table,
|
||||||
|
change_operation,
|
||||||
|
change_source,
|
||||||
|
NULL,
|
||||||
|
change_payload,
|
||||||
|
NOW()
|
||||||
|
)
|
||||||
|
RETURNING id INTO change_event_id;
|
||||||
|
|
||||||
|
change_payload := change_payload || jsonb_build_object(
|
||||||
|
'event_id', change_event_id
|
||||||
|
);
|
||||||
|
|
||||||
|
UPDATE earth_data_change_events
|
||||||
|
SET payload = change_payload
|
||||||
|
WHERE id = change_event_id;
|
||||||
|
|
||||||
|
PERFORM pg_notify(
|
||||||
|
'planet_earth_data_changes',
|
||||||
|
change_payload::text
|
||||||
|
);
|
||||||
|
END;
|
||||||
|
$$ LANGUAGE plpgsql;
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE OR REPLACE FUNCTION planet_emit_collected_data_changed_statement(
|
||||||
|
change_operation TEXT,
|
||||||
|
change_source TEXT,
|
||||||
|
source_record_count INTEGER,
|
||||||
|
source_entity_keys TEXT[]
|
||||||
|
)
|
||||||
|
RETURNS VOID AS $$
|
||||||
|
BEGIN
|
||||||
|
PERFORM planet_emit_earth_data_changed_statement(
|
||||||
|
'collected_data',
|
||||||
|
change_operation,
|
||||||
|
change_source,
|
||||||
|
source_record_count,
|
||||||
|
source_entity_keys
|
||||||
|
);
|
||||||
|
END;
|
||||||
|
$$ LANGUAGE plpgsql;
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE OR REPLACE FUNCTION planet_notify_earth_table_changed_statement()
|
||||||
|
RETURNS trigger AS $$
|
||||||
|
DECLARE
|
||||||
|
change_source TEXT;
|
||||||
|
source_record_count INTEGER;
|
||||||
|
source_entity_keys TEXT[];
|
||||||
|
BEGIN
|
||||||
|
IF TG_OP = 'INSERT' THEN
|
||||||
|
FOR change_source IN
|
||||||
|
SELECT DISTINCT COALESCE(NULLIF(row_data->>'source', ''), TG_TABLE_NAME)
|
||||||
|
FROM (SELECT to_jsonb(t) AS row_data FROM new_rows AS t) changed_rows
|
||||||
|
LOOP
|
||||||
|
SELECT
|
||||||
|
COUNT(*),
|
||||||
|
ARRAY(
|
||||||
|
SELECT DISTINCT COALESCE(
|
||||||
|
NULLIF(row_data->>'entity_key', ''),
|
||||||
|
NULLIF(row_data->>'source_id', ''),
|
||||||
|
NULLIF(row_data->>'incident_key', ''),
|
||||||
|
NULLIF(row_data->>'id', ''),
|
||||||
|
NULLIF(row_data->>'mmsi', '')
|
||||||
|
)
|
||||||
|
FROM (SELECT to_jsonb(t) AS row_data FROM new_rows AS t) rows_for_keys
|
||||||
|
WHERE COALESCE(NULLIF(row_data->>'source', ''), TG_TABLE_NAME) = change_source
|
||||||
|
LIMIT 20
|
||||||
|
)
|
||||||
|
INTO source_record_count, source_entity_keys
|
||||||
|
FROM (SELECT to_jsonb(t) AS row_data FROM new_rows AS t) rows_for_count
|
||||||
|
WHERE COALESCE(NULLIF(row_data->>'source', ''), TG_TABLE_NAME) = change_source;
|
||||||
|
|
||||||
|
PERFORM planet_emit_earth_data_changed_statement(
|
||||||
|
TG_TABLE_NAME,
|
||||||
|
TG_OP,
|
||||||
|
change_source,
|
||||||
|
source_record_count,
|
||||||
|
source_entity_keys
|
||||||
|
);
|
||||||
|
END LOOP;
|
||||||
|
ELSIF TG_OP = 'DELETE' THEN
|
||||||
|
FOR change_source IN
|
||||||
|
SELECT DISTINCT COALESCE(NULLIF(row_data->>'source', ''), TG_TABLE_NAME)
|
||||||
|
FROM (SELECT to_jsonb(t) AS row_data FROM old_rows AS t) changed_rows
|
||||||
|
LOOP
|
||||||
|
SELECT
|
||||||
|
COUNT(*),
|
||||||
|
ARRAY(
|
||||||
|
SELECT DISTINCT COALESCE(
|
||||||
|
NULLIF(row_data->>'entity_key', ''),
|
||||||
|
NULLIF(row_data->>'source_id', ''),
|
||||||
|
NULLIF(row_data->>'incident_key', ''),
|
||||||
|
NULLIF(row_data->>'id', ''),
|
||||||
|
NULLIF(row_data->>'mmsi', '')
|
||||||
|
)
|
||||||
|
FROM (SELECT to_jsonb(t) AS row_data FROM old_rows AS t) rows_for_keys
|
||||||
|
WHERE COALESCE(NULLIF(row_data->>'source', ''), TG_TABLE_NAME) = change_source
|
||||||
|
LIMIT 20
|
||||||
|
)
|
||||||
|
INTO source_record_count, source_entity_keys
|
||||||
|
FROM (SELECT to_jsonb(t) AS row_data FROM old_rows AS t) rows_for_count
|
||||||
|
WHERE COALESCE(NULLIF(row_data->>'source', ''), TG_TABLE_NAME) = change_source;
|
||||||
|
|
||||||
|
PERFORM planet_emit_earth_data_changed_statement(
|
||||||
|
TG_TABLE_NAME,
|
||||||
|
TG_OP,
|
||||||
|
change_source,
|
||||||
|
source_record_count,
|
||||||
|
source_entity_keys
|
||||||
|
);
|
||||||
|
END LOOP;
|
||||||
|
ELSIF TG_OP = 'UPDATE' THEN
|
||||||
|
FOR change_source IN
|
||||||
|
SELECT DISTINCT COALESCE(NULLIF(row_data->>'source', ''), TG_TABLE_NAME)
|
||||||
|
FROM (
|
||||||
|
SELECT to_jsonb(t) AS row_data FROM new_rows AS t
|
||||||
|
UNION ALL
|
||||||
|
SELECT to_jsonb(t) AS row_data FROM old_rows AS t
|
||||||
|
) changed_rows
|
||||||
|
LOOP
|
||||||
|
SELECT
|
||||||
|
COUNT(*),
|
||||||
|
ARRAY(
|
||||||
|
SELECT DISTINCT COALESCE(
|
||||||
|
NULLIF(row_data->>'entity_key', ''),
|
||||||
|
NULLIF(row_data->>'source_id', ''),
|
||||||
|
NULLIF(row_data->>'incident_key', ''),
|
||||||
|
NULLIF(row_data->>'id', ''),
|
||||||
|
NULLIF(row_data->>'mmsi', '')
|
||||||
|
)
|
||||||
|
FROM (
|
||||||
|
SELECT to_jsonb(t) AS row_data FROM new_rows AS t
|
||||||
|
UNION ALL
|
||||||
|
SELECT to_jsonb(t) AS row_data FROM old_rows AS t
|
||||||
|
) rows_for_keys
|
||||||
|
WHERE COALESCE(NULLIF(row_data->>'source', ''), TG_TABLE_NAME) = change_source
|
||||||
|
LIMIT 20
|
||||||
|
)
|
||||||
|
INTO source_record_count, source_entity_keys
|
||||||
|
FROM (
|
||||||
|
SELECT to_jsonb(t) AS row_data FROM new_rows AS t
|
||||||
|
UNION ALL
|
||||||
|
SELECT to_jsonb(t) AS row_data FROM old_rows AS t
|
||||||
|
) rows_for_count
|
||||||
|
WHERE COALESCE(NULLIF(row_data->>'source', ''), TG_TABLE_NAME) = change_source;
|
||||||
|
|
||||||
|
PERFORM planet_emit_earth_data_changed_statement(
|
||||||
|
TG_TABLE_NAME,
|
||||||
|
TG_OP,
|
||||||
|
change_source,
|
||||||
|
source_record_count,
|
||||||
|
source_entity_keys
|
||||||
|
);
|
||||||
|
END LOOP;
|
||||||
|
END IF;
|
||||||
|
|
||||||
|
RETURN NULL;
|
||||||
|
END;
|
||||||
|
$$ LANGUAGE plpgsql;
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE OR REPLACE FUNCTION planet_notify_collected_data_changed_statement()
|
||||||
|
RETURNS trigger AS $$
|
||||||
|
DECLARE
|
||||||
|
change_source TEXT;
|
||||||
|
source_record_count INTEGER;
|
||||||
|
source_entity_keys TEXT[];
|
||||||
|
BEGIN
|
||||||
|
IF TG_OP = 'INSERT' THEN
|
||||||
|
FOR change_source IN
|
||||||
|
SELECT DISTINCT source FROM new_rows WHERE source IS NOT NULL
|
||||||
|
LOOP
|
||||||
|
SELECT
|
||||||
|
COUNT(*),
|
||||||
|
ARRAY(
|
||||||
|
SELECT DISTINCT COALESCE(entity_key, source_id, id::text)
|
||||||
|
FROM new_rows
|
||||||
|
WHERE source = change_source
|
||||||
|
LIMIT 20
|
||||||
|
)
|
||||||
|
INTO source_record_count, source_entity_keys
|
||||||
|
FROM new_rows
|
||||||
|
WHERE source = change_source;
|
||||||
|
|
||||||
|
PERFORM planet_emit_collected_data_changed_statement(
|
||||||
|
TG_OP,
|
||||||
|
change_source,
|
||||||
|
source_record_count,
|
||||||
|
source_entity_keys
|
||||||
|
);
|
||||||
|
END LOOP;
|
||||||
|
ELSIF TG_OP = 'DELETE' THEN
|
||||||
|
FOR change_source IN
|
||||||
|
SELECT DISTINCT source FROM old_rows WHERE source IS NOT NULL
|
||||||
|
LOOP
|
||||||
|
SELECT
|
||||||
|
COUNT(*),
|
||||||
|
ARRAY(
|
||||||
|
SELECT DISTINCT COALESCE(entity_key, source_id, id::text)
|
||||||
|
FROM old_rows
|
||||||
|
WHERE source = change_source
|
||||||
|
LIMIT 20
|
||||||
|
)
|
||||||
|
INTO source_record_count, source_entity_keys
|
||||||
|
FROM old_rows
|
||||||
|
WHERE source = change_source;
|
||||||
|
|
||||||
|
PERFORM planet_emit_collected_data_changed_statement(
|
||||||
|
TG_OP,
|
||||||
|
change_source,
|
||||||
|
source_record_count,
|
||||||
|
source_entity_keys
|
||||||
|
);
|
||||||
|
END LOOP;
|
||||||
|
ELSIF TG_OP = 'UPDATE' THEN
|
||||||
|
FOR change_source IN
|
||||||
|
SELECT DISTINCT source FROM (
|
||||||
|
SELECT source FROM new_rows
|
||||||
|
UNION
|
||||||
|
SELECT source FROM old_rows
|
||||||
|
) changed_sources
|
||||||
|
WHERE source IS NOT NULL
|
||||||
|
LOOP
|
||||||
|
SELECT
|
||||||
|
COUNT(*),
|
||||||
|
ARRAY(
|
||||||
|
SELECT DISTINCT COALESCE(entity_key, source_id, id::text)
|
||||||
|
FROM (
|
||||||
|
SELECT id, source_id, entity_key, source FROM new_rows
|
||||||
|
UNION ALL
|
||||||
|
SELECT id, source_id, entity_key, source FROM old_rows
|
||||||
|
) changed_rows
|
||||||
|
WHERE source = change_source
|
||||||
|
LIMIT 20
|
||||||
|
)
|
||||||
|
INTO source_record_count, source_entity_keys
|
||||||
|
FROM (
|
||||||
|
SELECT id, source_id, entity_key, source FROM new_rows
|
||||||
|
UNION ALL
|
||||||
|
SELECT id, source_id, entity_key, source FROM old_rows
|
||||||
|
) changed_rows
|
||||||
|
WHERE source = change_source;
|
||||||
|
|
||||||
|
PERFORM planet_emit_collected_data_changed_statement(
|
||||||
|
TG_OP,
|
||||||
|
change_source,
|
||||||
|
source_record_count,
|
||||||
|
source_entity_keys
|
||||||
|
);
|
||||||
|
END LOOP;
|
||||||
|
END IF;
|
||||||
|
|
||||||
|
RETURN NULL;
|
||||||
|
END;
|
||||||
|
$$ LANGUAGE plpgsql;
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
for statement in (
|
||||||
|
"DROP TRIGGER IF EXISTS tr_planet_collected_data_changed ON collected_data",
|
||||||
|
"DROP TRIGGER IF EXISTS tr_planet_collected_data_changed_insert ON collected_data",
|
||||||
|
"DROP TRIGGER IF EXISTS tr_planet_collected_data_changed_update ON collected_data",
|
||||||
|
"DROP TRIGGER IF EXISTS tr_planet_collected_data_changed_delete ON collected_data",
|
||||||
|
"DROP FUNCTION IF EXISTS planet_notify_collected_data_changed()",
|
||||||
|
"""
|
||||||
|
CREATE TRIGGER tr_planet_collected_data_changed_insert
|
||||||
|
AFTER INSERT ON collected_data
|
||||||
|
REFERENCING NEW TABLE AS new_rows
|
||||||
|
FOR EACH STATEMENT
|
||||||
|
EXECUTE FUNCTION planet_notify_collected_data_changed_statement()
|
||||||
|
""",
|
||||||
|
"""
|
||||||
|
CREATE TRIGGER tr_planet_collected_data_changed_update
|
||||||
|
AFTER UPDATE ON collected_data
|
||||||
|
REFERENCING OLD TABLE AS old_rows NEW TABLE AS new_rows
|
||||||
|
FOR EACH STATEMENT
|
||||||
|
EXECUTE FUNCTION planet_notify_collected_data_changed_statement()
|
||||||
|
""",
|
||||||
|
"""
|
||||||
|
CREATE TRIGGER tr_planet_collected_data_changed_delete
|
||||||
|
AFTER DELETE ON collected_data
|
||||||
|
REFERENCING OLD TABLE AS old_rows
|
||||||
|
FOR EACH STATEMENT
|
||||||
|
EXECUTE FUNCTION planet_notify_collected_data_changed_statement()
|
||||||
|
""",
|
||||||
|
):
|
||||||
|
await conn.execute(text(statement))
|
||||||
|
for table_name in (
|
||||||
|
"bgp_observations",
|
||||||
|
"bgp_anomalies",
|
||||||
|
"bgp_incidents",
|
||||||
|
"bgp_collector_locations",
|
||||||
|
"vessel_static",
|
||||||
|
"vessel_position",
|
||||||
|
"ais_raw_observations",
|
||||||
|
"ais_source_health",
|
||||||
|
"compute_center_locations",
|
||||||
|
"earth_interactables",
|
||||||
|
"earth_news_items",
|
||||||
|
):
|
||||||
|
for statement in (
|
||||||
|
f"DROP TRIGGER IF EXISTS tr_planet_{table_name}_changed_insert ON {table_name}",
|
||||||
|
f"DROP TRIGGER IF EXISTS tr_planet_{table_name}_changed_update ON {table_name}",
|
||||||
|
f"DROP TRIGGER IF EXISTS tr_planet_{table_name}_changed_delete ON {table_name}",
|
||||||
|
f"""
|
||||||
|
CREATE TRIGGER tr_planet_{table_name}_changed_insert
|
||||||
|
AFTER INSERT ON {table_name}
|
||||||
|
REFERENCING NEW TABLE AS new_rows
|
||||||
|
FOR EACH STATEMENT
|
||||||
|
EXECUTE FUNCTION planet_notify_earth_table_changed_statement()
|
||||||
|
""",
|
||||||
|
f"""
|
||||||
|
CREATE TRIGGER tr_planet_{table_name}_changed_update
|
||||||
|
AFTER UPDATE ON {table_name}
|
||||||
|
REFERENCING OLD TABLE AS old_rows NEW TABLE AS new_rows
|
||||||
|
FOR EACH STATEMENT
|
||||||
|
EXECUTE FUNCTION planet_notify_earth_table_changed_statement()
|
||||||
|
""",
|
||||||
|
f"""
|
||||||
|
CREATE TRIGGER tr_planet_{table_name}_changed_delete
|
||||||
|
AFTER DELETE ON {table_name}
|
||||||
|
REFERENCING OLD TABLE AS old_rows
|
||||||
|
FOR EACH STATEMENT
|
||||||
|
EXECUTE FUNCTION planet_notify_earth_table_changed_statement()
|
||||||
|
""",
|
||||||
|
):
|
||||||
|
await conn.execute(text(statement))
|
||||||
await conn.execute(
|
await conn.execute(
|
||||||
text(
|
text(
|
||||||
"""
|
"""
|
||||||
ALTER TABLE collection_tasks
|
ALTER TABLE collection_tasks
|
||||||
ADD COLUMN IF NOT EXISTS phase VARCHAR(30) DEFAULT 'queued'
|
ADD COLUMN IF NOT EXISTS phase VARCHAR(30) DEFAULT 'queued',
|
||||||
|
ADD COLUMN IF NOT EXISTS phase_progress DOUBLE PRECISION,
|
||||||
|
ADD COLUMN IF NOT EXISTS phase_message VARCHAR(255),
|
||||||
|
ADD COLUMN IF NOT EXISTS phase_current BIGINT,
|
||||||
|
ADD COLUMN IF NOT EXISTS phase_total BIGINT,
|
||||||
|
ADD COLUMN IF NOT EXISTS phase_unit VARCHAR(30),
|
||||||
|
ADD COLUMN IF NOT EXISTS source VARCHAR(100),
|
||||||
|
ADD COLUMN IF NOT EXISTS task_type VARCHAR(30) NOT NULL DEFAULT 'collect',
|
||||||
|
ADD COLUMN IF NOT EXISTS payload JSONB NOT NULL DEFAULT '{}'::jsonb,
|
||||||
|
ADD COLUMN IF NOT EXISTS rollback_policy VARCHAR(40) NOT NULL DEFAULT 'keep_committed_batches',
|
||||||
|
ADD COLUMN IF NOT EXISTS dedupe_key VARCHAR(180),
|
||||||
|
ADD COLUMN IF NOT EXISTS worker_id VARCHAR(120),
|
||||||
|
ADD COLUMN IF NOT EXISTS locked_at TIMESTAMPTZ,
|
||||||
|
ADD COLUMN IF NOT EXISTS requested_cancel_at TIMESTAMPTZ,
|
||||||
|
ADD COLUMN IF NOT EXISTS cancel_reason TEXT
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
ALTER TABLE earth_news_items
|
||||||
|
ADD COLUMN IF NOT EXISTS content_language VARCHAR(32) NOT NULL DEFAULT 'en',
|
||||||
|
ADD COLUMN IF NOT EXISTS localizations JSONB NOT NULL DEFAULT '{}'::jsonb,
|
||||||
|
ADD COLUMN IF NOT EXISTS enrichment_status VARCHAR(80) NOT NULL DEFAULT 'pending',
|
||||||
|
ADD COLUMN IF NOT EXISTS enrichment_error TEXT,
|
||||||
|
ADD COLUMN IF NOT EXISTS enriched_at TIMESTAMPTZ
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
ALTER TABLE earth_interactables
|
||||||
|
ADD COLUMN IF NOT EXISTS altitude DOUBLE PRECISION,
|
||||||
|
ADD COLUMN IF NOT EXISTS revision INTEGER NOT NULL DEFAULT 1,
|
||||||
|
ADD COLUMN IF NOT EXISTS is_deleted BOOLEAN NOT NULL DEFAULT FALSE,
|
||||||
|
ADD COLUMN IF NOT EXISTS deleted_at TIMESTAMPTZ
|
||||||
"""
|
"""
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
@@ -156,6 +712,108 @@ async def init_db():
|
|||||||
"""
|
"""
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_earth_news_enrichment_status
|
||||||
|
ON earth_news_items (enrichment_status)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_earth_news_enriched_at
|
||||||
|
ON earth_news_items (enriched_at)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_earth_interactables_layer_deleted
|
||||||
|
ON earth_interactables (layer, is_deleted)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_earth_interactables_updated_at
|
||||||
|
ON earth_interactables (updated_at)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_collection_tasks_source_status
|
||||||
|
ON collection_tasks (source, status)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_collection_tasks_queue
|
||||||
|
ON collection_tasks (status, created_at, id)
|
||||||
|
WHERE status = 'queued'
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_collection_tasks_dedupe
|
||||||
|
ON collection_tasks (dedupe_key)
|
||||||
|
WHERE dedupe_key IS NOT NULL
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_collected_data_source_current_id
|
||||||
|
ON collected_data (source, is_current, id)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_collected_data_source_task_id
|
||||||
|
ON collected_data (source, task_id, id)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_ais_raw_schema_observed_entity
|
||||||
|
ON ais_raw_observations (target_schema, observed_at, entity_key)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_ais_raw_schema_observed_desc
|
||||||
|
ON ais_raw_observations (target_schema, observed_at DESC)
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await conn.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
CREATE INDEX IF NOT EXISTS idx_ais_raw_payload_lon_lat
|
||||||
|
ON ais_raw_observations (
|
||||||
|
((normalized_payload->>'lon')::double precision),
|
||||||
|
((normalized_payload->>'lat')::double precision)
|
||||||
|
)
|
||||||
|
WHERE target_schema = 'vessel_ais'
|
||||||
|
"""
|
||||||
|
)
|
||||||
|
)
|
||||||
await conn.execute(
|
await conn.execute(
|
||||||
text(
|
text(
|
||||||
"""
|
"""
|
||||||
@@ -176,5 +834,15 @@ async def init_db():
|
|||||||
)
|
)
|
||||||
|
|
||||||
async with async_session_factory() as session:
|
async with async_session_factory() as session:
|
||||||
|
from app.services.bgp_collector_locations import (
|
||||||
|
seed_default_bgp_collector_locations,
|
||||||
|
)
|
||||||
|
from app.services.compute_center_locations import (
|
||||||
|
seed_compute_center_locations_from_source_coords,
|
||||||
|
)
|
||||||
|
|
||||||
|
await seed_default_bgp_collector_locations(session)
|
||||||
|
await seed_compute_center_locations_from_source_coords(session)
|
||||||
await seed_default_datasources(session)
|
await seed_default_datasources(session)
|
||||||
|
await purge_legacy_earth_boundary_datasources(session)
|
||||||
await ensure_default_admin_user(session)
|
await ensure_default_admin_user(session)
|
||||||
|
|||||||
@@ -1,8 +1,10 @@
|
|||||||
from contextlib import asynccontextmanager
|
from contextlib import asynccontextmanager
|
||||||
|
from pathlib import Path
|
||||||
from uuid import uuid4
|
from uuid import uuid4
|
||||||
|
|
||||||
from fastapi import FastAPI
|
from fastapi import FastAPI
|
||||||
from fastapi.middleware.cors import CORSMiddleware
|
from fastapi.middleware.cors import CORSMiddleware
|
||||||
|
from fastapi.staticfiles import StaticFiles
|
||||||
from starlette.middleware.base import BaseHTTPMiddleware
|
from starlette.middleware.base import BaseHTTPMiddleware
|
||||||
|
|
||||||
from app.api.main import api_router
|
from app.api.main import api_router
|
||||||
@@ -18,6 +20,15 @@ from app.services.scheduler import (
|
|||||||
stop_scheduler,
|
stop_scheduler,
|
||||||
sync_scheduler_with_datasources,
|
sync_scheduler_with_datasources,
|
||||||
)
|
)
|
||||||
|
from app.services.earth_news_worker import (
|
||||||
|
start_earth_news_target_worker,
|
||||||
|
stop_earth_news_target_worker,
|
||||||
|
)
|
||||||
|
from app.services.earth_db_change_listener import (
|
||||||
|
start_earth_db_change_listener,
|
||||||
|
stop_earth_db_change_listener,
|
||||||
|
)
|
||||||
|
from app.services.data_jobs import start_data_job_worker, stop_data_job_worker
|
||||||
|
|
||||||
|
|
||||||
configure_logging()
|
configure_logging()
|
||||||
@@ -53,7 +64,13 @@ async def lifespan(app: FastAPI):
|
|||||||
start_scheduler()
|
start_scheduler()
|
||||||
await sync_scheduler_with_datasources()
|
await sync_scheduler_with_datasources()
|
||||||
broadcaster.start()
|
broadcaster.start()
|
||||||
|
start_data_job_worker()
|
||||||
|
start_earth_db_change_listener()
|
||||||
|
start_earth_news_target_worker()
|
||||||
yield
|
yield
|
||||||
|
await stop_earth_news_target_worker()
|
||||||
|
await stop_earth_db_change_listener()
|
||||||
|
await stop_data_job_worker()
|
||||||
broadcaster.stop()
|
broadcaster.stop()
|
||||||
stop_scheduler()
|
stop_scheduler()
|
||||||
|
|
||||||
@@ -82,6 +99,14 @@ app.add_middleware(WebSocketCORSMiddleware)
|
|||||||
app.include_router(api_router, prefix="/api/v1")
|
app.include_router(api_router, prefix="/api/v1")
|
||||||
app.include_router(websocket.router)
|
app.include_router(websocket.router)
|
||||||
|
|
||||||
|
EARTH_BRAND_ASSET_DIR = Path(__file__).resolve().parents[2] / "data" / "earth-brand"
|
||||||
|
EARTH_BRAND_ASSET_DIR.mkdir(parents=True, exist_ok=True)
|
||||||
|
app.mount(
|
||||||
|
"/earth-brand-assets",
|
||||||
|
StaticFiles(directory=str(EARTH_BRAND_ASSET_DIR)),
|
||||||
|
name="earth-brand-assets",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
@app.get("/health")
|
@app.get("/health")
|
||||||
async def health_check():
|
async def health_check():
|
||||||
|
|||||||
@@ -6,12 +6,18 @@ from app.models.datasource import DataSource
|
|||||||
from app.models.datasource_config import DataSourceConfig
|
from app.models.datasource_config import DataSourceConfig
|
||||||
from app.models.alert import Alert, AlertSeverity, AlertStatus
|
from app.models.alert import Alert, AlertSeverity, AlertStatus
|
||||||
from app.models.bgp_anomaly import BGPAnomaly
|
from app.models.bgp_anomaly import BGPAnomaly
|
||||||
|
from app.models.bgp_collector_location import BGPCollectorLocation
|
||||||
from app.models.bgp_incident import BGPIncident
|
from app.models.bgp_incident import BGPIncident
|
||||||
from app.models.bgp_observation import BGPObservation
|
from app.models.bgp_observation import BGPObservation
|
||||||
|
from app.models.compute_center_location import ComputeCenterLocationRecord
|
||||||
from app.models.system_setting import SystemSetting
|
from app.models.system_setting import SystemSetting
|
||||||
from app.models.playground_session import PlaygroundSession
|
from app.models.playground_session import PlaygroundSession
|
||||||
from app.models.playground_message import PlaygroundMessage
|
from app.models.playground_message import PlaygroundMessage
|
||||||
from app.models.system_log import SystemLog, AuditLog
|
from app.models.system_log import SystemLog, AuditLog
|
||||||
|
from app.models.vessel import AISConflictRecord, AISRawObservation, AISSourceHealth, VesselPosition, VesselStatic
|
||||||
|
from app.models.datasource_mapping import DataSourceMappingTemplate
|
||||||
|
from app.models.earth_news import EarthNewsItem
|
||||||
|
from app.models.earth_interactable import EarthInteractable
|
||||||
|
|
||||||
__all__ = [
|
__all__ = [
|
||||||
"User",
|
"User",
|
||||||
@@ -25,8 +31,20 @@ __all__ = [
|
|||||||
"AlertSeverity",
|
"AlertSeverity",
|
||||||
"AlertStatus",
|
"AlertStatus",
|
||||||
"BGPAnomaly",
|
"BGPAnomaly",
|
||||||
|
"BGPCollectorLocation",
|
||||||
"BGPIncident",
|
"BGPIncident",
|
||||||
"BGPObservation",
|
"BGPObservation",
|
||||||
|
"ComputeCenterLocationRecord",
|
||||||
"SystemLog",
|
"SystemLog",
|
||||||
"AuditLog",
|
"AuditLog",
|
||||||
|
"PlaygroundSession",
|
||||||
|
"PlaygroundMessage",
|
||||||
|
"VesselPosition",
|
||||||
|
"VesselStatic",
|
||||||
|
"AISRawObservation",
|
||||||
|
"AISConflictRecord",
|
||||||
|
"AISSourceHealth",
|
||||||
|
"DataSourceMappingTemplate",
|
||||||
|
"EarthNewsItem",
|
||||||
|
"EarthInteractable",
|
||||||
]
|
]
|
||||||
|
|||||||
52
backend/app/models/bgp_collector_location.py
Normal file
52
backend/app/models/bgp_collector_location.py
Normal file
@@ -0,0 +1,52 @@
|
|||||||
|
"""Stored BGP route-collector locations."""
|
||||||
|
|
||||||
|
from sqlalchemy import Boolean, Column, DateTime, Float, Integer, JSON, String, Text
|
||||||
|
from sqlalchemy.sql import func
|
||||||
|
|
||||||
|
from app.core.time import to_iso8601_utc
|
||||||
|
from app.db.session import Base
|
||||||
|
|
||||||
|
|
||||||
|
class BGPCollectorLocation(Base):
|
||||||
|
"""Current known location for a BGP route collector."""
|
||||||
|
|
||||||
|
__tablename__ = "bgp_collector_locations"
|
||||||
|
|
||||||
|
id = Column(Integer, primary_key=True, autoincrement=True)
|
||||||
|
collector_id = Column(String(100), nullable=False, unique=True, index=True)
|
||||||
|
operator = Column(String(255), nullable=True)
|
||||||
|
site = Column(String(255), nullable=True)
|
||||||
|
city = Column(String(255), nullable=True)
|
||||||
|
country = Column(String(255), nullable=True)
|
||||||
|
latitude = Column(Float, nullable=True)
|
||||||
|
longitude = Column(Float, nullable=True)
|
||||||
|
precision = Column(String(30), nullable=False, default="city")
|
||||||
|
confidence = Column(Float, nullable=True)
|
||||||
|
source = Column(String(80), nullable=False, default="legacy_seed", index=True)
|
||||||
|
source_url = Column(String(500), nullable=True)
|
||||||
|
source_note = Column(Text, nullable=True)
|
||||||
|
raw_payload = Column(JSON, nullable=False, default=dict)
|
||||||
|
needs_confirmation = Column(Boolean, nullable=False, default=True, index=True)
|
||||||
|
verification_status = Column(String(30), nullable=False, default="unverified", index=True)
|
||||||
|
verified_at = Column(DateTime(timezone=True), nullable=True)
|
||||||
|
created_at = Column(DateTime(timezone=True), server_default=func.now())
|
||||||
|
updated_at = Column(DateTime(timezone=True), server_default=func.now(), onupdate=func.now())
|
||||||
|
|
||||||
|
def to_location_dict(self) -> dict:
|
||||||
|
return {
|
||||||
|
"city": self.city,
|
||||||
|
"country": self.country,
|
||||||
|
"latitude": self.latitude,
|
||||||
|
"longitude": self.longitude,
|
||||||
|
"precision": self.precision,
|
||||||
|
"source": self.source,
|
||||||
|
"needs_confirmation": self.needs_confirmation,
|
||||||
|
"matched_location_name": self.site or self.collector_id,
|
||||||
|
"verified_at": to_iso8601_utc(self.verified_at),
|
||||||
|
"confidence": self.confidence,
|
||||||
|
"operator": self.operator,
|
||||||
|
"site": self.site,
|
||||||
|
"verification_status": self.verification_status,
|
||||||
|
"source_note": self.source_note,
|
||||||
|
"source_url": self.source_url,
|
||||||
|
}
|
||||||
@@ -48,6 +48,8 @@ class CollectedData(Base):
|
|||||||
# Indexes for common queries
|
# Indexes for common queries
|
||||||
__table_args__ = (
|
__table_args__ = (
|
||||||
Index("idx_collected_data_source_collected", "source", "collected_at"),
|
Index("idx_collected_data_source_collected", "source", "collected_at"),
|
||||||
|
Index("idx_collected_data_source_current_id", "source", "is_current", "id"),
|
||||||
|
Index("idx_collected_data_source_task_id", "source", "task_id", "id"),
|
||||||
Index("idx_collected_data_source_type", "source", "data_type"),
|
Index("idx_collected_data_source_type", "source", "data_type"),
|
||||||
Index("idx_collected_data_source_source_id", "source", "source_id"),
|
Index("idx_collected_data_source_source_id", "source", "source_id"),
|
||||||
)
|
)
|
||||||
|
|||||||
60
backend/app/models/compute_center_location.py
Normal file
60
backend/app/models/compute_center_location.py
Normal file
@@ -0,0 +1,60 @@
|
|||||||
|
"""Stored compute-center locations."""
|
||||||
|
|
||||||
|
from sqlalchemy import Boolean, Column, DateTime, Float, Integer, JSON, String, Text, UniqueConstraint
|
||||||
|
from sqlalchemy.sql import func
|
||||||
|
|
||||||
|
from app.core.time import to_iso8601_utc
|
||||||
|
from app.db.session import Base
|
||||||
|
|
||||||
|
|
||||||
|
class ComputeCenterLocationRecord(Base):
|
||||||
|
"""Current known location for a compute-center record."""
|
||||||
|
|
||||||
|
__tablename__ = "compute_center_locations"
|
||||||
|
__table_args__ = (
|
||||||
|
UniqueConstraint("source", "source_id", name="uq_compute_center_location_source_id"),
|
||||||
|
)
|
||||||
|
|
||||||
|
id = Column(Integer, primary_key=True, autoincrement=True)
|
||||||
|
source = Column(String(100), nullable=False, index=True)
|
||||||
|
source_id = Column(String(255), nullable=False, index=True)
|
||||||
|
name = Column(String(500), nullable=True)
|
||||||
|
operator = Column(String(255), nullable=True)
|
||||||
|
site = Column(String(255), nullable=True)
|
||||||
|
city = Column(String(255), nullable=True)
|
||||||
|
country = Column(String(255), nullable=True)
|
||||||
|
latitude = Column(Float, nullable=True)
|
||||||
|
longitude = Column(Float, nullable=True)
|
||||||
|
precision = Column(String(30), nullable=False, default="city")
|
||||||
|
confidence = Column(Float, nullable=True)
|
||||||
|
location_source = Column(String(80), nullable=False, default="stored_compute_center_location", index=True)
|
||||||
|
source_url = Column(String(500), nullable=True)
|
||||||
|
source_note = Column(Text, nullable=True)
|
||||||
|
raw_payload = Column(JSON, nullable=False, default=dict)
|
||||||
|
needs_confirmation = Column(Boolean, nullable=False, default=False, index=True)
|
||||||
|
verification_status = Column(String(30), nullable=False, default="verified", index=True)
|
||||||
|
verified_at = Column(DateTime(timezone=True), nullable=True)
|
||||||
|
created_at = Column(DateTime(timezone=True), server_default=func.now())
|
||||||
|
updated_at = Column(DateTime(timezone=True), server_default=func.now(), onupdate=func.now())
|
||||||
|
|
||||||
|
def to_location_dict(self) -> dict:
|
||||||
|
return {
|
||||||
|
"source": self.source,
|
||||||
|
"source_id": self.source_id,
|
||||||
|
"name": self.name,
|
||||||
|
"operator": self.operator,
|
||||||
|
"site": self.site,
|
||||||
|
"city": self.city,
|
||||||
|
"country": self.country,
|
||||||
|
"latitude": self.latitude,
|
||||||
|
"longitude": self.longitude,
|
||||||
|
"precision": self.precision,
|
||||||
|
"confidence": self.confidence,
|
||||||
|
"location_source": self.location_source,
|
||||||
|
"source_url": self.source_url,
|
||||||
|
"source_note": self.source_note,
|
||||||
|
"raw_payload": self.raw_payload or {},
|
||||||
|
"needs_confirmation": self.needs_confirmation,
|
||||||
|
"verification_status": self.verification_status,
|
||||||
|
"verified_at": to_iso8601_utc(self.verified_at),
|
||||||
|
}
|
||||||
32
backend/app/models/datasource_mapping.py
Normal file
32
backend/app/models/datasource_mapping.py
Normal file
@@ -0,0 +1,32 @@
|
|||||||
|
"""Mapping templates for user-defined data source payloads."""
|
||||||
|
|
||||||
|
from sqlalchemy import Boolean, Column, DateTime, ForeignKey, Integer, JSON, String
|
||||||
|
from sqlalchemy.sql import func
|
||||||
|
|
||||||
|
from app.db.session import Base
|
||||||
|
|
||||||
|
|
||||||
|
class DataSourceMappingTemplate(Base):
|
||||||
|
__tablename__ = "datasource_mapping_templates"
|
||||||
|
|
||||||
|
id = Column(Integer, primary_key=True, autoincrement=True)
|
||||||
|
datasource_config_id = Column(
|
||||||
|
Integer,
|
||||||
|
ForeignKey("datasource_configs.id"),
|
||||||
|
nullable=False,
|
||||||
|
index=True,
|
||||||
|
)
|
||||||
|
target_schema = Column(String(80), nullable=False, index=True)
|
||||||
|
mapping_json = Column(JSON, nullable=False, default={})
|
||||||
|
sample_payload_hash = Column(String(64), nullable=True)
|
||||||
|
validation_status = Column(String(30), nullable=False, default="draft")
|
||||||
|
version = Column(Integer, nullable=False, default=1)
|
||||||
|
is_active = Column(Boolean, nullable=False, default=False, index=True)
|
||||||
|
created_at = Column(DateTime(timezone=True), server_default=func.now())
|
||||||
|
updated_at = Column(DateTime(timezone=True), server_default=func.now(), onupdate=func.now())
|
||||||
|
|
||||||
|
def __repr__(self):
|
||||||
|
return (
|
||||||
|
f"<DataSourceMappingTemplate {self.id}: "
|
||||||
|
f"{self.datasource_config_id}/{self.target_schema}/v{self.version}>"
|
||||||
|
)
|
||||||
30
backend/app/models/earth_interactable.py
Normal file
30
backend/app/models/earth_interactable.py
Normal file
@@ -0,0 +1,30 @@
|
|||||||
|
"""Persistent Earth interactable objects."""
|
||||||
|
|
||||||
|
from sqlalchemy import Boolean, Column, DateTime, Float, Index, Integer, JSON, String, Text
|
||||||
|
from sqlalchemy.sql import func
|
||||||
|
|
||||||
|
from app.db.session import Base
|
||||||
|
|
||||||
|
|
||||||
|
class EarthInteractable(Base):
|
||||||
|
__tablename__ = "earth_interactables"
|
||||||
|
|
||||||
|
id = Column(String(160), primary_key=True)
|
||||||
|
layer = Column(String(80), nullable=False, default="interactables", index=True)
|
||||||
|
kind = Column(String(80), nullable=False, default="default", index=True)
|
||||||
|
label = Column(String(255), nullable=False, default="")
|
||||||
|
description = Column(Text, nullable=False, default="")
|
||||||
|
latitude = Column(Float, nullable=False)
|
||||||
|
longitude = Column(Float, nullable=False)
|
||||||
|
altitude = Column(Float, nullable=True)
|
||||||
|
revision = Column(Integer, nullable=False, default=1)
|
||||||
|
properties = Column(JSON, nullable=False, default=dict)
|
||||||
|
is_deleted = Column(Boolean, nullable=False, default=False, index=True)
|
||||||
|
created_at = Column(DateTime(timezone=True), server_default=func.now(), nullable=False)
|
||||||
|
updated_at = Column(DateTime(timezone=True), server_default=func.now(), onupdate=func.now())
|
||||||
|
deleted_at = Column(DateTime(timezone=True), nullable=True, index=True)
|
||||||
|
|
||||||
|
__table_args__ = (
|
||||||
|
Index("idx_earth_interactables_layer_deleted", "layer", "is_deleted"),
|
||||||
|
Index("idx_earth_interactables_updated_at", "updated_at"),
|
||||||
|
)
|
||||||
40
backend/app/models/earth_news.py
Normal file
40
backend/app/models/earth_news.py
Normal file
@@ -0,0 +1,40 @@
|
|||||||
|
from sqlalchemy import Boolean, Column, DateTime, Float, Index, JSON, String, Text
|
||||||
|
from sqlalchemy.sql import func
|
||||||
|
|
||||||
|
from app.db.session import Base
|
||||||
|
|
||||||
|
|
||||||
|
class EarthNewsItem(Base):
|
||||||
|
__tablename__ = "earth_news_items"
|
||||||
|
|
||||||
|
id = Column(String(160), primary_key=True)
|
||||||
|
title = Column(String(500), nullable=False)
|
||||||
|
summary = Column(Text, nullable=False, default="")
|
||||||
|
content_language = Column(String(32), nullable=False, default="en")
|
||||||
|
localizations = Column(JSON, nullable=False, default=dict)
|
||||||
|
url = Column(Text, nullable=False)
|
||||||
|
source = Column(String(255), nullable=False, default="")
|
||||||
|
feed_name = Column(String(255), nullable=False, default="")
|
||||||
|
region = Column(String(80), nullable=False, index=True)
|
||||||
|
homepage_url = Column(Text, nullable=False, default="")
|
||||||
|
published_at = Column(DateTime(timezone=True), nullable=True, index=True)
|
||||||
|
|
||||||
|
latitude = Column(Float, nullable=False)
|
||||||
|
longitude = Column(Float, nullable=False)
|
||||||
|
location_label = Column(String(255), nullable=False)
|
||||||
|
location_source = Column(String(80), nullable=False, default="region_anchor")
|
||||||
|
verified = Column(Boolean, nullable=False, default=False, index=True)
|
||||||
|
location_meta = Column(JSON, nullable=False, default=dict)
|
||||||
|
|
||||||
|
first_seen_at = Column(DateTime(timezone=True), server_default=func.now(), nullable=False)
|
||||||
|
last_seen_at = Column(DateTime(timezone=True), server_default=func.now(), nullable=False, index=True)
|
||||||
|
resolved_at = Column(DateTime(timezone=True), nullable=True, index=True)
|
||||||
|
enrichment_status = Column(String(80), nullable=False, default="pending", index=True)
|
||||||
|
enrichment_error = Column(Text, nullable=True)
|
||||||
|
enriched_at = Column(DateTime(timezone=True), nullable=True, index=True)
|
||||||
|
updated_at = Column(DateTime(timezone=True), server_default=func.now(), onupdate=func.now())
|
||||||
|
|
||||||
|
__table_args__ = (
|
||||||
|
Index("idx_earth_news_region_published", "region", "published_at"),
|
||||||
|
Index("idx_earth_news_region_seen", "region", "last_seen_at"),
|
||||||
|
)
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
"""Collection Task model"""
|
"""Datasource job model."""
|
||||||
|
|
||||||
from sqlalchemy import Column, DateTime, Integer, String, Text, Float
|
from sqlalchemy import BigInteger, Column, DateTime, Float, Integer, JSON, String, Text
|
||||||
from sqlalchemy.sql import func
|
from sqlalchemy.sql import func
|
||||||
|
|
||||||
from app.db.session import Base
|
from app.db.session import Base
|
||||||
@@ -11,14 +11,28 @@ class CollectionTask(Base):
|
|||||||
|
|
||||||
id = Column(Integer, primary_key=True, autoincrement=True)
|
id = Column(Integer, primary_key=True, autoincrement=True)
|
||||||
datasource_id = Column(Integer, nullable=False, index=True)
|
datasource_id = Column(Integer, nullable=False, index=True)
|
||||||
status = Column(String(20), nullable=False) # pending, running, success, failed, cancelled
|
source = Column(String(100), nullable=True, index=True)
|
||||||
|
task_type = Column(String(30), nullable=False, default="collect", index=True)
|
||||||
|
status = Column(String(20), nullable=False) # queued, running, cancelling, success, failed, cancelled
|
||||||
phase = Column(String(30), default="queued")
|
phase = Column(String(30), default="queued")
|
||||||
|
phase_progress = Column(Float)
|
||||||
|
phase_message = Column(String(255))
|
||||||
|
phase_current = Column(BigInteger)
|
||||||
|
phase_total = Column(BigInteger)
|
||||||
|
phase_unit = Column(String(30))
|
||||||
started_at = Column(DateTime(timezone=True))
|
started_at = Column(DateTime(timezone=True))
|
||||||
completed_at = Column(DateTime(timezone=True))
|
completed_at = Column(DateTime(timezone=True))
|
||||||
records_processed = Column(Integer, default=0)
|
records_processed = Column(Integer, default=0)
|
||||||
total_records = Column(Integer, default=0) # Total records to process
|
total_records = Column(Integer, default=0) # Total records to process
|
||||||
progress = Column(Float, default=0.0) # Progress percentage (0-100)
|
progress = Column(Float, default=0.0) # Progress percentage (0-100)
|
||||||
error_message = Column(Text)
|
error_message = Column(Text)
|
||||||
|
payload = Column(JSON, default=dict)
|
||||||
|
rollback_policy = Column(String(40), nullable=False, default="keep_committed_batches")
|
||||||
|
dedupe_key = Column(String(180), nullable=True, index=True)
|
||||||
|
worker_id = Column(String(120), nullable=True, index=True)
|
||||||
|
locked_at = Column(DateTime(timezone=True), nullable=True, index=True)
|
||||||
|
requested_cancel_at = Column(DateTime(timezone=True), nullable=True)
|
||||||
|
cancel_reason = Column(Text)
|
||||||
created_at = Column(DateTime(timezone=True), server_default=func.now())
|
created_at = Column(DateTime(timezone=True), server_default=func.now())
|
||||||
|
|
||||||
def __repr__(self):
|
def __repr__(self):
|
||||||
|
|||||||
@@ -1,4 +1,4 @@
|
|||||||
from sqlalchemy import Boolean, Column, Integer, String, DateTime
|
from sqlalchemy import Boolean, Column, DateTime, Integer, JSON, String
|
||||||
from sqlalchemy.sql import func
|
from sqlalchemy.sql import func
|
||||||
|
|
||||||
from app.db.session import Base
|
from app.db.session import Base
|
||||||
@@ -12,7 +12,10 @@ class User(Base):
|
|||||||
email = Column(String(255), unique=True, index=True, nullable=False)
|
email = Column(String(255), unique=True, index=True, nullable=False)
|
||||||
password_hash = Column(String(255), nullable=False)
|
password_hash = Column(String(255), nullable=False)
|
||||||
role = Column(String(20), default="viewer")
|
role = Column(String(20), default="viewer")
|
||||||
|
gatekeeper_groups = Column(JSON, default=list)
|
||||||
is_active = Column(Boolean, default=True)
|
is_active = Column(Boolean, default=True)
|
||||||
|
email_verified = Column(Boolean, default=False, nullable=False)
|
||||||
|
pending_email = Column(String(255), nullable=True)
|
||||||
last_login_at = Column(DateTime(timezone=True))
|
last_login_at = Column(DateTime(timezone=True))
|
||||||
created_at = Column(DateTime(timezone=True), server_default=func.now())
|
created_at = Column(DateTime(timezone=True), server_default=func.now())
|
||||||
updated_at = Column(
|
updated_at = Column(
|
||||||
|
|||||||
186
backend/app/models/vessel.py
Normal file
186
backend/app/models/vessel.py
Normal file
@@ -0,0 +1,186 @@
|
|||||||
|
"""Vessel AIS models for live maritime tracking."""
|
||||||
|
|
||||||
|
from sqlalchemy import BigInteger, Column, DateTime, Float, Index, Integer, JSON, SmallInteger, String
|
||||||
|
from sqlalchemy.sql import func
|
||||||
|
|
||||||
|
from app.core.time import to_iso8601_utc
|
||||||
|
from app.db.session import Base
|
||||||
|
|
||||||
|
|
||||||
|
class VesselStatic(Base):
|
||||||
|
"""Slow-changing vessel identity and dimensions."""
|
||||||
|
|
||||||
|
__tablename__ = "vessel_static"
|
||||||
|
|
||||||
|
mmsi = Column(BigInteger, primary_key=True)
|
||||||
|
name = Column(String(128), nullable=True)
|
||||||
|
callsign = Column(String(16), nullable=True)
|
||||||
|
vessel_type = Column(SmallInteger, nullable=True, index=True)
|
||||||
|
vessel_type_name = Column(String(64), nullable=True, index=True)
|
||||||
|
flag = Column(String(4), nullable=True, index=True)
|
||||||
|
length = Column(Float, nullable=True)
|
||||||
|
width = Column(Float, nullable=True)
|
||||||
|
draught = Column(Float, nullable=True)
|
||||||
|
imo = Column(BigInteger, nullable=True)
|
||||||
|
updated_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now())
|
||||||
|
|
||||||
|
def to_dict(self) -> dict:
|
||||||
|
return {
|
||||||
|
"mmsi": self.mmsi,
|
||||||
|
"name": self.name,
|
||||||
|
"callsign": self.callsign,
|
||||||
|
"vessel_type": self.vessel_type,
|
||||||
|
"vessel_type_name": self.vessel_type_name,
|
||||||
|
"flag": self.flag,
|
||||||
|
"length": self.length,
|
||||||
|
"width": self.width,
|
||||||
|
"draught": self.draught,
|
||||||
|
"imo": self.imo,
|
||||||
|
"updated_at": to_iso8601_utc(self.updated_at),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class VesselPosition(Base):
|
||||||
|
"""Append-only AIS positions retained for short history windows."""
|
||||||
|
|
||||||
|
__tablename__ = "vessel_position"
|
||||||
|
|
||||||
|
id = Column(Integer, primary_key=True, autoincrement=True)
|
||||||
|
mmsi = Column(BigInteger, nullable=False, index=True)
|
||||||
|
lat = Column(Float, nullable=False)
|
||||||
|
lon = Column(Float, nullable=False)
|
||||||
|
sog = Column(Float, nullable=True)
|
||||||
|
cog = Column(Float, nullable=True)
|
||||||
|
heading = Column(SmallInteger, nullable=True)
|
||||||
|
nav_status = Column(SmallInteger, nullable=True, index=True)
|
||||||
|
received_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now(), index=True)
|
||||||
|
|
||||||
|
__table_args__ = (
|
||||||
|
Index("idx_vessel_pos_mmsi_time", "mmsi", "received_at"),
|
||||||
|
Index("idx_vessel_pos_time", "received_at"),
|
||||||
|
Index("idx_vessel_pos_lat_lon", "lat", "lon"),
|
||||||
|
)
|
||||||
|
|
||||||
|
def to_dict(self) -> dict:
|
||||||
|
return {
|
||||||
|
"id": self.id,
|
||||||
|
"mmsi": self.mmsi,
|
||||||
|
"lat": self.lat,
|
||||||
|
"lon": self.lon,
|
||||||
|
"sog": self.sog,
|
||||||
|
"cog": self.cog,
|
||||||
|
"heading": self.heading,
|
||||||
|
"nav_status": self.nav_status,
|
||||||
|
"received_at": to_iso8601_utc(self.received_at),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class AISRawObservation(Base):
|
||||||
|
"""Source-level AIS fact before aggregation and conflict resolution."""
|
||||||
|
|
||||||
|
__tablename__ = "ais_raw_observations"
|
||||||
|
|
||||||
|
id = Column(Integer, primary_key=True, autoincrement=True)
|
||||||
|
target_schema = Column(String(64), nullable=False, default="vessel_ais", index=True)
|
||||||
|
source = Column(String(100), nullable=False, index=True)
|
||||||
|
entity_key = Column(String(64), nullable=False, index=True)
|
||||||
|
delivery_mode = Column(String(32), nullable=False, index=True)
|
||||||
|
transport = Column(String(32), nullable=False, index=True)
|
||||||
|
message_type = Column(String(64), nullable=True, index=True)
|
||||||
|
source_message_id = Column(String(128), nullable=True, index=True)
|
||||||
|
observation_hash = Column(String(64), nullable=False, unique=True, index=True)
|
||||||
|
observed_at = Column(DateTime(timezone=True), nullable=False, index=True)
|
||||||
|
collected_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now(), index=True)
|
||||||
|
normalized_payload = Column(JSON, default=dict)
|
||||||
|
raw_payload = Column(JSON, default=dict)
|
||||||
|
quality_flags = Column(JSON, default=list)
|
||||||
|
|
||||||
|
__table_args__ = (
|
||||||
|
Index("idx_ais_raw_entity_observed", "target_schema", "entity_key", "observed_at"),
|
||||||
|
Index("idx_ais_raw_schema_observed_entity", "target_schema", "observed_at", "entity_key"),
|
||||||
|
Index("idx_ais_raw_source_entity", "source", "entity_key"),
|
||||||
|
)
|
||||||
|
|
||||||
|
def to_dict(self) -> dict:
|
||||||
|
return {
|
||||||
|
"id": self.id,
|
||||||
|
"target_schema": self.target_schema,
|
||||||
|
"source": self.source,
|
||||||
|
"entity_key": self.entity_key,
|
||||||
|
"delivery_mode": self.delivery_mode,
|
||||||
|
"transport": self.transport,
|
||||||
|
"message_type": self.message_type,
|
||||||
|
"source_message_id": self.source_message_id,
|
||||||
|
"observation_hash": self.observation_hash,
|
||||||
|
"observed_at": to_iso8601_utc(self.observed_at),
|
||||||
|
"collected_at": to_iso8601_utc(self.collected_at),
|
||||||
|
"normalized_payload": self.normalized_payload or {},
|
||||||
|
"raw_payload": self.raw_payload or {},
|
||||||
|
"quality_flags": self.quality_flags or [],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class AISConflictRecord(Base):
|
||||||
|
"""Recorded field-level disagreement between AIS sources."""
|
||||||
|
|
||||||
|
__tablename__ = "ais_conflict_records"
|
||||||
|
|
||||||
|
id = Column(Integer, primary_key=True, autoincrement=True)
|
||||||
|
target_schema = Column(String(64), nullable=False, default="vessel_ais", index=True)
|
||||||
|
entity_key = Column(String(64), nullable=False, index=True)
|
||||||
|
field = Column(String(64), nullable=False, index=True)
|
||||||
|
candidates = Column(JSON, default=dict)
|
||||||
|
selected_source = Column(String(100), nullable=True, index=True)
|
||||||
|
selected_value = Column(JSON, nullable=True)
|
||||||
|
selected_reason = Column(String(64), nullable=True, index=True)
|
||||||
|
resolved_by = Column(String(32), nullable=False, default="system", index=True)
|
||||||
|
status = Column(String(32), nullable=False, default="open", index=True)
|
||||||
|
created_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now(), index=True)
|
||||||
|
updated_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now())
|
||||||
|
|
||||||
|
__table_args__ = (
|
||||||
|
Index("idx_ais_conflict_entity_field", "target_schema", "entity_key", "field"),
|
||||||
|
)
|
||||||
|
|
||||||
|
def to_dict(self) -> dict:
|
||||||
|
return {
|
||||||
|
"id": self.id,
|
||||||
|
"target_schema": self.target_schema,
|
||||||
|
"entity_key": self.entity_key,
|
||||||
|
"field": self.field,
|
||||||
|
"candidates": self.candidates or {},
|
||||||
|
"selected_source": self.selected_source,
|
||||||
|
"selected_value": self.selected_value,
|
||||||
|
"selected_reason": self.selected_reason,
|
||||||
|
"resolved_by": self.resolved_by,
|
||||||
|
"status": self.status,
|
||||||
|
"created_at": to_iso8601_utc(self.created_at),
|
||||||
|
"updated_at": to_iso8601_utc(self.updated_at),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class AISSourceHealth(Base):
|
||||||
|
"""Runtime health signal for an AIS collector source."""
|
||||||
|
|
||||||
|
__tablename__ = "ais_source_health"
|
||||||
|
|
||||||
|
source = Column(String(100), primary_key=True)
|
||||||
|
connection_state = Column(String(32), nullable=False, default="disconnected", index=True)
|
||||||
|
last_seen_at = Column(DateTime(timezone=True), nullable=True, index=True)
|
||||||
|
last_success_at = Column(DateTime(timezone=True), nullable=True, index=True)
|
||||||
|
last_error = Column(String(500), nullable=True)
|
||||||
|
message_rate = Column(Float, nullable=True)
|
||||||
|
lag_seconds = Column(Float, nullable=True)
|
||||||
|
updated_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now(), index=True)
|
||||||
|
|
||||||
|
def to_dict(self) -> dict:
|
||||||
|
return {
|
||||||
|
"source": self.source,
|
||||||
|
"connection_state": self.connection_state,
|
||||||
|
"last_seen_at": to_iso8601_utc(self.last_seen_at),
|
||||||
|
"last_success_at": to_iso8601_utc(self.last_success_at),
|
||||||
|
"last_error": self.last_error,
|
||||||
|
"message_rate": self.message_rate,
|
||||||
|
"lag_seconds": self.lag_seconds,
|
||||||
|
"updated_at": to_iso8601_utc(self.updated_at),
|
||||||
|
}
|
||||||
63
backend/app/models/vessel_enrichment.py
Normal file
63
backend/app/models/vessel_enrichment.py
Normal file
@@ -0,0 +1,63 @@
|
|||||||
|
"""Vessel enrichment cache tables (v5).
|
||||||
|
|
||||||
|
Profile and media enrichment are stored separately so cache TTLs can differ
|
||||||
|
and so the conflict-resolution + display layers can read either independently.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from sqlalchemy import BigInteger, Column, DateTime, Float, JSON, String
|
||||||
|
from sqlalchemy.sql import func
|
||||||
|
|
||||||
|
from app.core.time import to_iso8601_utc
|
||||||
|
from app.db.session import Base
|
||||||
|
|
||||||
|
|
||||||
|
class VesselProfileEnrichment(Base):
|
||||||
|
"""Cached static vessel profile (type, flag, dimensions, operator, etc.)."""
|
||||||
|
|
||||||
|
__tablename__ = "vessel_profile_enrichment"
|
||||||
|
|
||||||
|
mmsi = Column(BigInteger, primary_key=True)
|
||||||
|
source = Column(String(100), nullable=False, default="system")
|
||||||
|
payload = Column(JSON, nullable=False, default=dict)
|
||||||
|
fetched_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now())
|
||||||
|
expires_at = Column(DateTime(timezone=True), nullable=True)
|
||||||
|
confidence = Column(Float, nullable=True)
|
||||||
|
reference_url = Column(String(500), nullable=True)
|
||||||
|
updated_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now(), onupdate=func.now())
|
||||||
|
|
||||||
|
def to_dict(self) -> dict:
|
||||||
|
return {
|
||||||
|
"mmsi": self.mmsi,
|
||||||
|
"source": self.source,
|
||||||
|
"payload": self.payload or {},
|
||||||
|
"fetched_at": to_iso8601_utc(self.fetched_at),
|
||||||
|
"expires_at": to_iso8601_utc(self.expires_at),
|
||||||
|
"confidence": self.confidence,
|
||||||
|
"reference_url": self.reference_url,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class VesselMediaEnrichment(Base):
|
||||||
|
"""Cached vessel imagery / external detail references."""
|
||||||
|
|
||||||
|
__tablename__ = "vessel_media_enrichment"
|
||||||
|
|
||||||
|
mmsi = Column(BigInteger, primary_key=True)
|
||||||
|
source = Column(String(100), nullable=False, default="system")
|
||||||
|
payload = Column(JSON, nullable=False, default=dict)
|
||||||
|
fetched_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now())
|
||||||
|
expires_at = Column(DateTime(timezone=True), nullable=True)
|
||||||
|
confidence = Column(Float, nullable=True)
|
||||||
|
reference_url = Column(String(500), nullable=True)
|
||||||
|
updated_at = Column(DateTime(timezone=True), nullable=False, server_default=func.now(), onupdate=func.now())
|
||||||
|
|
||||||
|
def to_dict(self) -> dict:
|
||||||
|
return {
|
||||||
|
"mmsi": self.mmsi,
|
||||||
|
"source": self.source,
|
||||||
|
"payload": self.payload or {},
|
||||||
|
"fetched_at": to_iso8601_utc(self.fetched_at),
|
||||||
|
"expires_at": to_iso8601_utc(self.expires_at),
|
||||||
|
"confidence": self.confidence,
|
||||||
|
"reference_url": self.reference_url,
|
||||||
|
}
|
||||||
@@ -13,10 +13,11 @@ class AIContentBlock(BaseModel):
|
|||||||
|
|
||||||
class SituationalAnalysisRequest(BaseModel):
|
class SituationalAnalysisRequest(BaseModel):
|
||||||
title: str = Field(..., min_length=1, max_length=200)
|
title: str = Field(..., min_length=1, max_length=200)
|
||||||
objective: str = Field(..., min_length=1, max_length=1000)
|
objective: str = Field(..., min_length=1, max_length=20000)
|
||||||
context: dict[str, Any] = Field(default_factory=dict)
|
context: dict[str, Any] = Field(default_factory=dict)
|
||||||
observations: list[str] = Field(default_factory=list)
|
observations: list[str] = Field(default_factory=list)
|
||||||
constraints: list[str] = Field(default_factory=list)
|
constraints: list[str] = Field(default_factory=list)
|
||||||
|
system_prompt: str | None = Field(default=None, max_length=8000)
|
||||||
preferred_model: str | None = Field(default=None, max_length=200)
|
preferred_model: str | None = Field(default=None, max_length=200)
|
||||||
thinking: dict[str, Any] | None = None
|
thinking: dict[str, Any] | None = None
|
||||||
|
|
||||||
|
|||||||
@@ -12,17 +12,20 @@ class UserBase(BaseModel):
|
|||||||
class UserCreate(UserBase):
|
class UserCreate(UserBase):
|
||||||
password: str = Field(..., min_length=8)
|
password: str = Field(..., min_length=8)
|
||||||
role: str = "viewer"
|
role: str = "viewer"
|
||||||
|
gatekeeper_groups: list[str] = Field(default_factory=list)
|
||||||
|
|
||||||
|
|
||||||
class UserUpdate(BaseModel):
|
class UserUpdate(BaseModel):
|
||||||
email: Optional[EmailStr] = None
|
email: Optional[EmailStr] = None
|
||||||
role: Optional[str] = None
|
role: Optional[str] = None
|
||||||
|
gatekeeper_groups: Optional[list[str]] = None
|
||||||
is_active: Optional[bool] = None
|
is_active: Optional[bool] = None
|
||||||
|
|
||||||
|
|
||||||
class UserInDB(UserBase):
|
class UserInDB(UserBase):
|
||||||
id: int
|
id: int
|
||||||
role: str
|
role: str
|
||||||
|
gatekeeper_groups: list[str] = Field(default_factory=list)
|
||||||
is_active: bool
|
is_active: bool
|
||||||
last_login_at: Optional[datetime]
|
last_login_at: Optional[datetime]
|
||||||
created_at: datetime
|
created_at: datetime
|
||||||
@@ -34,8 +37,36 @@ class UserInDB(UserBase):
|
|||||||
class UserResponse(UserBase):
|
class UserResponse(UserBase):
|
||||||
id: int
|
id: int
|
||||||
role: str
|
role: str
|
||||||
|
gatekeeper_groups: list[str] = Field(default_factory=list)
|
||||||
is_active: bool
|
is_active: bool
|
||||||
|
email_verified: bool = False
|
||||||
created_at: datetime
|
created_at: datetime
|
||||||
|
|
||||||
class Config:
|
class Config:
|
||||||
from_attributes = True
|
from_attributes = True
|
||||||
|
|
||||||
|
|
||||||
|
class UserRegister(BaseModel):
|
||||||
|
username: str = Field(..., min_length=3, max_length=50)
|
||||||
|
email: EmailStr
|
||||||
|
password: str = Field(..., min_length=8, max_length=128)
|
||||||
|
|
||||||
|
|
||||||
|
class VerifyEmailRequest(BaseModel):
|
||||||
|
email: EmailStr
|
||||||
|
code: str = Field(..., min_length=6, max_length=6)
|
||||||
|
|
||||||
|
|
||||||
|
class ResendCodeRequest(BaseModel):
|
||||||
|
email: EmailStr
|
||||||
|
purpose: str = Field(default="register", pattern="^(register|verify_email|reset_password)$")
|
||||||
|
|
||||||
|
|
||||||
|
class ForgotPasswordRequest(BaseModel):
|
||||||
|
email: EmailStr
|
||||||
|
|
||||||
|
|
||||||
|
class ResetPasswordRequest(BaseModel):
|
||||||
|
email: EmailStr
|
||||||
|
code: str = Field(..., min_length=6, max_length=6)
|
||||||
|
new_password: str = Field(..., min_length=8, max_length=128)
|
||||||
|
|||||||
@@ -1,24 +1,49 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import asyncio
|
import asyncio
|
||||||
|
import json
|
||||||
|
from time import perf_counter
|
||||||
|
|
||||||
import httpx
|
import httpx
|
||||||
from fastapi import HTTPException, status
|
from fastapi import Depends, HTTPException, status
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
from app.core.config import settings
|
from app.core.config import settings
|
||||||
|
from app.core.logging import get_logger
|
||||||
|
from app.db.session import get_db
|
||||||
from app.schemas.ai import (
|
from app.schemas.ai import (
|
||||||
AIProviderStatusResponse,
|
AIProviderStatusResponse,
|
||||||
SituationalAnalysisRequest,
|
SituationalAnalysisRequest,
|
||||||
SituationalAnalysisResponse,
|
SituationalAnalysisResponse,
|
||||||
)
|
)
|
||||||
|
from app.services.business_logs import emit_business_log, exception_context
|
||||||
|
|
||||||
|
|
||||||
|
logger = get_logger(__name__, service="ai")
|
||||||
|
|
||||||
|
|
||||||
class AIProviderClient:
|
class AIProviderClient:
|
||||||
def __init__(self) -> None:
|
def __init__(
|
||||||
self.service_url = settings.AI_PROVIDER_SERVICE_URL.rstrip("/")
|
self,
|
||||||
self.service_token = settings.AI_PROVIDER_SERVICE_TOKEN
|
*,
|
||||||
self.timeout = settings.AI_PROVIDER_TIMEOUT_SECONDS
|
service_url: str | None = None,
|
||||||
self.retry_attempts = max(settings.AI_PROVIDER_RETRY_ATTEMPTS, 1)
|
service_token: str | None = None,
|
||||||
|
timeout: int | None = None,
|
||||||
|
retry_attempts: int | None = None,
|
||||||
|
llm_config: dict | None = None,
|
||||||
|
) -> None:
|
||||||
|
self.service_url = (
|
||||||
|
service_url if service_url is not None else settings.AI_PROVIDER_SERVICE_URL
|
||||||
|
).rstrip("/")
|
||||||
|
self.service_token = (
|
||||||
|
service_token if service_token is not None else settings.AI_PROVIDER_SERVICE_TOKEN
|
||||||
|
)
|
||||||
|
self.timeout = timeout if timeout is not None else settings.AI_PROVIDER_TIMEOUT_SECONDS
|
||||||
|
self.retry_attempts = max(
|
||||||
|
retry_attempts if retry_attempts is not None else settings.AI_PROVIDER_RETRY_ATTEMPTS,
|
||||||
|
1,
|
||||||
|
)
|
||||||
|
self.llm_config = llm_config or {}
|
||||||
|
|
||||||
def _headers(self, request_id: str | None = None) -> dict[str, str]:
|
def _headers(self, request_id: str | None = None) -> dict[str, str]:
|
||||||
headers = {"Content-Type": "application/json"}
|
headers = {"Content-Type": "application/json"}
|
||||||
@@ -26,10 +51,38 @@ class AIProviderClient:
|
|||||||
headers["X-Provider-Token"] = self.service_token
|
headers["X-Provider-Token"] = self.service_token
|
||||||
if request_id:
|
if request_id:
|
||||||
headers["X-Request-ID"] = request_id
|
headers["X-Request-ID"] = request_id
|
||||||
|
llm_header_map = {
|
||||||
|
"provider": "X-AI-Provider",
|
||||||
|
"provider_api": "X-AI-Provider-API",
|
||||||
|
"base_url": "X-AI-Base-URL",
|
||||||
|
"api_key": "X-AI-API-Key",
|
||||||
|
"model": "X-AI-Model",
|
||||||
|
"max_tokens": "X-AI-Max-Tokens",
|
||||||
|
"anthropic_version": "X-AI-Anthropic-Version",
|
||||||
|
}
|
||||||
|
for key, header_name in llm_header_map.items():
|
||||||
|
value = self.llm_config.get(key)
|
||||||
|
if value not in (None, ""):
|
||||||
|
headers[header_name] = str(value)
|
||||||
|
model_provider_apis = self.llm_config.get("model_provider_apis")
|
||||||
|
if isinstance(model_provider_apis, dict) and model_provider_apis:
|
||||||
|
headers["X-AI-Model-Provider-APIs"] = json.dumps(model_provider_apis)
|
||||||
return headers
|
return headers
|
||||||
|
|
||||||
async def get_status(self, request_id: str | None = None) -> AIProviderStatusResponse:
|
async def get_status(self, request_id: str | None = None) -> AIProviderStatusResponse:
|
||||||
|
context = self._base_log_context(operation="status")
|
||||||
if not self.service_url:
|
if not self.service_url:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.provider.status.failed",
|
||||||
|
message="AI provider status skipped because service URL is not configured",
|
||||||
|
category="ai",
|
||||||
|
level="warning",
|
||||||
|
service="ai",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
context={**context, "status": "unconfigured"},
|
||||||
|
)
|
||||||
return AIProviderStatusResponse(
|
return AIProviderStatusResponse(
|
||||||
provider="unconfigured",
|
provider="unconfigured",
|
||||||
enabled=False,
|
enabled=False,
|
||||||
@@ -38,27 +91,133 @@ class AIProviderClient:
|
|||||||
base_url=None,
|
base_url=None,
|
||||||
)
|
)
|
||||||
|
|
||||||
data = await self._request("GET", "/v1/provider/status", request_id=request_id)
|
started_at = perf_counter()
|
||||||
return AIProviderStatusResponse.model_validate(data)
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.provider.status.start",
|
||||||
|
message="AI provider status request started",
|
||||||
|
category="ai",
|
||||||
|
service="ai",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
context=context,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
data = await self._request("GET", "/v1/provider/status", request_id=request_id, operation="status")
|
||||||
|
result = AIProviderStatusResponse.model_validate(data)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.provider.status.success",
|
||||||
|
message="AI provider status request completed",
|
||||||
|
category="ai",
|
||||||
|
service="ai",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
context={
|
||||||
|
**context,
|
||||||
|
"status": "success",
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
"result_provider": result.provider,
|
||||||
|
"result_model": result.model,
|
||||||
|
"configured": result.configured,
|
||||||
|
"enabled": result.enabled,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return result
|
||||||
|
except Exception as exc:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.provider.status.failed",
|
||||||
|
message="AI provider status request failed",
|
||||||
|
category="ai",
|
||||||
|
level="error",
|
||||||
|
service="ai",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
context=exception_context(exc, {**context, "status": "failed", "duration_ms": self._duration_ms(started_at)}),
|
||||||
|
)
|
||||||
|
raise
|
||||||
|
|
||||||
async def analyze(
|
async def analyze(
|
||||||
self,
|
self,
|
||||||
payload: SituationalAnalysisRequest,
|
payload: SituationalAnalysisRequest,
|
||||||
request_id: str | None = None,
|
request_id: str | None = None,
|
||||||
) -> SituationalAnalysisResponse:
|
) -> SituationalAnalysisResponse:
|
||||||
|
context = self._base_log_context(
|
||||||
|
operation="analyze",
|
||||||
|
preferred_model=payload.preferred_model,
|
||||||
|
input_summary=self._summarize_analysis_payload(payload),
|
||||||
|
)
|
||||||
if not self.service_url:
|
if not self.service_url:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.provider.analyze.failed",
|
||||||
|
message="AI provider analyze skipped because service URL is not configured",
|
||||||
|
category="ai",
|
||||||
|
level="warning",
|
||||||
|
service="ai",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
context={**context, "status": "unconfigured"},
|
||||||
|
)
|
||||||
raise HTTPException(
|
raise HTTPException(
|
||||||
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
|
status_code=status.HTTP_503_SERVICE_UNAVAILABLE,
|
||||||
detail="AI provider service URL is not configured.",
|
detail="AI provider service URL is not configured.",
|
||||||
)
|
)
|
||||||
|
|
||||||
data = await self._request(
|
started_at = perf_counter()
|
||||||
"POST",
|
await emit_business_log(
|
||||||
"/v1/analyze",
|
logger,
|
||||||
json=payload.model_dump(),
|
event="ai.provider.analyze.start",
|
||||||
|
message="AI provider analyze request started",
|
||||||
|
category="ai",
|
||||||
|
service="ai",
|
||||||
|
module=__name__,
|
||||||
request_id=request_id,
|
request_id=request_id,
|
||||||
|
context=context,
|
||||||
)
|
)
|
||||||
return SituationalAnalysisResponse.model_validate(data)
|
try:
|
||||||
|
data = await self._request(
|
||||||
|
"POST",
|
||||||
|
"/v1/analyze",
|
||||||
|
json=payload.model_dump(),
|
||||||
|
request_id=request_id,
|
||||||
|
operation="analyze",
|
||||||
|
payload_summary=context["input_summary"],
|
||||||
|
)
|
||||||
|
result = SituationalAnalysisResponse.model_validate(data)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.provider.analyze.success",
|
||||||
|
message="AI provider analyze request completed",
|
||||||
|
category="ai",
|
||||||
|
service="ai",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
context={
|
||||||
|
**context,
|
||||||
|
"status": "success",
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
"result_provider": result.provider,
|
||||||
|
"result_model": result.model,
|
||||||
|
"content_block_count": len(result.content_blocks or []),
|
||||||
|
"thinking_block_count": len(result.thinking_blocks or []),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return result
|
||||||
|
except Exception as exc:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai.provider.analyze.failed",
|
||||||
|
message="AI provider analyze request failed",
|
||||||
|
category="ai",
|
||||||
|
level="error",
|
||||||
|
service="ai",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
context=exception_context(exc, {**context, "status": "failed", "duration_ms": self._duration_ms(started_at)}),
|
||||||
|
)
|
||||||
|
raise
|
||||||
|
|
||||||
async def _request(
|
async def _request(
|
||||||
self,
|
self,
|
||||||
@@ -66,9 +225,12 @@ class AIProviderClient:
|
|||||||
path: str,
|
path: str,
|
||||||
json: dict | None = None,
|
json: dict | None = None,
|
||||||
request_id: str | None = None,
|
request_id: str | None = None,
|
||||||
|
operation: str = "request",
|
||||||
|
payload_summary: dict | None = None,
|
||||||
) -> dict:
|
) -> dict:
|
||||||
last_error: Exception | None = None
|
last_error: Exception | None = None
|
||||||
for attempt in range(1, self.retry_attempts + 1):
|
for attempt in range(1, self.retry_attempts + 1):
|
||||||
|
attempt_started_at = perf_counter()
|
||||||
try:
|
try:
|
||||||
async with httpx.AsyncClient(timeout=self.timeout) as client:
|
async with httpx.AsyncClient(timeout=self.timeout) as client:
|
||||||
response = await client.request(
|
response = await client.request(
|
||||||
@@ -82,6 +244,15 @@ class AIProviderClient:
|
|||||||
except httpx.HTTPStatusError as exc:
|
except httpx.HTTPStatusError as exc:
|
||||||
last_error = exc
|
last_error = exc
|
||||||
if attempt < self.retry_attempts and exc.response.status_code >= 500:
|
if attempt < self.retry_attempts and exc.response.status_code >= 500:
|
||||||
|
await self._log_retry(
|
||||||
|
operation=operation,
|
||||||
|
request_id=request_id,
|
||||||
|
attempt=attempt,
|
||||||
|
status_code=exc.response.status_code,
|
||||||
|
duration_ms=self._duration_ms(attempt_started_at),
|
||||||
|
error=exc,
|
||||||
|
payload_summary=payload_summary,
|
||||||
|
)
|
||||||
await asyncio.sleep(0.3 * attempt)
|
await asyncio.sleep(0.3 * attempt)
|
||||||
continue
|
continue
|
||||||
detail = exc.response.text or "AI provider service returned an error"
|
detail = exc.response.text or "AI provider service returned an error"
|
||||||
@@ -92,6 +263,14 @@ class AIProviderClient:
|
|||||||
except httpx.HTTPError as exc:
|
except httpx.HTTPError as exc:
|
||||||
last_error = exc
|
last_error = exc
|
||||||
if attempt < self.retry_attempts:
|
if attempt < self.retry_attempts:
|
||||||
|
await self._log_retry(
|
||||||
|
operation=operation,
|
||||||
|
request_id=request_id,
|
||||||
|
attempt=attempt,
|
||||||
|
duration_ms=self._duration_ms(attempt_started_at),
|
||||||
|
error=exc,
|
||||||
|
payload_summary=payload_summary,
|
||||||
|
)
|
||||||
await asyncio.sleep(0.3 * attempt)
|
await asyncio.sleep(0.3 * attempt)
|
||||||
continue
|
continue
|
||||||
raise HTTPException(
|
raise HTTPException(
|
||||||
@@ -104,6 +283,80 @@ class AIProviderClient:
|
|||||||
detail=f"AI provider service request failed: {last_error}",
|
detail=f"AI provider service request failed: {last_error}",
|
||||||
)
|
)
|
||||||
|
|
||||||
|
def _base_log_context(self, **extra: object) -> dict:
|
||||||
|
llm_provider_apis = self.llm_config.get("model_provider_apis")
|
||||||
|
return {
|
||||||
|
"provider": self.llm_config.get("provider") or "",
|
||||||
|
"provider_api": self.llm_config.get("provider_api") or "",
|
||||||
|
"model": self.llm_config.get("model") or "",
|
||||||
|
"base_url_configured": bool(self.llm_config.get("base_url")),
|
||||||
|
"service_url_configured": bool(self.service_url),
|
||||||
|
"timeout_seconds": self.timeout,
|
||||||
|
"retry_attempts": self.retry_attempts,
|
||||||
|
"model_provider_api_count": len(llm_provider_apis or {}) if isinstance(llm_provider_apis, dict) else 0,
|
||||||
|
**extra,
|
||||||
|
}
|
||||||
|
|
||||||
def get_ai_provider_client() -> AIProviderClient:
|
@staticmethod
|
||||||
return AIProviderClient()
|
def _duration_ms(started_at: float) -> int:
|
||||||
|
return int((perf_counter() - started_at) * 1000)
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _summarize_analysis_payload(payload: SituationalAnalysisRequest) -> dict:
|
||||||
|
context = payload.context if isinstance(payload.context, dict) else {}
|
||||||
|
thinking = payload.thinking if isinstance(payload.thinking, dict) else payload.thinking
|
||||||
|
return {
|
||||||
|
"title_length": len(payload.title or ""),
|
||||||
|
"objective_length": len(payload.objective or ""),
|
||||||
|
"observation_count": len(payload.observations or []),
|
||||||
|
"constraint_count": len(payload.constraints or []),
|
||||||
|
"has_system_prompt": bool(payload.system_prompt),
|
||||||
|
"thinking_enabled": bool(thinking),
|
||||||
|
"context_keys": sorted(str(key) for key in context.keys()),
|
||||||
|
}
|
||||||
|
|
||||||
|
async def _log_retry(
|
||||||
|
self,
|
||||||
|
*,
|
||||||
|
operation: str,
|
||||||
|
request_id: str | None,
|
||||||
|
attempt: int,
|
||||||
|
duration_ms: int,
|
||||||
|
error: BaseException,
|
||||||
|
status_code: int | None = None,
|
||||||
|
payload_summary: dict | None = None,
|
||||||
|
) -> None:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event=f"ai.provider.{operation}.retry",
|
||||||
|
message="AI provider request will retry",
|
||||||
|
category="ai",
|
||||||
|
level="warning",
|
||||||
|
service="ai",
|
||||||
|
module=__name__,
|
||||||
|
request_id=request_id,
|
||||||
|
context=exception_context(
|
||||||
|
error,
|
||||||
|
{
|
||||||
|
**self._base_log_context(operation=operation),
|
||||||
|
"attempt": attempt,
|
||||||
|
"next_attempt": attempt + 1,
|
||||||
|
"status_code": status_code,
|
||||||
|
"duration_ms": duration_ms,
|
||||||
|
"input_summary": payload_summary,
|
||||||
|
},
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def get_ai_provider_client(db: AsyncSession = Depends(get_db)) -> AIProviderClient:
|
||||||
|
from app.api.v1.settings import get_runtime_ai_provider_config
|
||||||
|
|
||||||
|
runtime_config = await get_runtime_ai_provider_config(db)
|
||||||
|
return AIProviderClient(
|
||||||
|
service_url=runtime_config["service_url"],
|
||||||
|
service_token=runtime_config["service_token"],
|
||||||
|
timeout=runtime_config["timeout_seconds"],
|
||||||
|
retry_attempts=runtime_config["retry_attempts"],
|
||||||
|
llm_config=runtime_config.get("llm_config") or {},
|
||||||
|
)
|
||||||
|
|||||||
7
backend/app/services/ai_tools/__init__.py
Normal file
7
backend/app/services/ai_tools/__init__.py
Normal file
@@ -0,0 +1,7 @@
|
|||||||
|
"""Backend-owned AI tool services.
|
||||||
|
|
||||||
|
The services in this package are business tools used by Planet's backend
|
||||||
|
orchestrators. They intentionally live outside ``aiprovider`` so model transport
|
||||||
|
stays separate from evidence collection and domain policy.
|
||||||
|
"""
|
||||||
|
|
||||||
48
backend/app/services/ai_tools/evidence_store.py
Normal file
48
backend/app/services/ai_tools/evidence_store.py
Normal file
@@ -0,0 +1,48 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import hashlib
|
||||||
|
from typing import Iterable
|
||||||
|
|
||||||
|
from app.services.ai_tools.schemas import FetchedEvidence, SearchEvidence
|
||||||
|
|
||||||
|
|
||||||
|
def evidence_content_hash(text: str) -> str:
|
||||||
|
return hashlib.sha256(text.encode("utf-8")).hexdigest()
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_search_evidence(items: Iterable[SearchEvidence], *, limit: int = 5) -> list[dict]:
|
||||||
|
normalized: list[dict] = []
|
||||||
|
seen_urls: set[str] = set()
|
||||||
|
for item in items:
|
||||||
|
if not item.url or item.url in seen_urls:
|
||||||
|
continue
|
||||||
|
seen_urls.add(item.url)
|
||||||
|
normalized.append(
|
||||||
|
{
|
||||||
|
"title": item.title,
|
||||||
|
"url": item.url,
|
||||||
|
"snippet": item.snippet,
|
||||||
|
"content": item.compact_text(),
|
||||||
|
"score": item.score,
|
||||||
|
"source_provider": item.source_provider,
|
||||||
|
"retrieved_at": item.retrieved_at.isoformat(),
|
||||||
|
"metadata": item.metadata,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
if len(normalized) >= limit:
|
||||||
|
break
|
||||||
|
return normalized
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_fetched_evidence(item: FetchedEvidence, *, text_limit: int = 1200) -> dict:
|
||||||
|
text = " ".join(item.text.split())[:text_limit]
|
||||||
|
return {
|
||||||
|
"title": item.title,
|
||||||
|
"url": item.final_url or item.url,
|
||||||
|
"text": text,
|
||||||
|
"content_hash": item.content_hash or evidence_content_hash(item.text),
|
||||||
|
"extractor": item.extractor,
|
||||||
|
"retrieved_at": item.retrieved_at.isoformat(),
|
||||||
|
"metadata": item.metadata,
|
||||||
|
}
|
||||||
|
|
||||||
63
backend/app/services/ai_tools/schemas.py
Normal file
63
backend/app/services/ai_tools/schemas.py
Normal file
@@ -0,0 +1,63 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from pydantic import BaseModel, Field
|
||||||
|
|
||||||
|
|
||||||
|
class SearchEvidence(BaseModel):
|
||||||
|
title: str = ""
|
||||||
|
url: str = ""
|
||||||
|
snippet: str = ""
|
||||||
|
content: str = ""
|
||||||
|
score: float | None = None
|
||||||
|
source_provider: str = ""
|
||||||
|
retrieved_at: datetime = Field(default_factory=lambda: datetime.now(UTC))
|
||||||
|
metadata: dict[str, Any] = Field(default_factory=dict)
|
||||||
|
|
||||||
|
def compact_text(self, limit: int = 700) -> str:
|
||||||
|
text = " ".join((self.content or self.snippet or "").split())
|
||||||
|
return text[:limit]
|
||||||
|
|
||||||
|
|
||||||
|
class FetchedEvidence(BaseModel):
|
||||||
|
url: str
|
||||||
|
final_url: str = ""
|
||||||
|
title: str = ""
|
||||||
|
text: str = ""
|
||||||
|
content_hash: str = ""
|
||||||
|
extractor: str = "basic_html"
|
||||||
|
retrieved_at: datetime = Field(default_factory=lambda: datetime.now(UTC))
|
||||||
|
metadata: dict[str, Any] = Field(default_factory=dict)
|
||||||
|
|
||||||
|
|
||||||
|
class WebSearchProviderConfig(BaseModel):
|
||||||
|
provider: str = "tavily"
|
||||||
|
base_url: str = ""
|
||||||
|
api_key: str = ""
|
||||||
|
max_results: int = Field(default=5, ge=1, le=20)
|
||||||
|
timeout_seconds: int = Field(default=20, ge=3, le=120)
|
||||||
|
endpoint_path: str = ""
|
||||||
|
search_depth: str = "basic"
|
||||||
|
engine: str = "google"
|
||||||
|
include_answer: bool = False
|
||||||
|
include_raw_content: bool = False
|
||||||
|
include_text: bool = False
|
||||||
|
categories: str = "general"
|
||||||
|
engines: list[str] = Field(default_factory=list)
|
||||||
|
search_path: str = ""
|
||||||
|
scrape_path: str = ""
|
||||||
|
scrape_formats: list[str] = Field(default_factory=lambda: ["markdown"])
|
||||||
|
|
||||||
|
|
||||||
|
class WebSearchConfig(BaseModel):
|
||||||
|
enabled: bool = False
|
||||||
|
default_provider: str = "tavily"
|
||||||
|
provider: str = "tavily"
|
||||||
|
providers: dict[str, WebSearchProviderConfig] = Field(default_factory=dict)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def active_provider_config(self) -> WebSearchProviderConfig:
|
||||||
|
return self.providers.get(self.default_provider) or self.providers.get(self.provider) or WebSearchProviderConfig(provider=self.default_provider or self.provider)
|
||||||
|
|
||||||
122
backend/app/services/ai_tools/web_fetch.py
Normal file
122
backend/app/services/ai_tools/web_fetch.py
Normal file
@@ -0,0 +1,122 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import hashlib
|
||||||
|
from time import perf_counter
|
||||||
|
from urllib.parse import urlparse
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
from bs4 import BeautifulSoup
|
||||||
|
|
||||||
|
from app.core.logging import get_logger
|
||||||
|
from app.services.ai_tools.schemas import FetchedEvidence
|
||||||
|
from app.services.business_logs import emit_business_log, exception_context
|
||||||
|
|
||||||
|
|
||||||
|
logger = get_logger(__name__, service="ai_tool")
|
||||||
|
|
||||||
|
|
||||||
|
class WebFetchError(RuntimeError):
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def _extract_title_and_text(html: str) -> tuple[str, str]:
|
||||||
|
soup = BeautifulSoup(html, "html.parser")
|
||||||
|
for tag in soup(["script", "style", "noscript", "svg"]):
|
||||||
|
tag.decompose()
|
||||||
|
title = soup.title.get_text(" ", strip=True) if soup.title else ""
|
||||||
|
main = soup.find("main") or soup.find("article") or soup.body or soup
|
||||||
|
text = main.get_text("\n", strip=True)
|
||||||
|
lines = [line.strip() for line in text.splitlines() if line.strip()]
|
||||||
|
return title, "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
|
async def fetch_url_evidence(
|
||||||
|
url: str,
|
||||||
|
*,
|
||||||
|
timeout_seconds: int = 20,
|
||||||
|
max_bytes: int = 1_500_000,
|
||||||
|
) -> FetchedEvidence:
|
||||||
|
started_at = perf_counter()
|
||||||
|
if not url:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai_tool.web_fetch.failed",
|
||||||
|
message="WebFetch failed because URL is empty",
|
||||||
|
category="ai_tool",
|
||||||
|
level="warning",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context={"reason": "empty_url"},
|
||||||
|
)
|
||||||
|
raise WebFetchError("url is required")
|
||||||
|
request_host = urlparse(url).netloc
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai_tool.web_fetch.start",
|
||||||
|
message="WebFetch request started",
|
||||||
|
category="ai_tool",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"url_host": request_host,
|
||||||
|
"timeout_seconds": timeout_seconds,
|
||||||
|
"max_bytes": max_bytes,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
async with httpx.AsyncClient(
|
||||||
|
timeout=timeout_seconds,
|
||||||
|
follow_redirects=True,
|
||||||
|
headers={"User-Agent": "PlanetEvidenceFetcher/1.0"},
|
||||||
|
) as client:
|
||||||
|
response = await client.get(url)
|
||||||
|
response.raise_for_status()
|
||||||
|
content = response.content[:max_bytes]
|
||||||
|
except httpx.HTTPError as exc:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai_tool.web_fetch.failed",
|
||||||
|
message="WebFetch request failed",
|
||||||
|
category="ai_tool",
|
||||||
|
level="error",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context=exception_context(
|
||||||
|
exc,
|
||||||
|
{
|
||||||
|
"url_host": request_host,
|
||||||
|
"status": "failed",
|
||||||
|
"duration_ms": int((perf_counter() - started_at) * 1000),
|
||||||
|
},
|
||||||
|
),
|
||||||
|
)
|
||||||
|
raise WebFetchError(f"failed to fetch page: {exc}") from exc
|
||||||
|
|
||||||
|
title, text = _extract_title_and_text(content.decode(response.encoding or "utf-8", errors="ignore"))
|
||||||
|
content_hash = hashlib.sha256(text.encode("utf-8")).hexdigest()
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai_tool.web_fetch.success",
|
||||||
|
message="WebFetch request completed",
|
||||||
|
category="ai_tool",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"url_host": request_host,
|
||||||
|
"final_url_host": urlparse(str(response.url)).netloc,
|
||||||
|
"status": "success",
|
||||||
|
"status_code": response.status_code,
|
||||||
|
"bytes_read": len(content),
|
||||||
|
"content_hash": content_hash,
|
||||||
|
"duration_ms": int((perf_counter() - started_at) * 1000),
|
||||||
|
"extractor": "beautifulsoup_basic",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return FetchedEvidence(
|
||||||
|
url=url,
|
||||||
|
final_url=str(response.url),
|
||||||
|
title=title,
|
||||||
|
text=text,
|
||||||
|
content_hash=content_hash,
|
||||||
|
extractor="beautifulsoup_basic",
|
||||||
|
)
|
||||||
476
backend/app/services/ai_tools/web_search.py
Normal file
476
backend/app/services/ai_tools/web_search.py
Normal file
@@ -0,0 +1,476 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from copy import deepcopy
|
||||||
|
import hashlib
|
||||||
|
from time import perf_counter
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
|
||||||
|
from app.core.logging import get_logger
|
||||||
|
from app.services.ai_tools.schemas import SearchEvidence, WebSearchConfig, WebSearchProviderConfig
|
||||||
|
from app.services.business_logs import emit_business_log, exception_context
|
||||||
|
|
||||||
|
|
||||||
|
logger = get_logger(__name__, service="ai_tool")
|
||||||
|
|
||||||
|
|
||||||
|
WEB_SEARCH_PROVIDER_PRESETS: dict[str, dict[str, Any]] = {
|
||||||
|
"tavily": {
|
||||||
|
"provider": "tavily",
|
||||||
|
"label": "Tavily",
|
||||||
|
"api_key_env": "TAVILY_API_KEY",
|
||||||
|
"base_url": "https://api.tavily.com",
|
||||||
|
"endpoint_path": "/search",
|
||||||
|
"max_results": 5,
|
||||||
|
"timeout_seconds": 20,
|
||||||
|
"search_depth": "basic",
|
||||||
|
"include_answer": False,
|
||||||
|
"include_raw_content": False,
|
||||||
|
},
|
||||||
|
"brave": {
|
||||||
|
"provider": "brave",
|
||||||
|
"label": "Brave Search API",
|
||||||
|
"api_key_env": "BRAVE_SEARCH_API_KEY",
|
||||||
|
"base_url": "https://api.search.brave.com",
|
||||||
|
"endpoint_path": "/res/v1/web/search",
|
||||||
|
"max_results": 5,
|
||||||
|
"timeout_seconds": 20,
|
||||||
|
},
|
||||||
|
"serpapi": {
|
||||||
|
"provider": "serpapi",
|
||||||
|
"label": "SerpAPI",
|
||||||
|
"api_key_env": "SERPAPI_API_KEY",
|
||||||
|
"base_url": "https://serpapi.com",
|
||||||
|
"endpoint_path": "/search.json",
|
||||||
|
"engine": "google",
|
||||||
|
"max_results": 5,
|
||||||
|
"timeout_seconds": 20,
|
||||||
|
},
|
||||||
|
"exa": {
|
||||||
|
"provider": "exa",
|
||||||
|
"label": "Exa",
|
||||||
|
"api_key_env": "EXA_API_KEY",
|
||||||
|
"base_url": "https://api.exa.ai",
|
||||||
|
"endpoint_path": "/search",
|
||||||
|
"max_results": 5,
|
||||||
|
"timeout_seconds": 20,
|
||||||
|
"include_text": False,
|
||||||
|
},
|
||||||
|
"firecrawl": {
|
||||||
|
"provider": "firecrawl",
|
||||||
|
"label": "Firecrawl Search / Scrape",
|
||||||
|
"api_key_env": "FIRECRAWL_API_KEY",
|
||||||
|
"base_url": "https://api.firecrawl.dev",
|
||||||
|
"search_path": "/v2/search",
|
||||||
|
"scrape_path": "/v2/scrape",
|
||||||
|
"max_results": 5,
|
||||||
|
"timeout_seconds": 30,
|
||||||
|
"scrape_formats": ["markdown"],
|
||||||
|
},
|
||||||
|
"searxng": {
|
||||||
|
"provider": "searxng",
|
||||||
|
"label": "SearXNG",
|
||||||
|
"api_key_env": "SEARXNG_API_KEY",
|
||||||
|
"base_url": "http://localhost:8080",
|
||||||
|
"endpoint_path": "/",
|
||||||
|
"max_results": 5,
|
||||||
|
"timeout_seconds": 20,
|
||||||
|
"categories": "general",
|
||||||
|
"engines": [],
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class WebSearchError(RuntimeError):
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
class WebSearchConfigurationError(WebSearchError):
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_web_search_provider(provider: str | None) -> str:
|
||||||
|
return (provider or "tavily").strip().lower() or "tavily"
|
||||||
|
|
||||||
|
|
||||||
|
def get_web_search_provider_preset(provider: str) -> dict[str, Any]:
|
||||||
|
provider_id = normalize_web_search_provider(provider)
|
||||||
|
preset = WEB_SEARCH_PROVIDER_PRESETS.get(provider_id)
|
||||||
|
if not preset:
|
||||||
|
raise ValueError(f"Unsupported web search provider: {provider}")
|
||||||
|
return deepcopy(preset)
|
||||||
|
|
||||||
|
|
||||||
|
def list_web_search_provider_presets() -> list[dict[str, Any]]:
|
||||||
|
return [get_web_search_provider_preset(provider) for provider in WEB_SEARCH_PROVIDER_PRESETS]
|
||||||
|
|
||||||
|
|
||||||
|
def provider_defaults(provider: str) -> WebSearchProviderConfig:
|
||||||
|
preset = get_web_search_provider_preset(provider)
|
||||||
|
return WebSearchProviderConfig(**{
|
||||||
|
key: value
|
||||||
|
for key, value in preset.items()
|
||||||
|
if key in WebSearchProviderConfig.model_fields
|
||||||
|
})
|
||||||
|
|
||||||
|
|
||||||
|
class WebSearchClient:
|
||||||
|
def __init__(self, config: WebSearchConfig) -> None:
|
||||||
|
self.config = config
|
||||||
|
|
||||||
|
async def search(
|
||||||
|
self,
|
||||||
|
query: str,
|
||||||
|
*,
|
||||||
|
max_results: int | None = None,
|
||||||
|
domains: list[str] | None = None,
|
||||||
|
freshness_days: int | None = None,
|
||||||
|
) -> list[SearchEvidence]:
|
||||||
|
started_at = perf_counter()
|
||||||
|
if not self.config.enabled:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai_tool.web_search.unavailable",
|
||||||
|
message="WebSearch skipped because integration is disabled",
|
||||||
|
category="ai_tool",
|
||||||
|
level="warning",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context={"provider": self.config.default_provider, "reason": "disabled"},
|
||||||
|
)
|
||||||
|
raise WebSearchConfigurationError("WebSearch is disabled.")
|
||||||
|
provider_config = self.config.active_provider_config
|
||||||
|
provider = normalize_web_search_provider(provider_config.provider)
|
||||||
|
if provider != "searxng" and not provider_config.api_key:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai_tool.web_search.unavailable",
|
||||||
|
message="WebSearch skipped because API key is not configured",
|
||||||
|
category="ai_tool",
|
||||||
|
level="warning",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context={"provider": provider, "reason": "missing_api_key"},
|
||||||
|
)
|
||||||
|
raise WebSearchConfigurationError(f"{provider} API key is not configured.")
|
||||||
|
query = " ".join(str(query or "").split())
|
||||||
|
if not query:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai_tool.web_search.failed",
|
||||||
|
message="WebSearch failed because query is empty",
|
||||||
|
category="ai_tool",
|
||||||
|
level="warning",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context={"provider": provider, "reason": "empty_query"},
|
||||||
|
)
|
||||||
|
raise WebSearchConfigurationError("search query is required.")
|
||||||
|
limit = max_results or provider_config.max_results
|
||||||
|
context = {
|
||||||
|
"provider": provider,
|
||||||
|
"query_hash": hashlib.sha256(query.encode("utf-8")).hexdigest(),
|
||||||
|
"query_length": len(query),
|
||||||
|
"max_results": limit,
|
||||||
|
"domain_count": len(domains or []),
|
||||||
|
"freshness_days": freshness_days,
|
||||||
|
}
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai_tool.web_search.start",
|
||||||
|
message="WebSearch request started",
|
||||||
|
category="ai_tool",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context=context,
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
if provider == "tavily":
|
||||||
|
results = await self._search_tavily(provider_config, query, limit, domains, freshness_days)
|
||||||
|
elif provider == "brave":
|
||||||
|
results = await self._search_brave(provider_config, query, limit, domains)
|
||||||
|
elif provider == "serpapi":
|
||||||
|
results = await self._search_serpapi(provider_config, query, limit)
|
||||||
|
elif provider == "exa":
|
||||||
|
results = await self._search_exa(provider_config, query, limit, domains)
|
||||||
|
elif provider == "firecrawl":
|
||||||
|
results = await self._search_firecrawl(provider_config, query, limit)
|
||||||
|
elif provider == "searxng":
|
||||||
|
results = await self._search_searxng(provider_config, query, limit, domains)
|
||||||
|
else:
|
||||||
|
raise WebSearchConfigurationError(f"Unsupported web search provider: {provider}")
|
||||||
|
event = "ai_tool.web_search.success" if results else "ai_tool.web_search.empty"
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event=event,
|
||||||
|
message="WebSearch request completed" if results else "WebSearch returned no results",
|
||||||
|
category="ai_tool",
|
||||||
|
level="info" if results else "warning",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
**context,
|
||||||
|
"status": "success" if results else "empty",
|
||||||
|
"result_count": len(results),
|
||||||
|
"duration_ms": int((perf_counter() - started_at) * 1000),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return results
|
||||||
|
except Exception as exc:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="ai_tool.web_search.failed",
|
||||||
|
message="WebSearch request failed",
|
||||||
|
category="ai_tool",
|
||||||
|
level="error",
|
||||||
|
service="ai_tool",
|
||||||
|
module=__name__,
|
||||||
|
context=exception_context(exc, {**context, "status": "failed", "duration_ms": int((perf_counter() - started_at) * 1000)}),
|
||||||
|
)
|
||||||
|
raise
|
||||||
|
|
||||||
|
async def test_connection(self) -> list[SearchEvidence]:
|
||||||
|
return await self.search("Planet WebSearch connectivity test", max_results=1)
|
||||||
|
|
||||||
|
async def _request_json(
|
||||||
|
self,
|
||||||
|
method: str,
|
||||||
|
url: str,
|
||||||
|
*,
|
||||||
|
provider_config: WebSearchProviderConfig,
|
||||||
|
headers: dict[str, str] | None = None,
|
||||||
|
params: dict[str, Any] | None = None,
|
||||||
|
json: dict[str, Any] | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
try:
|
||||||
|
async with httpx.AsyncClient(timeout=provider_config.timeout_seconds) as client:
|
||||||
|
response = await client.request(
|
||||||
|
method,
|
||||||
|
url,
|
||||||
|
headers=headers,
|
||||||
|
params=params,
|
||||||
|
json=json,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
data = response.json()
|
||||||
|
except httpx.HTTPStatusError as exc:
|
||||||
|
detail = exc.response.text or exc.response.reason_phrase
|
||||||
|
raise WebSearchError(f"{provider_config.provider} request failed: {detail}") from exc
|
||||||
|
except httpx.HTTPError as exc:
|
||||||
|
raise WebSearchError(f"{provider_config.provider} request failed: {exc}") from exc
|
||||||
|
except ValueError as exc:
|
||||||
|
raise WebSearchError(f"{provider_config.provider} returned invalid JSON") from exc
|
||||||
|
return data if isinstance(data, dict) else {}
|
||||||
|
|
||||||
|
async def _search_tavily(
|
||||||
|
self,
|
||||||
|
config: WebSearchProviderConfig,
|
||||||
|
query: str,
|
||||||
|
max_results: int,
|
||||||
|
domains: list[str] | None,
|
||||||
|
freshness_days: int | None,
|
||||||
|
) -> list[SearchEvidence]:
|
||||||
|
body: dict[str, Any] = {
|
||||||
|
"api_key": config.api_key,
|
||||||
|
"query": query,
|
||||||
|
"max_results": max_results,
|
||||||
|
"search_depth": config.search_depth or "basic",
|
||||||
|
"include_answer": config.include_answer,
|
||||||
|
"include_raw_content": config.include_raw_content,
|
||||||
|
}
|
||||||
|
if domains:
|
||||||
|
body["include_domains"] = domains
|
||||||
|
if freshness_days:
|
||||||
|
body["days"] = freshness_days
|
||||||
|
data = await self._request_json(
|
||||||
|
"POST",
|
||||||
|
_join_url(config.base_url, config.endpoint_path or "/search"),
|
||||||
|
provider_config=config,
|
||||||
|
json=body,
|
||||||
|
)
|
||||||
|
return [
|
||||||
|
SearchEvidence(
|
||||||
|
title=str(item.get("title") or ""),
|
||||||
|
url=str(item.get("url") or ""),
|
||||||
|
snippet=str(item.get("content") or ""),
|
||||||
|
content=str(item.get("raw_content") or ""),
|
||||||
|
score=_float_or_none(item.get("score")),
|
||||||
|
source_provider="tavily",
|
||||||
|
metadata={"query": data.get("query") or query},
|
||||||
|
)
|
||||||
|
for item in data.get("results") or []
|
||||||
|
if isinstance(item, dict) and item.get("url")
|
||||||
|
]
|
||||||
|
|
||||||
|
async def _search_brave(
|
||||||
|
self,
|
||||||
|
config: WebSearchProviderConfig,
|
||||||
|
query: str,
|
||||||
|
max_results: int,
|
||||||
|
domains: list[str] | None,
|
||||||
|
) -> list[SearchEvidence]:
|
||||||
|
search_query = query
|
||||||
|
if domains:
|
||||||
|
search_query = f"{query} " + " ".join(f"site:{domain}" for domain in domains)
|
||||||
|
data = await self._request_json(
|
||||||
|
"GET",
|
||||||
|
_join_url(config.base_url, config.endpoint_path or "/res/v1/web/search"),
|
||||||
|
provider_config=config,
|
||||||
|
headers={"X-Subscription-Token": config.api_key},
|
||||||
|
params={"q": search_query, "count": max_results},
|
||||||
|
)
|
||||||
|
results = (data.get("web") or {}).get("results") or []
|
||||||
|
return [
|
||||||
|
SearchEvidence(
|
||||||
|
title=str(item.get("title") or ""),
|
||||||
|
url=str(item.get("url") or ""),
|
||||||
|
snippet=str(item.get("description") or ""),
|
||||||
|
source_provider="brave",
|
||||||
|
metadata={"age": item.get("age")},
|
||||||
|
)
|
||||||
|
for item in results
|
||||||
|
if isinstance(item, dict) and item.get("url")
|
||||||
|
]
|
||||||
|
|
||||||
|
async def _search_serpapi(
|
||||||
|
self,
|
||||||
|
config: WebSearchProviderConfig,
|
||||||
|
query: str,
|
||||||
|
max_results: int,
|
||||||
|
) -> list[SearchEvidence]:
|
||||||
|
data = await self._request_json(
|
||||||
|
"GET",
|
||||||
|
_join_url(config.base_url, config.endpoint_path or "/search.json"),
|
||||||
|
provider_config=config,
|
||||||
|
params={
|
||||||
|
"api_key": config.api_key,
|
||||||
|
"engine": config.engine or "google",
|
||||||
|
"q": query,
|
||||||
|
"num": max_results,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return [
|
||||||
|
SearchEvidence(
|
||||||
|
title=str(item.get("title") or ""),
|
||||||
|
url=str(item.get("link") or ""),
|
||||||
|
snippet=str(item.get("snippet") or ""),
|
||||||
|
source_provider="serpapi",
|
||||||
|
metadata={"position": item.get("position")},
|
||||||
|
)
|
||||||
|
for item in data.get("organic_results") or []
|
||||||
|
if isinstance(item, dict) and item.get("link")
|
||||||
|
]
|
||||||
|
|
||||||
|
async def _search_exa(
|
||||||
|
self,
|
||||||
|
config: WebSearchProviderConfig,
|
||||||
|
query: str,
|
||||||
|
max_results: int,
|
||||||
|
domains: list[str] | None,
|
||||||
|
) -> list[SearchEvidence]:
|
||||||
|
body: dict[str, Any] = {
|
||||||
|
"query": query,
|
||||||
|
"numResults": max_results,
|
||||||
|
}
|
||||||
|
if domains:
|
||||||
|
body["includeDomains"] = domains
|
||||||
|
if config.include_text:
|
||||||
|
body["contents"] = {"text": True}
|
||||||
|
data = await self._request_json(
|
||||||
|
"POST",
|
||||||
|
_join_url(config.base_url, config.endpoint_path or "/search"),
|
||||||
|
provider_config=config,
|
||||||
|
headers={"Authorization": f"Bearer {config.api_key}"},
|
||||||
|
json=body,
|
||||||
|
)
|
||||||
|
return [
|
||||||
|
SearchEvidence(
|
||||||
|
title=str(item.get("title") or ""),
|
||||||
|
url=str(item.get("url") or ""),
|
||||||
|
snippet=str(item.get("summary") or ""),
|
||||||
|
content=str(item.get("text") or ""),
|
||||||
|
score=_float_or_none(item.get("score")),
|
||||||
|
source_provider="exa",
|
||||||
|
metadata={"id": item.get("id")},
|
||||||
|
)
|
||||||
|
for item in data.get("results") or []
|
||||||
|
if isinstance(item, dict) and item.get("url")
|
||||||
|
]
|
||||||
|
|
||||||
|
async def _search_firecrawl(
|
||||||
|
self,
|
||||||
|
config: WebSearchProviderConfig,
|
||||||
|
query: str,
|
||||||
|
max_results: int,
|
||||||
|
) -> list[SearchEvidence]:
|
||||||
|
data = await self._request_json(
|
||||||
|
"POST",
|
||||||
|
_join_url(config.base_url, config.search_path or "/v2/search"),
|
||||||
|
provider_config=config,
|
||||||
|
headers={"Authorization": f"Bearer {config.api_key}"},
|
||||||
|
json={"query": query, "limit": max_results},
|
||||||
|
)
|
||||||
|
raw_results = data.get("data") or data.get("results") or []
|
||||||
|
return [
|
||||||
|
SearchEvidence(
|
||||||
|
title=str(item.get("title") or ""),
|
||||||
|
url=str(item.get("url") or item.get("sourceURL") or ""),
|
||||||
|
snippet=str(item.get("description") or item.get("markdown") or ""),
|
||||||
|
source_provider="firecrawl",
|
||||||
|
metadata={"status": item.get("status")},
|
||||||
|
)
|
||||||
|
for item in raw_results
|
||||||
|
if isinstance(item, dict) and (item.get("url") or item.get("sourceURL"))
|
||||||
|
]
|
||||||
|
|
||||||
|
async def _search_searxng(
|
||||||
|
self,
|
||||||
|
config: WebSearchProviderConfig,
|
||||||
|
query: str,
|
||||||
|
max_results: int,
|
||||||
|
domains: list[str] | None,
|
||||||
|
) -> list[SearchEvidence]:
|
||||||
|
search_query = query
|
||||||
|
if domains:
|
||||||
|
search_query = f"{query} " + " ".join(f"site:{domain}" for domain in domains)
|
||||||
|
params: dict[str, Any] = {
|
||||||
|
"q": search_query,
|
||||||
|
"format": "json",
|
||||||
|
"categories": config.categories or "general",
|
||||||
|
}
|
||||||
|
if config.engines:
|
||||||
|
params["engines"] = ",".join(config.engines)
|
||||||
|
headers = {"Authorization": f"Bearer {config.api_key}"} if config.api_key else None
|
||||||
|
data = await self._request_json(
|
||||||
|
"GET",
|
||||||
|
_join_url(config.base_url, config.endpoint_path or "/"),
|
||||||
|
provider_config=config,
|
||||||
|
headers=headers,
|
||||||
|
params=params,
|
||||||
|
)
|
||||||
|
results = data.get("results") or []
|
||||||
|
evidence = [
|
||||||
|
SearchEvidence(
|
||||||
|
title=str(item.get("title") or ""),
|
||||||
|
url=str(item.get("url") or ""),
|
||||||
|
snippet=str(item.get("content") or ""),
|
||||||
|
score=_float_or_none(item.get("score")),
|
||||||
|
source_provider="searxng",
|
||||||
|
metadata={"engine": item.get("engine")},
|
||||||
|
)
|
||||||
|
for item in results
|
||||||
|
if isinstance(item, dict) and item.get("url")
|
||||||
|
]
|
||||||
|
return evidence[:max_results]
|
||||||
|
|
||||||
|
|
||||||
|
def _join_url(base_url: str, path: str) -> str:
|
||||||
|
return f"{(base_url or '').rstrip('/')}/{(path or '').lstrip('/')}"
|
||||||
|
|
||||||
|
|
||||||
|
def _float_or_none(value: Any) -> float | None:
|
||||||
|
try:
|
||||||
|
return float(value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
@@ -8,6 +8,9 @@ from sqlalchemy.ext.asyncio import AsyncSession
|
|||||||
|
|
||||||
from app.models.alert import Alert, AlertSeverity, AlertStatus
|
from app.models.alert import Alert, AlertSeverity, AlertStatus
|
||||||
from app.schemas.ai import AlertBriefRequest, SituationalAnalysisRequest
|
from app.schemas.ai import AlertBriefRequest, SituationalAnalysisRequest
|
||||||
|
from app.ai_tasks.prompts import get_effective_prompt
|
||||||
|
|
||||||
|
ALERT_BRIEF_PROMPT_KEY = "alerts.brief"
|
||||||
|
|
||||||
|
|
||||||
def _format_counter(counter: Counter[str], empty_text: str = "无") -> str:
|
def _format_counter(counter: Counter[str], empty_text: str = "无") -> str:
|
||||||
@@ -84,11 +87,13 @@ async def build_alert_brief_request(
|
|||||||
"top_datasources": dict(datasource_counts.most_common(6)),
|
"top_datasources": dict(datasource_counts.most_common(6)),
|
||||||
"top_active_datasources": dict(active_datasource_counts.most_common(5)),
|
"top_active_datasources": dict(active_datasource_counts.most_common(5)),
|
||||||
}
|
}
|
||||||
|
prompt = await get_effective_prompt(db, ALERT_BRIEF_PROMPT_KEY)
|
||||||
|
|
||||||
return (
|
return (
|
||||||
SituationalAnalysisRequest(
|
SituationalAnalysisRequest(
|
||||||
title="告警态势 AI 简报",
|
title="告警态势 AI 简报",
|
||||||
objective="基于当前告警总量、严重度、状态、数据源分布与最近告警摘录,生成一份面向值班人员的简明告警态势简报,突出待处理风险、告警集中点和优先动作。",
|
objective=prompt.prompt,
|
||||||
|
system_prompt=prompt.system_prompt or None,
|
||||||
observations=facts,
|
observations=facts,
|
||||||
constraints=[
|
constraints=[
|
||||||
"明确区分事实、推断与建议。",
|
"明确区分事实、推断与建议。",
|
||||||
|
|||||||
209
backend/app/services/barentswatch.py
Normal file
209
backend/app/services/barentswatch.py
Normal file
@@ -0,0 +1,209 @@
|
|||||||
|
"""BarentsWatch AIS credential resolution and connectivity checks."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import os
|
||||||
|
import shlex
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
from sqlalchemy import select
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.data_sources import get_data_sources_config
|
||||||
|
from app.models.datasource_config import DataSourceConfig
|
||||||
|
|
||||||
|
|
||||||
|
BARENTSWATCH_LATEST_URL = "https://live.ais.barentswatch.no/v1/latest/combined"
|
||||||
|
BARENTSWATCH_TOKEN_URL = "https://id.barentswatch.no/connect/token"
|
||||||
|
BARENTSWATCH_DATASOURCE_NAME = "barentswatch_vessels"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class BarentsWatchConfig:
|
||||||
|
endpoint: str
|
||||||
|
client_id: str
|
||||||
|
client_secret: str
|
||||||
|
credential_source: str
|
||||||
|
endpoint_source: str
|
||||||
|
|
||||||
|
|
||||||
|
def _read_zshrc_env(path: Path | None = None) -> dict[str, str]:
|
||||||
|
zshrc_path = path or Path.home() / ".zshrc"
|
||||||
|
if not zshrc_path.exists():
|
||||||
|
return {}
|
||||||
|
|
||||||
|
values: dict[str, str] = {}
|
||||||
|
for raw_line in zshrc_path.read_text(encoding="utf-8", errors="ignore").splitlines():
|
||||||
|
line = raw_line.strip()
|
||||||
|
if not line or line.startswith("#"):
|
||||||
|
continue
|
||||||
|
if line.startswith("export "):
|
||||||
|
line = line[len("export ") :].strip()
|
||||||
|
if "=" not in line:
|
||||||
|
continue
|
||||||
|
|
||||||
|
key, value = line.split("=", 1)
|
||||||
|
key = key.strip()
|
||||||
|
if not key or not key.replace("_", "").isalnum() or not key[0].isalpha():
|
||||||
|
continue
|
||||||
|
|
||||||
|
try:
|
||||||
|
parsed = shlex.split(value, comments=True, posix=True)
|
||||||
|
except ValueError:
|
||||||
|
parsed = [value.strip().strip("'\"")]
|
||||||
|
if parsed:
|
||||||
|
values[key] = parsed[0]
|
||||||
|
return values
|
||||||
|
|
||||||
|
|
||||||
|
def _first_env_value(zshrc_env: dict[str, str], *keys: str) -> tuple[str, str]:
|
||||||
|
for key in keys:
|
||||||
|
value = os.getenv(key)
|
||||||
|
if value:
|
||||||
|
return value, "environment"
|
||||||
|
for key in keys:
|
||||||
|
value = zshrc_env.get(key)
|
||||||
|
if value:
|
||||||
|
return value, "~/.zshrc"
|
||||||
|
return "", ""
|
||||||
|
|
||||||
|
|
||||||
|
async def get_barentswatch_datasource_record(db: AsyncSession) -> DataSourceConfig | None:
|
||||||
|
result = await db.execute(
|
||||||
|
select(DataSourceConfig)
|
||||||
|
.where(DataSourceConfig.name == BARENTSWATCH_DATASOURCE_NAME)
|
||||||
|
.where(DataSourceConfig.is_active.is_(True))
|
||||||
|
)
|
||||||
|
return result.scalar_one_or_none()
|
||||||
|
|
||||||
|
|
||||||
|
async def resolve_barentswatch_config(db: AsyncSession | None = None) -> BarentsWatchConfig:
|
||||||
|
record = await get_barentswatch_datasource_record(db) if db else None
|
||||||
|
auth_config = dict(record.auth_config or {}) if record else {}
|
||||||
|
config = dict(record.config or {}) if record else {}
|
||||||
|
zshrc_env = _read_zshrc_env()
|
||||||
|
|
||||||
|
env_client_id, env_source = _first_env_value(
|
||||||
|
zshrc_env,
|
||||||
|
"BARENTSWATCH_CLIENT_ID",
|
||||||
|
"BARRENTSWATCH_CLIENT_ID",
|
||||||
|
)
|
||||||
|
env_client_secret, secret_env_source = _first_env_value(
|
||||||
|
zshrc_env,
|
||||||
|
"BARENTSWATCH_CLIENT_SECRET",
|
||||||
|
"BARRENTSWATCH_CLIENT_SECRET",
|
||||||
|
)
|
||||||
|
client_id = auth_config.get("client_id") or config.get("client_id") or env_client_id
|
||||||
|
client_secret = (
|
||||||
|
auth_config.get("client_secret") or config.get("client_secret") or env_client_secret
|
||||||
|
)
|
||||||
|
|
||||||
|
credential_source = ""
|
||||||
|
if auth_config.get("client_id") or auth_config.get("client_secret"):
|
||||||
|
credential_source = "datasource_config"
|
||||||
|
elif config.get("client_id") or config.get("client_secret"):
|
||||||
|
credential_source = "datasource_runtime_config"
|
||||||
|
elif env_source or secret_env_source:
|
||||||
|
credential_source = env_source or secret_env_source
|
||||||
|
|
||||||
|
yaml_endpoint = get_data_sources_config().get_yaml_url(BARENTSWATCH_DATASOURCE_NAME)
|
||||||
|
endpoint = record.endpoint if record and record.endpoint else yaml_endpoint
|
||||||
|
return BarentsWatchConfig(
|
||||||
|
endpoint=endpoint or BARENTSWATCH_LATEST_URL,
|
||||||
|
client_id=str(client_id or ""),
|
||||||
|
client_secret=str(client_secret or ""),
|
||||||
|
credential_source=credential_source or "missing",
|
||||||
|
endpoint_source="datasource_config" if record and record.endpoint else "default",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def fetch_barentswatch_access_token(
|
||||||
|
client: httpx.AsyncClient,
|
||||||
|
config: BarentsWatchConfig,
|
||||||
|
) -> str | None:
|
||||||
|
if not config.client_id or not config.client_secret:
|
||||||
|
return None
|
||||||
|
|
||||||
|
response = await client.post(
|
||||||
|
BARENTSWATCH_TOKEN_URL,
|
||||||
|
data={
|
||||||
|
"client_id": config.client_id,
|
||||||
|
"client_secret": config.client_secret,
|
||||||
|
"scope": "ais",
|
||||||
|
"grant_type": "client_credentials",
|
||||||
|
},
|
||||||
|
headers={"Content-Type": "application/x-www-form-urlencoded"},
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
payload = response.json()
|
||||||
|
token = payload.get("access_token")
|
||||||
|
return str(token) if token else None
|
||||||
|
|
||||||
|
|
||||||
|
async def check_barentswatch_connectivity(db: AsyncSession) -> dict[str, Any]:
|
||||||
|
config = await resolve_barentswatch_config(db)
|
||||||
|
return await check_barentswatch_config(config)
|
||||||
|
|
||||||
|
|
||||||
|
async def check_barentswatch_config(config: BarentsWatchConfig) -> dict[str, Any]:
|
||||||
|
if not config.client_id or not config.client_secret:
|
||||||
|
return {
|
||||||
|
"success": False,
|
||||||
|
"stage": "credentials",
|
||||||
|
"message": "未找到 BarentsWatch client id/client secret,请先配置采集器凭证。",
|
||||||
|
"endpoint": config.endpoint,
|
||||||
|
"credential_source": config.credential_source,
|
||||||
|
"settings_tab": "collector_credentials",
|
||||||
|
}
|
||||||
|
|
||||||
|
try:
|
||||||
|
async with httpx.AsyncClient(timeout=20.0) as client:
|
||||||
|
token = await fetch_barentswatch_access_token(client, config)
|
||||||
|
if not token:
|
||||||
|
return {
|
||||||
|
"success": False,
|
||||||
|
"stage": "token",
|
||||||
|
"message": "BarentsWatch token 响应中没有 access_token,请检查凭证。",
|
||||||
|
"endpoint": config.endpoint,
|
||||||
|
"credential_source": config.credential_source,
|
||||||
|
"settings_tab": "collector_credentials",
|
||||||
|
}
|
||||||
|
|
||||||
|
async with client.stream(
|
||||||
|
"GET",
|
||||||
|
config.endpoint,
|
||||||
|
headers={"Authorization": f"Bearer {token}"},
|
||||||
|
) as response:
|
||||||
|
response.raise_for_status()
|
||||||
|
|
||||||
|
return {
|
||||||
|
"success": True,
|
||||||
|
"stage": "endpoint",
|
||||||
|
"message": "BarentsWatch AIS token 和数据接口均可连通。",
|
||||||
|
"endpoint": config.endpoint,
|
||||||
|
"credential_source": config.credential_source,
|
||||||
|
"endpoint_source": config.endpoint_source,
|
||||||
|
}
|
||||||
|
except httpx.HTTPStatusError as exc:
|
||||||
|
status_code = exc.response.status_code
|
||||||
|
stage = "token" if str(exc.request.url) == BARENTSWATCH_TOKEN_URL else "endpoint"
|
||||||
|
return {
|
||||||
|
"success": False,
|
||||||
|
"stage": stage,
|
||||||
|
"message": f"BarentsWatch {stage} 请求返回 HTTP {status_code},请检查凭证或接口地址。",
|
||||||
|
"endpoint": config.endpoint,
|
||||||
|
"credential_source": config.credential_source,
|
||||||
|
"settings_tab": "collector_credentials",
|
||||||
|
}
|
||||||
|
except httpx.HTTPError as exc:
|
||||||
|
return {
|
||||||
|
"success": False,
|
||||||
|
"stage": "network",
|
||||||
|
"message": f"BarentsWatch 链路检查失败:{exc.__class__.__name__}",
|
||||||
|
"endpoint": config.endpoint,
|
||||||
|
"credential_source": config.credential_source,
|
||||||
|
"settings_tab": "collector_credentials",
|
||||||
|
}
|
||||||
@@ -11,9 +11,12 @@ from app.models.bgp_anomaly import BGPAnomaly
|
|||||||
from app.models.bgp_incident import BGPIncident
|
from app.models.bgp_incident import BGPIncident
|
||||||
from app.models.bgp_observation import BGPObservation
|
from app.models.bgp_observation import BGPObservation
|
||||||
from app.schemas.ai import SituationalAnalysisRequest
|
from app.schemas.ai import SituationalAnalysisRequest
|
||||||
|
from app.ai_tasks.prompts import get_effective_prompt
|
||||||
from app.services.bgp_collectors import build_bgp_collector_coverage
|
from app.services.bgp_collectors import build_bgp_collector_coverage
|
||||||
from app.services.bgp_enrichment import lookup_prefix_geography
|
from app.services.bgp_enrichment import lookup_prefix_geography
|
||||||
|
|
||||||
|
BGP_BRIEF_PROMPT_KEY = "bgp.brief"
|
||||||
|
|
||||||
|
|
||||||
def _format_counter(counter: dict[str, int], empty_text: str = "无") -> str:
|
def _format_counter(counter: dict[str, int], empty_text: str = "无") -> str:
|
||||||
if not counter:
|
if not counter:
|
||||||
@@ -243,12 +246,15 @@ async def build_bgp_brief_request(
|
|||||||
for prefix, item in list(prefix_geographies.items())[:8]
|
for prefix, item in list(prefix_geographies.items())[:8]
|
||||||
},
|
},
|
||||||
}
|
}
|
||||||
|
prompt = await get_effective_prompt(db, BGP_BRIEF_PROMPT_KEY)
|
||||||
|
|
||||||
return SituationalAnalysisRequest(
|
return SituationalAnalysisRequest(
|
||||||
title="BGP 态势 AI 简报",
|
title="BGP 态势 AI 简报",
|
||||||
objective="基于当前 BGP incidents、anomalies、原始观测事件、观测站覆盖与 prefix geography 证据,生成一份面向操作员的简明态势简报,突出区域热点、观测偏差、当前风险、证据和优先动作。",
|
objective=prompt.prompt,
|
||||||
|
system_prompt=prompt.system_prompt or None,
|
||||||
observations=observations_lines,
|
observations=observations_lines,
|
||||||
constraints=[
|
constraints=[
|
||||||
|
"直接输出中文 Markdown 简报正文,不要输出英文写作计划、提示词复述、字段说明或元评论。",
|
||||||
"明确区分事实、推断与建议。",
|
"明确区分事实、推断与建议。",
|
||||||
"优先指出需要立即关注的高严重度 incident 或异常模式。",
|
"优先指出需要立即关注的高严重度 incident 或异常模式。",
|
||||||
"需要单独指出哪些区域结论来自 prefix geography / affected regions,哪些可能受 collector coverage 偏差影响。",
|
"需要单独指出哪些区域结论来自 prefix geography / affected regions,哪些可能受 collector coverage 偏差影响。",
|
||||||
|
|||||||
324
backend/app/services/bgp_collector_locations.py
Normal file
324
backend/app/services/bgp_collector_locations.py
Normal file
@@ -0,0 +1,324 @@
|
|||||||
|
"""BGP route-collector location resolver.
|
||||||
|
|
||||||
|
Collector positions are stored in the ``bgp_collector_locations`` database
|
||||||
|
table. The old JSON registry is now only a seed payload used during database
|
||||||
|
initialization, not a runtime resolver or candidate source.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Iterator
|
||||||
|
|
||||||
|
from sqlalchemy import select
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.models.bgp_collector_location import BGPCollectorLocation
|
||||||
|
from app.services.location import (
|
||||||
|
LocationCandidate,
|
||||||
|
LocationPipeline,
|
||||||
|
LocationQuery,
|
||||||
|
NominatimResolver,
|
||||||
|
ResolutionResult,
|
||||||
|
ResolverOutput,
|
||||||
|
SourceCoordinatesResolver,
|
||||||
|
build_default_nominatim_geocoder,
|
||||||
|
coerce_str,
|
||||||
|
normalize_text,
|
||||||
|
)
|
||||||
|
|
||||||
|
SEED_PATH = (
|
||||||
|
Path(__file__).resolve().parents[1]
|
||||||
|
/ "data"
|
||||||
|
/ "seeds"
|
||||||
|
/ "ripe_ris_collector_locations_seed.json"
|
||||||
|
)
|
||||||
|
|
||||||
|
# ── Geocoder (kept at module level for monkeypatching + cache_clear) ──
|
||||||
|
|
||||||
|
_geocode_online = build_default_nominatim_geocoder()
|
||||||
|
|
||||||
|
|
||||||
|
# ── In-process compatibility cache ──────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
RIPE_RIS_COLLECTOR_COORDS: dict[str, dict[str, Any]] = {}
|
||||||
|
|
||||||
|
|
||||||
|
def _collector_record_to_dict(record: BGPCollectorLocation) -> dict[str, Any]:
|
||||||
|
return record.to_location_dict()
|
||||||
|
|
||||||
|
|
||||||
|
def set_bgp_collector_location_cache(
|
||||||
|
locations: dict[str, dict[str, Any]],
|
||||||
|
) -> None:
|
||||||
|
"""Replace the legacy compatibility cache in-place."""
|
||||||
|
RIPE_RIS_COLLECTOR_COORDS.clear()
|
||||||
|
RIPE_RIS_COLLECTOR_COORDS.update(
|
||||||
|
{coerce_str(key): dict(value) for key, value in locations.items()}
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def refresh_bgp_collector_location_cache(
|
||||||
|
session: AsyncSession,
|
||||||
|
) -> dict[str, dict[str, Any]]:
|
||||||
|
result = await session.execute(select(BGPCollectorLocation))
|
||||||
|
records = result.scalars().all()
|
||||||
|
cache = {
|
||||||
|
record.collector_id: _collector_record_to_dict(record)
|
||||||
|
for record in records
|
||||||
|
if record.collector_id
|
||||||
|
}
|
||||||
|
set_bgp_collector_location_cache(cache)
|
||||||
|
return cache
|
||||||
|
|
||||||
|
|
||||||
|
def _load_seed_payload() -> dict[str, Any]:
|
||||||
|
with SEED_PATH.open("r", encoding="utf-8") as handle:
|
||||||
|
return json.load(handle)
|
||||||
|
|
||||||
|
|
||||||
|
def _seed_entry_to_record_kwargs(entry: dict[str, Any], collector_id: str) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"collector_id": collector_id,
|
||||||
|
"operator": entry.get("operator") or "RIPE NCC",
|
||||||
|
"site": entry.get("site"),
|
||||||
|
"city": entry.get("city"),
|
||||||
|
"country": entry.get("country"),
|
||||||
|
"latitude": entry.get("latitude"),
|
||||||
|
"longitude": entry.get("longitude"),
|
||||||
|
"precision": entry.get("precision") or "city",
|
||||||
|
"confidence": entry.get("confidence"),
|
||||||
|
"source": "legacy_seed",
|
||||||
|
"source_url": None,
|
||||||
|
"source_note": entry.get("source_note")
|
||||||
|
or "Seeded from legacy RIPE RIS collector coordinates",
|
||||||
|
"raw_payload": entry,
|
||||||
|
"needs_confirmation": True,
|
||||||
|
"verification_status": "unverified",
|
||||||
|
"verified_at": None,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def seed_default_bgp_collector_locations(session: AsyncSession) -> None:
|
||||||
|
"""Seed default RIPE RIS collector locations without overwriting users."""
|
||||||
|
payload = _load_seed_payload()
|
||||||
|
for entry in payload.get("locations", []):
|
||||||
|
aliases = entry.get("aliases") or []
|
||||||
|
collector_ids = [
|
||||||
|
coerce_str(alias)
|
||||||
|
for alias in aliases
|
||||||
|
if coerce_str(alias).startswith("rrc")
|
||||||
|
]
|
||||||
|
if not collector_ids:
|
||||||
|
continue
|
||||||
|
collector_id = collector_ids[0]
|
||||||
|
existing = await session.scalar(
|
||||||
|
select(BGPCollectorLocation).where(
|
||||||
|
BGPCollectorLocation.collector_id == collector_id
|
||||||
|
)
|
||||||
|
)
|
||||||
|
if existing:
|
||||||
|
continue
|
||||||
|
session.add(
|
||||||
|
BGPCollectorLocation(
|
||||||
|
**_seed_entry_to_record_kwargs(entry, collector_id)
|
||||||
|
)
|
||||||
|
)
|
||||||
|
await session.commit()
|
||||||
|
await refresh_bgp_collector_location_cache(session)
|
||||||
|
|
||||||
|
|
||||||
|
def get_bgp_collector_location_dict(collector_name: str) -> dict[str, Any]:
|
||||||
|
"""Return the current cached collector location dict, or ``{}`` if unknown."""
|
||||||
|
return dict(RIPE_RIS_COLLECTOR_COORDS.get(coerce_str(collector_name), {}))
|
||||||
|
|
||||||
|
|
||||||
|
def iter_known_collector_names() -> Iterator[str]:
|
||||||
|
"""Yield every collector technical name (rrcXX) known in the cache."""
|
||||||
|
return iter(sorted(RIPE_RIS_COLLECTOR_COORDS.keys()))
|
||||||
|
|
||||||
|
|
||||||
|
# ── Pipeline construction ──────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
class StoredCollectorLocationResolver:
|
||||||
|
"""Resolve a collector through the DB-backed compatibility cache."""
|
||||||
|
|
||||||
|
name = "stored_collector_location"
|
||||||
|
|
||||||
|
def resolve(self, query: LocationQuery) -> ResolverOutput:
|
||||||
|
collector = coerce_str(query.name)
|
||||||
|
if not collector:
|
||||||
|
for alias in query.aliases:
|
||||||
|
collector = coerce_str(alias)
|
||||||
|
if collector:
|
||||||
|
break
|
||||||
|
if not collector:
|
||||||
|
return ResolverOutput()
|
||||||
|
location = get_bgp_collector_location_dict(collector)
|
||||||
|
if not location:
|
||||||
|
return ResolverOutput()
|
||||||
|
latitude = location.get("latitude")
|
||||||
|
longitude = location.get("longitude")
|
||||||
|
if latitude in (None, 0.0) or longitude in (None, 0.0):
|
||||||
|
return ResolverOutput()
|
||||||
|
return ResolverOutput(
|
||||||
|
candidates=(
|
||||||
|
LocationCandidate(
|
||||||
|
latitude=float(latitude),
|
||||||
|
longitude=float(longitude),
|
||||||
|
display_name=location.get("matched_location_name") or collector,
|
||||||
|
precision=location.get("precision") or "city",
|
||||||
|
confidence=float(location.get("confidence") or 0.85),
|
||||||
|
query=f"stored_collector_location::{collector}",
|
||||||
|
source=location.get("source") or self.name,
|
||||||
|
source_note=location.get("source_note"),
|
||||||
|
matched_fields=("collector",),
|
||||||
|
needs_confirmation=bool(location.get("needs_confirmation")),
|
||||||
|
city=location.get("city"),
|
||||||
|
region=None,
|
||||||
|
country=location.get("country"),
|
||||||
|
matched_location_name=(
|
||||||
|
location.get("matched_location_name") or collector
|
||||||
|
),
|
||||||
|
location_verified_at=location.get("verified_at"),
|
||||||
|
suggested_registry_entry=None,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _bgp_collector_query_plan(
|
||||||
|
query: LocationQuery,
|
||||||
|
) -> list[tuple[str, tuple[str, ...]]]:
|
||||||
|
"""Build the Nominatim query plan for a BGP collector."""
|
||||||
|
extra = query.extra or {}
|
||||||
|
site = str(extra.get("site") or "")
|
||||||
|
operator = str(extra.get("operator") or "")
|
||||||
|
city = query.city or ""
|
||||||
|
country = query.country or ""
|
||||||
|
|
||||||
|
plan: list[tuple[str, tuple[str, ...]]] = []
|
||||||
|
|
||||||
|
def add(parts: list[tuple[str, str]]) -> None:
|
||||||
|
non_empty = [(field, value) for field, value in parts if value]
|
||||||
|
if not non_empty:
|
||||||
|
return
|
||||||
|
seen: set[str] = set()
|
||||||
|
cleaned: list[str] = []
|
||||||
|
fields: list[str] = []
|
||||||
|
for field, value in non_empty:
|
||||||
|
key = normalize_text(value)
|
||||||
|
if not key or key in seen:
|
||||||
|
continue
|
||||||
|
seen.add(key)
|
||||||
|
cleaned.append(value)
|
||||||
|
fields.append(field)
|
||||||
|
if not cleaned:
|
||||||
|
return
|
||||||
|
composed = ", ".join(cleaned)
|
||||||
|
if not any(composed == existing for existing, _ in plan):
|
||||||
|
plan.append((composed, tuple(fields)))
|
||||||
|
|
||||||
|
add([("site", site), ("city", city), ("country", country)])
|
||||||
|
add([("site", site), ("country", country)])
|
||||||
|
add([("operator", operator), ("city", city), ("country", country)])
|
||||||
|
add([("city", city), ("country", country)])
|
||||||
|
return plan
|
||||||
|
|
||||||
|
|
||||||
|
BGP_COLLECTOR_PIPELINE = LocationPipeline(
|
||||||
|
[
|
||||||
|
SourceCoordinatesResolver(),
|
||||||
|
StoredCollectorLocationResolver(),
|
||||||
|
],
|
||||||
|
failure_reason=(
|
||||||
|
"Could not resolve BGP collector to renderable coordinates from"
|
||||||
|
" source coordinates or stored collector location."
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
BGP_COLLECTOR_COLLECTION_PIPELINE = LocationPipeline(
|
||||||
|
[
|
||||||
|
SourceCoordinatesResolver(),
|
||||||
|
NominatimResolver(
|
||||||
|
query_plan_builder=_bgp_collector_query_plan,
|
||||||
|
# Late-binding so tests can monkeypatch ``_geocode_online``.
|
||||||
|
geocoder=lambda q: _geocode_online(q),
|
||||||
|
),
|
||||||
|
],
|
||||||
|
failure_reason=(
|
||||||
|
"Could not resolve BGP collector to renderable coordinates from"
|
||||||
|
" source coordinates or online geocoding."
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ── Public API ─────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_bgp_collector_location(
|
||||||
|
collector_name: str,
|
||||||
|
*,
|
||||||
|
city: str | None = None,
|
||||||
|
country: str | None = None,
|
||||||
|
site: str | None = None,
|
||||||
|
operator: str | None = None,
|
||||||
|
) -> ResolutionResult:
|
||||||
|
"""Resolve a BGP collector to its best-known stored location."""
|
||||||
|
stored = get_bgp_collector_location_dict(collector_name)
|
||||||
|
name = coerce_str(collector_name) or None
|
||||||
|
query = LocationQuery(
|
||||||
|
name=name,
|
||||||
|
aliases=tuple(filter(None, (collector_name,))),
|
||||||
|
city=coerce_str(city or stored.get("city")) or None,
|
||||||
|
country=coerce_str(country or stored.get("country")) or None,
|
||||||
|
extra={
|
||||||
|
"site": coerce_str(site or stored.get("site")),
|
||||||
|
"operator": coerce_str(operator or stored.get("operator")) or "RIPE NCC",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return BGP_COLLECTOR_PIPELINE.resolve_best(query)
|
||||||
|
|
||||||
|
|
||||||
|
def collect_bgp_collector_location_candidates(
|
||||||
|
*,
|
||||||
|
collector: str | None = None,
|
||||||
|
city: str | None = None,
|
||||||
|
country: str | None = None,
|
||||||
|
site: str | None = None,
|
||||||
|
operator: str | None = None,
|
||||||
|
) -> tuple[list[LocationCandidate], list[str]]:
|
||||||
|
query = build_bgp_collector_location_query(
|
||||||
|
collector=collector,
|
||||||
|
city=city,
|
||||||
|
country=country,
|
||||||
|
site=site,
|
||||||
|
operator=operator,
|
||||||
|
)
|
||||||
|
return BGP_COLLECTOR_COLLECTION_PIPELINE.collect_candidates(query)
|
||||||
|
|
||||||
|
|
||||||
|
def build_bgp_collector_location_query(
|
||||||
|
*,
|
||||||
|
collector: str | None = None,
|
||||||
|
city: str | None = None,
|
||||||
|
country: str | None = None,
|
||||||
|
site: str | None = None,
|
||||||
|
operator: str | None = None,
|
||||||
|
) -> LocationQuery:
|
||||||
|
stored = get_bgp_collector_location_dict(collector or "")
|
||||||
|
name = coerce_str(collector) or None
|
||||||
|
return LocationQuery(
|
||||||
|
name=name,
|
||||||
|
aliases=tuple(filter(None, (collector,))),
|
||||||
|
city=coerce_str(city or stored.get("city")) or None,
|
||||||
|
country=coerce_str(country or stored.get("country")) or None,
|
||||||
|
extra={
|
||||||
|
"site": coerce_str(site or stored.get("site")),
|
||||||
|
"operator": coerce_str(operator or stored.get("operator")) or "RIPE NCC",
|
||||||
|
"collector": coerce_str(collector),
|
||||||
|
},
|
||||||
|
)
|
||||||
155
backend/app/services/bgp_event_locations.py
Normal file
155
backend/app/services/bgp_event_locations.py
Normal file
@@ -0,0 +1,155 @@
|
|||||||
|
"""BGP event location resolver.
|
||||||
|
|
||||||
|
A BGP event (announcement / withdrawal / RIB entry) is geographically tied to
|
||||||
|
the route collector that observed it. This module defines the pipeline that
|
||||||
|
turns an event payload into renderable coordinates.
|
||||||
|
|
||||||
|
Current resolver chain:
|
||||||
|
|
||||||
|
SourceCoordinates → event payload itself carries lat/lon (rare; some
|
||||||
|
enriched feeds do).
|
||||||
|
InheritFromCollector → look up the owning collector via
|
||||||
|
:func:`resolve_bgp_collector_location`.
|
||||||
|
|
||||||
|
Future plug-ins (no consumer changes required, just append to the list):
|
||||||
|
|
||||||
|
ASNFacilityResolver — origin/peer ASN → peeringdb facility.
|
||||||
|
PrefixGeoResolver — prefix → IP range geo lookup (iptoasn / opengeofeed).
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from app.services.bgp_collector_locations import (
|
||||||
|
get_bgp_collector_location_dict,
|
||||||
|
)
|
||||||
|
from app.services.location import (
|
||||||
|
InheritFromAnotherEntityResolver,
|
||||||
|
LocationCandidate,
|
||||||
|
LocationPipeline,
|
||||||
|
LocationQuery,
|
||||||
|
ResolutionResult,
|
||||||
|
SourceCoordinatesResolver,
|
||||||
|
coerce_str,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _inherit_from_owning_collector(
|
||||||
|
query: LocationQuery,
|
||||||
|
) -> LocationCandidate | None:
|
||||||
|
"""Look up the event's owning collector by exact name in the DB-backed cache."""
|
||||||
|
extra = query.extra or {}
|
||||||
|
collector_name = coerce_str(extra.get("collector"))
|
||||||
|
if not collector_name:
|
||||||
|
return None
|
||||||
|
legacy = get_bgp_collector_location_dict(collector_name)
|
||||||
|
if not legacy:
|
||||||
|
return None
|
||||||
|
latitude = legacy.get("latitude")
|
||||||
|
longitude = legacy.get("longitude")
|
||||||
|
if latitude in (None, 0.0) or longitude in (None, 0.0):
|
||||||
|
return None
|
||||||
|
return LocationCandidate(
|
||||||
|
latitude=float(latitude),
|
||||||
|
longitude=float(longitude),
|
||||||
|
display_name=legacy.get("matched_location_name") or collector_name,
|
||||||
|
precision=legacy.get("precision") or "city",
|
||||||
|
confidence=float(legacy.get("confidence") or 0.85),
|
||||||
|
query=f"inherit_from_collector::{collector_name}",
|
||||||
|
source="inherited_from_collector",
|
||||||
|
source_note=(
|
||||||
|
f"Inherited from owning collector {collector_name}"
|
||||||
|
),
|
||||||
|
matched_fields=("collector",),
|
||||||
|
needs_confirmation=bool(legacy.get("needs_confirmation")),
|
||||||
|
city=legacy.get("city"),
|
||||||
|
region=None,
|
||||||
|
country=legacy.get("country"),
|
||||||
|
matched_location_name=legacy.get("matched_location_name"),
|
||||||
|
location_verified_at=legacy.get("verified_at"),
|
||||||
|
suggested_registry_entry=None,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
BGP_EVENT_PIPELINE = LocationPipeline(
|
||||||
|
[
|
||||||
|
SourceCoordinatesResolver(),
|
||||||
|
InheritFromAnotherEntityResolver(
|
||||||
|
source_lookup=_inherit_from_owning_collector,
|
||||||
|
name="inherited_from_collector",
|
||||||
|
),
|
||||||
|
# Plug new resolvers (peeringdb / ASN facility / prefix-geo) here.
|
||||||
|
],
|
||||||
|
failure_reason=(
|
||||||
|
"Could not resolve BGP event coordinates: no source coords, owning"
|
||||||
|
" collector unknown, and no fallback resolver matched."
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_bgp_event_location(
|
||||||
|
*,
|
||||||
|
collector: str,
|
||||||
|
source_latitude: float | None = None,
|
||||||
|
source_longitude: float | None = None,
|
||||||
|
site: str | None = None,
|
||||||
|
operator: str | None = None,
|
||||||
|
peer_asn: int | None = None,
|
||||||
|
origin_asn: int | None = None,
|
||||||
|
prefix: str | None = None,
|
||||||
|
) -> ResolutionResult:
|
||||||
|
"""Resolve a BGP event to its renderable coordinates.
|
||||||
|
|
||||||
|
The ``peer_asn`` / ``origin_asn`` / ``prefix`` arguments are accepted
|
||||||
|
today so future resolvers (ASN→facility, prefix→geo) can consume them
|
||||||
|
without callers needing to change.
|
||||||
|
"""
|
||||||
|
query = LocationQuery(
|
||||||
|
name=collector or None,
|
||||||
|
aliases=tuple(filter(None, (collector,))),
|
||||||
|
source_latitude=source_latitude,
|
||||||
|
source_longitude=source_longitude,
|
||||||
|
extra={
|
||||||
|
"collector": collector or "",
|
||||||
|
"site": coerce_str(site),
|
||||||
|
"operator": coerce_str(operator),
|
||||||
|
"peer_asn": peer_asn,
|
||||||
|
"origin_asn": origin_asn,
|
||||||
|
"prefix": coerce_str(prefix),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return BGP_EVENT_PIPELINE.resolve_best(query)
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_bgp_event_geo_dict(
|
||||||
|
collector: str,
|
||||||
|
*,
|
||||||
|
source_latitude: float | None = None,
|
||||||
|
source_longitude: float | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
"""Convenience wrapper returning the legacy ``collector_geo`` dict shape.
|
||||||
|
|
||||||
|
Preserves ``city``/``country``/``latitude``/``longitude`` keys (consumed
|
||||||
|
by existing detectors / enrichment / DB serialization) and adds
|
||||||
|
``precision``/``source``/``needs_confirmation`` for richer downstream use.
|
||||||
|
"""
|
||||||
|
result = resolve_bgp_event_location(
|
||||||
|
collector=collector,
|
||||||
|
source_latitude=source_latitude,
|
||||||
|
source_longitude=source_longitude,
|
||||||
|
)
|
||||||
|
candidate = result.location
|
||||||
|
if candidate is None:
|
||||||
|
return {}
|
||||||
|
return {
|
||||||
|
"city": candidate.city,
|
||||||
|
"country": candidate.country,
|
||||||
|
"latitude": candidate.latitude,
|
||||||
|
"longitude": candidate.longitude,
|
||||||
|
"precision": candidate.precision,
|
||||||
|
"source": candidate.source,
|
||||||
|
"needs_confirmation": candidate.needs_confirmation,
|
||||||
|
"matched_location_name": candidate.matched_location_name,
|
||||||
|
"confidence": candidate.confidence,
|
||||||
|
}
|
||||||
117
backend/app/services/business_logs.py
Normal file
117
backend/app/services/business_logs.py
Normal file
@@ -0,0 +1,117 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
from collections.abc import Mapping
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from app.core.logging import PlanetLoggerAdapter, sanitize_log_value
|
||||||
|
from app.core.request_context import get_request_id
|
||||||
|
from app.services.persistent_logs import record_system_log
|
||||||
|
|
||||||
|
|
||||||
|
LEVEL_METHODS = {
|
||||||
|
"debug": "debug_event",
|
||||||
|
"info": "info_event",
|
||||||
|
"warning": "warning_event",
|
||||||
|
"error": "error_event",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_business_level(level: str | None) -> str:
|
||||||
|
normalized = str(level or "info").strip().lower()
|
||||||
|
if normalized in {"warn", "warning"}:
|
||||||
|
return "warning"
|
||||||
|
if normalized in {"err", "error", "critical", "fatal"}:
|
||||||
|
return "error"
|
||||||
|
if normalized == "debug":
|
||||||
|
return "debug"
|
||||||
|
return "info"
|
||||||
|
|
||||||
|
|
||||||
|
def build_business_context(
|
||||||
|
context: Mapping[str, Any] | None = None,
|
||||||
|
**fields: Any,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
payload = dict(context or {})
|
||||||
|
for key, value in fields.items():
|
||||||
|
if value is not None:
|
||||||
|
payload[key] = value
|
||||||
|
return sanitize_log_value(payload)
|
||||||
|
|
||||||
|
|
||||||
|
async def emit_business_log(
|
||||||
|
logger: PlanetLoggerAdapter,
|
||||||
|
*,
|
||||||
|
event: str,
|
||||||
|
message: str,
|
||||||
|
category: str,
|
||||||
|
level: str = "info",
|
||||||
|
source: str = "backend",
|
||||||
|
service: str | None = None,
|
||||||
|
module: str | None = None,
|
||||||
|
request_id: str | None = None,
|
||||||
|
user_id: int | None = None,
|
||||||
|
context: Mapping[str, Any] | None = None,
|
||||||
|
) -> None:
|
||||||
|
normalized_level = normalize_business_level(level)
|
||||||
|
safe_context = build_business_context(context)
|
||||||
|
log_method = getattr(logger, LEVEL_METHODS[normalized_level])
|
||||||
|
log_method(message, event=event, context=safe_context)
|
||||||
|
await record_system_log(
|
||||||
|
source=source,
|
||||||
|
level=normalized_level,
|
||||||
|
message=message,
|
||||||
|
service=service,
|
||||||
|
module=module,
|
||||||
|
event=event,
|
||||||
|
request_id=request_id or get_request_id(),
|
||||||
|
user_id=user_id,
|
||||||
|
category=category,
|
||||||
|
context=safe_context,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def emit_business_log_background(
|
||||||
|
logger: PlanetLoggerAdapter,
|
||||||
|
*,
|
||||||
|
event: str,
|
||||||
|
message: str,
|
||||||
|
category: str,
|
||||||
|
level: str = "info",
|
||||||
|
source: str = "backend",
|
||||||
|
service: str | None = None,
|
||||||
|
module: str | None = None,
|
||||||
|
request_id: str | None = None,
|
||||||
|
user_id: int | None = None,
|
||||||
|
context: Mapping[str, Any] | None = None,
|
||||||
|
) -> None:
|
||||||
|
normalized_level = normalize_business_level(level)
|
||||||
|
safe_context = build_business_context(context)
|
||||||
|
log_method = getattr(logger, LEVEL_METHODS[normalized_level])
|
||||||
|
log_method(message, event=event, context=safe_context)
|
||||||
|
try:
|
||||||
|
loop = asyncio.get_running_loop()
|
||||||
|
except RuntimeError:
|
||||||
|
return
|
||||||
|
loop.create_task(
|
||||||
|
record_system_log(
|
||||||
|
source=source,
|
||||||
|
level=normalized_level,
|
||||||
|
message=message,
|
||||||
|
service=service,
|
||||||
|
module=module,
|
||||||
|
event=event,
|
||||||
|
request_id=request_id or get_request_id(),
|
||||||
|
user_id=user_id,
|
||||||
|
category=category,
|
||||||
|
context=safe_context,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def exception_context(exc: BaseException, context: Mapping[str, Any] | None = None) -> dict[str, Any]:
|
||||||
|
return build_business_context(
|
||||||
|
context,
|
||||||
|
error_type=type(exc).__name__,
|
||||||
|
error=str(exc),
|
||||||
|
)
|
||||||
@@ -36,6 +36,9 @@ from app.services.collectors.iptoasn import IPtoASNPrefixGeoCollector
|
|||||||
from app.services.collectors.opengeofeed import OpenGeoFeedPrefixGeoCollector
|
from app.services.collectors.opengeofeed import OpenGeoFeedPrefixGeoCollector
|
||||||
from app.services.collectors.nro_delegated import NRODelegatedPrefixGeoCollector
|
from app.services.collectors.nro_delegated import NRODelegatedPrefixGeoCollector
|
||||||
from app.services.collectors.news_live_streams import NewsLiveStreamsCollector
|
from app.services.collectors.news_live_streams import NewsLiveStreamsCollector
|
||||||
|
from app.services.collectors.media_news_archive import MediaNewsArchiveCollector
|
||||||
|
from app.services.collectors.aisstream import AISStreamCollector
|
||||||
|
from app.services.collectors.vessel_ais import VesselAISCollector
|
||||||
|
|
||||||
collector_registry.register(TOP500Collector())
|
collector_registry.register(TOP500Collector())
|
||||||
collector_registry.register(EpochAIGPUCollector())
|
collector_registry.register(EpochAIGPUCollector())
|
||||||
@@ -63,3 +66,43 @@ collector_registry.register(IPtoASNPrefixGeoCollector())
|
|||||||
collector_registry.register(OpenGeoFeedPrefixGeoCollector())
|
collector_registry.register(OpenGeoFeedPrefixGeoCollector())
|
||||||
collector_registry.register(NRODelegatedPrefixGeoCollector())
|
collector_registry.register(NRODelegatedPrefixGeoCollector())
|
||||||
collector_registry.register(NewsLiveStreamsCollector())
|
collector_registry.register(NewsLiveStreamsCollector())
|
||||||
|
collector_registry.register(MediaNewsArchiveCollector())
|
||||||
|
collector_registry.register(VesselAISCollector())
|
||||||
|
collector_registry.register(AISStreamCollector())
|
||||||
|
|
||||||
|
__all__ = [
|
||||||
|
"BaseCollector",
|
||||||
|
"HTTPCollector",
|
||||||
|
"IntervalCollector",
|
||||||
|
"collector_registry",
|
||||||
|
"CollectorRegistry",
|
||||||
|
"TOP500Collector",
|
||||||
|
"EpochAIGPUCollector",
|
||||||
|
"HuggingFaceModelCollector",
|
||||||
|
"HuggingFaceDatasetCollector",
|
||||||
|
"HuggingFaceSpacesCollector",
|
||||||
|
"PeeringDBIXPCollector",
|
||||||
|
"PeeringDBNetworkCollector",
|
||||||
|
"PeeringDBFacilityCollector",
|
||||||
|
"TeleGeographyCableCollector",
|
||||||
|
"TeleGeographyLandingPointCollector",
|
||||||
|
"TeleGeographyCableSystemCollector",
|
||||||
|
"CloudflareRadarDeviceCollector",
|
||||||
|
"CloudflareRadarTrafficCollector",
|
||||||
|
"CloudflareRadarTopASCollector",
|
||||||
|
"ArcGISCableCollector",
|
||||||
|
"FAOLandingPointCollector",
|
||||||
|
"ArcGISLandingPointCollector",
|
||||||
|
"ArcGISCableLandingRelationCollector",
|
||||||
|
"SpaceTrackTLECollector",
|
||||||
|
"CelesTrakTLECollector",
|
||||||
|
"RISLiveCollector",
|
||||||
|
"BGPStreamBackfillCollector",
|
||||||
|
"IPtoASNPrefixGeoCollector",
|
||||||
|
"OpenGeoFeedPrefixGeoCollector",
|
||||||
|
"NRODelegatedPrefixGeoCollector",
|
||||||
|
"NewsLiveStreamsCollector",
|
||||||
|
"MediaNewsArchiveCollector",
|
||||||
|
"VesselAISCollector",
|
||||||
|
"AISStreamCollector",
|
||||||
|
]
|
||||||
|
|||||||
491
backend/app/services/collectors/aisstream.py
Normal file
491
backend/app/services/collectors/aisstream.py
Normal file
@@ -0,0 +1,491 @@
|
|||||||
|
"""AISStream WebSocket collector for realtime vessel AIS observations."""
|
||||||
|
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
import asyncio
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from sqlalchemy import select
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.data_sources import get_data_sources_config
|
||||||
|
from app.core.time import to_iso8601_utc
|
||||||
|
from app.core.websocket.broadcaster import broadcaster
|
||||||
|
from app.models.datasource_config import DataSourceConfig
|
||||||
|
from app.models.task import CollectionTask
|
||||||
|
from app.services.collectors.base import BaseCollector
|
||||||
|
from app.services.vessel_ais_aggregation import (
|
||||||
|
AISSTREAM_DELIVERY_MODE,
|
||||||
|
AISSTREAM_TRANSPORT,
|
||||||
|
record_vessel_ais_observation,
|
||||||
|
update_ais_source_health,
|
||||||
|
)
|
||||||
|
from app.services.vessel_types import normalize_vessel_type_name
|
||||||
|
|
||||||
|
DEFAULT_AISSTREAM_URL = "wss://stream.aisstream.io/v0/stream"
|
||||||
|
DEFAULT_BOUNDING_BOXES = [[[-90, -180], [90, 180]]]
|
||||||
|
DEFAULT_MESSAGE_TYPES = ["PositionReport", "ShipStaticData"]
|
||||||
|
|
||||||
|
|
||||||
|
class AISStreamCollector(BaseCollector):
|
||||||
|
"""Collect AISStream WebSocket messages into the raw AIS observation layer."""
|
||||||
|
|
||||||
|
name = "aisstream_vessels"
|
||||||
|
priority = "P1"
|
||||||
|
module = "L4"
|
||||||
|
frequency_hours = 1
|
||||||
|
data_type = "vessel_ais"
|
||||||
|
fail_on_empty = False
|
||||||
|
|
||||||
|
async def _load_datasource_config(self) -> DataSourceConfig | None:
|
||||||
|
if self._db_session is None:
|
||||||
|
return None
|
||||||
|
result = await self._db_session.execute(
|
||||||
|
select(DataSourceConfig)
|
||||||
|
.where(DataSourceConfig.name == self.name)
|
||||||
|
.where(DataSourceConfig.is_active.is_(True))
|
||||||
|
)
|
||||||
|
return result.scalar_one_or_none()
|
||||||
|
|
||||||
|
async def _get_effective_config(self) -> dict[str, Any]:
|
||||||
|
datasource_config = await self._load_datasource_config()
|
||||||
|
config = dict(datasource_config.config or {}) if datasource_config else {}
|
||||||
|
auth_config = dict(datasource_config.auth_config or {}) if datasource_config else {}
|
||||||
|
endpoint = (
|
||||||
|
(datasource_config.endpoint if datasource_config else None)
|
||||||
|
or self._resolved_url
|
||||||
|
or get_data_sources_config().get_yaml_url(self.name)
|
||||||
|
or DEFAULT_AISSTREAM_URL
|
||||||
|
)
|
||||||
|
api_key = (
|
||||||
|
auth_config.get("api_key")
|
||||||
|
or config.get("api_key")
|
||||||
|
or os.getenv("AISSTREAM_API_KEY")
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"endpoint": endpoint,
|
||||||
|
"api_key": api_key,
|
||||||
|
"bounding_boxes": config.get("bounding_boxes") or DEFAULT_BOUNDING_BOXES,
|
||||||
|
"message_types": config.get("message_types") or DEFAULT_MESSAGE_TYPES,
|
||||||
|
"max_messages": int(config.get("max_messages") or 500),
|
||||||
|
"streaming_enabled": config.get("streaming_enabled", True) is not False,
|
||||||
|
"streaming_commit_interval": int(config.get("streaming_commit_interval") or 1),
|
||||||
|
"streaming_max_messages": int(config.get("streaming_max_messages") or 0),
|
||||||
|
"reconnect_delay_seconds": float(config.get("reconnect_delay_seconds") or 5),
|
||||||
|
"receive_timeout_seconds": float(config.get("receive_timeout_seconds") or 30),
|
||||||
|
}
|
||||||
|
|
||||||
|
def _build_subscription(self, config: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"APIKey": config["api_key"],
|
||||||
|
"BoundingBoxes": config["bounding_boxes"],
|
||||||
|
"FilterMessageTypes": config["message_types"],
|
||||||
|
}
|
||||||
|
|
||||||
|
async def fetch(self) -> list[dict[str, Any]]:
|
||||||
|
config = await self._get_effective_config()
|
||||||
|
if not config["api_key"]:
|
||||||
|
raise RuntimeError("AISStream API key is not configured")
|
||||||
|
|
||||||
|
try:
|
||||||
|
import websockets
|
||||||
|
except ImportError as exc:
|
||||||
|
raise RuntimeError("Python package 'websockets' is required for AISStream") from exc
|
||||||
|
|
||||||
|
subscription = self._build_subscription(config)
|
||||||
|
|
||||||
|
messages: list[dict[str, Any]] = []
|
||||||
|
try:
|
||||||
|
async with websockets.connect(config["endpoint"]) as websocket:
|
||||||
|
await websocket.send(json.dumps(subscription))
|
||||||
|
while len(messages) < config["max_messages"]:
|
||||||
|
try:
|
||||||
|
raw_message = await asyncio.wait_for(
|
||||||
|
websocket.recv(),
|
||||||
|
timeout=config["receive_timeout_seconds"],
|
||||||
|
)
|
||||||
|
except TimeoutError:
|
||||||
|
break
|
||||||
|
payload = json.loads(raw_message)
|
||||||
|
if isinstance(payload, dict):
|
||||||
|
messages.append(payload)
|
||||||
|
except Exception as exc:
|
||||||
|
if self._db_session is not None:
|
||||||
|
await update_ais_source_health(
|
||||||
|
self._db_session,
|
||||||
|
source=self.name,
|
||||||
|
connection_state="disconnected",
|
||||||
|
last_error=f"{exc.__class__.__name__}: {exc}",
|
||||||
|
)
|
||||||
|
await self._db_session.commit()
|
||||||
|
raise
|
||||||
|
|
||||||
|
return messages
|
||||||
|
|
||||||
|
async def run(self, db: AsyncSession) -> dict[str, Any]:
|
||||||
|
"""Run AISStream as a long-lived streaming collector by default."""
|
||||||
|
config = await self._get_effective_config()
|
||||||
|
if not config.get("streaming_enabled", True):
|
||||||
|
return await super().run(db)
|
||||||
|
if not config["api_key"]:
|
||||||
|
return {"status": "failed", "error": "AISStream API key is not configured"}
|
||||||
|
|
||||||
|
from app.services.collectors.registry import collector_registry
|
||||||
|
|
||||||
|
if not collector_registry.is_active(self.name):
|
||||||
|
return {"status": "skipped", "reason": "Collector is disabled"}
|
||||||
|
|
||||||
|
try:
|
||||||
|
import websockets
|
||||||
|
except ImportError:
|
||||||
|
return {"status": "failed", "error": "Python package 'websockets' is required for AISStream"}
|
||||||
|
|
||||||
|
start_time = datetime.now(UTC)
|
||||||
|
task = CollectionTask(
|
||||||
|
datasource_id=getattr(self, "_datasource_id", 1),
|
||||||
|
status="running",
|
||||||
|
phase="connecting",
|
||||||
|
phase_message="正在连接 AISStream 实时流",
|
||||||
|
phase_unit="messages",
|
||||||
|
started_at=start_time,
|
||||||
|
)
|
||||||
|
db.add(task)
|
||||||
|
await db.commit()
|
||||||
|
self._current_task = task
|
||||||
|
self._db_session = db
|
||||||
|
self._last_broadcast_progress = None
|
||||||
|
await self.resolve_url(db)
|
||||||
|
await self._publish_task_update(force=True)
|
||||||
|
|
||||||
|
records_added = 0
|
||||||
|
messages_seen = 0
|
||||||
|
unique_mmsi: set[str] = set()
|
||||||
|
reconnect_delay = config["reconnect_delay_seconds"]
|
||||||
|
|
||||||
|
try:
|
||||||
|
while True:
|
||||||
|
config = await self._get_effective_config()
|
||||||
|
subscription = self._build_subscription(config)
|
||||||
|
try:
|
||||||
|
await update_ais_source_health(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
connection_state="connecting",
|
||||||
|
)
|
||||||
|
await self.set_phase("connecting", message="正在连接 AISStream 实时流")
|
||||||
|
await db.commit()
|
||||||
|
|
||||||
|
async with websockets.connect(config["endpoint"]) as websocket:
|
||||||
|
await websocket.send(json.dumps(subscription))
|
||||||
|
await update_ais_source_health(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
connection_state="connected",
|
||||||
|
last_success_at=datetime.now(UTC),
|
||||||
|
)
|
||||||
|
await self.set_phase(
|
||||||
|
"streaming",
|
||||||
|
message="正在接收 AISStream 实时消息",
|
||||||
|
reset_progress=False,
|
||||||
|
)
|
||||||
|
await db.commit()
|
||||||
|
|
||||||
|
while True:
|
||||||
|
try:
|
||||||
|
raw_message = await asyncio.wait_for(
|
||||||
|
websocket.recv(),
|
||||||
|
timeout=config["receive_timeout_seconds"],
|
||||||
|
)
|
||||||
|
except TimeoutError:
|
||||||
|
await update_ais_source_health(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
connection_state="connected",
|
||||||
|
last_success_at=datetime.now(UTC),
|
||||||
|
)
|
||||||
|
await db.commit()
|
||||||
|
continue
|
||||||
|
|
||||||
|
payload = json.loads(raw_message)
|
||||||
|
if not isinstance(payload, dict):
|
||||||
|
continue
|
||||||
|
messages_seen += 1
|
||||||
|
record = self._normalize_message(payload)
|
||||||
|
if not record:
|
||||||
|
continue
|
||||||
|
unique_mmsi.add(str(record["mmsi"]))
|
||||||
|
created = await self._save_stream_record(db, record)
|
||||||
|
if created:
|
||||||
|
records_added += 1
|
||||||
|
|
||||||
|
task.records_processed = messages_seen
|
||||||
|
task.total_records = None
|
||||||
|
task.progress = None
|
||||||
|
task.phase = "streaming"
|
||||||
|
task.phase_message = "正在接收 AISStream 实时消息"
|
||||||
|
task.phase_current = messages_seen
|
||||||
|
task.phase_total = None
|
||||||
|
task.phase_unit = "messages"
|
||||||
|
await self._publish_task_update(force=True)
|
||||||
|
|
||||||
|
if config["streaming_max_messages"] and messages_seen >= config["streaming_max_messages"]:
|
||||||
|
task.status = "success"
|
||||||
|
task.phase = "stopped"
|
||||||
|
task.phase_message = "AISStream 测试流已停止"
|
||||||
|
task.completed_at = datetime.now(UTC)
|
||||||
|
await db.commit()
|
||||||
|
await self._publish_task_update(force=True)
|
||||||
|
return {
|
||||||
|
"status": "success",
|
||||||
|
"task_id": task.id,
|
||||||
|
"records_processed": records_added,
|
||||||
|
"messages_seen": messages_seen,
|
||||||
|
"unique_mmsi": len(unique_mmsi),
|
||||||
|
"execution_time_seconds": (datetime.now(UTC) - start_time).total_seconds(),
|
||||||
|
}
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
raise
|
||||||
|
except Exception as exc:
|
||||||
|
await update_ais_source_health(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
connection_state="reconnecting",
|
||||||
|
last_error=f"{exc.__class__.__name__}: {exc}",
|
||||||
|
)
|
||||||
|
task.phase = "reconnecting"
|
||||||
|
task.phase_message = "AISStream 连接中断,正在重连"
|
||||||
|
task.error_message = f"{exc.__class__.__name__}: {exc}"
|
||||||
|
await db.commit()
|
||||||
|
await self._publish_task_update(force=True)
|
||||||
|
await asyncio.sleep(reconnect_delay)
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
task.status = "cancelled"
|
||||||
|
task.phase = "stopped"
|
||||||
|
task.phase_message = "AISStream 实时流已停止"
|
||||||
|
task.completed_at = datetime.now(UTC)
|
||||||
|
await update_ais_source_health(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
connection_state="disconnected",
|
||||||
|
last_error=None,
|
||||||
|
)
|
||||||
|
await db.commit()
|
||||||
|
await self._publish_task_update(force=True)
|
||||||
|
raise
|
||||||
|
|
||||||
|
def transform(self, raw_data: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||||
|
records = []
|
||||||
|
for item in raw_data:
|
||||||
|
record = self._normalize_message(item)
|
||||||
|
if record:
|
||||||
|
records.append(record)
|
||||||
|
return records
|
||||||
|
|
||||||
|
async def _save_data(
|
||||||
|
self,
|
||||||
|
db: AsyncSession,
|
||||||
|
data: list[dict[str, Any]],
|
||||||
|
task_id: int | None = None,
|
||||||
|
snapshot_id: int | None = None,
|
||||||
|
) -> int:
|
||||||
|
now = datetime.now(UTC)
|
||||||
|
records_added = 0
|
||||||
|
latest_observed_at = now
|
||||||
|
for index, item in enumerate(data):
|
||||||
|
observed_at = item.get("received_at") or now
|
||||||
|
observation = await record_vessel_ais_observation(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
normalized_payload=item,
|
||||||
|
raw_payload=item.get("_raw_payload") or item,
|
||||||
|
delivery_mode=AISSTREAM_DELIVERY_MODE,
|
||||||
|
transport=AISSTREAM_TRANSPORT,
|
||||||
|
message_type=item.get("_message_type") or "PositionReport",
|
||||||
|
source_message_id=item.get("_source_message_id"),
|
||||||
|
observed_at=observed_at,
|
||||||
|
collected_at=now,
|
||||||
|
)
|
||||||
|
if observation is not None:
|
||||||
|
records_added += 1
|
||||||
|
if isinstance(observed_at, datetime) and observed_at > latest_observed_at:
|
||||||
|
latest_observed_at = observed_at
|
||||||
|
if (index + 1) % 1000 == 0:
|
||||||
|
await self.update_progress(index + 1, commit=True)
|
||||||
|
|
||||||
|
await update_ais_source_health(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
connection_state="connected",
|
||||||
|
observed_count=len(data),
|
||||||
|
last_seen_at=latest_observed_at,
|
||||||
|
last_success_at=now if data else None,
|
||||||
|
lag_seconds=max((now - latest_observed_at).total_seconds(), 0),
|
||||||
|
)
|
||||||
|
await db.commit()
|
||||||
|
await self.update_progress(records_added, force=True)
|
||||||
|
return records_added
|
||||||
|
|
||||||
|
async def _save_stream_record(self, db: AsyncSession, item: dict[str, Any]) -> bool:
|
||||||
|
now = datetime.now(UTC)
|
||||||
|
observed_at = item.get("received_at") or now
|
||||||
|
observation = await record_vessel_ais_observation(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
normalized_payload=item,
|
||||||
|
raw_payload=item.get("_raw_payload") or item,
|
||||||
|
delivery_mode=AISSTREAM_DELIVERY_MODE,
|
||||||
|
transport=AISSTREAM_TRANSPORT,
|
||||||
|
message_type=item.get("_message_type") or "PositionReport",
|
||||||
|
source_message_id=item.get("_source_message_id"),
|
||||||
|
observed_at=observed_at,
|
||||||
|
collected_at=now,
|
||||||
|
)
|
||||||
|
await update_ais_source_health(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
connection_state="connected",
|
||||||
|
observed_count=1,
|
||||||
|
last_seen_at=observed_at if isinstance(observed_at, datetime) else now,
|
||||||
|
last_success_at=now,
|
||||||
|
lag_seconds=max((now - observed_at).total_seconds(), 0) if isinstance(observed_at, datetime) else None,
|
||||||
|
)
|
||||||
|
await db.commit()
|
||||||
|
await self._broadcast_vessel_delta(item, created=observation is not None)
|
||||||
|
return observation is not None
|
||||||
|
|
||||||
|
async def _broadcast_vessel_delta(self, item: dict[str, Any], *, created: bool) -> None:
|
||||||
|
await broadcaster.broadcast_custom(
|
||||||
|
"vessels",
|
||||||
|
{
|
||||||
|
"action": "upsert",
|
||||||
|
"source": self.name,
|
||||||
|
"created": created,
|
||||||
|
"vessels": [
|
||||||
|
{
|
||||||
|
"mmsi": item.get("mmsi"),
|
||||||
|
"mmsi_display": str(item.get("mmsi")) if item.get("mmsi") is not None else None,
|
||||||
|
"name": item.get("name"),
|
||||||
|
"lat": item.get("lat"),
|
||||||
|
"lon": item.get("lon"),
|
||||||
|
"sog": item.get("sog"),
|
||||||
|
"cog": item.get("cog"),
|
||||||
|
"heading": item.get("heading"),
|
||||||
|
"nav_status": item.get("nav_status"),
|
||||||
|
"vessel_type": item.get("vessel_type"),
|
||||||
|
"vessel_type_name": item.get("vessel_type_name"),
|
||||||
|
"received_at": to_iso8601_utc(item.get("received_at")),
|
||||||
|
}
|
||||||
|
],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
def _normalize_message(self, item: dict[str, Any]) -> dict[str, Any] | None:
|
||||||
|
message_type = str(item.get("MessageType") or item.get("message_type") or "")
|
||||||
|
metadata = item.get("MetaData") if isinstance(item.get("MetaData"), dict) else {}
|
||||||
|
message = item.get("Message") if isinstance(item.get("Message"), dict) else {}
|
||||||
|
body = message.get(message_type) if isinstance(message.get(message_type), dict) else message
|
||||||
|
if not isinstance(body, dict):
|
||||||
|
body = {}
|
||||||
|
|
||||||
|
mmsi = _as_int(_pick(metadata, "MMSI", "mmsi") or _pick(body, "MMSI", "mmsi"))
|
||||||
|
if mmsi is None:
|
||||||
|
return None
|
||||||
|
|
||||||
|
received_at = _parse_datetime(
|
||||||
|
_pick(metadata, "time_utc", "Time_UTC", "timestamp")
|
||||||
|
or _pick(body, "Timestamp", "timestamp", "time")
|
||||||
|
)
|
||||||
|
ship_name = _clean_text(
|
||||||
|
_pick(body, "Name", "ShipName", "name")
|
||||||
|
or _pick(metadata, "ShipName", "ship_name", "name")
|
||||||
|
)
|
||||||
|
record: dict[str, Any] = {
|
||||||
|
"mmsi": mmsi,
|
||||||
|
"received_at": received_at,
|
||||||
|
"_message_type": message_type or None,
|
||||||
|
"_source_message_id": item.get("MessageID") or item.get("message_id"),
|
||||||
|
"_raw_payload": item,
|
||||||
|
}
|
||||||
|
|
||||||
|
lat = _as_float(_pick(body, "Latitude", "lat", "latitude"))
|
||||||
|
lon = _as_float(_pick(body, "Longitude", "lon", "lng", "longitude"))
|
||||||
|
if lat is not None and lon is not None:
|
||||||
|
if not (-90 <= lat <= 90 and -180 <= lon <= 180):
|
||||||
|
return None
|
||||||
|
record.update(
|
||||||
|
{
|
||||||
|
"lat": lat,
|
||||||
|
"lon": lon,
|
||||||
|
"sog": _as_float(_pick(body, "Sog", "SOG", "speedOverGround")),
|
||||||
|
"cog": _as_float(_pick(body, "Cog", "COG", "courseOverGround")),
|
||||||
|
"heading": _as_int(_pick(body, "TrueHeading", "Heading", "heading")),
|
||||||
|
"nav_status": _as_int(_pick(body, "NavigationalStatus", "nav_status")),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
vessel_type = _as_int(_pick(body, "Type", "ShipType", "vessel_type"))
|
||||||
|
record.update(
|
||||||
|
{
|
||||||
|
"name": ship_name,
|
||||||
|
"callsign": _pick(body, "CallSign", "callsign"),
|
||||||
|
"imo": _as_int(_pick(body, "ImoNumber", "IMO", "imo")),
|
||||||
|
"vessel_type": vessel_type,
|
||||||
|
"vessel_type_name": _pick(body, "TypeName", "ShipTypeName", "vessel_type_name")
|
||||||
|
or normalize_vessel_type_name(vessel_type),
|
||||||
|
"length": _as_float(_pick(body, "DimensionToBow", "Length", "length")),
|
||||||
|
"width": _as_float(_pick(body, "DimensionToPort", "Width", "width")),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return record
|
||||||
|
|
||||||
|
|
||||||
|
def _pick(item: dict[str, Any], *keys: str) -> Any:
|
||||||
|
for key in keys:
|
||||||
|
if key in item and item[key] not in (None, ""):
|
||||||
|
return item[key]
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _clean_text(value: Any) -> str | None:
|
||||||
|
if value in (None, ""):
|
||||||
|
return None
|
||||||
|
text = str(value).strip()
|
||||||
|
return text or None
|
||||||
|
|
||||||
|
|
||||||
|
def _as_float(value: Any) -> float | None:
|
||||||
|
try:
|
||||||
|
if value in (None, ""):
|
||||||
|
return None
|
||||||
|
return float(value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _as_int(value: Any) -> int | None:
|
||||||
|
try:
|
||||||
|
if value in (None, ""):
|
||||||
|
return None
|
||||||
|
return int(float(value))
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_datetime(value: Any) -> datetime | None:
|
||||||
|
if isinstance(value, datetime):
|
||||||
|
return value if value.tzinfo else value.replace(tzinfo=UTC)
|
||||||
|
if not value:
|
||||||
|
return None
|
||||||
|
if isinstance(value, (int, float)):
|
||||||
|
timestamp = float(value)
|
||||||
|
if timestamp > 10_000_000_000:
|
||||||
|
timestamp /= 1000
|
||||||
|
return datetime.fromtimestamp(timestamp, UTC)
|
||||||
|
if isinstance(value, str):
|
||||||
|
try:
|
||||||
|
parsed = datetime.fromisoformat(value.replace("Z", "+00:00"))
|
||||||
|
return parsed if parsed.tzinfo else parsed.replace(tzinfo=UTC)
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
|
return None
|
||||||
@@ -4,15 +4,22 @@ import asyncio
|
|||||||
from abc import ABC, abstractmethod
|
from abc import ABC, abstractmethod
|
||||||
from typing import Dict, List, Any, Optional
|
from typing import Dict, List, Any, Optional
|
||||||
from datetime import UTC, datetime
|
from datetime import UTC, datetime
|
||||||
|
from time import perf_counter
|
||||||
|
from urllib.parse import urlparse
|
||||||
import httpx
|
import httpx
|
||||||
from sqlalchemy import select, text
|
from sqlalchemy import select, text
|
||||||
from sqlalchemy.ext.asyncio import AsyncSession
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
from app.core.collected_data_fields import build_dynamic_metadata, get_record_field
|
from app.core.collected_data_fields import build_dynamic_metadata, get_record_field
|
||||||
from app.core.config import settings
|
|
||||||
from app.core.countries import normalize_country
|
from app.core.countries import normalize_country
|
||||||
|
from app.core.logging import get_logger
|
||||||
from app.core.time import to_iso8601_utc
|
from app.core.time import to_iso8601_utc
|
||||||
from app.core.websocket.broadcaster import broadcaster
|
from app.core.websocket.broadcaster import broadcaster
|
||||||
|
from app.services.business_logs import emit_business_log, exception_context
|
||||||
|
from app.services.earth_layer_adapters import get_earth_update_layers_for_source
|
||||||
|
|
||||||
|
|
||||||
|
logger = get_logger(__name__, service="collector")
|
||||||
|
|
||||||
|
|
||||||
class BaseCollector(ABC):
|
class BaseCollector(ABC):
|
||||||
@@ -31,6 +38,7 @@ class BaseCollector(ABC):
|
|||||||
self._datasource_id = 1
|
self._datasource_id = 1
|
||||||
self._resolved_url: Optional[str] = None
|
self._resolved_url: Optional[str] = None
|
||||||
self._last_broadcast_progress: Optional[int] = None
|
self._last_broadcast_progress: Optional[int] = None
|
||||||
|
self._last_save_summary: dict[str, int] = {}
|
||||||
|
|
||||||
async def resolve_url(self, db: AsyncSession) -> None:
|
async def resolve_url(self, db: AsyncSession) -> None:
|
||||||
from app.core.data_sources import get_data_sources_config
|
from app.core.data_sources import get_data_sources_config
|
||||||
@@ -54,6 +62,11 @@ class BaseCollector(ABC):
|
|||||||
"task_id": self._current_task.id,
|
"task_id": self._current_task.id,
|
||||||
"status": self._current_task.status,
|
"status": self._current_task.status,
|
||||||
"phase": self._current_task.phase,
|
"phase": self._current_task.phase,
|
||||||
|
"phase_progress": self._current_task.phase_progress,
|
||||||
|
"phase_message": self._current_task.phase_message,
|
||||||
|
"phase_current": self._current_task.phase_current,
|
||||||
|
"phase_total": self._current_task.phase_total,
|
||||||
|
"phase_unit": self._current_task.phase_unit,
|
||||||
"progress": progress,
|
"progress": progress,
|
||||||
"records_processed": self._current_task.records_processed,
|
"records_processed": self._current_task.records_processed,
|
||||||
"total_records": self._current_task.total_records,
|
"total_records": self._current_task.total_records,
|
||||||
@@ -80,12 +93,52 @@ class BaseCollector(ABC):
|
|||||||
|
|
||||||
await self._publish_task_update(force=force)
|
await self._publish_task_update(force=force)
|
||||||
|
|
||||||
async def set_phase(self, phase: str):
|
async def set_phase(self, phase: str, *, message: str | None = None, reset_progress: bool = True):
|
||||||
if self._current_task and self._db_session:
|
if self._current_task and self._db_session:
|
||||||
self._current_task.phase = phase
|
self._current_task.phase = phase
|
||||||
|
self._current_task.phase_message = message
|
||||||
|
if reset_progress:
|
||||||
|
self._current_task.phase_progress = None
|
||||||
|
self._current_task.phase_current = None
|
||||||
|
self._current_task.phase_total = None
|
||||||
|
self._current_task.phase_unit = None
|
||||||
await self._db_session.commit()
|
await self._db_session.commit()
|
||||||
await self._publish_task_update(force=True)
|
await self._publish_task_update(force=True)
|
||||||
|
|
||||||
|
async def update_phase_progress(
|
||||||
|
self,
|
||||||
|
*,
|
||||||
|
current: int | None = None,
|
||||||
|
total: int | None = None,
|
||||||
|
unit: str | None = None,
|
||||||
|
message: str | None = None,
|
||||||
|
progress: float | None = None,
|
||||||
|
commit: bool = False,
|
||||||
|
force: bool = False,
|
||||||
|
):
|
||||||
|
"""Update progress for the current phase without changing task totals."""
|
||||||
|
if not self._current_task or not self._db_session:
|
||||||
|
return
|
||||||
|
|
||||||
|
if progress is None and current is not None and total and total > 0:
|
||||||
|
progress = (current / total) * 100
|
||||||
|
|
||||||
|
if progress is not None:
|
||||||
|
self._current_task.phase_progress = max(0.0, min(float(progress), 100.0))
|
||||||
|
if current is not None:
|
||||||
|
self._current_task.phase_current = max(0, int(current))
|
||||||
|
if total is not None:
|
||||||
|
self._current_task.phase_total = max(0, int(total))
|
||||||
|
if unit is not None:
|
||||||
|
self._current_task.phase_unit = unit
|
||||||
|
if message is not None:
|
||||||
|
self._current_task.phase_message = message
|
||||||
|
|
||||||
|
if commit:
|
||||||
|
await self._db_session.commit()
|
||||||
|
|
||||||
|
await self._publish_task_update(force=force)
|
||||||
|
|
||||||
@abstractmethod
|
@abstractmethod
|
||||||
async def fetch(self) -> List[Dict[str, Any]]:
|
async def fetch(self) -> List[Dict[str, Any]]:
|
||||||
"""Fetch raw data from source"""
|
"""Fetch raw data from source"""
|
||||||
@@ -141,7 +194,7 @@ class BaseCollector(ABC):
|
|||||||
|
|
||||||
result = await db.execute(
|
result = await db.execute(
|
||||||
select(DataSnapshot)
|
select(DataSnapshot)
|
||||||
.where(DataSnapshot.source == self.name, DataSnapshot.is_current == True)
|
.where(DataSnapshot.source == self.name, DataSnapshot.is_current.is_(True))
|
||||||
.order_by(DataSnapshot.completed_at.desc().nullslast(), DataSnapshot.id.desc())
|
.order_by(DataSnapshot.completed_at.desc().nullslast(), DataSnapshot.id.desc())
|
||||||
.limit(1)
|
.limit(1)
|
||||||
)
|
)
|
||||||
@@ -227,19 +280,39 @@ class BaseCollector(ABC):
|
|||||||
from app.models.data_snapshot import DataSnapshot
|
from app.models.data_snapshot import DataSnapshot
|
||||||
|
|
||||||
start_time = datetime.now(UTC)
|
start_time = datetime.now(UTC)
|
||||||
|
started_at = perf_counter()
|
||||||
datasource_id = getattr(self, "_datasource_id", 1)
|
datasource_id = getattr(self, "_datasource_id", 1)
|
||||||
snapshot_id: Optional[int] = None
|
snapshot_id: Optional[int] = None
|
||||||
|
|
||||||
if not collector_registry.is_active(self.name):
|
if not collector_registry.is_active(self.name):
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.run.skipped_disabled",
|
||||||
|
"Collector skipped because it is disabled",
|
||||||
|
level="info",
|
||||||
|
context={"status": "skipped", "reason": "disabled"},
|
||||||
|
)
|
||||||
return {"status": "skipped", "reason": "Collector is disabled"}
|
return {"status": "skipped", "reason": "Collector is disabled"}
|
||||||
|
|
||||||
task = CollectionTask(
|
task = self._current_task if isinstance(self._current_task, CollectionTask) else None
|
||||||
datasource_id=datasource_id,
|
if task is None:
|
||||||
status="running",
|
task = CollectionTask(
|
||||||
phase="queued",
|
datasource_id=datasource_id,
|
||||||
started_at=start_time,
|
source=self.name,
|
||||||
)
|
task_type="collect",
|
||||||
db.add(task)
|
status="running",
|
||||||
|
phase="queued",
|
||||||
|
started_at=start_time,
|
||||||
|
)
|
||||||
|
db.add(task)
|
||||||
|
else:
|
||||||
|
task.datasource_id = datasource_id
|
||||||
|
task.source = task.source or self.name
|
||||||
|
task.task_type = task.task_type or "collect"
|
||||||
|
task.status = "running"
|
||||||
|
task.phase = "queued"
|
||||||
|
task.started_at = task.started_at or start_time
|
||||||
|
task.completed_at = None
|
||||||
|
task.error_message = None
|
||||||
await db.commit()
|
await db.commit()
|
||||||
task_id = task.id
|
task_id = task.id
|
||||||
|
|
||||||
@@ -249,31 +322,103 @@ class BaseCollector(ABC):
|
|||||||
|
|
||||||
await self.resolve_url(db)
|
await self.resolve_url(db)
|
||||||
await self._publish_task_update(force=True)
|
await self._publish_task_update(force=True)
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.run.started",
|
||||||
|
"Collector run started",
|
||||||
|
context={"status": "running", "task_id": task_id},
|
||||||
|
)
|
||||||
|
|
||||||
try:
|
try:
|
||||||
await self.set_phase("fetching")
|
phase_started_at = perf_counter()
|
||||||
|
await self.set_phase("fetching", message="正在拉取原始数据")
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.phase.fetching.start",
|
||||||
|
"Collector fetch phase started",
|
||||||
|
context={"task_id": task_id, "snapshot_id": snapshot_id},
|
||||||
|
)
|
||||||
raw_data = await self.fetch()
|
raw_data = await self.fetch()
|
||||||
task.total_records = len(raw_data)
|
task.total_records = len(raw_data)
|
||||||
await db.commit()
|
await db.commit()
|
||||||
await self._publish_task_update(force=True)
|
await self._publish_task_update(force=True)
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.phase.fetching.success",
|
||||||
|
"Collector fetch phase completed",
|
||||||
|
context={
|
||||||
|
"task_id": task_id,
|
||||||
|
"raw_count": len(raw_data),
|
||||||
|
"duration_ms": self._duration_ms(phase_started_at),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
if self.fail_on_empty and not raw_data:
|
if self.fail_on_empty and not raw_data:
|
||||||
raise RuntimeError(f"Collector {self.name} returned no data")
|
raise RuntimeError(f"Collector {self.name} returned no data")
|
||||||
|
|
||||||
await self.set_phase("transforming")
|
phase_started_at = perf_counter()
|
||||||
|
await self.set_phase("transforming", message="正在转换采集数据")
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.phase.transforming.start",
|
||||||
|
"Collector transform phase started",
|
||||||
|
context={"task_id": task_id, "raw_count": len(raw_data)},
|
||||||
|
)
|
||||||
data = self.transform(raw_data)
|
data = self.transform(raw_data)
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.phase.transforming.success",
|
||||||
|
"Collector transform phase completed",
|
||||||
|
context={
|
||||||
|
"task_id": task_id,
|
||||||
|
"raw_count": len(raw_data),
|
||||||
|
"transformed_count": len(data),
|
||||||
|
"duration_ms": self._duration_ms(phase_started_at),
|
||||||
|
},
|
||||||
|
)
|
||||||
snapshot_id = await self._create_snapshot(db, task_id, data, start_time)
|
snapshot_id = await self._create_snapshot(db, task_id, data, start_time)
|
||||||
|
|
||||||
await self.set_phase("saving")
|
phase_started_at = perf_counter()
|
||||||
|
await self.set_phase("saving", message="正在保存采集数据")
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.phase.saving.start",
|
||||||
|
"Collector save phase started",
|
||||||
|
context={"task_id": task_id, "snapshot_id": snapshot_id, "transformed_count": len(data)},
|
||||||
|
)
|
||||||
records_count = await self._save_data(db, data, task_id=task_id, snapshot_id=snapshot_id)
|
records_count = await self._save_data(db, data, task_id=task_id, snapshot_id=snapshot_id)
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.phase.saving.success",
|
||||||
|
"Collector save phase completed",
|
||||||
|
context={
|
||||||
|
"task_id": task_id,
|
||||||
|
"snapshot_id": snapshot_id,
|
||||||
|
"saved_count": records_count,
|
||||||
|
**self._last_save_summary,
|
||||||
|
"duration_ms": self._duration_ms(phase_started_at),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
task.status = "success"
|
task.status = "success"
|
||||||
task.phase = "completed"
|
task.phase = "completed"
|
||||||
|
task.phase_progress = 100.0
|
||||||
|
task.phase_message = "采集完成"
|
||||||
|
task.phase_current = records_count
|
||||||
|
task.phase_total = records_count
|
||||||
|
task.phase_unit = "records"
|
||||||
task.records_processed = records_count
|
task.records_processed = records_count
|
||||||
task.progress = 100.0
|
task.progress = 100.0
|
||||||
task.completed_at = datetime.now(UTC)
|
task.completed_at = datetime.now(UTC)
|
||||||
await db.commit()
|
await db.commit()
|
||||||
await self._publish_task_update(force=True)
|
await self._publish_task_update(force=True)
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.run.completed",
|
||||||
|
"Collector run completed",
|
||||||
|
context={
|
||||||
|
"status": "success",
|
||||||
|
"task_id": task_id,
|
||||||
|
"snapshot_id": snapshot_id,
|
||||||
|
"raw_count": len(raw_data),
|
||||||
|
"transformed_count": len(data),
|
||||||
|
"saved_count": records_count,
|
||||||
|
**self._last_save_summary,
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
return {
|
return {
|
||||||
"status": "success",
|
"status": "success",
|
||||||
@@ -285,6 +430,7 @@ class BaseCollector(ABC):
|
|||||||
await db.rollback()
|
await db.rollback()
|
||||||
task.status = "cancelled"
|
task.status = "cancelled"
|
||||||
task.phase = "cancelled"
|
task.phase = "cancelled"
|
||||||
|
task.phase_message = "采集已取消"
|
||||||
task.error_message = "Collection cancelled by operator and rolled back"
|
task.error_message = "Collection cancelled by operator and rolled back"
|
||||||
task.completed_at = datetime.now(UTC)
|
task.completed_at = datetime.now(UTC)
|
||||||
if snapshot_id is not None:
|
if snapshot_id is not None:
|
||||||
@@ -296,11 +442,23 @@ class BaseCollector(ABC):
|
|||||||
)
|
)
|
||||||
await db.commit()
|
await db.commit()
|
||||||
await self._publish_task_update(force=True)
|
await self._publish_task_update(force=True)
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.run.cancelled",
|
||||||
|
"Collector run cancelled",
|
||||||
|
level="warning",
|
||||||
|
context={
|
||||||
|
"status": "cancelled",
|
||||||
|
"task_id": task_id,
|
||||||
|
"snapshot_id": snapshot_id,
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
)
|
||||||
raise
|
raise
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
await db.rollback()
|
await db.rollback()
|
||||||
task.status = "failed"
|
task.status = "failed"
|
||||||
task.phase = "failed"
|
task.phase = "failed"
|
||||||
|
task.phase_message = str(e)
|
||||||
task.error_message = str(e)
|
task.error_message = str(e)
|
||||||
task.completed_at = datetime.now(UTC)
|
task.completed_at = datetime.now(UTC)
|
||||||
if snapshot_id is not None:
|
if snapshot_id is not None:
|
||||||
@@ -311,6 +469,20 @@ class BaseCollector(ABC):
|
|||||||
snapshot.summary = {"error": str(e)}
|
snapshot.summary = {"error": str(e)}
|
||||||
await db.commit()
|
await db.commit()
|
||||||
await self._publish_task_update(force=True)
|
await self._publish_task_update(force=True)
|
||||||
|
await self._log_collection_event(
|
||||||
|
"collector.run.failed",
|
||||||
|
"Collector run failed",
|
||||||
|
level="error",
|
||||||
|
context=exception_context(
|
||||||
|
e,
|
||||||
|
{
|
||||||
|
"status": "failed",
|
||||||
|
"task_id": task_id,
|
||||||
|
"snapshot_id": snapshot_id,
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
return {
|
return {
|
||||||
"status": "failed",
|
"status": "failed",
|
||||||
@@ -331,6 +503,7 @@ class BaseCollector(ABC):
|
|||||||
from app.models.data_snapshot import DataSnapshot
|
from app.models.data_snapshot import DataSnapshot
|
||||||
|
|
||||||
if not data:
|
if not data:
|
||||||
|
self._last_save_summary = {"created": 0, "updated": 0, "unchanged": 0, "deleted": 0}
|
||||||
if snapshot_id is not None:
|
if snapshot_id is not None:
|
||||||
snapshot = await db.get(DataSnapshot, snapshot_id)
|
snapshot = await db.get(DataSnapshot, snapshot_id)
|
||||||
if snapshot:
|
if snapshot:
|
||||||
@@ -353,7 +526,7 @@ class BaseCollector(ABC):
|
|||||||
select(CollectedData)
|
select(CollectedData)
|
||||||
.where(
|
.where(
|
||||||
CollectedData.source == self.name,
|
CollectedData.source == self.name,
|
||||||
CollectedData.is_current == True,
|
CollectedData.is_current.is_(True),
|
||||||
)
|
)
|
||||||
.order_by(CollectedData.entity_key.asc(), CollectedData.collected_at.desc().nullslast(), CollectedData.id.desc())
|
.order_by(CollectedData.entity_key.asc(), CollectedData.collected_at.desc().nullslast(), CollectedData.id.desc())
|
||||||
)
|
)
|
||||||
@@ -477,11 +650,51 @@ class BaseCollector(ABC):
|
|||||||
"unchanged": unchanged_count,
|
"unchanged": unchanged_count,
|
||||||
"deleted": len(deleted_keys),
|
"deleted": len(deleted_keys),
|
||||||
}
|
}
|
||||||
|
self._last_save_summary = {
|
||||||
|
"created": created_count,
|
||||||
|
"updated": updated_count,
|
||||||
|
"unchanged": unchanged_count,
|
||||||
|
"deleted": len(deleted_keys),
|
||||||
|
}
|
||||||
|
else:
|
||||||
|
self._last_save_summary = {
|
||||||
|
"created": created_count,
|
||||||
|
"updated": updated_count,
|
||||||
|
"unchanged": unchanged_count,
|
||||||
|
"deleted": 0,
|
||||||
|
}
|
||||||
|
|
||||||
await db.commit()
|
await db.commit()
|
||||||
await self.update_progress(len(data), force=True)
|
await self.update_progress(len(data), force=True)
|
||||||
return records_added
|
return records_added
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _duration_ms(started_at: float) -> int:
|
||||||
|
return int((perf_counter() - started_at) * 1000)
|
||||||
|
|
||||||
|
async def _log_collection_event(
|
||||||
|
self,
|
||||||
|
event: str,
|
||||||
|
message: str,
|
||||||
|
*,
|
||||||
|
level: str = "info",
|
||||||
|
context: Dict[str, Any] | None = None,
|
||||||
|
) -> None:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event=event,
|
||||||
|
message=message,
|
||||||
|
category="collector",
|
||||||
|
level=level,
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
**(context or {}),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
async def save(self, db: AsyncSession, data: List[Dict[str, Any]]) -> int:
|
async def save(self, db: AsyncSession, data: List[Dict[str, Any]]) -> int:
|
||||||
"""Save data to database (legacy method, use _save_data instead)"""
|
"""Save data to database (legacy method, use _save_data instead)"""
|
||||||
return await self._save_data(db, data)
|
return await self._save_data(db, data)
|
||||||
@@ -494,10 +707,65 @@ class HTTPCollector(BaseCollector):
|
|||||||
headers: Dict[str, str] = {}
|
headers: Dict[str, str] = {}
|
||||||
|
|
||||||
async def fetch(self) -> List[Dict[str, Any]]:
|
async def fetch(self) -> List[Dict[str, Any]]:
|
||||||
|
started_at = perf_counter()
|
||||||
|
request_host = urlparse(self.base_url).netloc
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.http.fetch.start",
|
||||||
|
message="Collector HTTP request started",
|
||||||
|
category="collector",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"url_host": request_host,
|
||||||
|
},
|
||||||
|
)
|
||||||
async with httpx.AsyncClient(timeout=60.0) as client:
|
async with httpx.AsyncClient(timeout=60.0) as client:
|
||||||
response = await client.get(self.base_url, headers=self.headers)
|
try:
|
||||||
response.raise_for_status()
|
response = await client.get(self.base_url, headers=self.headers)
|
||||||
return self.parse_response(response.json())
|
response.raise_for_status()
|
||||||
|
payload = response.json()
|
||||||
|
parsed = self.parse_response(payload)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.http.fetch.success",
|
||||||
|
message="Collector HTTP request completed",
|
||||||
|
category="collector",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"url_host": request_host,
|
||||||
|
"status_code": response.status_code,
|
||||||
|
"response_bytes": len(response.content or b""),
|
||||||
|
"parsed_count": len(parsed),
|
||||||
|
"duration_ms": BaseCollector._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return parsed
|
||||||
|
except Exception as exc:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.http.fetch.failed",
|
||||||
|
message="Collector HTTP request failed",
|
||||||
|
category="collector",
|
||||||
|
level="error",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context=exception_context(
|
||||||
|
exc,
|
||||||
|
{
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"url_host": request_host,
|
||||||
|
"duration_ms": BaseCollector._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
),
|
||||||
|
)
|
||||||
|
raise
|
||||||
|
|
||||||
@abstractmethod
|
@abstractmethod
|
||||||
def parse_response(self, response: Dict[str, Any]) -> List[Dict[str, Any]]:
|
def parse_response(self, response: Dict[str, Any]) -> List[Dict[str, Any]]:
|
||||||
|
|||||||
@@ -13,7 +13,13 @@ from sqlalchemy.ext.asyncio import AsyncSession
|
|||||||
from app.models.bgp_anomaly import BGPAnomaly
|
from app.models.bgp_anomaly import BGPAnomaly
|
||||||
from app.models.bgp_observation import BGPObservation
|
from app.models.bgp_observation import BGPObservation
|
||||||
from app.models.collected_data import CollectedData
|
from app.models.collected_data import CollectedData
|
||||||
|
from app.services.bgp_collector_locations import (
|
||||||
|
RIPE_RIS_COLLECTOR_COORDS,
|
||||||
|
get_bgp_collector_location_dict,
|
||||||
|
)
|
||||||
|
from app.services.bgp_event_locations import resolve_bgp_event_geo_dict
|
||||||
from app.services.bgp_incidents import create_bgp_incidents_for_anomalies
|
from app.services.bgp_incidents import create_bgp_incidents_for_anomalies
|
||||||
|
from app.services.earth_layer_cache import invalidate_earth_layer_cache_for_source
|
||||||
from app.services.bgp_detectors import (
|
from app.services.bgp_detectors import (
|
||||||
detect_mass_withdrawal_anomalies,
|
detect_mass_withdrawal_anomalies,
|
||||||
detect_more_specific_burst_anomalies,
|
detect_more_specific_burst_anomalies,
|
||||||
@@ -23,32 +29,17 @@ from app.services.bgp_detectors import (
|
|||||||
)
|
)
|
||||||
from app.services.bgp_enrichment import enrich_bgp_events_for_batch, extract_bgp_network_fields
|
from app.services.bgp_enrichment import enrich_bgp_events_for_batch, extract_bgp_network_fields
|
||||||
|
|
||||||
|
# Re-exported for backward compatibility with anything that imports
|
||||||
RIPE_RIS_COLLECTOR_COORDS: dict[str, dict[str, Any]] = {
|
# ``RIPE_RIS_COLLECTOR_COORDS`` from this module. New code should call
|
||||||
"rrc00": {"city": "Amsterdam", "country": "Netherlands", "latitude": 52.3676, "longitude": 4.9041},
|
# ``app.services.bgp_collector_locations.get_bgp_collector_location_dict()``
|
||||||
"rrc01": {"city": "London", "country": "United Kingdom", "latitude": 51.5072, "longitude": -0.1276},
|
# or ``resolve_bgp_collector_location()`` instead — those use the DB-backed
|
||||||
"rrc03": {"city": "Amsterdam", "country": "Netherlands", "latitude": 52.3676, "longitude": 4.9041},
|
# collector-location cache.
|
||||||
"rrc04": {"city": "Geneva", "country": "Switzerland", "latitude": 46.2044, "longitude": 6.1432},
|
__all__ = [
|
||||||
"rrc05": {"city": "Vienna", "country": "Austria", "latitude": 48.2082, "longitude": 16.3738},
|
"RIPE_RIS_COLLECTOR_COORDS",
|
||||||
"rrc06": {"city": "Otemachi", "country": "Japan", "latitude": 35.686, "longitude": 139.7671},
|
"normalize_bgp_event",
|
||||||
"rrc07": {"city": "Stockholm", "country": "Sweden", "latitude": 59.3293, "longitude": 18.0686},
|
"save_bgp_observations_for_batch",
|
||||||
"rrc10": {"city": "Milan", "country": "Italy", "latitude": 45.4642, "longitude": 9.19},
|
"create_bgp_anomalies_for_batch",
|
||||||
"rrc11": {"city": "New York", "country": "United States", "latitude": 40.7128, "longitude": -74.006},
|
]
|
||||||
"rrc12": {"city": "Frankfurt", "country": "Germany", "latitude": 50.1109, "longitude": 8.6821},
|
|
||||||
"rrc13": {"city": "Moscow", "country": "Russia", "latitude": 55.7558, "longitude": 37.6173},
|
|
||||||
"rrc14": {"city": "Palo Alto", "country": "United States", "latitude": 37.4419, "longitude": -122.143},
|
|
||||||
"rrc15": {"city": "Sao Paulo", "country": "Brazil", "latitude": -23.5558, "longitude": -46.6396},
|
|
||||||
"rrc16": {"city": "Miami", "country": "United States", "latitude": 25.7617, "longitude": -80.1918},
|
|
||||||
"rrc18": {"city": "Barcelona", "country": "Spain", "latitude": 41.3874, "longitude": 2.1686},
|
|
||||||
"rrc19": {"city": "Johannesburg", "country": "South Africa", "latitude": -26.2041, "longitude": 28.0473},
|
|
||||||
"rrc20": {"city": "Zurich", "country": "Switzerland", "latitude": 47.3769, "longitude": 8.5417},
|
|
||||||
"rrc21": {"city": "Paris", "country": "France", "latitude": 48.8566, "longitude": 2.3522},
|
|
||||||
"rrc22": {"city": "Bucharest", "country": "Romania", "latitude": 44.4268, "longitude": 26.1025},
|
|
||||||
"rrc23": {"city": "Singapore", "country": "Singapore", "latitude": 1.3521, "longitude": 103.8198},
|
|
||||||
"rrc24": {"city": "Montevideo", "country": "Uruguay", "latitude": -34.9011, "longitude": -56.1645},
|
|
||||||
"rrc25": {"city": "Amsterdam", "country": "Netherlands", "latitude": 52.3676, "longitude": 4.9041},
|
|
||||||
"rrc26": {"city": "Dubai", "country": "United Arab Emirates", "latitude": 25.2048, "longitude": 55.2708},
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
def _safe_int(value: Any) -> int | None:
|
def _safe_int(value: Any) -> int | None:
|
||||||
@@ -131,7 +122,19 @@ def normalize_bgp_event(payload: dict[str, Any], *, project: str) -> dict[str, A
|
|||||||
)
|
)
|
||||||
source_id = hashlib.sha1(source_material.encode("utf-8")).hexdigest()[:24]
|
source_id = hashlib.sha1(source_material.encode("utf-8")).hexdigest()[:24]
|
||||||
|
|
||||||
collector_location = RIPE_RIS_COLLECTOR_COORDS.get(collector, {})
|
# Routes through the BGP event pipeline: source coords (if any) →
|
||||||
|
# collector inheritance. Returned dict keeps the legacy
|
||||||
|
# {city, country, latitude, longitude} keys plus richer
|
||||||
|
# {precision, source, needs_confirmation, matched_location_name, confidence}.
|
||||||
|
collector_location = resolve_bgp_event_geo_dict(
|
||||||
|
collector,
|
||||||
|
source_latitude=payload.get("latitude"),
|
||||||
|
source_longitude=payload.get("longitude"),
|
||||||
|
)
|
||||||
|
# Empty result (unknown collector & no source coords) — keep the
|
||||||
|
# downstream-expected dict shape so detectors / serializers don't crash.
|
||||||
|
if not collector_location:
|
||||||
|
collector_location = get_bgp_collector_location_dict(collector)
|
||||||
network_fields = extract_bgp_network_fields(prefix)
|
network_fields = extract_bgp_network_fields(prefix)
|
||||||
metadata = {
|
metadata = {
|
||||||
"project": project,
|
"project": project,
|
||||||
@@ -221,6 +224,8 @@ async def save_bgp_observations_for_batch(
|
|||||||
|
|
||||||
if created:
|
if created:
|
||||||
await db.commit()
|
await db.commit()
|
||||||
|
for source in {"ris_live_bgp", "bgpstream_bgp"}:
|
||||||
|
invalidate_earth_layer_cache_for_source(source)
|
||||||
|
|
||||||
return created
|
return created
|
||||||
|
|
||||||
|
|||||||
@@ -1,15 +1,39 @@
|
|||||||
"""CelesTrak TLE Collector
|
"""CelesTrak TLE Collector.
|
||||||
|
|
||||||
Collects satellite TLE (Two-Line Element) data from CelesTrak.org.
|
Collects the full active satellite GP element set from CelesTrak.
|
||||||
Free, no authentication required.
|
|
||||||
"""
|
"""
|
||||||
|
|
||||||
|
import asyncio
|
||||||
import json
|
import json
|
||||||
from typing import Dict, Any, List
|
from pathlib import Path
|
||||||
|
from time import perf_counter
|
||||||
|
from typing import Any, Dict, List
|
||||||
|
from urllib.parse import urlencode, urlparse
|
||||||
|
|
||||||
import httpx
|
import httpx
|
||||||
|
|
||||||
|
from app.core.logging import get_logger
|
||||||
from app.core.satellite_tle import build_tle_lines_from_elements
|
from app.core.satellite_tle import build_tle_lines_from_elements
|
||||||
|
from app.services.business_logs import emit_business_log, exception_context
|
||||||
from app.services.collectors.base import BaseCollector
|
from app.services.collectors.base import BaseCollector
|
||||||
|
from app.services.collectors.downloads import DownloadHTTPStatusError, ResumableFileDownloader
|
||||||
|
|
||||||
|
|
||||||
|
logger = get_logger(__name__, service="collector")
|
||||||
|
ACTIVE_GROUP = "active"
|
||||||
|
FALLBACK_GROUPS = (
|
||||||
|
"starlink",
|
||||||
|
"gps-ops",
|
||||||
|
"galileo",
|
||||||
|
"glonass",
|
||||||
|
"beidou",
|
||||||
|
"leo",
|
||||||
|
"geo",
|
||||||
|
"iridium-next",
|
||||||
|
)
|
||||||
|
FETCH_RETRY_ATTEMPTS = 3
|
||||||
|
FETCH_RETRY_BASE_DELAY_SECONDS = 0.8
|
||||||
|
CELESTRAK_NOT_UPDATED_MARKER = "GP data has not updated since your last successful"
|
||||||
|
|
||||||
|
|
||||||
class CelesTrakTLECollector(BaseCollector):
|
class CelesTrakTLECollector(BaseCollector):
|
||||||
@@ -18,55 +42,359 @@ class CelesTrakTLECollector(BaseCollector):
|
|||||||
module = "L3"
|
module = "L3"
|
||||||
frequency_hours = 24
|
frequency_hours = 24
|
||||||
data_type = "satellite_tle"
|
data_type = "satellite_tle"
|
||||||
|
_downloader = ResumableFileDownloader(
|
||||||
|
cache_namespace="celestrak",
|
||||||
|
default_accept="application/json",
|
||||||
|
)
|
||||||
|
|
||||||
@property
|
@property
|
||||||
def base_url(self) -> str:
|
def base_url(self) -> str:
|
||||||
return self._resolved_url or ""
|
return self._resolved_url or ""
|
||||||
|
|
||||||
|
def _active_url(self) -> str:
|
||||||
|
return self._group_url(ACTIVE_GROUP)
|
||||||
|
|
||||||
|
def _group_url(self, group: str) -> str:
|
||||||
|
if not self.base_url:
|
||||||
|
raise RuntimeError("CelesTrak base URL is not configured")
|
||||||
|
return f"{self.base_url}?{urlencode({'GROUP': group, 'FORMAT': 'json'})}"
|
||||||
|
|
||||||
async def fetch(self) -> List[Dict[str, Any]]:
|
async def fetch(self) -> List[Dict[str, Any]]:
|
||||||
satellite_groups = [
|
url = self._active_url()
|
||||||
"starlink",
|
last_error: Exception | None = None
|
||||||
"gps-ops",
|
|
||||||
"galileo",
|
|
||||||
"glonass",
|
|
||||||
"beidou",
|
|
||||||
"leo",
|
|
||||||
"geo",
|
|
||||||
"iridium-next",
|
|
||||||
]
|
|
||||||
|
|
||||||
all_satellites = []
|
async with httpx.AsyncClient(timeout=180.0, follow_redirects=True) as client:
|
||||||
|
for attempt in range(1, FETCH_RETRY_ATTEMPTS + 1):
|
||||||
async with httpx.AsyncClient(timeout=120.0) as client:
|
started_at = perf_counter()
|
||||||
for group in satellite_groups:
|
|
||||||
try:
|
try:
|
||||||
url = f"{self.base_url}?GROUP={group}&FORMAT=json"
|
await emit_business_log(
|
||||||
response = await client.get(url)
|
logger,
|
||||||
|
event="collector.celestrak.download.start",
|
||||||
|
message="CelesTrak active satellite download started",
|
||||||
|
category="collector",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"group": ACTIVE_GROUP,
|
||||||
|
"attempt": attempt,
|
||||||
|
"url_host": urlparse(url).netloc,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
body_path = await self._downloader.download_file(
|
||||||
|
client,
|
||||||
|
url,
|
||||||
|
extension=".json",
|
||||||
|
accept="application/json",
|
||||||
|
progress_callback=self._report_download_progress,
|
||||||
|
validate_existing=self._validate_json_file,
|
||||||
|
)
|
||||||
|
data = await self._load_downloaded_payload(body_path, url)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.celestrak.download.success",
|
||||||
|
message="CelesTrak active satellite download completed",
|
||||||
|
category="collector",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"group": ACTIVE_GROUP,
|
||||||
|
"attempt": attempt,
|
||||||
|
"record_count": len(data),
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return data
|
||||||
|
except DownloadHTTPStatusError as exc:
|
||||||
|
if self._is_not_updated_response(exc):
|
||||||
|
cached_path = self._downloader.get_cached_file(
|
||||||
|
url,
|
||||||
|
".json",
|
||||||
|
validate_existing=self._validate_json_file,
|
||||||
|
)
|
||||||
|
if cached_path is not None:
|
||||||
|
data = await self._load_downloaded_payload(cached_path, url)
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.celestrak.download.cached_not_updated",
|
||||||
|
message="CelesTrak active satellite data has not changed; using cached download",
|
||||||
|
category="collector",
|
||||||
|
level="warning",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"group": ACTIVE_GROUP,
|
||||||
|
"attempt": attempt,
|
||||||
|
"record_count": len(data),
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return data
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.celestrak.download.not_updated_no_cache",
|
||||||
|
message="CelesTrak active satellite data has not changed; trying fallback groups",
|
||||||
|
category="collector",
|
||||||
|
level="warning",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context=exception_context(
|
||||||
|
exc,
|
||||||
|
{
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"group": ACTIVE_GROUP,
|
||||||
|
"attempt": attempt,
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
),
|
||||||
|
)
|
||||||
|
return await self._fetch_fallback_groups(client, active_error=exc)
|
||||||
|
raise
|
||||||
|
except Exception as exc:
|
||||||
|
last_error = exc
|
||||||
|
is_final_attempt = attempt >= FETCH_RETRY_ATTEMPTS
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event=(
|
||||||
|
"collector.celestrak.download.failed"
|
||||||
|
if is_final_attempt
|
||||||
|
else "collector.celestrak.download.retry"
|
||||||
|
),
|
||||||
|
message=(
|
||||||
|
"CelesTrak active satellite download failed"
|
||||||
|
if is_final_attempt
|
||||||
|
else "CelesTrak active satellite download will retry"
|
||||||
|
),
|
||||||
|
category="collector",
|
||||||
|
level="error" if is_final_attempt else "warning",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context=exception_context(
|
||||||
|
exc,
|
||||||
|
{
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"group": ACTIVE_GROUP,
|
||||||
|
"attempt": attempt,
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
),
|
||||||
|
)
|
||||||
|
if not is_final_attempt:
|
||||||
|
await asyncio.sleep(FETCH_RETRY_BASE_DELAY_SECONDS * attempt)
|
||||||
|
|
||||||
if response.status_code == 200:
|
raise RuntimeError(f"CelesTrak active satellite download failed after retries: {last_error}")
|
||||||
data = response.json()
|
|
||||||
if isinstance(data, list):
|
|
||||||
for item in data:
|
|
||||||
if isinstance(item, dict):
|
|
||||||
item["_celestrak_group"] = group
|
|
||||||
all_satellites.extend(data)
|
|
||||||
print(f"CelesTrak: Fetched {len(data)} satellites from group '{group}'")
|
|
||||||
except Exception as e:
|
|
||||||
print(f"CelesTrak: Error fetching group '{group}': {e}")
|
|
||||||
|
|
||||||
if not all_satellites:
|
async def _fetch_fallback_groups(
|
||||||
return self._get_sample_data()
|
self,
|
||||||
|
client: httpx.AsyncClient,
|
||||||
|
*,
|
||||||
|
active_error: DownloadHTTPStatusError,
|
||||||
|
) -> List[Dict[str, Any]]:
|
||||||
|
started_at = perf_counter()
|
||||||
|
records_by_norad: dict[str, Dict[str, Any]] = {}
|
||||||
|
group_counts: dict[str, int] = {}
|
||||||
|
|
||||||
print(f"CelesTrak: Total satellites fetched: {len(all_satellites)}")
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.celestrak.fallback_groups.start",
|
||||||
|
message="CelesTrak fallback group download started",
|
||||||
|
category="collector",
|
||||||
|
level="warning",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"groups": list(FALLBACK_GROUPS),
|
||||||
|
"reason": "active_not_updated_without_cache",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
# Return raw data - base.run() will call transform()
|
try:
|
||||||
return all_satellites
|
for group in FALLBACK_GROUPS:
|
||||||
|
group_url = self._group_url(group)
|
||||||
|
try:
|
||||||
|
body_path = await self._downloader.download_file(
|
||||||
|
client,
|
||||||
|
group_url,
|
||||||
|
extension=".json",
|
||||||
|
accept="application/json",
|
||||||
|
validate_existing=self._validate_json_file,
|
||||||
|
)
|
||||||
|
except DownloadHTTPStatusError as exc:
|
||||||
|
if not self._is_not_updated_response(exc):
|
||||||
|
raise RuntimeError(f"CelesTrak fallback group '{group}' download failed: {exc}") from exc
|
||||||
|
cached_path = self._downloader.get_cached_file(
|
||||||
|
group_url,
|
||||||
|
".json",
|
||||||
|
validate_existing=self._validate_json_file,
|
||||||
|
)
|
||||||
|
if cached_path is None:
|
||||||
|
raise RuntimeError(
|
||||||
|
f"CelesTrak fallback group '{group}' has not updated and no local cached copy is available"
|
||||||
|
) from exc
|
||||||
|
body_path = cached_path
|
||||||
|
|
||||||
|
group_records = await self._load_downloaded_payload(
|
||||||
|
body_path,
|
||||||
|
group_url,
|
||||||
|
query_group=group,
|
||||||
|
constellation_group=group,
|
||||||
|
)
|
||||||
|
group_counts[group] = len(group_records)
|
||||||
|
for item in group_records:
|
||||||
|
norad_cat_id = item.get("NORAD_CAT_ID")
|
||||||
|
if norad_cat_id is None:
|
||||||
|
continue
|
||||||
|
records_by_norad.setdefault(str(norad_cat_id), item)
|
||||||
|
except Exception as exc:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.celestrak.fallback_groups.failed",
|
||||||
|
message="CelesTrak fallback group download failed",
|
||||||
|
category="collector",
|
||||||
|
level="error",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context=exception_context(
|
||||||
|
exc,
|
||||||
|
{
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"groups": list(FALLBACK_GROUPS),
|
||||||
|
"completed_groups": list(group_counts),
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
),
|
||||||
|
)
|
||||||
|
raise RuntimeError(
|
||||||
|
"CelesTrak active data has not updated since this network's last successful download, "
|
||||||
|
"no active cache is available, and fallback group mode failed. Wait until CelesTrak "
|
||||||
|
"publishes the next GP update, restore the Planet download cache, or use Space-Track."
|
||||||
|
) from active_error
|
||||||
|
|
||||||
|
records = list(records_by_norad.values())
|
||||||
|
if not records:
|
||||||
|
raise RuntimeError("CelesTrak fallback group mode produced no satellite records")
|
||||||
|
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.celestrak.fallback_groups.success",
|
||||||
|
message="CelesTrak fallback group download completed",
|
||||||
|
category="collector",
|
||||||
|
level="warning",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context={
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"groups": list(FALLBACK_GROUPS),
|
||||||
|
"group_counts": group_counts,
|
||||||
|
"record_count": len(records),
|
||||||
|
"duration_ms": self._duration_ms(started_at),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
return records
|
||||||
|
|
||||||
|
async def _load_downloaded_payload(
|
||||||
|
self,
|
||||||
|
body_path: Path,
|
||||||
|
url: str,
|
||||||
|
*,
|
||||||
|
query_group: str = ACTIVE_GROUP,
|
||||||
|
constellation_group: str | None = None,
|
||||||
|
) -> List[Dict[str, Any]]:
|
||||||
|
try:
|
||||||
|
data = self._load_active_payload(body_path)
|
||||||
|
except RuntimeError as exc:
|
||||||
|
await self._log_parse_failure(exc)
|
||||||
|
raise
|
||||||
|
for item in data:
|
||||||
|
item["_celestrak_query_group"] = query_group
|
||||||
|
item["_celestrak_source_url"] = url
|
||||||
|
if constellation_group:
|
||||||
|
item["_celestrak_group"] = constellation_group
|
||||||
|
return data
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _is_not_updated_response(exc: DownloadHTTPStatusError) -> bool:
|
||||||
|
return exc.status_code == 403 and CELESTRAK_NOT_UPDATED_MARKER in exc.body
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _duration_ms(started_at: float) -> int:
|
||||||
|
return int((perf_counter() - started_at) * 1000)
|
||||||
|
|
||||||
|
async def _report_download_progress(self, downloaded: int, total: int | None) -> None:
|
||||||
|
if total and total > 0:
|
||||||
|
await self.update_phase_progress(
|
||||||
|
current=min(downloaded, total),
|
||||||
|
total=total,
|
||||||
|
unit="bytes",
|
||||||
|
message=f"正在下载 CelesTrak active 卫星数据 {downloaded}/{total} bytes",
|
||||||
|
commit=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _validate_json_file(path: Path) -> bool:
|
||||||
|
try:
|
||||||
|
data = json.loads(path.read_text(encoding="utf-8"))
|
||||||
|
except (OSError, json.JSONDecodeError, UnicodeDecodeError):
|
||||||
|
return False
|
||||||
|
return isinstance(data, list)
|
||||||
|
|
||||||
|
async def _log_parse_failure(self, exc: Exception) -> None:
|
||||||
|
await emit_business_log(
|
||||||
|
logger,
|
||||||
|
event="collector.celestrak.parse.failed",
|
||||||
|
message="CelesTrak active satellite JSON parsing failed",
|
||||||
|
category="collector",
|
||||||
|
level="error",
|
||||||
|
service="collector",
|
||||||
|
module=__name__,
|
||||||
|
context=exception_context(
|
||||||
|
exc,
|
||||||
|
{
|
||||||
|
"collector_name": self.name,
|
||||||
|
"datasource_id": getattr(self, "_datasource_id", None),
|
||||||
|
"group": ACTIVE_GROUP,
|
||||||
|
},
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
def _load_active_payload(self, path: Path) -> List[Dict[str, Any]]:
|
||||||
|
try:
|
||||||
|
raw = json.loads(path.read_text(encoding="utf-8"))
|
||||||
|
except (OSError, json.JSONDecodeError, UnicodeDecodeError) as exc:
|
||||||
|
raise RuntimeError(f"CelesTrak active payload is not valid JSON: {exc}") from exc
|
||||||
|
if not isinstance(raw, list):
|
||||||
|
raise RuntimeError("CelesTrak active payload is not a JSON array")
|
||||||
|
|
||||||
|
records: List[Dict[str, Any]] = []
|
||||||
|
invalid_count = 0
|
||||||
|
for item in raw:
|
||||||
|
if isinstance(item, dict) and item.get("NORAD_CAT_ID") is not None:
|
||||||
|
records.append(item)
|
||||||
|
else:
|
||||||
|
invalid_count += 1
|
||||||
|
if invalid_count:
|
||||||
|
raise RuntimeError(f"CelesTrak active payload contains {invalid_count} invalid record(s)")
|
||||||
|
if not records:
|
||||||
|
raise RuntimeError("CelesTrak active payload contains no satellite records")
|
||||||
|
return records
|
||||||
|
|
||||||
def transform(self, raw_data: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
|
def transform(self, raw_data: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
|
||||||
transformed = []
|
transformed = []
|
||||||
for item in raw_data:
|
for item in raw_data:
|
||||||
|
norad_cat_id = item.get("NORAD_CAT_ID")
|
||||||
tle_line1, tle_line2 = build_tle_lines_from_elements(
|
tle_line1, tle_line2 = build_tle_lines_from_elements(
|
||||||
norad_cat_id=item.get("NORAD_CAT_ID"),
|
norad_cat_id=norad_cat_id,
|
||||||
epoch=item.get("EPOCH"),
|
epoch=item.get("EPOCH"),
|
||||||
inclination=item.get("INCLINATION"),
|
inclination=item.get("INCLINATION"),
|
||||||
raan=item.get("RA_OF_ASC_NODE"),
|
raan=item.get("RA_OF_ASC_NODE"),
|
||||||
@@ -75,14 +403,18 @@ class CelesTrakTLECollector(BaseCollector):
|
|||||||
mean_anomaly=item.get("MEAN_ANOMALY"),
|
mean_anomaly=item.get("MEAN_ANOMALY"),
|
||||||
mean_motion=item.get("MEAN_MOTION"),
|
mean_motion=item.get("MEAN_MOTION"),
|
||||||
)
|
)
|
||||||
|
constellation_group = self._infer_constellation_group(item)
|
||||||
|
|
||||||
transformed.append(
|
transformed.append(
|
||||||
{
|
{
|
||||||
|
"source_id": str(norad_cat_id),
|
||||||
"name": item.get("OBJECT_NAME", "Unknown"),
|
"name": item.get("OBJECT_NAME", "Unknown"),
|
||||||
"reference_date": item.get("EPOCH", ""),
|
"reference_date": item.get("EPOCH", ""),
|
||||||
"metadata": {
|
"metadata": {
|
||||||
"constellation_group": item.get("_celestrak_group"),
|
"constellation_group": constellation_group,
|
||||||
"norad_cat_id": item.get("NORAD_CAT_ID"),
|
"celestrak_query_group": item.get("_celestrak_query_group") or ACTIVE_GROUP,
|
||||||
|
"celestrak_source_url": item.get("_celestrak_source_url"),
|
||||||
|
"norad_cat_id": norad_cat_id,
|
||||||
"international_designator": item.get("OBJECT_ID"),
|
"international_designator": item.get("OBJECT_ID"),
|
||||||
"epoch": item.get("EPOCH"),
|
"epoch": item.get("EPOCH"),
|
||||||
"mean_motion": item.get("MEAN_MOTION"),
|
"mean_motion": item.get("MEAN_MOTION"),
|
||||||
@@ -105,6 +437,19 @@ class CelesTrakTLECollector(BaseCollector):
|
|||||||
)
|
)
|
||||||
return transformed
|
return transformed
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def _infer_constellation_group(item: Dict[str, Any]) -> str | None:
|
||||||
|
explicit_group = str(item.get("_celestrak_group") or "").strip().lower()
|
||||||
|
if explicit_group and explicit_group != ACTIVE_GROUP:
|
||||||
|
return explicit_group
|
||||||
|
|
||||||
|
name = str(item.get("OBJECT_NAME") or "").strip().upper()
|
||||||
|
if name.startswith("STARLINK"):
|
||||||
|
return "starlink"
|
||||||
|
if name.startswith("IRIDIUM"):
|
||||||
|
return "iridium-next"
|
||||||
|
return None
|
||||||
|
|
||||||
def _get_sample_data(self) -> List[Dict[str, Any]]:
|
def _get_sample_data(self) -> List[Dict[str, Any]]:
|
||||||
return [
|
return [
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -4,7 +4,7 @@ from __future__ import annotations
|
|||||||
|
|
||||||
import hashlib
|
import hashlib
|
||||||
import json
|
import json
|
||||||
import tempfile
|
import os
|
||||||
import time
|
import time
|
||||||
from datetime import UTC, datetime
|
from datetime import UTC, datetime
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
@@ -17,6 +17,31 @@ ProgressCallback = Callable[[int, int | None], Awaitable[None]]
|
|||||||
ValidateCallback = Callable[[Path], bool]
|
ValidateCallback = Callable[[Path], bool]
|
||||||
|
|
||||||
|
|
||||||
|
class DownloadHTTPStatusError(RuntimeError):
|
||||||
|
"""HTTP status error that keeps the upstream response body for caller-specific handling."""
|
||||||
|
|
||||||
|
def __init__(self, *, url: str, status_code: int, body: str) -> None:
|
||||||
|
self.url = url
|
||||||
|
self.status_code = status_code
|
||||||
|
self.body = body
|
||||||
|
preview = body.strip().replace("\r", " ").replace("\n", " ")[:240]
|
||||||
|
suffix = f": {preview}" if preview else ""
|
||||||
|
super().__init__(f"HTTP {status_code} while downloading {url}{suffix}")
|
||||||
|
|
||||||
|
|
||||||
|
def default_download_cache_root() -> Path:
|
||||||
|
configured = os.getenv("PLANET_DOWNLOAD_CACHE_DIR")
|
||||||
|
if configured:
|
||||||
|
return Path(configured).expanduser()
|
||||||
|
planet_cache = os.getenv("PLANET_CACHE_DIR")
|
||||||
|
if planet_cache:
|
||||||
|
return Path(planet_cache).expanduser() / "downloads"
|
||||||
|
xdg_cache = os.getenv("XDG_CACHE_HOME")
|
||||||
|
if xdg_cache:
|
||||||
|
return Path(xdg_cache).expanduser() / "planet" / "downloads"
|
||||||
|
return Path.home() / ".cache" / "planet" / "downloads"
|
||||||
|
|
||||||
|
|
||||||
class ResumableFileDownloader:
|
class ResumableFileDownloader:
|
||||||
"""Download files with cache validators and byte-range resume support."""
|
"""Download files with cache validators and byte-range resume support."""
|
||||||
|
|
||||||
@@ -26,8 +51,9 @@ class ResumableFileDownloader:
|
|||||||
cache_namespace: str,
|
cache_namespace: str,
|
||||||
user_agent: str = "Planet-Intelligence-System/1.0 (Python/collector)",
|
user_agent: str = "Planet-Intelligence-System/1.0 (Python/collector)",
|
||||||
default_accept: str = "*/*",
|
default_accept: str = "*/*",
|
||||||
|
cache_root: Path | None = None,
|
||||||
) -> None:
|
) -> None:
|
||||||
self._cache_dir = Path(tempfile.gettempdir()) / "planet-download-cache" / cache_namespace
|
self._cache_dir = (cache_root or default_download_cache_root()) / cache_namespace
|
||||||
self._user_agent = user_agent
|
self._user_agent = user_agent
|
||||||
self._default_accept = default_accept
|
self._default_accept = default_accept
|
||||||
|
|
||||||
@@ -43,6 +69,25 @@ class ResumableFileDownloader:
|
|||||||
meta_path = self._cache_dir / f"{key}.meta.json"
|
meta_path = self._cache_dir / f"{key}.meta.json"
|
||||||
return final_path, part_path, meta_path
|
return final_path, part_path, meta_path
|
||||||
|
|
||||||
|
def cached_file_path(self, url: str, extension: str) -> Path:
|
||||||
|
final_path, _, _ = self._cache_paths(url, extension)
|
||||||
|
return final_path
|
||||||
|
|
||||||
|
def get_cached_file(
|
||||||
|
self,
|
||||||
|
url: str,
|
||||||
|
extension: str,
|
||||||
|
*,
|
||||||
|
validate_existing: ValidateCallback | None = None,
|
||||||
|
) -> Path | None:
|
||||||
|
final_path = self.cached_file_path(url, extension)
|
||||||
|
if not final_path.exists():
|
||||||
|
return None
|
||||||
|
if validate_existing and not validate_existing(final_path):
|
||||||
|
final_path.unlink(missing_ok=True)
|
||||||
|
return None
|
||||||
|
return final_path
|
||||||
|
|
||||||
@staticmethod
|
@staticmethod
|
||||||
def _load_meta(meta_path: Path) -> dict[str, Any]:
|
def _load_meta(meta_path: Path) -> dict[str, Any]:
|
||||||
if not meta_path.exists():
|
if not meta_path.exists():
|
||||||
@@ -140,7 +185,9 @@ class ResumableFileDownloader:
|
|||||||
if progress_callback and expected_size and expected_size > 0:
|
if progress_callback and expected_size and expected_size > 0:
|
||||||
await progress_callback(expected_size, expected_size)
|
await progress_callback(expected_size, expected_size)
|
||||||
return final_path
|
return final_path
|
||||||
response.raise_for_status()
|
if response.status_code >= 400:
|
||||||
|
body = (await response.aread()).decode("utf-8", errors="replace")
|
||||||
|
raise DownloadHTTPStatusError(url=url, status_code=response.status_code, body=body)
|
||||||
|
|
||||||
if response.status_code == 206 and resume_from > 0:
|
if response.status_code == 206 and resume_from > 0:
|
||||||
mode = "ab"
|
mode = "ab"
|
||||||
|
|||||||
@@ -108,6 +108,11 @@ class IPtoASNPrefixGeoCollector(BaseCollector):
|
|||||||
self._current_task.total_records = total_expected
|
self._current_task.total_records = total_expected
|
||||||
self._current_task.records_processed = 0
|
self._current_task.records_processed = 0
|
||||||
self._current_task.progress = 0.0
|
self._current_task.progress = 0.0
|
||||||
|
self._current_task.phase_progress = 0.0
|
||||||
|
self._current_task.phase_message = "正在下载 IPtoASN 数据"
|
||||||
|
self._current_task.phase_current = 0
|
||||||
|
self._current_task.phase_total = total_expected
|
||||||
|
self._current_task.phase_unit = "bytes"
|
||||||
await self._db_session.commit()
|
await self._db_session.commit()
|
||||||
await self._publish_task_update(force=True)
|
await self._publish_task_update(force=True)
|
||||||
|
|
||||||
@@ -135,7 +140,14 @@ class IPtoASNPrefixGeoCollector(BaseCollector):
|
|||||||
return
|
return
|
||||||
last_emit["value"] = aggregated
|
last_emit["value"] = aggregated
|
||||||
last_emit["t"] = now
|
last_emit["t"] = now
|
||||||
await self.update_progress(min(aggregated, total_expected), commit=True)
|
current = min(aggregated, total_expected)
|
||||||
|
await self.update_phase_progress(
|
||||||
|
current=current,
|
||||||
|
total=total_expected,
|
||||||
|
unit="bytes",
|
||||||
|
message="正在下载 IPtoASN 数据",
|
||||||
|
)
|
||||||
|
await self.update_progress(current, commit=True)
|
||||||
|
|
||||||
batches = await asyncio.gather(
|
batches = await asyncio.gather(
|
||||||
*(
|
*(
|
||||||
@@ -148,6 +160,12 @@ class IPtoASNPrefixGeoCollector(BaseCollector):
|
|||||||
)
|
)
|
||||||
)
|
)
|
||||||
if total_expected > 0:
|
if total_expected > 0:
|
||||||
|
await self.update_phase_progress(
|
||||||
|
current=total_expected,
|
||||||
|
total=total_expected,
|
||||||
|
unit="bytes",
|
||||||
|
message="IPtoASN 数据下载完成",
|
||||||
|
)
|
||||||
await self.update_progress(total_expected, commit=True, force=True)
|
await self.update_progress(total_expected, commit=True, force=True)
|
||||||
|
|
||||||
rows: list[dict[str, Any]] = []
|
rows: list[dict[str, Any]] = []
|
||||||
|
|||||||
57
backend/app/services/collectors/media_news_archive.py
Normal file
57
backend/app/services/collectors/media_news_archive.py
Normal file
@@ -0,0 +1,57 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from app.services.collectors.base import BaseCollector
|
||||||
|
from app.services.earth_news_store import list_all_earth_news_records
|
||||||
|
|
||||||
|
|
||||||
|
class MediaNewsArchiveCollector(BaseCollector):
|
||||||
|
name = "media_news_archive"
|
||||||
|
priority = "P2"
|
||||||
|
module = "L4"
|
||||||
|
frequency_hours = 12
|
||||||
|
data_type = "news_item"
|
||||||
|
fail_on_empty = False
|
||||||
|
|
||||||
|
async def fetch(self) -> list[dict[str, Any]]:
|
||||||
|
if not self._db_session:
|
||||||
|
return []
|
||||||
|
|
||||||
|
records = await list_all_earth_news_records(self._db_session)
|
||||||
|
items: list[dict[str, Any]] = []
|
||||||
|
for record in records:
|
||||||
|
location_meta = dict(record.location_meta or {})
|
||||||
|
target = location_meta.get("target") if isinstance(location_meta.get("target"), dict) else {}
|
||||||
|
country = target.get("country")
|
||||||
|
city = target.get("city")
|
||||||
|
items.append(
|
||||||
|
{
|
||||||
|
"id": record.id,
|
||||||
|
"source_id": record.id,
|
||||||
|
"name": record.title,
|
||||||
|
"title": record.title,
|
||||||
|
"description": record.summary,
|
||||||
|
"country": country,
|
||||||
|
"city": city,
|
||||||
|
"latitude": record.latitude,
|
||||||
|
"longitude": record.longitude,
|
||||||
|
"reference_date": record.published_at,
|
||||||
|
"metadata": {
|
||||||
|
"url": record.url,
|
||||||
|
"source": record.source,
|
||||||
|
"feed_name": record.feed_name,
|
||||||
|
"region": record.region,
|
||||||
|
"homepage_url": record.homepage_url,
|
||||||
|
"published_at": record.published_at.isoformat() if record.published_at else None,
|
||||||
|
"location_label": record.location_label,
|
||||||
|
"location_source": record.location_source,
|
||||||
|
"verified": record.verified,
|
||||||
|
"location_meta": location_meta,
|
||||||
|
"first_seen_at": record.first_seen_at.isoformat() if record.first_seen_at else None,
|
||||||
|
"last_seen_at": record.last_seen_at.isoformat() if record.last_seen_at else None,
|
||||||
|
"resolved_at": record.resolved_at.isoformat() if record.resolved_at else None,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return items
|
||||||
@@ -39,12 +39,23 @@ class NRODelegatedPrefixGeoCollector(BaseCollector):
|
|||||||
self._current_task.total_records = total_expected
|
self._current_task.total_records = total_expected
|
||||||
self._current_task.records_processed = 0
|
self._current_task.records_processed = 0
|
||||||
self._current_task.progress = 0.0
|
self._current_task.progress = 0.0
|
||||||
|
self._current_task.phase_progress = 0.0
|
||||||
|
self._current_task.phase_message = "正在下载 NRO delegated 数据"
|
||||||
|
self._current_task.phase_current = 0
|
||||||
|
self._current_task.phase_total = total_expected
|
||||||
|
self._current_task.phase_unit = "bytes"
|
||||||
await self._db_session.commit()
|
await self._db_session.commit()
|
||||||
await self._publish_task_update(force=True)
|
await self._publish_task_update(force=True)
|
||||||
|
|
||||||
async def on_progress(downloaded: int, total: int | None) -> None:
|
async def on_progress(downloaded: int, total: int | None) -> None:
|
||||||
if not total or total <= 0:
|
if not total or total <= 0:
|
||||||
return
|
return
|
||||||
|
await self.update_phase_progress(
|
||||||
|
current=min(downloaded, total),
|
||||||
|
total=total,
|
||||||
|
unit="bytes",
|
||||||
|
message="正在下载 NRO delegated 数据",
|
||||||
|
)
|
||||||
await self.update_progress(min(downloaded, total), commit=True)
|
await self.update_progress(min(downloaded, total), commit=True)
|
||||||
|
|
||||||
body_path = await self._downloader.download_file(
|
body_path = await self._downloader.download_file(
|
||||||
|
|||||||
@@ -40,12 +40,23 @@ class OpenGeoFeedPrefixGeoCollector(BaseCollector):
|
|||||||
self._current_task.total_records = total_expected
|
self._current_task.total_records = total_expected
|
||||||
self._current_task.records_processed = 0
|
self._current_task.records_processed = 0
|
||||||
self._current_task.progress = 0.0
|
self._current_task.progress = 0.0
|
||||||
|
self._current_task.phase_progress = 0.0
|
||||||
|
self._current_task.phase_message = "正在下载 OpenGeoFeed 数据"
|
||||||
|
self._current_task.phase_current = 0
|
||||||
|
self._current_task.phase_total = total_expected
|
||||||
|
self._current_task.phase_unit = "bytes"
|
||||||
await self._db_session.commit()
|
await self._db_session.commit()
|
||||||
await self._publish_task_update(force=True)
|
await self._publish_task_update(force=True)
|
||||||
|
|
||||||
async def on_progress(downloaded: int, total: int | None) -> None:
|
async def on_progress(downloaded: int, total: int | None) -> None:
|
||||||
if not total or total <= 0:
|
if not total or total <= 0:
|
||||||
return
|
return
|
||||||
|
await self.update_phase_progress(
|
||||||
|
current=min(downloaded, total),
|
||||||
|
total=total,
|
||||||
|
unit="bytes",
|
||||||
|
message="正在下载 OpenGeoFeed 数据",
|
||||||
|
)
|
||||||
await self.update_progress(min(downloaded, total), commit=True)
|
await self.update_progress(min(downloaded, total), commit=True)
|
||||||
|
|
||||||
body_path = await self._downloader.download_file(
|
body_path = await self._downloader.download_file(
|
||||||
|
|||||||
@@ -163,32 +163,64 @@ class TeleGeographyLandingPointCollector(BaseCollector):
|
|||||||
data_type = "landing_point"
|
data_type = "landing_point"
|
||||||
|
|
||||||
async def fetch(self) -> List[Dict[str, Any]]:
|
async def fetch(self) -> List[Dict[str, Any]]:
|
||||||
"""Fetch landing point data from GitHub mirror"""
|
"""Fetch landing point data, falling back when the old mirror disappears."""
|
||||||
url = self._resolved_url or ""
|
config = get_data_sources_config()
|
||||||
|
sources = [
|
||||||
|
self._resolved_url or "",
|
||||||
|
str(config.get_yaml_value("telegeography.landing_point_url") or ""),
|
||||||
|
str(config.get_yaml_value("arcgis.landing_point_url") or ""),
|
||||||
|
]
|
||||||
|
|
||||||
async with httpx.AsyncClient(timeout=60.0) as client:
|
last_error: Exception | None = None
|
||||||
response = await client.get(url)
|
async with httpx.AsyncClient(timeout=60.0, follow_redirects=True) as client:
|
||||||
response.raise_for_status()
|
for url in dict.fromkeys(source for source in sources if source):
|
||||||
return self.parse_response(response.json())
|
try:
|
||||||
|
params = (
|
||||||
|
{"where": "1=1", "outFields": "*", "returnGeometry": "true", "f": "geojson"}
|
||||||
|
if "FeatureServer" in url or url.endswith("/query")
|
||||||
|
else None
|
||||||
|
)
|
||||||
|
response = await client.get(url, params=params)
|
||||||
|
response.raise_for_status()
|
||||||
|
records = self.parse_response(response.json())
|
||||||
|
if records:
|
||||||
|
return records
|
||||||
|
except Exception as exc:
|
||||||
|
last_error = exc
|
||||||
|
continue
|
||||||
|
|
||||||
def parse_response(self, data: List[Dict[str, Any]]) -> List[Dict[str, Any]]:
|
if last_error:
|
||||||
|
raise last_error
|
||||||
|
return self._get_sample_data()
|
||||||
|
|
||||||
|
def parse_response(self, data: List[Dict[str, Any]] | Dict[str, Any]) -> List[Dict[str, Any]]:
|
||||||
"""Parse landing point data"""
|
"""Parse landing point data"""
|
||||||
result = []
|
result = []
|
||||||
|
items = data.get("features", []) if isinstance(data, dict) else data
|
||||||
|
|
||||||
for item in data:
|
for item in items:
|
||||||
|
props = item.get("properties", {}) if isinstance(item, dict) else {}
|
||||||
|
geometry = item.get("geometry", {}) if isinstance(item, dict) else {}
|
||||||
|
source = props or item
|
||||||
|
coords = geometry.get("coordinates", []) if isinstance(geometry, dict) else []
|
||||||
|
longitude = coords[0] if len(coords) > 0 else source.get("longitude")
|
||||||
|
latitude = coords[1] if len(coords) > 1 else source.get("latitude")
|
||||||
|
source_id = source.get("id") or source.get("OBJECTID") or source.get("city_id") or ""
|
||||||
try:
|
try:
|
||||||
entry = {
|
entry = {
|
||||||
"source_id": f"telegeo_lp_{item.get('id', '')}",
|
"source_id": f"telegeo_lp_{source_id}",
|
||||||
"name": item.get("name", "Unknown"),
|
"name": source.get("name", source.get("Name", "Unknown")),
|
||||||
"country": item.get("country", "Unknown"),
|
"country": source.get("country", "Unknown"),
|
||||||
"city": item.get("city", item.get("name", "")),
|
"city": source.get("city", source.get("Name", source.get("name", ""))),
|
||||||
"latitude": str(item.get("latitude", "")),
|
"latitude": str(latitude or ""),
|
||||||
"longitude": str(item.get("longitude", "")),
|
"longitude": str(longitude or ""),
|
||||||
"value": "",
|
"value": "",
|
||||||
"unit": "",
|
"unit": "",
|
||||||
"metadata": {
|
"metadata": {
|
||||||
"cable_count": len(item.get("cables", [])),
|
"cable_count": len(source.get("cables", [])),
|
||||||
"url": item.get("url"),
|
"url": source.get("url"),
|
||||||
|
"objectid": source.get("OBJECTID"),
|
||||||
|
"city_id": source.get("city_id"),
|
||||||
},
|
},
|
||||||
"reference_date": datetime.now(UTC).strftime("%Y-%m-%d"),
|
"reference_date": datetime.now(UTC).strftime("%Y-%m-%d"),
|
||||||
}
|
}
|
||||||
|
|||||||
279
backend/app/services/collectors/vessel_ais.py
Normal file
279
backend/app/services/collectors/vessel_ais.py
Normal file
@@ -0,0 +1,279 @@
|
|||||||
|
"""BarentsWatch AIS collector for vessel tracking."""
|
||||||
|
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.time import to_iso8601_utc
|
||||||
|
from app.core.websocket.broadcaster import broadcaster
|
||||||
|
from app.services.barentswatch import (
|
||||||
|
BARENTSWATCH_LATEST_URL,
|
||||||
|
fetch_barentswatch_access_token,
|
||||||
|
resolve_barentswatch_config,
|
||||||
|
)
|
||||||
|
from app.services.collectors.base import BaseCollector
|
||||||
|
from app.services.vessel_ais_aggregation import (
|
||||||
|
BARENTSWATCH_DELIVERY_MODE,
|
||||||
|
BARENTSWATCH_TRANSPORT,
|
||||||
|
record_vessel_ais_observation,
|
||||||
|
update_ais_source_health,
|
||||||
|
)
|
||||||
|
from app.services.vessel_types import normalize_vessel_type_name
|
||||||
|
|
||||||
|
|
||||||
|
class VesselAISCollector(BaseCollector):
|
||||||
|
"""Collect latest AIS positions and append them to vessel tables."""
|
||||||
|
|
||||||
|
name = "barentswatch_vessels"
|
||||||
|
priority = "P1"
|
||||||
|
module = "L4"
|
||||||
|
frequency_hours = 1
|
||||||
|
data_type = "vessel_ais"
|
||||||
|
|
||||||
|
@property
|
||||||
|
def base_url(self) -> str:
|
||||||
|
return self._resolved_url or BARENTSWATCH_LATEST_URL
|
||||||
|
|
||||||
|
async def _get_access_token(self, client: httpx.AsyncClient) -> str | None:
|
||||||
|
config = await resolve_barentswatch_config(self._db_session)
|
||||||
|
return await fetch_barentswatch_access_token(client, config)
|
||||||
|
|
||||||
|
async def fetch(self) -> list[dict[str, Any]]:
|
||||||
|
async with httpx.AsyncClient(timeout=60.0) as client:
|
||||||
|
headers: dict[str, str] = {}
|
||||||
|
token = await self._get_access_token(client)
|
||||||
|
if token:
|
||||||
|
headers["Authorization"] = f"Bearer {token}"
|
||||||
|
|
||||||
|
response = await client.get(self.base_url, headers=headers)
|
||||||
|
if response.status_code == 401 and not token:
|
||||||
|
return self._get_sample_data()
|
||||||
|
response.raise_for_status()
|
||||||
|
payload = response.json()
|
||||||
|
|
||||||
|
if isinstance(payload, list):
|
||||||
|
return [item for item in payload if isinstance(item, dict)]
|
||||||
|
if isinstance(payload, dict):
|
||||||
|
for key in ("features", "data", "items", "vessels"):
|
||||||
|
value = payload.get(key)
|
||||||
|
if isinstance(value, list):
|
||||||
|
if key == "features":
|
||||||
|
return [
|
||||||
|
{
|
||||||
|
**(item.get("properties") or {}),
|
||||||
|
"geometry": item.get("geometry"),
|
||||||
|
}
|
||||||
|
for item in value
|
||||||
|
if isinstance(item, dict)
|
||||||
|
]
|
||||||
|
return [item for item in value if isinstance(item, dict)]
|
||||||
|
return self._get_sample_data()
|
||||||
|
|
||||||
|
def transform(self, raw_data: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||||
|
transformed = []
|
||||||
|
for item in raw_data:
|
||||||
|
record = self._normalize_record(item)
|
||||||
|
if record:
|
||||||
|
transformed.append(record)
|
||||||
|
return transformed
|
||||||
|
|
||||||
|
async def _save_data(
|
||||||
|
self,
|
||||||
|
db: AsyncSession,
|
||||||
|
data: list[dict[str, Any]],
|
||||||
|
task_id: int | None = None,
|
||||||
|
snapshot_id: int | None = None,
|
||||||
|
) -> int:
|
||||||
|
now = datetime.now(UTC)
|
||||||
|
records_added = 0
|
||||||
|
|
||||||
|
for index, item in enumerate(data):
|
||||||
|
observed_at = item.get("received_at") or now
|
||||||
|
await record_vessel_ais_observation(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
normalized_payload=item,
|
||||||
|
raw_payload=item,
|
||||||
|
delivery_mode=BARENTSWATCH_DELIVERY_MODE,
|
||||||
|
transport=BARENTSWATCH_TRANSPORT,
|
||||||
|
observed_at=observed_at,
|
||||||
|
collected_at=now,
|
||||||
|
)
|
||||||
|
records_added += 1
|
||||||
|
|
||||||
|
if (index + 1) % 1000 == 0:
|
||||||
|
await self.update_progress(index + 1, commit=True)
|
||||||
|
|
||||||
|
latest_observed_at = max(
|
||||||
|
(item.get("received_at") for item in data if item.get("received_at")),
|
||||||
|
default=now,
|
||||||
|
)
|
||||||
|
await update_ais_source_health(
|
||||||
|
db,
|
||||||
|
source=self.name,
|
||||||
|
connection_state="connected",
|
||||||
|
observed_count=len(data),
|
||||||
|
last_seen_at=latest_observed_at,
|
||||||
|
last_success_at=now if data else None,
|
||||||
|
lag_seconds=max((now - latest_observed_at).total_seconds(), 0),
|
||||||
|
)
|
||||||
|
await db.commit()
|
||||||
|
await self._broadcast_vessel_snapshot(data)
|
||||||
|
await self.update_progress(records_added, force=True)
|
||||||
|
return records_added
|
||||||
|
|
||||||
|
async def _broadcast_vessel_snapshot(self, data: list[dict[str, Any]]) -> None:
|
||||||
|
"""Push REST collector updates through the same realtime vessel channel."""
|
||||||
|
if not data:
|
||||||
|
return
|
||||||
|
|
||||||
|
batch_size = 500
|
||||||
|
for offset in range(0, len(data), batch_size):
|
||||||
|
batch = data[offset : offset + batch_size]
|
||||||
|
await broadcaster.broadcast_custom(
|
||||||
|
"vessels",
|
||||||
|
{
|
||||||
|
"action": "upsert",
|
||||||
|
"source": self.name,
|
||||||
|
"created": True,
|
||||||
|
"vessels": [
|
||||||
|
{
|
||||||
|
"mmsi": item.get("mmsi"),
|
||||||
|
"mmsi_display": str(item.get("mmsi")) if item.get("mmsi") is not None else None,
|
||||||
|
"name": item.get("name"),
|
||||||
|
"callsign": item.get("callsign"),
|
||||||
|
"lat": item.get("lat"),
|
||||||
|
"lon": item.get("lon"),
|
||||||
|
"sog": item.get("sog"),
|
||||||
|
"cog": item.get("cog"),
|
||||||
|
"heading": item.get("heading"),
|
||||||
|
"nav_status": item.get("nav_status"),
|
||||||
|
"vessel_type": item.get("vessel_type"),
|
||||||
|
"vessel_type_name": item.get("vessel_type_name"),
|
||||||
|
"received_at": to_iso8601_utc(item.get("received_at")),
|
||||||
|
}
|
||||||
|
for item in batch
|
||||||
|
],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
def _normalize_record(self, item: dict[str, Any]) -> dict[str, Any] | None:
|
||||||
|
mmsi = _as_int(_pick(item, "mmsi", "MMSI", "Mmsi"))
|
||||||
|
lat = _as_float(_pick(item, "lat", "latitude", "Latitude"))
|
||||||
|
lon = _as_float(_pick(item, "lon", "lng", "longitude", "Longitude"))
|
||||||
|
|
||||||
|
geometry = item.get("geometry")
|
||||||
|
coordinates = geometry.get("coordinates") if isinstance(geometry, dict) else None
|
||||||
|
if (lat is None or lon is None) and isinstance(coordinates, list) and len(coordinates) >= 2:
|
||||||
|
lon = _as_float(coordinates[0])
|
||||||
|
lat = _as_float(coordinates[1])
|
||||||
|
|
||||||
|
if mmsi is None or lat is None or lon is None:
|
||||||
|
return None
|
||||||
|
if not (-90 <= lat <= 90 and -180 <= lon <= 180):
|
||||||
|
return None
|
||||||
|
|
||||||
|
vessel_type = _as_int(_pick(item, "vessel_type", "shipType", "ship_type", "ShipType"))
|
||||||
|
vessel_type_name = (
|
||||||
|
_pick(item, "vessel_type_name", "shipTypeName", "ship_type_name", "VesselTypeName")
|
||||||
|
or normalize_vessel_type_name(vessel_type)
|
||||||
|
)
|
||||||
|
received_at = _parse_datetime(_pick(item, "received_at", "timestamp", "time", "msgtime"))
|
||||||
|
|
||||||
|
return {
|
||||||
|
"mmsi": mmsi,
|
||||||
|
"name": _pick(item, "name", "shipName", "ship_name", "Name"),
|
||||||
|
"callsign": _pick(item, "callsign", "callSign", "CallSign"),
|
||||||
|
"vessel_type": vessel_type,
|
||||||
|
"vessel_type_name": vessel_type_name,
|
||||||
|
"flag": _pick(item, "flag", "country", "Flag"),
|
||||||
|
"length": _as_float(_pick(item, "length", "shipLength", "Length")),
|
||||||
|
"width": _as_float(_pick(item, "width", "shipWidth", "Width")),
|
||||||
|
"draught": _as_float(_pick(item, "draught", "draft", "Draught")),
|
||||||
|
"imo": _as_int(_pick(item, "imo", "IMO", "imoNumber")),
|
||||||
|
"lat": lat,
|
||||||
|
"lon": lon,
|
||||||
|
"sog": _as_float(_pick(item, "sog", "speedOverGround", "SOG")),
|
||||||
|
"cog": _as_float(_pick(item, "cog", "courseOverGround", "COG")),
|
||||||
|
"heading": _as_int(_pick(item, "heading", "trueHeading", "Heading")),
|
||||||
|
"nav_status": _as_int(_pick(item, "nav_status", "navStatus", "NavigationalStatus")),
|
||||||
|
"received_at": received_at,
|
||||||
|
}
|
||||||
|
|
||||||
|
def _get_sample_data(self) -> list[dict[str, Any]]:
|
||||||
|
return [
|
||||||
|
{
|
||||||
|
"mmsi": 257123000,
|
||||||
|
"name": "OSLO TRADER",
|
||||||
|
"lat": 59.91,
|
||||||
|
"lon": 10.73,
|
||||||
|
"sog": 12.4,
|
||||||
|
"cog": 214,
|
||||||
|
"heading": 215,
|
||||||
|
"nav_status": 0,
|
||||||
|
"vessel_type": 70,
|
||||||
|
"vessel_type_name": "Cargo",
|
||||||
|
"flag": "NO",
|
||||||
|
"length": 185,
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"mmsi": 257456000,
|
||||||
|
"name": "NORDIC FJORD",
|
||||||
|
"lat": 60.39,
|
||||||
|
"lon": 5.32,
|
||||||
|
"sog": 0.2,
|
||||||
|
"cog": 82,
|
||||||
|
"heading": 80,
|
||||||
|
"nav_status": 1,
|
||||||
|
"vessel_type": 60,
|
||||||
|
"vessel_type_name": "Passenger",
|
||||||
|
"flag": "NO",
|
||||||
|
"length": 126,
|
||||||
|
},
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def _pick(item: dict[str, Any], *keys: str) -> Any:
|
||||||
|
for key in keys:
|
||||||
|
if key in item and item[key] not in (None, ""):
|
||||||
|
return item[key]
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _as_float(value: Any) -> float | None:
|
||||||
|
try:
|
||||||
|
if value in (None, ""):
|
||||||
|
return None
|
||||||
|
return float(value)
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _as_int(value: Any) -> int | None:
|
||||||
|
try:
|
||||||
|
if value in (None, ""):
|
||||||
|
return None
|
||||||
|
return int(float(value))
|
||||||
|
except (TypeError, ValueError):
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_datetime(value: Any) -> datetime | None:
|
||||||
|
if isinstance(value, datetime):
|
||||||
|
return value if value.tzinfo else value.replace(tzinfo=UTC)
|
||||||
|
if not value:
|
||||||
|
return None
|
||||||
|
if isinstance(value, (int, float)):
|
||||||
|
timestamp = float(value)
|
||||||
|
if timestamp > 10_000_000_000:
|
||||||
|
timestamp /= 1000
|
||||||
|
return datetime.fromtimestamp(timestamp, UTC)
|
||||||
|
if isinstance(value, str):
|
||||||
|
try:
|
||||||
|
parsed = datetime.fromisoformat(value.replace("Z", "+00:00"))
|
||||||
|
return parsed if parsed.tzinfo else parsed.replace(tzinfo=UTC)
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
|
return None
|
||||||
886
backend/app/services/compute_center_locations.py
Normal file
886
backend/app/services/compute_center_locations.py
Normal file
@@ -0,0 +1,886 @@
|
|||||||
|
"""Compute-center location resolver, built on the shared location pipeline.
|
||||||
|
|
||||||
|
This module is a thin domain wrapper that wires up
|
||||||
|
:mod:`app.services.location` for compute centers:
|
||||||
|
|
||||||
|
SourceCoordinates
|
||||||
|
|
||||||
|
The online Nominatim step is intentionally reserved for the user-triggered
|
||||||
|
``collect-location`` flow. The regular GeoJSON endpoint runs during Earth
|
||||||
|
startup, so it must stay local and deterministic.
|
||||||
|
|
||||||
|
For the full design and the reason behind the abstraction (compute centers,
|
||||||
|
BGP collectors, BGP events, and future entities all share one pipeline),
|
||||||
|
see ``docs/technical/zh/location-pipeline-development.md``.
|
||||||
|
|
||||||
|
The ``ComputeCenterLocation`` dataclass and the public function signatures are
|
||||||
|
preserved verbatim so existing callers and tests do not need to change.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from functools import lru_cache
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
from sqlalchemy import select
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.collected_data_fields import get_record_field
|
||||||
|
from app.models.collected_data import CollectedData
|
||||||
|
from app.models.compute_center_location import ComputeCenterLocationRecord
|
||||||
|
|
||||||
|
from app.services.location import (
|
||||||
|
LocationCandidate,
|
||||||
|
LocationPipeline,
|
||||||
|
LocationQuery,
|
||||||
|
NominatimResolver,
|
||||||
|
ResolverOutput,
|
||||||
|
SourceCoordinatesResolver,
|
||||||
|
build_default_nominatim_geocoder,
|
||||||
|
coerce_str,
|
||||||
|
normalize_country_text,
|
||||||
|
normalize_text,
|
||||||
|
parse_float,
|
||||||
|
)
|
||||||
|
|
||||||
|
ROR_SEARCH_URL = "https://api.ror.org/v2/organizations"
|
||||||
|
DEFAULT_ROR_USER_AGENT = "planet-earth-location-resolver/1.0"
|
||||||
|
DEFAULT_ROR_TIMEOUT_SECONDS = 8.0
|
||||||
|
RENDERABLE_PRECISIONS: tuple[str, ...] = ("precise", "site", "city")
|
||||||
|
FORBIDDEN_PRECISIONS: tuple[str, ...] = (
|
||||||
|
"country",
|
||||||
|
"estimated_country",
|
||||||
|
"country_major_compute_city",
|
||||||
|
"region",
|
||||||
|
"unknown",
|
||||||
|
)
|
||||||
|
|
||||||
|
# ── Public dataclasses ──────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ComputeCenterLocation:
|
||||||
|
latitude: float | None
|
||||||
|
longitude: float | None
|
||||||
|
location_precision: str
|
||||||
|
geography_mode: str
|
||||||
|
is_estimated: bool
|
||||||
|
estimated_reason: str | None = None
|
||||||
|
location_confidence: float | None = None
|
||||||
|
location_source: str | None = None
|
||||||
|
location_source_note: str | None = None
|
||||||
|
location_verified_at: str | None = None
|
||||||
|
matched_location_name: str | None = None
|
||||||
|
needs_confirmation: bool = False
|
||||||
|
city: str | None = None
|
||||||
|
region: str | None = None
|
||||||
|
country: str | None = None
|
||||||
|
|
||||||
|
@property
|
||||||
|
def is_renderable(self) -> bool:
|
||||||
|
if self.latitude in (None, 0.0) or self.longitude in (None, 0.0):
|
||||||
|
return False
|
||||||
|
return self.location_precision in RENDERABLE_PRECISIONS
|
||||||
|
|
||||||
|
def to_geojson_properties(self) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"latitude": self.latitude,
|
||||||
|
"longitude": self.longitude,
|
||||||
|
"location_precision": self.location_precision,
|
||||||
|
"geography_mode": self.geography_mode,
|
||||||
|
"is_estimated": self.is_estimated,
|
||||||
|
"estimated_reason": self.estimated_reason,
|
||||||
|
"location_confidence": self.location_confidence,
|
||||||
|
"location_source": self.location_source,
|
||||||
|
"location_source_note": self.location_source_note,
|
||||||
|
"location_verified_at": self.location_verified_at,
|
||||||
|
"matched_location_name": self.matched_location_name,
|
||||||
|
"needs_confirmation": self.needs_confirmation,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ResolutionDiagnostic:
|
||||||
|
failure_reason: str
|
||||||
|
attempted_queries: tuple[str, ...] = ()
|
||||||
|
record_id: int | None = None
|
||||||
|
source: str | None = None
|
||||||
|
source_id: str | None = None
|
||||||
|
name: str | None = None
|
||||||
|
country: str | None = None
|
||||||
|
city: str | None = None
|
||||||
|
site: str | None = None
|
||||||
|
operator: str | None = None
|
||||||
|
|
||||||
|
def to_dict(self) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"failure_reason": self.failure_reason,
|
||||||
|
"attempted_queries": list(self.attempted_queries),
|
||||||
|
"record_id": self.record_id,
|
||||||
|
"source": self.source,
|
||||||
|
"source_id": self.source_id,
|
||||||
|
"name": self.name,
|
||||||
|
"country": self.country,
|
||||||
|
"city": self.city,
|
||||||
|
"site": self.site,
|
||||||
|
"operator": self.operator,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ResolutionResult:
|
||||||
|
location: ComputeCenterLocation | None
|
||||||
|
diagnostic: ResolutionDiagnostic | None
|
||||||
|
|
||||||
|
@property
|
||||||
|
def is_resolved(self) -> bool:
|
||||||
|
return bool(self.location and self.location.is_renderable)
|
||||||
|
|
||||||
|
|
||||||
|
# ── Geocoder (kept at module level so tests can monkeypatch + cache_clear) ──
|
||||||
|
|
||||||
|
_geocode_online = build_default_nominatim_geocoder()
|
||||||
|
|
||||||
|
|
||||||
|
# ── Stored location cache ───────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
COMPUTE_CENTER_LOCATION_CACHE: dict[str, dict[str, Any]] = {}
|
||||||
|
|
||||||
|
|
||||||
|
def _cache_key(source: str | None, source_id: str | None) -> str:
|
||||||
|
return f"{coerce_str(source)}:{coerce_str(source_id)}"
|
||||||
|
|
||||||
|
|
||||||
|
def set_compute_center_location_cache(
|
||||||
|
locations: dict[str, dict[str, Any]],
|
||||||
|
) -> None:
|
||||||
|
COMPUTE_CENTER_LOCATION_CACHE.clear()
|
||||||
|
COMPUTE_CENTER_LOCATION_CACHE.update(
|
||||||
|
{coerce_str(key): dict(value) for key, value in locations.items()}
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def refresh_compute_center_location_cache(
|
||||||
|
session: AsyncSession,
|
||||||
|
) -> dict[str, dict[str, Any]]:
|
||||||
|
result = await session.execute(select(ComputeCenterLocationRecord))
|
||||||
|
records = result.scalars().all()
|
||||||
|
cache = {}
|
||||||
|
for record in records:
|
||||||
|
if not hasattr(record, "to_location_dict"):
|
||||||
|
continue
|
||||||
|
if not record.source or not record.source_id:
|
||||||
|
continue
|
||||||
|
cache[_cache_key(record.source, record.source_id)] = record.to_location_dict()
|
||||||
|
set_compute_center_location_cache(cache)
|
||||||
|
return cache
|
||||||
|
|
||||||
|
|
||||||
|
def get_compute_center_location_dict(
|
||||||
|
source: str | None,
|
||||||
|
source_id: str | None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
return dict(COMPUTE_CENTER_LOCATION_CACHE.get(_cache_key(source, source_id), {}))
|
||||||
|
|
||||||
|
|
||||||
|
# ── Pipeline construction ──────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
@lru_cache(maxsize=512)
|
||||||
|
def _lookup_ror_organization(query: str) -> dict[str, Any] | None:
|
||||||
|
"""Lookup a research organization in ROR for user-triggered candidates."""
|
||||||
|
if not query:
|
||||||
|
return None
|
||||||
|
response = httpx.get(
|
||||||
|
ROR_SEARCH_URL,
|
||||||
|
params={"query": query},
|
||||||
|
headers={"User-Agent": DEFAULT_ROR_USER_AGENT},
|
||||||
|
timeout=DEFAULT_ROR_TIMEOUT_SECONDS,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
payload = response.json()
|
||||||
|
items = payload.get("items") if isinstance(payload, dict) else None
|
||||||
|
if not isinstance(items, list) or not items:
|
||||||
|
return None
|
||||||
|
first = items[0]
|
||||||
|
if not isinstance(first, dict):
|
||||||
|
return None
|
||||||
|
organization = first.get("organization")
|
||||||
|
if isinstance(organization, dict):
|
||||||
|
return organization
|
||||||
|
return first
|
||||||
|
|
||||||
|
|
||||||
|
def _compute_center_ror_query_plan(
|
||||||
|
query: LocationQuery,
|
||||||
|
) -> list[tuple[str, tuple[str, ...]]]:
|
||||||
|
extra = query.extra or {}
|
||||||
|
raw_parts: list[tuple[str, str]] = [
|
||||||
|
("site", coerce_str(extra.get("site"))),
|
||||||
|
("operator", coerce_str(extra.get("operator"))),
|
||||||
|
("organization", coerce_str(extra.get("organization"))),
|
||||||
|
]
|
||||||
|
for field, value in tuple(raw_parts):
|
||||||
|
if "/" not in value:
|
||||||
|
continue
|
||||||
|
raw_parts.extend(
|
||||||
|
(field, part.strip())
|
||||||
|
for part in value.split("/")
|
||||||
|
if len(part.strip()) >= 3
|
||||||
|
)
|
||||||
|
|
||||||
|
plan: list[tuple[str, tuple[str, ...]]] = []
|
||||||
|
seen: set[str] = set()
|
||||||
|
for field, value in raw_parts:
|
||||||
|
key = normalize_text(value)
|
||||||
|
if not key or key in seen:
|
||||||
|
continue
|
||||||
|
seen.add(key)
|
||||||
|
plan.append((value, (field,)))
|
||||||
|
return plan
|
||||||
|
|
||||||
|
|
||||||
|
def _organization_label(organization: dict[str, Any], fallback: str) -> str:
|
||||||
|
names = organization.get("names")
|
||||||
|
if isinstance(names, list):
|
||||||
|
for name in names:
|
||||||
|
if not isinstance(name, dict):
|
||||||
|
continue
|
||||||
|
types = name.get("types")
|
||||||
|
if isinstance(types, list) and "ror_display" in types:
|
||||||
|
value = coerce_str(name.get("value"))
|
||||||
|
if value:
|
||||||
|
return value
|
||||||
|
for name in names:
|
||||||
|
if isinstance(name, dict):
|
||||||
|
value = coerce_str(name.get("value"))
|
||||||
|
if value:
|
||||||
|
return value
|
||||||
|
return fallback
|
||||||
|
|
||||||
|
|
||||||
|
class ROROrganizationResolver:
|
||||||
|
"""Resolve source-provided organization/site text through the open ROR API."""
|
||||||
|
|
||||||
|
name = "ror_organization_registry"
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
*,
|
||||||
|
query_plan_builder=_compute_center_ror_query_plan,
|
||||||
|
lookup=lambda q: _lookup_ror_organization(q),
|
||||||
|
confidence: float = 0.68,
|
||||||
|
) -> None:
|
||||||
|
self._query_plan_builder = query_plan_builder
|
||||||
|
self._lookup = lookup
|
||||||
|
self._confidence = confidence
|
||||||
|
|
||||||
|
def resolve(self, query: LocationQuery):
|
||||||
|
from app.services.location import ResolverOutput
|
||||||
|
from app.services.location.text import parse_float
|
||||||
|
|
||||||
|
attempted: list[str] = []
|
||||||
|
candidates: list[LocationCandidate] = []
|
||||||
|
context_country = normalize_text(normalize_country_text(query.country))
|
||||||
|
|
||||||
|
for ror_query, matched_fields in self._query_plan_builder(query):
|
||||||
|
attempted.append(f"ror:{ror_query}")
|
||||||
|
try:
|
||||||
|
organization = self._lookup(ror_query)
|
||||||
|
except Exception:
|
||||||
|
continue
|
||||||
|
if not isinstance(organization, dict):
|
||||||
|
continue
|
||||||
|
locations = organization.get("locations")
|
||||||
|
if not isinstance(locations, list) or not locations:
|
||||||
|
continue
|
||||||
|
location = locations[0]
|
||||||
|
if not isinstance(location, dict):
|
||||||
|
continue
|
||||||
|
details = location.get("geonames_details")
|
||||||
|
if not isinstance(details, dict):
|
||||||
|
continue
|
||||||
|
latitude = parse_float(details.get("lat"))
|
||||||
|
longitude = parse_float(details.get("lng"))
|
||||||
|
if latitude in (None, 0.0) or longitude in (None, 0.0):
|
||||||
|
continue
|
||||||
|
|
||||||
|
country = normalize_country_text(details.get("country_name"))
|
||||||
|
if context_country and normalize_text(country) != context_country:
|
||||||
|
continue
|
||||||
|
|
||||||
|
city = coerce_str(details.get("name")) or None
|
||||||
|
region = coerce_str(details.get("country_subdivision_name")) or None
|
||||||
|
display_name = _organization_label(organization, ror_query)
|
||||||
|
ror_id = coerce_str(organization.get("id"))
|
||||||
|
geonames_id = location.get("geonames_id")
|
||||||
|
source_note = (
|
||||||
|
f"ROR organization match: {display_name}"
|
||||||
|
+ (f" ({ror_id})" if ror_id else "")
|
||||||
|
+ (f"; GeoNames {geonames_id}" if geonames_id else "")
|
||||||
|
)
|
||||||
|
candidates.append(
|
||||||
|
LocationCandidate(
|
||||||
|
latitude=latitude,
|
||||||
|
longitude=longitude,
|
||||||
|
display_name=display_name,
|
||||||
|
precision="city",
|
||||||
|
confidence=self._confidence,
|
||||||
|
query=ror_query,
|
||||||
|
source=self.name,
|
||||||
|
source_note=source_note,
|
||||||
|
matched_fields=matched_fields,
|
||||||
|
needs_confirmation=True,
|
||||||
|
city=city,
|
||||||
|
region=region,
|
||||||
|
country=country or query.country,
|
||||||
|
matched_location_name=display_name,
|
||||||
|
location_verified_at=None,
|
||||||
|
suggested_registry_entry=None,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
return ResolverOutput(
|
||||||
|
candidates=tuple(candidates),
|
||||||
|
attempted_queries=tuple(attempted),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class StoredComputeCenterLocationResolver:
|
||||||
|
"""Resolve a compute center through the DB-backed current-location cache."""
|
||||||
|
|
||||||
|
name = "stored_compute_center_location"
|
||||||
|
|
||||||
|
def resolve(self, query: LocationQuery) -> ResolverOutput:
|
||||||
|
extra = query.extra or {}
|
||||||
|
stored = get_compute_center_location_dict(
|
||||||
|
coerce_str(extra.get("source")),
|
||||||
|
coerce_str(extra.get("source_id")),
|
||||||
|
)
|
||||||
|
if not stored:
|
||||||
|
return ResolverOutput()
|
||||||
|
latitude = parse_float(stored.get("latitude"))
|
||||||
|
longitude = parse_float(stored.get("longitude"))
|
||||||
|
if latitude in (None, 0.0) or longitude in (None, 0.0):
|
||||||
|
return ResolverOutput()
|
||||||
|
return ResolverOutput(
|
||||||
|
candidates=(
|
||||||
|
LocationCandidate(
|
||||||
|
latitude=latitude,
|
||||||
|
longitude=longitude,
|
||||||
|
display_name=stored.get("name") or query.name or "Compute center",
|
||||||
|
precision=stored.get("precision") or "city",
|
||||||
|
confidence=float(stored.get("confidence") or 0.85),
|
||||||
|
query=f"stored_compute_center_location::{stored.get('source')}:{stored.get('source_id')}",
|
||||||
|
source=self.name,
|
||||||
|
source_note=stored.get("source_note"),
|
||||||
|
matched_fields=("source", "source_id"),
|
||||||
|
needs_confirmation=bool(stored.get("needs_confirmation")),
|
||||||
|
city=stored.get("city") or query.city,
|
||||||
|
region=None,
|
||||||
|
country=stored.get("country") or query.country,
|
||||||
|
matched_location_name=stored.get("site") or stored.get("name") or query.name,
|
||||||
|
location_verified_at=stored.get("verified_at"),
|
||||||
|
suggested_registry_entry=None,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _short_system_name(name: Any) -> str:
|
||||||
|
"""Strip vendor/system suffix from TOP500 names like ``"El Capitan - HPE Cray ..."``."""
|
||||||
|
text = coerce_str(name)
|
||||||
|
if not text:
|
||||||
|
return ""
|
||||||
|
head = text.split(" - ", 1)[0].strip()
|
||||||
|
return head or text
|
||||||
|
|
||||||
|
|
||||||
|
def _record_context(record: Any, metadata: dict[str, Any]) -> dict[str, str]:
|
||||||
|
name = coerce_str(getattr(record, "name", None))
|
||||||
|
return {
|
||||||
|
"source": coerce_str(getattr(record, "source", None)),
|
||||||
|
"source_id": coerce_str(getattr(record, "source_id", None)),
|
||||||
|
"name": name,
|
||||||
|
"name_short": _short_system_name(name),
|
||||||
|
"city": coerce_str(get_record_field(record, "city")),
|
||||||
|
"country": coerce_str(get_record_field(record, "country")),
|
||||||
|
"site": coerce_str(metadata.get("site") or metadata.get("organization")),
|
||||||
|
"operator": coerce_str(
|
||||||
|
metadata.get("operator")
|
||||||
|
or metadata.get("organization")
|
||||||
|
or metadata.get("owner")
|
||||||
|
or metadata.get("manufacturer")
|
||||||
|
),
|
||||||
|
"organization": coerce_str(metadata.get("organization")),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _context_to_query(
|
||||||
|
context: dict[str, str],
|
||||||
|
*,
|
||||||
|
source_lat: float | None = None,
|
||||||
|
source_lon: float | None = None,
|
||||||
|
) -> LocationQuery:
|
||||||
|
name = context.get("name") or None
|
||||||
|
name_short = context.get("name_short") or ""
|
||||||
|
aliases: tuple[str, ...] = ()
|
||||||
|
if name_short and name_short != name:
|
||||||
|
aliases = (name_short,)
|
||||||
|
return LocationQuery(
|
||||||
|
name=name,
|
||||||
|
aliases=aliases,
|
||||||
|
city=context.get("city") or None,
|
||||||
|
country=context.get("country") or None,
|
||||||
|
source_latitude=source_lat,
|
||||||
|
source_longitude=source_lon,
|
||||||
|
extra={
|
||||||
|
"source": context.get("source") or "",
|
||||||
|
"source_id": context.get("source_id") or "",
|
||||||
|
"site": context.get("site") or "",
|
||||||
|
"operator": context.get("operator") or "",
|
||||||
|
"organization": context.get("organization") or "",
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _compute_center_query_plan(
|
||||||
|
query: LocationQuery,
|
||||||
|
) -> list[tuple[str, tuple[str, ...]]]:
|
||||||
|
"""Build the Nominatim query plan for a compute-center query.
|
||||||
|
|
||||||
|
Mirrors the legacy ``_build_online_query_plan`` ordering exactly.
|
||||||
|
"""
|
||||||
|
name = query.name or ""
|
||||||
|
name_short = (query.aliases[0] if query.aliases else "") or name
|
||||||
|
extra = query.extra or {}
|
||||||
|
site = str(extra.get("site") or "")
|
||||||
|
operator = str(extra.get("operator") or "")
|
||||||
|
city = query.city or ""
|
||||||
|
country = query.country or ""
|
||||||
|
|
||||||
|
plan: list[tuple[str, tuple[str, ...]]] = []
|
||||||
|
|
||||||
|
def add(parts: list[tuple[str, str]]) -> None:
|
||||||
|
non_empty = [(field, value) for field, value in parts if value]
|
||||||
|
if not non_empty:
|
||||||
|
return
|
||||||
|
seen: set[str] = set()
|
||||||
|
cleaned: list[str] = []
|
||||||
|
fields: list[str] = []
|
||||||
|
for field, value in non_empty:
|
||||||
|
key = normalize_text(value)
|
||||||
|
if not key or key in seen:
|
||||||
|
continue
|
||||||
|
seen.add(key)
|
||||||
|
cleaned.append(value)
|
||||||
|
fields.append(field)
|
||||||
|
if not cleaned:
|
||||||
|
return
|
||||||
|
composed = ", ".join(cleaned)
|
||||||
|
if not any(composed == existing for existing, _ in plan):
|
||||||
|
plan.append((composed, tuple(fields)))
|
||||||
|
|
||||||
|
add([("site", site), ("country", country)])
|
||||||
|
add([("operator", operator), ("city", city), ("country", country)])
|
||||||
|
add([("name", name_short), ("operator", operator), ("country", country)])
|
||||||
|
add([("name", name_short), ("site", site)])
|
||||||
|
add([("name", name_short), ("country", country)])
|
||||||
|
add([("name", name_short), ("city", city), ("country", country)])
|
||||||
|
add([("city", city), ("country", country)])
|
||||||
|
if name and name != name_short:
|
||||||
|
add([("name", name), ("country", country)])
|
||||||
|
return plan
|
||||||
|
|
||||||
|
|
||||||
|
COMPUTE_CENTER_PIPELINE = LocationPipeline(
|
||||||
|
[
|
||||||
|
SourceCoordinatesResolver(),
|
||||||
|
StoredComputeCenterLocationResolver(),
|
||||||
|
],
|
||||||
|
failure_reason=(
|
||||||
|
"Could not resolve to city-level coordinates from source coords"
|
||||||
|
" or stored compute-center location."
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
COMPUTE_CENTER_COLLECTION_PIPELINE = LocationPipeline(
|
||||||
|
[
|
||||||
|
SourceCoordinatesResolver(),
|
||||||
|
ROROrganizationResolver(),
|
||||||
|
NominatimResolver(
|
||||||
|
query_plan_builder=_compute_center_query_plan,
|
||||||
|
# Late-binding so test monkeypatching of ``_geocode_online`` works.
|
||||||
|
geocoder=lambda q: _geocode_online(q),
|
||||||
|
),
|
||||||
|
],
|
||||||
|
failure_reason=(
|
||||||
|
"Could not resolve to city-level coordinates from source coords"
|
||||||
|
", ROR organization lookup, or online geocoding."
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ── Candidate → ComputeCenterLocation conversion ───────────────────
|
||||||
|
|
||||||
|
|
||||||
|
_GEOGRAPHY_MODE_BY_SOURCE = {
|
||||||
|
"source_coordinates": "source_coordinates",
|
||||||
|
"stored_compute_center_location": "stored_compute_center_location",
|
||||||
|
"ror_organization_registry": "ror_organization",
|
||||||
|
"nominatim_online_geocode": "online_geocode",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _candidate_to_location(
|
||||||
|
candidate: LocationCandidate,
|
||||||
|
*,
|
||||||
|
context: dict[str, str],
|
||||||
|
) -> ComputeCenterLocation:
|
||||||
|
geography_mode = _GEOGRAPHY_MODE_BY_SOURCE.get(candidate.source, "online_geocode")
|
||||||
|
is_estimated = candidate.needs_confirmation or candidate.source.startswith(
|
||||||
|
"nominatim"
|
||||||
|
)
|
||||||
|
estimated_reason: str | None
|
||||||
|
if candidate.source == "source_coordinates":
|
||||||
|
estimated_reason = None
|
||||||
|
elif candidate.source == "stored_compute_center_location":
|
||||||
|
estimated_reason = candidate.source_note
|
||||||
|
elif candidate.source == "ror_organization_registry":
|
||||||
|
fields_summary = ", ".join(candidate.matched_fields) or "organization"
|
||||||
|
estimated_reason = (
|
||||||
|
f"Resolved by ROR organization lookup '{candidate.query}' "
|
||||||
|
f"(matched fields: {fields_summary})"
|
||||||
|
)
|
||||||
|
elif candidate.source == "nominatim_online_geocode":
|
||||||
|
fields_summary = ", ".join(candidate.matched_fields) or "name"
|
||||||
|
estimated_reason = (
|
||||||
|
f"Resolved by online geocoding query '{candidate.query}' "
|
||||||
|
f"(matched fields: {fields_summary})"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
estimated_reason = candidate.source_note
|
||||||
|
|
||||||
|
country = (
|
||||||
|
candidate.country
|
||||||
|
or normalize_country_text(context.get("country"))
|
||||||
|
or context.get("country")
|
||||||
|
or None
|
||||||
|
)
|
||||||
|
return ComputeCenterLocation(
|
||||||
|
latitude=candidate.latitude,
|
||||||
|
longitude=candidate.longitude,
|
||||||
|
location_precision=candidate.precision,
|
||||||
|
geography_mode=geography_mode,
|
||||||
|
is_estimated=is_estimated,
|
||||||
|
estimated_reason=estimated_reason,
|
||||||
|
location_confidence=candidate.confidence,
|
||||||
|
location_source=candidate.source,
|
||||||
|
location_source_note=candidate.source_note,
|
||||||
|
location_verified_at=candidate.location_verified_at,
|
||||||
|
matched_location_name=candidate.matched_location_name
|
||||||
|
or context.get("name")
|
||||||
|
or None,
|
||||||
|
needs_confirmation=candidate.needs_confirmation,
|
||||||
|
city=candidate.city or context.get("city") or None,
|
||||||
|
region=candidate.region,
|
||||||
|
country=country,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _diagnostic_for(
|
||||||
|
record: Any,
|
||||||
|
context: dict[str, str],
|
||||||
|
*,
|
||||||
|
failure_reason: str,
|
||||||
|
attempted_queries: tuple[str, ...] = (),
|
||||||
|
) -> ResolutionDiagnostic:
|
||||||
|
return ResolutionDiagnostic(
|
||||||
|
failure_reason=failure_reason,
|
||||||
|
attempted_queries=attempted_queries,
|
||||||
|
record_id=getattr(record, "id", None),
|
||||||
|
source=getattr(record, "source", None),
|
||||||
|
source_id=getattr(record, "source_id", None),
|
||||||
|
name=context.get("name") or getattr(record, "name", None),
|
||||||
|
country=context.get("country") or None,
|
||||||
|
city=context.get("city") or None,
|
||||||
|
site=context.get("site") or None,
|
||||||
|
operator=context.get("operator") or None,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ── Public API ─────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_compute_center_location(
|
||||||
|
record: Any,
|
||||||
|
metadata: dict[str, Any] | None = None,
|
||||||
|
) -> ComputeCenterLocation:
|
||||||
|
"""Backwards-compatible thin wrapper returning the renderable location only.
|
||||||
|
|
||||||
|
Records that cannot be resolved to city-level get a placeholder
|
||||||
|
:class:`ComputeCenterLocation` with ``location_precision='unknown'``.
|
||||||
|
Callers should generally prefer :func:`resolve_compute_center_location_full`.
|
||||||
|
"""
|
||||||
|
full = resolve_compute_center_location_full(record, metadata)
|
||||||
|
return full.location or ComputeCenterLocation(
|
||||||
|
latitude=None,
|
||||||
|
longitude=None,
|
||||||
|
location_precision="unknown",
|
||||||
|
geography_mode="unresolved",
|
||||||
|
is_estimated=True,
|
||||||
|
estimated_reason="No resolvable location hints",
|
||||||
|
location_confidence=0.0,
|
||||||
|
location_source="unknown",
|
||||||
|
location_source_note=(
|
||||||
|
"No source coordinates, ROR organization match, or online"
|
||||||
|
" geocoding result."
|
||||||
|
),
|
||||||
|
matched_location_name=None,
|
||||||
|
needs_confirmation=False,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_compute_center_location_full(
|
||||||
|
record: Any,
|
||||||
|
metadata: dict[str, Any] | None = None,
|
||||||
|
*,
|
||||||
|
allow_online: bool = False,
|
||||||
|
) -> ResolutionResult:
|
||||||
|
metadata = metadata or {}
|
||||||
|
context = _record_context(record, metadata)
|
||||||
|
|
||||||
|
from app.services.location.text import parse_float as _parse_float
|
||||||
|
|
||||||
|
source_lat = _parse_float(get_record_field(record, "latitude"))
|
||||||
|
source_lon = _parse_float(get_record_field(record, "longitude"))
|
||||||
|
if source_lat in (None, 0.0):
|
||||||
|
source_lat = None
|
||||||
|
if source_lon in (None, 0.0):
|
||||||
|
source_lon = None
|
||||||
|
|
||||||
|
query = _context_to_query(
|
||||||
|
context, source_lat=source_lat, source_lon=source_lon
|
||||||
|
)
|
||||||
|
pipeline = (
|
||||||
|
COMPUTE_CENTER_COLLECTION_PIPELINE
|
||||||
|
if allow_online
|
||||||
|
else COMPUTE_CENTER_PIPELINE
|
||||||
|
)
|
||||||
|
pipeline_result = pipeline.resolve_best(query)
|
||||||
|
|
||||||
|
if pipeline_result.location and pipeline_result.location.precision in RENDERABLE_PRECISIONS:
|
||||||
|
location = _candidate_to_location(pipeline_result.location, context=context)
|
||||||
|
return ResolutionResult(location=location, diagnostic=None)
|
||||||
|
|
||||||
|
return ResolutionResult(
|
||||||
|
location=None,
|
||||||
|
diagnostic=_diagnostic_for(
|
||||||
|
record,
|
||||||
|
context,
|
||||||
|
failure_reason=(
|
||||||
|
"Could not resolve to city-level coordinates from source coords"
|
||||||
|
", ROR organization lookup, or online geocoding."
|
||||||
|
if allow_online
|
||||||
|
else (
|
||||||
|
"Could not resolve to city-level coordinates from source coords"
|
||||||
|
" or stored compute-center location."
|
||||||
|
)
|
||||||
|
),
|
||||||
|
attempted_queries=pipeline_result.attempted_queries,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def collect_location_candidates(
|
||||||
|
*,
|
||||||
|
name: str | None = None,
|
||||||
|
source: str | None = None,
|
||||||
|
source_id: str | None = None,
|
||||||
|
operator: str | None = None,
|
||||||
|
site: str | None = None,
|
||||||
|
city: str | None = None,
|
||||||
|
country: str | None = None,
|
||||||
|
organization: str | None = None,
|
||||||
|
record_id: int | None = None,
|
||||||
|
) -> tuple[list[LocationCandidate], list[str]]:
|
||||||
|
"""Run the full resolution chain and return ranked candidates with attempted queries.
|
||||||
|
|
||||||
|
The unused ``source`` / ``source_id`` / ``record_id`` arguments are kept
|
||||||
|
for backward compatibility with the API handler that calls this function.
|
||||||
|
"""
|
||||||
|
query = build_compute_center_location_query(
|
||||||
|
name=name,
|
||||||
|
source=source,
|
||||||
|
source_id=source_id,
|
||||||
|
operator=operator,
|
||||||
|
site=site,
|
||||||
|
city=city,
|
||||||
|
country=country,
|
||||||
|
organization=organization,
|
||||||
|
)
|
||||||
|
return COMPUTE_CENTER_COLLECTION_PIPELINE.collect_candidates(query)
|
||||||
|
|
||||||
|
|
||||||
|
def build_compute_center_location_query(
|
||||||
|
*,
|
||||||
|
name: str | None = None,
|
||||||
|
source: str | None = None,
|
||||||
|
source_id: str | None = None,
|
||||||
|
operator: str | None = None,
|
||||||
|
site: str | None = None,
|
||||||
|
city: str | None = None,
|
||||||
|
country: str | None = None,
|
||||||
|
organization: str | None = None,
|
||||||
|
) -> LocationQuery:
|
||||||
|
name_value = coerce_str(name)
|
||||||
|
context: dict[str, str] = {
|
||||||
|
"source": coerce_str(source),
|
||||||
|
"source_id": coerce_str(source_id),
|
||||||
|
"name": name_value,
|
||||||
|
"name_short": _short_system_name(name_value),
|
||||||
|
"city": coerce_str(city),
|
||||||
|
"country": coerce_str(country),
|
||||||
|
"site": coerce_str(site or organization),
|
||||||
|
"operator": coerce_str(operator or organization),
|
||||||
|
"organization": coerce_str(organization),
|
||||||
|
}
|
||||||
|
return _context_to_query(context)
|
||||||
|
|
||||||
|
|
||||||
|
def _record_operator(metadata: dict[str, Any]) -> str | None:
|
||||||
|
return coerce_str(
|
||||||
|
metadata.get("operator")
|
||||||
|
or metadata.get("organization")
|
||||||
|
or metadata.get("owner")
|
||||||
|
or metadata.get("manufacturer")
|
||||||
|
) or None
|
||||||
|
|
||||||
|
|
||||||
|
async def seed_compute_center_locations_from_source_coords(
|
||||||
|
session: AsyncSession,
|
||||||
|
) -> None:
|
||||||
|
"""Seed stored compute-center locations only from real source coordinates."""
|
||||||
|
stmt = (
|
||||||
|
select(CollectedData)
|
||||||
|
.where(CollectedData.source.in_(["top500", "epoch_ai_gpu"]))
|
||||||
|
.where(CollectedData.is_current.is_(True))
|
||||||
|
)
|
||||||
|
result = await session.execute(stmt)
|
||||||
|
records = result.scalars().all()
|
||||||
|
changed = False
|
||||||
|
|
||||||
|
for record in records:
|
||||||
|
source_value = coerce_str(getattr(record, "source", None))
|
||||||
|
source_id = coerce_str(getattr(record, "source_id", None))
|
||||||
|
if not source_value or not source_id:
|
||||||
|
continue
|
||||||
|
latitude = parse_float(get_record_field(record, "latitude"))
|
||||||
|
longitude = parse_float(get_record_field(record, "longitude"))
|
||||||
|
if latitude in (None, 0.0) or longitude in (None, 0.0):
|
||||||
|
continue
|
||||||
|
existing = await session.scalar(
|
||||||
|
select(ComputeCenterLocationRecord)
|
||||||
|
.where(ComputeCenterLocationRecord.source == source_value)
|
||||||
|
.where(ComputeCenterLocationRecord.source_id == source_id)
|
||||||
|
)
|
||||||
|
if existing:
|
||||||
|
continue
|
||||||
|
metadata = record.extra_data or {}
|
||||||
|
session.add(
|
||||||
|
ComputeCenterLocationRecord(
|
||||||
|
source=source_value,
|
||||||
|
source_id=source_id,
|
||||||
|
name=getattr(record, "name", None),
|
||||||
|
operator=_record_operator(metadata),
|
||||||
|
site=coerce_str(metadata.get("site") or metadata.get("organization")) or None,
|
||||||
|
city=coerce_str(get_record_field(record, "city")) or None,
|
||||||
|
country=coerce_str(get_record_field(record, "country")) or None,
|
||||||
|
latitude=latitude,
|
||||||
|
longitude=longitude,
|
||||||
|
precision="precise",
|
||||||
|
confidence=1.0,
|
||||||
|
location_source="source_coordinates",
|
||||||
|
source_note="Seeded from source-provided compute-center coordinates",
|
||||||
|
raw_payload={
|
||||||
|
"record_id": getattr(record, "id", None),
|
||||||
|
"source": source_value,
|
||||||
|
"source_id": source_id,
|
||||||
|
},
|
||||||
|
needs_confirmation=False,
|
||||||
|
verification_status="source_provided",
|
||||||
|
verified_at=None,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
changed = True
|
||||||
|
|
||||||
|
if changed:
|
||||||
|
await session.commit()
|
||||||
|
await refresh_compute_center_location_cache(session)
|
||||||
|
|
||||||
|
|
||||||
|
async def upsert_compute_center_location(
|
||||||
|
session: AsyncSession,
|
||||||
|
*,
|
||||||
|
source: str,
|
||||||
|
source_id: str,
|
||||||
|
name: str | None = None,
|
||||||
|
operator: str | None = None,
|
||||||
|
site: str | None = None,
|
||||||
|
city: str | None = None,
|
||||||
|
country: str | None = None,
|
||||||
|
latitude: float,
|
||||||
|
longitude: float,
|
||||||
|
precision: str = "city",
|
||||||
|
confidence: float | None = None,
|
||||||
|
location_source: str = "manual_selection",
|
||||||
|
source_url: str | None = None,
|
||||||
|
source_note: str | None = None,
|
||||||
|
raw_payload: dict[str, Any] | None = None,
|
||||||
|
needs_confirmation: bool = False,
|
||||||
|
verification_status: str = "verified",
|
||||||
|
) -> ComputeCenterLocationRecord:
|
||||||
|
existing = await session.scalar(
|
||||||
|
select(ComputeCenterLocationRecord)
|
||||||
|
.where(ComputeCenterLocationRecord.source == source)
|
||||||
|
.where(ComputeCenterLocationRecord.source_id == source_id)
|
||||||
|
)
|
||||||
|
verified_at = None if needs_confirmation else datetime.now(UTC)
|
||||||
|
values = {
|
||||||
|
"name": name,
|
||||||
|
"operator": operator,
|
||||||
|
"site": site,
|
||||||
|
"city": city,
|
||||||
|
"country": country,
|
||||||
|
"latitude": latitude,
|
||||||
|
"longitude": longitude,
|
||||||
|
"precision": precision,
|
||||||
|
"confidence": confidence,
|
||||||
|
"location_source": location_source,
|
||||||
|
"source_url": source_url,
|
||||||
|
"source_note": source_note,
|
||||||
|
"raw_payload": raw_payload or {},
|
||||||
|
"needs_confirmation": needs_confirmation,
|
||||||
|
"verification_status": verification_status,
|
||||||
|
"verified_at": verified_at,
|
||||||
|
}
|
||||||
|
if existing:
|
||||||
|
for key, value in values.items():
|
||||||
|
setattr(existing, key, value)
|
||||||
|
record = existing
|
||||||
|
else:
|
||||||
|
record = ComputeCenterLocationRecord(
|
||||||
|
source=source,
|
||||||
|
source_id=source_id,
|
||||||
|
**values,
|
||||||
|
)
|
||||||
|
session.add(record)
|
||||||
|
|
||||||
|
await session.commit()
|
||||||
|
await session.refresh(record)
|
||||||
|
await refresh_compute_center_location_cache(session)
|
||||||
|
return record
|
||||||
314
backend/app/services/credential_guides.py
Normal file
314
backend/app/services/credential_guides.py
Normal file
@@ -0,0 +1,314 @@
|
|||||||
|
"""Credential setup guides for collector integrations."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from copy import deepcopy
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from sqlalchemy import select
|
||||||
|
from sqlalchemy.orm.attributes import flag_modified
|
||||||
|
|
||||||
|
from app.ai_tasks.prompts import get_effective_prompt
|
||||||
|
from app.models.system_setting import SystemSetting
|
||||||
|
from app.schemas.ai import SituationalAnalysisRequest
|
||||||
|
from app.services.ai_client import AIProviderClient
|
||||||
|
from app.services.ai_tools.evidence_store import normalize_search_evidence
|
||||||
|
from app.services.ai_tools.web_search import WebSearchClient, WebSearchError
|
||||||
|
|
||||||
|
|
||||||
|
CREDENTIAL_GUIDES_CATEGORY = "collector_credential_guides"
|
||||||
|
CREDENTIAL_GUIDE_PROMPT_KEY = "credential.guide"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class CredentialGuideDefault:
|
||||||
|
provider: str
|
||||||
|
title: str
|
||||||
|
prompt: str
|
||||||
|
markdown: str
|
||||||
|
|
||||||
|
|
||||||
|
BARENTSWATCH_DEFAULT_GUIDE = CredentialGuideDefault(
|
||||||
|
provider="barentswatch",
|
||||||
|
title="BarentsWatch AIS 凭证获取教程",
|
||||||
|
prompt=(
|
||||||
|
"请生成一份中文教程,指导开发者获取 BarentsWatch Live AIS API 的 "
|
||||||
|
"OAuth client credentials。教程要面向已经有本地开发环境的人,包含注册/登录、"
|
||||||
|
"创建 client、申请或确认 ais scope、复制 client id 和 client secret、"
|
||||||
|
"在系统设置中填写并验证连接、常见失败排查。不要编造具体页面按钮文案,"
|
||||||
|
"必须参考官方 tutorial:https://developer.barentswatch.no/docs/tutorial 。"
|
||||||
|
"必须强调 Live AIS 要选择 AIS-client / AIS - API,而不是普通 API-client。"
|
||||||
|
"如果步骤可能变化,要提醒以 BarentsWatch developer portal 当前页面为准。"
|
||||||
|
),
|
||||||
|
markdown="""## BarentsWatch AIS 凭证获取
|
||||||
|
|
||||||
|
官方教程:https://developer.barentswatch.no/docs/tutorial
|
||||||
|
|
||||||
|
1. 先打开上面的 BarentsWatch 官方 tutorial,按官方流程登录或注册开发者账号。
|
||||||
|
2. 在 Developer access 页面选择 `AIS - API`,不要选择普通的 `BarentsWatch - API`。
|
||||||
|
3. 在 `AIS - API` 下创建用于 Planet 的 AIS client。
|
||||||
|
4. 创建时记下你设置的 password / client secret。
|
||||||
|
5. 回到 My Page 复制完整 `Client ID`。它通常长得像 `your.email@example.com:client-name`。
|
||||||
|
6. 回到 Planet 的 `设置 -> 采集器设置 -> BarentsWatch AIS`,填入 `Client ID` 和 `Client Secret`。
|
||||||
|
7. 点击 `连接` 验证 token 和 AIS endpoint 是否可访问。
|
||||||
|
8. 连接成功后保存凭证。
|
||||||
|
|
||||||
|
### 请求规则
|
||||||
|
|
||||||
|
- Token 地址:`https://id.barentswatch.no/connect/token`
|
||||||
|
- 请求方式:`POST`
|
||||||
|
- Content-Type:`application/x-www-form-urlencoded`
|
||||||
|
- Body 必须包含:`grant_type=client_credentials`、`client_id`、`client_secret`、`scope=ais`
|
||||||
|
- `client_id`、`client_secret`、`scope`、`grant_type` 都要放在 body,不要放在 header。
|
||||||
|
- AIS 数据请求使用 header:`Authorization: Bearer <access_token>`
|
||||||
|
|
||||||
|
### 常见排查
|
||||||
|
|
||||||
|
- `未找到凭证`:确认 `Client ID` 和 `Client Secret` 已填写,或已经写入 `~/.zshrc`。
|
||||||
|
- `HTTP 401/403`:通常是选成了普通 `BarentsWatch - API` client、client secret 错误,或 token 请求没有使用 `scope=ais`。
|
||||||
|
- `network` 错误:检查本机是否能访问 `id.barentswatch.no` 和 `live.ais.barentswatch.no`。
|
||||||
|
- Endpoint 建议保持默认:`https://live.ais.barentswatch.no/v1/latest/combined`。
|
||||||
|
""",
|
||||||
|
)
|
||||||
|
|
||||||
|
AISSTREAM_DEFAULT_GUIDE = CredentialGuideDefault(
|
||||||
|
provider="aisstream",
|
||||||
|
title="AISStream API Key 获取教程",
|
||||||
|
prompt=(
|
||||||
|
"请生成一份中文教程,指导开发者获取 AISStream 的 API Key 并配置到 Planet。"
|
||||||
|
"教程要面向已经有本地开发环境的人,包含注册/登录 AISStream、获取 API Key、"
|
||||||
|
"理解免费额度和订阅范围、在 Planet 设置中心填写 API Key、配置 bounding boxes "
|
||||||
|
"和 message types、验证连接、常见失败排查。必须提醒用户以 AISStream 当前官网和"
|
||||||
|
"服务条款为准,不要编造具体页面按钮文案。"
|
||||||
|
),
|
||||||
|
markdown="""## AISStream API Key 获取
|
||||||
|
|
||||||
|
官方入口:https://aisstream.io/
|
||||||
|
|
||||||
|
1. 打开 AISStream 官网,按当前页面指引注册或登录账号。
|
||||||
|
2. 在账号/API 管理页面创建或复制你的 API Key。
|
||||||
|
3. 先确认当前账号额度、使用条款和可订阅区域。实时 AIS 流量可能很大,不建议一开始订阅全球范围。
|
||||||
|
4. 回到 Planet 的 `设置 -> 采集器设置 -> AISStream 实时船舶`。
|
||||||
|
5. 在 `AISStream 凭证` 中填入 API Key。
|
||||||
|
6. Endpoint 通常保持默认:`wss://stream.aisstream.io/v0/stream`。
|
||||||
|
7. 按需配置 `Bounding Boxes JSON` 和 `消息类型`。
|
||||||
|
8. 点击连接测试,确认系统能读取凭证且 WebSocket endpoint 格式有效。
|
||||||
|
9. 保存采集器设置后再运行 `aisstream_vessels` collector。
|
||||||
|
|
||||||
|
### 推荐配置
|
||||||
|
|
||||||
|
默认消息类型:
|
||||||
|
|
||||||
|
```json
|
||||||
|
["PositionReport", "ShipStaticData"]
|
||||||
|
```
|
||||||
|
|
||||||
|
默认 Bounding Boxes 示例:
|
||||||
|
|
||||||
|
```json
|
||||||
|
[[[-90, -180], [90, 180]]]
|
||||||
|
```
|
||||||
|
|
||||||
|
这个示例表示全球范围。实际使用时建议先改成较小区域,降低消息量和处理压力。
|
||||||
|
|
||||||
|
### 请求规则
|
||||||
|
|
||||||
|
- Endpoint:`wss://stream.aisstream.io/v0/stream`
|
||||||
|
- 传输方式:WebSocket
|
||||||
|
- API Key 放在订阅 payload 中,不放在 HTTP header。
|
||||||
|
- Planet 会把 AISStream 标记为 `delivery_mode = realtime_stream`、`transport = websocket`。
|
||||||
|
- AISStream collector 只写入 AIS raw observations,不直接覆盖最终船只展示表。
|
||||||
|
|
||||||
|
### 常见排查
|
||||||
|
|
||||||
|
- `未找到凭证`:确认 API Key 已保存到采集器设置,或设置了 `AISSTREAM_API_KEY` 环境变量 / `~/.zshrc`。
|
||||||
|
- `endpoint 必须是 ws:// 或 wss://`:AISStream 是 WebSocket 流接口,不要填普通 `https://` API 地址。
|
||||||
|
- 采集量过大:缩小 `Bounding Boxes JSON`,减少 `message_types`,或降低单次最大消息数。
|
||||||
|
- 没有船只数据:确认订阅区域内确实有 AIS 活动,并检查 API Key 当前额度和权限。
|
||||||
|
- 连接中断:实时流可能受网络和上游限流影响,collector 会记录源健康状态供聚合服务回退。
|
||||||
|
""",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_CREDENTIAL_GUIDES = {
|
||||||
|
BARENTSWATCH_DEFAULT_GUIDE.provider: BARENTSWATCH_DEFAULT_GUIDE,
|
||||||
|
AISSTREAM_DEFAULT_GUIDE.provider: AISSTREAM_DEFAULT_GUIDE,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _normalize_provider(provider: str) -> str:
|
||||||
|
return provider.strip().lower().replace(" ", "_")
|
||||||
|
|
||||||
|
|
||||||
|
def _credential_guide_default(provider: str) -> CredentialGuideDefault:
|
||||||
|
normalized = _normalize_provider(provider)
|
||||||
|
known = DEFAULT_CREDENTIAL_GUIDES.get(normalized)
|
||||||
|
if known is not None:
|
||||||
|
return known
|
||||||
|
title = f"{normalized or 'collector'} 凭证配置教程"
|
||||||
|
return CredentialGuideDefault(
|
||||||
|
provider=normalized,
|
||||||
|
title=title,
|
||||||
|
prompt=(
|
||||||
|
f"请生成一份中文教程,指导开发者为 Planet 采集器配置 {normalized} 凭证。"
|
||||||
|
"教程要面向已经有本地开发环境的人,包含官方入口或文档查找方式、"
|
||||||
|
"获取 API Key / Token / Client credentials 的通用步骤、在 Planet 采集器配置中"
|
||||||
|
"填写凭证字段、连接测试、保存、常见失败排查。不要编造具体页面按钮文案;"
|
||||||
|
"如果公开资料不足,必须明确提醒以 provider 官方文档和当前控制台页面为准。"
|
||||||
|
),
|
||||||
|
markdown="",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def _get_guide_store(db) -> tuple[SystemSetting | None, dict[str, Any]]:
|
||||||
|
result = await db.execute(
|
||||||
|
select(SystemSetting).where(SystemSetting.category == CREDENTIAL_GUIDES_CATEGORY)
|
||||||
|
)
|
||||||
|
record = result.scalar_one_or_none()
|
||||||
|
payload = deepcopy(record.payload) if record and isinstance(record.payload, dict) else {}
|
||||||
|
return record, payload
|
||||||
|
|
||||||
|
|
||||||
|
async def get_credential_guide(db, provider: str) -> dict[str, Any]:
|
||||||
|
provider = _normalize_provider(provider)
|
||||||
|
default = _credential_guide_default(provider)
|
||||||
|
|
||||||
|
_record, store = await _get_guide_store(db)
|
||||||
|
custom = store.get(provider) if isinstance(store.get(provider), dict) else None
|
||||||
|
has_default_markdown = bool(default.markdown.strip())
|
||||||
|
return {
|
||||||
|
"provider": provider,
|
||||||
|
"title": custom.get("title") if custom else default.title,
|
||||||
|
"markdown": custom.get("markdown") if custom else default.markdown,
|
||||||
|
"prompt": default.prompt,
|
||||||
|
"source": "ai" if custom else "default" if has_default_markdown else "missing",
|
||||||
|
"sources": custom.get("sources", []) if custom else [],
|
||||||
|
"verification_status": (
|
||||||
|
custom.get("verification_status", "verified_with_search_evidence")
|
||||||
|
if custom
|
||||||
|
else "default_unverified" if has_default_markdown else "missing"
|
||||||
|
),
|
||||||
|
"verification_error": custom.get("verification_error") if custom else None,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def save_credential_guide(
|
||||||
|
db,
|
||||||
|
provider: str,
|
||||||
|
title: str,
|
||||||
|
markdown: str,
|
||||||
|
*,
|
||||||
|
sources: list[dict[str, Any]] | None = None,
|
||||||
|
verification_status: str = "verified_with_search_evidence",
|
||||||
|
verification_error: str | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
provider = _normalize_provider(provider)
|
||||||
|
default = _credential_guide_default(provider)
|
||||||
|
|
||||||
|
record, store = await _get_guide_store(db)
|
||||||
|
store[provider] = {
|
||||||
|
"title": title or default.title,
|
||||||
|
"markdown": markdown,
|
||||||
|
"sources": sources or [],
|
||||||
|
"verification_status": verification_status,
|
||||||
|
"verification_error": verification_error,
|
||||||
|
}
|
||||||
|
if record is None:
|
||||||
|
db.add(SystemSetting(category=CREDENTIAL_GUIDES_CATEGORY, payload=store))
|
||||||
|
else:
|
||||||
|
record.payload = deepcopy(store)
|
||||||
|
flag_modified(record, "payload")
|
||||||
|
await db.commit()
|
||||||
|
return await get_credential_guide(db, provider)
|
||||||
|
|
||||||
|
|
||||||
|
async def reset_credential_guide(db, provider: str) -> dict[str, Any]:
|
||||||
|
provider = _normalize_provider(provider)
|
||||||
|
|
||||||
|
record, store = await _get_guide_store(db)
|
||||||
|
if provider in store:
|
||||||
|
store.pop(provider, None)
|
||||||
|
if record is not None:
|
||||||
|
record.payload = deepcopy(store)
|
||||||
|
flag_modified(record, "payload")
|
||||||
|
await db.commit()
|
||||||
|
return await get_credential_guide(db, provider)
|
||||||
|
|
||||||
|
|
||||||
|
async def generate_credential_guide(
|
||||||
|
db,
|
||||||
|
provider: str,
|
||||||
|
ai_client: AIProviderClient,
|
||||||
|
web_search_client: WebSearchClient | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
provider = _normalize_provider(provider)
|
||||||
|
default = _credential_guide_default(provider)
|
||||||
|
|
||||||
|
search_evidence: list[dict[str, Any]] = []
|
||||||
|
search_error: str | None = None
|
||||||
|
if web_search_client is not None:
|
||||||
|
try:
|
||||||
|
evidence = await web_search_client.search(
|
||||||
|
_credential_guide_search_query(default),
|
||||||
|
max_results=5,
|
||||||
|
)
|
||||||
|
search_evidence = normalize_search_evidence(evidence, limit=5)
|
||||||
|
except WebSearchError as exc:
|
||||||
|
search_error = str(exc)
|
||||||
|
except Exception as exc:
|
||||||
|
search_error = f"WebSearch unavailable: {exc}"
|
||||||
|
|
||||||
|
if not search_evidence:
|
||||||
|
guide = await get_credential_guide(db, provider)
|
||||||
|
guide["verification_status"] = "unverified_no_search_evidence"
|
||||||
|
guide["verification_error"] = search_error
|
||||||
|
guide["sources"] = []
|
||||||
|
return guide
|
||||||
|
|
||||||
|
prompt = await get_effective_prompt(db, CREDENTIAL_GUIDE_PROMPT_KEY)
|
||||||
|
response = await ai_client.analyze(
|
||||||
|
SituationalAnalysisRequest(
|
||||||
|
title=f"Generate credential guide for {provider}",
|
||||||
|
objective=f"{default.prompt}\n{prompt.prompt}",
|
||||||
|
system_prompt=prompt.system_prompt or None,
|
||||||
|
context={
|
||||||
|
"provider": provider,
|
||||||
|
"current_default_guide": default.markdown,
|
||||||
|
"product_context": "Planet collector credential settings",
|
||||||
|
"search_evidence": search_evidence,
|
||||||
|
},
|
||||||
|
observations=[
|
||||||
|
"Use concise Chinese markdown.",
|
||||||
|
"Prefer stable concepts over brittle UI labels.",
|
||||||
|
"Include verification and troubleshooting steps.",
|
||||||
|
"Include a short sources section with the provided URLs.",
|
||||||
|
],
|
||||||
|
constraints=[
|
||||||
|
"Do not ask the user for secrets.",
|
||||||
|
"Do not include fabricated screenshots.",
|
||||||
|
"Do not invent source URLs or product UI labels.",
|
||||||
|
"Use only the provided search_evidence as factual support.",
|
||||||
|
"Return markdown only.",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
markdown = response.content.strip()
|
||||||
|
if not markdown:
|
||||||
|
markdown = default.markdown
|
||||||
|
return await save_credential_guide(
|
||||||
|
db,
|
||||||
|
provider,
|
||||||
|
default.title,
|
||||||
|
markdown,
|
||||||
|
sources=search_evidence,
|
||||||
|
verification_status="verified_with_search_evidence",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _credential_guide_search_query(default: CredentialGuideDefault) -> str:
|
||||||
|
if default.provider == "barentswatch":
|
||||||
|
return "BarentsWatch developer tutorial AIS API OAuth client credentials"
|
||||||
|
if default.provider == "aisstream":
|
||||||
|
return "AISStream API key documentation websocket stream"
|
||||||
|
return f"{default.provider} API credentials documentation"
|
||||||
391
backend/app/services/custom_datasource_runtime.py
Normal file
391
backend/app/services/custom_datasource_runtime.py
Normal file
@@ -0,0 +1,391 @@
|
|||||||
|
"""Runtime helpers for mapped custom data sources."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
import base64
|
||||||
|
import json
|
||||||
|
from datetime import UTC, datetime
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
from sqlalchemy import select
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.target_schema_registry import TARGET_SCHEMAS
|
||||||
|
from app.db.session import async_session_factory
|
||||||
|
from app.models.datasource_config import DataSourceConfig
|
||||||
|
from app.models.datasource_mapping import DataSourceMappingTemplate
|
||||||
|
from app.services.datasource_mapping import (
|
||||||
|
MappingError,
|
||||||
|
execute_mapping,
|
||||||
|
extract_path,
|
||||||
|
persist_mapped_records,
|
||||||
|
)
|
||||||
|
|
||||||
|
DEFAULT_MAPPING_TEMPLATES: dict[str, dict[str, Any]] = {
|
||||||
|
"vessel_ais": {
|
||||||
|
"source": {"items_path": "$"},
|
||||||
|
"fields": {
|
||||||
|
"mmsi": {"path": "$.mmsi", "type": "integer"},
|
||||||
|
"name": {"path": "$.name", "type": "string", "default": None},
|
||||||
|
"lat": {"path": "$.lat", "type": "float"},
|
||||||
|
"lon": {"path": "$.lon", "type": "float"},
|
||||||
|
"sog": {"path": "$.sog", "type": "float", "default": None},
|
||||||
|
"cog": {"path": "$.cog", "type": "float", "default": None},
|
||||||
|
"heading": {"path": "$.heading", "type": "integer", "default": None},
|
||||||
|
"nav_status": {"path": "$.nav_status", "type": "integer", "default": None},
|
||||||
|
"callsign": {"path": "$.callsign", "type": "string", "default": None},
|
||||||
|
"vessel_type": {"path": "$.vessel_type", "type": "string", "default": None},
|
||||||
|
"vessel_type_name": {"path": "$.vessel_type_name", "type": "string", "default": None},
|
||||||
|
"received_at": {"path": "$.received_at", "type": "datetime", "default": None},
|
||||||
|
},
|
||||||
|
"meta": {"generated_by": "default_template", "requires_review": False},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
RUNNING_CUSTOM_STREAM_TASKS: dict[int, asyncio.Task[Any]] = {}
|
||||||
|
|
||||||
|
|
||||||
|
class CustomDatasourceRuntimeError(RuntimeError):
|
||||||
|
"""Raised when a custom datasource cannot run."""
|
||||||
|
|
||||||
|
|
||||||
|
def build_request_headers(auth_type: str, auth_config: dict, headers: dict) -> dict[str, str]:
|
||||||
|
request_headers = {str(key): str(value) for key, value in (headers or {}).items()}
|
||||||
|
auth_type = str(auth_type or "none").lower()
|
||||||
|
auth_config = auth_config or {}
|
||||||
|
|
||||||
|
if auth_type == "bearer" and auth_config.get("token"):
|
||||||
|
request_headers["Authorization"] = f"Bearer {auth_config['token']}"
|
||||||
|
elif auth_type == "api_key" and auth_config.get("api_key"):
|
||||||
|
location = str(auth_config.get("in") or auth_config.get("location") or "header").lower()
|
||||||
|
if location != "query":
|
||||||
|
key_name = auth_config.get("key_name", "X-API-Key")
|
||||||
|
request_headers[str(key_name)] = str(auth_config["api_key"])
|
||||||
|
elif auth_type == "basic":
|
||||||
|
username = auth_config.get("username", "")
|
||||||
|
password = auth_config.get("password", "")
|
||||||
|
credentials = f"{username}:{password}"
|
||||||
|
encoded = base64.b64encode(credentials.encode()).decode()
|
||||||
|
request_headers["Authorization"] = f"Basic {encoded}"
|
||||||
|
return request_headers
|
||||||
|
|
||||||
|
|
||||||
|
def build_query_params(auth_type: str, auth_config: dict, config: dict) -> dict[str, Any]:
|
||||||
|
params: dict[str, Any] = {}
|
||||||
|
candidate = (config or {}).get("params") or (config or {}).get("query_params")
|
||||||
|
if isinstance(candidate, dict):
|
||||||
|
params.update(candidate)
|
||||||
|
|
||||||
|
auth_type = str(auth_type or "none").lower()
|
||||||
|
auth_config = auth_config or {}
|
||||||
|
if auth_type == "api_key" and auth_config.get("api_key"):
|
||||||
|
location = str(auth_config.get("in") or auth_config.get("location") or "header").lower()
|
||||||
|
if location == "query":
|
||||||
|
key_name = auth_config.get("key_name") or auth_config.get("param_name") or "api_key"
|
||||||
|
params[str(key_name)] = auth_config["api_key"]
|
||||||
|
return params
|
||||||
|
|
||||||
|
|
||||||
|
async def load_active_mapping(
|
||||||
|
db: AsyncSession,
|
||||||
|
datasource_config_id: int,
|
||||||
|
) -> DataSourceMappingTemplate:
|
||||||
|
result = await db.execute(
|
||||||
|
select(DataSourceMappingTemplate)
|
||||||
|
.where(DataSourceMappingTemplate.datasource_config_id == datasource_config_id)
|
||||||
|
.where(DataSourceMappingTemplate.is_active.is_(True))
|
||||||
|
.order_by(DataSourceMappingTemplate.version.desc())
|
||||||
|
.limit(1)
|
||||||
|
)
|
||||||
|
mapping = result.scalar_one_or_none()
|
||||||
|
if mapping is not None:
|
||||||
|
return mapping
|
||||||
|
|
||||||
|
datasource = await db.get(DataSourceConfig, datasource_config_id)
|
||||||
|
if datasource is None:
|
||||||
|
raise CustomDatasourceRuntimeError("Configuration not found")
|
||||||
|
target_schema = (datasource.config or {}).get("target_schema")
|
||||||
|
template_body = DEFAULT_MAPPING_TEMPLATES.get(str(target_schema or "")) if target_schema else None
|
||||||
|
if not template_body or target_schema not in TARGET_SCHEMAS:
|
||||||
|
raise CustomDatasourceRuntimeError(
|
||||||
|
"No active mapping template found and no default template available for this target schema"
|
||||||
|
)
|
||||||
|
|
||||||
|
mapping = DataSourceMappingTemplate(
|
||||||
|
datasource_config_id=datasource_config_id,
|
||||||
|
target_schema=str(target_schema),
|
||||||
|
mapping_json=template_body,
|
||||||
|
sample_payload_hash=None,
|
||||||
|
validation_status="valid",
|
||||||
|
version=1,
|
||||||
|
is_active=True,
|
||||||
|
)
|
||||||
|
db.add(mapping)
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(mapping)
|
||||||
|
return mapping
|
||||||
|
|
||||||
|
|
||||||
|
async def fetch_rest_payload(config: DataSourceConfig, limit_bytes: int) -> Any:
|
||||||
|
request_config = config.config or {}
|
||||||
|
method = str(request_config.get("method") or request_config.get("request_method") or "GET").upper()
|
||||||
|
if method not in {"GET", "POST"}:
|
||||||
|
raise CustomDatasourceRuntimeError("Only GET and POST sample requests are supported.")
|
||||||
|
|
||||||
|
headers = build_request_headers(config.auth_type, config.auth_config or {}, config.headers or {})
|
||||||
|
params = build_query_params(config.auth_type, config.auth_config or {}, request_config)
|
||||||
|
timeout = float(request_config.get("timeout", 30))
|
||||||
|
json_body = request_config.get("json_body")
|
||||||
|
if json_body is None and str(request_config.get("body_type") or "").lower() in {"json", ""}:
|
||||||
|
candidate = request_config.get("body")
|
||||||
|
if isinstance(candidate, (dict, list)):
|
||||||
|
json_body = candidate
|
||||||
|
|
||||||
|
async with httpx.AsyncClient(timeout=timeout, follow_redirects=True) as client:
|
||||||
|
response = await client.request(
|
||||||
|
method,
|
||||||
|
config.endpoint,
|
||||||
|
headers=headers,
|
||||||
|
params=params or None,
|
||||||
|
json=json_body,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
content = response.content[:limit_bytes]
|
||||||
|
if "application/json" in response.headers.get("content-type", ""):
|
||||||
|
return json.loads(content.decode(response.encoding or "utf-8"))
|
||||||
|
return {"text": content.decode(response.encoding or "utf-8", errors="replace")}
|
||||||
|
|
||||||
|
|
||||||
|
async def run_mapped_rest_config(
|
||||||
|
db: AsyncSession,
|
||||||
|
datasource: DataSourceConfig,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
mapping = await load_active_mapping(db, datasource.id)
|
||||||
|
sample = await fetch_rest_payload(datasource, 5_000_000)
|
||||||
|
mapped = execute_mapping(sample, mapping.mapping_json, mapping.target_schema)
|
||||||
|
if mapped["failed_count"] > 0:
|
||||||
|
return {
|
||||||
|
"status": "failed",
|
||||||
|
"datasource_config_id": datasource.id,
|
||||||
|
"mapping_id": mapping.id,
|
||||||
|
"mapping_version": mapping.version,
|
||||||
|
"target_schema": mapping.target_schema,
|
||||||
|
"mapped_count": mapped["mapped_count"],
|
||||||
|
"failed_count": mapped["failed_count"],
|
||||||
|
"errors": mapped["errors"][:20],
|
||||||
|
}
|
||||||
|
|
||||||
|
request_config = datasource.config or {}
|
||||||
|
written_count = await persist_mapped_records(
|
||||||
|
db,
|
||||||
|
datasource_name=datasource.name,
|
||||||
|
datasource_config_id=datasource.id,
|
||||||
|
target_schema=mapping.target_schema,
|
||||||
|
records=mapped["records"],
|
||||||
|
mapping_version=mapping.version,
|
||||||
|
delivery_mode=request_config.get("delivery_mode") or "polling",
|
||||||
|
transport="http",
|
||||||
|
)
|
||||||
|
return {
|
||||||
|
"status": "success",
|
||||||
|
"datasource_config_id": datasource.id,
|
||||||
|
"mapping_id": mapping.id,
|
||||||
|
"mapping_version": mapping.version,
|
||||||
|
"target_schema": mapping.target_schema,
|
||||||
|
"fetched_count": mapped["total_items"],
|
||||||
|
"mapped_count": mapped["mapped_count"],
|
||||||
|
"written_count": written_count,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _items_from_ws_message(payload: Any, config: dict) -> Any:
|
||||||
|
message_path = config.get("ws_message_path")
|
||||||
|
items_path = config.get("ws_items_path")
|
||||||
|
value = extract_path(payload, message_path) if message_path else payload
|
||||||
|
return extract_path(value, items_path) if items_path else value
|
||||||
|
|
||||||
|
|
||||||
|
async def _connect_websocket(endpoint: str, headers: dict[str, str]):
|
||||||
|
import websockets
|
||||||
|
|
||||||
|
try:
|
||||||
|
return await websockets.connect(endpoint, additional_headers=headers or None)
|
||||||
|
except TypeError:
|
||||||
|
return await websockets.connect(endpoint, extra_headers=headers or None)
|
||||||
|
|
||||||
|
|
||||||
|
async def test_websocket_config(config: DataSourceConfig) -> dict[str, Any]:
|
||||||
|
if not str(config.endpoint or "").startswith(("ws://", "wss://")):
|
||||||
|
raise CustomDatasourceRuntimeError("WebSocket datasource endpoint must start with ws:// or wss://")
|
||||||
|
|
||||||
|
runtime_config = config.config or {}
|
||||||
|
headers = build_request_headers(config.auth_type, config.auth_config or {}, config.headers or {})
|
||||||
|
receive_timeout = float(runtime_config.get("receive_timeout_seconds") or runtime_config.get("timeout") or 10)
|
||||||
|
async with await _connect_websocket(config.endpoint, headers) as websocket:
|
||||||
|
subscribe_message = runtime_config.get("ws_subscribe_message")
|
||||||
|
if isinstance(subscribe_message, (dict, list)):
|
||||||
|
await websocket.send(json.dumps(subscribe_message))
|
||||||
|
elif isinstance(subscribe_message, str) and subscribe_message.strip():
|
||||||
|
await websocket.send(subscribe_message)
|
||||||
|
raw_message = await asyncio.wait_for(websocket.recv(), timeout=receive_timeout)
|
||||||
|
return {
|
||||||
|
"success": True,
|
||||||
|
"message_preview": raw_message[:1000] if isinstance(raw_message, str) else str(raw_message)[:1000],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def run_mapped_websocket_config(
|
||||||
|
db: AsyncSession,
|
||||||
|
datasource: DataSourceConfig,
|
||||||
|
*,
|
||||||
|
debug_max_messages: int | None = None,
|
||||||
|
use_config_debug_max_messages: bool = True,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
if not str(datasource.endpoint or "").startswith(("ws://", "wss://")):
|
||||||
|
raise CustomDatasourceRuntimeError("WebSocket datasource endpoint must start with ws:// or wss://")
|
||||||
|
|
||||||
|
mapping = await load_active_mapping(db, datasource.id)
|
||||||
|
runtime_config = datasource.config or {}
|
||||||
|
max_messages = debug_max_messages
|
||||||
|
if max_messages is None and use_config_debug_max_messages:
|
||||||
|
max_messages = runtime_config.get("debug_max_messages")
|
||||||
|
max_messages = int(max_messages) if max_messages else None
|
||||||
|
receive_timeout = float(runtime_config.get("receive_timeout_seconds") or runtime_config.get("timeout") or 30)
|
||||||
|
reconnect = bool(runtime_config.get("ws_reconnect", True))
|
||||||
|
reconnect_delay = float(runtime_config.get("reconnect_delay_seconds") or 3)
|
||||||
|
headers = build_request_headers(datasource.auth_type, datasource.auth_config or {}, datasource.headers or {})
|
||||||
|
|
||||||
|
messages_seen = 0
|
||||||
|
mapped_count = 0
|
||||||
|
failed_count = 0
|
||||||
|
written_count = 0
|
||||||
|
errors: list[dict[str, Any]] = []
|
||||||
|
started_at = datetime.now(UTC)
|
||||||
|
|
||||||
|
while True:
|
||||||
|
try:
|
||||||
|
async with await _connect_websocket(datasource.endpoint, headers) as websocket:
|
||||||
|
subscribe_message = runtime_config.get("ws_subscribe_message")
|
||||||
|
if isinstance(subscribe_message, (dict, list)):
|
||||||
|
await websocket.send(json.dumps(subscribe_message))
|
||||||
|
elif isinstance(subscribe_message, str) and subscribe_message.strip():
|
||||||
|
await websocket.send(subscribe_message)
|
||||||
|
|
||||||
|
while True:
|
||||||
|
raw_message = await asyncio.wait_for(websocket.recv(), timeout=receive_timeout)
|
||||||
|
messages_seen += 1
|
||||||
|
try:
|
||||||
|
payload = json.loads(raw_message)
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
failed_count += 1
|
||||||
|
errors.append({"message": "invalid_json", "error": str(exc)})
|
||||||
|
continue
|
||||||
|
|
||||||
|
extracted = _items_from_ws_message(payload, runtime_config)
|
||||||
|
try:
|
||||||
|
mapped = execute_mapping(extracted, mapping.mapping_json, mapping.target_schema)
|
||||||
|
except (MappingError, ValueError) as exc:
|
||||||
|
failed_count += 1
|
||||||
|
errors.append({"message": "mapping_failed", "error": str(exc)})
|
||||||
|
continue
|
||||||
|
|
||||||
|
mapped_count += mapped["mapped_count"]
|
||||||
|
failed_count += mapped["failed_count"]
|
||||||
|
if mapped["errors"]:
|
||||||
|
errors.extend(mapped["errors"][:5])
|
||||||
|
if mapped["records"]:
|
||||||
|
written_count += await persist_mapped_records(
|
||||||
|
db,
|
||||||
|
datasource_name=datasource.name,
|
||||||
|
datasource_config_id=datasource.id,
|
||||||
|
target_schema=mapping.target_schema,
|
||||||
|
records=mapped["records"],
|
||||||
|
mapping_version=mapping.version,
|
||||||
|
delivery_mode=runtime_config.get("delivery_mode") or "realtime_stream",
|
||||||
|
transport="websocket",
|
||||||
|
)
|
||||||
|
|
||||||
|
if max_messages and messages_seen >= max_messages:
|
||||||
|
return {
|
||||||
|
"status": "success",
|
||||||
|
"datasource_config_id": datasource.id,
|
||||||
|
"mapping_id": mapping.id,
|
||||||
|
"mapping_version": mapping.version,
|
||||||
|
"target_schema": mapping.target_schema,
|
||||||
|
"messages_seen": messages_seen,
|
||||||
|
"mapped_count": mapped_count,
|
||||||
|
"failed_count": failed_count,
|
||||||
|
"written_count": written_count,
|
||||||
|
"errors": errors[:20],
|
||||||
|
"execution_time_seconds": (datetime.now(UTC) - started_at).total_seconds(),
|
||||||
|
}
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
raise
|
||||||
|
except Exception as exc:
|
||||||
|
failed_count += 1
|
||||||
|
errors.append({"message": "websocket_error", "error": f"{exc.__class__.__name__}: {exc}"})
|
||||||
|
if not reconnect or max_messages:
|
||||||
|
return {
|
||||||
|
"status": "failed" if written_count == 0 else "partial",
|
||||||
|
"datasource_config_id": datasource.id,
|
||||||
|
"mapping_id": mapping.id,
|
||||||
|
"mapping_version": mapping.version,
|
||||||
|
"target_schema": mapping.target_schema,
|
||||||
|
"messages_seen": messages_seen,
|
||||||
|
"mapped_count": mapped_count,
|
||||||
|
"failed_count": failed_count,
|
||||||
|
"written_count": written_count,
|
||||||
|
"errors": errors[:20],
|
||||||
|
}
|
||||||
|
await asyncio.sleep(reconnect_delay)
|
||||||
|
|
||||||
|
|
||||||
|
async def run_custom_stream_by_id(config_id: int) -> dict[str, Any]:
|
||||||
|
async with async_session_factory() as db:
|
||||||
|
datasource = await db.get(DataSourceConfig, config_id)
|
||||||
|
if not datasource:
|
||||||
|
raise CustomDatasourceRuntimeError("Configuration not found")
|
||||||
|
return await run_mapped_websocket_config(
|
||||||
|
db,
|
||||||
|
datasource,
|
||||||
|
use_config_debug_max_messages=False,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def start_custom_stream(config_id: int) -> bool:
|
||||||
|
existing = RUNNING_CUSTOM_STREAM_TASKS.get(config_id)
|
||||||
|
if existing is not None and not existing.done():
|
||||||
|
return False
|
||||||
|
task = asyncio.create_task(run_custom_stream_by_id(config_id), name=f"custom-stream:{config_id}")
|
||||||
|
RUNNING_CUSTOM_STREAM_TASKS[config_id] = task
|
||||||
|
|
||||||
|
def _cleanup(done_task: asyncio.Task[Any]) -> None:
|
||||||
|
if RUNNING_CUSTOM_STREAM_TASKS.get(config_id) is done_task:
|
||||||
|
RUNNING_CUSTOM_STREAM_TASKS.pop(config_id, None)
|
||||||
|
|
||||||
|
task.add_done_callback(_cleanup)
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
async def stop_custom_stream(config_id: int) -> bool:
|
||||||
|
task = RUNNING_CUSTOM_STREAM_TASKS.get(config_id)
|
||||||
|
if task is None or task.done():
|
||||||
|
RUNNING_CUSTOM_STREAM_TASKS.pop(config_id, None)
|
||||||
|
return False
|
||||||
|
task.cancel()
|
||||||
|
try:
|
||||||
|
await task
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
return True
|
||||||
|
return task.cancelled()
|
||||||
|
|
||||||
|
|
||||||
|
def get_custom_stream_status(config_id: int) -> dict[str, Any]:
|
||||||
|
task = RUNNING_CUSTOM_STREAM_TASKS.get(config_id)
|
||||||
|
return {
|
||||||
|
"config_id": config_id,
|
||||||
|
"running": bool(task and not task.done()),
|
||||||
|
"done": bool(task and task.done()),
|
||||||
|
}
|
||||||
597
backend/app/services/data_jobs.py
Normal file
597
backend/app/services/data_jobs.py
Normal file
@@ -0,0 +1,597 @@
|
|||||||
|
"""Kafka-ready datasource job queue backed by PostgreSQL for v1."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import asyncio
|
||||||
|
from datetime import UTC, datetime, timedelta
|
||||||
|
from uuid import uuid4
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from sqlalchemy import bindparam, select, text
|
||||||
|
from sqlalchemy.ext.asyncio import AsyncSession
|
||||||
|
|
||||||
|
from app.core.cache import cache
|
||||||
|
from app.core.config import settings
|
||||||
|
from app.core.logging import get_logger
|
||||||
|
from app.core.time import to_iso8601_utc
|
||||||
|
from app.core.websocket.broadcaster import broadcaster
|
||||||
|
from app.db.session import async_session_factory
|
||||||
|
from app.models.collected_data import CollectedData
|
||||||
|
from app.models.data_snapshot import DataSnapshot
|
||||||
|
from app.models.datasource import DataSource
|
||||||
|
from app.models.task import CollectionTask
|
||||||
|
from app.services.collectors.registry import collector_registry
|
||||||
|
from app.services.datasource_connectivity import (
|
||||||
|
build_builtin_connectivity_checksum,
|
||||||
|
get_builtin_effective_candidate,
|
||||||
|
save_connectivity_success,
|
||||||
|
)
|
||||||
|
from app.services.earth_layer_adapters import (
|
||||||
|
clear_derived_datasource_data,
|
||||||
|
get_earth_refresh_strategy_for_change,
|
||||||
|
get_earth_update_layers_for_source,
|
||||||
|
)
|
||||||
|
from app.services.earth_layer_cache import invalidate_earth_layer_cache_for_source
|
||||||
|
from app.services.scheduler import sync_datasource_job
|
||||||
|
|
||||||
|
logger = get_logger(__name__)
|
||||||
|
|
||||||
|
JOB_TYPE_COLLECT = "collect"
|
||||||
|
JOB_TYPE_CLEAR_DATA = "clear_data"
|
||||||
|
JOB_TYPE_CLEAR_CACHE = "clear_cache"
|
||||||
|
JOB_TYPE_EARTH_REFRESH = "earth_refresh"
|
||||||
|
|
||||||
|
JOB_STATUS_QUEUED = "queued"
|
||||||
|
JOB_STATUS_RUNNING = "running"
|
||||||
|
JOB_STATUS_CANCELLING = "cancelling"
|
||||||
|
JOB_STATUS_SUCCESS = "success"
|
||||||
|
JOB_STATUS_FAILED = "failed"
|
||||||
|
JOB_STATUS_CANCELLED = "cancelled"
|
||||||
|
|
||||||
|
ACTIVE_JOB_STATUSES = (JOB_STATUS_QUEUED, JOB_STATUS_RUNNING, JOB_STATUS_CANCELLING)
|
||||||
|
TERMINAL_JOB_STATUSES = (JOB_STATUS_SUCCESS, JOB_STATUS_FAILED, JOB_STATUS_CANCELLED)
|
||||||
|
DATA_WRITE_JOB_TYPES = (JOB_TYPE_COLLECT, JOB_TYPE_CLEAR_DATA, JOB_TYPE_CLEAR_CACHE)
|
||||||
|
SOURCE_LOCK_JOB_STATUSES = (JOB_STATUS_RUNNING, JOB_STATUS_CANCELLING)
|
||||||
|
QUEUE_POLL_SECONDS = 0.35
|
||||||
|
JOB_STALE_LOCK_MINUTES = 90
|
||||||
|
DEFAULT_WORKER_CONCURRENCY = 2
|
||||||
|
|
||||||
|
RUNNING_DATA_JOB_TASKS: dict[int, asyncio.Task[Any]] = {}
|
||||||
|
|
||||||
|
|
||||||
|
def _utcnow() -> datetime:
|
||||||
|
return datetime.now(UTC)
|
||||||
|
|
||||||
|
|
||||||
|
def _job_worker_id() -> str:
|
||||||
|
return f"{settings.PROJECT_NAME}:data-job-worker:{uuid4().hex[:8]}"
|
||||||
|
|
||||||
|
|
||||||
|
def is_terminal_job_status(status: str | None) -> bool:
|
||||||
|
return status in TERMINAL_JOB_STATUSES
|
||||||
|
|
||||||
|
|
||||||
|
async def enqueue_datasource_job(
|
||||||
|
db: AsyncSession,
|
||||||
|
datasource: DataSource,
|
||||||
|
task_type: str,
|
||||||
|
*,
|
||||||
|
payload: dict[str, Any] | None = None,
|
||||||
|
rollback_policy: str = "keep_committed_batches",
|
||||||
|
dedupe_key: str | None = None,
|
||||||
|
) -> CollectionTask:
|
||||||
|
if dedupe_key:
|
||||||
|
existing = await _get_active_job_by_dedupe_key(db, dedupe_key)
|
||||||
|
if existing is not None:
|
||||||
|
return existing
|
||||||
|
|
||||||
|
task = CollectionTask(
|
||||||
|
datasource_id=datasource.id,
|
||||||
|
source=datasource.source,
|
||||||
|
task_type=task_type,
|
||||||
|
status=JOB_STATUS_QUEUED,
|
||||||
|
phase="queued",
|
||||||
|
phase_message="任务已进入队列",
|
||||||
|
payload=payload or {},
|
||||||
|
rollback_policy=rollback_policy,
|
||||||
|
dedupe_key=dedupe_key,
|
||||||
|
)
|
||||||
|
db.add(task)
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(task)
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
return task
|
||||||
|
|
||||||
|
|
||||||
|
async def enqueue_earth_refresh_job(
|
||||||
|
db: AsyncSession,
|
||||||
|
*,
|
||||||
|
source: str,
|
||||||
|
payload: dict[str, Any] | None = None,
|
||||||
|
) -> CollectionTask | None:
|
||||||
|
layers = list((payload or {}).get("layers") or get_earth_update_layers_for_source(source))
|
||||||
|
if not layers:
|
||||||
|
return None
|
||||||
|
|
||||||
|
datasource = await _get_or_create_virtual_datasource(db, source)
|
||||||
|
refresh_payload = {
|
||||||
|
"source": source,
|
||||||
|
"layers": layers,
|
||||||
|
"refresh_strategy": (payload or {}).get("refresh_strategy")
|
||||||
|
or get_earth_refresh_strategy_for_change((payload or {}).get("table"), source)
|
||||||
|
or "clear_then_reload",
|
||||||
|
**(payload or {}),
|
||||||
|
}
|
||||||
|
return await enqueue_datasource_job(
|
||||||
|
db,
|
||||||
|
datasource,
|
||||||
|
JOB_TYPE_EARTH_REFRESH,
|
||||||
|
payload=refresh_payload,
|
||||||
|
dedupe_key=f"earth_refresh:{source}",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
async def enqueue_earth_refresh_from_update(payload: dict[str, Any]) -> None:
|
||||||
|
source = str(payload.get("source") or "").strip()
|
||||||
|
if not source:
|
||||||
|
return
|
||||||
|
async with async_session_factory() as db:
|
||||||
|
await enqueue_earth_refresh_job(db, source=source, payload=payload)
|
||||||
|
|
||||||
|
|
||||||
|
async def request_cancel_datasource_task(
|
||||||
|
db: AsyncSession,
|
||||||
|
task: CollectionTask,
|
||||||
|
*,
|
||||||
|
reason: str = "cancelled_by_operator",
|
||||||
|
) -> CollectionTask:
|
||||||
|
if is_terminal_job_status(task.status):
|
||||||
|
return task
|
||||||
|
|
||||||
|
running_task = RUNNING_DATA_JOB_TASKS.get(task.id)
|
||||||
|
if task.status == JOB_STATUS_QUEUED or (
|
||||||
|
running_task is None
|
||||||
|
and (
|
||||||
|
task.status == JOB_STATUS_CANCELLING
|
||||||
|
or (task.status == JOB_STATUS_RUNNING and task.task_type != JOB_TYPE_COLLECT)
|
||||||
|
)
|
||||||
|
):
|
||||||
|
return await _cancel_task_without_runner(db, task, reason=reason)
|
||||||
|
|
||||||
|
task.status = JOB_STATUS_CANCELLING
|
||||||
|
task.phase = JOB_STATUS_CANCELLING
|
||||||
|
task.phase_message = "正在停止任务"
|
||||||
|
task.requested_cancel_at = _utcnow()
|
||||||
|
task.cancel_reason = reason
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(task)
|
||||||
|
|
||||||
|
if running_task is not None and not running_task.done():
|
||||||
|
running_task.cancel()
|
||||||
|
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
return task
|
||||||
|
|
||||||
|
|
||||||
|
async def _cancel_task_without_runner(
|
||||||
|
db: AsyncSession,
|
||||||
|
task: CollectionTask,
|
||||||
|
*,
|
||||||
|
reason: str,
|
||||||
|
) -> CollectionTask:
|
||||||
|
if task.task_type == JOB_TYPE_COLLECT:
|
||||||
|
await db.execute(CollectedData.__table__.delete().where(CollectedData.task_id == task.id))
|
||||||
|
snapshot_result = await db.execute(select(DataSnapshot).where(DataSnapshot.task_id == task.id))
|
||||||
|
for snapshot in snapshot_result.scalars().all():
|
||||||
|
snapshot.status = JOB_STATUS_CANCELLED
|
||||||
|
snapshot.completed_at = _utcnow()
|
||||||
|
snapshot.error_message = "Cancelled after operator stop request; no active worker handle remained"
|
||||||
|
datasource = await db.get(DataSource, task.datasource_id)
|
||||||
|
if datasource is not None:
|
||||||
|
datasource.last_status = JOB_STATUS_CANCELLED
|
||||||
|
|
||||||
|
task.status = JOB_STATUS_CANCELLED
|
||||||
|
task.phase = JOB_STATUS_CANCELLED
|
||||||
|
task.phase_message = "任务已停止"
|
||||||
|
task.completed_at = _utcnow()
|
||||||
|
task.requested_cancel_at = task.requested_cancel_at or _utcnow()
|
||||||
|
task.cancel_reason = reason
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(task)
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
return task
|
||||||
|
|
||||||
|
|
||||||
|
async def get_active_datasource_job(
|
||||||
|
db: AsyncSession,
|
||||||
|
datasource_id: int,
|
||||||
|
*,
|
||||||
|
task_types: tuple[str, ...] = DATA_WRITE_JOB_TYPES,
|
||||||
|
) -> CollectionTask | None:
|
||||||
|
result = await db.execute(
|
||||||
|
select(CollectionTask)
|
||||||
|
.where(CollectionTask.datasource_id == datasource_id)
|
||||||
|
.where(CollectionTask.task_type.in_(task_types))
|
||||||
|
.where(CollectionTask.status.in_(ACTIVE_JOB_STATUSES))
|
||||||
|
.order_by(CollectionTask.created_at.desc().nullslast(), CollectionTask.id.desc())
|
||||||
|
.limit(1)
|
||||||
|
)
|
||||||
|
return result.scalar_one_or_none()
|
||||||
|
|
||||||
|
|
||||||
|
async def _get_active_job_by_dedupe_key(db: AsyncSession, dedupe_key: str) -> CollectionTask | None:
|
||||||
|
result = await db.execute(
|
||||||
|
select(CollectionTask)
|
||||||
|
.where(CollectionTask.dedupe_key == dedupe_key)
|
||||||
|
.where(CollectionTask.status.in_(ACTIVE_JOB_STATUSES))
|
||||||
|
.order_by(CollectionTask.created_at.desc().nullslast(), CollectionTask.id.desc())
|
||||||
|
.limit(1)
|
||||||
|
)
|
||||||
|
return result.scalar_one_or_none()
|
||||||
|
|
||||||
|
|
||||||
|
async def _get_or_create_virtual_datasource(db: AsyncSession, source: str) -> DataSource:
|
||||||
|
result = await db.execute(select(DataSource).where(DataSource.source == source))
|
||||||
|
datasource = result.scalar_one_or_none()
|
||||||
|
if datasource is not None:
|
||||||
|
return datasource
|
||||||
|
|
||||||
|
datasource = DataSource(
|
||||||
|
name=f"Earth refresh: {source}",
|
||||||
|
source=source,
|
||||||
|
module="SYS",
|
||||||
|
collector_class="EarthRefreshJob",
|
||||||
|
is_active=True,
|
||||||
|
)
|
||||||
|
db.add(datasource)
|
||||||
|
await db.commit()
|
||||||
|
await db.refresh(datasource)
|
||||||
|
return datasource
|
||||||
|
|
||||||
|
|
||||||
|
async def _broadcast_task_update(task: CollectionTask) -> None:
|
||||||
|
await broadcaster.broadcast_datasource_task_update(
|
||||||
|
{
|
||||||
|
"datasource_id": task.datasource_id,
|
||||||
|
"collector_name": task.source,
|
||||||
|
"task_id": task.id,
|
||||||
|
"task_type": task.task_type,
|
||||||
|
"status": task.status,
|
||||||
|
"phase": task.phase,
|
||||||
|
"phase_progress": task.phase_progress,
|
||||||
|
"phase_message": task.phase_message,
|
||||||
|
"phase_current": task.phase_current,
|
||||||
|
"phase_total": task.phase_total,
|
||||||
|
"phase_unit": task.phase_unit,
|
||||||
|
"progress": task.progress,
|
||||||
|
"records_processed": task.records_processed,
|
||||||
|
"total_records": task.total_records,
|
||||||
|
"started_at": to_iso8601_utc(task.started_at),
|
||||||
|
"completed_at": to_iso8601_utc(task.completed_at),
|
||||||
|
"requested_cancel_at": to_iso8601_utc(task.requested_cancel_at),
|
||||||
|
"error_message": task.error_message,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class DataJobWorker:
|
||||||
|
def __init__(self, *, concurrency: int = DEFAULT_WORKER_CONCURRENCY) -> None:
|
||||||
|
self.worker_id = _job_worker_id()
|
||||||
|
self.concurrency = max(1, concurrency)
|
||||||
|
self._task: asyncio.Task[None] | None = None
|
||||||
|
self._stop_event: asyncio.Event | None = None
|
||||||
|
self._running: set[asyncio.Task[Any]] = set()
|
||||||
|
|
||||||
|
def start(self) -> None:
|
||||||
|
if self._task and not self._task.done():
|
||||||
|
return
|
||||||
|
self._stop_event = asyncio.Event()
|
||||||
|
self._task = asyncio.create_task(self._run(), name="data-job-worker")
|
||||||
|
|
||||||
|
async def stop(self) -> None:
|
||||||
|
if self._stop_event:
|
||||||
|
self._stop_event.set()
|
||||||
|
for task in list(self._running):
|
||||||
|
task.cancel()
|
||||||
|
if self._task:
|
||||||
|
await asyncio.gather(self._task, return_exceptions=True)
|
||||||
|
if self._running:
|
||||||
|
await asyncio.gather(*self._running, return_exceptions=True)
|
||||||
|
|
||||||
|
async def _run(self) -> None:
|
||||||
|
assert self._stop_event is not None
|
||||||
|
await self._recover_stale_running_jobs()
|
||||||
|
while not self._stop_event.is_set():
|
||||||
|
self._running = {task for task in self._running if not task.done()}
|
||||||
|
if len(self._running) >= self.concurrency:
|
||||||
|
await asyncio.sleep(QUEUE_POLL_SECONDS)
|
||||||
|
continue
|
||||||
|
|
||||||
|
task_id = await self._claim_next_job()
|
||||||
|
if task_id is None:
|
||||||
|
await asyncio.sleep(QUEUE_POLL_SECONDS)
|
||||||
|
continue
|
||||||
|
|
||||||
|
runner = asyncio.create_task(self._run_claimed_job(task_id), name=f"data-job:{task_id}")
|
||||||
|
self._running.add(runner)
|
||||||
|
|
||||||
|
async def _recover_stale_running_jobs(self) -> None:
|
||||||
|
cutoff = _utcnow() - timedelta(minutes=JOB_STALE_LOCK_MINUTES)
|
||||||
|
async with async_session_factory() as db:
|
||||||
|
result = await db.execute(
|
||||||
|
select(CollectionTask)
|
||||||
|
.where(CollectionTask.status.in_((JOB_STATUS_RUNNING, JOB_STATUS_CANCELLING)))
|
||||||
|
.where(CollectionTask.locked_at.is_not(None))
|
||||||
|
.where(CollectionTask.locked_at < cutoff)
|
||||||
|
)
|
||||||
|
stale_jobs = list(result.scalars().all())
|
||||||
|
for job in stale_jobs:
|
||||||
|
job.status = JOB_STATUS_FAILED
|
||||||
|
job.phase = JOB_STATUS_FAILED
|
||||||
|
job.completed_at = _utcnow()
|
||||||
|
job.error_message = "Marked failed after stale data job lock timeout"
|
||||||
|
if stale_jobs:
|
||||||
|
await db.commit()
|
||||||
|
|
||||||
|
async def _claim_next_job(self) -> int | None:
|
||||||
|
async with async_session_factory() as db:
|
||||||
|
row = await db.execute(
|
||||||
|
text(
|
||||||
|
"""
|
||||||
|
SELECT queued.id
|
||||||
|
FROM collection_tasks AS queued
|
||||||
|
WHERE queued.status = :queued_status
|
||||||
|
AND NOT EXISTS (
|
||||||
|
SELECT 1
|
||||||
|
FROM collection_tasks AS active
|
||||||
|
WHERE active.source = queued.source
|
||||||
|
AND active.id <> queued.id
|
||||||
|
AND active.status IN :active_statuses
|
||||||
|
)
|
||||||
|
ORDER BY queued.created_at ASC NULLS FIRST, queued.id ASC
|
||||||
|
LIMIT 1
|
||||||
|
FOR UPDATE SKIP LOCKED
|
||||||
|
"""
|
||||||
|
).bindparams(bindparam("active_statuses", expanding=True)),
|
||||||
|
{
|
||||||
|
"queued_status": JOB_STATUS_QUEUED,
|
||||||
|
"active_statuses": SOURCE_LOCK_JOB_STATUSES,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
task_id = row.scalar_one_or_none()
|
||||||
|
if task_id is None:
|
||||||
|
return None
|
||||||
|
|
||||||
|
task = await db.get(CollectionTask, int(task_id))
|
||||||
|
if task is None:
|
||||||
|
return None
|
||||||
|
task.status = JOB_STATUS_RUNNING
|
||||||
|
task.phase = "starting"
|
||||||
|
task.phase_message = "任务开始执行"
|
||||||
|
task.started_at = task.started_at or _utcnow()
|
||||||
|
task.worker_id = self.worker_id
|
||||||
|
task.locked_at = _utcnow()
|
||||||
|
await db.commit()
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
return int(task_id)
|
||||||
|
|
||||||
|
async def _run_claimed_job(self, task_id: int) -> None:
|
||||||
|
current_task = asyncio.current_task()
|
||||||
|
if current_task is not None:
|
||||||
|
RUNNING_DATA_JOB_TASKS[task_id] = current_task
|
||||||
|
try:
|
||||||
|
async with async_session_factory() as db:
|
||||||
|
task = await db.get(CollectionTask, task_id)
|
||||||
|
if task is None:
|
||||||
|
return
|
||||||
|
await self._execute_job(db, task)
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
async with async_session_factory() as db:
|
||||||
|
task = await db.get(CollectionTask, task_id)
|
||||||
|
if task is not None and not is_terminal_job_status(task.status):
|
||||||
|
task.status = JOB_STATUS_CANCELLED
|
||||||
|
task.phase = JOB_STATUS_CANCELLED
|
||||||
|
task.phase_message = "任务已停止"
|
||||||
|
task.completed_at = _utcnow()
|
||||||
|
await db.commit()
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
raise
|
||||||
|
except Exception as exc:
|
||||||
|
logger.exception_event(
|
||||||
|
"Data job failed",
|
||||||
|
event="data_jobs.job_failed",
|
||||||
|
context={"task_id": task_id, "error": str(exc)},
|
||||||
|
)
|
||||||
|
async with async_session_factory() as db:
|
||||||
|
task = await db.get(CollectionTask, task_id)
|
||||||
|
if task is not None:
|
||||||
|
task.status = JOB_STATUS_FAILED
|
||||||
|
task.phase = JOB_STATUS_FAILED
|
||||||
|
task.phase_message = str(exc)
|
||||||
|
task.error_message = str(exc)
|
||||||
|
task.completed_at = _utcnow()
|
||||||
|
await db.commit()
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
finally:
|
||||||
|
RUNNING_DATA_JOB_TASKS.pop(task_id, None)
|
||||||
|
|
||||||
|
async def _execute_job(self, db: AsyncSession, task: CollectionTask) -> None:
|
||||||
|
if task.task_type == JOB_TYPE_COLLECT:
|
||||||
|
await _run_collect_job(db, task)
|
||||||
|
elif task.task_type == JOB_TYPE_CLEAR_DATA:
|
||||||
|
await _run_clear_data_job(db, task)
|
||||||
|
elif task.task_type == JOB_TYPE_CLEAR_CACHE:
|
||||||
|
await _run_clear_cache_job(db, task)
|
||||||
|
elif task.task_type == JOB_TYPE_EARTH_REFRESH:
|
||||||
|
await _run_earth_refresh_job(db, task)
|
||||||
|
else:
|
||||||
|
raise RuntimeError(f"Unsupported data job type: {task.task_type}")
|
||||||
|
|
||||||
|
|
||||||
|
async def _run_collect_job(db: AsyncSession, task: CollectionTask) -> None:
|
||||||
|
datasource = await db.get(DataSource, task.datasource_id)
|
||||||
|
if datasource is None:
|
||||||
|
raise RuntimeError("Data source not found")
|
||||||
|
|
||||||
|
collector = collector_registry.get(datasource.source)
|
||||||
|
if collector is None:
|
||||||
|
raise RuntimeError(f"Collector '{datasource.source}' not found")
|
||||||
|
if not datasource.is_active:
|
||||||
|
raise RuntimeError("Data source is disabled")
|
||||||
|
|
||||||
|
collector._datasource_id = datasource.id
|
||||||
|
collector._current_task = task
|
||||||
|
collector._db_session = db
|
||||||
|
result = await collector.run(db)
|
||||||
|
|
||||||
|
datasource.last_run_at = _utcnow()
|
||||||
|
datasource.last_status = result.get("status")
|
||||||
|
if datasource.last_status == JOB_STATUS_SUCCESS:
|
||||||
|
effective_candidate = await get_builtin_effective_candidate(db, datasource.source)
|
||||||
|
checksum, _credential_context = await build_builtin_connectivity_checksum(
|
||||||
|
datasource.source,
|
||||||
|
effective_candidate["endpoint"],
|
||||||
|
effective_candidate["auth_type"],
|
||||||
|
effective_candidate["headers"],
|
||||||
|
effective_candidate["config"],
|
||||||
|
db,
|
||||||
|
)
|
||||||
|
await save_connectivity_success(
|
||||||
|
db,
|
||||||
|
datasource.source,
|
||||||
|
checksum,
|
||||||
|
{"status_code": None},
|
||||||
|
connected_by="collection",
|
||||||
|
)
|
||||||
|
await db.commit()
|
||||||
|
await sync_datasource_job(datasource.id)
|
||||||
|
|
||||||
|
|
||||||
|
async def _run_clear_data_job(db: AsyncSession, task: CollectionTask) -> None:
|
||||||
|
source = str(task.source or (task.payload or {}).get("source") or "").strip()
|
||||||
|
if not source:
|
||||||
|
raise RuntimeError("Clear data job has no source")
|
||||||
|
|
||||||
|
task.phase = "clearing_data"
|
||||||
|
task.phase_message = "正在删除数据库数据"
|
||||||
|
await db.commit()
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
|
||||||
|
count_result = await db.execute(
|
||||||
|
select(CollectedData.id).where(CollectedData.source == source)
|
||||||
|
)
|
||||||
|
collected_ids = [row[0] for row in count_result.all()]
|
||||||
|
derived_deleted_counts = await clear_derived_datasource_data(db, source)
|
||||||
|
if collected_ids:
|
||||||
|
await db.execute(CollectedData.__table__.delete().where(CollectedData.id.in_(collected_ids)))
|
||||||
|
deleted_count = len(collected_ids)
|
||||||
|
derived_deleted_count = sum(derived_deleted_counts.values())
|
||||||
|
|
||||||
|
task.records_processed = deleted_count + derived_deleted_count
|
||||||
|
task.total_records = task.records_processed
|
||||||
|
task.progress = 100.0
|
||||||
|
task.phase_progress = 100.0
|
||||||
|
task.phase_current = task.records_processed
|
||||||
|
task.phase_total = task.records_processed
|
||||||
|
task.phase_unit = "records"
|
||||||
|
task.payload = {
|
||||||
|
**(task.payload or {}),
|
||||||
|
"deleted_count": deleted_count,
|
||||||
|
"derived_deleted_count": derived_deleted_count,
|
||||||
|
"derived_deleted_counts": derived_deleted_counts,
|
||||||
|
}
|
||||||
|
task.status = JOB_STATUS_SUCCESS
|
||||||
|
task.phase = "completed"
|
||||||
|
task.phase_message = "数据库数据已清理"
|
||||||
|
task.completed_at = _utcnow()
|
||||||
|
await db.commit()
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
|
||||||
|
|
||||||
|
async def _run_clear_cache_job(db: AsyncSession, task: CollectionTask) -> None:
|
||||||
|
source = str(task.source or (task.payload or {}).get("source") or "").strip()
|
||||||
|
if not source:
|
||||||
|
raise RuntimeError("Clear cache job has no source")
|
||||||
|
|
||||||
|
earth_deleted_count = invalidate_earth_layer_cache_for_source(source)
|
||||||
|
dashboard_deleted_count = int(cache.delete("dashboard:stats")) + int(cache.delete("dashboard:summary"))
|
||||||
|
deleted_count = earth_deleted_count + dashboard_deleted_count
|
||||||
|
|
||||||
|
task.records_processed = deleted_count
|
||||||
|
task.total_records = deleted_count
|
||||||
|
task.progress = 100.0
|
||||||
|
task.phase_progress = 100.0
|
||||||
|
task.phase = "completed"
|
||||||
|
task.phase_message = "缓存已清理"
|
||||||
|
task.payload = {
|
||||||
|
**(task.payload or {}),
|
||||||
|
"earth_layer_deleted_count": earth_deleted_count,
|
||||||
|
"dashboard_deleted_count": dashboard_deleted_count,
|
||||||
|
}
|
||||||
|
task.status = JOB_STATUS_SUCCESS
|
||||||
|
task.completed_at = _utcnow()
|
||||||
|
await db.commit()
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
await enqueue_earth_refresh_job(db, source=source, payload={"operation": "CACHE_INVALIDATED"})
|
||||||
|
|
||||||
|
|
||||||
|
async def _run_earth_refresh_job(db: AsyncSession, task: CollectionTask) -> None:
|
||||||
|
payload = task.payload or {}
|
||||||
|
source = str(payload.get("source") or task.source or "").strip()
|
||||||
|
layers = list(payload.get("layers") or get_earth_update_layers_for_source(source))
|
||||||
|
if not source or not layers:
|
||||||
|
task.status = JOB_STATUS_SUCCESS
|
||||||
|
task.phase = "completed"
|
||||||
|
task.phase_message = "没有需要刷新的 Earth 图层"
|
||||||
|
task.completed_at = _utcnow()
|
||||||
|
await db.commit()
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
return
|
||||||
|
|
||||||
|
deleted_cache_entries = invalidate_earth_layer_cache_for_source(source)
|
||||||
|
update_payload = {
|
||||||
|
"event": "earth.layer.changed",
|
||||||
|
"action": "database_changed",
|
||||||
|
"source": source,
|
||||||
|
"table": payload.get("table"),
|
||||||
|
"data_type": source,
|
||||||
|
"layers": layers,
|
||||||
|
"refresh_strategy": payload.get("refresh_strategy") or "clear_then_reload",
|
||||||
|
"records_processed": payload.get("records_processed", 0),
|
||||||
|
"operations": payload.get("operations") or [payload.get("operation") or "CHANGE"],
|
||||||
|
"operation": payload.get("operation"),
|
||||||
|
"cache_entries_invalidated": deleted_cache_entries,
|
||||||
|
"timestamp": to_iso8601_utc(_utcnow()),
|
||||||
|
}
|
||||||
|
if payload.get("entity") == "interactable":
|
||||||
|
update_payload.update(
|
||||||
|
{
|
||||||
|
"entity": "interactable",
|
||||||
|
"action": payload.get("action") or "changed",
|
||||||
|
"ids": payload.get("ids") or payload.get("entity_keys") or [],
|
||||||
|
"item": payload.get("item"),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
await broadcaster.broadcast_earth_update(update_payload)
|
||||||
|
|
||||||
|
task.records_processed = int(payload.get("records_processed") or 0)
|
||||||
|
task.progress = 100.0
|
||||||
|
task.phase_progress = 100.0
|
||||||
|
task.phase = "completed"
|
||||||
|
task.phase_message = "Earth 图层刷新通知已发送"
|
||||||
|
task.status = JOB_STATUS_SUCCESS
|
||||||
|
task.completed_at = _utcnow()
|
||||||
|
task.payload = {**payload, "cache_entries_invalidated": deleted_cache_entries}
|
||||||
|
await db.commit()
|
||||||
|
await _broadcast_task_update(task)
|
||||||
|
|
||||||
|
|
||||||
|
_worker = DataJobWorker()
|
||||||
|
|
||||||
|
|
||||||
|
def start_data_job_worker() -> None:
|
||||||
|
_worker.start()
|
||||||
|
|
||||||
|
|
||||||
|
async def stop_data_job_worker() -> None:
|
||||||
|
await _worker.stop()
|
||||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user