Files
planet/docs/technical/en/ops-runbook.md
rayd1o 58671e7bc3
Some checks failed
ci / backend (push) Has been cancelled
ci / frontend (push) Has been cancelled
ci / delivery (push) Has been cancelled
release / images (push) Has been cancelled
release: bump version to 0.74.4
2026-09-13 10:27:00 +08:00

19 KiB

Intelligent Planet Ops Runbook

This runbook is for deployment, on-call, and maintenance engineers. End-user UI flows live in the Intelligent Planet Manual; this document only covers shell, Docker, logs, environment variables, and troubleshooting.

Docker Initialization and Access

Initialize a new machine before starting the application services:

zsh ./planet.sh init --non-motion-agent && zsh ./planet.sh start --non-motion-agent

The entry point still requires zsh, curl, and reachable package repositories. Before synchronizing Python and frontend dependencies, init prepares Docker:

  • Reuse working Docker, Compose v2, and Buildx (at least 0.17.0).
  • On Ubuntu / Ubuntu WSL, use apt to install the missing parts of docker.io, docker-compose-v2, and docker-buildx. When Docker CE CLI is already installed, use the configured Docker CE repository and corresponding plugin packages to keep the package family consistent.
  • If the local daemon is unavailable, check that docker.service exists, then enable and start it. WSL must have systemd enabled; unavailable service management produces an explicit Docker preparation error.
  • If the current user cannot read and write the Docker socket, check for usermod, install its passwd package when needed, and add the user to the docker group. This group grants privileged control of the local Docker engine. The script uses sudo to refresh group access as the original user and continue the original command with its arguments preserved. It does not depend on sg or run application processes as root.

When elevation is needed, sudo authentication runs in the foreground. Missing sudo for an unprivileged user, authentication failure, repository errors, or insufficient versions after installation stop initialization with a specific error.

Subsequent planet.sh start and other service commands in the same old terminal also refresh Docker group membership when it has been granted but is not yet active. Open a new Ubuntu session to use docker directly in the terminal.

When Docker Desktop is present but its WSL integration is unavailable, the script asks the operator to start Desktop and enable WSL Integration for the distribution. Unreachable remote or rootless endpoints produce a diagnostic for that environment; neither case installs a second local engine automatically. Automatic installation on other operating systems is not currently supported.

planet.sh calls scripts/lib/docker-bootstrap.zsh for preparation. Missing CLI, missing service units, socket permissions, and stopped daemons receive separate diagnostics. Advice to start docker.socket is shown only after confirming that the unit exists. Verify the result with:

docker info
docker compose version
docker buildx version

Database Initialization and Connection Checks

init and start reconcile PostgreSQL / Redis containers through Compose, including port configuration on existing containers. A plain docker start cannot apply configuration changes. Compose failures retain their specific errors, such as an occupied port, instead of falling back to an old container and reporting success.

The container's pg_isready check only establishes that the server accepts connections; it does not validate the host backend's address and credentials. Once containers are healthy, both initialization and backend startup run scripts/check_database_connection.py using the backend's effective DATABASE_URL. It checks the local PostgreSQL published port and executes a read-only SELECT 1. Startup performs this check before preparing the AI Provider image and stops immediately on failure. Initialization only creates tables and seed data after the check passes.

  • If the actual local port mapping is still missing or mismatched, the script recreates PostgreSQL once from Compose while preserving its data volume, then checks again. A second failure stops initialization.
  • Authentication, database-name, and network failures stop before schema changes. Diagnostics show the host, port, and database name without passwords, full connection strings, or raw driver exceptions.
  • A process-level DATABASE_URL overrides backend/.env. Changing POSTGRES_PASSWORD alone updates neither the connection string nor the password stored in an existing data volume. Existing environment files are retained and their effective configuration must be checked.
  • Explicit external databases do not require a local container mapping. Host networking also does not require published ports. Both still require the real connection check.

For port is already allocated or address already in use, inspect docker ps port information and ss -ltnp '( sport = :5432 )'. With WSL mirrored networking, also inspect Windows listeners. Initialization does not kill other database services to acquire a port, delete data volumes, or reset passwords.

First Startup

./planet.sh start

Default behavior:

  • Starts PostgreSQL and Redis
  • Starts AI Provider
  • Starts the backend API
  • Starts the frontend Vite dev server
  • Prints Earth, console, Playground, and backend API doc URLs

First startup seeds two default accounts (see DEFAULT_LOGIN_USERS in backend/app/db/session.py):

Username Password Role
admin admin123 super_admin
linkong LK12345678 super_admin

Both seed accounts are created with email_verified = TRUE and can log into the console immediately. Any other account must either go through the public registration flow described in the Manual, or be created via ./planet.sh createuser.

Specify custom ports:

./planet.sh start -b 8001 -f 3001 -a 8101
Flag Meaning
-b <port> Backend port
-f <port> Frontend port
-a <port> AI Provider port
--allow-lan Enable LAN access
--verbose Show extra command output

Stop and Per-Module Restart

Stop everything:

./planet.sh stop

Stops backend, AI Provider, frontend, PostgreSQL, Redis.

Per-module restart:

./planet.sh restart        # full
./planet.sh restart -b     # backend
./planet.sh restart -f     # frontend
./planet.sh restart -a     # AI Provider
./planet.sh restart -d     # database

Per-module restart is preferred during development to avoid interrupting unrelated services.

Destructive Reset

./planet.sh destroy

destroy returns a local development environment to a near-empty project state. It requires typing Y before it runs; source files and existing .env files are preserved.

Cleanup order and boundaries:

  • If planet_postgres is running, the script first clears the public schema in planet_db. This prevents old collected_data.is_current = true rows from making Earth OOBE report ready=true if Docker volume removal later fails.
  • Docker cleanup targets resources whose Compose project is planet, plus the explicit volumes planet_postgres_data, planet_redis_data, postgres_data, and redis_data; do not delete unlabeled volumes by a broad planet_* pattern, because another local project could own them.
  • Local build state removes .venv, frontend node_modules / dist, Planet state, and scattered Python / Vite cache directories. $PLANET_CACHE_DIR/downloads is preserved so upstream raw downloads such as CelesTrak can survive database resets and local rebuild cleanup.

After the reset, run ./planet.sh init again to recreate tables and default seed data. Old collected records are not restored, and Earth OOBE is evaluated from the backend's real collection state on the next visit. When CelesTrak later returns its "GP data has not updated" HTTP 403, the backend first reuses the preserved download cache to repopulate the database; if no active cache exists, it tries valid CelesTrak fallback group caches; if no download cache exists at all, wait for the next CelesTrak update window or use Space-Track. Datasource Clear Data and Clear Cache actions in the console do not delete $PLANET_CACHE_DIR/downloads/celestrak.

Health Check

./planet.sh health

Checks:

  • planet_* container status
  • Backend /health
  • AI Provider /health
  • Frontend reachability

If anything reports offline, check the corresponding logs first.

Logs

Recent logs:

./planet.sh log

Follow:

./planet.sh log -f   # frontend: ~/.local/state/planet/frontend.log
./planet.sh log -b   # backend:  ~/.local/state/planet/backend.log
./planet.sh log -a   # AI Provider: planet_aiprovider container logs

CLI User Creation

./planet.sh createuser

Interactively prompts for username, password, and role; writes the user with email_verified = TRUE directly.

Use when:

  • SMTP is not yet configured but an admin account is needed now
  • Pre-seeding internal test accounts
  • Public registration is unavailable for any reason and a fallback is required

For ordinary user onboarding, configure SMTP at /settings -> SMTP Email first and let users self-register at /register.

LAN / WSL Access

./planet.sh start --allow-lan

Useful for:

  • Starting in WSL, accessing from Windows browser
  • Demoing Earth from a phone or tablet
  • Other LAN machines reaching the same dev instance

On Windows, the repository-root planet.cmd can be used as a one-click entrypoint. It requests Administrator privileges, enters the Ubuntu WSL distribution at /home/linkong/planet, runs ./planet.sh restart --allow-lan, opens http://localhost:3000/earth after a successful restart, and leaves the terminal inside a WSL shell for log inspection. If the local WSL distribution name or checkout path differs, adjust the wsl.exe -d ... --cd ... arguments in planet.cmd first.

On a new Windows machine, check the WSL generation first:

wsl -l -v

Planet development should use WSL2. WSL1 has different networking, filesystem, and process behavior, and can surface as Bun package-manager commands returning only An unknown error occurred (Unexpected), unstable port release, or LAN behavior that does not match the script's assumptions. Convert the distribution if it still runs as WSL1:

wsl --set-version Ubuntu 2

--allow-lan directly exposes the frontend, backend, and AI Provider from the development machine: frontend 3000, backend 8000, and AI Provider 8010. Before startup, the script checks all three ports. If WSL/Linux cannot release a port and a Windows-side listener or stale portproxy rule owns it, the script requests Administrator PowerShell cleanup. When Planet runs in WSL, Windows can usually reach it through localhost; other LAN machines reaching the Windows LAN IP still need Windows Firewall allow rules.

Diagnose in this order:

# From the shell running Planet
curl http://localhost:3000
curl http://localhost:8000/health
curl http://localhost:8010/health
ss -ltnp | grep -E ':3000|:8000|:8010'

If the services are running but the LAN IP still fails, first remove stale portproxy rules and confirm Windows Firewall allows the ports. The script checks this automatically and requests Administrator PowerShell when needed. Manual fallback commands:

netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=3000
netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=8000
netsh interface portproxy delete v4tov4 listenaddress=0.0.0.0 listenport=8010

New-NetFirewallRule -DisplayName "WSL Planet 3000" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 3000
New-NetFirewallRule -DisplayName "WSL Planet 8000" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 8000
New-NetFirewallRule -DisplayName "WSL Planet 8010" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 8010

LAN devices should use the Windows external ports, for example http://<Windows LAN IP>:3000/earth, http://<Windows LAN IP>:8000/health, and http://<Windows LAN IP>:8010/health.

AI Provider Environment and Builds

AI Provider runtime configuration lives in two places:

Location Best for Notes
aiprovider/.env Team-shared local defaults Read by Docker Compose as env_file
~/.zshrc Personal provider/model/key/proxy planet.sh reads common AI_*, SERVICE_*, PYTHON_IMAGE, UV_IMAGE lines

Recommended form:

export AI_PROVIDER=minimax
export AI_PROVIDER_API=anthropic-messages
export AI_BASE_URL=https://api.example.com/anthropic
export AI_API_KEY=sk-change-me
export AI_MODEL=MiniMax-M2.7
export AI_PROVIDER_SERVICE_TOKEN=change_me

By default planet.sh only statically parses simple export KEY=value lines from ~/.zshrc. When complex shell expansion is required, opt in explicitly:

PLANET_LOAD_ZSHRC_ENV=source ./planet.sh start -a

To ignore ~/.zshrc entirely:

PLANET_LOAD_ZSHRC_ENV=0 ./planet.sh start -a

The AI Provider image only rebuilds when code, Dockerfile, Compose config, or Python dependencies change. After changing keys or base URL, restarting the container is enough:

./planet.sh restart -a

Rebuild detection is based on a content fingerprint rather than only file mtimes. planet.sh hashes the aiprovider/ files, aiprovider/Dockerfile, pyproject.toml, uv.lock, PYTHON_IMAGE, UV_IMAGE, and the dependency fingerprint into AI_PROVIDER_BUILD_FINGERPRINT; Docker writes it into the image label planet.aiprovider.build-fingerprint. If the existing planet-aiprovider:latest image has a matching label, the script skips rebuild and refreshes the local stamp. Older images without the label fall back to the state/cache stamp.

Docker builds use uv sync --frozen. To make container builds reuse the host uv mirror configuration, the script resolves the first config file in this order and mounts it into the build as a BuildKit secret at /root/.config/uv/uv.toml:

  1. The current UV_CONFIG_FILE
  2. Repository-root uv.toml
  3. ${XDG_CONFIG_HOME:-~/.config}/uv/uv.toml
  4. ~/.uv/uv.toml

If none exists, the script creates an empty state file for the secret so Compose does not fail on a missing file. Before Docker build it unsets UV_DEFAULT_INDEX, UV_INDEX_URL, and UV_EXTRA_INDEX_URL, keeping the build tied to the explicit UV_CONFIG_FILE. For a temporary Tsinghua mirror, place this in repository-root uv.toml:

[[index]]
name = "tsinghua"
url = "https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple/"
default = true

Diagnose slow builds:

Symptom Common cause Fix
Large transferring context build context includes unrelated frontend / data files .dockerignore ships only required files
uv sync --frozen is slow first build, cold cache, or missing uv mirror config wait for the first build; later runs reuse BuildKit cache; configure uv.toml when needed
Old keys still in effect after edit container not restarted ./planet.sh restart -a
Code changed but the image did not rebuild fingerprint still matches the image label Confirm the change is under aiprovider/, Dockerfile, or Python dependency inputs; delete planet-aiprovider:latest and retry if needed

SMTP Email (Required for Public Registration)

Public registration and email verification depend on SMTP. Administrators configure host, port, username, password, from-address, and TLS mode at /settings -> SMTP Email in the console, then use the "Send Test Email" button to verify. Settings are persisted in the system_settings.smtp row.

When SMTP is unset, POST /api/v1/auth/register returns 503 EMAIL_PROVIDER_NOT_CONFIGURED and the frontend surfaces a clear error. The operational fallback is ./planet.sh createuser.

One-time codes are stored in Redis under otp:{purpose}:{email} with a 600-second TTL. The key is invalidated after 5 invalid attempts. Resend cooldown is 60 seconds, enforced via otp_rate:{purpose}:{email}.

Troubleshooting Order

./planet.sh health        # 1. service state
./planet.sh log           # 2. recent logs
./planet.sh log -f        # 3. per-module logs
./planet.sh log -b
./planet.sh log -a
./planet.sh restart -f    # 4. restart only the affected module
./planet.sh restart -b
./planet.sh restart -a
./planet.sh restart -d    # 5. database / cache issues
./planet.sh restart       # 6. full restart if still broken

Development Command Conventions

Backend setup and script initialization use the lockfile:

uv python install 3.14
uv sync --frozen --group dev

--frozen rejects implicit uv.lock rewrites, which is the desired behavior on new machines, CI, and Docker builds. Dependency upgrades should explicitly update pyproject.toml / uv.lock on a development machine and commit the lockfile.

Frontend must use Bun:

cd frontend
bun install
bun run dev
bun run build

Do not use npm run .... In the WSL / Windows mixed environment Bun avoids Node/npm path inconsistencies.

./planet.sh start / init now runs bun install before startup instead of only checking whether the Vite entry file exists. This keeps new devices, cleaned node_modules, and lockfile changes synchronized before the console loads, avoiding dynamic-import 500s caused by missing frontend dependencies.

Validate the frontend build:

source ~/.zshrc && bun run build

Use bun run dev during development; Vite HMR refreshes the browser after source saves. bun run build only writes the dist artifact and does not refresh an already-open dev page.

To inspect the production bundle with automatic reload after successful builds:

bun run preview:auto

To watch sources and rebuild continuously without starting the preview server:

bun run build:watch

Backend dependencies are managed with uv:

uv sync
uv run pytest backend/tests/test_otp_service.py

Earth Boundary PMTiles Operations

  1. In the console, open Operations and Configuration -> Earth Content -> Boundary Precision to save boundary source configuration. The local config is written to config/earth-boundary-sources.local.json; do not commit it.
  2. Click "Build high precision boundaries", or switch the Earth toolbar settings gear to High Precision for the first build. The backend downloads the three source packages to data/earth-boundary-sources/, writes the source manifest, and invokes the PMTiles build script.
  3. The builder requires tippecanoe and pmtiles on PATH. Missing tools return a clear API error and do not write data-source collection records.
  4. A successful production build outputs frontend/public/earth/data/boundaries/earth-boundaries-china-pov-v1.pmtiles and its manifest.
  5. After deployment, open Earth, enable "Border Lines", and inspect China's southeast coast, Taiwan, Hainan, the South China Sea, Zangnan, Kosovo, and Gaza for hover behavior and boundary policy.
  6. If no high-precision manifest/PMTiles exists locally, Earth uses the bundled frontend/public/earth/data/countries-admin0.min.geojson fallback. If high-precision assets exist but tile requests fail, troubleshoot PMTiles range requests, manifest provider, Nginx .pmtiles static serving, and sha256 consistency.