6.8 KiB
Data Jobs and Outbox Architecture
This document records the technical boundary for Planet v1 data jobs, database outbox, and Earth refresh. For each data product's business purpose and end-to-end flow, see Business Architecture and Data Flows.
Architecture Boundary
- PostgreSQL is the durable v1 job ledger and outbox. Kafka, Celery, and RQ are intentionally not part of v1.
collection_tasksrecords collection, data clearing, cache clearing, and non-database Earth refresh jobs.earth_data_change_eventsrecords fact-table or derived-table changes and is the reliable source for Earth sync.LISTEN/NOTIFYis only the low-latency wakeup path; the listener still polls unconsumed outbox rows.- Redis is mainly cache, auth helper, OTP / rate limit, temporary logs, and WebSocket support. It is not the durable queue.
Database Change Sync
flowchart LR
Write["Fact or derived table write"] --> Trigger["PostgreSQL trigger"]
Trigger --> Outbox["earth_data_change_events"]
Trigger --> Notify["planet_earth_data_changes"]
Outbox --> Listener["earth_db_change_listener"]
Notify --> Listener
Listener --> Adapter["earth_layer_adapters"]
Adapter --> Cache["cache invalidation"]
Cache --> WS["earth_updates"]
- Collection, clearing, location resolution, or projection jobs write fact or derived tables.
- Statement triggers write outbox rows for
INSERT/UPDATE/DELETE. - The listener wakes through notify or finds pending rows through polling.
- The listener maps
table + sourceto Earth layers, refresh strategy, and cache patterns throughearth_layer_adapters.py. - The listener merges short-window same-layer events, invalidates cache, and broadcasts
earth_updates. - The outbox row is marked consumed only after successful broadcast; failed rows stay retryable.
DB changes no longer create default earth_refresh jobs, so they are not blocked by long same-source collection or clearing jobs. earth_refresh remains for manual cache clearing and non-DB refresh hints.
Data Job Queue
collection_tasks is the unified job ledger. Workers claim queued jobs with PostgreSQL FOR UPDATE SKIP LOCKED; write jobs for the same source run serially, while different sources may run in parallel.
| task_type | Purpose |
|---|---|
collect |
Run a built-in datasource collector |
clear_data |
Delete collected rows and declared derived rows for the source |
clear_cache |
Delete Earth / dashboard cache for the source |
earth_refresh |
Invalidate Earth layer cache and broadcast a refresh hint for non-DB changes |
API handlers only create jobs and return task_id. Execution, progress, cancellation, and terminal state are written back by workers and pushed to the frontend through the datasource_tasks channel.
Cancellation means “keep committed batches”: clicking stop marks the job as cancelling and cancels the in-memory coroutine. Already committed batches remain; unfinished batches follow the collector or cleanup rollback path.
Earth Sync Event Model
The unified event model is earth.layer.changed:
{
"event": "earth.layer.changed",
"action": "database_changed",
"source": "celestrak_tle",
"table": "collected_data",
"operation": "DELETE",
"layers": ["satellites"],
"refresh_strategy": "clear_then_reload",
"records_processed": 11125,
"occurred_at": "2026-05-25T10:20:30Z"
}
| strategy | Purpose |
|---|---|
clear_then_reload |
Clear local frontend layer objects first, then force a refetch. Prefer this for deletes. |
reload |
Keep old objects until fresh data returns. Use it for location, metadata, or non-destructive updates. |
delta |
Used only for earth_interactables; upsert or remove objects by id. |
APIs must return HTTP 200 with an empty collection for real zero-data states; 5xx is reserved for real endpoint failures. After a delete event, if refetch fails, the frontend should keep the cleared state and show a lightweight error instead of restoring stale objects.
Layer Adapter Contract
earth_layer_adapters.py is the single registry for sources, derived tables, Earth layers, cache patterns, and refresh strategy. New layers should be added through an adapter entry, not through one-off button handlers, collector branches, or frontend special cases.
Each adapter must declare:
- Which source or table feeds which Earth layer.
- Which Earth cache key patterns must be invalidated.
- Which owned derived tables must be removed during
clear_data. - The default refresh strategy for that layer.
When a source is cleared, the clear_data job first deletes collected_data.source = <source>, then deletes adapter-owned derived rows. Direct derived-table edits also trigger the outbox, so background jobs, admin APIs, and SQL repair scripts reach Earth as long as they mutate fact or derived tables.
Operations and Troubleshooting
Check whether outbox rows are piling up:
SELECT id, table_name, operation, source, occurred_at
FROM earth_data_change_events
WHERE consumed_at IS NULL
ORDER BY id
LIMIT 20;
Check that triggers exist:
SELECT tgname, tgrelid::regclass
FROM pg_trigger
WHERE tgname LIKE 'tr_planet_%_changed_%'
ORDER BY 2, 1;
Useful log events:
earth.db_changes.connected: the listener connected to PostgreSQL and started listening.earth.db_changes.outbox_polled: polling found unconsumed outbox rows.earth.db_changes.broadcasted: an Earth refresh broadcast was produced.data_job.started/data_job.completed: job execution state.
If Earth does not update, check in order: fact table changed, outbox was consumed, adapter covers the source/table, the listener is online, frontend WebSocket is connected, and the visualization API returns HTTP 200 with either an empty collection or fresh data.
Kafka-ready Boundaries
Business code avoids depending on a concrete queue implementation by preserving these boundaries:
JobQueue: submit, claim, cancel, and complete data jobs.DataChangeBus: publish database fact changes.EarthLayerAdapterRegistry: declare source, layer, cache, and derived-data relationships.
Kafka becomes appropriate when:
- Several independent services must consume the same data-change stream.
- AIS, BGP, or sensor streams become sustained high-throughput inputs.
- Consumer groups, replay, and service decoupling are required.
Spark becomes appropriate when:
- Historical data reaches tens or hundreds of millions of rows and PostgreSQL aggregation becomes expensive.
- Cross-source, long-window, spatiotemporal analysis is needed.
- Raw data lands in Parquet / Iceberg / Delta and the system starts producing offline derived data products.
For second-level continuous stream processing, evaluate Flink first. Spark is a better fit for batch or micro-batch analytics.