Files
planet/docs/technical/en/backend-datasources-api-performance.md
linkong e1984c7a35 release: bump version to 0.49.0
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-08 17:42:27 +08:00

3.7 KiB

DataSources List API Performance Optimization

Background

GET /api/v1/datasources is the core API for the Data Sources page. Slow responses directly block page rendering.

Query Path Before Optimization

_load_datasource_list_context used to run these queries sequentially:

Order Function Query Bottleneck
1 _load_latest_running_tasks collection_tasks window query; stale check depends on this result Must be serial
2 _load_latest_completed_tasks collection_tasks window query for latest completed tasks Serial wait
3 _load_datasource_data_counts COUNT(*) GROUP BY source on collected_data Slow full-table scan
4 _load_datasource_endpoint_overrides Simple datasource_configs SELECT Serial wait

Phase 1: Parallelization

The independent queries 2, 3, and 4 were moved to asyncio.gather with separate sessions:

async def _fetch_completed():
    async with async_session_factory() as s:
        return await _load_latest_completed_tasks(s, datasource_ids)

async def _fetch_counts():
    async with async_session_factory() as s:
        return await _load_datasource_data_counts(s, sources)

async def _fetch_overrides():
    async with async_session_factory() as s:
        return await _load_datasource_endpoint_overrides(s, sources)

completed_tasks, data_counts, endpoint_overrides = await asyncio.gather(
    _fetch_completed(), _fetch_counts(), _fetch_overrides(),
)

SQLAlchemy AsyncSession does not support concurrent use from multiple coroutines, so every parallel branch needs its own session.

Phase 2: Remove Heavy Queries

Remove _load_datasource_data_counts

data_count was only used by the frontend to show an edge-case (0 records) hint in the latest collection column. It was not worth keeping a COUNT(*) GROUP BY full-table scan.

  • Frontend (0 records) display logic was removed.
  • data_count was removed from the BuiltInDataSource interface.

Remove _load_latest_completed_tasks

last_status and last_run_at are already written to the DataSource model when collectors finish, so the list endpoint no longer needs to join collection_tasks:

# Before: completed_tasks query required
last_run_at = datasource.last_run_at or (last_task.completed_at if last_task else None)
last_status = datasource.last_status or (last_task.status if last_task else None)

# After: read model fields directly
last_run_at = datasource.last_run_at
last_status = datasource.last_status

last_records_processed was removed as well because it came from completed task rows and is not displayed in the list.

Query Path After Optimization

datasources SELECT                     -> required primary data
_load_latest_running_tasks             -> required for running state and stale check
_load_datasource_endpoint_overrides    -> required for endpoint overrides and collector settings display

The endpoint now runs three queries instead of five. The last two run sequentially because running tasks are needed for stale checks and endpoint overrides are lightweight.

Frontend triggerDatasource Double Refresh Fix

triggerDatasource previously called fetchData() twice:

// Before
} else {
    window.setTimeout(() => { fetchData() }, 800)
}
fetchData()

// After: mutually exclusive
if (res.data.task_id) {
    fetchData()
} else {
    window.setTimeout(fetchData, 800)
}