3.7 KiB
DataSources List API Performance Optimization
Background
GET /api/v1/datasources is the core API for the Data Sources page. Slow responses directly block page rendering.
Query Path Before Optimization
_load_datasource_list_context used to run these queries sequentially:
| Order | Function | Query | Bottleneck |
|---|---|---|---|
| 1 | _load_latest_running_tasks |
collection_tasks window query; stale check depends on this result |
Must be serial |
| 2 | _load_latest_completed_tasks |
collection_tasks window query for latest completed tasks |
Serial wait |
| 3 | _load_datasource_data_counts |
COUNT(*) GROUP BY source on collected_data |
Slow full-table scan |
| 4 | _load_datasource_endpoint_overrides |
Simple datasource_configs SELECT |
Serial wait |
Phase 1: Parallelization
The independent queries 2, 3, and 4 were moved to asyncio.gather with separate sessions:
async def _fetch_completed():
async with async_session_factory() as s:
return await _load_latest_completed_tasks(s, datasource_ids)
async def _fetch_counts():
async with async_session_factory() as s:
return await _load_datasource_data_counts(s, sources)
async def _fetch_overrides():
async with async_session_factory() as s:
return await _load_datasource_endpoint_overrides(s, sources)
completed_tasks, data_counts, endpoint_overrides = await asyncio.gather(
_fetch_completed(), _fetch_counts(), _fetch_overrides(),
)
SQLAlchemy AsyncSession does not support concurrent use from multiple coroutines, so every parallel branch needs its own session.
Phase 2: Remove Heavy Queries
Remove _load_datasource_data_counts
data_count was only used by the frontend to show an edge-case (0 records) hint in the latest collection column. It was not worth keeping a COUNT(*) GROUP BY full-table scan.
- Frontend
(0 records)display logic was removed. data_countwas removed from theBuiltInDataSourceinterface.
Remove _load_latest_completed_tasks
last_status and last_run_at are already written to the DataSource model when collectors finish, so the list endpoint no longer needs to join collection_tasks:
# Before: completed_tasks query required
last_run_at = datasource.last_run_at or (last_task.completed_at if last_task else None)
last_status = datasource.last_status or (last_task.status if last_task else None)
# After: read model fields directly
last_run_at = datasource.last_run_at
last_status = datasource.last_status
last_records_processed was removed as well because it came from completed task rows and is not displayed in the list.
Query Path After Optimization
datasources SELECT -> required primary data
_load_latest_running_tasks -> required for running state and stale check
_load_datasource_endpoint_overrides -> required for endpoint overrides and collector settings display
The endpoint now runs three queries instead of five. The last two run sequentially because running tasks are needed for stale checks and endpoint overrides are lightweight.
Frontend triggerDatasource Double Refresh Fix
triggerDatasource previously called fetchData() twice:
// Before
} else {
window.setTimeout(() => { fetchData() }, 800)
}
fetchData()
// After: mutually exclusive
if (res.data.task_id) {
fetchData()
} else {
window.setTimeout(fetchData, 800)
}
Related Files
- datasources.py:
_load_datasource_list_context,list_datasources - DataSources.tsx:
BuiltInDataSource,triggerDatasource