Files
planet/docs/plans/agents-datasource-health-plan.md
2026-04-21 22:49:39 +08:00

10 KiB

Datasource Health Plan

Overview

This document defines a phased plan for datasource health governance.

The goal is to make collectors observable, diagnosable, and recoverable when upstream APIs change, while avoiding unsafe automatic mutation of repository defaults.

The key principle is:

  • do not let runtime automation rewrite repository default config

Instead, split responsibilities across:

  • default config
  • runtime overrides
  • health check records
  • agent-generated repair proposals

Problem Statement

Collectors currently depend on third-party APIs, data downloads, mirrored JSON files, archive links, and web pages.

These upstream dependencies can fail in several ways:

  • endpoint becomes unreachable
  • endpoint still responds but schema changes
  • content-type changes
  • website shuts down or moves
  • mirror link disappears
  • HTML structure changes and scraping fails
  • endpoint requires a new path or new host

We want a system that can:

  • detect datasource health degradation early
  • identify likely cause
  • search for updated endpoints when reasonable
  • apply safe runtime fixes without polluting default repo config
  • preserve auditability and rollback

Design Principles

  1. Default config is stable
  • backend/app/core/data_sources.yaml remains the repository baseline.
  • It should be changed intentionally through normal development flow, not by autonomous runtime agents.
  1. Runtime fixes are isolated
  • Emergency or adaptive fixes should live in a runtime override layer.
  • Overrides should be reversible and auditable.
  1. Deterministic checks come first
  • Use normal programmatic health checks before using LLMs.
  • Only call an agent when deterministic checks indicate a meaningful failure.
  1. Agents suggest before they mutate
  • Agents should produce proposals with evidence and confidence.
  • Application of a proposal should be controlled by policy.
  1. Every repair is attributable
  • Store what changed, why, who or what suggested it, and when it was applied.

Configuration Layers

Recommended runtime precedence:

  1. datasource endpoint override
  2. datasource DB endpoint override
  3. repository default YAML
  4. collector internal fallback logic

Definitions:

  • repository default YAML:
    • backend/app/core/data_sources.yaml
    • versioned baseline
  • datasource DB endpoint override:
    • existing DataSourceConfig.endpoint
    • current runtime override entrypoint
  • datasource endpoint override:
    • a dedicated new override table
    • used for health-repair and proposal application
  • collector internal fallback logic:
    • final defensive fallback
    • should be minimized over time

1. Deterministic Health Checks

Each collector gets a health profile with checks such as:

  • endpoint resolves
  • HTTP request succeeds
  • status code is acceptable
  • content-type is expected
  • body parses successfully
  • minimum structural fields exist
  • sample item count is plausible
  • latency is within threshold

Output states:

  • healthy
  • degraded
  • failed
  • schema_changed
  • rate_limited
  • auth_required

2. Agent-Assisted Repair Discovery

Only triggered when deterministic health checks fail or return suspicious structure.

Agent responsibilities:

  • search for current official endpoint or replacement path
  • inspect likely upstream documentation or landing pages
  • compare candidate endpoint output to collector expectations
  • produce a repair proposal with confidence and evidence

Agent should not directly modify repository defaults.

3. Safe Runtime Repair Application

Repair proposals can be:

  • reviewed manually
  • auto-applied only under strict low-risk policy

Auto-apply should be limited to cases like:

  • same trusted domain
  • highly similar response structure
  • repeated successful verification
  • confidence above threshold

Phased Delivery Plan

Phase 1: Deterministic Health MVP

Goal:

  • build health observability without automated repair

Scope:

  • datasource health check task runner
  • datasource health result persistence
  • endpoint reachability + parse checks
  • dashboard or API visibility into health status

Deliverables:

  • health check service
  • health check record table
  • status endpoint
  • scheduled or manual check trigger

No agent usage yet.

Phase 2: Agent Repair Proposals

Goal:

  • let agent investigate failing sources and propose updated endpoints

Scope:

  • invoke agent only when datasource health is failed or schema_changed
  • web search + page inspection
  • candidate endpoint extraction
  • proposal persistence

Deliverables:

  • repair proposal schema
  • proposal generation pipeline
  • confidence and evidence model
  • operator review view or API

Still no automatic config mutation.

Phase 3: Runtime Overrides

Goal:

  • allow approved proposals to take effect safely at runtime

Scope:

  • add dedicated override storage
  • runtime resolution prefers override over default config
  • proposal application writes override only

Deliverables:

  • endpoint override table
  • override-aware resolution logic
  • apply/reject endpoints
  • rollback endpoint

Repository default YAML remains untouched.

Phase 4: Limited Auto-Apply

Goal:

  • safely automate a narrow slice of low-risk repairs

Scope:

  • policy engine for auto-apply
  • same-domain or trusted-domain checks
  • structure validation
  • staged verification after apply

Deliverables:

  • auto-apply rules
  • audit logs
  • automatic post-apply health verification
  • auto-disable or rollback on regression

Data Model Draft

datasource_health_checks

Purpose:

  • store each health evaluation result

Suggested fields:

  • id
  • datasource_id
  • collector_name
  • endpoint_checked
  • status
  • http_status
  • content_type
  • latency_ms
  • sample_count
  • error_message
  • details
  • checked_at

details can store structured diagnostic data such as:

  • parsed fields
  • schema mismatch summary
  • retry count
  • exception class

datasource_repair_proposals

Purpose:

  • store agent-generated repair suggestions

Suggested fields:

  • id
  • datasource_id
  • collector_name
  • old_endpoint
  • candidate_endpoint
  • reason
  • confidence
  • evidence_urls
  • evidence_summary
  • status
  • created_by
  • created_at
  • reviewed_at

Suggested status values:

  • proposed
  • approved
  • rejected
  • applied
  • expired

datasource_endpoint_overrides

Purpose:

  • runtime endpoint override layer

Suggested fields:

  • id
  • datasource_id
  • collector_name
  • endpoint
  • reason
  • source
  • proposal_id
  • enabled
  • created_at
  • updated_at

Suggested source values:

  • manual
  • health-agent
  • migration

API Draft

Health

  • GET /api/v1/datasources/health
  • GET /api/v1/datasources/{id}/health
  • POST /api/v1/datasources/{id}/health-check
  • POST /api/v1/datasources/health-check-all

Repair proposals

  • GET /api/v1/datasources/{id}/repair-proposals
  • POST /api/v1/datasources/{id}/repair-proposals/generate
  • POST /api/v1/datasources/{id}/repair-proposals/{proposal_id}/approve
  • POST /api/v1/datasources/{id}/repair-proposals/{proposal_id}/reject
  • POST /api/v1/datasources/{id}/repair-proposals/{proposal_id}/apply

Overrides

  • GET /api/v1/datasources/{id}/overrides
  • POST /api/v1/datasources/{id}/overrides
  • PUT /api/v1/datasources/{id}/overrides/{override_id}
  • DELETE /api/v1/datasources/{id}/overrides/{override_id}

Agent Contract Draft

When deterministic health fails, the agent should receive:

  • datasource name
  • collector name
  • current endpoint
  • current failure mode
  • expected response shape summary
  • known trusted domains

Expected output:

{
  "status": "proposal",
  "candidate_endpoint": "https://example.com/api/v2/data",
  "confidence": 0.86,
  "reason": "Official docs now point to v2 endpoint",
  "evidence_urls": [
    "https://example.com/docs/api",
    "https://example.com/changelog"
  ],
  "notes": "Response shape appears compatible after light field remapping"
}

The agent should never output "rewrite the default yaml" as its primary action.

Risk Analysis

Risk: wrong endpoint chosen by agent

Mitigation:

  • use trusted-domain allowlists
  • require evidence URLs
  • require confidence threshold
  • add manual review for medium-risk sources

Risk: endpoint responds but schema silently changed

Mitigation:

  • deterministic schema checks
  • parse and sample validation
  • content-type checks
  • collector-specific required fields

Risk: automatic runtime override causes hidden drift

Mitigation:

  • store all overrides explicitly
  • mark source of override
  • keep default YAML unchanged
  • expose active overrides in API/UI

Risk: persistent bad override breaks data collection

Mitigation:

  • allow rollback
  • keep parent/default endpoint visible
  • re-run verification after apply
  • auto-disable override on repeated failure

Operational Policy Recommendations

  1. Do not auto-apply for high-value or high-fragility sources initially.

  2. Use manual approval for:

  • scraped HTML sources
  • unofficial mirrors
  • sources with auth or rate-limit complexity
  • sources with legal or trust ambiguity
  1. Allow auto-apply only for:
  • same-domain version bumps
  • obvious official migration paths
  • repeated passing verification
  1. Expose health + proposal + override state together in one operator view.

Suggested Implementation Order

  1. Phase 1
  • health result table
  • deterministic checks
  • API and UI visibility
  1. Phase 2
  • proposal table
  • agent prompt/output contract
  • proposal generation job
  1. Phase 3
  • runtime override table
  • resolver precedence update
  • apply/reject endpoints
  1. Phase 4
  • auto-apply rules
  • rollback policy
  • operator automation

Out Of Scope For The First Iteration

  • direct automatic mutation of repository default YAML
  • automatic git commits by repair agents
  • unrestricted autonomous endpoint replacement
  • fully generalized schema remapping engine

The first milestone should be:

  • deterministic datasource health checks
  • persisted results
  • manual visibility
  • no automatic repair

This gives immediate operational value with low risk, and prepares clean inputs for the later agent phase.