# Agent Runtime Roadmap ## Overview This document connects three existing planning threads into one implementation roadmap: - `aiprovider` as the model gateway - datasource health governance as the first practical agent use case - situational awareness as the broader long-term target Related documents: - [aiprovider](/home/ray/dev/linkong/planet/docs/aiprovider.md) - [datasource-health-plan](/home/ray/dev/linkong/planet/docs/datasource-health-plan.md) - [agent-architecture-plan](/home/ray/dev/linkong/planet/docs/agent-architecture-plan.md) ## Big Picture Planet should evolve in layers: 1. stable model gateway 2. deterministic health and evidence collection 3. agent runtime for reasoning and proposal generation 4. situational-awareness assessments and controlled actions This prevents the system from collapsing into a single giant "AI feature" with unclear boundaries. ## Architecture Overview ```mermaid flowchart TD U["Frontend / Backend APIs / Operators"] --> B["Planet Backend"] B --> H["Datasource Health Services"] B --> R["Agent Runtime"] R --> P["aiprovider"] P --> M["OpenAI / Anthropic / MiniMax / Ollama / Local Models"] C["Collectors / Snapshots / Logs / Alerts / BGP Signals"] --> S["Signal Store"] H --> S S --> E["Evaluation Layer"] E --> F["Findings"] F --> R W["Web Search / Page Fetch / Docs Fetch"] --> R R --> PR["Proposals"] R --> AS["Assessments"] PR --> O["Runtime Overrides / Review Queue / Tasks"] AS --> SA["Situational Awareness APIs / UI"] O --> V["Verification Loop"] V --> S ``` ## Role Boundaries ### `aiprovider` Responsibilities: - provider compatibility - protocol adaptation - auth and model transport - request/response normalization Not responsible for: - agent orchestration - business workflows - datasource repair policy - situational-awareness domain logic ### Backend Responsibilities: - stable business APIs - auth and permissions - task orchestration - health records - proposal and override persistence - assessment exposure ### Agent Runtime Responsibilities: - consume findings and context - invoke LLMs via `aiprovider` - invoke tools such as web search - create proposals - create assessments - route to policy-controlled action paths ## Delivery Sequence ## Stage 1: Gateway Foundation Status: - already in place Delivered by current work: - `aiprovider` - multi-provider compatibility - backend AI facade - MiniMax / Anthropic-compatible support - request-id propagation Primary outcome: - the system already has a stable way to call models ## Stage 2: Datasource Health MVP Goal: - establish deterministic health observability Key work: - health check task runner - health result table - datasource health APIs - UI visibility - collector endpoint override precedence cleanup Primary outcome: - Planet knows which collectors are healthy before asking an LLM anything ## Stage 3: Health Agent Goal: - let the first agent role operate on health failures Key work: - convert health failures into signals/findings - invoke agent only for failed or suspicious cases - produce repair proposals with evidence and confidence Primary outcome: - Planet can suggest endpoint repairs without mutating defaults ## Stage 4: Runtime Repair Application Goal: - safely apply approved datasource repair proposals Key work: - override storage - policy-gated apply flow - verification after apply - rollback path Primary outcome: - datasource repair becomes operationally useful without polluting repository defaults ## Stage 5: Situational Awareness Assessments Goal: - reuse the same runtime for broader operator-facing assessment Key work: - normalize telemetry and incident evidence into signals/findings - build Assessment Agent - expose structured assessments through backend APIs and UI Primary outcome: - LLM output becomes evidence-backed situational summary, not just ad hoc chat output ## Stage 6: Correlation and Controlled Actions Goal: - connect multiple sources into higher-level posture and event groupings Key work: - event correlation - incident grouping - recommendation scoring - controlled action routing Primary outcome: - Planet becomes a true agent-assisted situational-awareness system ## Implementation Tracks These tracks can progress in parallel, but they should stay loosely coupled. ### Track A: Config and Runtime Resolution Scope: - datasource defaults - overrides - runtime precedence - audit trails First milestone: - health-safe override layer ### Track B: Health and Evidence Scope: - deterministic checks - failure categorization - signal and finding persistence First milestone: - datasource health record system ### Track C: Agent Runtime Scope: - shared object model - orchestration flow - prompt/tool pipeline - policy integration First milestone: - Health Agent proposal pipeline ### Track D: Situational Awareness Scope: - assessment schema - multi-source context assembly - operator-facing outputs First milestone: - structured assessment API ## Shared Artifacts To avoid fragmentation, these artifacts should be shared across all future agent work. ### Shared object model - `Signal` - `Finding` - `Proposal` - `Assessment` ### Shared orchestration flow - collect - validate - classify - reason - propose or assess - review or apply - verify - archive ### Shared policy model - read-only - propose-only - apply-limited ## Recommended Next Concrete Steps 1. Build Stage 2 first - datasource health records - deterministic checks - no automatic repair 2. Then build Stage 3 - Health Agent - proposal generation only 3. Then Stage 4 - override apply flow - rollback and verification 4. Only after that start Stage 5 - broader situational-awareness assessment workflows ## Why This Order Because situational-awareness quality depends on reliable upstream data. If datasource health is weak: - agent reasoning quality will degrade - false explanations will increase - assessment trust will drop So datasource health is not a side task. It is the first operational foundation for the later situational-awareness system. ## Summary Planet should be built as: - `aiprovider` for model access - backend services for orchestration and persistence - datasource health as the first evidence-governance layer - agent runtime as the reusable reasoning core - situational awareness as the long-term application layer That path keeps the architecture coherent and lets each phase produce useful functionality without forcing a rewrite later.