Files
planet/docs/plans/agents-agent-runtime-roadmap.md
2026-04-21 22:49:39 +08:00

6.5 KiB

Agent Runtime Roadmap

Overview

This document connects three existing planning threads into one implementation roadmap:

  • aiprovider as the model gateway
  • datasource health governance as the first practical agent use case
  • situational awareness as the broader long-term target

Related documents:

Big Picture

Planet should evolve in layers:

  1. stable model gateway
  2. deterministic health and evidence collection
  3. agent runtime for reasoning and proposal generation
  4. situational-awareness assessments and controlled actions

This prevents the system from collapsing into a single giant "AI feature" with unclear boundaries.

Architecture Overview

flowchart TD
    U["Frontend / Backend APIs / Operators"] --> B["Planet Backend"]
    B --> H["Datasource Health Services"]
    B --> R["Agent Runtime"]
    R --> P["aiprovider"]
    P --> M["OpenAI / Anthropic / MiniMax / Ollama / Local Models"]

    C["Collectors / Snapshots / Logs / Alerts / BGP Signals"] --> S["Signal Store"]
    H --> S
    S --> E["Evaluation Layer"]
    E --> F["Findings"]
    F --> R

    W["Web Search / Page Fetch / Docs Fetch"] --> R
    R --> PR["Proposals"]
    R --> AS["Assessments"]

    PR --> O["Runtime Overrides / Review Queue / Tasks"]
    AS --> SA["Situational Awareness APIs / UI"]

    O --> V["Verification Loop"]
    V --> S

Role Boundaries

aiprovider

Responsibilities:

  • provider compatibility
  • protocol adaptation
  • auth and model transport
  • request/response normalization

Not responsible for:

  • agent orchestration
  • business workflows
  • datasource repair policy
  • situational-awareness domain logic

Backend

Responsibilities:

  • stable business APIs
  • auth and permissions
  • task orchestration
  • health records
  • proposal and override persistence
  • assessment exposure

Agent Runtime

Responsibilities:

  • consume findings and context
  • invoke LLMs via aiprovider
  • invoke tools such as web search
  • create proposals
  • create assessments
  • route to policy-controlled action paths

Delivery Sequence

Stage 1: Gateway Foundation

Status:

  • already in place

Delivered by current work:

  • aiprovider
  • multi-provider compatibility
  • backend AI facade
  • MiniMax / Anthropic-compatible support
  • request-id propagation

Primary outcome:

  • the system already has a stable way to call models

Stage 2: Datasource Health MVP

Goal:

  • establish deterministic health observability

Key work:

  • health check task runner
  • health result table
  • datasource health APIs
  • UI visibility
  • collector endpoint override precedence cleanup

Primary outcome:

  • Planet knows which collectors are healthy before asking an LLM anything

Stage 3: Health Agent

Goal:

  • let the first agent role operate on health failures

Key work:

  • convert health failures into signals/findings
  • invoke agent only for failed or suspicious cases
  • produce repair proposals with evidence and confidence

Primary outcome:

  • Planet can suggest endpoint repairs without mutating defaults

Stage 4: Runtime Repair Application

Goal:

  • safely apply approved datasource repair proposals

Key work:

  • override storage
  • policy-gated apply flow
  • verification after apply
  • rollback path

Primary outcome:

  • datasource repair becomes operationally useful without polluting repository defaults

Stage 5: Situational Awareness Assessments

Goal:

  • reuse the same runtime for broader operator-facing assessment

Key work:

  • normalize telemetry and incident evidence into signals/findings
  • build Assessment Agent
  • expose structured assessments through backend APIs and UI

Primary outcome:

  • LLM output becomes evidence-backed situational summary, not just ad hoc chat output

Stage 6: Correlation and Controlled Actions

Goal:

  • connect multiple sources into higher-level posture and event groupings

Key work:

  • event correlation
  • incident grouping
  • recommendation scoring
  • controlled action routing

Primary outcome:

  • Planet becomes a true agent-assisted situational-awareness system

Implementation Tracks

These tracks can progress in parallel, but they should stay loosely coupled.

Track A: Config and Runtime Resolution

Scope:

  • datasource defaults
  • overrides
  • runtime precedence
  • audit trails

First milestone:

  • health-safe override layer

Track B: Health and Evidence

Scope:

  • deterministic checks
  • failure categorization
  • signal and finding persistence

First milestone:

  • datasource health record system

Track C: Agent Runtime

Scope:

  • shared object model
  • orchestration flow
  • prompt/tool pipeline
  • policy integration

First milestone:

  • Health Agent proposal pipeline

Track D: Situational Awareness

Scope:

  • assessment schema
  • multi-source context assembly
  • operator-facing outputs

First milestone:

  • structured assessment API

Shared Artifacts

To avoid fragmentation, these artifacts should be shared across all future agent work.

Shared object model

  • Signal
  • Finding
  • Proposal
  • Assessment

Shared orchestration flow

  • collect
  • validate
  • classify
  • reason
  • propose or assess
  • review or apply
  • verify
  • archive

Shared policy model

  • read-only
  • propose-only
  • apply-limited
  1. Build Stage 2 first
  • datasource health records
  • deterministic checks
  • no automatic repair
  1. Then build Stage 3
  • Health Agent
  • proposal generation only
  1. Then Stage 4
  • override apply flow
  • rollback and verification
  1. Only after that start Stage 5
  • broader situational-awareness assessment workflows

Why This Order

Because situational-awareness quality depends on reliable upstream data.

If datasource health is weak:

  • agent reasoning quality will degrade
  • false explanations will increase
  • assessment trust will drop

So datasource health is not a side task.

It is the first operational foundation for the later situational-awareness system.

Summary

Planet should be built as:

  • aiprovider for model access
  • backend services for orchestration and persistence
  • datasource health as the first evidence-governance layer
  • agent runtime as the reusable reasoning core
  • situational awareness as the long-term application layer

That path keeps the architecture coherent and lets each phase produce useful functionality without forcing a rewrite later.