347 lines
6.5 KiB
Markdown
347 lines
6.5 KiB
Markdown
# Agent Runtime Roadmap
|
|
|
|
## Overview
|
|
|
|
This document connects three existing planning threads into one implementation roadmap:
|
|
|
|
- `aiprovider` as the model gateway
|
|
- datasource health governance as the first practical agent use case
|
|
- situational awareness as the broader long-term target
|
|
|
|
Related documents:
|
|
|
|
- [aiprovider](/home/ray/dev/linkong/planet/docs/agents/aiprovider.md)
|
|
- [datasource-health-plan](/home/ray/dev/linkong/planet/docs/agents/datasource-health-plan.md)
|
|
- [agent-architecture-plan](/home/ray/dev/linkong/planet/docs/agents/agent-architecture-plan.md)
|
|
|
|
|
|
## Big Picture
|
|
|
|
Planet should evolve in layers:
|
|
|
|
1. stable model gateway
|
|
2. deterministic health and evidence collection
|
|
3. agent runtime for reasoning and proposal generation
|
|
4. situational-awareness assessments and controlled actions
|
|
|
|
This prevents the system from collapsing into a single giant "AI feature" with unclear boundaries.
|
|
|
|
|
|
## Architecture Overview
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
U["Frontend / Backend APIs / Operators"] --> B["Planet Backend"]
|
|
B --> H["Datasource Health Services"]
|
|
B --> R["Agent Runtime"]
|
|
R --> P["aiprovider"]
|
|
P --> M["OpenAI / Anthropic / MiniMax / Ollama / Local Models"]
|
|
|
|
C["Collectors / Snapshots / Logs / Alerts / BGP Signals"] --> S["Signal Store"]
|
|
H --> S
|
|
S --> E["Evaluation Layer"]
|
|
E --> F["Findings"]
|
|
F --> R
|
|
|
|
W["Web Search / Page Fetch / Docs Fetch"] --> R
|
|
R --> PR["Proposals"]
|
|
R --> AS["Assessments"]
|
|
|
|
PR --> O["Runtime Overrides / Review Queue / Tasks"]
|
|
AS --> SA["Situational Awareness APIs / UI"]
|
|
|
|
O --> V["Verification Loop"]
|
|
V --> S
|
|
```
|
|
|
|
|
|
## Role Boundaries
|
|
|
|
### `aiprovider`
|
|
|
|
Responsibilities:
|
|
|
|
- provider compatibility
|
|
- protocol adaptation
|
|
- auth and model transport
|
|
- request/response normalization
|
|
|
|
Not responsible for:
|
|
|
|
- agent orchestration
|
|
- business workflows
|
|
- datasource repair policy
|
|
- situational-awareness domain logic
|
|
|
|
|
|
### Backend
|
|
|
|
Responsibilities:
|
|
|
|
- stable business APIs
|
|
- auth and permissions
|
|
- task orchestration
|
|
- health records
|
|
- proposal and override persistence
|
|
- assessment exposure
|
|
|
|
|
|
### Agent Runtime
|
|
|
|
Responsibilities:
|
|
|
|
- consume findings and context
|
|
- invoke LLMs via `aiprovider`
|
|
- invoke tools such as web search
|
|
- create proposals
|
|
- create assessments
|
|
- route to policy-controlled action paths
|
|
|
|
|
|
## Delivery Sequence
|
|
|
|
## Stage 1: Gateway Foundation
|
|
|
|
Status:
|
|
|
|
- already in place
|
|
|
|
Delivered by current work:
|
|
|
|
- `aiprovider`
|
|
- multi-provider compatibility
|
|
- backend AI facade
|
|
- MiniMax / Anthropic-compatible support
|
|
- request-id propagation
|
|
|
|
Primary outcome:
|
|
|
|
- the system already has a stable way to call models
|
|
|
|
|
|
## Stage 2: Datasource Health MVP
|
|
|
|
Goal:
|
|
|
|
- establish deterministic health observability
|
|
|
|
Key work:
|
|
|
|
- health check task runner
|
|
- health result table
|
|
- datasource health APIs
|
|
- UI visibility
|
|
- collector endpoint override precedence cleanup
|
|
|
|
Primary outcome:
|
|
|
|
- Planet knows which collectors are healthy before asking an LLM anything
|
|
|
|
|
|
## Stage 3: Health Agent
|
|
|
|
Goal:
|
|
|
|
- let the first agent role operate on health failures
|
|
|
|
Key work:
|
|
|
|
- convert health failures into signals/findings
|
|
- invoke agent only for failed or suspicious cases
|
|
- produce repair proposals with evidence and confidence
|
|
|
|
Primary outcome:
|
|
|
|
- Planet can suggest endpoint repairs without mutating defaults
|
|
|
|
|
|
## Stage 4: Runtime Repair Application
|
|
|
|
Goal:
|
|
|
|
- safely apply approved datasource repair proposals
|
|
|
|
Key work:
|
|
|
|
- override storage
|
|
- policy-gated apply flow
|
|
- verification after apply
|
|
- rollback path
|
|
|
|
Primary outcome:
|
|
|
|
- datasource repair becomes operationally useful without polluting repository defaults
|
|
|
|
|
|
## Stage 5: Situational Awareness Assessments
|
|
|
|
Goal:
|
|
|
|
- reuse the same runtime for broader operator-facing assessment
|
|
|
|
Key work:
|
|
|
|
- normalize telemetry and incident evidence into signals/findings
|
|
- build Assessment Agent
|
|
- expose structured assessments through backend APIs and UI
|
|
|
|
Primary outcome:
|
|
|
|
- LLM output becomes evidence-backed situational summary, not just ad hoc chat output
|
|
|
|
|
|
## Stage 6: Correlation and Controlled Actions
|
|
|
|
Goal:
|
|
|
|
- connect multiple sources into higher-level posture and event groupings
|
|
|
|
Key work:
|
|
|
|
- event correlation
|
|
- incident grouping
|
|
- recommendation scoring
|
|
- controlled action routing
|
|
|
|
Primary outcome:
|
|
|
|
- Planet becomes a true agent-assisted situational-awareness system
|
|
|
|
|
|
## Implementation Tracks
|
|
|
|
These tracks can progress in parallel, but they should stay loosely coupled.
|
|
|
|
### Track A: Config and Runtime Resolution
|
|
|
|
Scope:
|
|
|
|
- datasource defaults
|
|
- overrides
|
|
- runtime precedence
|
|
- audit trails
|
|
|
|
First milestone:
|
|
|
|
- health-safe override layer
|
|
|
|
|
|
### Track B: Health and Evidence
|
|
|
|
Scope:
|
|
|
|
- deterministic checks
|
|
- failure categorization
|
|
- signal and finding persistence
|
|
|
|
First milestone:
|
|
|
|
- datasource health record system
|
|
|
|
|
|
### Track C: Agent Runtime
|
|
|
|
Scope:
|
|
|
|
- shared object model
|
|
- orchestration flow
|
|
- prompt/tool pipeline
|
|
- policy integration
|
|
|
|
First milestone:
|
|
|
|
- Health Agent proposal pipeline
|
|
|
|
|
|
### Track D: Situational Awareness
|
|
|
|
Scope:
|
|
|
|
- assessment schema
|
|
- multi-source context assembly
|
|
- operator-facing outputs
|
|
|
|
First milestone:
|
|
|
|
- structured assessment API
|
|
|
|
|
|
## Shared Artifacts
|
|
|
|
To avoid fragmentation, these artifacts should be shared across all future agent work.
|
|
|
|
### Shared object model
|
|
|
|
- `Signal`
|
|
- `Finding`
|
|
- `Proposal`
|
|
- `Assessment`
|
|
|
|
### Shared orchestration flow
|
|
|
|
- collect
|
|
- validate
|
|
- classify
|
|
- reason
|
|
- propose or assess
|
|
- review or apply
|
|
- verify
|
|
- archive
|
|
|
|
### Shared policy model
|
|
|
|
- read-only
|
|
- propose-only
|
|
- apply-limited
|
|
|
|
|
|
## Recommended Next Concrete Steps
|
|
|
|
1. Build Stage 2 first
|
|
|
|
- datasource health records
|
|
- deterministic checks
|
|
- no automatic repair
|
|
|
|
2. Then build Stage 3
|
|
|
|
- Health Agent
|
|
- proposal generation only
|
|
|
|
3. Then Stage 4
|
|
|
|
- override apply flow
|
|
- rollback and verification
|
|
|
|
4. Only after that start Stage 5
|
|
|
|
- broader situational-awareness assessment workflows
|
|
|
|
|
|
## Why This Order
|
|
|
|
Because situational-awareness quality depends on reliable upstream data.
|
|
|
|
If datasource health is weak:
|
|
|
|
- agent reasoning quality will degrade
|
|
- false explanations will increase
|
|
- assessment trust will drop
|
|
|
|
So datasource health is not a side task.
|
|
|
|
It is the first operational foundation for the later situational-awareness system.
|
|
|
|
|
|
## Summary
|
|
|
|
Planet should be built as:
|
|
|
|
- `aiprovider` for model access
|
|
- backend services for orchestration and persistence
|
|
- datasource health as the first evidence-governance layer
|
|
- agent runtime as the reusable reasoning core
|
|
- situational awareness as the long-term application layer
|
|
|
|
That path keeps the architecture coherent and lets each phase produce useful functionality without forcing a rewrite later.
|