648 lines
12 KiB
Markdown
648 lines
12 KiB
Markdown
# Agent Architecture Plan
|
|
|
|
## Overview
|
|
|
|
This document defines the agent architecture for Planet.
|
|
|
|
The architecture is intentionally broader than datasource health checking.
|
|
|
|
It is designed to support both:
|
|
|
|
- datasource health governance
|
|
- future situational-awareness workflows
|
|
|
|
The core idea is to avoid building a one-off "repair broken API links" agent.
|
|
|
|
Instead, Planet should grow a reusable agent runtime that can:
|
|
|
|
- collect evidence
|
|
- evaluate signals
|
|
- reason over incomplete information
|
|
- generate proposals
|
|
- produce assessments
|
|
- execute limited actions under policy
|
|
|
|
|
|
## Design Goal
|
|
|
|
Build an agent foundation that can evolve in this order:
|
|
|
|
1. datasource health checks
|
|
2. datasource repair proposals
|
|
3. signal correlation
|
|
4. situational assessments
|
|
5. controlled runtime actions
|
|
|
|
This means the architecture should treat datasource health as one use case of the larger agent system, not as the whole system.
|
|
|
|
|
|
## Core Principles
|
|
|
|
1. Separate evidence from reasoning
|
|
|
|
- raw signals should be gathered first
|
|
- deterministic checks should run before LLM reasoning
|
|
|
|
2. Agents do not own the defaults
|
|
|
|
- repository defaults remain human-owned
|
|
- agents operate on runtime state, proposals, and overrides
|
|
|
|
3. Reasoning and action are different responsibilities
|
|
|
|
- many agents should be read-only or propose-only
|
|
- only tightly controlled flows may apply changes
|
|
|
|
4. Shared runtime, specialized roles
|
|
|
|
- multiple agent roles should share the same object model and orchestration patterns
|
|
- health and situational-awareness agents should not invent incompatible payloads
|
|
|
|
5. Auditability is mandatory
|
|
|
|
- every proposal, assessment, and applied action should be attributable
|
|
|
|
|
|
## System Layers
|
|
|
|
Planet agent architecture should be split into four layers.
|
|
|
|
### 1. Signal Layer
|
|
|
|
Purpose:
|
|
|
|
- gather raw evidence from internal and external systems
|
|
|
|
Example sources:
|
|
|
|
- collector outputs
|
|
- datasource health checks
|
|
- logs
|
|
- snapshots
|
|
- alerts
|
|
- web search results
|
|
- scraped pages
|
|
- external APIs
|
|
- operator inputs
|
|
|
|
Responsibilities:
|
|
|
|
- fetch
|
|
- normalize
|
|
- timestamp
|
|
- tag with source and trust level
|
|
|
|
This layer should not make high-level judgments.
|
|
|
|
|
|
### 2. Evaluation Layer
|
|
|
|
Purpose:
|
|
|
|
- perform deterministic analysis
|
|
|
|
Examples:
|
|
|
|
- reachability checks
|
|
- schema validation
|
|
- threshold checks
|
|
- time-window comparisons
|
|
- anomaly counters
|
|
- completeness checks
|
|
|
|
Responsibilities:
|
|
|
|
- classify signals into machine-readable findings
|
|
- attach deterministic evidence
|
|
|
|
This layer should avoid LLM dependency whenever possible.
|
|
|
|
|
|
### 3. Reasoning Layer
|
|
|
|
Purpose:
|
|
|
|
- use LLMs when semantic interpretation or incomplete-information reasoning is needed
|
|
|
|
Examples:
|
|
|
|
- endpoint migration inference
|
|
- multi-source event correlation
|
|
- causality hypotheses
|
|
- ambiguity reduction
|
|
- assessment narrative generation
|
|
- action recommendation generation
|
|
|
|
Responsibilities:
|
|
|
|
- synthesize evidence
|
|
- produce hypotheses
|
|
- rank confidence
|
|
- explain reasoning boundaries
|
|
|
|
This is the main place where `aiprovider` and web search are used.
|
|
|
|
|
|
### 4. Action Layer
|
|
|
|
Purpose:
|
|
|
|
- convert proposals or assessments into controlled system actions
|
|
|
|
Examples:
|
|
|
|
- create runtime override
|
|
- create proposal
|
|
- publish alert
|
|
- update operator task queue
|
|
- generate summary artifact
|
|
- trigger follow-up verification
|
|
|
|
Responsibilities:
|
|
|
|
- enforce policy
|
|
- enforce approval requirements
|
|
- verify post-action outcomes
|
|
- record audit trails
|
|
|
|
|
|
## Architecture Sketch
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
A["Collectors / Logs / Snapshots / External APIs"] --> B["Signal Layer"]
|
|
W["Web Search / Page Fetch / Docs"] --> B
|
|
B --> C["Evaluation Layer"]
|
|
C --> D["Findings"]
|
|
D --> E["Reasoning Layer (LLM + Tools)"]
|
|
E --> F["Proposals"]
|
|
E --> G["Assessments"]
|
|
F --> H["Action Layer"]
|
|
H --> I["Runtime Overrides / Alerts / Tasks"]
|
|
H --> J["Verification Loop"]
|
|
J --> B
|
|
|
|
K["Policy Engine"] --> H
|
|
L["Audit / History Store"] --> H
|
|
L --> E
|
|
L --> C
|
|
```
|
|
|
|
|
|
## Agent Roles
|
|
|
|
The first version should define these logical roles.
|
|
|
|
### 1. Health Agent
|
|
|
|
Primary use case:
|
|
|
|
- datasource health governance
|
|
|
|
Inputs:
|
|
|
|
- datasource metadata
|
|
- current endpoint
|
|
- latest health records
|
|
- latest failures
|
|
- deterministic findings
|
|
|
|
Outputs:
|
|
|
|
- health interpretation
|
|
- repair proposal
|
|
- confidence
|
|
- evidence references
|
|
|
|
Typical action level:
|
|
|
|
- propose-only
|
|
|
|
|
|
### 2. Correlation Agent
|
|
|
|
Primary use case:
|
|
|
|
- identify whether multiple signals describe the same event or related events
|
|
|
|
Inputs:
|
|
|
|
- findings from multiple collectors
|
|
- time windows
|
|
- region / ASN / prefix / cable relationships
|
|
- prior incidents
|
|
|
|
Outputs:
|
|
|
|
- grouped event candidates
|
|
- correlation rationale
|
|
- confidence per relationship
|
|
|
|
Typical action level:
|
|
|
|
- read-only
|
|
|
|
|
|
### 3. Assessment Agent
|
|
|
|
Primary use case:
|
|
|
|
- produce situational-awareness outputs
|
|
|
|
Inputs:
|
|
|
|
- grouped events
|
|
- findings
|
|
- current context
|
|
- historical context
|
|
- operator constraints
|
|
|
|
Outputs:
|
|
|
|
- structured assessment
|
|
- risk summary
|
|
- evidence-backed recommendations
|
|
- missing-information list
|
|
|
|
Typical action level:
|
|
|
|
- read-only or propose-only
|
|
|
|
|
|
### 4. Recovery Agent
|
|
|
|
Primary use case:
|
|
|
|
- carry low-risk proposals into controlled runtime actions
|
|
|
|
Inputs:
|
|
|
|
- approved proposal
|
|
- policy constraints
|
|
- trusted-domain rules
|
|
- verification checks
|
|
|
|
Outputs:
|
|
|
|
- applied override
|
|
- failed application
|
|
- rollback request
|
|
|
|
Typical action level:
|
|
|
|
- apply-limited
|
|
|
|
|
|
## Shared Object Model
|
|
|
|
All agents should work on a shared object model.
|
|
|
|
That prevents the health subsystem and situational-awareness subsystem from drifting into incompatible payloads.
|
|
|
|
### Signal
|
|
|
|
Represents a raw observed fact.
|
|
|
|
Examples:
|
|
|
|
- a datasource returned HTTP 404
|
|
- a collector returned empty results
|
|
- BGP updates spiked in one region
|
|
- a known endpoint now redirects elsewhere
|
|
|
|
Suggested shape:
|
|
|
|
```json
|
|
{
|
|
"id": "sig_123",
|
|
"type": "datasource.http_failure",
|
|
"source": "ris_live_bgp",
|
|
"occurred_at": "2026-04-08T10:00:00Z",
|
|
"severity": "medium",
|
|
"payload": {},
|
|
"trust": 0.95
|
|
}
|
|
```
|
|
|
|
|
|
### Finding
|
|
|
|
Represents a deterministic or semi-deterministic interpretation of one or more signals.
|
|
|
|
Examples:
|
|
|
|
- `schema_changed`
|
|
- `endpoint_unreachable`
|
|
- `data_volume_abnormally_low`
|
|
- `event_cluster_detected`
|
|
|
|
Suggested shape:
|
|
|
|
```json
|
|
{
|
|
"id": "find_123",
|
|
"type": "datasource.schema_changed",
|
|
"source_ids": ["sig_123"],
|
|
"confidence": 0.92,
|
|
"evidence": [],
|
|
"details": {}
|
|
}
|
|
```
|
|
|
|
|
|
### Proposal
|
|
|
|
Represents a recommended action, not an already-applied action.
|
|
|
|
Examples:
|
|
|
|
- switch endpoint to new URL
|
|
- disable bad override
|
|
- escalate issue for manual review
|
|
|
|
Suggested shape:
|
|
|
|
```json
|
|
{
|
|
"id": "prop_123",
|
|
"kind": "endpoint_override",
|
|
"target": "telegeography_cables",
|
|
"confidence": 0.84,
|
|
"reason": "Official docs now point to a new API path",
|
|
"payload": {},
|
|
"evidence_urls": [],
|
|
"status": "proposed"
|
|
}
|
|
```
|
|
|
|
|
|
### Assessment
|
|
|
|
Represents a structured situational-awareness output for operators or downstream systems.
|
|
|
|
Examples:
|
|
|
|
- current network posture summary
|
|
- incident impact assessment
|
|
- risk and response recommendations
|
|
|
|
Suggested shape:
|
|
|
|
```json
|
|
{
|
|
"id": "assess_123",
|
|
"scope": "regional-network",
|
|
"risk_level": "high",
|
|
"summary": "Regional routing instability is increasing.",
|
|
"key_risks": [],
|
|
"evidence": [],
|
|
"recommendations": [],
|
|
"missing_data": []
|
|
}
|
|
```
|
|
|
|
|
|
## State Machine
|
|
|
|
The shared orchestration flow should look like this:
|
|
|
|
```mermaid
|
|
stateDiagram-v2
|
|
[*] --> Collect
|
|
Collect --> Validate
|
|
Validate --> Classify
|
|
Classify --> Reason
|
|
Reason --> Propose
|
|
Reason --> Assess
|
|
Propose --> Review
|
|
Review --> Apply
|
|
Apply --> Verify
|
|
Verify --> Archive
|
|
Assess --> Archive
|
|
Archive --> [*]
|
|
```
|
|
|
|
Definitions:
|
|
|
|
- `Collect`: gather signals
|
|
- `Validate`: run deterministic checks
|
|
- `Classify`: create findings
|
|
- `Reason`: invoke LLM reasoning when needed
|
|
- `Propose`: create change proposals
|
|
- `Review`: policy or human approval
|
|
- `Apply`: perform limited runtime action
|
|
- `Verify`: confirm action effect
|
|
- `Archive`: store artifacts and decisions
|
|
|
|
|
|
## Permission Model
|
|
|
|
Each agent role should be assigned one of these action levels.
|
|
|
|
### `read-only`
|
|
|
|
Allowed:
|
|
|
|
- read signals
|
|
- search web
|
|
- fetch pages
|
|
- read internal state
|
|
- generate findings and assessments
|
|
|
|
Not allowed:
|
|
|
|
- mutate config
|
|
- write overrides
|
|
- change live runtime behavior
|
|
|
|
|
|
### `propose-only`
|
|
|
|
Allowed:
|
|
|
|
- everything in `read-only`
|
|
- create proposals
|
|
- create review tasks
|
|
|
|
Not allowed:
|
|
|
|
- apply live changes
|
|
|
|
|
|
### `apply-limited`
|
|
|
|
Allowed:
|
|
|
|
- everything in `propose-only`
|
|
- write approved runtime overrides
|
|
- trigger verification checks
|
|
|
|
Not allowed:
|
|
|
|
- mutate repository defaults
|
|
- make destructive data changes
|
|
- bypass policy engine
|
|
|
|
|
|
## Runtime Components
|
|
|
|
The first durable architecture should introduce these components.
|
|
|
|
### 1. Signal Store
|
|
|
|
Stores normalized evidence and health outputs.
|
|
|
|
|
|
### 2. Finding Store
|
|
|
|
Stores deterministic classifications that can be reused by multiple agents.
|
|
|
|
|
|
### 3. Proposal Store
|
|
|
|
Stores recommended actions with evidence and confidence.
|
|
|
|
|
|
### 4. Assessment Store
|
|
|
|
Stores structured situational-awareness outputs.
|
|
|
|
|
|
### 5. Policy Engine
|
|
|
|
Decides:
|
|
|
|
- whether agent may run
|
|
- whether proposal requires review
|
|
- whether proposal may auto-apply
|
|
- whether post-apply verification passed
|
|
|
|
|
|
### 6. Override Store
|
|
|
|
Stores runtime-only configuration changes.
|
|
|
|
This is where endpoint repairs should live.
|
|
|
|
|
|
## Relation To `aiprovider`
|
|
|
|
`aiprovider` should remain the model gateway.
|
|
|
|
It should not become the full agent runtime.
|
|
|
|
Recommended split:
|
|
|
|
- `aiprovider`
|
|
- provider adaptation
|
|
- prompt transport
|
|
- model execution
|
|
- protocol compatibility
|
|
|
|
- agent runtime
|
|
- orchestration
|
|
- signal handling
|
|
- tool selection
|
|
- proposal generation
|
|
- policy and audit
|
|
|
|
This keeps provider concerns and agent behavior concerns separate.
|
|
|
|
|
|
## Relation To Datasource Health
|
|
|
|
Datasource health becomes one vertical slice of this architecture.
|
|
|
|
Mapping:
|
|
|
|
- signal:
|
|
- endpoint unreachable
|
|
- schema mismatch
|
|
- bad content type
|
|
- finding:
|
|
- `failed`
|
|
- `schema_changed`
|
|
- `moved_endpoint_suspected`
|
|
- proposal:
|
|
- runtime override suggestion
|
|
- assessment:
|
|
- datasource health summary for operators
|
|
|
|
|
|
## Relation To Situational Awareness
|
|
|
|
Future situational-awareness capabilities should reuse the same flow:
|
|
|
|
- raw telemetry becomes signals
|
|
- anomaly detection becomes findings
|
|
- LLM correlation becomes reasoning
|
|
- operator-facing output becomes assessments
|
|
- policy-approved mitigations become actions
|
|
|
|
This lets the platform evolve from operational health governance into broader cyber/network posture workflows without changing the architecture.
|
|
|
|
|
|
## Suggested Delivery Sequence
|
|
|
|
### Phase A
|
|
|
|
- finalize shared object model
|
|
- implement health-oriented signal and finding storage
|
|
|
|
### Phase B
|
|
|
|
- implement Health Agent
|
|
- generate proposals only
|
|
|
|
### Phase C
|
|
|
|
- implement Assessment Agent
|
|
- expose structured assessments via API
|
|
|
|
### Phase D
|
|
|
|
- implement Correlation Agent
|
|
- support multi-source incident grouping
|
|
|
|
### Phase E
|
|
|
|
- implement Recovery Agent with policy-gated runtime actions
|
|
|
|
|
|
## Recommended First Build
|
|
|
|
The first build should not try to implement every agent role.
|
|
|
|
Recommended initial slice:
|
|
|
|
- shared object model
|
|
- health signals
|
|
- health findings
|
|
- Health Agent
|
|
- proposal generation only
|
|
|
|
This gives immediate value while preserving the longer-term architecture.
|
|
|
|
|
|
## Non-Goals For The First Iteration
|
|
|
|
- repository YAML auto-rewrites
|
|
- unrestricted autonomous action
|
|
- full incident graph reasoning
|
|
- automatic large-scale remediation
|
|
- agent-owned configuration source of truth
|
|
|
|
|
|
## Summary
|
|
|
|
Planet should treat agents as a reusable runtime for evidence, reasoning, proposals, and assessments.
|
|
|
|
The datasource health use case is the first practical entrypoint, but the architecture should already assume future situational-awareness expansion.
|
|
|
|
The safest path is:
|
|
|
|
- deterministic checks first
|
|
- agent reasoning second
|
|
- proposals before actions
|
|
- runtime overrides instead of default mutation
|