408 lines
8.9 KiB
Markdown
408 lines
8.9 KiB
Markdown
# Lightweight Agent Orchestrator and WebSearch Evidence Plan
|
|
|
|
## Overview
|
|
|
|
Planet should not turn `aiprovider` into a general-purpose agent runtime.
|
|
|
|
`aiprovider` should remain the model gateway:
|
|
|
|
- provider compatibility
|
|
- protocol adaptation
|
|
- model authentication
|
|
- request and response normalization
|
|
|
|
Agent behavior belongs in the backend, where Planet already owns business state,
|
|
permissions, persistence, evidence records, and operator workflows.
|
|
|
|
The recommended direction is a lightweight backend Agent Orchestrator with a
|
|
controlled tool layer. The first version should use fixed workflows instead of a
|
|
free-form tool-calling loop.
|
|
|
|
|
|
## Architecture Decision
|
|
|
|
Use this boundary:
|
|
|
|
```text
|
|
aiprovider = model adapter only
|
|
backend Agent = task orchestration + tools + evidence + policy + business rules
|
|
```
|
|
|
|
This keeps model transport separate from Planet-specific behavior. It also lets
|
|
OpenAI, MiniMax, Anthropic-compatible providers, Ollama, and later providers all
|
|
reuse the same backend tools.
|
|
|
|
Recommended module shape:
|
|
|
|
```text
|
|
backend/app/services/
|
|
ai/
|
|
agent_orchestrator.py
|
|
tool_registry.py
|
|
prompts.py
|
|
schemas.py
|
|
ai_tools/
|
|
web_search.py
|
|
web_fetch.py
|
|
geo_resolve.py
|
|
internal_data_query.py
|
|
incident_query.py
|
|
evidence_store.py
|
|
situation/
|
|
bgp_analyzer.py
|
|
risk_scoring.py
|
|
event_correlator.py
|
|
alert_policy.py
|
|
|
|
aiprovider/
|
|
provider_service.py
|
|
main.py
|
|
```
|
|
|
|
|
|
## Phase 1: Controlled Workflow Agent
|
|
|
|
The first implementation should not be a full OpenClaw/Codex-style agent loop.
|
|
Planet's immediate needs are better served by explicit workflows:
|
|
|
|
1. `tutorial_refresh`
|
|
2. `geo_correction`
|
|
3. `situation_brief`
|
|
|
|
Each workflow should:
|
|
|
|
1. collect evidence with backend tools
|
|
2. normalize and store evidence
|
|
3. call `AIProviderClient` through the configured global provider/model/key
|
|
4. validate the result with Pydantic schemas
|
|
5. return a proposal, candidate, or brief instead of directly mutating critical state
|
|
|
|
For location correction, the flow should be:
|
|
|
|
```text
|
|
object name / type / current coordinate / description
|
|
-> web_search
|
|
-> web_fetch for selected results
|
|
-> geo_resolve for city/site coordinates
|
|
-> LLM structured extraction
|
|
-> schema validation and confidence scoring
|
|
-> pending review candidate
|
|
```
|
|
|
|
The LLM output must be constrained to a schema such as:
|
|
|
|
```json
|
|
{
|
|
"object_id": "string",
|
|
"object_type": "datacenter|ixp|submarine_cable|asn|city|facility|satellite",
|
|
"current_location": {
|
|
"lat": 0,
|
|
"lon": 0
|
|
},
|
|
"suggested_location": {
|
|
"lat": 0,
|
|
"lon": 0
|
|
},
|
|
"confidence": 0.82,
|
|
"reason": "short evidence-backed explanation",
|
|
"evidence": [
|
|
{
|
|
"title": "source title",
|
|
"url": "https://example.com/source",
|
|
"quote": "short supporting excerpt",
|
|
"retrieved_at": "2026-05-10T00:00:00Z"
|
|
}
|
|
],
|
|
"needs_human_review": true
|
|
}
|
|
```
|
|
|
|
The LLM may generate a suggestion, but it must not directly write final
|
|
coordinates into the dimension tables.
|
|
|
|
|
|
## Phase 2: Backend Tool Registry
|
|
|
|
Add a small Python tool interface in the backend:
|
|
|
|
```python
|
|
class ToolResult(BaseModel):
|
|
ok: bool
|
|
data: Any = None
|
|
error: str | None = None
|
|
evidence: list[dict] = []
|
|
```
|
|
|
|
Register tools through a backend registry:
|
|
|
|
```text
|
|
web_search
|
|
web_fetch
|
|
geo_resolve
|
|
internal_data_query
|
|
incident_query
|
|
evidence_store
|
|
```
|
|
|
|
Do not put WebSearch inside `aiprovider`.
|
|
|
|
Reasons:
|
|
|
|
- search is a business tool, not a model-provider feature
|
|
- search evidence must be stored and audited by the backend
|
|
- different LLM providers should share the same search pipeline
|
|
- Planet may switch between Tavily, Brave, Exa, SearXNG, or MiniMax MCP without
|
|
changing model transport
|
|
|
|
The first WebSearch implementation should be an HTTP evidence provider. Tavily is
|
|
the recommended first default because it is simple to call from the existing
|
|
`httpx` backend stack and returns LLM/RAG-friendly search results. The interface
|
|
should remain provider-neutral so Brave, Exa, SearXNG, or MiniMax MCP can be
|
|
added later.
|
|
|
|
WebSearch configuration should live under PostgreSQL `system_settings` with the
|
|
rest of external integrations:
|
|
|
|
```text
|
|
external_integrations.web_search
|
|
enabled
|
|
provider
|
|
api_key
|
|
base_url
|
|
max_results
|
|
timeout_seconds
|
|
```
|
|
|
|
Secret resolution should follow the existing settings pattern:
|
|
|
|
1. saved PostgreSQL secret
|
|
2. provider-specific environment variable, for example `TAVILY_API_KEY`
|
|
3. generic fallback `WEB_SEARCH_API_KEY`
|
|
|
|
|
|
## Phase 3: Limited Agent Loop
|
|
|
|
After the fixed workflows are stable, the backend can add a limited agent loop:
|
|
|
|
```text
|
|
LLM sees an allowed tool list
|
|
-> LLM requests a tool call
|
|
-> backend validates and executes the tool
|
|
-> tool result is added to context
|
|
-> LLM continues
|
|
-> final structured output after at most N steps
|
|
```
|
|
|
|
Guardrails:
|
|
|
|
- max tool steps: 3 to 5
|
|
- only read-only tools may run automatically
|
|
- writes go to pending review first
|
|
- all web evidence must be persisted
|
|
- all final outputs must pass schema validation
|
|
- prompts must include explicit evidence boundaries
|
|
|
|
Permission levels:
|
|
|
|
```text
|
|
L0: pure analysis, no tools
|
|
L1: read-only tools, web_search / web_fetch / internal_query
|
|
L2: proposal generation, write pending review records
|
|
L3: low-risk notifications and briefs
|
|
L4: database mutation or alert triggering, human confirmation required
|
|
```
|
|
|
|
|
|
## Situational Awareness Boundary
|
|
|
|
Planet's situational-awareness layer should not rely on the LLM as the primary
|
|
risk engine.
|
|
|
|
Use deterministic analysis for:
|
|
|
|
- anomaly type
|
|
- affected prefixes
|
|
- affected ASNs
|
|
- geographic scope
|
|
- duration
|
|
- severity score
|
|
- confidence
|
|
- related events
|
|
- raw evidence
|
|
|
|
Use the LLM for:
|
|
|
|
- readable summaries
|
|
- risk explanation
|
|
- likely impact narrative
|
|
- next recommended actions
|
|
- missing data requests
|
|
|
|
In short:
|
|
|
|
```text
|
|
deterministic services compute the score
|
|
LLM explains the evidence and options
|
|
```
|
|
|
|
Proactive alerts should be triggered by deterministic rules or scheduled jobs,
|
|
then optionally summarized by the Agent Orchestrator.
|
|
|
|
|
|
## Persistence Model
|
|
|
|
Add lightweight persistence for auditability:
|
|
|
|
```text
|
|
ai_tasks
|
|
id
|
|
task_type
|
|
status
|
|
input_json
|
|
output_json
|
|
model
|
|
created_at
|
|
finished_at
|
|
error
|
|
|
|
ai_evidence
|
|
id
|
|
task_id
|
|
source_type
|
|
title
|
|
url
|
|
snippet
|
|
content_hash
|
|
retrieved_at
|
|
credibility_score
|
|
|
|
ai_briefs
|
|
id
|
|
brief_type
|
|
severity
|
|
title
|
|
summary
|
|
evidence_ids
|
|
related_entity_ids
|
|
created_at
|
|
acknowledged_at
|
|
|
|
ai_location_suggestions
|
|
id
|
|
object_type
|
|
object_id
|
|
old_lat
|
|
old_lon
|
|
new_lat
|
|
new_lon
|
|
confidence
|
|
reason
|
|
evidence_ids
|
|
status
|
|
```
|
|
|
|
The tables can be introduced incrementally. The first implementation may start
|
|
with `ai_tasks` and `ai_evidence`, then add specialized tables when the UI needs
|
|
review queues and acknowledgement state.
|
|
|
|
|
|
## MVP Scope
|
|
|
|
The MVP should deliver three fixed capabilities:
|
|
|
|
### 1. Tutorial Refresh
|
|
|
|
Input:
|
|
|
|
- provider or tutorial topic
|
|
- current tutorial text
|
|
- known stale point, when available
|
|
|
|
Tools:
|
|
|
|
- `web_search`
|
|
- `web_fetch`
|
|
|
|
Output:
|
|
|
|
- updated Markdown
|
|
- source list
|
|
- verification status
|
|
|
|
### 2. Geo Correction
|
|
|
|
Input:
|
|
|
|
- object id
|
|
- object name
|
|
- object type
|
|
- current coordinates
|
|
- source description
|
|
|
|
Tools:
|
|
|
|
- `web_search`
|
|
- `web_fetch`
|
|
- `geo_resolve`
|
|
|
|
Output:
|
|
|
|
- `LocationCorrection` JSON
|
|
- evidence list
|
|
- pending review candidate
|
|
|
|
### 3. Situation Brief
|
|
|
|
Input:
|
|
|
|
- anomaly event
|
|
- deterministic findings
|
|
- internal data summary
|
|
|
|
Tools:
|
|
|
|
- `internal_data_query`
|
|
- optional `web_search`
|
|
|
|
Output:
|
|
|
|
- `SituationBrief` JSON
|
|
- risk explanation
|
|
- recommended actions
|
|
- missing evidence list
|
|
|
|
|
|
## Test Plan
|
|
|
|
Backend tests:
|
|
|
|
- WebSearch settings persist to `system_settings` and mask secrets in API responses.
|
|
- Env fallback resolves provider-specific keys before `WEB_SEARCH_API_KEY`.
|
|
- WebSearch provider normalizes success, empty results, 401, 429, and timeout responses.
|
|
- `tutorial_refresh` uses evidence when available and marks output unverified when no evidence exists.
|
|
- `geo_correction` returns pending review candidates and never writes final coordinates directly.
|
|
- `situation_brief` accepts deterministic findings and returns schema-valid summaries.
|
|
- Agent outputs fail closed when schema validation fails.
|
|
|
|
Frontend tests:
|
|
|
|
- WebSearch settings card shows configured state, masked key, connection test result, and save feedback.
|
|
- Candidate review UI can display evidence links and pending location suggestions.
|
|
- Situation brief UI can show evidence-backed summaries without exposing raw secrets.
|
|
|
|
Regression tests:
|
|
|
|
- existing `aiprovider` status and analysis calls remain unchanged
|
|
- current LLM provider configuration remains the global model source
|
|
- location pipeline tests continue to pass
|
|
- datasource credential guide tests continue to pass
|
|
|
|
|
|
## Assumptions
|
|
|
|
- `aiprovider` remains model-adapter-only.
|
|
- Backend tools are implemented directly in Python first; MCP support is optional and later.
|
|
- Search is evidence collection, not model transport.
|
|
- Writes to important domain tables require human confirmation.
|
|
- Deterministic analysis owns risk scores; LLM output is explanatory and evidence-backed.
|