8.9 KiB
Lightweight Agent Orchestrator and WebSearch Evidence Plan
Overview
Planet should not turn aiprovider into a general-purpose agent runtime.
aiprovider should remain the model gateway:
- provider compatibility
- protocol adaptation
- model authentication
- request and response normalization
Agent behavior belongs in the backend, where Planet already owns business state, permissions, persistence, evidence records, and operator workflows.
The recommended direction is a lightweight backend Agent Orchestrator with a controlled tool layer. The first version should use fixed workflows instead of a free-form tool-calling loop.
Architecture Decision
Use this boundary:
aiprovider = model adapter only
backend Agent = task orchestration + tools + evidence + policy + business rules
This keeps model transport separate from Planet-specific behavior. It also lets OpenAI, MiniMax, Anthropic-compatible providers, Ollama, and later providers all reuse the same backend tools.
Recommended module shape:
backend/app/services/
ai/
agent_orchestrator.py
tool_registry.py
prompts.py
schemas.py
ai_tools/
web_search.py
web_fetch.py
geo_resolve.py
internal_data_query.py
incident_query.py
evidence_store.py
situation/
bgp_analyzer.py
risk_scoring.py
event_correlator.py
alert_policy.py
aiprovider/
provider_service.py
main.py
Phase 1: Controlled Workflow Agent
The first implementation should not be a full OpenClaw/Codex-style agent loop. Planet's immediate needs are better served by explicit workflows:
tutorial_refreshgeo_correctionsituation_brief
Each workflow should:
- collect evidence with backend tools
- normalize and store evidence
- call
AIProviderClientthrough the configured global provider/model/key - validate the result with Pydantic schemas
- return a proposal, candidate, or brief instead of directly mutating critical state
For location correction, the flow should be:
object name / type / current coordinate / description
-> web_search
-> web_fetch for selected results
-> geo_resolve for city/site coordinates
-> LLM structured extraction
-> schema validation and confidence scoring
-> pending review candidate
The LLM output must be constrained to a schema such as:
{
"object_id": "string",
"object_type": "datacenter|ixp|submarine_cable|asn|city|facility|satellite",
"current_location": {
"lat": 0,
"lon": 0
},
"suggested_location": {
"lat": 0,
"lon": 0
},
"confidence": 0.82,
"reason": "short evidence-backed explanation",
"evidence": [
{
"title": "source title",
"url": "https://example.com/source",
"quote": "short supporting excerpt",
"retrieved_at": "2026-05-10T00:00:00Z"
}
],
"needs_human_review": true
}
The LLM may generate a suggestion, but it must not directly write final coordinates into the dimension tables.
Phase 2: Backend Tool Registry
Add a small Python tool interface in the backend:
class ToolResult(BaseModel):
ok: bool
data: Any = None
error: str | None = None
evidence: list[dict] = []
Register tools through a backend registry:
web_search
web_fetch
geo_resolve
internal_data_query
incident_query
evidence_store
Do not put WebSearch inside aiprovider.
Reasons:
- search is a business tool, not a model-provider feature
- search evidence must be stored and audited by the backend
- different LLM providers should share the same search pipeline
- Planet may switch between Tavily, Brave, Exa, SearXNG, or MiniMax MCP without changing model transport
The first WebSearch implementation should be an HTTP evidence provider. Tavily is
the recommended first default because it is simple to call from the existing
httpx backend stack and returns LLM/RAG-friendly search results. The interface
should remain provider-neutral so Brave, Exa, SearXNG, or MiniMax MCP can be
added later.
WebSearch configuration should live under PostgreSQL system_settings with the
rest of external integrations:
external_integrations.web_search
enabled
provider
api_key
base_url
max_results
timeout_seconds
Secret resolution should follow the existing settings pattern:
- saved PostgreSQL secret
- provider-specific environment variable, for example
TAVILY_API_KEY - generic fallback
WEB_SEARCH_API_KEY
Phase 3: Limited Agent Loop
After the fixed workflows are stable, the backend can add a limited agent loop:
LLM sees an allowed tool list
-> LLM requests a tool call
-> backend validates and executes the tool
-> tool result is added to context
-> LLM continues
-> final structured output after at most N steps
Guardrails:
- max tool steps: 3 to 5
- only read-only tools may run automatically
- writes go to pending review first
- all web evidence must be persisted
- all final outputs must pass schema validation
- prompts must include explicit evidence boundaries
Permission levels:
L0: pure analysis, no tools
L1: read-only tools, web_search / web_fetch / internal_query
L2: proposal generation, write pending review records
L3: low-risk notifications and briefs
L4: database mutation or alert triggering, human confirmation required
Situational Awareness Boundary
Planet's situational-awareness layer should not rely on the LLM as the primary risk engine.
Use deterministic analysis for:
- anomaly type
- affected prefixes
- affected ASNs
- geographic scope
- duration
- severity score
- confidence
- related events
- raw evidence
Use the LLM for:
- readable summaries
- risk explanation
- likely impact narrative
- next recommended actions
- missing data requests
In short:
deterministic services compute the score
LLM explains the evidence and options
Proactive alerts should be triggered by deterministic rules or scheduled jobs, then optionally summarized by the Agent Orchestrator.
Persistence Model
Add lightweight persistence for auditability:
ai_tasks
id
task_type
status
input_json
output_json
model
created_at
finished_at
error
ai_evidence
id
task_id
source_type
title
url
snippet
content_hash
retrieved_at
credibility_score
ai_briefs
id
brief_type
severity
title
summary
evidence_ids
related_entity_ids
created_at
acknowledged_at
ai_location_suggestions
id
object_type
object_id
old_lat
old_lon
new_lat
new_lon
confidence
reason
evidence_ids
status
The tables can be introduced incrementally. The first implementation may start
with ai_tasks and ai_evidence, then add specialized tables when the UI needs
review queues and acknowledgement state.
MVP Scope
The MVP should deliver three fixed capabilities:
1. Tutorial Refresh
Input:
- provider or tutorial topic
- current tutorial text
- known stale point, when available
Tools:
web_searchweb_fetch
Output:
- updated Markdown
- source list
- verification status
2. Geo Correction
Input:
- object id
- object name
- object type
- current coordinates
- source description
Tools:
web_searchweb_fetchgeo_resolve
Output:
LocationCorrectionJSON- evidence list
- pending review candidate
3. Situation Brief
Input:
- anomaly event
- deterministic findings
- internal data summary
Tools:
internal_data_query- optional
web_search
Output:
SituationBriefJSON- risk explanation
- recommended actions
- missing evidence list
Test Plan
Backend tests:
- WebSearch settings persist to
system_settingsand mask secrets in API responses. - Env fallback resolves provider-specific keys before
WEB_SEARCH_API_KEY. - WebSearch provider normalizes success, empty results, 401, 429, and timeout responses.
tutorial_refreshuses evidence when available and marks output unverified when no evidence exists.geo_correctionreturns pending review candidates and never writes final coordinates directly.situation_briefaccepts deterministic findings and returns schema-valid summaries.- Agent outputs fail closed when schema validation fails.
Frontend tests:
- WebSearch settings card shows configured state, masked key, connection test result, and save feedback.
- Candidate review UI can display evidence links and pending location suggestions.
- Situation brief UI can show evidence-backed summaries without exposing raw secrets.
Regression tests:
- existing
aiproviderstatus and analysis calls remain unchanged - current LLM provider configuration remains the global model source
- location pipeline tests continue to pass
- datasource credential guide tests continue to pass
Assumptions
aiproviderremains model-adapter-only.- Backend tools are implemented directly in Python first; MCP support is optional and later.
- Search is evidence collection, not model transport.
- Writes to important domain tables require human confirmation.
- Deterministic analysis owns risk scores; LLM output is explanatory and evidence-backed.