21 KiB
Agent Runtime, Earth LLM Command, And Speech Entry Plan
Summary
Build an auditable backend Agent Runtime and upgrade the existing Earth search panel into a combined search, AI command, and voice wake entry. The first version is a runtime foundation, not the full multi-role simulation product yet.
Typical user goals:
- In Earth, type "高亮所有北斗卫星" and have the system open the satellite layer and highlight matching Beidou satellites.
- Type "中国大陆的算力中心" and have the system open the compute-center layer, match mainland China compute centers, and highlight them.
- When microphone permission is granted, use a configurable wake word, then speak an Earth command.
- Save every Earth AI command as an agent run so operators can review the original input, speech transcription, tool steps, entity matches, final action plan, and frontend execution result.
Core boundaries:
aiproviderremains the model gateway. It must not own business tools, database access, Earth actions, or agent policy.- The backend owns agent orchestration, tools, evidence storage, permission policy, and proposal application.
- Earth v1 executes visualization actions only. It does not mutate business data.
- Speech recognition uses a provider-neutral ASR API first, defaulting to OpenAI/Whisper-compatible transcription APIs, with local
whisper.cppstyle providers as later adapters.
Architecture
Backend Agent Runtime
Recommended module shape:
backend/app/models/
agent_run.py
agent_step.py
agent_evidence.py
agent_proposal.py
backend/app/schemas/
agents.py
speech.py
backend/app/services/agents/
runtime.py
orchestrator.py
tool_protocol.py
tool_registry.py
policy.py
proposals.py
earth_command.py
entity_query.py
speech.py
backend/app/api/v1/
agents.py
Database conventions should follow the current project style:
- Use SQLAlchemy models.
- Import new models from
init_db(). - Let
Base.metadata.create_allcreate tables. - Add required indexes with
CREATE INDEX IF NOT EXISTS. - Do not introduce Alembic for this feature.
aiprovider Boundary
aiprovider should continue to provide model transport only:
- Keep the existing
/v1/analyzeendpoint. - If schemas are extended, only pass through model/provider request fields and normalize responses.
- Do not add WebSearch, database queries, Earth entity lookup, or business action execution inside
aiprovider. - If a provider does not support native tool calling, the backend must use JSON tool-call fallback.
Agent Tool Protocol
The first version should support two protocols:
- Default path: backend JSON tool-call loop.
- Optional path: provider-native tools when the configured provider supports them.
- Fallback path: if provider-native tools are unavailable or unstable, automatically use JSON tool-call.
Example JSON tool call:
{
"type": "tool_call",
"tool": "earth.find_entities",
"arguments": {
"domain": "satellites",
"filters": {
"constellation": "beidou"
},
"limit": 500
}
}
Example final response:
{
"type": "final",
"summary": "已找到并高亮北斗卫星。",
"result": {
"actions": []
}
}
Hard limits:
- One run may execute at most 8 tool steps by default.
- One tool call times out after 30 seconds by default.
- Only registered whitelist tools may run.
- Tool arguments must pass Pydantic validation.
- Illegal tools, invalid arguments, and denied actions must be recorded as step errors.
- LLM output may not directly write business state. Writes are either proposals or backend-executed policy-approved actions.
Data Model
AgentRun
Suggested fields:
id: integer primary key.public_id: public string id.run_type:earth_command,situational_awareness,config_proposal, ordiagnostic.status:queued,running,waiting_approval,completed,failed, orstopped.title: short display title.objective: text objective.input: JSONB original input, including text and audio metadata.context: JSONB run context.result_markdown: final human-readable output.result_json: structured output, including Earth action plans.provider: nullable provider id.model: nullable model id.request_id: nullable propagated request id.created_by: user id.created_at,updated_at,completed_at.error: nullable text error.
AgentStep
Suggested fields:
id.run_id.step_index.step_type:llm,tool,policy,action, ortranscription.status:pending,running,completed,failed, orskipped.name: step name, for exampleearth.find_entities.input: JSONB.output: JSONB.error: nullable text.started_at,completed_at.duration_ms.
AgentEvidence
Suggested fields:
id.run_id.step_id: nullable.evidence_type:internal_record,web_search,web_fetch,entity_match, ortranscription.source: source id.title: display title.url: nullable source URL.content: text evidence content.content_hash: nullable hash.metadata: JSONB.retrieved_at.
AgentProposal
Suggested fields:
id.run_id.proposal_type:datasource_config,ai_prompt, orexternal_integration.status:pending,approved,applied,rejected, orfailed.risk_level:low,medium, orhigh.target: JSONB target descriptor.before_payload: JSONB.after_payload: JSONB.rationale: text.policy_result: JSONB.applied_by: nullable user id.applied_at: nullable timestamp.error: nullable text.
Public APIs
Agent Runs
Add:
POST /api/v1/agents/runs
GET /api/v1/agents/runs
GET /api/v1/agents/runs/{run_id}
POST /api/v1/agents/runs/{run_id}/stop
POST /api/v1/agents/proposals/{proposal_id}/apply
Earth Command
Add:
POST /api/v1/agents/earth/command
Request shape:
{
"input_text": "高亮所有北斗卫星",
"source": "text",
"transcription_id": null,
"client_context": {
"visible_layers": ["satellites"],
"locale": "zh-CN",
"viewport": {
"is_mobile": false
}
}
}
Response shape:
{
"run_id": "agent_xxx",
"summary": "已找到并高亮北斗卫星。",
"actions": [
{
"type": "show_layer",
"layer": "satellites",
"enabled": true
},
{
"type": "highlight_entities",
"domain": "satellites",
"entity_ids": ["satellite:norad:12345"],
"style": {
"color": "#7dd3fc"
}
},
{
"type": "open_result_panel",
"title": "北斗卫星",
"items": []
}
],
"matched_entities": [],
"confidence": 0.86,
"missing_data": []
}
Speech / ASR
Add:
POST /api/v1/agents/speech/transcriptions
Request should use multipart form data:
file: audio blob.language: defaultzh.provider: optional provider override.source: defaultearth_command.
Response shape:
{
"text": "高亮所有北斗卫星",
"provider": "openai_compatible",
"model": "whisper-1",
"duration_ms": 1234,
"confidence": null,
"metadata": {}
}
Earth Action Plan
Allowed Action Types
The first version may only return:
show_layer
highlight_entities
filter_entities
focus_view
open_result_panel
clear_highlight
Allowed Domains
The first version supports:
satellites
compute_centers
bgp
news
vessels
cables
EarthAction Shape
Example:
{
"type": "highlight_entities",
"domain": "satellites",
"entity_ids": ["satellite:norad:12345"],
"style": {
"color": "#7dd3fc",
"mode": "glow"
},
"reason": "用户要求高亮北斗卫星"
}
Safety Rules
The backend must validate action plans before returning them:
typemust be in the allowed action list.domainmust be in the allowed domain list.entity_idsmust come from backend or current Earth candidate data. The model may not invent ids.- One highlight action should include at most 500 entities by default. If more match, return a truncation note in
missing_dataorsummary. focus_viewmust include valid coordinates, a valid region, or a matched entity.- No action may return executable JavaScript, arbitrary CSS, or arbitrary URLs to fetch.
Earth Entity Query
Add a backend entity query service used by both deterministic resolvers and LLM tools.
Satellites
Sources:
/api/v1/visualization/geo/satellites.- Current TLE collected data.
- GeoJSON feature properties.
Filters:
constellation:beidou,gps,galileo,glonass,starlink,iridium,geo,leo.- Name contains.
- NORAD id.
- Country/operator when present in data.
- Orbital class when inferable from existing fields.
Beidou matching:
- Prefer
constellation == beidou. - Then match names containing
BEIDOU,BDS,BEIDOU-, or北斗. - Stable entity id format should be
satellite:norad:{norad_id}when possible, with fallbacksatellite:index:{index}.
Compute Centers
Sources:
/api/v1/visualization/geo/compute-centers.- Unified TOP500 and Epoch AI GPU GeoJSON properties.
Filters:
- Country/region.
site_type:supercomputerorgpu_cluster.- Source:
top500orepoch_ai_gpu. - Name contains.
needs_confirmation.location_precision.
Mainland China matching:
- Match country values such as
China,中国, orPeople's Republic of China. - Exclude obvious non-mainland records when fields identify Hong Kong, Macau, or Taiwan.
- If records do not expose enough region detail to separate mainland China from Hong Kong, Macau, or Taiwan, return a
missing_datanote and conservatively matchcountry=China.
BGP, News, Vessels, And Cables
First-version basic support:
- BGP: severity, status, region, collector, prefix, ASN.
- News: region, source, localized title, localized summary.
- Vessels: vessel type, country/area, name, status.
- Cables: cable name, landing point, country/region.
Earth Frontend Integration
Search Panel Merge
Reuse the existing Earth search panel:
- Default behavior remains normal local search.
- Add an AI command state for natural-language commands.
- Add a "use AI" command button.
- Add a microphone button.
- Show running status, result summary, and clear-highlight action.
- Visually distinguish ordinary search results from AI action results.
Natural-language routing:
- If the user clicks the AI command button, always call the AI command endpoint.
- If input contains action words such as
高亮,显示,找出,聚焦,打开,筛选,隐藏, or清除, suggest AI command mode. - Short ordinary keywords continue to use local search.
Earth Action Executor
Add:
frontend/public/earth/js/earth-command-actions.js
Responsibilities:
- Receive backend
EarthActionPlan. - Toggle required layers.
- Apply highlights and filters.
- Focus view when requested.
- Open or update the result panel.
- Clear the previous AI command highlight.
Execution order:
- Clear previous AI command highlights unless the plan is explicitly additive.
- Show required layers.
- Apply filters and highlights.
- Focus the view if requested.
- Open the result panel.
- Record the execution summary in UI state.
Satellite Highlight
Existing highlightRelatedSatellites(indices, color) can be reused, but Earth command support needs an entity-id-to-index map.
Add:
- Satellite id map built from current satellite data.
- Highlight by
entity_ids. - Clear AI command satellite highlights without clearing manual locked selection.
- Auto-enable the satellite layer before highlighting.
Compute Center Highlight
Existing compute center markers support marker state, but need a dedicated batch AI highlight state.
Add:
- Batch set compute center markers to AI highlighted state.
- Keep AI highlight compatible with hover and locked states.
- Clear AI command compute highlights without clearing manual locked selection.
- Auto-enable the compute-center layer before highlighting.
Result Panel
First version can show results inside the merged search panel:
- Title, for example
北斗卫星. - Count, for example
已高亮 32 个对象. - List with the first 20 entities.
- Actions: clear highlight, rerun, view Agent Run.
Voice And Wake Word
ASR Configuration
Add Speech/ASR under the AI settings tool tab.
Fields:
enabled.provider.base_url.api_key.model.language.timeout_seconds.max_audio_size_mb.
First-version providers:
openai_whisper.openai_compatible.local_whisper.
Recommended defaults:
provider:openai_compatible.model:whisper-1or user-configured equivalent.language:zh.
Secret handling:
- Use the existing masked secret pattern.
- Do not allow agent proposals to write API keys.
- Support environment fallback such as
ASR_API_KEYandOPENAI_API_KEY.
Wake Word
The wake word is configurable.
Implementation rules:
- Treat wake word as a browser-local preference in v1.
- Store it in
localStorage. - Default suggestion:
小星球. - Provide a wake-word setting inside the Earth search/command panel.
- Do not listen until the user explicitly enables voice wake.
- Do not upload audio before wake.
- After wake, record one instruction audio segment and upload it to backend ASR.
Fallbacks:
- Browser does not support continuous local recognition: fall back to click-to-record.
- Microphone permission denied: show
麦克风权限未开启,仍可输入文字指令. - ASR not configured: show
语音识别未配置; text commands remain available. - User can disable wake listening and keep manual microphone recording.
Wake Word Technical Choice
Do not require local Whisper in v1.
Recommended path:
- Use browser Web Speech API for local wake-word detection when available.
- Fall back to click-to-record when unavailable.
- Upload only the post-wake instruction audio to backend ASR.
- Add local
whisper.cppstreaming wake-word or ASR provider later.
Prompt Registry
Add default prompt keys:
agents.runtime.system
agents.earth.command
agents.situational.assessment
agents.config.proposal
agents.roles.network
agents.roles.bgp
agents.roles.platform_ops
agents.roles.business_impact
agents.earth.command must require:
- Strict JSON output only.
- Use only provided candidate entities.
- Do not invent objects.
- Do not return arbitrary code.
- Do not modify business data.
- Return
clarification_neededwhen intent is unclear. - Return action plans matching the schema.
- Use Chinese summary for Chinese user input.
Policy
Tool Policy
- The LLM can request tools, but the backend decides whether to execute.
- Every tool must declare:
- name.
- description.
- input schema.
- output schema.
- permission.
- side-effect level.
- First-version side-effect levels:
read: can run directly.proposal: can only generate proposals.write: can only execute during proposal apply.
Proposal Policy
First version supports configuration proposal application with role-based automatic/manual gating:
super_adminmay enable automatic application for low-risk proposals.- Normal admins must manually confirm.
- High-risk proposals always require manual confirmation.
- Proposals containing secret fields are rejected.
- Proposals failing schema validation are rejected.
- Before and after payloads must be saved.
- Failed applications must save errors.
Low-risk scope:
- Datasource endpoint/config non-secret fields.
- AI prompt overrides.
- External integration non-secret fields.
Out of scope for v1:
- Alert acknowledge/resolve.
- Data deletion.
- User permission changes.
- Authentication configuration changes.
- Database schema changes by agent.
- Mutation of original Earth collected data.
Agent Operations UI
Add an Agent page under 运维与配置.
Run list:
- Status.
- Type.
- Title.
- Creator.
- Model.
- Time.
- Duration.
Run detail:
- Input.
- Status timeline.
- LLM steps.
- Tool steps.
- Evidence.
- Final result.
- Proposals.
Proposal apply:
- Before/after diff.
- Risk level.
- Policy result.
- Apply button.
- Reject button.
Earth command run detail:
- Original text or speech transcription.
- Matched entities.
- Action plan.
- Link back to Earth or copy run id.
Frontend Types
Add or extend:
interface AgentRun {}
interface AgentStep {}
interface AgentEvidence {}
interface AgentProposal {}
interface EarthCommandRequest {}
interface EarthActionPlan {}
interface EarthAction {}
interface SpeechTranscriptionResponse {}
Earth action executor API:
executeEarthActionPlan(plan, context)
clearEarthCommandHighlights()
getEarthCommandExecutionState()
Testing
Backend Tests
Use uv.
Suggested files:
backend/tests/test_agents_runtime.py
backend/tests/test_agent_tool_protocol.py
backend/tests/test_agent_policy.py
backend/tests/test_earth_agent_command.py
backend/tests/test_speech_transcription.py
Coverage:
- Create run.
- Persist ordered run steps.
- Parse valid JSON tool calls.
- Reject invalid JSON tool calls without executing tools.
- Reject unregistered tools.
- Reject invalid tool arguments.
- Pass through provider-native tools fields when configured.
- Fall back to JSON tool calls when native tools are unavailable.
- When WebSearch is not configured, record
missing_dataand do not fail Earth command. - "高亮所有北斗卫星" returns a satellite highlight action.
- "中国大陆的算力中心" returns a compute-center highlight action.
- Mainland China matching excludes Hong Kong, Macau, and Taiwan when fields permit it.
- Field granularity limitations are returned in
missing_data. - ASR not configured returns a clear error.
- ASR provider success returns transcription.
- Normal admins cannot auto-apply proposals.
super_admincan apply low-risk proposals.- Secret-field proposals are rejected.
Frontend Tests
Use bun.
Required build check:
cd frontend && bun run build
Suggested Earth JS tests:
- Action executor opens the satellite layer.
- Action executor highlights satellite entity ids.
- Action executor highlights compute-center entity ids.
- Clear highlight does not clear manual locked selection.
- Search panel ordinary search still uses local search.
- AI command button calls
/api/v1/agents/earth/command. - ASR missing configuration and microphone permission denial show clear UI feedback.
Manual Acceptance
-
Type
高亮所有北斗卫星.- Satellite layer opens.
- Beidou satellites are highlighted.
- Panel shows count and summary.
- Agent run is reviewable.
-
Type
显示中国大陆的算力中心.- Compute-center layer opens.
- Mainland China related compute centers are highlighted.
- If data cannot separate mainland China from Hong Kong/Macau/Taiwan, the UI shows a missing-data note.
-
Click microphone.
- Permission denial is clear.
- Permission grant allows recording.
- Configured ASR transcribes and executes the command.
- If ASR is not configured, text input still works.
-
Enable voice wake.
- Wake word is configurable.
- Audio is not uploaded before wake.
- After wake, one instruction audio segment is uploaded.
- Wake listening can be disabled.
-
Open Agent operations UI.
- Earth command run is listed.
- Transcription step is visible when speech was used.
- Entity query step is visible.
- Final action plan is visible.
- User can return to Earth or copy the run id.
Documentation
When implemented, update:
docs/technical/zh/agents-aiprovider.md
docs/technical/en/agents-aiprovider.md
docs/technical/zh/manual.md
docs/technical/en/manual.md
docs/technical/zh/earth-frontend-context.md
docs/technical/en/earth-frontend-context.md
Document:
aiproviderand backend agent boundaries.- Earth AI command entry.
- Speech/ASR configuration.
- Wake-word privacy behavior.
- Agent run/evidence/proposal review.
- V1 capability boundary: Earth visualization actions only, no business data mutation.
Implementation Order
- Add backend Agent data models,
init_db()imports, and indexes. - Add Agent schemas and run CRUD API.
- Add tool registry, JSON tool-call parser, and policy skeleton.
- Add Earth entity query service.
- Add
agents.earth.commandprompt and backend command endpoint. - Add frontend Earth action executor.
- Merge AI command entry into Earth search panel.
- Add satellite and compute-center batch highlight support.
- Add Speech/ASR settings and transcription API.
- Add Earth microphone recording, wake-word local setting, and fallback behavior.
- Add Agent operations UI.
- Add proposal apply policy and low-risk configuration application.
- Add tests and documentation updates.
Assumptions
- Earth LLM v1 only executes visualization actions.
- Earth commands create agent runs but do not automatically create configuration proposals.
- Configuration proposals are applied only from the Agent operations UI.
- Wake word is a device-local preference in v1 and is stored in
localStorage. - ASR is API-first and Whisper-compatible by default; local Whisper is a later provider.
- Multi-role simulation only reserves schemas and prompt keys in v1.
- Palantir-style situational workflows should grow from the shared evidence, entity, action, and assessment model instead of a separate isolated system.