Mnemose Platform Re-Alignment & Execution Roadmap
Comprehensive architecture specification and execution roadmap for transforming Mnemose into a vendor-agnostic, edge-native platform for AI and agent-augmented work.
1. Vision & Core Value Proposition
Mnemose is a unified, cloud-native platform providing the essential substrate for human-AI collaborative workflows across software engineering, cloud operations, incident remediation, and automated evaluation.
MNEMOSE PLATFORM ┌─────────────────────────────────────────────────────────────────────────────┐ │ CONSOLIDATED WEB CONSOLE │ │ [Studio] [Models] [Agents & Fleets] [Evals] [Observe] │ ├─────────────────┬─────────────────┬───────────────────┬─────────────────────┤ │ MODEL ROUTER │ AGENT PLATFORM │ AGENT SANDBOX │ EVALS & MONITOR │ │ │ │ PLATFORM │ │ │ • Unified proxy │ • Durable Object│ • MicroVM runners │ • LLM-as-a-judge │ │ • Fallback tiers│ runtimes (DO) │ • Cowrie deception│ • Benchmark suites │ │ • Token budget │ • Mnemosyne mem │ honeypots │ • OTEL distributed │ │ • Streaming │ • Tool registry │ • Facet policies │ trace attributes │ └─────────────────┴─────────────────┴───────────────────┴─────────────────────┘2. Five Pillars Architecture Breakdown
Pillar 1: Model Router
Vendor-agnostic LLM gateway, fallback routing, quota/cost management, streaming.
- Dynamic Provider Registry: Configurable OpenAI, Anthropic, Google Gemini, Cloudflare Workers AI, and custom OpenAI-compatible endpoints stored in D1 (
llm_provider_configs). - Streaming Proxy & Normalization: Unified SSE/chunk stream protocol powered by Cloudflare Worker edge runtime with minimal time-to-first-token.
- Failover & Cascade Logic: Automatic fallback to secondary/tertiary model tiers upon rate-limit (
429), context length overflow, or provider downtime. - Tenant Token Accounting: Per-tenant and per-workspace token quota limits, real-time usage metering, and spend alerting.
Pillar 2: Agent Platform
Edge-native agent execution, memory, tool orchestration, human-in-the-loop.
- Durable Object Execution Loop (
AgentRuntimeDO): State machine handling agent reasoning steps (thought -> tool_call -> observation -> response) with zero cold starts and SQLite-backed durable hibernation. - Mnemosyne Memory Integration: Long-term episodic and semantic memory retrieval (
vendor/mnemosyne-memory-system) injected per agent turn. - Tool Registry (
@mnemose/tool-registry): Automatic tool schema discovery and invocation over MCP servers, internal GraphQL operations, and custom plugins. - Human-in-the-Loop (HITL) Interventions: Durable pause/resume workflows waiting on explicit human approval via Webhook, Console, or Slack/Email notifications.
Pillar 3: Agent Sandbox Platform
Safe code execution, tool containment, deception honeypots, kernel fleets.
- Dual-Engine Execution:
- Legitimate Workloads: MicroVM / remote container runners for secure multi-tenant code compilation, linting, testing, and script execution.
- Adversarial / Honeypot Defense:
@mnemose/deception-engineinfinite-regression sandbox trapping untrusted agents or adversarial code in simulated environments.
- Facet Security Controls: Network egress allowlists, file system mounts, redaction filters, and ephemeral volume lifecycle management (
@mnemose/sandbox-runtime).
Pillar 4: Agent Monitoring Platform
Token spend observability, execution tracing, audit trails, telemetry.
- Agent-Centric Distributed Tracing: OpenTelemetry instrumentation capturing exact prompt tokens, completion tokens, model latency, tool execution time, and error rates.
- Immutable Audit Trail: Append-only event store (
domain_events) documenting every state change, command dispatched, and tool executed. - Execution Replay & Transcript Inspector: Interactive step-by-step transcript inspection in the web console for live and historical agent runs.
Pillar 5: Model / Agent Evaluation Platform
Task benchmark suites, persona regression testing, output evaluation.
- Automated Benchmark Runner: Batch execution engine running parameterized task prompts across multiple model configurations.
- Multi-Method Scoring:
- Deterministic assertions (regex, exit code, test pass/fail).
- LLM-as-a-judge evaluation with customizable scoring rubrics.
- Cost vs quality Pareto frontier analysis.
- Model Leaderboard: Tenant-scoped evaluation dashboard tracking model performance trends over time.
3. Four-Phase Phased Execution Roadmap
PHASE 1 (Core Engine) PHASE 2 (Sandboxes & Fleet) PHASE 3 (Monitoring & UX) PHASE 4 (Eval Engine)┌────────────────────────┐ ┌─────────────────────────┐ ┌─────────────────────────┐ ┌─────────────────────────┐│ • AgentRuntimeDO Port │ │ • Unified Sandbox Mgr │ │ • Console UX Pivot │ │ • Benchmark Dataset Reg ││ • Model Router Gateway │ │ • Remote MicroVM Runner │ │ • Transcript Inspector │ │ • LLM-Judge Scorer ││ • Tool Registry Hook │ │ • Deception Facet Wire │ │ • OTEL Token Attributes │ │ • Leaderboard Dash │└────────────────────────┘ └─────────────────────────┘ └─────────────────────────┘ └─────────────────────────┘Phase 1: Core Engine Foundation
- Refactor
AgentRuntimeDOto run full agent execution loop (porting logic from legacyservices/agent). - Wire
@mnemose/llm-providerand@mnemose/tool-registryintoAgentRuntimeDO. - Implement edge streaming Model Router endpoint (
/v1/chat/completions) in Gateway Worker with fallback handling.
Phase 2: Sandbox & Fleet Consolidation
- Build unified
SandboxManagerorchestrating microVM environments and@mnemose/deception-engine. - Connect
packages/sandbox-executionandpackages/sandbox-runtimeto the gateway GraphQL layer. - Enable dynamic file injection and artifact retrieval across active sandboxes.
Phase 3: Observability & Console UX Pivot
- Restructure Web Console (
apps/console) into 5 core workspaces: Studio, Models, Fleets/Sandboxes, Telemetry, and Evals. - Build interactive Agent Transcript Inspector with step-by-step replay and token spend breakdown.
- Export LLM metrics and tool execution spans to OpenTelemetry collectors.
Phase 4: Model / Agent Evaluation Platform
- Create Eval Runner service with support for synthetic datasets and regression suites.
- Implement deterministic and LLM-as-a-judge scoring pipelines.
- Add Model/Agent Leaderboard with cost-performance comparison visualizations.