Skip to content

Mnemose Platform Re-Alignment & Execution Roadmap

Comprehensive architecture specification and execution roadmap for transforming Mnemose into a vendor-agnostic, edge-native platform for AI and agent-augmented work.


1. Vision & Core Value Proposition

Mnemose is a unified, cloud-native platform providing the essential substrate for human-AI collaborative workflows across software engineering, cloud operations, incident remediation, and automated evaluation.

MNEMOSE PLATFORM
┌─────────────────────────────────────────────────────────────────────────────┐
│ CONSOLIDATED WEB CONSOLE │
│ [Studio] [Models] [Agents & Fleets] [Evals] [Observe] │
├─────────────────┬─────────────────┬───────────────────┬─────────────────────┤
│ MODEL ROUTER │ AGENT PLATFORM │ AGENT SANDBOX │ EVALS & MONITOR │
│ │ │ PLATFORM │ │
│ • Unified proxy │ • Durable Object│ • MicroVM runners │ • LLM-as-a-judge │
│ • Fallback tiers│ runtimes (DO) │ • Cowrie deception│ • Benchmark suites │
│ • Token budget │ • Mnemosyne mem │ honeypots │ • OTEL distributed │
│ • Streaming │ • Tool registry │ • Facet policies │ trace attributes │
└─────────────────┴─────────────────┴───────────────────┴─────────────────────┘

2. Five Pillars Architecture Breakdown

Pillar 1: Model Router

Vendor-agnostic LLM gateway, fallback routing, quota/cost management, streaming.

  • Dynamic Provider Registry: Configurable OpenAI, Anthropic, Google Gemini, Cloudflare Workers AI, and custom OpenAI-compatible endpoints stored in D1 (llm_provider_configs).
  • Streaming Proxy & Normalization: Unified SSE/chunk stream protocol powered by Cloudflare Worker edge runtime with minimal time-to-first-token.
  • Failover & Cascade Logic: Automatic fallback to secondary/tertiary model tiers upon rate-limit (429), context length overflow, or provider downtime.
  • Tenant Token Accounting: Per-tenant and per-workspace token quota limits, real-time usage metering, and spend alerting.

Pillar 2: Agent Platform

Edge-native agent execution, memory, tool orchestration, human-in-the-loop.

  • Durable Object Execution Loop (AgentRuntimeDO): State machine handling agent reasoning steps (thought -> tool_call -> observation -> response) with zero cold starts and SQLite-backed durable hibernation.
  • Mnemosyne Memory Integration: Long-term episodic and semantic memory retrieval (vendor/mnemosyne-memory-system) injected per agent turn.
  • Tool Registry (@mnemose/tool-registry): Automatic tool schema discovery and invocation over MCP servers, internal GraphQL operations, and custom plugins.
  • Human-in-the-Loop (HITL) Interventions: Durable pause/resume workflows waiting on explicit human approval via Webhook, Console, or Slack/Email notifications.

Pillar 3: Agent Sandbox Platform

Safe code execution, tool containment, deception honeypots, kernel fleets.

  • Dual-Engine Execution:
    • Legitimate Workloads: MicroVM / remote container runners for secure multi-tenant code compilation, linting, testing, and script execution.
    • Adversarial / Honeypot Defense: @mnemose/deception-engine infinite-regression sandbox trapping untrusted agents or adversarial code in simulated environments.
  • Facet Security Controls: Network egress allowlists, file system mounts, redaction filters, and ephemeral volume lifecycle management (@mnemose/sandbox-runtime).

Pillar 4: Agent Monitoring Platform

Token spend observability, execution tracing, audit trails, telemetry.

  • Agent-Centric Distributed Tracing: OpenTelemetry instrumentation capturing exact prompt tokens, completion tokens, model latency, tool execution time, and error rates.
  • Immutable Audit Trail: Append-only event store (domain_events) documenting every state change, command dispatched, and tool executed.
  • Execution Replay & Transcript Inspector: Interactive step-by-step transcript inspection in the web console for live and historical agent runs.

Pillar 5: Model / Agent Evaluation Platform

Task benchmark suites, persona regression testing, output evaluation.

  • Automated Benchmark Runner: Batch execution engine running parameterized task prompts across multiple model configurations.
  • Multi-Method Scoring:
    • Deterministic assertions (regex, exit code, test pass/fail).
    • LLM-as-a-judge evaluation with customizable scoring rubrics.
    • Cost vs quality Pareto frontier analysis.
  • Model Leaderboard: Tenant-scoped evaluation dashboard tracking model performance trends over time.

3. Four-Phase Phased Execution Roadmap

PHASE 1 (Core Engine) PHASE 2 (Sandboxes & Fleet) PHASE 3 (Monitoring & UX) PHASE 4 (Eval Engine)
┌────────────────────────┐ ┌─────────────────────────┐ ┌─────────────────────────┐ ┌─────────────────────────┐
│ • AgentRuntimeDO Port │ │ • Unified Sandbox Mgr │ │ • Console UX Pivot │ │ • Benchmark Dataset Reg │
│ • Model Router Gateway │ │ • Remote MicroVM Runner │ │ • Transcript Inspector │ │ • LLM-Judge Scorer │
│ • Tool Registry Hook │ │ • Deception Facet Wire │ │ • OTEL Token Attributes │ │ • Leaderboard Dash │
└────────────────────────┘ └─────────────────────────┘ └─────────────────────────┘ └─────────────────────────┘

Phase 1: Core Engine Foundation

  • Refactor AgentRuntimeDO to run full agent execution loop (porting logic from legacy services/agent).
  • Wire @mnemose/llm-provider and @mnemose/tool-registry into AgentRuntimeDO.
  • Implement edge streaming Model Router endpoint (/v1/chat/completions) in Gateway Worker with fallback handling.

Phase 2: Sandbox & Fleet Consolidation

  • Build unified SandboxManager orchestrating microVM environments and @mnemose/deception-engine.
  • Connect packages/sandbox-execution and packages/sandbox-runtime to the gateway GraphQL layer.
  • Enable dynamic file injection and artifact retrieval across active sandboxes.

Phase 3: Observability & Console UX Pivot

  • Restructure Web Console (apps/console) into 5 core workspaces: Studio, Models, Fleets/Sandboxes, Telemetry, and Evals.
  • Build interactive Agent Transcript Inspector with step-by-step replay and token spend breakdown.
  • Export LLM metrics and tool execution spans to OpenTelemetry collectors.

Phase 4: Model / Agent Evaluation Platform

  • Create Eval Runner service with support for synthetic datasets and regression suites.
  • Implement deterministic and LLM-as-a-judge scoring pipelines.
  • Add Model/Agent Leaderboard with cost-performance comparison visualizations.