Mnemose Platform — Functional State Audit (2026-08-19)
Scope: Complete evaluation of what is genuinely wired end-to-end versus what is
mocked, stubbed, or backed by ephemeral state. Baseline: main @ e6df91e.
Verdict: The legacy console surface is real. The five-pillar surface added in the
recent modernization waves is UI without application: rendered screens, hardcoded
sample data, zero API calls. Additionally, three backend gaps make even the existing
APIs unable to support a real UI (missing D1 migration, in-memory-only persistence,
missing honeypot alert pipeline).
1. Verdict matrix
Legend: ✅ REAL (wired end-to-end) · 🟡 PARTIAL (works but ephemeral/incomplete) · ❌ MOCK (fake data, no calls)
Console views
| Area | View | State | Evidence |
|---|---|---|---|
| Dashboard | dashboard.tsx | ✅ | GraphQL DASHBOARD_QUERY |
| Fleet / Kernels | Fleet/* | ✅ | GraphQL queries/mutations + subscriptions |
| Tenants / Users / IAM | tenants, users, iam-bindings, iam-escalations | ✅ | GraphQL, D1-backed |
| Cloud projects / Ownership | cloud-projects, ownership | ✅ | GraphQL |
| Provisioning / Remediation / Reports / Resources | provisioning, remediation, reports, resources | ✅ | GraphQL |
| Encryption keys | encryption-keys.tsx | ✅ | GraphQL |
| SEO suite | seo/* | ✅ | GraphQL, full CRUD |
| Login / Settings | login, settings | ✅ | Auth flow real |
| Sandboxes — list/table/policy | sandboxes/index, SandboxTable, NetworkPolicyInspector | ✅ | GraphQL SANDBOXES_QUERY + DO mutations |
| Sandboxes — terminal | sandboxes/SandboxTerminal | ❌ MOCK | 0 API calls; backend executeCommand mutation + /sandbox/* proxy exist |
| Sandboxes — file/artifact explorer | sandboxes/FilesystemArtifactExplorer | ❌ MOCK | 0 API calls; artifact download fabricates a blob (mockContent) instead of fetching bytes; REST /v1/sandboxes/:id/files + /artifacts exist |
| Studio — index + HITL banner | studio/index, HitlApprovalBanner | 🟡 | Queries wired; HITL reuses IAM-escalation/kernel-session deny mutations (real, but not the AgentRuntimeDO approval queue) |
| Studio — workbench / live steps / transcript | studio/AgentWorkbench, LiveStepExecution, TranscriptInspector | ❌ MOCK | 0 API calls; backend /agent/* proxy + chat schema exist |
| Models — router manager / provider grid / fallback tiers / spend panel | models/* | ❌ MOCK | 0 API calls; llmProviderConfigs GraphQL + /v1/models REST exist |
| Store — browse / detail / installed / permission matrix | store/* | ❌ MOCK | 0 API calls; full tools/apps GraphQL + REST exist (but backend broken — see §2.1) |
| Evaluations — all views | evaluations/* | ❌ MOCK | 0 API calls; SAMPLE_RUNS, SAMPLE_JUDGE_RESULTS, SAMPLE_DIFF_TESTS; backend runEvaluation mutation + REST exist |
| Observability — trace timeline | observability/OtelTraceTimeline | ❌ MOCK | SAMPLE_TRACES; REST /v1/observability/traces exists but is ephemeral (§2.2) |
| Observability — token charts | observability/TokenUsageCharts | ❌ MOCK | 0 API calls; REST /v1/observability/metrics/tokens exists but ephemeral |
| Observability — honeypot alerts | observability/CanaryHoneypotAlerts | ❌ MOCK | SAMPLE_ALERTS; no backend exists at all (§2.3) |
Score: 10 pillar views fully mocked, 2 partially wired, ~20 legacy views real.
Backend (gateway Worker — deployed, health 200, auth enforcing)
| Surface | State | Notes |
|---|---|---|
| GraphQL core (tenants, users, IAM, kernels, sandboxes, SEO, chat, infra) | ✅ | D1-backed, migrated |
GraphQL tools/apps (searchTools, availableApps, installApp, …) | 🟡 | Code real, storage tables never migrated (§2.1) |
GraphQL evaluations (runEvaluation, evaluationRuns) | 🟡 | Executes real benchmarks via ModelRouter; results in-memory only (§2.2) |
GraphQL llm-providers (llmProviderConfigs, persona overrides) | ✅ | D1-backed |
REST /v1/models, /v1/chat/completions | ✅ | Real ModelRouter + provider cascade |
REST /v1/sandboxes/:id/files, /artifacts | 🟡 | Real reads; no artifact-content download route |
REST /v1/observability/* | 🟡 | Real code, in-memory only (§2.2) |
DO proxies /agent/*, /sandbox/*, /api/kernels/* | ✅ | AgentRuntimeDO / SandboxDO / KernelBrokerDO real |
| Honeypot alert pipeline | ❌ | Does not exist anywhere in gateway |
| Runtime app-registry loading | ❌ | schema/index.ts TODO: “Wave 5”, never done |
2. Backend gaps that block a real UI
2.1 CRITICAL — Missing D1 migration (Tool/App Store is dead on arrival)
Drizzle defines tool_catalog_entries, tool_permissions, tool_audit_entries
(packages/db/src/schema/tools.ts) but packages/db/src/migrations/ contains only
0000 and 0001. Every /v1/tools*, /v1/apps* REST call and every tools/apps
GraphQL resolver will throw “no such table” against live D1. The entire pillar-3
backend is paper.
2.2 CRITICAL — Ephemeral state masquerading as product features
All in-memory Maps on Cloudflare Workers die on isolate eviction (minutes, or any
deploy). These currently hold:
historicalEvaluationRuns(router/evaluations.ts:22) — all benchmark results.AgentTracer.recordedSpans(telemetry/agent-tracer.ts:120) — all OTEL spans.TokenSpendAccumulatormaps (telemetry/token-spend.ts:104-105) — all token/cost metrics.
Consequence: even if the console were wired today, Observability and Evaluations would show near-empty results most of the time.
2.3 Honeypot/deception alerts — no backend
CanaryHoneypotAlerts renders fabricated alerts. The deception engine package exists,
but nothing transports its tripwire/escape events to gateway storage or API. Zero
tables, zero routes, zero GraphQL fields.
2.4 Smaller gaps
- No artifact-content download route (explorer has no real bytes to fetch).
- No provider health/status query (ProviderStatusGrid target).
- Runtime app-registry module loading in gateway: TODO since inception.
- Studio HITL banner reuses IAM/kernel deny mutations instead of the AgentRuntimeDO
approval queue (
resolveApprovalexists on the DO but has no console path).
3. Plan — 100% functionally wired, zero mocks
Rule: backend durability first, then console wiring, then a CI gate that makes mocks impossible to reintroduce. Each wave uses the established stripe mechanics (disjoint file ownership, push-per-commit, independent QA wave).
Wave F-0 — Backend durability & completeness (4 stripes, parallel)
| Stripe | Deliverable | Files owned |
|---|---|---|
| F-0-A | Migration 0002_tool_registry.sql (3 tables + indexes); verify dev D1 apply; tools/apps REST+GraphQL live-verified | packages/db/src/migrations/*, packages/db/src/schema/tools.ts fixes only |
| F-0-B | Migration 0003_evaluation_runs.sql (evaluation_runs, evaluation_cases incl. judge output, per-case transcripts); replace historicalEvaluationRuns Map with D1 store (write-through, Map optional cache) | router/evaluations.ts, schema/evaluations.ts, migration |
| F-0-C | Observability persistence: otel_spans + token_spend_metrics D1 tables (migration in same stripe), write-through batching in AgentTracer/TokenSpendAccumulator, REST reads from D1, TTL/retention policy | packages/telemetry/*, router/observability.ts, migration |
| F-0-D | Honeypot alert pipeline: honeypot_alerts table, alert write path from SandboxDO deception volume tripwires, REST GET /v1/observability/alerts + GraphQL honeypotAlerts; artifact-content route GET /v1/sandboxes/:id/artifacts/:artifactId/content; provider status query (from router telemetry counters) | do/sandbox-do.ts alert emit, router/observability.ts, router/sandbox-files.ts, schema/llm-providers.ts |
Exit criteria: every new REST route returns real persisted data on the dev Worker across a deploy boundary (data survives redeploy).
Wave F-1 — Console wiring (6 stripes, parallel, per pillar)
Every stripe deletes its SAMPLE_* constants first, then wires:
| Stripe | Pillar | Wiring |
|---|---|---|
| F-1-A | Store | StoreBrowseView→availableApps, StoreAppDetailView→app+appHealth, InstalledAppsView→installedApps+installApp/uninstallApp/enableApp/disableApp, ToolPermissionMatrix→getToolPermissions/setToolPermission/deleteToolPermission/evaluateToolPermission |
| F-1-B | Models | ModelRouterManager→llmProviderConfigs+setLlmProviderConfig, FallbackTierList→tier config mutation, ProviderStatusGrid→provider status query (F-0-D), SpendTelemetryPanel→/v1/observability/metrics/tokens |
| F-1-C | Evaluations | Launcher→runEvaluation mutation, RunsTable→evaluationRuns (poll 5s while running), Leaderboard+Scorecard+Diff→evaluationRun detail incl. case-level results |
| F-1-D | Observability | OtelTraceTimeline→/v1/observability/traces, TokenUsageCharts→metrics endpoint, CanaryHoneypotAlerts→alerts query (F-0-D) |
| F-1-E | Sandboxes | SandboxTerminal→executeCommand mutation (poll-based result fetch), FilesystemArtifactExplorer→files REST + real artifact-content download (kill mockContent) |
| F-1-F | Studio | AgentWorkbench→createChatSession/appendChatMessage + /agent/* session endpoints, LiveStepExecution→runtime step polling/SSE, TranscriptInspector→session messages, HITL banner→AgentRuntimeDO approval queue (approve/deny/feedback) |
UI rule for all stripes: empty states and loading states replace fabricated data; never render a value the API did not return.
Wave F-2 — Governance & QA (2 stripes)
| Stripe | Deliverable |
|---|---|
| F-2-A | Anti-mock CI gate: ESLint rule (or ci.yml grep gate) failing on SAMPLE_[A-Z]/MOCK_[A-Z] constants and hardcoded object-array useState seeds in apps/console/src/views; docs: per-view wiring status matrix appended to docs/architecture.md |
| F-2-B | Live smoke suite: console integration tests against the dev gateway (dev-auth bypass) hitting each pillar’s real query; QA pass over all F-1 stripes (independent model) |
Sequencing
F-0 (parallel) ──► F-1 (parallel, after F-0 merges) ──► F-2 (QA + gates)Estimated stripe size: F-0 stripes ~1 session each; F-1 stripes ~1 session each. Total: 12 stripes + QA.
4. Definition of done (platform-wide)
- Zero
SAMPLE_*/MOCK_*constants in console; CI gate enforces it. - Every console view issues real API calls; empty/loading/error states are honest.
- All product data survives Worker redeploy (D1 persistence verified across deploy).
- Live smoke suite green against dev gateway on every push.
docs/architecture.mdwiring matrix matches reality.