What this is: a sovereign single-file WebLLM host. Pair it with the bundled MCP server (
mcp-server/server.mjs) and Claude Code can call this tab via local_complete instead of spawning Sonnet subagents — burning zero Claude API tokens. The model runs in your browser's WebGPU, weights cached in IndexedDB after first download (~1.2GB Llama 3.2 1B / ~2.4GB Phi-3-mini). Keep this tab open while coding.
0
total calls
0
tokens generated
TBA
est sonnet cost saved
—
avg latency ms
◐ Model + Bridge
awaiting load…
◯ State
model—
webgpuprobing…
bridgedisconnected
queue depth0
active callidle
last call ts—
errors (session)0
konomi prime1009
cacheindexeddb
△ Quick test (no bridge needed)
—
◯ Call log
—no calls yet · open this tab + run the MCP server + tell Claude to use local_complete