The on-prem brain that learns from your work. Anthropic-API-compatible local proxy + RAG + LoRA fine-tune loop. Burn frontier tokens today? In 12 months, most of those calls answer locally — and the model that ships at year 3 is yours, trained on your work.
Every Fall* tool — and any app hitting Anthropic's API — points at FallCore instead. The cascade decides per-query: local first, frontier only when truly needed. Every frontier call logged for next week's fine-tune.
POST /v1/messages — same Anthropic shape| Anthropic / OpenAI | FallCore | |
|---|---|---|
| Where the model runs | Their servers | Your hardware |
| Where your data goes | To them | Stays on your network |
| Pricing model | Pay per token (scales with usage) | Flat annual fee (decoupled from usage) |
| Whose reviewer corrections train the model | Theirs (you gift them labels) | Yours (the adapter is your IP) |
| Compliance posture (HIPAA / GDPR / SOX) | Customer is liable | Data never moves → liability collapses |
| Vendor lock-in | Total · model owned by them | None · model + adapter on your disk |
| If they go out of business | You stop working | Nothing changes |
| Year-3 economics | Bill grew with usage | Bill is 5% of year 1 |
Tier = which model you're running. Lite for laptops and small teams. Pro for departmental GPUs. Sovereign for regulated industries running their own hardware. No paywall — open source, MIT.
For laptops and small teams. Runs on consumer GPUs (or CPU only at low volume).
For departmental deployments. Single workstation-class GPU.
For regulated industries · finance, healthcare, legal. Server-class hardware.
Multi-region, custom certifications, white-label.
Pick a tier, set the .env, docker compose up. The proxy listens on port 11434 (same as Ollama by default). Point your existing Anthropic clients at http://your-host:11434 and they Just Work.
# 1 · clone git clone https://github.com/sjgant80-hub/fallcore cd fallcore # 2 · configure cp .env.example .env # edit OLLAMA_MODEL, ANTHROPIC_FALLBACK_KEY, CONFIDENCE_THRESHOLD # 3 · bring it up docker compose up -d # 4 · pull the model (one time) docker compose exec ollama ollama pull qwen2.5:32b # 5 · point your Anthropic clients at it ANTHROPIC_BASE_URL=http://your-host:11434 # done. existing apps now use local-first cascade. # 6 · after a week, run the eval to see your ROI node eval/replay.js --days 7 # 7 · extract preference pairs for fine-tune node train/extract-preferences.js
Express server, single file. Anthropic-API-shape. OpenAI-shape bonus. Cascade logic + confidence scoring + frontier fallthrough + JSONL logging. No dependencies beyond express + cors.
Ollama + Qdrant + proxy. One docker compose up. GPU passthrough commented in for NVIDIA. Persistent volumes for models, vectors, and logs.
Replay last 30 days of frontier calls through local · score equivalence · output ROI report. Tells you when to lower the confidence threshold (more local · less spend).
Extract preference pairs from proxy logs + Fall* reviewer corrections → standard JSONL → train with axolotl / unsloth / TRL DPO. Deploy adapter into Ollama. Repeat weekly.
Pairs with the Fall* estate: FallForce CRM · GymOps · Apex Procurement · FallAccount · 30+ sovereign tools all share BroadcastChannel('fallmesh') + fall-kcc + Konomi licence signing.
Sovereign tier: every fine-tune cycle, every reviewer correction, every model deployment can be anchored on Bitcoin SV via OnlyBrains. Court-defensible audit trail. Nothing phones home.
Anthropic charges by the token. OpenAI charges by the token. Google charges by the token. Every vendor's incentive is for you to use more. None of them want you off their platform.
FallCore inverts it. We sell the layer that reduces your dependence on those vendors over time. Our growth = your bill shrinking. The model that ships at the end of year 3 is yours — trained on your work, fine-tuned by your reviewers, sitting on your hardware. You can fire us and it still works.
This is the first instance of cognitive sovereignty becoming infrastructure. Built in 18 months on no capital by one person and his AI in a gaming café in Pattaya. Part of the Fall* estate · 30+ sovereign tools, all open source, all aligned with the same Konomi / KCC architecture.