FallCore on-prem · sovereign
◊·κ=1 · the inversion

Anthropic charges per token.
FallCore is a flat fee that gets cheaper
as your bill collapses.

The on-prem brain that learns from your work. Anthropic-API-compatible local proxy + RAG + LoRA fine-tune loop. Burn frontier tokens today? In 12 months, most of those calls answer locally — and the model that ships at year 3 is yours, trained on your work.

Install on your infra How it works

how it worksSame UX. Opposite economics.

Every Fall* tool — and any app hitting Anthropic's API — points at FallCore instead. The cascade decides per-query: local first, frontier only when truly needed. Every frontier call logged for next week's fine-tune.

1
Customer queryApp hits POST /v1/messages — same Anthropic shape
~50ms
2
Local Ollama (Qwen2.5-32B / 72B)Runs on customer's GPU · data never leaves the network
2-8s
3
Confidence checkIf local ≥ threshold (default 0.6) → return locally · saved $0.03 · skip step 4
local stays · 75%+ of cases
4
Frontier fallthroughOnly when local genuinely insufficient → real Anthropic with customer's key
~25% initial · drops over time
5
Log every cascade-fire(prompt · local-answer · frontier-answer) → JSONL → next LoRA cycle
training data accrues
6
Weekly LoRA fine-tuneTrain base model on accumulated preference pairs · deploy adapter
local accuracy ↑
7
Frontier ratio dropsMonth 1: 25% frontier · Month 6: 8% · Month 12: 3%
frontier spend collapses

why this is differentThe inversion of every AI vendor's incentive.

Anthropic / OpenAIFallCore
Where the model runsTheir serversYour hardware
Where your data goesTo themStays on your network
Pricing modelPay per token (scales with usage)Flat annual fee (decoupled from usage)
Whose reviewer corrections train the modelTheirs (you gift them labels)Yours (the adapter is your IP)
Compliance posture (HIPAA / GDPR / SOX)Customer is liableData never moves → liability collapses
Vendor lock-inTotal · model owned by themNone · model + adapter on your disk
If they go out of businessYou stop workingNothing changes
Year-3 economicsBill grew with usageBill is 5% of year 1

tiersPick by hardware. Free trial. Always.

Tier = which model you're running. Lite for laptops and small teams. Pro for departmental GPUs. Sovereign for regulated industries running their own hardware. No paywall — open source, MIT.

·Lite

For laptops and small teams. Runs on consumer GPUs (or CPU only at low volume).

  • Llama-3.1-8B
  • 8GB VRAM
  • RAG over your docs
  • Monthly fine-tune cadence
  • ~40% frontier-call reduction

◯Sovereign

For regulated industries · finance, healthcare, legal. Server-class hardware.

  • Qwen2.5-72B
  • 48GB VRAM
  • Bi-weekly LoRA
  • BSV audit anchor
  • ~95% frontier-call reduction

⬡Enterprise

Multi-region, custom certifications, white-label.

  • Multi-region failover
  • Custom model selection
  • Bespoke certifications
  • White-label / OEM option
  • ~99% frontier-call reduction
Free during launch. The source is MIT-licensed. The factory mints branded stacks at FallCore Factory — no card, no commitment. When commercial managed-tier pricing is set, it lives here. Until then: clone, run, ship.

installOne command on your infra.

Pick a tier, set the .env, docker compose up. The proxy listens on port 11434 (same as Ollama by default). Point your existing Anthropic clients at http://your-host:11434 and they Just Work.

# 1 · clone
git clone https://github.com/sjgant80-hub/fallcore
cd fallcore

# 2 · configure
cp .env.example .env
#  edit OLLAMA_MODEL, ANTHROPIC_FALLBACK_KEY, CONFIDENCE_THRESHOLD

# 3 · bring it up
docker compose up -d

# 4 · pull the model (one time)
docker compose exec ollama ollama pull qwen2.5:32b

# 5 · point your Anthropic clients at it
ANTHROPIC_BASE_URL=http://your-host:11434
#  done. existing apps now use local-first cascade.

# 6 · after a week, run the eval to see your ROI
node eval/replay.js --days 7

# 7 · extract preference pairs for fine-tune
node train/extract-preferences.js

what's includedOpen source, MIT licenced. Yours to inspect, modify, fork.

⚡Proxy server

Express server, single file. Anthropic-API-shape. OpenAI-shape bonus. Cascade logic + confidence scoring + frontier fallthrough + JSONL logging. No dependencies beyond express + cors.

🐳Docker stack

Ollama + Qdrant + proxy. One docker compose up. GPU passthrough commented in for NVIDIA. Persistent volumes for models, vectors, and logs.

📊Eval harness

Replay last 30 days of frontier calls through local · score equivalence · output ROI report. Tells you when to lower the confidence threshold (more local · less spend).

🎯LoRA pipeline

Extract preference pairs from proxy logs + Fall* reviewer corrections → standard JSONL → train with axolotl / unsloth / TRL DPO. Deploy adapter into Ollama. Repeat weekly.

🌐Mesh interop

Pairs with the Fall* estate: FallForce CRM · GymOps · Apex Procurement · FallAccount · 30+ sovereign tools all share BroadcastChannel('fallmesh') + fall-kcc + Konomi licence signing.

⛓BSV audit anchor

Sovereign tier: every fine-tune cycle, every reviewer correction, every model deployment can be anchored on Bitcoin SV via OnlyBrains. Court-defensible audit trail. Nothing phones home.

the v18 readThis isn't a product. It's the proof.

Anthropic charges by the token. OpenAI charges by the token. Google charges by the token. Every vendor's incentive is for you to use more. None of them want you off their platform.

FallCore inverts it. We sell the layer that reduces your dependence on those vendors over time. Our growth = your bill shrinking. The model that ships at the end of year 3 is yours — trained on your work, fine-tuned by your reviewers, sitting on your hardware. You can fire us and it still works.

This is the first instance of cognitive sovereignty becoming infrastructure. Built in 18 months on no capital by one person and his AI in a gaming café in Pattaya. Part of the Fall* estate · 30+ sovereign tools, all open source, all aligned with the same Konomi / KCC architecture.

Read the source Talk to us