Distributed LLM inference
Nobody here owns a datacenter.
Vinculum pools idle consumer hardware into one shared inference network. Lend a machine and earn credits. Spend them on models no single machine here could run alone.
BitTorrent proved strangers will pool bandwidth to give each other things no one peer could serve. This is the same bargain, for GPU time.
Two ways in
Bring a machine, or bring a question.
Both are welcome and they are not the same deal. Read the one that's you.
For the machine you already own
Lend a machine
Your GPU is idle most of the day. Install the app, flip one switch, and it serves requests from the network in the background — earning credits the whole time. Close the window and it carries on from the menu bar, where the icon fills in while your machine is actually serving something.
- What it does
- Detects your hardware, picks the largest model tier it can actually hold, and serves whole requests. Bundled llama.cpp — no CUDA afternoon.
- What you get
- Credits per token you actually serve, weighted by tier, plus a trickle for staying reachable.
- What it asks of you
- A switch, and the electricity your machine was already plugged into. Full disclosure.
For the question you have now
Ask a question
The same app is a chat client. With no credits and no network it still works: replies run on your own machine, instantly and privately.
- Free, always
- Local Mode runs a small model on your hardware. Nothing leaves the machine, and it never expires.
- What credits unlock
- The big tiers — the models your laptop cannot hold, served by machines that can.
- Always tagged
- Every reply says where it ran — your machine or the Collective — so you always know. Full disclosure.
The Interlink built, and used in anger
Your agent already speaks this.
The Interlink is an OpenAI-compatible endpoint the app serves on your own machine. Point a coding agent or any client that speaks that dialect at it, change nothing else, and the replies come from Local Mode or from the Collective depending on what you asked for.
One base URL
http://127.0.0.1:1919/v1 — a fixed port, so the line you paste into a
harness is still true after a restart.
The parts that matter
Streaming, stop sequences, tool calling and structured output. An unmodified agent harness has completed real turns through it, tool round trip included.
Shut until you open it
There is no port on a fresh install. Issuing the first token opens it and revoking the last one closes it — the same rule as sharing being off until you flip it.
The loop no exit, by design
Contribute, earn, spend, repeat.
One app holds both halves. The sequence matters, and it closes — step four is why step one is worth doing.
Install
The app makes a keypair on first run. That's your identity — no email, no signup, nothing to verify.
Share
Flip the switch and your machine joins the Collective as a Drone, serving what its hardware can hold. Credits accrue while you work on something else.
Spend
Ask a big model something. Credits pay the machines that answer, per token, at the tier you chose.
Always answering
Local Mode keeps going for free while your machine earns the next batch. The network is never a paywall.
What's next on the way
Where this goes next.
Decided and designed, and being built in the open. Each one moves up the page as it lands.
Every Drone gets a designation
Drones join the Collective in pods of nine, and the ninth to arrive is Nine of Nine. It is assigned, never chosen, and it is yours for good — the number stays with your machine whoever else comes and goes.
Standing, and a board
A lifetime level that only ever ratchets upward, and a separate thirty-day board split by tier so it ranks what you contributed rather than how much RAM you own. A name is opt-in — a board publishes roughly when you are active, so it asks first.
The queue stays fair
Standing unlocks recognition and capability, never a place further up the queue. Everyone's request is worth the same here, whatever their level says.
Replies that finish
A model is never told how long it may answer for, so it plans freely and the limit lands mid-sentence. Telling it the budget is a steer, not a bind — the promise is not that it never runs out, but that it rarely does and never silently.
More when you're away
A machine nobody is sitting at can afford to lend more of itself than one in use. How much it offers should follow that, instead of being one number chosen once.
Almost ready
Be first through the door.
Builds for macOS and Windows are being signed and notarised right now — an app that asks to use your hardware should arrive properly introduced, not with the operating system warning you about it. Leave an address and you'll get one email the day it's ready.
One email, at launch — an address and a date, nothing else, and nothing else ever sent to it. Full disclosure.
Vinculum