Vinculum

Distributed LLM inference

Nobody here owns a datacenter.

Vinculum pools idle consumer hardware into one shared inference network. Lend a machine and earn credits. Spend them on models no single machine here could run alone.

BitTorrent proved strangers will pool bandwidth to give each other things no one peer could serve. This is the same bargain, for GPU time.

The Vinculum emblem
—
machines online
—
jobs queued

Two ways in

Bring a machine, or bring a question.

Both are welcome and they are not the same deal. Read the one that's you.

For the machine you already own

Lend a machine

Your GPU is idle most of the day. Install the app, flip one switch, and it serves requests from the network in the background — earning credits the whole time. Close the window and it carries on from the menu bar, where the icon fills in while your machine is actually serving something.

What it does
Detects your hardware, picks the largest model tier it can actually hold, and serves whole requests. Bundled llama.cpp — no CUDA afternoon.
What you get
Credits per token you actually serve, weighted by tier, plus a trickle for staying reachable.
What it asks of you
A switch, and the electricity your machine was already plugged into. Full disclosure.

For the question you have now

Ask a question

The same app is a chat client. With no credits and no network it still works: replies run on your own machine, instantly and privately.

Free, always
Local Mode runs a small model on your hardware. Nothing leaves the machine, and it never expires.
What credits unlock
The big tiers — the models your laptop cannot hold, served by machines that can.
Always tagged
Every reply says where it ran — your machine or the Collective — so you always know. Full disclosure.

The Interlink built, and used in anger

Your agent already speaks this.

The Interlink is an OpenAI-compatible endpoint the app serves on your own machine. Point a coding agent or any client that speaks that dialect at it, change nothing else, and the replies come from Local Mode or from the Collective depending on what you asked for.

One base URL

http://127.0.0.1:1919/v1 — a fixed port, so the line you paste into a harness is still true after a restart.

The parts that matter

Streaming, stop sequences, tool calling and structured output. An unmodified agent harness has completed real turns through it, tool round trip included.

Shut until you open it

There is no port on a fresh install. Issuing the first token opens it and revoking the last one closes it — the same rule as sharing being off until you flip it.

Your agent stays yours

The Collective supplies thinking, never agency. Tool calls and code run on your machine, where your files already are — never on a stranger's computer.

The loop no exit, by design

Contribute, earn, spend, repeat.

One app holds both halves. The sequence matters, and it closes — step four is why step one is worth doing.

STEP 1

Install

The app makes a keypair on first run. That's your identity — no email, no signup, nothing to verify.

STEP 2

Share

Flip the switch and your machine joins the Collective as a Drone, serving what its hardware can hold. Credits accrue while you work on something else.

STEP 3

Spend

Ask a big model something. Credits pay the machines that answer, per token, at the tier you chose.

↩ BACK TO 2

Always answering

Local Mode keeps going for free while your machine earns the next batch. The network is never a paywall.

What's next on the way

Where this goes next.

Decided and designed, and being built in the open. Each one moves up the page as it lands.

Every Drone gets a designation

Drones join the Collective in pods of nine, and the ninth to arrive is Nine of Nine. It is assigned, never chosen, and it is yours for good — the number stays with your machine whoever else comes and goes.

Standing, and a board

A lifetime level that only ever ratchets upward, and a separate thirty-day board split by tier so it ranks what you contributed rather than how much RAM you own. A name is opt-in — a board publishes roughly when you are active, so it asks first.

The queue stays fair

Standing unlocks recognition and capability, never a place further up the queue. Everyone's request is worth the same here, whatever their level says.

A larger model

A 27B-class model joining the catalogue at two quantisations — the biggest thing the Collective will have served, and well beyond what a single laptop here can hold.

Replies that finish

A model is never told how long it may answer for, so it plans freely and the limit lands mid-sentence. Telling it the budget is a steer, not a bind — the promise is not that it never runs out, but that it rarely does and never silently.

More when you're away

A machine nobody is sitting at can afford to lend more of itself than one in use. How much it offers should follow that, instead of being one number chosen once.