MDMA

CrystalMeet got rejected, so this time I am doing science.
MDMA stands for Model Dispatch and Memory Allocator. The idea is a cluster of inference engines that you spawn at will on the machines you already own, serving each other without a cloud in between. The repository went public on October 10.
The name is the plan, one half at a time.
Six processes, one card¶
Until last month my workstation served models the usual way. Each model was a llama-server process under systemd, listening on its own port. Six of them shared one GPU, and none of them knew what the other five had loaded.
In late September the GLM-4 32B unit got stuck in an OOM restart loop. It ran out of memory, and systemd restarted it, over and over.
llama.cpp ran each model well. The card as a whole belonged to nobody.
The memory allocator¶
MDMA serves every model from one process, and that process keeps the books. Each model and each request reserves its memory before allocating any of it. A request that can’t fit gets refused up front with a clear error, and the server stays up. A model that doesn’t fit at startup gets skipped, and the rest are served.
Two names that point at the same model file share a single copy. On a Mac, where the GPU and the CPU share memory, that memory is counted once.
On a laptop it watches the power supply. Unplugged, it shrinks its limits, and at 10% battery it stops taking new work.
The first version sat in front of the existing llama-server units, so I could move models over one at a time. The rule since then is simple: any model that can run on MDMA goes on MDMA.
The machines you already own¶
Most people who run models at home own more than one machine: a desktop, a laptop, a Mac, maybe a mini PC in a closet. Each has a different GPU or none at all, and each GPU vendor wants its own stack. My own list goes further: I have models running on every PlayStation from the 3 to the 5 and on old computers, with kernels I will open-source later.
MDMA has six backends so every one of those machines is useful: CPU, ROCm, Vulkan, Level Zero, CUDA and Metal. CUDA works with the NVIDIA driver alone, without installing the toolkit. Vulkan picks up the older AMD APUs through the open-source RADV driver. mdma doctor tells you which backends actually work on the machine in front of you.
I learned the processing side from llama.cpp, by reading its kernels and the machine code they compile to. It is also the yardstick. Every GPU kernel is tested against the CPU backend, and a model counts as supported when it produces the exact greedy tokens llama.cpp produces from the same file. When a fast path fails that check, it stays marked as failing. Nobody loosens the tolerance to make it pass.
A cluster you spawn at will¶
Dispatch is the other half of the name, and the reason MDMA exists. At heart MDMA is a networking layer: it decides which engine, on which machine, gets each request.
You start an engine when you need one, on whichever machine has room, and stop it when you don’t.
That only works if an engine can land on a machine without hurting it. mdma plan estimates whether a model fits before anything loads. Each machine runs one MDMA at a time, enforced by a machine-wide lock, so two engines never fight over one GPU the way my six llama-server units did.
Engines on the same network can already find each other. mdma node serve runs a small daemon on each machine that advertises its CPU and memory over mDNS, and mdma node discover lists what is out there. That part is still a proof of concept. It only advertises resources, and it does so in plaintext, so it belongs on a LAN you trust.
Sending work between engines goes through federation, and MDMA machines can form one today. The shape comes from Matrix, the chat protocol: independent servers, each run by its own operator, that agree to talk to each other. Here the servers are your machines, plus machines that belong to people you choose.
Every peer is enrolled by hand. Peers talk over mutual TLS with pinned keys, and each one signs the list of models it offers. There is no central scheduler and no shared certificate authority.
The encryption protects the wire. The person who runs the remote machine still sees the prompts it processes, so permissions come as four separate grants: seeing a model, running it, routing to it, and disclosing a caller’s input to that operator.
A request runs whole on one peer. If that peer fails, the caller gets the error, and nothing quietly retries it on a machine you didn’t pick. Federation is off by default.
The public half¶
MDMA is being extracted from a larger private system, a supply chain platform that needs privacy and decentralization as much as security and transparency. That system reached about 1,000 tokens per second on off-the-shelf hardware. The public version doesn’t reach that yet. Pieces land in the repository as they get rebuilt and tested.
Each backend has been validated on one or two devices. The most useful contribution right now is a report from hardware it hasn’t seen, such as an RDNA3 card or an Intel Arc GPU. Run mdma doctor and send the output, whether it passes or fails.
Phones joining the federation are on the roadmap. So is a Coral Edge TPU backend: mdma info already finds the device, though nothing computes on it yet.
MDMA is BSD 2-Clause. Parts of its Vulkan kernels are adapted from llama.cpp, under MIT.
MDMA runs on its own. LSD, the daemon from the Rust Revolution post, is a separate project, and it launches next.
🔗Interstellar Communications
No transmissions detected yet.Be the first to establish contact!
Related Posts
Kaunta Is Rust Now. Postgres Did Not Notice.
In March I shipped Kaunta as a Go binary. It is Rust now, in the same repository, on the same database. The port was cheap because the logic already lived in PostgreSQL. What changed is who owns the layers underneath.
CrystalMeet
CrystalMeet turned Ruby into Crystal with a pile of regular expressions. It half worked, and the other half lives in Ruby's runtime, where no regex can reach. /r/rails deleted it in 2024. Its scenarios are landing in spinel now.
Helmsman: Stop Writing AGENTS.md That Lies to Half Your Models
Your static instruction file works for Claude Opus and breaks for Claude Haiku. Helmsman serves model-aware instructions that adapt to capability tiers, environment, and project context.