fork(), but for somebody else's process

Copy-on-write between two unrelated processes, over one mapping, with the destination already running as its own program. Linux has no call for that, so we added import_cow_range(2).

Laurentiu CiobanuLaurentiu Ciobanu6 min readDeep dive
A creature carries one thin sheet of tracing paper between two identical machines while a second creature stands rigidly still; behind them an enormous crate sits untouched, fine threads running to it from both machines.
On this page · 6 min

Deep dive is a highly technical series where I go deeper into the technology that makes boxd tick. This is not for the faint of heart!

The last post was about a forked machine waking up and discovering that its clock, its timer and its name all belong to somebody else. It skipped the obvious question, which is how the copy gets made at all.

Here is the shape of the problem. A machine is running, with several gigabytes of guest RAM. We want a second machine, running, holding the same RAM, in about a tenth of a second, and we want the two to diverge from that point without either being able to see the other's writes.

Copy-on-write, in other words. The oldest trick in Unix. The difficulty is that the mechanism Unix gives you for it is welded to a use case we are not in.

Why not the obvious things

Copy the memory. Several gigabytes at memory bandwidth into a fresh allocation. Hundreds of milliseconds at best, and you have doubled your RAM usage for two machines that in practice share nearly all of their pages forever. The whole point of forking a machine rather than booting one is that the copy is free until it isn't.

Snapshot to a file, restore from the file. This is what we do for hibernation, and it is fine there, because hibernation is allowed to be slow. Here it means writing gigabytes to disk and reading them back to produce a machine that already exists in RAM one process over.

fork() the VMM process. Tempting, and wrong for a specific reason worth stating: the child has to be a different program. It needs its own identity, its own argv, its own control socket, its own supervisor relationship. fork() gives you the address space, and exec() is what makes the process a different program, and exec() throws the address space away. The two halves cancel out. You cannot keep the memory and become somebody else.

userfaultfd. We shipped this for a long time. Guest memory is registered with UFFD, a handler process resolves faults, and copy-on-write is implemented in userspace with write-protect mode. It works, and it means every first touch of a page by either machine is a round trip out to a userspace handler and back. You are re-implementing, in a process, the thing the MMU and the kernel's fault path already do in hardware and in C.

What we actually wanted was fork()'s semantics (copy the page tables, mark both sides read-only, let the kernel resolve faults natively) but between two processes that are not related, over one specific mapping, with the destination already running as its own program.

Linux has no call for that. So we added one.

C
SYSCALL_DEFINE4(import_cow_range,
		int, src_pidfd,
		unsigned long, addr,
		unsigned long, len,
		unsigned int, flags)

Take a range out of another process's address space and install it in mine, at the same address, with fork-style copy-on-write. The source keeps its mapping. Both sides come out write-protected. The kernel does the rest.

How it works

The child VMM starts, reserves an empty placeholder where its guest RAM will go, opens a pidfd to the parent, and makes one syscall. What comes back is a page table pointing at the parent's memory.

No bytes move. What is copied is the index: the page table entries, which are a few thousand times smaller than the pages they describe. Both processes end up write-protected against the shared pages, so the first time either one writes, the CPU takes a fault, the kernel copies that single page, and the writer gets a private copy. Everything neither has touched stays shared forever.

That is exactly what fork() does, and deliberately so. The implementation is Linux's own dup_mmap(), the code fork() itself runs, narrowed to a single mapping and pointed at a different destination. Reusing that path rather than walking page tables by hand is what makes huge pages work: guest RAM is backed by 2 MiB pages, so a hand-rolled version would have needed its own huge-page path, then its own path for the multi-size ones, then its own path for the ones that had been swapped out, each a subtly wrong reimplementation of code sitting two files away.

Three things in that signature are decisions rather than details.

A pidfd, not a pid. A pid is a number that can refer to a different process by the time you use it. This syscall's entire job is to reach into another process's memory, which makes it the worst possible place to lose a race to pid reuse.

No destination address. The range lands at the same address it occupied in the source. That is only tolerable because we can guarantee it:

Rust
/// Fixed host virtual address where every boxd-vmm process pins its
/// guest RAM mapping. The same VA in parent and child is what makes
/// `import_cow_range(2)` work: the syscall requires `src_addr == dst_addr`.
pub const GUEST_RAM_HOST_VA: usize = 0x0000_4000_0000_0000;

Every VMM on the host pins guest RAM at the same hardcoded address, reserved before anything else can claim it. Deciding this once, in a constant, removes an entire negotiation protocol from the fork path.

Permissions are ptrace's. The kernel already has a well-litigated answer to "may this process touch that process's memory". The worst thing a new syscall can do is invent a second one that is subtly more permissive.

What it costs to use

One constraint leaks out of the kernel and into the product: the source must be alive and paused.

Guest RAM lives in the parent's page tables, so the parent is not a file we can read. It is a process that has to still be there, and has to be holding still while its entries are copied. The fork path pauses the parent's vCPUs, imports, and only then resumes it. That pause is most of what a fork costs, and it is why forking is fast but not free.

It also means a fork is a thing you do to a running machine, not to an image. There is no artifact in between, nothing to store, nothing to garbage collect, and no version of this that works if the parent has already exited.

What it buys

A machine that is running, cloned into a second machine that is running, in under 200 milliseconds, with no memory copied and no bytes read from disk. The two share every page they have not written since the split, resolved by the MMU at hardware speed, with no handler process anywhere in the path.

The reason that is worth a system call is what it does to the unit of work. A machine stops being a thing you provision and becomes a thing you branch. An agent that wants to try three approaches gets three machines that all begin from the same running state, with the same processes at the same instruction, rather than three machines that boot from the same image and then have to be dragged back to where the interesting part started. A template stops being a disk image and becomes a live machine, already booted, already warm, that new machines are cut from.

Branching costs about as much as an HTTP request. Once that is true, you use it for things you would never have provisioned a VM for.

And the child wakes up with a dead timer, a lying clock and somebody else's IP address, which is the previous post, and now you know what happened to it a few hundred microseconds earlier.

Laurentiu CiobanuLaurentiu Ciobanu
Post
Published
Sep 21, 2026
Reading time
6 min
Words
1,255
Topic
Deep dive

Read next

Field notes

Subscribe for release notes and architecture write-ups

No spam, ever. Unsubscribe anytime.

Your inbox