Feng Zhang
← Writing

Is the Future of Databases Zero-Copy?

An introduction to devmem/TCP and NVMe Controller Memory Buffers (CMB), and whether combining them is the future of zero-copy storage

Last summer, I had the exciting opportunity to work under Prof. Martin Karsten in the Systems and Networking group at the University of Waterloo as an Undergraduate Research Assistant. I learned a ton about operating systems and networking over the summer, and wanted to share one of the most exciting projects I got to work on: zero-copy networking in Linux with devmem/TCP (devmem). This is my first blog post of hopefully many, and I’m writing this post to

  1. strengthen my understanding of this topic, as I believe that teaching a subject makes you stronger at it .
  2. share some of the work I did, because I genuinely think it is some of the most interesting work I have done.
  3. build/learn in in public.
  4. help others learn about zero-copy networking and devmem, since online documentation is extremely sparse, and I found it super difficult and time consuming to learn about this very interesting topic!

This post will dive deep into the operating systems, networking, and storage background necessary to understand devmem and zero-copy networking, and hopefully provide an opinionated view of whether this technology might become mainstream in the future.

0. Intro / Motivation

The problem: consider a web server that accepts HTTP requests over TCP to stores the same data in a database on disk. Lets analyze what actually happens under the hood when this server receives a request:

  1. A special piece of hardware called the network interface card (NIC) on the computer receives the packets and places them into a ring buffer in kernel memory over PCIe
  2. The data travels up the PCIe hierarchy to be copied into a userspace buffer where the data can be processed by the user space process
  3. The user space process sends the same data back down the PCIe hierarchy to a secondary storage device such as an SSD to be stored

This traditional path has two main limitations which can be optimized. Firstly, notice that we’re traveling across the PCIe hierarchy twice without really changing any of the packet data. This consumes PCIe bandwidth unnecessarily. Furthermore, the data has to bounce through an intermediate buffer in kernel memory before it travels to it's final destination in an user space buffer. If we can eliminate this extra data copy, we can speed the entire process up.

As you’ve seen, every payload byte in this path crosses the PCIe switch twice and is copied by the CPU, even though the CPU never actually looks at it. So can we modify this data movement pipeline to get the NIC to write our packet data directly into the SSD to make our storage system faster?

1. Background

This post assumes little to no background in Operating Systems and Networking, since I started this research at the same point, so this section will explain some of the background knowledge necessary to understand the solution to the above problem.

1.1 How the OS sees Memory and does Data Movement

Almost all modern operating systems have the concept of user space and kernel space. User space is the region of physical memory where untrusted user applications run. Kernel space as you can probably guess, is the region of physical memory where the OS kernel runs. The reason for the separation is to provide security for the system; user space applications have restricted access to the CPU so that a bug in one process can’t corrupt the kernel or other processes. The kernel runs with full privileged access, which is necessary to manage hardware, but also means a kernel bug is far more dangerous, since it can take down the entire system.

System Calls

System calls (syscalls) are how user space applications interact with the kernel in a secure way. There are many times where user applications need to interact with the kernel. For example, when writing a to a file, under the hood, the write() syscall is called and the program transfers control to the kernel in whats known as a trap to the kernel. The kernel then handles the task of writing the data to the hardware.

Direct Memory Access (DMA):

Direct Memory Access (DMA) is a mechanism used by modern operating systems to allow hardware devices to transfer data directly to and from system memory without involving the CPU.

Before DMA, whenever packets would arrive in the NIC, the CPU would waste clock cycles manually copying all packet data from the NIC to a kernel buffer for processing. This is a clear waste of CPU time and can be optimized with hardware called a DMA engine. The DMA engine takes in parameters from the CPU (location of data in memory, how much data to copy, and which device to send it to from the operating system), and takes control of the internal bus (bus mastering) to transfer the data without the CPU’s involvement. It notifies the CPU when it's done with a signal called an interrupt.

DMA is supported and used in almost every modern operating system nowadays, and solves the problem of data movement from a hardware device to system memory. What it doesn’t solve however, is the example scenario I listed at the beginning of the blog, where we want to transfer data from one hardware device to another.

Why this matters later: Devmem TCP avoids the kernel-to-userspace copy by having the application share a buffer with the device ahead of time. When data arrives, recvmsg() returns a location in that buffer (an offset and a length) instead of a copy of the bytes.

1.2 The Linux Networking Stack

Let’s dive a bit into the basics of the Linux networking stack.

The NIC, RX/TX queues, and descriptors

The NIC is a special piece of hardware in computers that handles the processing of network packets. It maintains many different queues for receiving (RX) and transmitting (TX) packets to provide parallelism for packet processing. Each one of these queues is backed by a ring buffer in kernel space with descriptors.

A descriptor is a small struct placed in RX/TX ring buffers that contains a 64-bit physical memory address that is a DMA-able location where the NIC should write incoming data, packet length, and status flags.

On the receive side for example, the kernel NIC driver first fills the RX ring with descriptors that point to empty buffers and tells the NIC how far it has filled the ring with head/tail pointers. When a new packet arrives, the NIC takes the next empty descriptor, DMAs the packet into the buffer, and marks it as completed. The driver takes completed descriptors and refills the ring with newly allocated buffers.

Interrupts vs NAPI Polling

Once the packet is in RAM, the kernel needs to be notified about it. An obvious approach would be to issue a CPU interrupt (IRQ) each time the NIC finishes processing a packet, however, this would raise serious performance issues under high packet rates as the CPU would spend most of it's time processing packet IRQs.

Linux solves this with NAPI (New API). The first packet is processed with an IRQ. Then, NAPI interface drivers disable packet RX IRQs and provide a poll method to the kernel. The kernel polls the ring using SoftIRQ context and processes a fixed number of packets per pass. If the kernel finishes processing all packets in the ring, we switch back to using IRQs. Under light load, you have low latency from IRQs, and under heavy load, you avoid thrashing via polling under SoftIRQ contexts.

sk_buff and fragments

For each packet, the NIC driver builds a struct sk_buff (SKB) to represent the packet as it moves up the packet processing stack. SKBs hold head and tail pointers to packet data, not the data it'self. SKBs have a small linear area - it holds the header data (Ethernet, TCP/UDP, IP headers) and pointers to the packet data. However, when packet data is really big, we don’t want to allocate a huge amount of contiguous physical memory to this one packet. Thus, an array of pointers to fragments (skb_frag_t) is used. Each fragment is a page (4KB) long, and the struct contains a page offset and the size of the packet data inside that fragment.

The RX Packet Path

Putting all the previous steps together, a packet arriving to a Linux machine via TCP goes through the following steps:

  1. The NIC DMAs the packet into a buffer from a pointer in an RX descriptor in an RX queue
  2. An interrupt fires and NAPI begins to poll the RX ring
  3. The driver wraps the buffer in an SKB and passes it up the processing stack
  4. The IP and TCP layers validate headers, handle ACKs, and queues the data on the socket buffer queue
    1. A socket buffer queue is a temporary kernel-managed buffer used to hold data packets while they transit between the NIC and a user space process
    2. These are collections of SKBs
  5. The application calls recv() or read() and the kernel copies the payload from the from the SKBs into the user space application buffer

Step 5 is the only step where the CPU touches the payload bytes.

There is a much more detailed explaination of the journey of a packet in the linux stack that i have attached as part of the sources.

Scatter-Gather DMA and IO Vectors (IOV)

As mentioned earlier, the fragment array in an SKB exists because packet data typically isn’t in one contiguous blog. It contains a list of (page, offset, length) entries. It turns out that two other places in the network processing stack basically uses the same kind of list.

On the hardware side, the frag array is what scatter-gather DMA produces and consumes.

  • RX: each RX descriptor points at one buffer, so when a packet spans several buffers, the NIC fills them across several different descriptors. The driver then attaches each filled buffer to the skb as one skb_frag_t fragment entry
  • TX: ****the NIC driver traverses the SKB frags and gives the NIC one descriptor per entry

On the user space side, an iovec is the same kind of list**.** recvmsg() takes an array of struct iovec, each entry a pointer and a length, so one call can fill several separate application buffers in order. An app might pass one iovec for a 16-byte record header and another for the payload.

The inefficiency lies in the CPU copy between the SKB frag list and the IOV. When recv() is called, the kernel traverses the frag list and copies bytes into IOV entries in user space.

This copy is what devmem changes; it replaces the page in each frag with a chunk of exported device memory, and skips the frag to IOV copy for payloads. The application gets offsets into the device buffer (e.g. GPU buffer) instead of bytes in IOVs.

Receive Side Scaling (RSS) and Flow Steering

To provide parallelism in packet processing, modern NICs have many RX and TX queues, each with it's own interrupt and NAPI context, usually pinned to different CPU cores. This brings us into the topic of load balancing the different queues and which queue a packet should be routed to:

  • RSS: the NIC hashes each packet and uses the hash (source IP, destination IP, source port, destination port, protocol) to map to a queue. All packets of a flow land on the same queue, and different flows are spread across different queues
  • Flow Steering: Some NIC can support flow steering to install explicit rules to direct certain packets to certain queues. These rules take priority over RSS
    • ethtool -N <if> src-ip <ip> dst-ip <ip> ... queue 15

RSS and flow steering are critical parts of devmem, as we need them to direct packets from our application into specific queue(s) for special processing by our application.

To learn more about RSS and other packet steering options (RSS, RPS, RFS, aRFS), this is a great blog I found.

Header Split

Header split is essential for zero-copy networking; it is a feature supported in some NICs that split's the header of the packet into one buffer, and the payload into another. The kernel’s TCP stack only needs the header buffer for validations, so the packet can be processed without the payload remaining in the same place.

This split allows us to remove the copy between SKB frags in kernel space to IOVs in user space, as we can simply point to the physical memory page directly inside the virtual address space of the user space process.

Tying it all Together

In order for devmem to work, it needs to resolve one isolated RX queue and configure RSS and flow steering to ensure only the flow from you application lands on this queue (isolate traffic). Additionally, it also needs header split so the kernel can process TCP headers normally in host RAM while the payload goes straight to device memory without being copied an additional time.

1.3 Sharing Memory Between Devices

So far, every buffer we’ve seen has been host DRAM that the kernel allocated. Devmem requires the NIC to write payloads into memory that belongs in a different device to avoid the double copy and PCIe traversal, so we need a way for two drivers to export/import a region of device memory.

dmabuf

dmabuf (Direct Memory Access Buffer) is a framework in the Linux kernel for sharing memory buffers across different devices (GPU, NIC, etc.) and driver subsystems without copying data.

There are two roles:

  • The exporter is the driver that owns the buffer, e.g. a GPU exporting it’s VRAM
  • The importer is the driver that wants to DMA into or out of that memory. The importer attaches to the buffer and smaps it, giving it addresses it's own device can use

To user space applications, a dmabuf is a Linux file descriptor (fd), so an application can hold one and bound a queue to it later. Common exporters in the kernel today are mostly GPU devices, since this work has mostly been tested on GPUs for distributed training.

page_pool

Recall that the NIC driver always refills the RX ring with newly allocated buffers. Allocating a page and DMA mapping it for every packet is really expensive, so most NIC drivers use something called a page_pool which is an allocator for DMA mapped pages. The page pool receives used up pages after packet processing is finished, and reuses it. The flow looks like this:

  1. The driver takes a page from the pool and puts it in an RX descriptor.
  2. The NIC fills it, and the page becomes a fragment in an SKB.
  3. The TCP processing stack in the kernel, and eventually the application, finish processing it.
  4. The page is recycled back to the pool.

In devmem, we modify the driver’s page pool by serving chunks of dmabuf instead of kernel pages so that the NIC directly fills the packet payload into another device’s memory. The rest of the RX path barely changes.

Tying it all Together

We can use dmabuf backed memory in the NIC queue to point at device memory. We have to modify the exporter driver to support dmabufs, the NIC driver to import the same dmabuf. This has interesting use cases in speeding up distributed training (direct DMA into GPU DRAM), and in our case, in distributed database systems (direct DMA into SSD DRAM).

1.4 PCIe

So far, we have only talked about the software side of things. However, there are some hardware limitations on whether a NIC can DMA directly into an SSD. It depends on how the devices are physically connected and how the connections are configured. There are many different types of buses inside a computer that connect the various different hardware components inside like SATA, PCI, PCIe, etc. The most common and dominant today is PCIe, so we will be talking about PCIe.

Topology: endpoints, switches, root complex

PCIe is a tree of point to point links between devices. A point to point topology means that two devices are connected directly to each other with a dedicated link. Endpoints are the devices themselves (NIC, SSD, GPU, etc.), switches connect to an upstream port out to several downstream ones, meaning multiple endpoints hang off a single port. The root complex sit at the top of the tree, connecting the PCIe hierarchy to the CPU and it's memory controllers, allowing devices to reach host RAM.

Data travels via Transaction Layer Packets (TLPs) each carrying an address to specify which device to route to. For example, a DMA write from the NIC to host ram is a memory write TLP that travels up any switches to the root complex and into DRAM.

Address routing and BARs

To route TLPs in PCIe, every device is assigned a range of addresses from the systems address space. Each switch knows which address spaces sit below each of it's ports, so switches can forward TLPs toward whichever port has it's destination address.

Each PCIe device has a base address register (BAR) that tells the CPU how much memory address space the device needs and maps that memory into system memory for Memory Mapped IO (MMIO).

When the system boots, firmware or the kernel reads each device’s BARs and assigns them an address space. After that the device’s internal memory / registers are reachable via ordinary loads (LDUR in ARM64) and stores (STUR) to those addresses (via MMIO).

PCIe topology is what makes device to device DMA possible. For example, a GPU’s DRAM can be exposed through a BAR, so it has an address and the NIC can DMA to that address like any other memory write TLP. To resolve the PCIe bandwidth issue I talked about earlier, if both devices sit under the same switch, the switch can route the write straight to the GPU without ever reaching the root complex and consuming extra bandwidth. However, this is only possible if these two devices sit under a common switch.

IOMMU and IOVAs

In most modern systems, devices rarely access physical RAM addresses directly for security reasons. Instead a piece of hardware called the I/O Memory Management Unit (IOMMU) sit's inside the PCIe root complex translates I/O virtual addresses (IOVAs) into physical RAM addresses. This is very similar to how the MMU in the CPU translates virtual addresses of processes into physical addresses. When a device does DMA it uses IOVAs which are translated per device using tables inside the IOMMU into physical RAM addresses. This enforces better security, since a device can only read or write memory explicitly mapped for it by the OS.

Access Control Services (ACS) Why P2P Traffic is Forced UYpstream

As mentioned previously, the IOMMU exists for security, so if a PCIe switch routes TLPs between downstream ports without accessing the root complex, it can bypass the IOMMU directly. To prevent this from happening by default, PCIe has something called ACS. With ACS request redirect enabled, a switch will not forward a P2P TLP downstream until it has reached the root complex, where the IOMMU can validate the access.

Many kernels have ACS enabled by default, for security purposes. The result is that device P2P traffic that could have been routed downwards at the switch instead travels up to the root complex, consuming extra bandwidth. Most of the time, the bandwidth isn’t fully utilized, as the PCIe 5.0 specification with 4 lanes has 15.754 GB/s throughput, however, in high throughput applications like distributed training in GPUs or extremely high throughput distributed systems, this might become a slight bottleneck in the system. Disabling ACS on the relevant switch ports or disabling it overall is necessary for the flow we want.

Tying it all Together

Whether devmem will reduce PCIe bandwidth also depends on the PCIe topology of your system, not just the software. The NIC and SSD need to be under a common switch, the SSD needs a DRAM buffer that can be set via a BAR that the NIC can address, and ACS and IOMMU have to allow the transfer to happen.

1.5 NVMe

I’ve covered how a NIC can DMA to an address that lands in the buffer of another device other than host RAM. The last piece of background information is on the storage side, as my primary research focus was if zero-copy networking with devmem could be used in database applications (the creators of devmem have used it successfully at Google Cloud for distributed training on GPUs instead of SSDs).

NVMe Protocol Basics: Queues, Doorbells, Commands, and Data Pointers

NVMe is a high speed storage access and transport protocol used specifically for SSDs. Most modern SSDs today run on NVMe so it’s important to understand the basics of the specification and how it works.

NVMe connects storage to the CPU via the PCIe bus, and is build around pairs of queues in memory. The submission queue (SQ) is where the host places commands, and the completion queue (CQ) is where the controller places results. NVMe supports 64K independent queues, with each queue holding up to 64K commands for massive parallelism.

The submission-completion flow is as follows:

  1. The host writes a command into the next free slot in the SQ.
  2. The host writes to the SQ's doorbell register, an MMIO register that tells the controller a new command is ready.
  3. The controller fetches the command, executes it, and writes a completion entry into the CQ.
  4. The controller either sends an IRQ to the host or the host polls the CQ, and the host writes to the CQ's doorbell to ack how far it has consumed.

A command in the SQ doesn’t carry it's data payload inline. It carries a data pointer which is either a Physical Region Page (PRP) list (list of physical pages in host RAM) or an Scatter Gather List (SGL), which tells the controller where in memory to read from or write to. This is the same scatter-gather idea as before; the data can live in several non-contiguous chunks.

Normally, that data pointer references host RAM. The interesting question for our use case would be if that pointer pointed to another device buffer.

Controller Memory Buffer (CMB) in NVMe SSDs

A CMB is DRAM that lives on the SSD’s controller and can be exposed to the host to use through a BAR.

CMBs were first officially introduced into the NVMe 1.2 standard back in 2014. There are two main use cases for CMBs:

  1. Placing SQ and CQs in CMBs rather than host memory
  2. Providing a DMA target for data buffers

Once a CMB is mapped, it can be accessed via MMIO as described earlier.

For our purposes, we care about it's capability to provide a DMA target. If a NIC can DMA a payload into the CMB, and the same region can be the target of a PRP in a write command, then you can build a data movement path where the data lands straight from the NIC into NVMe without any intermediate host buffer copy, a “zero-copy” data movement pipeline.

Persistent Memory Buffer (PMR) in NVMe SSDs

CMBs are typically volatile, meaning a power loss can lose whatever is sitting in the buffer before committed to flash, raising durability concerns. The NVMe spec also has PMRs, which is a separate BAR exposed region that is backed by persistent media. The only difference is that because PMR has stronger durability guarantees, it may be slower for write workloads than a CMB.

Tying it All Together

The distinction between CMB and PMR is important when considering durability and transactions in databases. If a TCP payload lands in the CMB, you must wait until it's written to flash to acknowledge the transaction and mark it as durable; PMR does not have this restriction but slows down writes.

1.6 Other zero-copy approaches (short)

devmem isn’t the only way to achieve zero-copy data movement. A few other well established approaches exist, and I will mention each briefly.

Kernel Bypass (DPDK, SPDK)

With kernel bypass, the user space process talks to the NIC or SSD directly, with the kernel driver abandoned entirely. This is optimal for high performance applications, but the application now has to reimplement what the kernel driver used to do, such as network stack, memory management, garbage collection for pages, etc. This also severely changes the libraries and development approach of applications.

AF_XDP

The application registers a shared memory region with the kernel, and the RX/TX rings live inside it, meaning the NIC can DMA directly into the shared memory with the application. The kernel driver is still involved and owns the device, but the packet data skips the kernel bounce buffer.

RDMA (RoCEv2, InfiniBand)

RDMA is an extremely popular way of doing zero-copy networking, and used in many enterprise data centres. It requires an RDMA capable NIC, RDMA specific network switches (InfiniBand switches, RoCEv2 ethernet, etc.), and OS RDMA driver support. In RDMA, the NIC writes directly into a remote application’s memory or another device’s memory. The tradeoff is the extra hardware required to setup RDMA (does not work over traditional TCP like devmem does).

CXL

CXL is an interconnect that lets different devices share memory. CXL was developed by intel in 2019, and tooling and development are still relatively early. The CPU, endpoint devices, and the motherboard must all support CXL in order for it to work.

All of these other methods either require kernel bypass or specialized hardware/fabrics (except AF_XDP). devmem does not require either, since it keeps the standard Linux TCP/IP stack, and only changes where the payload lands. Most modern NICs support the features that devmem requires (must have RSS, header split, flow steering) with some kernel configuration required (ACS settings). With that being said, there are other hardware and software challenges to supporting devmem fully that I will address later on.

2. devmem/TCP Main Flow

I have now explained all the background information necessary to understand what the devmem modifications do.

NIC Setup

Before any packets arrive, you must configure the flow that you want to use devmem on to a separate queue of it's own using Linux’s ethtool command. We must install explicit flow steering rules to override RSS and make sure nothing but our desired flow lands in a queue:

ethtool -G eth1 tcp-data-split on         # enable header split
ethtool -K eth1 ntuple on                 # enable hardware flow steering
ethtool --set-rxfh-indir eth1 equal 15    # keep RSS off queue 15

Queue 15 is now empty and reserved, header split is on, and an ntuple rule can later point a specific flow at it.

Binding the Queue to a dmabuf

With the queue isolated from traffic, user space must allocate a dmabuf from our exporting device (e.g. SSD CMB or GPU DRAM) and bind it to that queue’s page_pool to make sure it hands out chunks of dmabuf backed memory rather than host pages.

Where Headers and Payloads go

For a packet in the bound queue, the NIC split's the packet into a header and a payload buffer. The header buffer goes into a small buffer in host RAM to go through the TCP/IP processing stack. The payload buffer is backed by the dmabuf and never ends up in host RAM at all.

net_iov

Recall that an SKB’s fragments are a list of struct (page, offset, length) entries. The problem is that dmabuf memory has no backing page struct. We solve this problem by piping the stack with a new net_iov struct which describes a fragment of dmabuf memory the same way a page would. net_iov is what now flows through the rest of the networking stack, and code now has to check which kind of memory it's looking at. (All of this is implemented and shipped in the Linux kernel already).

The resulting SKB s marked with a devmem flag, which tells the stack to not dereference the payload, since it lives in device memory.

Userspace API

In the user space process, recvmsg() doesn’t return the packet’s payload bytes directly. It will deliver a control message containing a byte offset and length for the dmabuf memory region that we previously mmap’d. The application adds the offset to the base buffer address to read the incoming data directly.

Once the processing is completed, the application issues a setsockopt() call to mark the fragment as reusable. The memory chunk is recycled back into the pool, ready to receive the next packet on that queue.

Previous Case Study

Google ran this in production on Cloud A3 VMs, moving data directly between NICs and GPUs. They reported roughly 96% of line rate, 192 Gbps bidirectional per NIC/GPU pair, about 3x the throughput of a regular TCP-based NCCL transport, and close to RDMA-based NCCL once messages are large enough to amortize the setup cost. All of that on plain TCP/IP, with no change to switches, cabling, or congestion control. This is a GPU deployment, not an SSD one, but confirms that this data movement path is possible and efficient.

3. Devmem TCP into NVMe CMB/PMR

We know from the previous section that devmem works for moving payloads from NIC to GPU directly. However, any device that can export a dmabuf backed memory region is a valid target for devmem, so a question I focused on in my research is if an SSD’s CMB or PMR could be a valid destination for database related tasks.

The idea

If we can export a section of an SSD’s CMB or PMR as a dmabuf, then we can use devmem to DMA straight from NIC to SSD CMB. Unfortunately, dmabuf support is extremely limited for many SSD drivers and whats worse is that SSDs that support CMBs or PMRs on the market go for thousands of dollars. This means that in order for this to work, we would have to modify the Linux driver code for that specific SSD to export dmabufs. Setting all of the drawbacks aside, lets take a look at the updated datapath.

Updated Datapath

[ Traditional Path ]
1. NIC → PCIe Switch → Root Complex → Host RAM (DMA write)
2. Host RAM → CPU Cache → Host RAM (Kernel-to-Userspace copy)
3. Host RAM → Root Complex → PCIe Switch → NVMe SSD (Storage DMA write)

[ Devmem TCP → CMB Path ]
1. Header:  NIC → PCIe Switch → Root Complex → Host RAM (kernel TCP/IP processing)
2. Payload: NIC → PCIe Switch → NVMe SSD CMB

Compared to the traditional path, the payload:

  • crosses the PCIe switch once instead of twice,
  • never occupies host DRAM bandwidth,
  • and is never touched by the CPU.

The header still takes the original route, since the kernel TCP stack needs it in host RAM to do it's job.

A Packet in CMB to Flash

Landing a payload into CMB is not the same as writing it to flash. The follow up necessary is that after recvmsg() returns the cmsg, the application needs to issue a normal NVMe write command that references the CMB offset. The SSD then should read from the CMB and commit to flash, the same as it would from host RAM. Acknowledging the write to the sender has to wait for both these two events to finish, not for the network arrival.

Applications to Databases

Replicated Logs, Storage Nodes, Block Storage

In distributed systems, replication is a super common pattern that appears almost everywhere in order to keep data durable. Some common examples are replicated logs for keeping the state of multiple nodes synchronized, storage node replication in a database, and block storage, where the node takes the raw payload and streams them onto disk in block chunks. These flows benefit a lot from our described zero-copy flow as we completely eliminate the CPU’s only job in copying bytes from NIC to SSD through the PCIe hierarchy with NIC to SSD DMA.

OLTP Databases (MySQL, PostgreSQL, etc.)

An OLTP database has many different workflows. A write to the database could be to update a specific row, for a specific user, in a specific table, and update any indexes on the tables. Many different parts on disk could change as the result of a single write. This means that the data that actually ends up durable is not the payload that arrived over the network directly, it’s usually something that the database computed from it.

This means that:

  • the payload needs interpreting, not just storing it’s raw data into the SSD. This parsing has to happen in host RAM
  • write requests can modify multiple parts on disk (row data, indexes, WAL, etc) which are all computed by the CPU

Therefore, landing the raw bytes in CMB doesn’t make much sense for traditional OLTP databases since you’d still need to the data into host RAM to process them.

  • Note: it is technically possible to read and modify CMB data I think, but bringing it into host RAM would be faster in this case since traversing the PCIe hierarchy takes more time

The zero-copy datapath only makes sense when the bytes you get from the payload are exaclty the bytes you keep.

4. Challenges

  • Software: no NVMe dma-buf exporter, so you'd write one (using P2PDMA); the devmem authors haven't tested against a real SSD.
  • Hardware: few SSDs expose a CMB or PMR, and those that do are expensive enterprise drives; the NIC needs header split and page_pool support; NIC and SSD must share a switch.
  • Config: ACS, IOMMU, and no multi-queue or header split in virtio-net, so bare metal or SR-IOV. Include your VM anecdote.
  • Protocol/system: app-level framing, encryption, and durability semantics.
  • Alternatives: FPGAs, and the CXL outlook.

Software

To my knowledge, there is no driver that supports exporting dmabufs based off CMBs in Linux. Anyone who wishes to test the devmem zero-copy flow using an SSD would have to modify the existing Linux kernel driver for that SSD to export dmabufs.

Hardware

Even with the software written, the hardware has to satisfy:

  • The SSD needs a CMB or PMR. To the best of my knowledge, there are only a few enterprise SSDs that expose a CMB or PMR, and they all cost at minimum thousands of dollars.
  • The NIC needs header split, RSS, and flow steering. The entire mechanism depends on the NIC being able to split headers from payloads, and direct specific application flows to a specific queue for devmem.
  • The NIC and SSD need to share a switch. We know that a P2P write only avoids the root complex if both devices sit under a common switch. If eliminating PCIe bottleneck is not a concern, this can be avoided.

Configuration

Assuming the hardware is right, the platform configuration still has to cooperate:

  • ACS, needs to be configured to allow the peer-to-peer route rather than forcing traffic up through the Root Complex.
  • The IOMMU needs to be set up so the P2P path is actually permitted rather than blocked or silently redirected.

Virtualized NICs don't support any of this**.** The virtio-net interface exposed a single queue, with no RSS and no header-split support at all. There was no way to even reserve an isolated queue for devmem TCP. Multi-queue RSS and hardware header split are properties of a real NIC and it's driver, and virtio-net is not a sufficient replacement, so testing must be done on bare metal, which rules out a lot of cloud VMs.

FPGAs

Since there are almost no SSDs that support CMBs at an affordable price on the market, one solution I’ve seen online is an FPGA based NVMe SSD with CMB support. If this would be something you’re interested in, reach out to me by email.

5. Conclusion

So, is the future of databases zero-copy? Right now, I think no, on both counts.

For general-purpose OLTP, no, for the reason covered in section 3.

For distributed replication, also no, at least for now, despite being the workload where the zero-copy networking might be best fit. This is because there are simply too many challenges in regards to hardware support, software support (requires kernel development), ACS configuration, etc. to make it worthwhile. Additionally, RDMA has been the industry standard for HPC and high performance storage networking for many years now, so devmem would need to bring a lot of advantages over RDMA to make adopting it worth it in my opinion.

I would like to thank Prof. Martin Karsten for supervising this work over the summer. Devmem TCP was only one part of what I got to explore; I also spent time benchmarking memcached and playing with kernel bypass. I went into the summer without having taken Operating Systems or Networking yet, so the amount I learned in a short window was substantial. I'm extremely grateful for the opportunity.

If you find any errors in this post, feel free to let me know by email! Operating system kernels are super complicated, and I may have gotten some small details wrong.

6. Sources / further reading

My full research notes, including everything below and a lot more, are here:

Background

devmem/TCP