The CXL (Compute Express Link) protocol stack represents a fundamental departure from traditional PCIe-based device communication by introducing a three-layered architecture that fundamentally changes how memory and I/O devices interact with host systems. Understanding this stack is essential for comprehending why multi-tenant memory pooling introduces latency variability that standard hypervisors cannot adequately control.
The Three-Layer CXL Architecture
CXL operates across three distinct semantic domains, each with different consistency guarantees and latency characteristics. The PCI Express (PCIe) layer at the bottom maintains backward compatibility with traditional device discovery and configuration mechanisms. Above this sits the CXL.io layer, which handles device-to-host communication using PCIe transactions. The CXL.cache and CXL.mem layers introduce the revolutionary capabilities that enable memory pooling but simultaneously create the contention challenges this course addresses.
The CXL.io semantics preserve traditional device semantics: a device initiates a request, the host processes it, and a response returns. This layer maintains strict ordering guarantees and uses interrupt-driven signaling. However, CXL.cache and CXL.mem layers operate under fundamentally different rules that blur the boundary between device and memory subsystem behavior.
Device Semantics in CXL.io
Device semantics in CXL.io closely mirror PCIe behavior but with extended capabilities. When a device sends a request across CXL.io, it follows a request-response model where the device initiates, and the host-side CPU or fabric responds. This creates a natural serialization point: the host's request queue. In multi-tenant scenarios, this queue becomes a critical bottleneck.
Consider a practical example: two virtual machines on the same host both have CXL devices attached. VM-A's device initiates a configuration read while VM-B's device initiates a memory write. The host-side CXL root complex must arbitrate between these requests. Traditional hypervisors use FIFO scheduling or simple priority schemes, meaning a low-priority VM's device request can block a high-priority VM's request. This creates device-level head-of-line blocking, where microsecond-scale delays accumulate rapidly in high-frequency workloads.
Memory Semantics and Coherency Models
CXL.mem and CXL.cache layers introduce memory semantics that fundamentally differ from device semantics. In CXL.mem, external memory devices (like CXL-attached DRAM) participate directly in the host's memory coherency domain. This is revolutionary but dangerous: the device no longer sends discrete "requests" but instead behaves like another NUMA node.
The coherency model used in CXL is critical here. Most CXL implementations employ write-through coherency or snoop-based coherency where the device's memory operations must be validated against the host's L3 cache and other devices' caches. When a CXL memory pool receives a write from VM-A targeting address 0x4000000, the CXL fabric must:
1. Check if any CPU cache holds that line (snoop phase)
2. Invalidate or update those caches (coherency enforcement)
3. Complete the write to the CXL memory device
4. Acknowledge back to VM-A
Each of these steps introduces latency, but more critically, they introduce contention points. If VM-B is simultaneously reading from nearby addresses in the same CXL pool, their coherency traffic competes for the same snoop bus and fabric bandwidth.
The Coherency Nightmare in Multi-Tenant Pooling
Standard hypervisors treat CXL memory pools as simple NUMA nodes and apply traditional NUMA balancing heuristics. These heuristics assume that cache coherency traffic is proportional to memory access patterns, but in CXL, coherency traffic is multiplicative. A single write generates snoop traffic to all CPU sockets, all other CXL devices, and all other memory pools.
Real-world measurement shows that in a 2-socket system with 4 CXL memory pools, a single write to pooled memory can generate 8-12 coherency transactions (2 sockets × 4 pools, plus additional traffic for invalidation). When multiple VMs contend for the same pool, this creates a coherency storm: each memory operation triggers exponential snoop traffic, and the snoop bus becomes the limiting resource, not the CXL fabric itself.
The critical insight hypervisors miss is that coherency models are not transparent to performance. Two VMs accessing the same CXL pool with different coherency patterns (one using write-intensive patterns, one using read-intensive patterns) experience dramatically different latencies due to coherency enforcement overhead. A VM performing sequential writes experiences 40-60% higher latency than a VM performing random reads, even when both access the same physical memory bandwidth.