Understanding CXL 3.0 Core Architecture
Compute Express Link (CXL) 3.0 represents a fundamental shift in how heterogeneous computing systems achieve coherence across fabric-attached accelerators, memory expansion devices, and switching infrastructure. Unlike traditional PCIe-based topologies, CXL 3.0 introduces a three-layer protocol stack that enables cache-coherent access patterns: the physical layer (PHY), the link layer (LL), and the protocol layer (PL). The protocol layer itself subdivides into three critical components: CXL.io (I/O semantics), CXL.cache (cache coherence), and CXL.mem (memory access semantics).
In multi-tenant deployments, the CXL 3.0 fabric must maintain coherence guarantees across multiple isolated workload domains while sharing physical interconnect resources. This creates a fundamental tension: the coherency protocol assumes a unified address space and consistent memory ordering, yet multi-tenancy demands strong isolation boundaries. The fabric topology becomes the mediating layer where this conflict manifests most acutely.
Coherency Semantics in CXL 3.0
CXL 3.0 implements a directory-based coherence protocol rather than snooping-based coherence. Each memory location maintains a directory entry tracking which caches hold copies and their states (Exclusive, Shared, Invalid). When a processor or accelerator requests data, the directory responds with precise information about cache state rather than broadcasting queries across the fabric.
However, CXL 3.0 introduces optional snoop-based coherence paths for performance optimization. This hybrid approach allows hot data paths to use snooping (faster, lower latency) while cold data paths use directory coherence (more scalable, lower bandwidth). The protocol stack specifies:
- CXL.cache transactions: Requests and responses for cache-line coherence operations
- CXL.mem transactions: Direct memory access operations that may or may not trigger coherence actions
- CXL.io transactions: Legacy I/O semantics that bypass coherence entirely
In multi-tenant environments, this creates a coherency visibility problem. Tenant A's snoop transactions may traverse the same fabric links as Tenant B's memory operations, creating implicit information leakage and potential coherence ordering violations if not carefully managed.
Fabric Topology Considerations for Multi-Tenancy
CXL 3.0 supports multiple fabric topologies: point-to-point links, switched fabrics with hierarchical routing, and mesh-based interconnects. In multi-tenant deployments, the chosen topology dramatically affects coherence behavior.
Hierarchical Switch Fabric: A common topology uses a root switch connecting multiple sub-switches, each serving 4-8 CXL endpoints (processors, accelerators, or memory expanders). This topology naturally partitions the fabric into regions. Tenant A might occupy switches S1 and S2, while Tenant B occupies S3 and S4. Coherence requests within a tenant's region can use fast snooping, but cross-tenant requests must traverse through the root switch.
Mesh Topology: Some deployments use direct mesh connections between endpoints. This maximizes bandwidth but complicates coherence routing. A snoop request from Tenant A's processor must somehow avoid reaching Tenant B's caches while still reaching all relevant copies of data.
The CXL 3.0 specification addresses this through fabric-level address translation and tenant-aware routing tables, but these mechanisms are optional and implementation-dependent. Many silicon implementations rely on static provisioning of tenant boundaries at boot time, creating rigid isolation that conflicts with dynamic workload consolidation.
Real-World Multi-Tenant Scenario
Consider a hyperscaler operating a CXL 3.0 fabric with 64 endpoints: 32 are CPUs, 16 are GPU accelerators, and 16 are CXL memory expanders. The operator wants to consolidate two customer workloads: Customer X (8 CPUs + 4 GPUs + 4 memory expanders) and Customer Y (8 CPUs + 4 GPUs + 4 memory expanders).
The fabric routing layer must ensure:
1. Coherence correctness: When Customer X's CPU writes to shared memory, all Customer X caches see the update
2. Isolation: Customer Y's caches never receive Customer X's snoop requests
3. Performance: Hot data accessed by Customer X should not trigger unnecessary directory lookups
The CXL 3.0 protocol stack supports this through virtual channels and tenant ID tagging in transaction headers, but the snoop filter—the hardware structure that optimizes snoop routing—often lacks tenant-aware filtering logic, creating the mismatch that drives hot-page pinning failures.