This is a GDRCopy project ported for Moore Threads GPU platform.
.
├── src/ # Source code directory
│ ├── gdrdrv/ # Kernel driver module
│ └── *.c # User-space API implementation
├── include/ # Header files directory
├── tests/ # Test programs directory
├── packages/ # Package configuration files directory
├── scripts/ # Build and installation scripts directory
├── build/ # Build output directory
└── Makefile # Build script
While GPUDirect RDMA is meant for direct access to GPU memory from third-party devices, it is possible to use these same APIs to create perfectly valid CPU mappings of the GPU memory.
The advantage of a CPU driven copy is the very small overhead involved. That might be useful when low latencies are required.
GDRCopy offers the infrastructure to create user-space mappings of GPU memory, which can then be manipulated as if it was plain host memory (caveats apply here).
A simple by-product of it is a copy library with the following characteristics:
-
An initial memory pinning phase is required, which is potentially expensive,
-
Fast H-D, because of write-combining.
-
Slow D-H, because the GPU BAR, which backs the mappings, can't be prefetched and so burst reads transactions are not generated through PCIE
The library comes with a few tests like:
- mtgdrcopy_sanity, which contains unit tests for the library and the driver.
- mtgdrcopy_copybw, a minimal application which calculates the R/W bandwidth for a specific buffer size.
- mtgdrcopy_copylat, a benchmark application which calculates the R/W copy latency for a range of buffer sizes.
- mtgdrcopy_apiperf, an application for benchmarking the latency of each GDRCopy API call.
GDRCopy is a low-latency GPU memory copy library based on NVIDIA GDRCopy driver, designed for Moore Threads MTGPU GPUDirect RDMA technology. It enables CPU to directly map and access GPU memory for low-latency data transfers.
Current version components:
- mtgdrdrv kernel module: 0.1.0
- libgdrapi user-space library: 1.0.0
libgdrapi 1.x.x versions are designed to be backward compatible with mtgdrdrv 0.x.x kernel module versions. This means users can update the user-space library independently without upgrading the kernel module.
make all# Build all componentsmake driver# Build kernel driver modulemake lib# Build user-space API librarymake exes# Build test programsmake clean# Clean build artifacts
make install# Install all componentsmake lib_install# Install user-space API librarymake exes_install# Install test programsmake drv_install# Install kernel driver module
make build# Build directory structure for packagingmake deb# Create single DEB package (traditional way)make deb-bundle# Create single DEB package with all components (bundle way)
make build-kmod# Build kernel module package structuremake build-lib# Build API library package structuremake build-tests# Build test tools package structuremake build-meta# Build meta package structuremake deb-kmod# Create kernel module DEB packagemake deb-lib# Create API library DEB packagemake deb-tests# Create test tools DEB packagemake deb-meta# Create meta package DEB packagemake deb-all# Create all DEB packages
make build
make debmake build
make deb-bundle# Build each package separately
make build-kmod
make build-lib
make build-tests
make build-meta
# Or build all packages at once
make deb-allAfter splitting, there will be 4 packages:
- mtgdrdrv-dkms - Kernel module package, containing DKMS kernel module and related configurations
- libgdrapi - API library package, containing user-space API library and header files
- mtgdrcopy-tests - Test tools package, containing test tools and related dependencies
- mtgdrcopy - Meta package, depending on the above three packages for backward compatibility
To verify package building, generate all packages:
make deb-allThis creates the four .deb files in the build/ directory.
GDRCopy has been tested on S80, S4000 and S5000.
For CPU compatibility, GDRCopy has primarily been tested on Intel, AMD and Hygon platforms. Functional compatibility does not imply identical performance. Performance can vary with the GPU model, PCIe topology, CPU cache policy, IOMMU configuration and the instruction set selected for the CPU.
ARM builds are expected to compile successfully, but compatibility across different ARM platforms has not been validated and ARM packages are not currently provided. ARM support is experimental. Other architectures supported by the upstream NVIDIA GDRCopy project are outside the supported scope of this project.
MUSA DRIVER is required in all deployment environments. For host-based use, install the complete MUSA DRIVER stack on the host before building or installing GDRCopy. This includes the GPU driver, the driver development headers, and the compatible MUSA toolkit required by the build and tests.
For container-based use, install the mtgdrdrv-dkms package on the host, and install the libgdrapi and mtgdrcopy-tests DEB packages inside the container. Ensure that the container has the required permissions and that the GPU devices and driver interfaces are mapped correctly into the container.
We recommend using the official MUSA containers and following the official container workflow. For details, see the MUSA container documentation and the official container image directory.
Root privileges are necessary to load or install the kernel-mode device driver on the host.
$ make
$ sudo ./insmod.sh
$ export LD_LIBRARY_PATH=$(readlink -f ./src):${LD_LIBRARY_PATH}After making from source successfully, you can run tests
$ cd tests
$ ./mtgdrcopy_copybwAll tests should be able to compile and run on latest MUSA platform. For old version of MUSA, some subtests of mtgdrcopy_sanity will fail. The performance varies by CPU and GPU.
GDRCopy works with regular MUSA device memory only, as returned by musaMalloc. In particular, it does not work with MUSA managed memory.
gdr_pin_buffer() accepts any addresses returned by musaMalloc and its family.
In contrast, gdr_map() requires that the pinned address is aligned to the GPU page.
Neither MUSA Runtime nor Driver APIs guarantees that GPU memory allocation
functions return aligned addresses. Users are responsible for proper alignment
of addresses passed to the library. This is a restriction from the original NVIDIA driver. Unaligned address may be supported by further updates.
Multiple musaMalloc allocations may appear contiguous in the GPU virtual address space. However, passing an address range that spans multiple allocations to gdr_pin_buffer() or gdr_map() is not generally guaranteed to work. Support depends on the MUSA DRIVER version and hardware platform, and applications should not rely on this behavior.
For reporting issues, please open an issue in the public repository:
https://github.com/MooreThreads/mtgdrcpy/issues.
The project is maintained by Moore Threads and is distributed through the public repository. Please review the platform requirements and known limitations before use.
If you find this software useful in your work, please cite: R. Shi et al., "Designing efficient small message transfer mechanism for inter-node MPI communication on InfiniBand GPU clusters," 2014 21st International Conference on High Performance Computing (HiPC), Dona Paula, 2014, pp. 1-10, doi: 10.1109/HiPC.2014.7116873.