Overview#
Introduction
This application note provides a practical overview of deploying 20 Triton10 10GigE cameras streaming concurrently to a single host using RDMA for GPU-based image processing. It covers two validated host configurations: Option 1, a discrete GPU workstation, and Option 2, an NVIDIA DGX Spark system. The guide details hardware selection, network topology, VLAN segmentation, host configuration, and camera IP management using the Arena SDK. These configurations support sustained full-frame acquisition with minimal CPU load for inspection, metrology, and real-time image processing.
Topology Overview#
Hardware Specifications#
Note: The hardware listed below represents a reference configuration. Alternative components may be used as long as similar performance, bandwidth, and feature requirements are maintained.
Host Platform: Discrete GPU or DGX Spark (GB10)#
We validated this system on two host platforms. The camera, switch, and VLAN setup is the same for the two platforms. The difference is where the image buffers are and how the NIC writes to them.
Option 1: Discrete GPU workstation (x86-64). The GPU has its own VRAM. The NIC writes image data directly into VRAM through PCIe (GPUDirect RDMA with nvidia-peermem). This is the configuration in the Discrete GPU System Components (Option 1) section.
Option 2: DGX Spark class mini AI computer (NVIDIA GB10, ARM64). The CPU and the GPU share one memory pool. The NIC writes image data into pinned system memory and the GPU reads the same memory. No copy occurs. nvidia-peermem is not used.
How to choose: Option 1 gives dedicated VRAM and the highest GPU memory bandwidth. Option 2 is a compact system with a simpler software setup, and the CPU can read the images with no copy.
| Option 1: Discrete GPU (See full system specs in next section) | Option 2: DGX Spark (GB10) | |
|---|---|---|
| CPU architecture | x86-64 | ARM64 |
| GPU | RTX 6000 Ada | Blackwell GPU in the GB10 package |
| Memory | 128 GB DDR5 and 48 GB VRAM, separate | 128 GB shared CPU/GPU memory |
| Image buffer allocation | cudaMalloc (VRAM) | cudaHostAlloc with cudaHostAllocMapped |
| nvidia-peermem and IOMMU steps | Necessary | Not necessary |
| CPU access to the image buffer | No. Copy to host memory first | Yes, direct |
| NIC | ConnectX-5 MCX516A-CDAT, add-in card | Built-in ConnectX-7 |
| NIC host interface | One PCIe Gen4 x16 link | Two PCIe Gen5 x4 links |
| NIC driver | MLNX_OFED | DOCA-OFED |
| Arena SDK package | Linux x64 | Linux ARM64 |
| Validated result | 20 cameras, RGB8 at 28 fps | 20 cameras, RGB8 at 28.8 fps and BayerRG8 at 87 fps |
Discrete GPU System Components (Option 1)#
| Component | Model | Specifications |
|---|---|---|
| CPU | AMD Ryzen Threadripper 7960X | 24-core, sTR5 Socket |
| Motherboard | ASUS Pro WS TRX50-SAGE WiFi | PCIe 5.0 x16, DDR5 ECC support |
| RAM | 128GB DDR5 ECC R-DIMM | 4x 32GB sticks, 5600MHz |
| GPU | NVIDIA RTX 6000 Ada Generation | 48GB GDDR6, 18,176 CUDA cores |
| NIC | Mellanox ConnectX-5 MCX516A-CDAT | Dual-port 100GbE, RoCEv2, GPUDirect |
| Network Switch | FS S5850-24XMG-U | 24x 10GBASE-T PoE++ (RJ45), 2x 100G QSFP28 |
| Cables | Mellanox MCP1600 QSFP28 DAC | 2x 100GbE QSFP28 passive DAC, one for each VLAN uplink |
| Cameras | TRX124S-CC (x20) | 12.4MP, 10GigE, 28 fps |
| Storage | NVMe PCIe 5.0 SSD | 2TB minimum |
| PSU | 1300W+ 80+ Platinum |
Camera Specifications#
| Parameter | Value |
|---|---|
| Model | TRX124S-CC |
| Resolution | 4096 x 3000px (12.3 Megapixels) |
| Sensor | Sony IMX535 CMOS Global Shutter |
| Frame Rate - Validated RGB8 streaming | 28 fps on Option 1; 28.8 fps on Option 2 |
| Maximum frame rate by pixel format | 29.16 fps RGB8; 87.33 fps BayerRG8 |
| Frame Rate - Validated BayerRG8 streaming | 87 fps on Option 2 |
| Interface | 10 Gigabit Ethernet (10GigE) |
| Pixel Format | Mono8 / BayerRG8 / RGB8 |
| Per-Camera Bandwidth | ~8.3 Gbps @ 28 fps RGB8, ~2.75 Gbps @ 28 fps Mono8 or BayerRG8 |
| Firmware Required | v1.64.0.0 or higher |
Pixel Format: RGB8 from the Camera or BayerRG8 with GPU Debayering#
BayerRG8 sends 1 byte for each pixel and RGB8 sends 3. With BayerRG8 the maximum camera frame rate increases from 29.16 fps to 87.33 fps, because the 10GigE link is the limit and not the sensor. The GPU then does the debayering. The Option 2 reference system operates 20 cameras at 87 fps with BayerRG8.
| RGB8 (camera debayers) | BayerRG8 (GPU debayers) | |
|---|---|---|
| Image size at 4096 x 3000 | 36.86 MB | 12.29 MB |
| Maximum frame rate for each camera | 29.16 fps | 87.33 fps |
| Bandwidth for each camera at 28 fps | 8.3 Gbps | 2.75 Gbps |
| Buffer memory, 20 cameras x 16 buffers | 11.8 GB | 3.9 GB |
| Debayering and colour processing | Camera | CUDA kernel |
| GPU debayering code required | No | Yes |
Use RGB8 if 28 fps is sufficient and you want the camera’s colour processing without GPU debayering code. Use BayerRG8 if you need more than 28 fps, more network margin, or less buffer memory.
The reference system uses a CUDA kernel with the Malvar-He-Cutler 5 x 5 filter. The kernel reads the image buffer directly on both host options.
Network Topology Setup#
VLAN Configuration (Both Options)#
Both host options use the same switch VLAN configuration. Divide the 20 cameras into two groups of 10, each on a separate VLAN with a dedicated 100GbE uplink to one physical host port. Both host ports connect to the same switch.
VLAN isolation is required in this reference configuration to prevent duplicate camera discovery. Use the switch-port and IP assignments below for either host option. The Linux interface selection differs between platforms and is covered in the host configuration sections.
| VLAN | Cameras | Switch Ports | Physical Host Port | Host IP / Prefix |
|---|---|---|---|---|
| VLAN 10 | Cameras 1-10 | Ports 1-10, Uplink 25 | Port 1 | 172.16.10.1/24 |
| VLAN 20 | Cameras 11-20 | Ports 11-20, Uplink 26 | Port 2 | 172.16.20.1/24 |
ConnectX-7 Interface Selection (Option 2 Only)#
For DGX Spark, keep the VLAN and physical cable connections shown above, but select the active Linux interfaces according to the table below.
The built-in ConnectX-7 connects to the GB10 through two PCIe Gen5 x4 links. Each physical network port appears as two Linux interfaces, one for each PCIe link. Use enp1s0f0np0 for VLAN 10 and enP2p1s0f1np1 for VLAN 20 so the two camera groups use different PCIe links.
Why Option 2 uses both PCIe links
At 4096 × 3000 resolution, 20 cameras generate approximately 21 GB/s of image data when streaming RGB8 at 28.8 fps or BayerRG8 at 87 fps. This exceeds the theoretical bandwidth of one PCIe Gen5 x4 link, approximately 15.75 GB/s per direction before protocol overhead. DGX Spark therefore uses one camera VLAN on each PCIe link.
The discrete workstation’s ConnectX-5 uses a single PCIe Gen4 x16 connection with approximately 31.5 GB/s of theoretical bandwidth per direction before protocol overhead. Both network ports share this wider connection, so Option 1 does not require the additional PCIe-link selection used on DGX Spark. Both options still use one physical network port per camera VLAN.
| Camera group | Physical port | Linux interface | PCIe link | Host IP | Interface Status |
|---|---|---|---|---|---|
| VLAN 10: Cameras 1 to 10 | Port 1 | enp1s0f0np0 | A | 172.16.10.1/24 | UP |
| VLAN 20: Cameras 11 to 20 | Port 2 | enP2p1s0f1np1 | B | 172.16.20.1/24 | UP |
| Not used | Port 2 | enp1s0f1np1 | A | No address | DOWN |
| Not used | Port 1 | enP2p1s0f0np0 | B | No address | DOWN |
Important: In this configuration, both unused interfaces must be administratively down. Leaving them enabled can cause the Arena SDK to discover the same cameras through multiple interfaces during broadcast discovery. Leaving these interfaces without an IP address is not sufficient.
Apply the commands for your selected host option. For Option 2, disable both unused interfaces, assign the host IP addresses to the two active interfaces, and set their MTU to 9000.
# Disable the unused interfaces$ sudo ip link set dev enp1s0f1np1 down$ sudo ip link set dev enP2p1s0f0np0 down
Camera IP Assignments#
Cameras must have static IP addresses persisted to firmware to ensure consistent addressing after power cycles.
| VLAN 10 Cameras | IP Address |
|---|---|
| Camera 1 | 172.16.10.11 |
| Camera 2 | 172.16.10.12 |
| Camera 3 | 172.16.10.13 |
| Camera 4 | 172.16.10.14 |
| Camera 5 | 172.16.10.15 |
| Camera 6 | 172.16.10.16 |
| Camera 7 | 172.16.10.17 |
| Camera 8 | 172.16.10.18 |
| Camera 9 | 172.16.10.19 |
| Camera 10 | 172.16.10.20 |
| VLAN 20 Cameras | IP Address |
|---|---|
| Camera 11 | 172.16.20.11 |
| Camera 12 | 172.16.20.12 |
| Camera 13 | 172.16.20.13 |
| Camera 14 | 172.16.20.14 |
| Camera 15 | 172.16.20.15 |
| Camera 16 | 172.16.20.16 |
| Camera 17 | 172.16.20.17 |
| Camera 18 | 172.16.20.18 |
| Camera 19 | 172.16.20.19 |
| Camera 20 | 172.16.20.20 |
Network Switch Configuration#
Important: Disable IGMP snooping on both camera VLANs, VLAN 10 and VLAN 20, for this reference configuration. In testing, IGMP snooping interfered with PTP multicast delivery and prevented cameras from synchronizing.
SSH into the FS S5850-24XMG-U network switch and apply the following configuration:
enableconfigure terminal# Create VLANsvlan 10name CAMERAS_VLAN10exitvlan 20name CAMERAS_VLAN20exit# Assign camera ports to VLANs (access mode)interface eth-0-1switchport mode accessswitchport access vlan 10exit... repeat for ports 2-10 (VLAN 10) and 11-20 (VLAN 20)# Configure QSFP28 uplinks (access mode)interface eth-0-25switchport mode accessswitchport access vlan 10exitinterface eth-0-26switchport mode accessswitchport access vlan 20exitendwrite memory
Software Requirements & Installation#
Operating System#
-
Ubuntu 24.04.3 LTS (Recommended)
-
Kernel 6.8+ with IOMMU support
Required Software Downloads#
Use software packages that match the host architecture: x86-64 for the discrete GPU workstation and ARM64 for DGX Spark. The demo is built from the same source code on each platform, with platform-specific buffer allocation and build settings. Compile it separately for each host architecture.
The validated workstation uses MLNX_OFED, while DGX Spark uses DOCA-OFED. Follow the platform-specific requirements in the table below.
| Software | Discrete GPU workstation (x86-64) | DGX Spark (ARM64) | Purpose |
|---|---|---|---|
| Ubuntu 24.04 LTS | Install | Preinstalled as DGX OS | Operating system |
| NVIDIA driver | Install, 550 series or newer, with nvidia-peermem | Preinstalled, 580 series | GPU driver |
| CUDA Toolkit | Install, 12.x | Preinstalled, 13.0 | Debayer kernel, GPU memory allocation |
| MLNX_OFED | Install, 24.10 or newer, x86-64 package | Not used | ConnectX-5 driver and RDMA stack |
| DOCA-OFED and rdma-core | Not used | Preinstalled, 25.10 | ConnectX-7 driver and RDMA stack |
| nvidia-peermem | Required, load after OFED at every boot | Not used, no discrete VRAM | Lets the NIC write directly into GPU memory |
| Arena SDK | Install, Linux x64 1.0.10.11 | Install, Linux ARM64 0.8.10155 | Camera control, GVRSP streaming, IpConfigUtility |
| linuxptp 4.0 | apt install | apt install | ptp4l and phc2sys, PTP master on the host |
| GLFW 3.3 | apt install | apt install | OpenGL window |
| GLEW 2.2 | apt install | apt install | OpenGL extensions |
| FreeType 2.13 | apt install | apt install | On-screen text |
| libpng 1.6 | apt install | apt install | Wait screen and flash frames |
| OpenGL headers, g++, make | apt install | apt install | Compile and render |
| ethtool, iptables, NetworkManager | Preinstalled | Preinstalled | NIC tuning, VLAN relay rule, persistent addresses |
| ffmpeg | apt install, preparation only | apt install, preparation only | Convert videos to frame folders, not needed at run time |
cuDNN, OpenCV, and GLM are not required to build or run the demo described in this article.
Mellanox OFED Installation (Option 1 Only)#
# Download MLNX_OFED for Ubuntu 24.04$ wget https://content.mellanox.com/ofed/MLNX_OFED-24.10-1.1.4.0/\MLNX_OFED_LINUX-24.10-1.1.4.0-ubuntu24.04-x86_64.tgz# Extract and install with GPUDirect support$ tar -xzf MLNX_OFED_LINUX-24.10-1.1.4.0-ubuntu24.04-x86_64.tgz$ cd MLNX_OFED_LINUX-24.10-1.1.4.0-ubuntu24.04-x86_64$ sudo ./mlnxofedinstall --add-kernel-support --with-nvmf \--with-nfsrdma --enable-gds# Restart OFED services$ sudo /etc/init.d/openibd restart
CUDA Installation#
# Install CUDA Toolkit$ sudo apt install nvidia-cuda-toolkit# Verify installation$ nvcc --version$ nvidia-smi
Host System Configuration#
This section outlines the required host system configuration to ensure reliable high-throughput streaming and RDMA operation.
IOMMU Configuration (Option 1 Only)#
The tested AMD workstation uses amd_iommu=on iommu=pt to enable IOMMU in passthrough mode for GPUDirect RDMA. The following settings apply to this Option 1 reference system. These kernel-command-line changes are not required for the DGX Spark configuration described in this article.
$ sudo nano /etc/default/grub
# Add to GRUB_CMDLINE_LINUX_DEFAULT:GRUB_CMDLINE_LINUX_DEFAULT="quiet splash amd_iommu=on iommu=pt"
# Update GRUB and reboot$ sudo update-grub$ sudo reboot# Verify after reboot$ cat /proc/cmdline | grep iommu# Should show: amd_iommu=on iommu=pt
Network Interface Card Configuration (Option 1)#
Enter the following terminal commands to configure both NIC ports:
# Set IP addresses for each NIC port$ sudo ip addr add 172.16.10.1/24 dev enp1s0f0np0$ sudo ip addr add 172.16.20.1/24 dev enp1s0f1np1# Bring interfaces up$ sudo ip link set enp1s0f0np0 up$ sudo ip link set enp1s0f1np1 up# Set MTU to 9000 (jumbo frames)$ sudo ip link set enp1s0f0np0 mtu 9000$ sudo ip link set enp1s0f1np1 mtu 9000# Verify configuration$ ip addr show enp1s0f0np0$ ip addr show enp1s0f1np1
Network Interface Card Configuration (Option 2)#
# Disable the two unused interfaces$ sudo ip link set dev enp1s0f1np1 down$ sudo ip link set dev enP2p1s0f0np0 down# Set IP addresses on the active interfaces$ sudo ip addr add 172.16.10.1/24 dev enp1s0f0np0$ sudo ip addr add 172.16.20.1/24 dev enP2p1s0f1np1# Set MTU to 9000 (jumbo frames)$ sudo ip link set dev enp1s0f0np0 mtu 9000$ sudo ip link set dev enP2p1s0f1np1 mtu 9000# Bring the active interfaces up$ sudo ip link set dev enp1s0f0np0 up$ sudo ip link set dev enP2p1s0f1np1 up# Verify IP addresses, MTU, and interface state$ ip addr show dev enp1s0f0np0$ ip addr show dev enP2p1s0f1np1$ ip link show dev enp1s0f1np1$ ip link show dev enP2p1s0f0np0
Verify that both active interfaces show mtu 9000 and are administratively UP. The two unused interfaces must remain administratively DOWN to prevent duplicate camera discovery. These commands apply to the current session; configure the same settings persistently so they are retained after reboot.
NVIDIA-peermem Configuration (Option 1: Discrete GPU Only)#
This section applies only to the discrete GPU workstation (Option 1). DGX Spark (Option 2) does not use nvidia-peermem. Skip this section and continue to RoCE v2 Mode Configuration if you’re on Option 2.
The nvidia-peermem module enables direct RDMA access to GPU memory, allowing image data to be transferred from the network interface to the GPU without unnecessary copies. Correct module loading and startup order are required for reliable operation.
Loading the nvidia-peermem module#
# Check if module is available$ find /lib/modules/$(uname -r) -name "nvidia*peermem*"# Load the module$ sudo modprobe nvidia-peermem# Verify it is loaded$ lsmod | grep nvidia_peermem# Should show: nvidia_peermem with ib_core as dependency
nvidia-peermem Start Sequence#
Every time the computer is started, nvidia-peermem must be loaded AFTER the OFED services start. If you experience RDMA failures, run this sequence:
# Unload nvidia-peermem if loaded$ sudo modprobe -r nvidia_peermem# Restart OFED services$ sudo /etc/init.d/openibd restart# Reload nvidia-peermem$ sudo modprobe nvidia-peermem# Verify RDMA can see GPU memory$ lsmod | grep nvidia_peermem$ ls /sys/class/infiniband/mlx5_*/device/peer_mem/
RoCE v2 Mode Configuration#
# Set RoCE v2 mode on both portssudo cma_roce_mode -d mlx5_0 -p 1 -m 2sudo cma_roce_mode -d mlx5_1 -p 1 -m 2
Camera IP Address Assignment#
The following steps describe how to assign and persist static IP addresses on each camera to ensure consistent network discovery and reliable operation across reboots and power cycles.
Force and Persist IP Addresses#
Use the Arena SDK IpConfigUtility to assign and persist camera IP addresses by MAC address:
(Replace the example MAC address with your camera’s MAC address and use the corresponding IP from the Camera IP Assignments table.)
$ cd ~/ArenaSDK_Linux_x64/Utilities# Force IP (temporary, immediate effect)$ echo "" | ./IpConfigUtility /force -m 0x1C0FAF20DD68 \-a 172.16.10.11 -s 255.255.255.0 -g 0.0.0.0# Persist IP (survives power cycle)$ echo "" | ./IpConfigUtility /persist -m 0x1C0FAF20DD68 \-a 172.16.10.11 -s 255.255.255.0 -g 0.0.0.0
PTP#
PTP Configuration (Both Options)#
Both the ConnectX-5 in Option 1 and the ConnectX-7 in Option 2 support hardware timestamping. Both options use the same PTP settings, with different Linux interface names. Hardware timestamping records packet timing at the NIC, reducing timing uncertainty introduced by operating system scheduling and software processing.
Set the host’s priority1 and priority2 values to 10 so it is preferred over cameras using the default priority of 128. Lower numerical values indicate higher priority. This helps keep the host as the grandmaster for both camera VLANs.
Create a PTP configuration file named: grandmaster.cfg
[global]verbose 1boundary_clock_jbod 1time_stamping hardwaretx_timestamp_timeout 50priority1 10priority2 10domainNumber 0logAnnounceInterval 1logSyncInterval 0logMinDelayReqInterval 0announceReceiptTimeout 3delay_mechanism E2Enetwork_transport UDPv4# Include the interface pair for your selected option# Option 1: Discrete GPU:[enp1s0f0np0][enp1s0f1np1]# Option 2: DGX Spark:[enp1s0f0np0][enP2p1s0f1np1]
Run PTP grandmaster
$ sudo ptp4l -f grandmaster.cfg -m
Keep ptp4l running and open a second terminal.
# Run in a second terminal while ptp4l remains running$ sudo phc2sys -a -r
phc2sys -a will obtain the clock configuration from the running ptp4l process and follow its port states. The -r option also includes the system clock in synchronization. This is needed for our existing boundary_clock_jbod 1 configuration, where separate port hardware clocks need synchronization.
eTrigger: Hardware Pulse to Scheduled Action Command#
With eTrigger, a sender camera receives an electrical pulse on its GPIO and broadcasts a scheduled action command. The command specifies a shared PTP time at which all cameras, including the sender, are scheduled to begin exposure.
All cameras must have PTP lock before the pulse.
- The host configures and arms all cameras, including the sender, with
TriggerSelector = FrameStart,TriggerSource = Action0, andTriggerMode = On. - The host sets the sender camera to
ActionMode = Sendand the remaining cameras toActionMode = Receive. - A button or PLC sends a pulse to Line3 of the sender camera.
- The sender camera broadcasts an action command with execution time T, calculated as the PTP time at the pulse plus a lead time. The sender also schedules its own exposure for time T.
- The host relays the broadcast to the second VLAN.
- All cameras are scheduled to begin exposure at PTP time T and transfer their captured images to the host through RDMA.
The host relays the command between VLANs but does not originate the trigger. The acquisition application detects the triggered acquisition when the captured images arrive.
In this configuration, the sender camera also acquires images. All 20 cameras each capture one image per pulse.
Relay to the second VLAN (Option 2: DGX Spark example)#
iptables -t mangle -A PREROUTING -i enP2p1s0f1np1 -s 172.16.20.15 -d 255.255.255.255 -p udp --dport 3956 -j TEE --gateway 172.16.10.255
Lead Time#
The lead time must be longer than the time for the command to get to the last receiver, which includes the relay. Start with a conservative lead time that exceeds the worst-case command delivery time to all receivers, including the relay to the second VLAN. Allow additional margin for delivery-time variation and differences between camera clocks. Reduce the lead time gradually while comparing the image timestamps from all receivers for each pulse. At each step, send 20 pulses and compare the timestamps of the image captured by each receiver for the same pulse. When the difference increases suddenly, the lead time is too short. A lead time that is too short does not cause missing images. This creates a larger timestamp difference because GenICam standards require devices to execute scheduled action commands unconditionally.
Application Configuration: Buffers, Triggering, and GPU Debayering#
After completing the host, network, and camera setup, configure the application to allocate image buffers, receive triggered images, and process Bayer data on the GPU.
Image Buffer Allocation#
Read the required buffer size from the camera’s PayloadSize node and allocate each buffer accordingly:
- Option 1: Discrete GPU: Allocate GPU memory using
cudaMalloc. - Option 2: DGX Spark: Allocate mapped, pinned host memory using
cudaHostAllocwithcudaHostAllocMapped.
Pass the allocated buffers to StartStream as user-supplied buffers using UserSuppliedBuffer. Images retrieved through GetImage reference data in these buffers without an additional image-data copy.
eTrigger Configuration#
Enable PTP on all cameras and wait for PTP lock before sending a trigger pulse.
Configure the sender camera:
| Parameter | Setting |
|---|---|
ActionMode |
Send |
ActionSendTriggerSource |
Line3 |
ActionPTPTimestampMode |
Relative |
ActionSendPTPTimestamp |
Lead time (refer to Lead Time Section above) |
TriggerSelector |
FrameStart |
TriggerSource |
Action0 |
TriggerMode |
On |
Choose a lead time that allows the scheduled action command to reach all receivers, including those on the second VLAN, before its execution time.
Configure the receiver cameras:
| Parameter | Setting |
|---|---|
ActionMode |
Receive |
TriggerSelector |
FrameStart |
TriggerSource |
Action0 |
TriggerMode |
On |
GPU Debayering#
When streaming BayerRG8, perform Bayer-to-RGB conversion on the GPU. The demo uses a custom CUDA kernel implementing Malvar-He-Cutler demosaicing, a 5 × 5 linear filter that combines bilinear interpolation with gradient correction. The kernel reads the acquisition buffer directly and writes the converted output into an OpenGL texture, avoiding an intermediate image copy.
The demosaicing is implemented from the published paper: Malvar, He and Cutler, ‘High-Quality Linear Interpolation for Demosaicing of Bayer-Patterned Color Images’, IEEE ICASSP 2004. The paper gives five 5×5 filter kernels; our CUDA kernel is a direct transcription of those coefficients for the RGGB pattern, written in-house. No third-party code was used. A peer-reviewed reference implementation of the same paper is available at IPOL (Getreuer, 2011) if you want to compare. Microsoft held a patent on the method (US 7,502,505); it expired in March 2024. https://ieeexplore.ieee.org/document/1326587
As an alternative to implementing a custom debayering kernel, NVIDIA NPP provides nppiCFAToRGB_8u_C1C3R for Bayer-to-RGB conversion. Integration with the application’s output buffers and OpenGL display remains a separate step.
Troubleshooting#
This system involves multiple high-performance components and low-level configuration, which means careful setup is important. The following troubleshooting guidance is provided to help quickly identify and resolve common issues if adjustments are needed during deployment or tuning.
Common Issues & Solutions#
| Symptom | Cause | Solution |
|---|---|---|
| RDMA buffer registration fails | nvidia-peermem not loaded | Run: sudo modprobe nvidia-peermem |
| Cameras not discovered | Wrong subnet or VLAN config | Verify IP addresses and switch ports |
| Low fps / dropped frames | Insufficient kernel buffers | Increase rmem_max to 128MB |
| HALGvspRDMA::Start() failed | RoCE mode incorrect | Set RoCE v2: cma_roce_mode -m 2 |
| Black screen on 2nd camera | CUDA context conflict | Check nvidia-smi for GPU usage |
| Link not coming up | Auto-negotiation issue | Disable auto-neg on NIC |
Diagnostic Commands#
nvidia-peermem status
$ lsmod | grep nvidia_peermem
Check RDMA devices
Check network interface stats$ ibv_devinfo$ rdma link show
$ ethtool -S enp1s0f0np0 | grep -i "drop\|error"
Check GPU status
Check IOMMU$ nvidia-smi$ nvidia-smi topo -m
$ cat /proc/cmdline | grep iommu
Check Arena SDK camera discovery
$ cd ~/ArenaSDK_Linux_x64/Examples/Arena/Cpp_Enumeration$ ./Cpp_Enumeration
Startup Checklist#
cat /proc/cmdline | grep iommu2. Restart OFED:
sudo /etc/init.d/openibd restart3. Load nvidia-peermem:
sudo modprobe nvidia-peermem4. Verify peermem:
lsmod | grep nvidia_peermem5. Configure NIC IPs:
sudo ip addr add 172.16.10.1/24 dev enp1s0f0np0 && sudo ip addr add 172.16.20.1/24 dev enp1s0f1np16. Bring interfaces up with MTU:
sudo ip link set enp1s0f0np0 up mtu 9000 && sudo ip link set enp1s0f1np1 up mtu 90007. Set ring buffers:
sudo ethtool -G enp1s0f0np0 rx 8192 tx 8192 && sudo ethtool -G enp1s0f1np1 rx 8192 tx 81928. Ping cameras:
ping 172.16.10.11 && ping 172.16.20.119. Run camera enumeration to verify that all 20 cameras are visible.
10. Confirm IGMP snooping is disabled on VLAN 10 and VLAN 20.
11. Start
ptp4l with the updated priorities, then start phc2sys -a -r in a second terminal.12. Verify that the host is the PTP grandmaster and all cameras have PTP lock across both VLANs before sending an eTrigger pulse.