Skip to content
AI Primer
workflow

Zach Mueller compares direct and cascaded PCIe connections for an eight-GPU setup

Zach Mueller compares direct and cascaded PCIe connections for an eight-GPU expansion setup. His updated explanation corrects CPU-bypass claims and details bandwidth, airflow, noise, and maintenance constraints.

4 min read
Zach Mueller compares direct and cascaded PCIe connections for an eight-GPU setup
Zach Mueller compares direct and cascaded PCIe connections for an eight-GPU setup

TL;DR

  • Eight GPUs can share one host x16 connection: Zach Mueller's wiring explanation uses two PCIe switches with four primary GPU slots apiece.
  • Cascade's expected CPU-bypass advantage is uncertain: Mueller revised his explanation after discussion of slower paths through additional switches.
  • A ninth GPU needs additional host cabling, Mueller learned from his distributor, correcting his one-slot, nine-card plan.

Miwin sells twelve-slot boards with either two Broadcom 89104 chips or two Broadcom 89144 chips. Eight GPUs behind one motherboard slot is a lovely hardware hack; Mueller's existing “H100 node at home” post contains a terminal listing for eight RTX PRO 6000 cards.

Mueller sourced the expansion board directly from Miwin through AliExpress, according to his purchase reply. He estimates roughly $4,000 for the board and another $200–300 for a retimer.

The hardware is an MCIO-connected backplane. Mueller describes the minimum connection in a reply: one full x16 host slot feeding PCIe-to-MCIO cabling.

He puts each switch at 104 lanes in a follow-up. That switch-port budget is separate from the x16 or dual-x16 links carrying traffic back to the host.

Direct connect and cascade

Eight primary GPU slots divide into two four-card groups in Mueller's topology breakdown. His two wiring arrangements are:

  1. Direct connect: Two host x16 interfaces feed the switches through four MCIO x8 connections in total. Each four-GPU group has its own switch and host uplink; traffic between groups follows the host-side PCIe path in his description.
  2. Cascade: One host x16 interface feeds SW1 through two MCIO x8 connections. SW1's downlink feeds SW2's uplink through another pair, leaving an x16 connection between the switches.

The GPU-facing x16 links and the shared host uplinks have different bandwidth budgets. Eight attached GPUs therefore share the available host connection, even when transfers within a switch group can remain local.

Cascade hops and root routing

Mueller revised his initial CPU-bypass explanation after discussion of the additional switch hops. He said cascade could make communication much slower, with root-connected groups likely faster.

Single-slot cross-group transfers would be poor, he warned in a later reply, citing other people's data.

Separately, NVIDIA's NCCL documentation describes how I/O virtualization and PCI Access Control Services can redirect point-to-point PCIe traffic to the CPU root complex, reducing performance or even causing hangs.

Auxiliary slots and the ninth GPU

Mueller's distributor clarified that the auxiliary direct-connected slots require their host connectors to be connected. That withdrew his proposed ninth GPU fed solely from SW2's downlink while retaining one host slot.

Keeping cascade would require an additional retimer and another direct host connection, according to his follow-up. He subsequently said he expected to use two host slots in direct mode, which would also make a ninth GPU slot available.

Airflow and rack maintenance

Mueller dismantled his server rack in favor of open frames and cases. His reasons for leaving the rack included:

  • Motherboard failures.
  • Two different GPU airflow patterns.
  • Noise.
  • More difficult access for maintenance.
  • A likely 12U requirement for just two PCs, with little space saved.

His revised physical layout puts five server cards directly on the backplane, with cooling for the closely spaced cards, and moves the Pro cards to a second level with more spacing, as he described in the same update.

Nine-GPU benchmark matrix

The eight-card terminal listing predates the new backplane's tests. After the shipping thread, Mueller said the board had not arrived and that he would share a purchase link after reviewing it.

He proposed AIPerf baselines at PCIe 5.0 x8, comparing four Pro cards against four Max-Q cards, with stable NVFP4 quantizations preferred:

  • GLM 5.3 Flash.
  • DeepSeek v4 Flash.
  • Qwen Flash Next.

Serving runs use SGLang, which he said he uses exclusively. With nine available GPUs, his planned backplane tests cover:

  • Individual four-GPU switch groups.
  • Both switch groups combined.
  • One switch group plus one to four direct-connected GPUs.
  • Both switch groups plus one direct-connected GPU.

He also left open the comparison between eight-way tensor parallelism (TP8) and four-way tensor parallelism with two pipeline stages (TP4/PP2) in the original explanation.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 7 threads
TL;DR2 posts
Direct connect and cascade2 posts
Cascade hops and root routing1 post
Auxiliary slots and the ninth GPU2 posts
Airflow and rack maintenance2 posts
Nine-GPU benchmark matrix7 posts
Share on X