R740xd – Dual V100 Build, refurbished
Warranty: Lifetime
Dell PowerEdge R740xd 24SFF Rack Server – Dual NVIDIA Tesla V100 Build (Refurbished)
2U Rack Server with Dual NVIDIA Tesla V100 32GB GPUs – Built for AI Training, Deep Learning & Parallel Compute
This fully configured Dell PowerEdge R740xd pairs dual Intel Xeon Silver processors with two NVIDIA Tesla V100 32GB GPUs (64GB total GPU memory), making it a strong choice for multi-GPU AI/ML training, HPC, and other parallel workloads that outgrow a single-GPU box.
Key Specifications:
- Server Model: Dell PowerEdge R740xd 24SFF (24x 2.5" Small Form Factor bays)
- Processors: 2x Intel Xeon Silver 4210R (2.4GHz base, 10 cores / 20 threads each) – 20 Cores / 40 Threads total
- Memory: 128GB DDR4 ECC Registered (expandable)
- GPU: 2x NVIDIA Tesla V100 32GB (64GB total GPU memory), with GPU Enablement Kit
- Storage: 1x Dell 960GB 6G SATA 2.5" SSD
- RAID/Controller: PERC HBA330 (Non-RAID / Pass-Through)
- Networking: Intel X710 2x 10Gb SFP+ + I350 2x 1Gb Base-T rNDC
- Remote Management: iDRAC Enterprise (with dedicated management port)
- Power: 2x 1600W AC Hot-Plug Power Supplies (redundant)
- Cooling: High Performance Fan Kit
- Rails: Included Rail Kit
- Condition: Professionally refurbished and tested
- OS: No OS included (supports Windows Server, Linux, VMware ESXi, etc.)
Highlights:
- Dual NVIDIA Tesla V100 32GB GPUs – 64GB combined GPU memory for multi-GPU training, large-model inference, and compute-intensive HPC workloads
- 24-bay 2.5" SFF chassis with room for storage expansion
- Strong dual-socket performance with 20 cores / 40 threads
- Enterprise remote management via iDRAC Enterprise
- High-bandwidth 10Gb SFP+ networking included
- Redundant 1600W power supplies + high-performance cooling for 24/7 reliability
Perfect For:
- Multi-GPU AI / Machine Learning training
- Deep learning research and large-batch inference
- High Performance Computing (HPC) and scientific computing
- GPU-accelerated virtualization and VDI
- Rendering and 3D graphics workloads
This server has been fully tested and is ready to deploy. Ships with all necessary power cables, rails, and bezel (where applicable).
Note: This is a refurbished unit in excellent cosmetic and functional condition. Minor signs of previous rack use may be present.
Compare Our R740xd Builds:
All three configurations below share the same base platform (24SFF chassis, 2x Intel Xeon Silver 4210R, 128GB RAM, PERC HBA330, Intel X710/I350 networking, iDRAC Enterprise) — they differ only in GPU configuration and price, so you can pick the compute level you need:
- R740xd – GPU Ready (No GPU Included) – $2,995 – Same base server with the GPU Enablement Kit pre-installed but no GPU — ideal if you already own GPUs or want to add your own later.
- R740xd – Dual V100 Build (this listing) – $5,849 – 2x NVIDIA Tesla V100 32GB (64GB total) — best value for multi-GPU training and parallel workloads.
- R740xd – A40 Build – $10,959 – 1x NVIDIA A40 48GB — best single-GPU performance for AI inference and professional visualization.
Need a custom configuration, more storage, or a different GPU pairing? Message us — we can build to your requirements.
What Two V100 32GB Cards Actually Give You
Worth being precise about the "64GB" number, because it is the thing buyers most often misread. The Tesla V100 PCIe has no NVLink connector — NVLink was exclusive to the SXM2 module version. These two cards talk to each other over PCIe Gen3 x16, not a high-speed GPU-to-GPU link. That means:
- 64GB is two pools of 32GB, not one pool of 64GB. A model that fits in 32GB runs beautifully on one card. A model larger than 32GB has to be sharded across both, and pays PCIe bandwidth on every layer boundary.
- Where two cards shine: data-parallel training (two independent copies chewing through different batches), running two different models side by side, per-user or per-tenant GPU assignment, batch inference, and traditional CUDA/HPC work that was written for multi-GPU nodes.
- Raw bandwidth per card is excellent. Each V100 carries 32GB of HBM2 at up to 900 GB/s — meaningfully faster than the GDDR6 on newer 48GB cards. For anything that fits inside 32GB, a V100 is still a quick card.
- Specs per card: 5,120 CUDA cores, 640 Tensor cores, 250W, dual-slot, passively cooled, PCIe Gen3 x16. Two cards means 500W of GPU load in the chassis.
Software Notes Before You Buy
The V100 is Volta (compute capability sm70), and that has real consequences for modern LLM tooling that are easy to discover the hard way:
- FP16 tensor cores, but no BF16 and no FP8. Volta predates both. Training recipes written for BF16 need adapting to FP16 with loss scaling.
- Several modern inference kernels target Ampere (sm80) and newer. FlashAttention-2 and the Marlin INT4 kernels are the common examples; on Volta you run FP16 paths, older kernel implementations, or community sm70 builds of vLLM. It works — people do run quantized models on V100s — but expect to spend time on the stack rather than pulling a container and going.
- Excellent for FP16 training, CUDA and HPC. Simulation, scientific computing, rendering, classic deep-learning training and CUDA-native codebases all run without ceremony.
Put plainly: if the plan is serving modern INT4-quantized language models with off-the-shelf tooling, our A40 build is the better fit despite the higher price — 48GB in a single Ampere pool with full modern kernel support. If the plan is FP16 training, multi-tenant GPU workloads, HPC, or rendering, two V100s at this price is hard to beat. Not sure which describes your workload? Tell us what you intend to run and we will say so honestly.
Configuration Requirements — Already Handled
Both V100s are 250W passively cooled cards with no fans of their own. Dropping them into a stock R740xd does not work — Dell's installation guidance requires a specific set of parts, all of which are installed and tested in this build:
- GPU enablement kit (risers, GPU power cables, GPU air shroud, mylar insulation) — Dell 330-BBMO / 85CYN in the US, 490-BEIX / JKFGX in EMEA
- Low-profile heat sinks in both CPU sockets — standard heat sinks physically foul the GPU air shroud, and Dell requires both sockets populated
- High-performance fan kit — mandatory for any GPU configuration in this chassis, and doubly so with 500W of passive GPU in the air path
- Dedicated GPU power cables — required for any card above 75W; each V100 draws 250W
- 1600W power supplies — Dell sets an 1100W floor for GPU configurations; 1600W supplies provide real redundancy headroom rather than a nominal pair
- 30°C maximum inlet temperature — the thermal ceiling once GPUs are installed alongside 150W/8C, 165W/12C, 200W or 205W processors
Assembled from separate secondary-market purchases, that is seven line items, seven condition claims and seven return windows. Here it is one part number, one warranty and one place to call. The full breakdown, including the 330-BBMO vs 490-BEIX part-number confusion and the two mistakes that stop most self-assembled builds: Putting a passive datacenter GPU in a PowerEdge R740xd
Frequently Asked Questions
Is this really 64GB of usable GPU memory?
It is 2 x 32GB, not a single 64GB pool. The Tesla V100 PCIe has no NVLink connector, so the cards communicate over PCIe Gen3 x16. Models that fit in 32GB run on one card at full speed; larger models must be sharded across both and pay PCIe bandwidth at each layer boundary.
Can I run a 70B language model on this server?
Yes, sharded across both cards at 4-bit quantization, but expect to invest effort in the software stack — several modern INT4 and attention kernels target Ampere and newer. For straightforward 70B serving, our A40 build holds the model in one 48GB pool with full modern kernel support.
Does the V100 support BF16 or FP8?
Neither. Volta tensor cores are FP16. BF16 arrived with Ampere and FP8 with Ada Lovelace and Hopper. FP16 training with loss scaling works well.
What is the V100's memory bandwidth compared to newer cards?
Up to 900 GB/s of HBM2 per card, which is higher than the 696 GB/s GDDR6 on a 48GB A40. For workloads that fit inside 32GB, the V100 is genuinely fast.
Can I add a third GPU?
The R740xd supports up to three double-width accelerators, and Dell requires all installed GPUs to be identical in type and model — so a third V100, not a mixed pair. Contact us first; power and riser assignments change.
Does it come with an operating system?
No OS is included. The server supports Windows Server, Linux distributions, VMware ESXi and Proxmox. iDRAC Enterprise is licensed for remote installation and management.
Can I expand the storage?
Yes. The chassis has 24 x 2.5" SFF bays and ships with one 960GB SATA SSD, leaving 23 bays open. The PERC HBA330 runs in pass-through mode, which suits ZFS, Ceph and software-defined storage.
What condition is it in and what warranty applies?
Professionally refurbished, fully tested and burned in, in excellent cosmetic and functional condition with minor signs of previous rack use possible. Backed by the Resilient Tec lifetime warranty, reachable directly rather than through a manufacturer ticket queue.
Further Reading
- Putting a passive datacenter GPU in a PowerEdge R740xd — every part the cards need, the enablement kit part numbers, and why a stock chassis throttles them
- What can you actually run on 48GB of GPU memory? — benchmarked model sizes and the KV cache math, useful for sizing either build
- Retrofitting legacy Dell & HPE servers for GPU workloads — the cable checklist for an existing chassis
- The cloud sent you a bill — on-prem compute versus rented GPU time
- GPU & AI infrastructure guide and enterprise servers guide
Dell PowerEdge R740xd 24SFF Rack Server – Dual NVIDIA Tesla V100 Build (Refurbished)
2U Rack Server with Dual NVIDIA Tesla V100 32GB GPUs – Built for AI Training, Deep Learning & Parallel Compute
This fully configured Dell PowerEdge R740xd pairs dual Intel Xeon Silver processors with two NVIDIA Tesla V100 32GB GPUs (64GB total GPU memory), making it a strong choice for multi-GPU AI/ML training, HPC, and other parallel workloads that outgrow a single-GPU box.
Key Specifications:
- Server Model: Dell PowerEdge R740xd 24SFF (24x 2.5" Small Form Factor bays)
- Processors: 2x Intel Xeon Silver 4210R (2.4GHz base, 10 cores / 20 threads each) – 20 Cores / 40 Threads total
- Memory: 128GB DDR4 ECC Registered (expandable)
- GPU: 2x NVIDIA Tesla V100 32GB (64GB total GPU memory), with GPU Enablement Kit
- Storage: 1x Dell 960GB 6G SATA 2.5" SSD
- RAID/Controller: PERC HBA330 (Non-RAID / Pass-Through)
- Networking: Intel X710 2x 10Gb SFP+ + I350 2x 1Gb Base-T rNDC
- Remote Management: iDRAC Enterprise (with dedicated management port)
- Power: 2x 1600W AC Hot-Plug Power Supplies (redundant)
- Cooling: High Performance Fan Kit
- Rails: Included Rail Kit
- Condition: Professionally refurbished and tested
- OS: No OS included (supports Windows Server, Linux, VMware ESXi, etc.)
Highlights:
- Dual NVIDIA Tesla V100 32GB GPUs – 64GB combined GPU memory for multi-GPU training, large-model inference, and compute-intensive HPC workloads
- 24-bay 2.5" SFF chassis with room for storage expansion
- Strong dual-socket performance with 20 cores / 40 threads
- Enterprise remote management via iDRAC Enterprise
- High-bandwidth 10Gb SFP+ networking included
- Redundant 1600W power supplies + high-performance cooling for 24/7 reliability
Perfect For:
- Multi-GPU AI / Machine Learning training
- Deep learning research and large-batch inference
- High Performance Computing (HPC) and scientific computing
- GPU-accelerated virtualization and VDI
- Rendering and 3D graphics workloads
This server has been fully tested and is ready to deploy. Ships with all necessary power cables, rails, and bezel (where applicable).
Note: This is a refurbished unit in excellent cosmetic and functional condition. Minor signs of previous rack use may be present.
Compare Our R740xd Builds:
All three configurations below share the same base platform (24SFF chassis, 2x Intel Xeon Silver 4210R, 128GB RAM, PERC HBA330, Intel X710/I350 networking, iDRAC Enterprise) — they differ only in GPU configuration and price, so you can pick the compute level you need:
- R740xd – GPU Ready (No GPU Included) – $2,995 – Same base server with the GPU Enablement Kit pre-installed but no GPU — ideal if you already own GPUs or want to add your own later.
- R740xd – Dual V100 Build (this listing) – $5,849 – 2x NVIDIA Tesla V100 32GB (64GB total) — best value for multi-GPU training and parallel workloads.
- R740xd – A40 Build – $10,959 – 1x NVIDIA A40 48GB — best single-GPU performance for AI inference and professional visualization.
Need a custom configuration, more storage, or a different GPU pairing? Message us — we can build to your requirements.
What Two V100 32GB Cards Actually Give You
Worth being precise about the "64GB" number, because it is the thing buyers most often misread. The Tesla V100 PCIe has no NVLink connector — NVLink was exclusive to the SXM2 module version. These two cards talk to each other over PCIe Gen3 x16, not a high-speed GPU-to-GPU link. That means:
- 64GB is two pools of 32GB, not one pool of 64GB. A model that fits in 32GB runs beautifully on one card. A model larger than 32GB has to be sharded across both, and pays PCIe bandwidth on every layer boundary.
- Where two cards shine: data-parallel training (two independent copies chewing through different batches), running two different models side by side, per-user or per-tenant GPU assignment, batch inference, and traditional CUDA/HPC work that was written for multi-GPU nodes.
- Raw bandwidth per card is excellent. Each V100 carries 32GB of HBM2 at up to 900 GB/s — meaningfully faster than the GDDR6 on newer 48GB cards. For anything that fits inside 32GB, a V100 is still a quick card.
- Specs per card: 5,120 CUDA cores, 640 Tensor cores, 250W, dual-slot, passively cooled, PCIe Gen3 x16. Two cards means 500W of GPU load in the chassis.
Software Notes Before You Buy
The V100 is Volta (compute capability sm70), and that has real consequences for modern LLM tooling that are easy to discover the hard way:
- FP16 tensor cores, but no BF16 and no FP8. Volta predates both. Training recipes written for BF16 need adapting to FP16 with loss scaling.
- Several modern inference kernels target Ampere (sm80) and newer. FlashAttention-2 and the Marlin INT4 kernels are the common examples; on Volta you run FP16 paths, older kernel implementations, or community sm70 builds of vLLM. It works — people do run quantized models on V100s — but expect to spend time on the stack rather than pulling a container and going.
- Excellent for FP16 training, CUDA and HPC. Simulation, scientific computing, rendering, classic deep-learning training and CUDA-native codebases all run without ceremony.
Put plainly: if the plan is serving modern INT4-quantized language models with off-the-shelf tooling, our A40 build is the better fit despite the higher price — 48GB in a single Ampere pool with full modern kernel support. If the plan is FP16 training, multi-tenant GPU workloads, HPC, or rendering, two V100s at this price is hard to beat. Not sure which describes your workload? Tell us what you intend to run and we will say so honestly.
Configuration Requirements — Already Handled
Both V100s are 250W passively cooled cards with no fans of their own. Dropping them into a stock R740xd does not work — Dell's installation guidance requires a specific set of parts, all of which are installed and tested in this build:
- GPU enablement kit (risers, GPU power cables, GPU air shroud, mylar insulation) — Dell 330-BBMO / 85CYN in the US, 490-BEIX / JKFGX in EMEA
- Low-profile heat sinks in both CPU sockets — standard heat sinks physically foul the GPU air shroud, and Dell requires both sockets populated
- High-performance fan kit — mandatory for any GPU configuration in this chassis, and doubly so with 500W of passive GPU in the air path
- Dedicated GPU power cables — required for any card above 75W; each V100 draws 250W
- 1600W power supplies — Dell sets an 1100W floor for GPU configurations; 1600W supplies provide real redundancy headroom rather than a nominal pair
- 30°C maximum inlet temperature — the thermal ceiling once GPUs are installed alongside 150W/8C, 165W/12C, 200W or 205W processors
Assembled from separate secondary-market purchases, that is seven line items, seven condition claims and seven return windows. Here it is one part number, one warranty and one place to call. The full breakdown, including the 330-BBMO vs 490-BEIX part-number confusion and the two mistakes that stop most self-assembled builds: Putting a passive datacenter GPU in a PowerEdge R740xd
Frequently Asked Questions
Is this really 64GB of usable GPU memory?
It is 2 x 32GB, not a single 64GB pool. The Tesla V100 PCIe has no NVLink connector, so the cards communicate over PCIe Gen3 x16. Models that fit in 32GB run on one card at full speed; larger models must be sharded across both and pay PCIe bandwidth at each layer boundary.
Can I run a 70B language model on this server?
Yes, sharded across both cards at 4-bit quantization, but expect to invest effort in the software stack — several modern INT4 and attention kernels target Ampere and newer. For straightforward 70B serving, our A40 build holds the model in one 48GB pool with full modern kernel support.
Does the V100 support BF16 or FP8?
Neither. Volta tensor cores are FP16. BF16 arrived with Ampere and FP8 with Ada Lovelace and Hopper. FP16 training with loss scaling works well.
What is the V100's memory bandwidth compared to newer cards?
Up to 900 GB/s of HBM2 per card, which is higher than the 696 GB/s GDDR6 on a 48GB A40. For workloads that fit inside 32GB, the V100 is genuinely fast.
Can I add a third GPU?
The R740xd supports up to three double-width accelerators, and Dell requires all installed GPUs to be identical in type and model — so a third V100, not a mixed pair. Contact us first; power and riser assignments change.
Does it come with an operating system?
No OS is included. The server supports Windows Server, Linux distributions, VMware ESXi and Proxmox. iDRAC Enterprise is licensed for remote installation and management.
Can I expand the storage?
Yes. The chassis has 24 x 2.5" SFF bays and ships with one 960GB SATA SSD, leaving 23 bays open. The PERC HBA330 runs in pass-through mode, which suits ZFS, Ceph and software-defined storage.
What condition is it in and what warranty applies?
Professionally refurbished, fully tested and burned in, in excellent cosmetic and functional condition with minor signs of previous rack use possible. Backed by the Resilient Tec lifetime warranty, reachable directly rather than through a manufacturer ticket queue.
Further Reading
- Putting a passive datacenter GPU in a PowerEdge R740xd — every part the cards need, the enablement kit part numbers, and why a stock chassis throttles them
- What can you actually run on 48GB of GPU memory? — benchmarked model sizes and the KV cache math, useful for sizing either build
- Retrofitting legacy Dell & HPE servers for GPU workloads — the cable checklist for an existing chassis
- The cloud sent you a bill — on-prem compute versus rented GPU time
- GPU & AI infrastructure guide and enterprise servers guide