Independent NVIDIA Hardware Reseller

AI infrastructure, explained before it's sold.

From a single desktop GPU to a full datacenter rack — every price, every watt, and every spec comes with a plain-language explanation of what it means and when it matters.

Browse the catalog Find my build
8
Products, desktop to rack-scale
3
Open-source models sized
2
Guided buying paths
Catalog
Hardware, compared plainly
Every product below shows its specs, price, and power draw — with a plain-language explanation of what each number means and when it matters.
Product GPU Memory Power (TDP) Form Factor Interconnect Price
NVIDIA GeForce RTX 4090Consumer / Prosumer
24 GB GDDR6X 450 W PCIe x16 (Desktop) PCIe 4.0 $2,400
NVIDIA RTX 6000 Ada GenerationProfessional Workstation
48 GB GDDR6 ECC 300 W PCIe x16 (Workstation) PCIe 4.0 + NVLink Bridge (2-GPU) $6,800
NVIDIA H100 PCIeData Center
80 GB HBM2e 350 W PCIe x16 (Server) PCIe 5.0 $25,000
NVIDIA H100 SXM5Data Center
80 GB HBM2e 700 W SXM5 (Needs HGX/DGX board) NVLink 4.0 (900 GB/s) $35,000
NVIDIA H200 SXMData Center
141 GB HBM3e 700 W SXM5 (Needs HGX/DGX board) NVLink 4.0 (900 GB/s) $40,000
NVIDIA DGX H100Complete AI Server
640 GB HBM2e (8 × 80 GB) 10,200 W 6U Rack Server NVLink 4.0 Full Mesh + 8× InfiniBand 400G $400,000
NVIDIA DGX H200Complete AI Server
1128 GB HBM3e (8 × 141 GB) 11,000 W 6U Rack Server NVLink 4.0 Full Mesh + 8× InfiniBand 400G $600,000
NVIDIA DGX GB200 NVL72AI Supercomputer Rack
13824 GB HBM3e (72 × 192 GB) 120,000 W Full Rack (72 B200 GPUs + 36 Grace CPUs) NVLink 5.0 (1.8 TB/s per GPU) $6,500,000
NVIDIA GeForce RTX 4090
Entry Point
Consumer / Prosumer
$2,400USD, list price
GPU Memory (VRAM)
24 GB GDDR6X
24 GB is enough to run a quantized 13–20B model in full. Not enough for serious 70B+ work.
Power Draw (TDP)
450 W
450 W is the peak draw — about the same as four gaming consoles running at once. Plan for a 750W power supply per GPU.
≈ 0.38 average homes  ·  drains 1 EV battery every 200.0 hrs
Form Factor
PCIe x16 (Desktop)
Interconnect
PCIe 4.0
The fastest consumer GPU ever made. Great for AI prototyping and running quantized models up to ~20B parameters at comfortable speed.
NVIDIA RTX 6000 Ada Generation
Best Value Pro
Professional Workstation
$6,800USD, list price
GPU Memory (VRAM)
48 GB GDDR6 ECC
48 GB ECC allows reliable inference of 34B models and, with 3 cards in one workstation, enough space for Llama 3.1 405B at Q4 quantization.
Power Draw (TDP)
300 W
300 W at full load — lower than RTX 4090 because workstation cards are optimized for sustained throughput, not peak bursts.
≈ 0.25 average homes  ·  drains 1 EV battery every 300.0 hrs
Form Factor
PCIe x16 (Workstation)
Interconnect
PCIe 4.0 + NVLink Bridge (2-GPU)
NVIDIA's flagship professional workstation card. Double the VRAM of RTX 4090 at lower power. ECC memory means no silent data corruption — critical for inference in production.
NVIDIA H100 PCIe
Pro Inference
Data Center
$25,000USD, list price
GPU Memory (VRAM)
80 GB HBM2e
80 GB HBM2e fits a quantized 65B model comfortably, or a full-precision 34B model. High bandwidth is what makes token generation feel fast.
Power Draw (TDP)
350 W
350 W in PCIe form means it fits in a server without special liquid cooling. Needs a datacenter-grade server rack though, not a desktop.
≈ 0.29 average homes  ·  drains 1 EV battery every 257.1 hrs
Form Factor
PCIe x16 (Server)
Interconnect
PCIe 5.0
NVIDIA's professional AI accelerator in a standard PCIe card. Fits in any modern server. HBM2e memory delivers 2 TB/s bandwidth — crucial for fast transformer inference.
NVIDIA H100 SXM5
Max Performance
Data Center
$35,000USD, list price
GPU Memory (VRAM)
80 GB HBM2e
Same 80 GB as PCIe H100, but the NVLink interconnect means multiple GPUs can share data so fast they nearly act as one large unified pool.
Power Draw (TDP)
700 W
700 W requires liquid cooling or specialized server infrastructure. You cannot put this in a standard desktop chassis.
≈ 0.58 average homes  ·  drains 1 EV battery every 128.6 hrs
Form Factor
SXM5 (Needs HGX/DGX board)
Interconnect
NVLink 4.0 (900 GB/s)
The SXM version of H100 requires a special server board but runs at 700W and connects to other H100s via NVLink at 900 GB/s — about 14× faster than PCIe. This is what makes a DGX server special.
NVIDIA H200 SXM
Large Model Hero
Data Center
$40,000USD, list price
GPU Memory (VRAM)
141 GB HBM3e
141 GB per GPU is the game-changer. A single DGX H200 (8 cards) gives you 1,128 GB — enough to run the world's largest open-source models with room to spare.
Power Draw (TDP)
700 W
Same 700 W power envelope as H100 SXM, but delivering far more memory and bandwidth per watt.
≈ 0.58 average homes  ·  drains 1 EV battery every 128.6 hrs
Form Factor
SXM5 (Needs HGX/DGX board)
Interconnect
NVLink 4.0 (900 GB/s)
The successor to H100 with 76% more memory and HBM3e delivering 4.8 TB/s of bandwidth. Each H200 alone can hold a quantized DeepSeek-R1 671B — 6 of them run it at full FP16 precision.
NVIDIA DGX H100
Enterprise Ready
Complete AI Server
$400,000USD, list price
GPU Memory (VRAM)
640 GB HBM2e (8 × 80 GB)
640 GB total across 8 GPUs, all reachable at NVLink speeds. Can run DeepSeek-R1 671B at Q4 quantization (needs ~201 GB) with abundant headroom.
Power Draw (TDP)
10200 W
10,200 W — about 8.5 average homes' worth of continuous power. Requires a dedicated 3-phase 30A circuit. Not for a regular office.
≈ 8.50 average homes  ·  drains 1 EV battery every 8.8 hrs
Form Factor
6U Rack Server
Interconnect
NVLink 4.0 Full Mesh + 8× InfiniBand 400G
A complete turnkey AI server from NVIDIA: 8 H100 SXM5 GPUs, a dual-socket Xeon CPU, 2 TB of system RAM, and a full NVLink mesh. Plug it in, and it's an AI powerhouse.
NVIDIA DGX H200
700B Model Ready
Complete AI Server
$600,000USD, list price
GPU Memory (VRAM)
1128 GB HBM3e (8 × 141 GB)
1,128 GB is more memory than the entire DeepSeek-R1 671B model needs at FP16 (805 GB). First time a single server can run 700B models without quantization.
Power Draw (TDP)
11000 W
11,000 W — slightly more than the DGX H100 due to the HBM3e subsystem. Plan for liquid cooling infrastructure.
≈ 9.17 average homes  ·  drains 1 EV battery every 8.2 hrs
Form Factor
6U Rack Server
Interconnect
NVLink 4.0 Full Mesh + 8× InfiniBand 400G
The DGX H200 can run DeepSeek-R1 671B and DeepSeek-V3 685B at full FP16 precision in a single box — no cluster needed. 1,128 GB of HBM3e is simply the most memory in one server available today.
NVIDIA DGX GB200 NVL72
Frontier AI
AI Supercomputer Rack
$6,500,000USD, list price
GPU Memory (VRAM)
13824 GB HBM3e (72 × 192 GB)
13,824 GB — enough to run over 17 copies of DeepSeek-R1 simultaneously, or to train models with trillions of parameters. This is the ceiling of commercial AI hardware today.
Power Draw (TDP)
120000 W
120,000 W (120 kW) — equivalent to 100 average homes running simultaneously. Needs a dedicated data hall with liquid cooling infrastructure, specialized power distribution, and raised flooring.
≈ 100.00 average homes  ·  drains 1 EV battery every 0.8 hrs
Form Factor
Full Rack (72 B200 GPUs + 36 Grace CPUs)
Interconnect
NVLink 5.0 (1.8 TB/s per GPU)
NVIDIA's Blackwell-generation AI supercomputer rack. 72 B200 GPUs linked by NVLink 5.0 into one 13,824 GB memory fabric. Used by hyperscalers to train and serve frontier models at scale.

Model Advisor
Open-source models, sized to hardware
Top open-source models (100B+ parameters, MIT or Apache-compatible license). For each, we show the exact memory math so you know what hardware you actually need.
Parameters (B) × 1 GB × 1.2 = Minimum GPU RAM at FP16 precision
Why 1.2×? The 20% overhead covers the KV cache (attention memory), activations, and system overhead during inference.
Q4 quantization stores weights in 4-bit instead of 16-bit, reducing memory by ~4×: divide the FP16 result by 4 for the Q4 estimate.
FP16 = full quality, large memory footprint.   Q4 = ~98% quality, 4× smaller footprint.
DeepSeek-R1
DeepSeek AI
MIT
Currently one of the top-ranked open-source reasoning models on the AI agent leaderboard. Matches or beats GPT-4 class models on math, coding, and science benchmarks. Released January 2025.
671B params × 1 GB × 1.2 =
→ 805.2 GB minimum GPU RAM (FP16)
→ 201.3 GB minimum GPU RAM (Q4 quantized, ÷4)
Minimum hardware (FP16)
10× H100 80GB (800 GB total) or 6× H200 141GB (846 GB total)
Minimum hardware (Q4 — budget option)
3× H100 80GB (240 GB total) or 2× H200 141GB (282 GB total)
DeepSeek-V3
DeepSeek AI
MIT
A Mixture-of-Experts (MoE) model — 685B total parameters but only ~37B activated per token. Top open-source chat model as of early 2025, strong at coding and reasoning.
685B params × 1 GB × 1.2 =
→ 822.0 GB minimum GPU RAM (FP16)
→ 205.5 GB minimum GPU RAM (Q4 quantized, ÷4)
Minimum hardware (FP16)
DGX H200 (1,128 GB) — only single-box option. Or 11× H100 PCIe clustered.
Minimum hardware (Q4 — budget option)
3× H100 80GB (240 GB total) or 2× H200 141GB (282 GB total)
Llama 3.1 405B
Meta AI
Meta Llama Community License (commercial use free under 700M MAU)
Meta's largest open model, competitive with GPT-4 on most benchmarks. The Apache-style community license makes it free to use commercially for virtually all companies.
405B params × 1 GB × 1.2 =
→ 486.0 GB minimum GPU RAM (FP16)
→ 121.5 GB minimum GPU RAM (Q4 quantized, ÷4)
Minimum hardware (FP16)
7× H100 80GB (560 GB total) or 4× H200 141GB (564 GB total)
Minimum hardware (Q4 — budget option)
3× RTX 6000 Ada 48GB (144 GB total) — this is our Small Startup path!

Guided Paths
Find the right build for your size
Two honest recommendations based on company size. Each ends in a concrete build with a real total price.
Small Startup
Small Startup Path
Cheapest honest way to run ONE big open-source model
Goal: Run Llama 3.1 405B at Q4 quantization
Q4 quantization reduces memory by ~4×, trading a small quality dip for dramatic cost savings. For most business use cases the quality difference is negligible.
Recommended Build
3× NVIDIA RTX 6000 Ada Generation (48GB each)
$20,400
High-end workstation chassis (e.g. Lenovo ThinkStation P920)
$5,000
Combined GPU Memory 144 GB
Memory needed (Q4 Llama 405B) 121.5 GB
Memory headroom +22.5 GB free
Total power draw 1,100 W
≈ average homes' power 0.92 homes
Drains 1 EV battery (90 kWh) in 81.8 hours
TOTAL PRICE $25,400
Cluster vs. stacking desktop cards: A cluster is multiple computers working together as one system to divide a task that's too large for any single machine.

These 3 GPUs live in ONE workstation connected via PCIe. This is NOT NVLink — data transfers between GPUs at ~64 GB/s (PCIe) vs 900 GB/s (NVLink). For inference, this is slow but functional. For training, you'll feel the bottleneck.
Mid-Size Company
Mid-Size Company Path
Capacity + room to grow — run the world's best open-source models
Goal: DeepSeek-R1 671B at full FP16 precision (no quality compromise)
Mid-size companies serving customers need consistent quality. Full precision matters. A DGX H100 gives you real NVLink speed, a proper NVIDIA software stack, and headroom to run 671B models with room for growth.
Recommended Build
NVIDIA DGX H100 (8× H100 SXM5, complete server)
$400,000
Combined GPU Memory 640 GB
Model needs at Q4 (DeepSeek-R1 671B) ~201 GB
Total power draw 10,200 W
≈ average homes' power 8.5 homes
Drains 1 EV battery (90 kWh) in 8.8 hours
TOTAL PRICE $400,000
Cluster vs. stacking desktop cards: A cluster is multiple computers working together as one system to divide a task that's too large for any single machine.

The DGX H100 has 8 GPUs connected by NVLink 4.0 in a full mesh — 900 GB/s between any pair of GPUs. This IS real clustering. For DeepSeek-R1 at Q4 (201 GB needed, 640 GB available), you'll have 439 GB of headroom to run multiple models or serve more users in parallel.
With 640 GB available and 671B FP16 needing 805 GB, you'd need Q4 quantization (~201 GB) for this box. For full FP16, step up to DGX H200 (1,128 GB, $600K). Still plenty of room to grow.

Get a Quote
Request a Quote
No payment. No account. Just tell us what you want and we'll get back to you.