Demand for accelerators still outruns supply. Here is how allocation really works this year — and how buyers get burned when they try to shortcut it.
How allocation works
OEMs receive quarterly GPU allocations from NVIDIA and AMD, and distribute them across validated orders. Your place in that queue is determined by three things: a complete bill of materials, a signed PO, and a delivery site that is actually ready. Incomplete orders quietly slide to the next quarter.
Realistic lead times right now
- H100 / H200 HGX systems — 6–10 weeks through major OEMs.
- B200 platforms — 12–20 weeks, allocation-gated; Q4 windows are filling now.
- L40S / RTX 6000 Ada — near stock; days, not months.
- AMD MI300X — improving fast; often the shortest path to large memory per GPU.
The grey market: what it really costs
Cards without OEM provenance arrive without enterprise warranty, without firmware support contracts, and increasingly without driver entitlements for cluster-scale features. We have audited grey-market batches with re-flashed consumer parts and mismatched steppings that never ran a stable NCCL ring. A 15% discount on hardware that fails burn-in is not a discount.
Buy the queue position, not the rumor: a validated order with a real OEM beats a container of promises.
What we do for buyers
PEXON holds standing allocations with Supermicro, Dell and Lenovo, verifies provenance on every unit, and burns in complete systems before shipment — so the GPUs you pay for are the GPUs that train.