Companies can rent NVIDIA GPU capacity from cloud providers by the hour, but that is just one of six routes: hyperscalers, large neoclouds, smaller specialist clouds, customer-owned GPUs operated by Oracle, negotiated capacity deals, and servers in a company’s own data center. The choice comes down to how quickly and at what scale capacity is needed, how much control the team wants, and who will handle the hardware.
For US buyers comparing listed rates, Lambda advertises one-GPU H100 SXM at $4.29 per GPU-hour, H100 PCIe at $3.29, A100 40GB at $1.99, and B200 SXM6 at $6.99. Applicable taxes may apply. Those rates are tied to distinct GPU configurations, not a like-for-like performance ranking.
Six ways to get NVIDIA GPU capacity
A hyperscaler or cloud provider supplies and operates the GPUs as part of its cloud service. A neocloud is a specialist cloud provider focused on GPU-heavy computing. At the other end of the spectrum, a company can buy and run servers in its own data center. Between those options are specialist rentals, customer-owned GPUs run by an operator, and large negotiated capacity agreements.
| Access route | Who supplies or operates the hardware | What the route offers | Main practical trade-off |
| Hyperscaler cloud | The cloud provider supplies and operates the GPUs. | GPU access alongside managed cloud services and familiar enterprise infrastructure. | Capacity may not match a customer’s timing or scale. |
| Flagship neocloud | The provider operates GPU-focused infrastructure. | Large blocks of specialist GPU capacity. | Large deployments can require upfront payments and months to come online. |
| Smaller or regional specialist cloud | The provider supplies GPU capacity; some offers provide bare-metal access. | Potentially flexible options and more direct access to the machine. | Customers may take on more technical work. |
| Customer-owned hardware operated by Oracle | The customer supplies the GPUs; Oracle operates the infrastructure. | A way to use owned GPUs when the customer lacks power, data-center space, or skilled staff. | The customer must first buy the GPUs. |
| Negotiated capacity deal | A provider commits capacity under a large agreement. | Capacity at a scale that can serve very large requirements. | The scale and commitment can exceed what many startups can afford. |
| On-premises servers | The company buys and operates its own GPUs. | Direct control over its hardware and deployment. | Requires capital, power, space, and staff to run the systems. |
Hyperscalers and large neoclouds
Major cloud platforms such as Amazon, Microsoft, and Google combine GPU access with broader managed services. NVIDIA describes cloud access through major cloud platforms and NVIDIA Cloud Partners. For teams already using those ecosystems, this route can keep compute close to other cloud services and familiar enterprise infrastructure.
The trade-off is capacity: large cloud providers may not have enough GPUs available for every customer at the required time or scale. Large neoclouds offer another source of GPU-heavy infrastructure, but major deployments can involve upfront payments and months of lead time. That makes them a different proposition from starting a single hourly instance.
Specialist clouds, marketplaces, and hourly rates
Smaller or regional providers may offer flexible access, including bare-metal GPUs. Bare metal means the customer works more directly with the server rather than relying on a fully managed software layer. That can suit teams that want control over their environment, but it also puts more setup and operating work on their own engineers.
Lambda’s listed US prices provide a concrete hourly reference for four single-GPU configurations:
| Provider | One-GPU configuration | Listed price | Billing unit and condition |
| Lambda | H100 SXM | $4.29 | Per GPU-hour; applicable taxes may apply. |
| Lambda | H100 PCIe | $3.29 | Per GPU-hour; applicable taxes may apply. |
| Lambda | A100 40GB | $1.99 | Per GPU-hour; applicable taxes may apply. |
| Lambda | B200 SXM6 | $6.99 | Per GPU-hour; applicable taxes may apply. |
The model and configuration matter: an H100 SXM, an H100 PCIe, and a B200 SXM6 are different offers, and the hourly rate alone does not show how a particular workload will perform or what its total cost will be.
A provider-comparison walkthrough also discusses bandwidth charges on Vast.ai and contrasts the features and workload focus of Vast.ai, Runpod, and Modal. Its provider judgments are qualitative rather than standardized performance measurements.
Bring your own GPUs to an operator
Oracle’s bring-your-own-hardware model separates GPU ownership from infrastructure operations: the customer supplies the GPUs, and Oracle operates them. It is aimed at organizations that can purchase hardware but lack the power, data-center space, or skilled staff to run it themselves.
This approach is not a way around the hardware purchase. It shifts the operating work to a provider while leaving the customer responsible for supplying the GPUs.
Large capacity agreements and on-premises servers
Negotiated capacity deals serve a different scale from ordinary hourly rental. A reported SpaceX arrangement with Anthropic was valued at $1.25 billion per month through mid-2029. That example illustrates the scale of a major capacity commitment, rather than a typical way for a small team to rent one GPU.
Owning and operating GPUs on premises gives a company direct control, but demands more than buying the accelerators. The organization also needs adequate power and data-center space, plus people who can install and operate the infrastructure. Companies such as Dropbox and Everpure have been cited as examples of businesses running their own GPUs; their cases do not establish that on-premises systems are cheaper for every organization.
Choose by workload and operating capacity
For a team that needs managed cloud services and already works with a major cloud platform, a hyperscaler is a natural route to consider. If the immediate need is GPU-heavy capacity beyond what that provider can supply, a neocloud may be a better fit, with larger deployments bringing longer lead times and possible upfront commitments.
A smaller specialist cloud can suit a team prepared to take on more of the technical environment, while Oracle’s model fits a different situation: the company owns GPUs but needs an operator and data-center resources. On-premises infrastructure is most relevant when an organization can support the hardware, facilities, and operations itself. A large negotiated deal belongs to requirements measured at a much bigger scale than a short hourly rental.
For a short job or an uncertain workload, an hourly listing makes the unit cost easy to see. It does not, by itself, answer questions about capacity timing, total workload cost, or whether that GPU configuration fits the task.