Signature experience
The Floor
A GPU or compute marketplace matches buyers who need accelerator work with sellers of GPU-hours, token streams, live host bids, or paid HTTP calls. It is not one catalog. Training buys synchronized capacity. Inference often buys a finished token stream. Peer markets discover a bid. Agents settle a bounded call. The Floor routes four tickets through those pits. This is an explainer, not procurement advice, and it does not invent prices.
The Floor
Four pits. Drop a ticket. Watch which market sells that unit.
examples dated 2026-07-12
Select a pit or drop a ticket. Dated catalog examples stay dated. Nothing here is a live quote.
Tickets
Capacity Pit
Allocated accelerator time
Synchronized accelerator time: one GPU, a node, or a cluster, usually billed by the hour and sometimes by the second.
- Delivers
- Allocated hardware. Interconnect, local disk, region, and the right to keep the job running are part of the purchase only when they are in the SKU.
- The quote excludes
- Tokens, latency SLAs, and finished answers. A cheap GPU-hour that sits idle, waits on storage, or cannot talk to its neighbors is not cheap work.
- Failure mode
- Quota, a missing interconnect, an idle reserved block, or a restore path that was never tested.
- Dated example
- On 12 July 2026, Runpod listed H100 PCIe pods at $2.89/hr and H100 SXM at $2.99/hr; Lambda listed 1x H100 PCIe 80 GB at $3.29/hr; CoreWeave Classic listed H100 PCIe at $4.25/hr. Those are dated catalog examples, not live quotes.
Keys when the board is focused: 1–4 select pits, Q W E R drop tickets, space walks or pauses, Esc clears the ticket. Respects reduced motion. Nothing is uploaded.
What a ticket proves
Routing a ticket shows which pit sells that unit, which pit is a possible second lane, and which pits would misprice the job by selling the wrong product. It does not pick a vendor.
What it does not prove
It does not prove a current price, a cheapest provider, or that a secondary pit is cheaper. Catalog numbers on this page are dated 2026-07-12 and must be re-read on the source page with region and SKU recorded.
What to do next
Take the surviving pit to a measured trial. Use the worksheets to normalize GPU-hour or token inputs. Then open the live provider page. The Floor stops before that step on purpose.
Four pits, not four vendors
The compute market is becoming tradable, which makes it look like a single board. It is not. A hyperscaler SKU, a GPU-first node, a Vast.ai listing, an Akash bid, a Claude or Gemini token meter, and an x402-paid HTTP resource can all be described as “buying compute.” They are not substitutes. They sell different claims about what will still be true an hour later.
Capacity Pit: $/GPU-hr
Synchronized accelerator time: one GPU, a node, or a cluster, usually billed by the hour and sometimes by the second.
Delivers. Allocated hardware. Interconnect, local disk, region, and the right to keep the job running are part of the purchase only when they are in the SKU.
Excludes. Tokens, latency SLAs, and finished answers. A cheap GPU-hour that sits idle, waits on storage, or cannot talk to its neighbors is not cheap work.
Fails when. Quota, a missing interconnect, an idle reserved block, or a restore path that was never tested.
On 12 July 2026, Runpod listed H100 PCIe pods at $2.89/hr and H100 SXM at $2.99/hr; Lambda listed 1x H100 PCIe 80 GB at $3.29/hr; CoreWeave Classic listed H100 PCIe at $4.25/hr. Those are dated catalog examples, not live quotes.
Read How compute is priced.
Token Window: $/1M tokens
A finished token stream from a managed model API. Input, cached input, output, batch, flex, and priority are often separate meters.
Delivers. An answer, not a named GPU. The provider absorbs hardware mix, scheduling, caches, and routing. The buyer pays for a product.
Excludes. Control of the accelerator, a guaranteed card generation, and any conversion from GPU-hour without a measured serving stack.
Fails when. Output-heavy prompts, retries, tool calls, cache misses, paid grounding, or a quality gap that forces a second model.
OpenAI, Anthropic, Google Gemini, and Amazon Bedrock all publish inference in per-million-token units. Gemini and Bedrock both document batch lanes that can cut token prices versus standard on-demand inference. Exact model rates move; open the linked pages.
Bid Board: live bid
Price discovery across hosts and providers: on-demand, interruptible, and reserved listings, plus specialized networks that price completed work.
Delivers. A listing, not a guarantee. The same GPU model can span a wide spread because reliability, region, verification, disk, and interruption policy differ.
Excludes. A durable median. Vast.ai describes rates as live, set by supply and demand. A static table should not freeze a marketplace median as a quote.
Fails when. An unverified host, a reclaim during a job that cannot checkpoint, sensitive data on the wrong trust tier, or a cheap card with the wrong CPU, disk, or network.
Vast.ai frames GPU prices as real-time platform rates across 40+ data centers, with per-second billing. Akash presents transparent hourly GPU marketplace pricing. Render Network prices rendering in OctaneBench-hours and states that 200 OBh equals one RTX 2070 for one hour. Checked 12 July 2026.
Read The marketplaces.
Call Desk: $/call
A machine-readable, bounded compute resource: a paid HTTP request, an MCP tool, a small inference call, an embedding batch, a render job, or an evaluation.
Delivers. Settlement inside the request flow. x402 revives HTTP 402 so a server can require payment and a client, including an agent, can pay programmatically.
Excludes. An eight-GPU cluster for a month. The first fit is a small unit. Larger autonomous procurement still needs spend caps, allowlists, identity, audit, and result checks.
Fails when. A loop that buys the same call, a malicious endpoint, plausible-but-wrong results, or data leaving a policy envelope.
x402.org displayed 75.41M transactions, $24.24M volume, 94.06K buyers, and 22K sellers for the last 30 days when checked 12 July 2026. Treat those as dated network metrics, not a GPU price.
Read Agent-native compute.
Why the tickets are shaped this way
Each ticket is a workload class, not a shopping list. The Floor sends it to the pit whose unit matches the job, lights a second pit when a real overlapping lane exists, and dims the pits that would charge for the wrong product. That last part is the teaching move. A cheaper token price cannot buy an eight-GPU interconnect. A live bid cannot, by itself, be a p95 SLA. A per-call resource cannot be a week of cluster time.
Eight-GPU training job
Several days on a synchronized eight-GPU node. NVLink or equivalent interconnect, local NVMe, checkpoint and restore, and a way to keep the job fed.
Training is a capacity-block market. The meter says GPU-hour. The purchase is a system that can fail and resume as one system.
Still measure: Cost per successful checkpoint, not sticker price per GPU-hour. Force an interruption in the trial.
Interactive chat, p95 bound
User-facing generation with a latency target. Burst traffic, cold starts, and retries are visible to a person. Data may be sensitive.
The buyer is asking for a latency-bounded product. Managed APIs sell that as tokens. Dedicated GPUs can sell it as capacity you operate. Interruptible listings are the wrong sole lane.
Still measure: Cost per accepted request at the target p95, including retries, tools, and cache misses. Average latency hides the bill.
Overnight embedding batch
Checkpointed, delay-tolerant embeddings or offline inference. The job can pause, move hosts, and resume. No user is waiting on p95.
Interruption is acceptable, so price discovery and interruptible discounts are rational. The Bid Board is where that trade is explicit.
Still measure: Cost per processed item after reclaim, restore, storage, and transfer. The lowest listing is not the denominator.
Agent, bounded paid call
Software needs to discover a compute resource, read a price, pay a cap, run one job, and continue. No human billing ceremony per provider.
The buyer is a machine. It needs a small unit, a machine-readable offer, and a policy envelope. That is the Call Desk.
Still measure: Cost per completed task, including retries, plus whether the agent can evaluate the output and the buyer can audit the spend.
Five findings
F1. The unit names the market
GPU-hour, token, live bid, and per-call settlement are not interchangeable stickers for the same product. If you cannot say which unit you are buying, you are not yet in a market. You are shopping a catalog.
Continue in Pricing units.
F2. Spot is a risk discount
Interruptible capacity is cheaper because reclaim is allowed. That is rational for checkpointed batch work. It is usually the wrong default for interactive inference unless another lane can absorb the reclaim without a user seeing it.
Continue in Spot and reserved.
F3. Token price hides hardware
A managed API price says nothing about which GPU served the request. That is the point. Compare cost per accepted task, not nominal dollars per million tokens, because output length, retries, tools, and quality change the denominator.
Continue in Per-token economics.
F4. Cheap listing is not cheap work
Host trust, interruption, region, VRAM, CPU, disk, network, storage, and egress all move the effective price. Normalize every quote to completed work: checkpoint, accepted request, processed item, or finished agent task.
Continue in How to read the table.
F5. Autonomy needs a policy envelope
x402-style rails make small compute purchases machine-payable. They do not make unbounded procurement safe. Spend caps, allowlists, identity, audit logs, data classification, and output verification are the market, not extras.
Continue in Agent-native risks.
How to read a compute marketplace
The walk is a rehearsal. The method transfers to any pricing page, including ones that did not exist when this site’s examples were checked.
- Name the workload. Write the job in operational terms: training, fine-tune, batch inference, interactive inference, render, or a single agent call. Include GPU count, latency, interruption tolerance, region, and data sensitivity.
- Read the unit. Find the meter the seller actually uses: GPU-hour, node-hour, token, request, live bid, or benchmark-hour. If two quotes use different units, they are not yet comparable.
- Name the risk. Quota, latency, host trust, interruption, and spend policy are not footnotes. They are why two rows with similar stickers belong in different pits.
- Drop the mismatched pits. A token API cannot sell NVLink. A live bid is usually the wrong sole lane for p95 chat. A per-call resource is the wrong product for a week of cluster time.
- Normalize to completed work. Convert the surviving quotes to cost per checkpoint, accepted request, processed item, or finished agent task. Use a measured trial, not a generic benchmark. The calculators on this site are worksheets, not quotes.
The GPU Price Compare table, the GPU Cost Estimator, and the Inference Throughput Cost Calculator are the worksheets that follow this read. They use the same dated snapshot. They still are not quotes.
Not procurement advice
ComputeMarket.io is an independent educational publication. It does not sell compute, take affiliate fees, or rank vendors. Provider order on this site is explanatory. The Floor will not tell you where to rent a GPU, which model API to call, or whether an agent should be allowed to spend. Those are procurement and policy decisions. This page exists so the market structure is legible before anyone makes them.
Prices move. Marketplace medians move faster. x402 network totals move daily. If a number appears here, it appears with an access date of 2026-07-12 and a first-party source. Re-open the source, set the region and SKU, and record a new date before treating the figure as current.
Answers
How does the GPU or compute marketplace work?
A compute marketplace matches buyers who need accelerator work with sellers of GPU-hours, token streams, live host bids, or paid HTTP calls. Hyperscalers and GPU-first clouds sell allocated capacity. Managed model APIs sell finished inference. Peer and decentralized networks discover a bid. Agent-native services settle a bounded request. The unit tells you which market you are in.
Where can I rent GPUs cheaply?
The lowest public listings are usually in peer marketplaces, community clouds, and interruptible tiers. Cheap is not a provider name. It is a workload, a trust tier, a region, and a measured cost per finished unit. The Floor does not rank vendors or invent a current cheapest GPU.
What is decentralized compute?
Decentralized compute aggregates hardware from many independent operators and exposes it through a marketplace or protocol. Akash is a bid-style deployment marketplace. io.net presents a DePIN-style GPU aggregation network. Render Network is a rendering market priced in OctaneBench-hours, not Render.com. Broader supply is the pitch. Security, completion, and region control remain the buyer’s problem.
How is inference priced?
Inference can be billed by GPU-hour, serverless worker time, request, input token, cached input, output token, batch job, priority lane, or paid tool call. Managed LLM APIs usually use per-million-token units. Self-hosted inference converts GPU-hour into cost per request only after you measure utilization, batching, caching, and latency.
Can agents buy compute autonomously?
Small purchases are becoming practical where a service can require payment over HTTP and a machine client can pay inside the same flow. Larger autonomous procurement still needs guardrails: spend limits, identity, provider allowlists, audit logs, and output verification. Autonomous means inside a policy, not unbounded.
Is The Floor procurement advice or a live price feed?
Neither. It is an explainer of market structure. Dated catalog examples on this site were checked 12 July 2026 and must be re-verified on the provider page, with region and SKU recorded, before anyone buys. ComputeMarket.io does not rank vendors and does not take procurement decisions.