

Baseten
Baseten pricing meters a minute of instance time, and almost every minute a container is alive counts: image builds, cold starts and idle warm replicas all bill.
Updated on:
Baseten pricing: billed by the minute, cold start included
Baseten pricing meters a minute of instance time, and almost every minute a container is alive counts: image builds, cold starts and idle warm replicas all bill. A second meter charges per token on Model APIs, where GPT-OSS 120B costs $0.10 per million input and $0.50 output, the lowest rate for that model in this index. Those cheap tokens push the GPU decision further out.
Key takeaways
GPT-OSS 120B costs $0.10 per million input tokens and $0.50 output, undercutting the $0.15 and $0.60 Groq, Together AI, Fireworks and Bedrock all charge by 33% and 17%.
An H100 costs $0.10833 a minute, or $6.50 an hour, between Together AI's $5.49 dedicated rate and Fireworks' $8.00. Multi-GPU nodes scale linearly, with no volume break.
Image builds, cold starts and idle warm replicas all bill, and partial minutes round up. A replica that never boots is free, and so is one scaled to zero.
The billing API returns a "model distribution surcharge" beside compute cost, and neither the pricing page nor the docs publish its rate.
Baseten pricing in 2026
Meter | Unit | Rate | Notes |
|---|---|---|---|
Model APIs | Per 1M tokens | GPT-OSS $0.10/$0.50, GLM-5.3 $1.40/$0.14/$4.40 | GPT-OSS: no cache |
Dedicated GPU | Per minute | H100 $0.10833, A100 $0.06667, B200 $0.16633 | $6.50/$4.00/$9.98 hr |
Fractional GPU | Per minute | H100 MIG 40 GiB $0.0625 | $3.75 an hour |
CPU instances | Per minute | $0.00058 to $0.01382 | $0.03 to $0.83 hourly |
Basic plan | $0 a month | Everything, plus SOC 2 Type II and HIPAA | Pay as you go |
Pro and Enterprise | Custom | Priority GPUs, higher limits, self-hosting | Volume discounts |
What Baseten actually meters
Baseten meters the minute an instance spends running on a node, and the published per-minute figure is the hourly rate divided by 60, not the reverse: $0.10833 times 60 is $6.4998 against the $6.50 Baseten publishes. Price from the hourly number.
What counts as a running minute is where this gets specific, and Baseten documents it better than anyone else in this index. Image builds bill, because the build is its own workload. Cold starts and model loading bill, because the replica is already up. Idle warm replicas bill whenever min_replica is 1 or higher. Three things don't: image pulls, a failed boot, and a deployment scaled to zero.
The second meter is tokens, billed per million input and output, with cached input at roughly a tenth of input on most models and GPT-OSS carved out entirely. Two things live only in the docs. The instance reference carries an H200 at $0.125 a minute ($7.50 an hour) and an RTX-PRO-6000 at $4.00, neither of which appears on the pricing page. And the billing API returns surcharge_cost, a model distribution surcharge, on every dedicated line, which Baseten's own example puts at exactly 10% of compute.
How credits work
Credits exist, but they're a signup grant rather than a system. New workspaces receive credits for testing, Baseten applies them to the invoice before charging the card, and nothing needs redeeming. The amount isn't published in the docs or on the pricing page.
Past that grant, nothing credit-shaped is left. Baseten states plainly that it offers no separate free tier and no perpetual free plan. There's no expiry rule, no rollover, no top-up pack and no deduction order, because the balance isn't a wallet you refill. Running out means your card pays, and running out with no payment method means Baseten deactivates your models.
What happens when you hit the limit
Baseten charges you and invoices, which makes it the exception among the prepaid vendors here. An invoice issues when usage passes $50 or at month end, whichever comes first.
Budgets carry the sharp edge. A monthly budget emails at 75%, 90% and 100%, and by default it only notifies. Turning on Enforce budget rejects Model API requests once spend reaches it, but it never stops dedicated deployments or training jobs. An enforced budget caps the token meter only.
Rate limits run per account tier: an unverified Basic account gets 15 requests and 100,000 tokens a minute, a verified one 120 and 500,000, Pro 120 and 1,000,000. A 429 means you hit your limit, a 529 means Baseten has no capacity, and the second can happen inside the first.
How Baseten's pricing has changed
Date | Milestone | Source |
|---|---|---|
17 Apr 2026 | Cached input billed at a discount on every Model API except GPT-OSS | Vendor |
4 Mar 2026 | Billing usage API ships, splitting dedicated, Model API and training spend | Vendor |
1 Dec 2025 | Invoices move to the first of each month | Vendor |
21 May 2025 | Model APIs launch, adding a per-token meter beside per-minute compute | Vendor |
21 Mar 2024 | H100 MIG launches at $0.0825 a minute, $4.95 an hour. Now $3.75 | Vendor |
6 Feb 2024 | H100 arrives at $9.984 an hour, A100 at $6.15. Now $6.50 and $4.00 | Vendor |
1 Jul 2023 | Rates cut 40%. A10G goes to $1.207 an hour, unmoved since | Vendor |
Customer
Sentiment Highlights
"Now Baseten's pricing is cheaper than the official one? Probably won't last, but still interesting."
Developer comparing open-weight model hosts, Hacker News, August 2026
"on AA openai gets 117tps. baseten gets 284tps. so 18% more expensive but 142% more tps."
Developer benchmarking price against throughput, Hacker News, September 2026
Explore other providers

Firecrawl
Developer Tool
Firecrawl pricing runs on a single credit, and one credit buys one page on a basic scrape, crawl or map.
Modal
Infrastructure Platform
Modal pricing charges by the second for three resources at once: CPU cores, memory and GPU each carry their own rate, and a single function's bill is the sum.

Replicate
Infrastructure Platform
Replicate pricing charges per second of compute, and the per-second rate depends on which GPU your model runs on.
How much does Baseten cost per hour?
Does Baseten have a free tier?
Does Baseten charge for cold starts?
Is Baseten cheaper than Fireworks AI?
























