Usage-based pricing · start free

Pricing

Per-token
usage-based billing
No limits
practically no rate limits
200 req
free every month

Join the teams behind 400+ production agents.

Read customer stories
JetBrains
Vercel
Onlook
Webflow
Databutton
Warp
Zo Computer
Anything
Inference at work today
Private tokens / dayProduction agentsFast Apply throughputDedicated availability
Private tokens / dayProduction agentsFast Apply throughputDedicated availability

Open Source Models

Frontier open weights, served on custom kernels for codegen.

Chat
Kimi K3 2.8TFeatured
morph-kimik3
Input$2.60/1M
Cache read$0.29/1M
Output$14.00/1M
Context1M
Kimi K3 2.8T, latency-tuned
morph-kimik3-fast
Input$6.00/1M
Cache read$0.60/1M
Output$22.50/1M
Context1M
GLM-5.3 744B
morph-glm53-744b
Input$1.25/1M
Cache read$0.26/1M
Output$4.40/1M
Context1M
GLM-5.3-Flash
morph-glm53flash
Input$0.10/1M
Cache read$0.02/1M
Output$0.34/1M
Context1M
DeepSeek V4 Flash 0731
morph-dsv4flash
Input$0.12/1M
Cache read$0.03/1M
Output$0.35/1M
Context1M

Specialized Models

Purpose-built APIs that offload the bottlenecks: code editing, search, context management, and per-turn classification.

Fast Apply
What is Fast Apply?
Fastest model
morph-v3-fast
Price
$0.80/1M in$1.20/1M out
Context262k
Most diversePopular
morph-v3-large
Price
$0.90/1M in$1.90/1M out
Context262k
Code Search
What is WarpGrep?
Fast context for agentsNew
morph-warp-grep-v2
Price
$0.80/100K
Context100K (1M for Pro)
Compaction
What is Compact?
Verbatim context compaction
morph-compact
Price
$0.20/1M in$0.50/1M out
Context1M
Reflex
What is Reflex?
Realtime per-turn classifiers
morph-reflex-v1
Price
$0.001/event$0.0005 over 1M/mo
Context64K
Batch (offline)
morph-reflex-v1
Price
$0.0005/event$0.00025 over 1M/mo
Context64K
Difficulty-based model routing
morph-router
Price$0.005/request
Context

Billing

Pay per token, or go flat-rate with Scale.

Pay as you go

Top up credits, use any model

From $10
First top-up+$5 free
LimitsPractically no rate limits
Get Started

Scale

For individuals who can't stop coding

$200/month
Credits40M($400)
LimitsPractically no rate limits

Dedicated Endpoints

Reserve B200 or B300 capacity by the GPU-hour, from $9.06 per reserved B200-hour, behind a 99.9% monthly availability SLA. Paying for GPU time instead of tokens runs up to 7× cheaper than DeepSeek API pricing.

2× B200

$8.64/GPU-hr

up to 6.6x cheaper than token-based pricing

  • 2× B200-equivalent reserved capacity
  • Dedicated endpoint with zero data retention
  • 99.9% availability SLA
  • Migration services
  • Slack support
  • Dedicated customer success
  • 24/7 Incident monitoring

4× B200

$8.43/GPU-hr

up to 6.8x cheaper than token-based pricing

  • Everything in 2× B200
  • 2× the reserved capacity
  • Lower GPU-hour rate
  • ~2× modeled token capacity

8× B200

$8.16/GPU-hr

up to 7.0x cheaper than token-based pricing

  • Everything in 4× B200
  • 2× the reserved capacity
  • Lowest GPU-hour rate
  • ~2× modeled token capacity
  • Full-node-equivalent capacity envelope

Compare DeepSeek V4 Flash 0731

Compare reserved capacity with token pricing.

Workload assumptions

Set utilization and cache share.

Capacity utilization100%
Input served from cache65%
Capacity vs. API

API reference: $0.14/M input, $0.0028/M cached, $0.28/M output.

Reserved capacityTokens / hrHourlyAPI pricingSavings
2× B200
$19.20/endpoint-hr
815.6M
725M in · 90.6M out
$19.20/hr
$0.02354/M
$0.07628/M
$62.22/hr
3.2× cheaper
4× B200
$37.48/endpoint-hr
1.63B
1.45B in · 181.2M out
$37.48/hr
$0.02298/M
$0.07628/M
$124.44/hr
3.3× cheaper
8× B200
$72.50/endpoint-hr
3.26B
2.9B in · 362.5M out
$72.50/hr
$0.02222/M
$0.07628/M
$248.88/hr
3.4× cheaper

Capacity or tokens?

Morph Dedicated vs. OpenRouter.

Economics
Billing unit
Reserved GPU-hour
Cost at sustained utilization
Falls with utilization
Full-use economics
Up to 7× cheaper
Idle capacity
Idle time billed
Credit purchase fee
None
Initial commitment
90 days
Best fit
Steady traffic
Platform
Model catalog
7 dedicated models
Serving path
Morph operated
OpenAI-compatible API
Spend controls
Reliability and privacy
Zero data retention
Default
Prompt and response logging
Off by default
Included with Morph Dedicated
Reserved B200 or B300 capacity
Isolated endpoint
No credit purchase fee
99.9% availability SLA
Migration services
Slack support
Dedicated customer success
24/7 endpoint monitoring

Dedicated FAQ

Billing, activation, ownership, and retention details.

How does hourly billing work?+

GPU hours set the monthly price. Each month is invoiced upfront, including idle capacity. Tokens are not billed.

Are tokens included or capped?+

No. You buy GPU time, not tokens. Calculator figures are estimates, not guarantees.

How is the GPU-hour price calculated?+

GPU rate times reserved GPUs times 730 hours. The month is billed upfront.

Is the physical hardware exclusive to me?+

No. Your endpoint, credentials, and capacity are isolated. Idle hardware may serve other traffic.

How quickly is an endpoint activated?+

Provisioning starts after payment. If activation exceeds two hours, Morph cancels and refunds the order.

What is the term and cancellation policy?+

The initial term is 90 days, then monthly with 30 days’ notice. No partial month refunds.

Does Morph retain prompts or responses?+

No. Morph stores only operational and billing metadata.

Can a coding agent deploy and manage an endpoint?+

Yes. Agents can choose, buy, track, list, and inspect endpoints through the Morph CLI. Purchases require user approval.

What does Morph manage?+

Morph handles provisioning, serving, monitoring, incidents, tuning, and migration. You own the app.

How can I monitor the endpoint?+

Use the dashboard for metrics and the CLI for status and metadata logs.

Is the API compatible with OpenAI clients?+

Yes. Change the base URL, key, and model name.

What does the availability SLA cover?+

It includes 99.9 percent monthly availability. Your agreement defines measurement and remedies.

Our fastest endpoints are private deployments

Over 100 billion tokens per day served on dedicated capacity with custom speculators and caching, at large discounts over the public pricing above. Includes SSO, custom rate limits, and priority support.

Get in touch