PRICING

Priced per token, committed once.

Every plan runs the same 300+ model catalog at the same rack rates, with end-to-end encryption on every connection — commitment is the only variable. Start at rack rate with nothing to sign. Or commit annual enterprise spend for a discount on every token: the balance decreases as your teams consume it, with dedicated deployments, customer-managed keys, and custom contracts when you need them. No platform fees, no markups, and no seats.

01PAY AS YOU GO

Rack rate on every model. Nothing to sign.

$0/ MO COMMIT
Metered · prepaid credits
None — pay for what you use
Closed · listOpen · list
Includes
  • 300+ models through one endpoint
  • API, CLI, and VS Code extension
  • Smart routing, failover, caching
  • End-to-end encryption on every connection
  • Zero data retention
  • Standard rate limits
02ENTERPRISE

Your model, your keys, your contract.

Custom
Annual · PO burn-down
Sized with your team
Closed −5%+Open −10%+
Everything in Pay As You Go, plus
  • Dedicated forward-deployed engineer
  • Implementation included, $0
  • Dedicated single-tenant deployment
  • Data residency
  • SAML SSO, SCIM, RBAC · audit logs
  • Zero data retention — contractual + DPA
  • PII removed before closed models
  • End-to-end encryption on every connection
  • Guaranteed TPM · custom rate limits
  • Dedicated Remote Agent runners · SSO
  • Agents API — cloud coding agents
  • Custom SLAs · volume discounts

Every plan is quoted against the per-token rates below as a reference point. The rate on your contract improves as committed spend grows.

03COMPARE

Everything, by plan.

Pay As You GoEnterprise
Products & surfaces
Core surfaces — API, CLI, VS Code
Agents API — cloud coding agents
/multi-agent + Chairman LLM orchestration✓ metered✓ metered
Remote Agent — cloud sandboxes
Remote Agent — dedicated runners, SSO
Platform
300+ models, one endpoint, one bill
End-to-end encryption on every connection
Smart routing, failover, caching
Rate limitsStandardCustom
Trust & operations
Zero data retention✓ contractual
PII removed before closed models
SAML SSO, SCIM, RBAC
Audit logs
Uptime SLACustom
Engineering & support
Forward-deployed engineeringDedicated FDE
ImplementationSelf-serveIncluded, $0
SupportDocsCustom SLAs
Deployment
Cloud API
Dedicated single-tenant deployment — Enterprise Inference
Data residency

We suppress retention and training on routed traffic through provider terms and per-request flags, wherever the provider API supports it. Blackbox runs dedicated deployments on single-tenant infrastructure, isolated from every other customer.

04PER-MODEL RATES

Transparent token pricing

Priced per 1M tokens. We bill input, output, and cached reads separately. These are list rates — we quote contracts against them, and your rate improves as committed spend grows.

MODELTYPECONTEXTINPUT / 1MOUTPUT / 1MCACHE READ / 1M
zai
glm-5.2
Text1M$1.40$4.40$0.26
moonshotai
kimi-k2.7-code-highspeed
Code262K$1.90$8.00$0.38
moonshotai
kimi-k2.7-code
Code262K$0.74$3.50$0.15
nvidia
nemotron-3-ultra-550b-a55b
Text1M$0.32$0.80$0.08
alibaba
qwen3.7-plus
Text1M$0.32$1.28
minimax
minimax-m3
Text1M$0.60$2.40$0.12
$0.44
BLENDED / 1M ON NEMOTRON 3 ULTRA · #1 SPEED
2.7×
CHEAPER THAN THE #2 FASTEST PROVIDER
PER TOKEN
METERED PER TOKEN, NOT PER SEAT
300+
MODELS, ONE ENDPOINT
05FAQ

Questions, answered

How does the commit work?

You commit to a token volume through a purchase order, and the balance decreases as your teams consume it, metered at the published per-model rates above. Commits are sized so that a director or VP can sign without an approval chain, and then expand. You can watch the balance in real time in the dashboard — and intelligent routing and prompt caching extend the same commit 10–20%, even on closed models.

What does the commit cover?

Every surface uses the same balance: Enterprise Inference, the Blackbox Router, the API, the CLI, the Agents API, and the VS Code extension. Metering is per token, not per seat, and the contract specifies how seats are treated.

What happens to unused commitment?

It expires at the end of the billing period. We send spend alerts at 75% and 90% of your commit, so nothing expires without warning. Size the commit to your minimum usage.

Is there a platform fee?

No. There is no credit-purchase fee and no per-seat charge. You pay for tokens at or below the list rates.

Can I connect my own provider accounts?

On Enterprise, yes. Connect your OpenAI, Anthropic, or Google accounts and route through your existing provider relationships and negotiated rates. Routing, failover, caching, and the dashboard remain available, while usage bills through your provider.

Why is the open-weight discount bigger?

Open-weight models run on infrastructure that Blackbox operates end to end, which includes the Nemotron deployment that Artificial Analysis benchmarked #1 on output speed (July 2026). Closed models route to their providers, and we pass our scale discounts to you.

What is included beyond the tokens?

The trust stack: end-to-end encryption, encryption at rest, the PII anonymization layer, custom SLAs, and a forward-deployed engineer who configures the account with you at contract start. See the security overview for the full detail.

What does a forward-deployed engineer actually do?

The engineer writes code in your stack: migration from your current provider, production integrations, eval harnesses, and routing and latency tuning. On Enterprise, the engineer is dedicated to your account, and implementation is included at no added cost.

When do I have to talk to sales?

Only for a commit — pay as you go needs nothing signed and no conversation. Enterprise exists for committed discounts, dedicated single-tenant deployments, data residency, and custom contracts.

How does an engagement start?

With a conversation, not a signup. We scope the workload and the volume with you and agree a commit. A forward-deployed engineer then configures the account, with any controls, such as PII removal, that your security review requires. Talk to sales to start one.

Scope a deployment with our team.

Tell us the workload, the controls that it must satisfy, and the volume that you expect. We reply with a commit and a deployment plan — not a form reply.