Pricing

Run agents all day, on a flat monthly plan.

Pick a daily token allocation and stop metering every request. Use it yourself or share it across your team, from $100 a month.

Monthly token plans

A flat monthly fee, for you or your whole team

Pick a daily allocation and share it across as many people as you want. No per-seat pricing, no per-request metering.

Pro

One developer, or a couple of people sharing the pool

$100/ month

20Mtokens a day

Tokens per month600M
Concurrent requests4
  • Every model on the API
  • Drop-in for Claude Code, Codex, OpenCode, or any OpenAI-compatible harness
Get Pro

Cancel any time

Most popular

Max

A small team, or one developer running agents hard

$500/ month

100Mtokens a day

Tokens per month3B
Concurrent requests20
  • Everything in Pro, with 5x the daily allocation
  • Enough concurrency to fan out agents across a large codebase
Get Max

Cancel any time

Team

A full engineering team on one allocation

$5,000/ month

1Btokens a day

Tokens per month30B
Concurrent requests200
  • Everything in Max, with 10x the daily allocation
  • Per-member API keys and a shared usage dashboard
  • Direct line to our engineering team
Get Team

Cancel any time

Every plan includes

  • Use it solo, or invite your team and share the daily allocation between you
  • Hard daily ceiling by default, so a runaway agent cannot surprise you
  • Add credits to keep going past the ceiling at the per-token rates below
  • Switch tiers or cancel at any time from the dashboard

Ceilings reset daily, so the monthly figure assumes a 30 day month. Need more than 1B tokens a day, or a dedicated endpoint? Talk to us about a deployment in your cloud below.

Pay per token

Or pay only for the tokens you use

Top up a balance and pay per token, with no plan and no commitment. These are also the rates that apply if you choose to run past a plan ceiling. We'll double your first top-up, up to $100.

ModelCached tokensInput tokensOutput tokens
GLM-5.2 Marathon

Open-source frontier coding model

subconscious/glm-5.2

$0.26$1.40$4.40
GLM-5.3 Marathon

Open-source frontier coding model

subconscious/glm-5.3-marathon

$0.26$1.40$4.40
DeepSeek V4 Flash Marathon

High-throughput DeepSeek V4

subconscious/deepseek-v4-flash-marathon

$0.0028$0.14$0.28
Qwen3.8 27B Marathon

Multimodal post-trained model

subconscious/tim-qwen3.6-27b

$0.15$0.30$3.00

Prices in USD per 1M tokens. With efficient caching, we see upwards of a 95% cache hit rate on agentic coding workloads, so most input tokens bill at the cached rate.

Pay 6.9× less than Opus 5, and 3.0× less than standard GLM-5.2.

Standard inference vs. Subconscious, with runtime context compression.

Less Tokens

More Tokens

Short

Med

Long

XL

Max

75K

375K

750K

1.5M

5M

Opus 5

$136.76

GLM-5.2

$59.65

GLM-5.2 Marathon

$19.88

Standard inferenceSubconscious
Retained context per step0500KContext window limitContext compression ceiling1500 stepsRetained context tokens

Effective context

750K tokens

Compaction stalls

0

Tokens billed

187.9M → 62.4M

Context ceiling

150K

Frontier closed models, for comparison

ModelCached tokensInput tokensOutput tokens
Claude Opus 5$0.501.9x$5.003.6x$25.005.7x
GPT-5.6 Sol$0.501.9x$5.003.6x$30.006.8x

Published list prices in USD per 1M tokens. The factor under each price is how much more expensive it is than our GLM-5.2 rate above.

Enterprise

Deploy in your cloud

Past a certain scale it is cheaper to own the serving. We install our inference system inside your VPC, on your GPUs, and charge a monthly fee per node instead of per token.

For enterprises

Half the GPUs, faster throughput, better capability, better economics

Our serving infrastructure runs more concurrent agents on the same hardware, so a node does the work that used to take two. You keep the savings, we charge a fraction of it, with hands-on setup by our team of world-class AI researchers.

PricingMonthly fee per GPU node
DeploymentInstalled in your VPC
ModelsAny open model
SupportWorld-class AI expert support

50%

Same work with half the GPUs

3.5x

Faster token throughput

10x

Longer usable context

Expert

Support from our team

If you have GPUs

We install in your cloud

Already running GPUs? We deploy our inference system directly onto the fleet you have today, and your team is serving open models in your own cloud within a week.

If you need GPUs

We get you the hardware first

No spare GPUs? We connect you to our compute partners or work with your cloud provider of choice, secure the lowest cost per GPU we can, then install and run the system for you.

Estimate your savings

GPU price, per hour

≈ H100
≈ B200 / B300

Deployment size, GPUs

Generic serving

$186,880/mo

 

With Subconscious

$116,800/mo

50% the GPUs + our per GPU pricing*

You save

$70,080/mo

38% lower per month

* An estimate of our cost per node using our inference system, which covers everything we provide: our inference system, a routing gateway and supporting infrastructure, a frontend for API key management and usage monitoring, and dedicated setup and support from our team of world-class AI researchers.

Questions

Before you pick a plan

What happens when I hit my daily ceiling?
Requests stop for that model until the ceiling resets, so a runaway agent can never run up a surprise bill. If you would rather keep going, add credits and anything past the allowance bills at the per-token rates above.
Is the ceiling shared across models?
No. Every model on your plan gets its own daily ceiling, so running GLM-5.2 all day does not eat into what you can spend on the others.
Can I change or cancel my plan?
Yes, from the billing page in the dashboard. Upgrades take effect on your next request, and a cancellation runs to the end of the period you already paid for.
Should I pick a plan or pay per token?
Pick a plan if you run agents most days; the flat fee works out far cheaper per token than metered rates. Pay per token if your usage is spiky or you are still evaluating. You can start metered and move onto a plan whenever you want.
Which harnesses does this work with?
Anything that speaks the OpenAI or Anthropic Messages format, including Claude Code, Codex, OpenCode, Cline, and your own agent loops. Point the base URL at api.subconscious.dev and use your API key.
What if I need more than the Team plan?
We install the same inference system inside your own cloud, priced per GPU node. That path is covered above.

Get started

Start on a plan today, in your cloud tomorrow.

Pick a monthly plan and point your harness at our API in a couple of minutes, or talk to us about installing the whole system on your own GPUs.

© 2026 Subconscious Systems Technologies, Inc.

Subconscious