Pricing
Run agents all day, on a flat monthly plan.
Pick a daily token allocation and stop metering every request. Use it yourself or share it across your team, from $100 a month.
Monthly token plans
A flat monthly fee, for you or your whole team
Pick a daily allocation and share it across as many people as you want. No per-seat pricing, no per-request metering.
Pro
One developer, or a couple of people sharing the pool
20Mtokens a day
- Every model on the API
- Drop-in for Claude Code, Codex, OpenCode, or any OpenAI-compatible harness
Cancel any time
Most popular
Max
A small team, or one developer running agents hard
100Mtokens a day
- Everything in Pro, with 5x the daily allocation
- Enough concurrency to fan out agents across a large codebase
Cancel any time
Team
A full engineering team on one allocation
1Btokens a day
- Everything in Max, with 10x the daily allocation
- Per-member API keys and a shared usage dashboard
- Direct line to our engineering team
Cancel any time
Every plan includes
- Use it solo, or invite your team and share the daily allocation between you
- Hard daily ceiling by default, so a runaway agent cannot surprise you
- Add credits to keep going past the ceiling at the per-token rates below
- Switch tiers or cancel at any time from the dashboard
Ceilings reset daily, so the monthly figure assumes a 30 day month. Need more than 1B tokens a day, or a dedicated endpoint? Talk to us about a deployment in your cloud below.
Pay per token
Or pay only for the tokens you use
Top up a balance and pay per token, with no plan and no commitment. These are also the rates that apply if you choose to run past a plan ceiling. We'll double your first top-up, up to $100.
| Model | Cached tokens | Input tokens | Output tokens |
|---|---|---|---|
| GLM-5.2 Marathon Open-source frontier coding model subconscious/glm-5.2 | $0.26 | $1.40 | $4.40 |
| GLM-5.3 Marathon Open-source frontier coding model subconscious/glm-5.3-marathon | $0.26 | $1.40 | $4.40 |
| DeepSeek V4 Flash Marathon High-throughput DeepSeek V4 subconscious/deepseek-v4-flash-marathon | $0.0028 | $0.14 | $0.28 |
| Qwen3.8 27B Marathon Multimodal post-trained model subconscious/tim-qwen3.6-27b | $0.15 | $0.30 | $3.00 |
Prices in USD per 1M tokens. With efficient caching, we see upwards of a 95% cache hit rate on agentic coding workloads, so most input tokens bill at the cached rate.
Pay 6.9× less than Opus 5,
and 3.0× less than standard GLM-5.2.
Standard inference vs. Subconscious, with runtime context compression.
Less Tokens
More Tokens
ShortShort
MedMedium
LongLong
XLExtra long
MaxMax
75K
375K
750K
1.5M
5M
Opus 5
$136.76
GLM-5.2
$59.65
GLM-5.2 Marathon
$19.88
Effective context
750K tokens
Compaction stalls
0
Tokens billed
187.9M → 62.4M
Context ceiling
150K
Frontier closed models, for comparison
| Model | Cached tokens | Input tokens | Output tokens |
|---|---|---|---|
| Claude Opus 5 | $0.501.9x | $5.003.6x | $25.005.7x |
| GPT-5.6 Sol | $0.501.9x | $5.003.6x | $30.006.8x |
Published list prices in USD per 1M tokens. The factor under each price is how much more expensive it is than our GLM-5.2 rate above.
Enterprise
Deploy in your cloud
Past a certain scale it is cheaper to own the serving. We install our inference system inside your VPC, on your GPUs, and charge a monthly fee per node instead of per token.
For enterprises
Half the GPUs, faster throughput, better capability, better economics
Our serving infrastructure runs more concurrent agents on the same hardware, so a node does the work that used to take two. You keep the savings, we charge a fraction of it, with hands-on setup by our team of world-class AI researchers.
50%
Same work with half the GPUs
3.5x
Faster token throughput
10x
Longer usable context
Expert
Support from our team
If you have GPUs
We install in your cloud
Already running GPUs? We deploy our inference system directly onto the fleet you have today, and your team is serving open models in your own cloud within a week.
If you need GPUs
We get you the hardware first
No spare GPUs? We connect you to our compute partners or work with your cloud provider of choice, secure the lowest cost per GPU we can, then install and run the system for you.
Estimate your savings
GPU price, per hour
Deployment size, GPUs
Generic serving
$186,880/mo
With Subconscious
$116,800/mo
50% the GPUs + our per GPU pricing*
You save
$70,080/mo
38% lower per month
* An estimate of our cost per node using our inference system, which covers everything we provide: our inference system, a routing gateway and supporting infrastructure, a frontend for API key management and usage monitoring, and dedicated setup and support from our team of world-class AI researchers.
Questions
Before you pick a plan
- What happens when I hit my daily ceiling?
- Requests stop for that model until the ceiling resets, so a runaway agent can never run up a surprise bill. If you would rather keep going, add credits and anything past the allowance bills at the per-token rates above.
- Is the ceiling shared across models?
- No. Every model on your plan gets its own daily ceiling, so running GLM-5.2 all day does not eat into what you can spend on the others.
- Can I change or cancel my plan?
- Yes, from the billing page in the dashboard. Upgrades take effect on your next request, and a cancellation runs to the end of the period you already paid for.
- Should I pick a plan or pay per token?
- Pick a plan if you run agents most days; the flat fee works out far cheaper per token than metered rates. Pay per token if your usage is spiky or you are still evaluating. You can start metered and move onto a plan whenever you want.
- Which harnesses does this work with?
- Anything that speaks the OpenAI or Anthropic Messages format, including Claude Code, Codex, OpenCode, Cline, and your own agent loops. Point the base URL at api.subconscious.dev and use your API key.
- What if I need more than the Team plan?
- We install the same inference system inside your own cloud, priced per GPU node. That path is covered above.