Pay for what you use. Nothing else.
Per-call billing for documents, images, audio, and video. No platform fees, no per-seat charges, no minimums.
Loved by leading
AI companies






Plans that scale with your usage.
Start free, upgrade when you need more throughput, custom models, or compliance.
Starter
- Pay-as-you-go
- Usage billed in USD
- $10 free signup balance
- Up to 10 requests/min
- Community Discord
- Basic Usage Logs
Pro
- Usage-based
- Usage billed in USD
- $1,000 usage included / month
- Up to 100 requests/min
- Dedicated Slack support
- Zero-Data Retention (ZDR)
- Business Associate Agreement (BAA)
Enterprise
- Invoiced Billing
- Tier-based pricing
- Volume discounts
- Custom rate-limits
- Dedicated Slack support
- In-VPC deployments
- Zero-Data Retention (ZDR)
- SOC 2, HIPAA, BAA
- Custom SLAs
Token Rates
Model usage priced per million tokens.
Rates are in USD per 1M tokens. Agent reasoning bills at the rate for the model ID you select; LLM-backed tools bill at the rate of the model that ran them, itemized per span.
Model | Inputper 1M | Outputper 1M | Cached inputof input rate |
|---|---|---|---|
vlmrun-orion-2:fastSpeed and cost-efficiency | $0.30 | $2.50 | 10% |
vlmrun-orion-2:autoRecommended default | $0.30 | $2.50 | 10% |
vlmrun-orion-2:proMaximum quality | $1.00 | $10.00 | 10% |
Fast and Auto share the same token rates; Pro is higher. Fixed-price tools follow the same tiers; see Agent Pricing below.
Examples
10K in / 2K out, auto
$0.008
(10K × $0.30 + 2K × $2.50) / 1M
10K in / 2K out, pro
$0.03
(10K × $1.00 + 2K × $10.00) / 1M
1M cached input, auto
$0.03
$0.30 × 10%
Single-tool run, auto
~$0.01-$0.03
Reasoning and tool dispatch
Single-tool run, pro
~$0.03-$0.08
Reasoning and tool dispatch
Model | Inputper 1M | Outputper 1M | Cached inputof input rate |
|---|---|---|---|
vlmrun-orion-2:qwen3.6-35b-a3bQwen3.6 35B A3B | $0.30 | $2.50 | 10% |
vlmrun-orion-2:gemma4-26b-a4bGemma4 26B A4B | $0.30 | $2.50 | 10% |
vlmrun-orion-2:cosmos3-nanoNVIDIA Cosmos3 Nano | $0.30 | $2.50 | 10% |
Open-weight backbones bill at the Fast / Auto token rates. Pin one by passing its model ID on chat completions or agent execute.
Examples
10K in / 2K out, any pin
$0.008
(10K × $0.30 + 2K × $2.50) / 1M
1M cached input
$0.03
$0.30 × 10%
Model | Inputper 1M | Outputper 1M | Cached inputof input rate | Thinkingper 1M |
|---|---|---|---|---|
vlmrun-orion-2:kimi-2.6Moonshot Kimi K2.6 | $0.66 | $3.41 | 22% | N/A |
vlmrun-orion-2:muse-spark-1.1Muse Spark 1.1 | $1.25 | $4.25 | 10% | N/A |
vlmrun-orion-2:gemini-flash-3.6Gemini Flash 3.6 | $1.50 | $7.50 | 10% | $7.50 |
vlmrun-orion-2:grok-4.5xAI Grok 4.5 | $2.00 | $6.00 | 25% | N/A |
vlmrun-orion-2:opus-4.8Claude Opus 4.8 | $5.00 | $25.00 | 10% | N/A |
vlmrun-orion-2:gpt-5.5OpenAI GPT-5.5 | $5.00 | $30.00 | 10% | N/A |
Hosted frontier backbones bill at model-specific pass-through rates rather than the tier rates. Fixed-price tools still bill at the resolved quality tier.
Examples
10K in / 2K out, kimi-2.6
$0.013
(10K × $0.66 + 2K × $3.41) / 1M
10K in / 2K out, opus-4.8
$0.10
(10K × $5.00 + 2K × $25.00) / 1M
10K in / 2K out, gpt-5.5
$0.11
(10K × $5.00 + 2K × $30.00) / 1M
1M cached input, grok-4.5
$0.50
$2.00 × 25%
1M thinking tokens, gemini-flash-3.6
$7.50
Billed at the Thinking rate
Cached input bills at the listed percentage of the input rate. Thinking tokens, where a model charges for them separately, bill at the Thinking rate.
Agent Pricing
Tools priced by the unit or by tokens.
GPU and API tools carry a fixed dollar price per unit. LLM-backed tools bill by the tokens they consume, at the rate of the model that runs them. Utility tools (image I/O, page navigation, video sampling) run free.
Capability | orion-2:fast / autoCost-efficient | orion-2:proMax quality |
|---|---|---|
SegmentObject segmentation, GPU-backed | $0.01 / image | $0.02 / image |
Generate, EditGenerate or edit images from prompts | $0.04 / image | $0.24 / image |
Caption, TagGenerate captions and tags for images | Token-billed | Token-billed |
Detect, PointLocate objects and points of interest | Token-billed | Token-billed |
UI ParsingParse UI elements from screenshots | Free | Free |
ToolsI/O, rotation, cropping, utilities | Free | Free |
Every run returns a per-span cost ledger: one line item per billable action, plus cost_dollars for the run total.
Examples
Segment 100 images
Fast $1.00 / Pro $2.00
Generate 10 hero images
Fast $0.40 / Pro $2.40
Edit 25 product images
Fast $1.00 / Pro $6.00
Segment then generate 1 image
$0.01 + $0.04 + ~$0.01 tokens ≈ $0.06
Caption 500 catalog images
Token-billed at the model rates
Parse 200 UI screenshots
Free
Capability | orion-2:fast / autoCost-efficient | orion-2:proMax quality |
|---|---|---|
OCR, LayoutExtract text and detect page structure | $0.01 / page | $0.04 / page |
Grounding, ConfidenceSurcharge when enabled on document.extract | $0.001 / page | $0.001 / page |
Parse, Extract, VQAExtract structured data and answer questions | Token-billed | Token-billed |
ToolsI/O, page navigation, utilities | Free | Free |
Every run returns a per-span cost ledger: one line item per billable action, plus cost_dollars for the run total.
Examples
OCR a 200-page report
Fast $2.00 / Pro $8.00
OCR a 1,000-page archive
Fast $10.00 / Pro $40.00
Extract 100 invoices
Token-billed at the model rates
Add grounding to 100 pages
Tokens + $0.10 surcharge
50-page doc, OCR + extract
$0.50 + tokens ≈ $0.52-$0.55
Same run at flex
≈ $0.26-$0.28
Capability | orion-2:fast / autoCost-efficient | orion-2:proMax quality |
|---|---|---|
Generate, EditGenerate from text prompts and input images | $0.15 / sec | $0.40 / sec |
Caption, Summary, TranscribeDescribe, summarize, and transcribe video | Token-billed | Token-billed |
ToolsI/O, sampling, trimming, utilities | Free | Free |
Every run returns a per-span cost ledger: one line item per billable action, plus cost_dollars for the run total.
Examples
Generate a 6-sec clip
Fast $0.90 / Pro $2.40
Generate a 15-sec ad
Fast $2.25 / Pro $6.00
Edit a 10-sec clip
Fast $1.50 / Pro $4.00
1-hour video, summarized
Token-billed at the model rates
Transcribe a 30-min webinar
Token-billed at the model rates
Capability | orion-2:fast / autoCost-efficient | orion-2:proMax quality |
|---|---|---|
Code executionOne sandboxed program run in code mode | $0.001 / call | $0.001 / call |
OrchestrationReasoning turns, tool dispatch, response synthesis | Token-billed | Token-billed |
ToolsMedia I/O, navigation, other utilities | Free | Free |
Every run returns a per-span cost ledger: one line item per billable action, plus cost_dollars for the run total.
Examples
Program-mode run, 3 executions
$0.001 × 3 = $0.003
Sandbox + extract + grounding
$0.001 + $0.002 + $0.001 = $0.004
Utility tool calls
Free
Service Tiers
Route each request for cost or latency.
The service tier controls delivery and scales the total dollar cost of a run. It is independent of the model quality tier (Fast / Auto / Pro).
Tier | Multiplier | When to use |
|---|---|---|
standarddefault | 1.0× | Base USD price. Applied by default when no tier is specified. |
flex50% off | 0.5× | Background agent runs and batch processing that tolerate variable latency. |
prioritypremium | 1.8× | User-facing interactions that need the lowest latency. |
The multiplier applies once, to the run total (model tokens + tools). Every response reports the effective cost, the standard-tier baseline, and the amount saved. For full details, refer to the documentation.
Examples
$1.00 run @ standard
$1.00
$1.00 × 1.0
$1.00 run @ flex
$0.50
$1.00 × 0.5
$1.00 run @ priority
$1.80
$1.00 × 1.8
Segment + generate @ flex
$0.03
$0.06 × 0.5
50-page OCR + extract @ flex
~$0.27
~$0.53 × 0.5
6-sec video generate @ priority
$1.62
$0.90 × 1.8
Pricing FAQs
Yes. Every account starts with a $10 USD balance and no card required. Upgrade to Pro or Enterprise whenever you need more throughput or compliance.
