Pricing

Pay for what you use. Nothing else.

Per-call billing for documents, images, audio, and video. No platform fees, no per-seat charges, no minimums.

Valerie Health
Scan.com
basata
Peak Health
York IE
Voxel51
Langflow
Activepieces
MCP.run
Make
Aurochs
MongoDB
n8n
Zapier
Google Cloud
Valerie Health
Scan.com
basata
Peak Health
York IE
Voxel51
Langflow
Activepieces
MCP.run
Make
Aurochs
MongoDB
n8n
Zapier
Google Cloud

Plans that scale with your usage.

Start free, upgrade when you need more throughput, custom models, or compliance.

Starter

$0/mo
  • Pay-as-you-go
  • Usage billed in USD
  • $10 free signup balance
Get Started for Free
  • Up to 10 requests/min
  • Community Discord
  • Basic Usage Logs

Pro

$799/mo
  • Usage-based
  • Usage billed in USD
  • $1,000 usage included / month
Sign Up
  • Up to 100 requests/min
  • Dedicated Slack support
  • Zero-Data Retention (ZDR)
  • Business Associate Agreement (BAA)

Enterprise

Custom
  • Invoiced Billing
  • Tier-based pricing
  • Volume discounts
Book a Demo
  • Custom rate-limits
  • Dedicated Slack support
  • In-VPC deployments
  • Zero-Data Retention (ZDR)
  • SOC 2, HIPAA, BAA
  • Custom SLAs

Token Rates

Model usage priced per million tokens.

Rates are in USD per 1M tokens. Agent reasoning bills at the rate for the model ID you select; LLM-backed tools bill at the rate of the model that ran them, itemized per span.

Model
vlmrun-orion-2:fastSpeed and cost-efficiency
Input · per 1M$0.30
Output · per 1M$2.50
Cached input · of input rate10%
Model
vlmrun-orion-2:autoRecommended default
Input · per 1M$0.30
Output · per 1M$2.50
Cached input · of input rate10%
Model
vlmrun-orion-2:proMaximum quality
Input · per 1M$1.00
Output · per 1M$10.00
Cached input · of input rate10%

Fast and Auto share the same token rates; Pro is higher. Fixed-price tools follow the same tiers; see Agent Pricing below.

Examples

10K in / 2K out, auto

$0.008

(10K × $0.30 + 2K × $2.50) / 1M

10K in / 2K out, pro

$0.03

(10K × $1.00 + 2K × $10.00) / 1M

1M cached input, auto

$0.03

$0.30 × 10%

Single-tool run, auto

~$0.01-$0.03

Reasoning and tool dispatch

Single-tool run, pro

~$0.03-$0.08

Reasoning and tool dispatch

Cached input bills at the listed percentage of the input rate. Thinking tokens, where a model charges for them separately, bill at that model output rate.

Agent Pricing

Tools priced by the unit or by tokens.

GPU and API tools carry a fixed dollar price per unit. LLM-backed tools bill by the tokens they consume, at the rate of the model that runs them. Utility tools (image I/O, page navigation, video sampling) run free.

Capability
SegmentObject segmentation, GPU-backed
orion-2:fast / auto · Cost-efficient$0.01 / image
orion-2:pro · Max quality$0.02 / image
Capability
Generate, EditGenerate or edit images from prompts
orion-2:fast / auto · Cost-efficient$0.04 / image
orion-2:pro · Max quality$0.24 / image
Capability
Caption, TagGenerate captions and tags for images
orion-2:fast / auto · Cost-efficientToken-billed
orion-2:pro · Max qualityToken-billed
Capability
Detect, PointLocate objects and points of interest
orion-2:fast / auto · Cost-efficientToken-billed
orion-2:pro · Max qualityToken-billed
Capability
UI ParsingParse UI elements from screenshots
orion-2:fast / auto · Cost-efficientFree
orion-2:pro · Max qualityFree
Capability
ToolsI/O, rotation, cropping, utilities
orion-2:fast / auto · Cost-efficientFree
orion-2:pro · Max qualityFree

Every run returns a per-span cost ledger: one line item per billable action, plus cost_dollars for the run total.

Examples

Segment 100 images

Fast $1.00 / Pro $2.00

Generate 10 hero images

Fast $0.40 / Pro $2.40

Edit 25 product images

Fast $1.00 / Pro $6.00

Segment then generate 1 image

$0.01 + $0.04 + ~$0.01 tokens ≈ $0.06

Caption 500 catalog images

Token-billed at the model rates

Parse 200 UI screenshots

Free

Service Tiers

Route each request for cost or latency.

The service tier controls delivery and scales the total dollar cost of a run. It is independent of the model quality tier (Fast / Auto / Pro).

Tier
standarddefault
Multiplier1.0×
When to useBase USD price. Applied by default when no tier is specified.
Tier
flex50% off
Multiplier0.5×
When to useBackground agent runs and batch processing that tolerate variable latency.
Tier
prioritypremium
Multiplier1.8×
When to useUser-facing interactions that need the lowest latency.

The multiplier applies once, to the run total (model tokens + tools). Every response reports the effective cost, the standard-tier baseline, and the amount saved. For full details, refer to the documentation.

Examples

$1.00 run @ standard

$1.00

$1.00 × 1.0

$1.00 run @ flex

$0.50

$1.00 × 0.5

$1.00 run @ priority

$1.80

$1.00 × 1.8

Segment + generate @ flex

$0.03

$0.06 × 0.5

50-page OCR + extract @ flex

~$0.27

~$0.53 × 0.5

6-sec video generate @ priority

$1.62

$0.90 × 1.8

Pricing FAQs

Yes. Every account starts with a $10 USD balance and no card required. Upgrade to Pro or Enterprise whenever you need more throughput or compliance.

Your first $10 is on us.