The Unified Gateway for Visual Intelligence.

Understand, reason over, and act on images, video, and documents with Orion, our flagship visual agent.

  • 0:00Dual-arm robot at a tabletop workspace with a padlock and key.
  • 0:04Right gripper closes on the key; left arm waits at the fixture.
  • 0:16Left arm steadies the padlock while the key is brought into line.
  • 0:21Key seated in the keyhole — unlock sequence begins.
Valerie Health
Scan.com
basata
Peak Health
York IE
Voxel51
Langflow
Activepieces
MCP.run
Make
Aurochs
MongoDB
n8n
Zapier
Google Cloud
Valerie Health
Scan.com
basata
Peak Health
York IE
Voxel51
Langflow
Activepieces
MCP.run
Make
Aurochs
MongoDB
n8n
Zapier
Google Cloud

Documents

Document intelligence for the agentic stack.

Parse, extract, cite, and redact. Schema-validated JSON with grounding, ready for agents to act on.

Learn more

Parse

Read documents like a human would, capturing layout, structure, and meaning. Tables, figures, multi-page layouts, and handwritten notes come back as structured content, not a wall of text.

Extract

Extract fields into typed JSON against a Hub domain or your own schema. Invoices, forms, and disclosures come back ready for your systems, with no cleanup pass.

Cite

Ground every field with a bounding box and confidence score. Reviewers click straight to the source on the page, with no need to re-read the file.

Redact

Blur PII and PHI in place, or swap them for consistent dummy values when the file still needs to parse. Names, dates, account and member numbers, wherever they appear. Built for HIPAA, GDPR, PCI DSS, and more.

Videos

Every video searchable, clip-ready, understood.

Understand, summarize, segment, and search video — switch capabilities and click results to seek.

Learn more
0:00 / 0:53
Pick up key0:00

Use cases

Built for real documents and footage.

Faxes, drawing sheets, robot demos, unlabeled archives. Same call, same structured output.

Healthcare

Handle low-quality faxes, handwritten notes, and complex medical forms with ease. Parse rotated, skewed, or noisy scans into structured records. HIPAA-ready and trusted in production.

Learn more

Pricing calculator

Document and video, priced by the unit.

Pick a workload and volume. See savings versus closed APIs and frontier models — pages for documents, hours for video.

Page volume / month

Gateway OCR tokens (in / out)

120M - 250M / 100M

Frontier billed tokens (in / out)

250M / 200M

2.5K image tokens per page; output includes reasoning at 2× OCR text.

Total costwhat the gateway bills for this workload

$23 - $90

Total cost savingsvs the cheapest frontier model

$848 - $915

Documents / month

100Kpages

Cost comparison (log scale)

Cheapest - most expensive
VLM Run Gateway
Document OCR VLMs
$23 - $90
OCR APIs
Textract, Azure Doc AI
$150 - $1K
Document AI APIs
Reducto, LlamaParse, Extend
$1K - $6K
Frontier VLMs
Gemini, Claude, GPT
$938 - $7.3K
$10$100$1K$10K

Estimates vs typical OCR, Document AI, and frontier VLM pricing.Source: llm-prices.com

12345678910
from vlmrun.client import VLMRun

client = VLMRun(api_key="<VLMRUN_API_KEY>")
response = client.agent.execute(
    prompt="Parse this doc in JSON {'invoice_no': ..., 'total': ...}",
    inputs=["invoice.pdf"],
)

print(response.response)
# {"invoice_no": "1042", "total": 1200.00}

One call, whether you are a developer or an agent.

Parse documents and understand video from your app, or hand the same catalog to Pydantic AI, Mastra, Claude Code, or Codex.

For Enterprises

The new visual intelligence layer for your enterprise.

Deploy securely inside your VPC or private cloud – bringing visual intelligence directly to your infrastructure. Power document, image, and video understanding across teams. SOC 2 Type II and HIPAA-ready.

HIPAA Compliant
AICPA SOC

Frequently asked questions

An OpenAI-compatible API for document and video intelligence. Parse, extract, cite, and redact documents; summarize, transcribe, and search video — then hand the same catalog to agents over MCP.

Try VLM Run free today.