Skip to content
Pricing  /  Usage-based

Pricing built like the runtime: per-token, per-eval, no surprises.

Three tiers, transparent unit economics, and a live estimator so you can model the line items before you talk to anyone. When your usage moves into enterprise territory — custom MSAs, SOC 2 paperwork, BYOC — the fastest path is a 30-minute call with solutions engineering.

  • SOC 2 Type IIsince Mar 2024
  • 2.1MSDK downloads
  • 99.97%trailing uptime
  • a16z-led$19M Series A

Tier matrix

Three tiers. One runtime. Pick the row your GPUs would.

Developer is free up to a working monthly quota. Team is the default for production teams shipping LLM features in earnest. Enterprise is a contracted rate card with audit, procurement, and deployment knobs you actually need.

Developer

$0/ month, then usage

For solo engineers, prototypes, side projects, and CI smoke tests.

  • Up to 1M tokens / month free — inference + eval combined quota
  • 2 seats included, additional seats $29 / seat
  • Community support, GitHub-issues SLA, public status page
  • Single region (us-west-2), 7-day metric retention
  • Pay-as-you-go overage: $0.0000049 / 1K inference tokens
  • ×SOC 2 audit pack
  • ×BYOC / hybrid deployment
Start free →

Enterprise

Customannual commit

For org-wide rollouts: BYOC, audit, MSA, and a named solutions engineer.

  • Volume-priced inference, locked unit rates for the contract term
  • Unlimited seats & regions, custom retention windows
  • BYOC / hybrid deployment on your VPC, our control plane
  • Custom MSA, DPA, BAA, security review support
  • Named solutions engineer, 99.95% uptime SLA, 1h P1
  • Annual commit discount up to 18% on blended usage
  • HIPAA, PCI, FedRAMP-ready deployment paths
Talk to solutions engineering →
Inference unit rates $0.0000049 per 1K tokens · $1.31 per eval run · $0.18 per region-day Same unit economics on every tier — only the included quota and support ceiling change.

Cost estimator

Estimate your monthly bill in four inputs — no email required.

Use the worked example as a sanity check, then plug in your own numbers. The estimator uses the same rate card our billing engine uses — if the math here surprises you on a demo call, that’s our problem, not yours.

01

Monthly tokens

Sum of prompt + completion tokens across all inference calls in a 30-day window. The estimator auto-bins you into the right tier once you cross 250M.

Example 240,000,000 tokens — 12M prompts avg × 20 calls × 1k req/day
02

Model class

Helios prices inference by token, but model class affects the latency tier and whether eval pricing is per-run or per-step.

Example llama-3-70b on H100 — $0.0000049 / 1K tokens, p99 < 180ms
03

Eval runs / month

Each eval run bundles scoring + judge-model calls + trace capture. 5,000 runs are bundled on Team; above that it’s $1.31 per run.

Example 142 eval runs × $1.31 = $186.02 (all on Team quota)
04

Regions & retention

Active regions drive observability ingest at $0.18 / region-day. Retention over 30 days is $0.0042 / GB-day on cold storage.

Example 4 regions × 30 days = $21.60/day ingest baseline

Worked example · 240M tokens / month, llama-3-70b

How the line items stack

Inference 240M tokens × $0.0000049 / 1K $1,182.40
Eval runs 142 runs, all inside Team quota $186.02
Observability ingest 4 regions × 30 days × $0.18 $21.60
Seats 8 engineers, all inside Team quota $0.00
Estimated monthly bill Team tier · pre-annual-commit $1,616.40

Annual commit on Team unlocks a 12% blended discount — bring your forecast to the demo and we’ll model it live.

If your forecast exceeds 250M tokens / month

You’re an Enterprise conversation. Let’s get on a call.

Book a 30-minute demo

Consolidation math

Four vendors → one invoice. Show the line items.

The 2024 MLOps Community Migration Report verified that Helios replaces W&B + LangSmith + BentoML + Helicone for teams running LLM workloads. Here’s what a typical mid-market team was actually paying, vs. the unified Helios bill.

Before

Stack of four tools

W&B + LangSmith + BentoML + Helicone

  • Weights & Biases — team plan$1,650
  • LangSmith — Plus seats$1,470
  • BentoML — inference cluster$2,840
  • Helicone — observability$640
  • Glue engineering (FTE fraction)$1,800
  • Vendor management overhead$300
Stack total / month $8,700

Plus 4 procurement cycles, 4 security reviews, 4 on-call rotations.

After

Helios unified runtime

Train · Evaluate · Observe · Serve

  • Inference — 240M tokens @ $0.0000049 / 1K$1,182.40
  • Evaluations — 142 runs$186.02
  • Observability — 4 regions$21.60
  • Seats — 8 engineers, included$0.00
  • Glue engineering$0
  • Vendor management overhead$0
Helios total / month $1,390

One invoice. One security review. One on-call rotation.

4 → 1

vendors consolidated

84%

lower monthly run-rate

4.7×

faster ship velocity (verified)

24h

SOC 2 report turnaround

Procurement FAQ

The questions your finance team will ask before the call.

Written so your controller and your CISO can pre-validate without a synchronous meeting.

How do overages work, and what happens when we burst past our tier quota?

Overage is metered at the same per-token and per-eval-run rate you see on this page — no surge multipliers, no “contact sales” thresholds. We send a soft alert at 80% of monthly quota and a hard alert at 100%. You can set a hard spend cap per workspace and pause inference automatically when it trips. Team tier includes 250M tokens / 5,000 eval runs; usage above that is billed monthly in arrears.

What discounts are available for annual commits, and how is the commit sized?

Annual commits are sized off your highest-traffic 30-day window from the trailing 90 days, with a 1.2× multiplier for headroom. Discounts scale with commit size — 8% on Team-tier commits, up to 18% on Enterprise blended commits above 5B tokens / year. Unused commit can be rolled forward one billing cycle; overage beyond commit is billed at standard rates.

Can we deploy Helios inside our own VPC or on-prem, and what does BYOC actually mean?

BYOC (Enterprise tier) runs our control plane in our SOC 2 boundary and the inference + observability data plane inside your AWS, GCP, or Azure VPC. Your data never leaves your account; Helios operators access it via short-lived, just-in-time credentials. Hybrid is also supported — control plane in our tenant, data plane on-prem on your H100s. Typical provisioning is 2–3 weeks for a standard VPC layout.

Is HIPAA available, and what about custom MSAs, DPAs, and BAA paperwork?

HIPAA attestation was completed in November 2024 and is available on Team and Enterprise tiers via a signed BAA — typically a 3-business-day turnaround once redlines are exchanged. SOC 2 Type II reports (since March 2024) are available under NDA within 24 hours via the trust portal. Custom MSAs, DPAs, and security addenda are handled by solutions engineering and our standard redline-to-sign turnaround is 5 business days.

Still need to validate something specific for your security review?

Book a 30-minute demo

Solutions engineering

Need scale, SOC 2 paperwork, or a custom MSA? Talk to solutions engineering.

If your forecast pushes past 250M tokens / month, if you need a BAA, or if procurement needs a custom MSA — skip the pricing back-and-forth. 30 minutes, a screen share, and a real engineer who has shipped this before.

  • Solutions engineer responds within 1 business day
  • SOC 2 report under NDA within 24h
  • Custom MSA redline turnaround in 5 business days