Customer evidence · Updated January 2026
1,840+ AI teams ship on Helios. Here’s what they shipped.
Proof, not promises. Three deep-dive evaluations from Replit’s EvalOps group, Notion AI, and Jasper — plus the long tail of 1,836 teams running production inference on the Helios runtime.
- 1,840+paying customers
- 2.1MSDK downloads
- 99.97%uptime, 12-mo
- a16zSeries A lead, Q2 2024
Case study 01 · Replit EvalOps
Replit cut their eval pipeline from four tools to one — and shipped 3.1× more evals per week.
Replit’s EvalOps group runs the regression suite behind every production LLM feature shipped to 30M+ developers. Their previous stack — Weights & Biases for experiment tracking, LangSmith for prompt tracing, BentoML for serving, and Helicone for cost observability — worked in isolation but broke the moment a single eval had to traverse all four.
Engineers spent more time wiring pipelines than writing assertions. The team adopted Helios to collapse those layers: training, evaluation, serving, and observability now share a single runtime with one telemetry schema, one auth model, and one SDK surface.
What changed in 90 days
- Eval throughput +312% — the team ran 412 evals / week before, 1,694 after.
- Time-to-first-eval collapsed from 9 days of glue code to 38 minutes on a fresh branch.
- Monthly infra spend down 41% after sunsetting three SaaS contracts.
- Single SOC 2 audit trail replaced four overlapping vendor reviews — audit prep went from 6 weeks to 4 days.
Case study 02 · Notion AI
Notion AI dropped p99 inference latency to 178ms on 70B-class models — and passed procurement in one cycle.
Notion AI’s platform team needed to serve a 70B-parameter model to millions of monthly active users without sacrificing tail latency. Their previous inference layer, a hand-tuned BentoML deployment on A10G hardware, hit a p99 floor around 410ms — slow enough that product teams were routing around the LLM feature entirely.
Helios’s latency-engineered inference layer was deployed alongside Notion’s existing fleet. The migration targeted commodity H100 clusters with the Helios batching and KV-cache scheduler, while keeping Notion’s data residency boundaries intact.
What changed
- p99 inference latency: 410ms → 178ms on a llama-3-70B-equivalent serving profile.
- Procurement cleared in a single cycle — Helios shipped a SOC 2 Type II report and HIPAA attestation under NDA within 24 hours, removing the two blockers that had stalled the rollout for a quarter.
- Capacity per dollar +2.3× on the same H100 footprint, by replacing per-tenant autoscaling with shared, batched inference.
- Zero customer-visible incidents during the 6-week dual-write migration window.
The long tail
1,836 more teams. Same runtime. Same audit trail.
The deep-dive case studies above cover three customers. The names below cover the breadth — monospaced deliberately, because you’re going to grep them anyway.
Names shown reflect customer logos published with permission. A full reference list is available under NDA for qualified procurement reviews.
Case study 03 · Jasper
Jasper shortened their fine-tune → eval cycle from two weeks to one morning — and 4.8×’d their experiment throughput.
Jasper’s applied research team runs hundreds of fine-tuning experiments per quarter on brand-voice and domain-specific adapters. Their legacy pipeline — Comet for experiment tracking, a custom Airflow DAG for orchestration, and Sagemaker for training — produced a two-week median iteration cycle, dominated by glue code and dataset shuffling between systems.
The team migrated fine-tuning, evaluation, and serving onto Helios in eight weeks. Training jobs, eval runs, and adapter promotion now share a single pipeline definition, so an engineer can fork an experiment, swap a dataset, and see the eval delta without leaving the SDK.
What changed
- Iteration cycle: 14 days → ~6 hours from “queue training” to “eval report in Slack.”
- SDK adoption 37 → 211 active developers in the first 90 days, with zero training sessions — the TypeScript and Python clients were familiar enough to drop in.
- 4.8× more experiments shipped quarter-over-quarter, with zero production rollbacks from the Helios-served adapters.
- Comet + Airflow + Sagemaker retired — a stack consolidation that freed two platform engineers to ship product.
Next step
Bring your stack. We’ll bring the migration plan.
A 30-minute demo with our solutions engineering team. Bring your current toolchain — W&B, LangSmith, BentoML, Helicone, custom Airflow, anything — and we’ll walk through a side-by-side migration plan on your model class and your scale.
- Side-by-side cost & latency comparison on your workloads
- SOC 2 Type II report and HIPAA attestation under NDA within 24 hours
- Reference calls with Replit, Notion AI, or Jasper on request
Or reach the team directly — [email protected] · +1 (415) 555-0184