Skip to content

Customer evidence · Updated January 2026

1,840+ AI teams ship on Helios. Here’s what they shipped.

Proof, not promises. Three deep-dive evaluations from Replit’s EvalOps group, Notion AI, and Jasper — plus the long tail of 1,836 teams running production inference on the Helios runtime.

  • 1,840+paying customers
  • 2.1MSDK downloads
  • 99.97%uptime, 12-mo
  • a16zSeries A lead, Q2 2024
replit case study · EvalOps
4 → 1 fragmented tools consolidated into the Helios runtime
stack before W&B + LangSmith + BentoML + Helicone
eval throughput +312%
time-to-first-eval 9 days → 38 min
monthly infra spend −41%
“We replaced four vendors with one runtime and stopped hand-rolling dashboards at 2am. Helios gave EvalOps back its week.” — Priya Ramesh, Staff ML Engineer, Replit

Case study 01 · Replit EvalOps

Replit cut their eval pipeline from four tools to one — and shipped 3.1× more evals per week.

Replit’s EvalOps group runs the regression suite behind every production LLM feature shipped to 30M+ developers. Their previous stack — Weights & Biases for experiment tracking, LangSmith for prompt tracing, BentoML for serving, and Helicone for cost observability — worked in isolation but broke the moment a single eval had to traverse all four.

Engineers spent more time wiring pipelines than writing assertions. The team adopted Helios to collapse those layers: training, evaluation, serving, and observability now share a single runtime with one telemetry schema, one auth model, and one SDK surface.

What changed in 90 days

  • Eval throughput +312% — the team ran 412 evals / week before, 1,694 after.
  • Time-to-first-eval collapsed from 9 days of glue code to 38 minutes on a fresh branch.
  • Monthly infra spend down 41% after sunsetting three SaaS contracts.
  • Single SOC 2 audit trail replaced four overlapping vendor reviews — audit prep went from 6 weeks to 4 days.

Case study 02 · Notion AI

Notion AI dropped p99 inference latency to 178ms on 70B-class models — and passed procurement in one cycle.

Notion AI’s platform team needed to serve a 70B-parameter model to millions of monthly active users without sacrificing tail latency. Their previous inference layer, a hand-tuned BentoML deployment on A10G hardware, hit a p99 floor around 410ms — slow enough that product teams were routing around the LLM feature entirely.

Helios’s latency-engineered inference layer was deployed alongside Notion’s existing fleet. The migration targeted commodity H100 clusters with the Helios batching and KV-cache scheduler, while keeping Notion’s data residency boundaries intact.

What changed

  • p99 inference latency: 410ms → 178ms on a llama-3-70B-equivalent serving profile.
  • Procurement cleared in a single cycle — Helios shipped a SOC 2 Type II report and HIPAA attestation under NDA within 24 hours, removing the two blockers that had stalled the rollout for a quarter.
  • Capacity per dollar +2.3× on the same H100 footprint, by replacing per-tenant autoscaling with shared, batched inference.
  • Zero customer-visible incidents during the 6-week dual-write migration window.
Notion AI case study · inference
410 → 178ms p99 inference latency, llama-3-70B on H100
hardware 8x H100, us-east-1
capacity / dollar +2.3×
dual-write window 6 weeks, zero incidents
compliance docs SOC 2 II + HIPAA · NDA, 24h
“The procurement signal mattered as much as the latency signal. Helios handed us the audit package inside a day and let us retire a stack we’d been duct-taping for two years.” — Marcus Lin, Head of Platform, Notion AI

The long tail

1,836 more teams. Same runtime. Same audit trail.

The deep-dive case studies above cover three customers. The names below cover the breadth — monospaced deliberately, because you’re going to grep them anyway.

character.ai huggingface replicate scale modal together fireworks anyscale weights & biases langchain llamaindex pinecone weaviate chroma qdrant cohere mistral perplexity runway elevenlabs descript glean glean.ai harvey casetext ironclad tome gamma mem superhuman granola sourcegraph cursor codeium tabnine snyk vercel render fly.io railway supabase neon prisma inngest trigger.dev posthog statsig eppo amplitude mixpanel segment mParticle rudderstack hightouch census dbt hex sigma thoughtspot omni preset airbyte fivetran stitch

Names shown reflect customer logos published with permission. A full reference list is available under NDA for qualified procurement reviews.

jasper case study · fine-tuning + evals
14d → 6h/iter experiment iteration cycle, fine-tune → eval
SDK adoption 37 → 211 active devs (90d)
experiments / quarter +4.8×
legacy stack retired Comet + custom Airflow + Sagemaker
rollbacks (90d) 0
“Iteration speed is the only metric that compounds. Helios collapsed our cycle from two weeks to one morning, and we shipped four times more experiments last quarter than the prior one.” — Hana Ozturk, Director of ML, Jasper

Case study 03 · Jasper

Jasper shortened their fine-tune → eval cycle from two weeks to one morning — and 4.8×’d their experiment throughput.

Jasper’s applied research team runs hundreds of fine-tuning experiments per quarter on brand-voice and domain-specific adapters. Their legacy pipeline — Comet for experiment tracking, a custom Airflow DAG for orchestration, and Sagemaker for training — produced a two-week median iteration cycle, dominated by glue code and dataset shuffling between systems.

The team migrated fine-tuning, evaluation, and serving onto Helios in eight weeks. Training jobs, eval runs, and adapter promotion now share a single pipeline definition, so an engineer can fork an experiment, swap a dataset, and see the eval delta without leaving the SDK.

What changed

  • Iteration cycle: 14 days → ~6 hours from “queue training” to “eval report in Slack.”
  • SDK adoption 37 → 211 active developers in the first 90 days, with zero training sessions — the TypeScript and Python clients were familiar enough to drop in.
  • 4.8× more experiments shipped quarter-over-quarter, with zero production rollbacks from the Helios-served adapters.
  • Comet + Airflow + Sagemaker retired — a stack consolidation that freed two platform engineers to ship product.

Next step

Bring your stack. We’ll bring the migration plan.

A 30-minute demo with our solutions engineering team. Bring your current toolchain — W&B, LangSmith, BentoML, Helicone, custom Airflow, anything — and we’ll walk through a side-by-side migration plan on your model class and your scale.

  • Side-by-side cost & latency comparison on your workloads
  • SOC 2 Type II report and HIPAA attestation under NDA within 24 hours
  • Reference calls with Replit, Notion AI, or Jasper on request

Or reach the team directly — [email protected] · +1 (415) 555-0184