Match the output
Use the same model version, quality target, precision, video resolution and duration. A base model and an enhanced hosted pipeline are different products.
Your inference service should open possibilities. Compare usage-based API spending with dedicated inference capacity operated by Perchy.
$30,000 in monthly API spend. $1,000 in estimated monthly service costs. That is 1/30 of the spend—and an illustration of what to investigate, not a promised result.
Compare usage-based API spending with an estimated monthly fee for equivalent managed inference and reserved capacity.
of your current API spend
* Illustrative scenario, not a quote or measured benchmark. Adjust all inputs to your workload. No savings or equivalent performance are guaranteed. Perchy operates the servers; reserved capacity is delivered through authenticated inference APIs.
Get a workload-specific assessmentA measured 5.17-second video took 93.19 seconds of generation. At the provider’s reference output-only rate of $0.08/second for 768p, the output would cost about $0.4133.
Reaching 1/30 of that requires an effective generation-compute cost of about $0.532 per GPU-hour or less. Explore the assumptions below.
Provider pricing referenceThis models internal inference-serving costs, not a customer GPU rental offer or a Perchy API service price. Excludes operations, idle/cold-start overhead beyond the selected utilisation, storage, transfer, licences, failed jobs and margin. Local INT8 + Turbo8 quality is not assumed equivalent to the full hosted pipeline.
Monthly service estimate = reserved inference capacity + service operations and licences. Annual difference = (current API spend − service estimate) × 12. The ratio is rounded for display. All amounts are in USD, exclude taxes and are user-controlled estimates. Capacity is delivered through Perchy's managed inference APIs; these inputs are not GPU rental prices or server access charges.
A useful cost comparison accounts for the entire workload—not just a GPU-hour.
Use the same model version, quality target, precision, video resolution and duration. A base model and an enhanced hosted pipeline are different products.
Account for reserved inference capacity, service operations, storage, transfer, model licences and support in the agreed API service fee.
Utilisation matters. Batch throughput and interactive latency describe different workloads; neither should be substituted for the other.
Perchy-operated H3-Base inference and the official hosted 2K pipeline have different component boundaries. A fair comparison must specify which generation, preprocessing, reference-input and regeneration costs are included. No measured 1/30 H3 saving is asserted on this website.
See the model provider’s current pricing ↗One conversation.
A whole new trajectory.