+ 03 / THE ECONOMICS

More intelligence.
Less overhead.

Your inference service should open possibilities. Compare usage-based API spending with dedicated inference capacity operated by Perchy.

THE 1/30 SCENARIO

$30,000 in monthly API spend. $1,000 in estimated monthly service costs. That is 1/30 of the spend—and an illustration of what to investigate, not a promised result.

MAKE THE NUMBERS YOURS

Your workload.
A different equation.

Compare usage-based API spending with an estimated monthly fee for equivalent managed inference and reserved capacity.

$
$
$
YOUR COST SCENARIOUSD / MONTH
1/30*

of your current API spend

Current API spend$30,000
Managed inference scenario$1,000
Potential annual savings$348,000
96.7%

* Illustrative scenario, not a quote or measured benchmark. Adjust all inputs to your workload. No savings or equivalent performance are guaranteed. Perchy operates the servers; reserved capacity is delivered through authenticated inference APIs.

Get a workload-specific assessment
MINIMAX H3 / INTERNAL SERVING-COST MODEL

What would 1/30 require?

A measured 5.17-second video took 93.19 seconds of generation. At the provider’s reference output-only rate of $0.08/second for 768p, the output would cost about $0.4133.

Reaching 1/30 of that requires an effective generation-compute cost of about $0.532 per GPU-hour or less. Explore the assumptions below.

Provider pricing reference
Modelled compute cost / clip$0.0129API / compute ratio31.9×

This models internal inference-serving costs, not a customer GPU rental offer or a Perchy API service price. Excludes operations, idle/cold-start overhead beyond the selected utilisation, storage, transfer, licences, failed jobs and margin. Local INT8 + Turbo8 quality is not assumed equivalent to the full hosted pipeline.

THE MATH, IN PLAIN SIGHT

Monthly service estimate = reserved inference capacity + service operations and licences. Annual difference = (current API spend − service estimate) × 12. The ratio is rounded for display. All amounts are in USD, exclude taxes and are user-controlled estimates. Capacity is delivered through Perchy's managed inference APIs; these inputs are not GPU rental prices or server access charges.

COMPARE LIKE WITH LIKE

The details
make the difference.

A useful cost comparison accounts for the entire workload—not just a GPU-hour.

01

Match the output

Use the same model version, quality target, precision, video resolution and duration. A base model and an enhanced hosted pipeline are different products.

02

Count the full cost

Account for reserved inference capacity, service operations, storage, transfer, model licences and support in the agreed API service fee.

03

Measure actual demand

Utilisation matters. Batch throughput and interactive latency describe different workloads; neither should be substituted for the other.

A note on MiniMax H3.

Perchy-operated H3-Base inference and the official hosted 2K pipeline have different component boundaries. A fair comparison must specify which generation, preprocessing, reference-input and regeneration costs are included. No measured 1/30 H3 saving is asserted on this website.

See the model provider’s current pricing ↗
THE NEXT MOVE IS YOURS

Find your
unfair advantage.

Let’s build your direction

One conversation.
A whole new trajectory.