Your traffic. Your deployment.
Dedicated capacity, agreed data processing and retention, and an inference stack optimised for your workload alone.
How we protect your dataPrivate, fast and affordable inference.
Tuned to the traffic you actually serve.
A custom KV backend and our optimised DSpark speculative decoding turn the patterns in your workload into more tokens per GPU, and fewer dollars per token.
Public APIs and stock serving stacks are tuned for the average request. But your prompts repeat, your contexts overlap and your outputs follow habits. We tune the whole inference path for your traffic, and measure the difference on your own workload.
Dedicated capacity, agreed data processing and retention, and an inference stack optimised for your workload alone.
How we protect your dataDSpark drafts several tokens at once and the target model verifies them in a single pass. A KV backend that remembers your context cuts the wait for the first token.
Inside the technologyWhen every GPU produces more tokens per second, every token costs less. Same hardware, same latency target, a smaller bill.
Your token economicsStandard decoding produces one token per forward pass. DSpark drafts a block and the target model verifies it at once. Ours is optimised for your workload, so more of every draft is accepted, with no change to the output.
Standard decoding: about 1 tokens per forward pass. Stock DSpark: about 2.4 tokens per forward pass. Our DSpark: about 6.2 tokens per forward pass.
At high concurrency, GPUs are already busy and extra drafts cost more than they save. We schedule the verify budget against your concurrency and latency target, so the speed-up survives production traffic.
| Concurrent requests | Our DSpark | DFlash | EAGLE-3 |
|---|---|---|---|
| 1 | 5.9 | 5.1 | 1.6 |
| 4 | 5.6 | 4.5 | — |
| 8 | 5.3 | 4.5 | — |
| 16 | 4.9 | 3.9 | — |
| 32 | 4.3 | 2.8 | 1 |
Agents, assistants and document pipelines resend the same prompts, tools and histories every turn. Our KV backend keeps that context across GPU memory, host memory and local NVMe, including for hybrid-attention models where stock caches miss.
6K-token system and tool prefix · about 5K tokens added per turn · 12K tokens/s prefill · 400K tokens/s KV load from the cache tiers.
Average time to first token across 32 concurrent coding-agent sessions with 100K-token contexts.
The same models on the same GPUs, at the same latency target. Only the serving stack changes.
Creative writing defeats generic speculative decoding and prefix caching: every reply is new prose, and readers pause long enough between turns for ordinary caches to forget the story. It is also where optimising for the workload pays off most.
Long stories and character dialogue at high temperature. Every reply is new prose, and readers pause long enough between turns for ordinary caches to forget the story.
See what higher throughput does to your cost per token and to the number of GPUs you need.
Cost per token is GPU time divided by the tokens it produces. Raise throughput on your workload and every token gets cheaper, on the same hardware.
Throughput uplift ×6.8 measured on a creative writing workload. Standard serving: SGLang without speculative decoding, 8K in / 1K out, 64 concurrent requests (SemiAnalysis InferenceX). GPU counts assume 60% average utilisation. List prices as published on 24 September 2026. Public APIs for open-weight models can list lower prices; they are shared services, not tuned to your workload.
Connect through an API. Perchy runs the tuned stack on capacity reserved for your workload, and keeps server administration.
Send authenticated requests and receive model outputs. Perchy runs the models and the tuned inference stack, including server administration and maintenance.
Perchy provides managed AI inference through authenticated APIs. Customers do not receive SSH, root, operating-system or bare-metal access. The service does not provide model-training or fine-tuning functionality. Server administration remains with Perchy.
Recurrent state meets sparse attention. A different way to carry context through an intelligent workflow.
A causal encoder–decoder, compressed sparse attention and conditional memory. Explore the architecture behind the next token.
Kimi Delta Attention and KPool-DSA share the work: recurrent state for continuity, selective attention for recall.
Language and image tokens meet inside a single-stream diffusion transformer, turning latent noise into deliberate composition.
A joint audio–video latent sequence. One single-stream transformer. A generation path designed to keep sound and motion together.
Model examples illustrate managed inference options. Availability, licences and performance are confirmed in your proposal. Model names belong to their respective owners; listing does not imply endorsement or partnership.
One benchmark.
A whole new cost curve.