A prompt with structure and depth.
- Text features are computed once and reused during sampling.
- Optional Magic Prompt expansion is upstream of the denoiser.
structured caption → text-only encoder → 13 feature tapsAn image model for typography, composition and visual expression. Explore the text encoder, joint transformer stream and flow-matching path behind its generation process.
Structured language and image latents meet inside a flow-matching transformer.
Structured captions describe objects, literal text and composition. A frozen Qwen3-VL-8B-Instruct encoder runs in text-only mode; thirteen hidden-state taps are combined for conditioning.
structured caption → text-only encoder → 13 feature tapsThe same latent canvas is updated repeatedly. Text features condition the joint stream; the VAE decodes the final image. The animation is explanatory; no model inference is executed.
The 34 blocks belong to the joint DiT. The text encoder, guidance branch and VAE are separate components.
Configuration, memory and computation—each with a specific role in the model.
structured caption → text-only encoder → 13 feature tapstext + image tokens → attention → SwiGLU → next blockzₙ₊₁ = zₙ + v_guided × (tₙ₊₁ − tₙ)image H×W → latent H/8×W/8 → tokens H/16×W/16Developer papers, released configurations and reference implementations. Reviewed 10 September 2026.
Text encoder, denoiser and VAE decoder
Configuration, attention and transformer blocks
Turbo 12, Default 20 and Quality 48 presets
Text features, guidance branches and Euler update
The released architecture is documented in code and pipeline guides; no undisclosed training-corpus claims are made here.
Open weights retain their model-specific licence. The code licence is not a substitute for weight-use rights.
Explore art direction, compositions and campaign assets around your product.
Generate concepts that put words and visual composition in the same frame.
Bring repeatable settings and predictable workflow integration to the creative process.
Perchy provides managed AI inference through authenticated APIs. Customers do not receive SSH, root, operating-system or bare-metal access. The service does not provide model-training or fine-tuning functionality. Server administration remains with Perchy.
Managed API availability requires appropriate model and commercial-use rights. Licence costs are included when assessing service economics. Outputs require review before publication. Image generation uses a latent denoising workflow; it is not an autoregressive chat KV-cache workload.
Read the model developer’s documentationOne conversation.
A whole new trajectory.