Models/Qwen3.8-Flash-Next
LANGUAGE / QWEN3.8-FLASH-NEXT

Fast thinking.
Your direction.

A hybrid language and vision model that combines Gated DeltaNet, Qwen Sparse Attention and routed experts. Explore how the architecture balances compact recurrent state with selective access to a long context.

QWEN3.8-FLASH-NEXT
BACKBONE125B parameters
ACTIVE / TOKEN6B parameters
ATTENTIONGDN + QSA
NATIVE CONTEXT262,144 tokens
INSIDE THE MODEL / 3D ARCHITECTURE ATLAS

A memory.
And a way to recall.

Recurrent state carries continuity. Sparse attention brings selected history back into focus.

PC / QWEN / SYSTEMINTERACTIVE SCHEMATIC
Drag to rotate · Arrow keys supported
MODULE 01 / RECURRENT MEMORY
36
GATED DELTANET LAYERS

Keep state. Write the difference.

Gated DeltaNet updates a bounded recurrent matrix rather than appending every token to a KV list. It decays the previous state, estimates an existing association and writes a gated correction.

previous state → decay → delta write → query read
READING THE DIAGRAM

Each slab is a complete decoder block, including its own MoE FFN. The 3× GDN → 1× QSA pattern repeats 12 times. The side drawing is an FFN detail inset. The animation is explanatory; no model inference is executed.

BACKBONE MAP

48 layers. A deliberate pattern.

36 GDN12 QSA

Attention schedule shown. Every attention mixer is followed by an MoE; four residual branches run through the backbone.

A CLOSER LOOK / QWEN

The details behind
the diagram.

Configuration, memory and computation—each with a specific role in the model.

01 / RECURRENT MEMORY

Keep state. Write the difference.

  • 48 value heads; 128 × 128 recurrent matrices per head.
  • Short convolution state is retained alongside the matrix.
previous state → decay → delta write → query read
02 / SPARSE RECALL

Select a block. Read its tokens.

  • 12 QSA layers; each head has 256 dimensions.
  • Selection reduces the attention read set, not the retained history.
index keys → pool 4 → top 512 blocks → sparse GQA
03 / CONDITIONAL COMPUTE

A small selection. A large capacity.

  • Channelwise read gates blend the four branches.
  • Scalar write gates return the sublayer result; no Sinkhorn residual mixer.
4 branches → gated read → top-10 + shared → gated write
04 / N-GRAM MEMORY

Some capacity starts with a lookup.

  • The 125B backbone and 4B MTP module are separate allocations.
  • This learned table is not a database of customer conversations.
token IDs → 2/3-gram hashes → table → context gate
FOLLOW THE RESEARCH

Go straight
to the source.

Developer papers, released configurations and reference implementations. Reviewed 10 September 2026.

WHAT THIS GUIDE COVERS

125B describes the language backbone; additional n-gram and MTP allocations are listed separately.

262,144 tokens is the native context. Extended-context serving needs a validated configuration.

MADE MEANINGFUL IN YOUR BUSINESS

From capability
to possibility.

01

Customer conversations

Bring product knowledge and conversational assistance closer to your customers.

02

Agentic workflows

Connect reasoning to approved tools, with application-level control and human oversight.

03

Knowledge in context

Connect retrieval to managed inference with agreed API access and data processing terms.

Know your inference service scope.

Perchy provides managed AI inference through authenticated APIs. Customers do not receive SSH, root, operating-system or bare-metal access. The service does not provide model-training or fine-tuning functionality. Server administration remains with Perchy.

API availability, model licensing, context length and serving configuration are confirmed for your workload. Parameter counts and context specifications describe the developer’s model, not a promise of Perchy service capacity. Extended context requires a separately validated configuration.

Read the model developer’s documentation
KEEP EXPLORING

Another model.
Another way to compute.

All models
LANGUAGE

DeepSeek-V4.1-Flash

Explore architecture
LANGUAGE

GLM-5.3-Flash

Explore architecture
IMAGE

Ideogram 4.0

Explore architecture
VIDEO

MiniMax H3

Explore architecture
THE NEXT MOVE IS YOURS

Find your
unfair advantage.

Let’s build your direction

One conversation.
A whole new trajectory.