Live now · beta round

GLM-5.3-Flash Uncensored

glm-5.3-flash-unquant · served unquantized

The 328B-parameter sparse Mixture-of-Experts flagship: frontier reasoning at a fraction of dense-model serving cost, which is what makes genuinely unlimited flat-rate tiers possible. Uncensored at the tensor level, so refusals are physically removed rather than suppressed, and served at full training precision with nothing quantized away.

Unquantized262K contextVisionTool callingReasoning~100 tok/s0% refusals

At a glance

Parameters
328B total · sparse activation
Context window
262,144 tokens native
Modality
Text · vision · tool calling
Decode speed
~100 tok/s per stream
Precision
Unquantized, full training precision
Uncensoring
Tensor-level · 0% refusals on the A/B suite
License
Openly licensed (MIT family)
Endpoint
api.unquant.io/v1 · OpenAI-compatible

Use it in one call

# OpenAI-compatible: point your SDK at us
from openai import OpenAI

client = OpenAI(
    base_url="https://api.unquant.io/v1",
    api_key="UQ-...",
)

resp = client.chat.completions.create(
    model="glm-5.3-flash-unquant",
    messages=[{"role": "user",
               "content": "Hello!"}],
)

Works with SillyTavern, Hermes, OpenWebUI, LangChain, raw curl, anything that speaks the OpenAI API.

Unlimited tiers on this model

GLM STANDARD
$109/mo

Full 262K context · 1 stream · unlimited tokens

GLM PRO · MOST POPULAR
$199/mo

2 streams · priority capacity

GLM UNLIMITED
$399/mo

4 streams · top priority · first access

Founding seats: this tier at $99/mo for life (75% below the forever price, 10 seats). All GLM tiers and the Dedicated whole-box option on the pricing page.