Beta open · Unquantized GLM-5.3-Flash & Qwen 3.8 27B · both uncensored · founding round

The unquantized uncensored frontier.
One flat rate. No meters.

Unquant serves two frontier models at full precision - the 328B GLM-5.3-Flash flagship with 262K context, vision, and tool calling, and the famously fast Qwen 3.8 27B for roleplay and high-volume work - both uncensored at the tensor level, both genuinely unlimited on every tier. Unlimited tokens. Unlimited messages. One monthly price.

10 GLM founding seats at $99/mo-for-life (75% below the $399 forever price) and 10 Qwen founding seats at $59/mo-for-life (46% below the $109 forever price). When they're gone, they're gone.
~100 tok/sper-stream decode at full precision
262K contextfull native window, every GLM tier
0 refusalstensor-level uncensoring, not a prompt trick
Full precisionevery weight served unquantized

Genuinely unlimited. That is the business model.

Every "unlimited" tier in this industry quietly meters you with rate limits, throttles, or tiny-context traps. Ours is built the other way around: we never count your tokens, meter your speed, or queue your requests.

The promise

No rate limits, no speed caps, no fair-share queues

Your usage never triggers throttling of any kind. Capacity management happens entirely on our side, and it's priced into the flat rate. That's the whole catch. Concurrency is the only tier difference.

The only constraint is capacity, and capacity is ours to manage
Daily token capnone
Monthly token capnone
Per-request context limitfull 262,144
Concurrent streams1 to 4 by tier · never queued
Precision servedUnquantized - full precision

Unlimited, unquantized, flat-rate.

Two frontier models, one flat-rate principle: every tier is genuinely unlimited. GLM tiers differ by concurrent streams - the full 262K window on all of them. Qwen tiers ladder up the context window, from 64K to the full 262K.

GLM founding round: 10 of 10 seats left - the $399 Unlimited tier for $99/mo, locked for life. 75% below the forever price - never offered again after these 10 seats.

GLM-5.3-Flash · the flagship - unquantized · 262K context · vision · tools · ~100 tok/s

Standard

$109/mo
GLM-5.3-Flash · unquantized
  • Unlimited tokens & messages
  • Full 262K context
  • 1 concurrent stream
  • ~100 tok/s, vision, tool calls
Choose Standard
Most popular

Pro

$199/mo
GLM-5.3-Flash · unquantized
  • Unlimited tokens & messages
  • Full 262K context
  • 2 concurrent streams
  • Priority capacity at peak
Choose Pro

Unlimited

$399/mo
GLM-5.3-Flash · unquantized · top of stack
  • Unlimited tokens & messages
  • Full 262K context
  • 4 concurrent streams, top priority
  • First access to new models
Choose Unlimited
Founding · 75% off forever

GLM Founding

$99/mo · locked for life
The Unlimited tier ($399) at $99 - 75% below the forever price, and your rate never rises as long as you stay subscribed. 10 seats, then gone.
  • The Unlimited tier at $99 - forever
  • Full 262K context
  • 4 concurrent streams, top priority
  • Your rate never rises as long as you stay subscribed
Claim a founding seat

Dedicated

$5,999/mo
your own Blackwell backend · zero neighbors
  • A whole serving node - nobody else on it
  • Full 262K context, all 8 streams yours
  • No fair-share, no throttle, ever
  • Sleep/wake control of your box
Talk to us

Qwen 3.8 27B · uncensored & fast - the community's favorite uncensored model · ~100+ tok/s · roleplay-tuned · 64K → 262K window ladder

Qwen founding round: 10 of 10 seats left - the $109 Qwen Unlimited tier for $59/mo, locked for life. 46% below the forever price - never offered again after these 10 seats.
Cheapest uncensored

Roleplay

$34.99/mo
Qwen 3.8 27B · uncensored · roleplay-tuned
  • Unlimited roleplay & chat
  • 64K context window
  • 1 stream
  • Vision, tool calling & thinking included
Choose Roleplay
Sweet spot

Qwen Plus

$49/mo
Qwen 3.8 27B · uncensored
  • Unlimited usage
  • 128K context window
  • 1 stream
  • Vision, tool calling & thinking included
Choose Qwen Plus

Qwen Pro

$79/mo
Qwen 3.8 27B · uncensored
  • Unlimited usage
  • Full 262K context window
  • 2 streams
  • Vision, tool calling & thinking
Choose Qwen Pro

Qwen Unlimited

$109/mo
Qwen 3.8 27B · uncensored · top Qwen tier
  • Unlimited usage
  • Full 262K context window
  • 4 streams, top priority
  • Vision, tool calling & thinking - agents, coding & research
Choose Qwen Unlimited
Founding · 46% off forever

Qwen Founding

$59/mo · locked for life
The Qwen Unlimited tier ($109) at $59 - 46% below the forever price, locked in writing as long as you stay subscribed. 10 seats, then gone.
  • Qwen Unlimited tier at $59 - forever
  • Full 262K context window
  • 4 streams, top priority
  • Vision, tool calling & thinking
Claim a founding seat

Qwen Dedicated

$549/mo
your own RTX 5090 backend · zero neighbors
  • A whole serving node - nobody else on it
  • Full 262K context, all 8 streams yours
  • No fair-share, no throttle, ever
  • Great for studios & heavy agents
Talk to us

Full details and comparisons on the pricing page. Founding prices are grandfathered for as long as the subscription stays active; later tiers cost more.

Everyone else serves you a compressed copy.
We serve the original.

Unquantized

Quantized models are photographs of a painting - convenient, smaller, and subtly wrong. Unquant serves the original: every weight at its trained precision, none of the compression loss.

Full-precision weightsZero quantization loss

Uncensored, in the weights

Refusal behavior is removed at the tensor level, 0% refusals on the A/B suite. No jailbreak prompts, no filter dodging, no sudden refusals mid-story.

Tensor-level uncensoring0% refusal suite

Fast, at full precision

~100 tokens per second per stream, real-time-shape streaming with prefix caching. Full 262,144-token requests on every tier, uncapped.

~100 tok/s262K window

Claim your position

We onboard in application order, founding seats first. A sentence about what you're building helps us place you in the right wave.

GLM founding seats10/10 left
Qwen founding seats10/10 left
Onboardingin order, days not weeks

Prefer email? luke.moran103@gmail.com

No spam. One email when your slot is ready.

Questions, answered.

What exactly is the model?

Two openly-licensed frontier models, both uncensored at the tensor level so refusals are physically removed rather than suppressed, and both served at full training precision (no quantization): GLM-5.3-Flash, a 328B-parameter sparse flagship with 262K native context, vision, reasoning, and tool calling, and Qwen 3.8 27B, the community's favorite uncensored model, blazingly fast, with vision, tool calling, and thinking intact. Every tier is genuinely unlimited on either model.

Is it genuinely unlimited? What's the catch?

No token meters, no rate limits, no speed caps, no fair-share algorithms. Your usage never triggers throttling of any kind, capacity management happens entirely on our side, and it's priced into the flat rate. That's the whole catch.

What does "unquantized" actually mean for me?

Most services hand you a compressed copy that saves them money and costs you quality. Unquant serves the original weights at full precision so you get the model exactly as trained, including on very long documents where quantization error compounds.

How fast is it, really?

Streamed decode runs at roughly 100 tokens per second per stream, fast enough to read along in real time, at full precision. Prompt prefill is heavily optimized with prefix caching, so repeat sessions start responding near-instantly.

Do you log my prompts?

We store your email, plan selection, and application notes on signup. We do not retain your prompts or completions beyond short-lived operational logs needed for abuse prevention and capacity planning, and we never sell data. Deletion requests are honored.

When do I get access?

Beta onboarding is processed in application order as capacity comes online. Founding members are onboarded first, in groups, beginning with the first wave within days of acceptance.

Which clients does it work with?

Anything that speaks the OpenAI API: point your base URL at https://api.unquant.io/v1 with your key, and it works, SillyTavern, Hermes, OpenWebUI, LangChain, raw curl, anything.

What does "for life" founding pricing mean?

GLM Founding members pay $99/mo for the Unlimited tier (normally $399), a 75% discount locked for as long as their subscription stays active, even as standard prices rise. Qwen Founding members pay $59/mo for Qwen Unlimited (normally $109), 46% locked for life. Cancel and rejoin later, and you rejoin at the then-current price.

Can I really run 262K-token prompts?

Yes, the full native context is served on every tier, unlimited tier included. No hidden per-request cap below 262,144 tokens.

Refunds?

Monthly plans are cancellable anytime. If the service doesn't work for you in your first week, we'll refund it.