LLMPrivate

Gemma 4 26B A4B Uncensored

An uncensored, privacy-first Mixture-of-Experts (MoE) model based on Google's Gemma 4, optimized for unbiased reasoning, coding, and search.

Get API key
Provider
Google DeepMind (Base) / Phala (Fine-tune)
Price
$0.19 in · $0.88 out / 1M
Context window
64K tokens
Released
May 23, 2026
License
Apache 2.0 (Base model)

What is Gemma 4 26B A4B Uncensored?

Gemma 4 26B A4B Uncensored is a community-modified, de-aligned version of Google's 26B Mixture-of-Experts model. Fine-tuned using the 'Heretic' ablation method, it bypasses standard safety filters to deliver uncensored, highly objective reasoning, coding, and multilingual text generation without corporate refusals.

Use it privately on Venice

On Venice, this model runs with absolute privacy inside a hardware-isolated Trusted Execution Environment (TEE) with end-to-end encryption. Because Venice enforces zero retention, your prompts are never stored, logged, or used to train future models, ensuring complete sovereignty over your uncensored workflows.

Private (zero retention)
No prompt training
TEE · hardware enclave
End-to-end encrypted

What can it do?

Strengths
  • Uncensored 'Heretic' fine-tune drastically reduces refusals for unbiased, raw creative and analytical tasks.
  • Highly efficient Mixture-of-Experts (MoE) architecture utilizing 8 active experts out of 128 total.
  • Hardware-enforced privacy running in TDX-attested TEE enclaves with end-to-end encryption.
  • Integrated Web Search capability for real-time, grounded information retrieval.
Limitations
  • May generate highly controversial, graphic, or unsafe content due to the removal of alignment guardrails.
  • Context window on Venice is limited to 64K tokens, compared to the base model's native 256K.
  • Lacks native image or multimodal input support on this specific Venice text pipeline.

Gemma 4 26B A4B Uncensored capabilities

How to use it via API

Venice exposes an OpenAI-compatible API. Swap your base URL and call e2ee-gemma-4-26b-a4b-uncensored-p.

curl https://api.venice.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "e2ee-gemma-4-26b-a4b-uncensored-p",
    "messages": [{ "role": "user", "content": "Explain quantum tunneling simply." }]
  }'

Specifications

MakerGoogle DeepMind (Base) / Phala (Fine-tune)
ReleasedMay 23, 2026
ArchitectureSparse Mixture-of-Experts (MoE)
Parameters25.2B total (3.8B active)
Open weightsNo (Hosted pipeline)
Context window64K tokens
Max output4.096K tokens
CapabilitiesWeb search
Privacy on VenicePrivate — zero retention
Available on Venice sinceMay 2026

Pricing

Billed per token on Venice: $0.19 per 1M input tokens and $0.88 per 1M output tokens.

Input / 1M tokens
$0.19
Output / 1M tokens
$0.88

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

Gemma 4 26B A4B Uncensored vs alternatives

ModelContext windowStrongest atOpen weightsPrice (Venice)
Gemma 4 26B A4B Uncensored64K tokensUncensored reasoning & TEE privacyNo$0.19 in · $0.88 out / 1M
Google Gemma 4 31B Instruct256K tokensAligned reasoning & codingYes$0.12 in · $0.36 out / 1M
DeepSeek V3.2160K tokensCoding & mathYes$0.33 in · $0.48 out / 1M
GLM 5.1200K tokensMultilingual reasoningYes$1.10 in · $4.15 out / 1M

The uncensored, MoE-powered privacy champion.

What is it good for?

  • Uncensored creative writing, roleplay, and brainstorming free from corporate safety filters.
  • Objective analysis of controversial historical, political, or philosophical topics.
  • Coding assistance and debugging where standard models refuse due to security-related keywords.
  • Real-time research and fact-checking using integrated web search.

Prompting tips

  • Use a system prompt to define the exact persona; this model natively supports system prompts across 35+ languages.
  • Be direct and explicit—no need to use euphemisms or bypass phrasing since the model is uncensored.
  • Enable web search when asking about events after its April/May 2026 cutoff.

Version history

Gemma 4 26B A4B
2026-04-03

Official base MoE model released by Google DeepMind.

Gemma 4 26B A4B Uncensored (Heretic)
2026-05-23

CurrentCommunity de-aligned variant fine-tuned by Phala.

Frequently asked questions

Gemma 4 26B A4B Uncensored is a community-modified, de-aligned version of Google's 26B Mixture-of-Experts (MoE) model. Fine-tuned using the 'Heretic' ablation method, it bypasses standard safety filters to deliver uncensored, highly objective reasoning, coding, and multilingual text generation without corporate refusals.

On Venice, the model is billed per token at $0.19 per 1M input tokens and $0.88 per 1M output tokens, making it highly cost-effective.

The base model is open-weights (Apache 2.0), but this specific hosted pipeline on Venice is served via a third-party provider. You can try it on Venice using free daily credits or a premium subscription.

It means the model has undergone Arbitrary-Rank Ablation (ARA) to remove safety alignment guardrails. It will answer sensitive, controversial, or complex prompts that standard Google models refuse.

The 31B Instruct is Google's official, aligned dense model with a 256K context. The 26B Uncensored is a community MoE variant with a 64K context on Venice that does not refuse prompts.

Yes, this model supports integrated web search on Venice, allowing it to fetch real-time information to ground its answers.

It runs in a hardware-isolated Trusted Execution Environment (TEE) with end-to-end encryption. Venice maintains a zero-retention policy, meaning your prompts are never stored or used for training.

Run Gemma 4 26B A4B Uncensored privately.

No prompt logging. No data used for training. Free to start — no credit card.

Room