LLMPrivate

GLM 4.7 Flash Heretic

A community-abliterated, open-weights variant of GLM-4.7-Flash built for fast inference, reasoning, and tool use with relaxed refusal behavior.

Get API key
Provider
Olafangensan (community mod; Z.AI base)
Price
$0.07 in · $0.40 out / 1M
Context window
200K tokens
Released
February 1, 2026
License
MIT

What is GLM 4.7 Flash Heretic?

GLM 4.7 Flash Heretic is a community-modified, open-weights text model derived from Z.AI's GLM-4.7-Flash. It uses a mixture-of-experts architecture and is optimized for fast reasoning, function calling, and web search. The 'Heretic' variant applies abliteration to reduce refusals while preserving core capabilities, and it runs privately on Venice with zero prompt retention.

Use it privately on Venice

On Venice, GLM 4.7 Flash Heretic runs under a private, zero-retention tier — your prompts are not stored or used for training. It is an open-weights, FP8-quantized model that supports tool use, reasoning, and structured JSON output, making it a permissionless choice for agentic workflows without Big-Tech surveillance. You pay only for tokens consumed, with no subscription lock-in.

Private (zero retention)
No prompt training
TEE · hardware enclave
End-to-end encrypted

What can it do?

Strengths
  • Extremely low API cost ($0.07 in / $0.40 out per 1M tokens) for a capable reasoning and tool-use model.
  • Open-weights MIT license enables self-hosting, fine-tuning, and full stack sovereignty outside closed APIs.
  • Native support for tool use / function calling, reasoning, web search, and structured JSON output for agentic workflows.
  • 200K context window and 24K max output handle long documents and extended generations.
  • Community abliteration strips automated refusal behavior, improving utility for creative and sensitive prompts compared to the base model.
Limitations
  • Community-modified rather than officially released by Z.AI, so updates, safety patches, and support depend on the contributor.
  • Exact parameter counts and active-expert ratios vary across third-party reports, complicating precise hardware planning for self-hosting.
  • FP8 quantization on Venice trades a small amount of precision for inference speed versus full-precision weights.
  • Text-only — no vision, audio, or multimodal input support.
  • Abliteration can occasionally weaken instruction-following or remove useful guardrails on edge-case prompts.

GLM 4.7 Flash Heretic capabilities

How to use it via API

Venice exposes an OpenAI-compatible API. Swap your base URL and call olafangensan-glm-4.7-flash-heretic.

curl https://api.venice.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "olafangensan-glm-4.7-flash-heretic",
    "messages": [{ "role": "user", "content": "Explain quantum tunneling simply." }]
  }'

Specifications

MakerOlafangensan (community modification of Z.AI GLM-4.7-Flash)
ReleasedFebruary 2026
ArchitectureMixture-of-Experts (MoE)
Parameters~30B total / ~3B active per token
ModalityText input, text output
LicenseMIT
Open weightsYes — downloadable from Hugging Face
Context window200K tokens
Max output24K tokens
CapabilitiesFunction calling, Reasoning, Web search
Privacy on VenicePrivate — zero retention
Available on Venice sinceFeb 2026

Pricing

Billed per token on Venice: $0.07 per 1M input tokens and $0.40 per 1M output tokens.

Input / 1M tokens
$0.07
Output / 1M tokens
$0.40
Cached input / 1M
$0.04

New Venice accounts include a free daily allowance and 500 welcome credits — no credit card required.

GLM 4.7 Flash Heretic vs alternatives

ModelContext windowOpen weightsPrice (Venice)Best for
GLM 4.7 Flash Heretic200K tokensYes$0.07 in · $0.40 out / 1MFast agentic coding & reasoning
DeepSeek V3.2160K tokensYes$0.33 in · $0.48 out / 1MGeneral reasoning & coding
GLM 5.1200K tokensYes$1.10 in · $4.15 out / 1MLong-context GLM flagship
Google Gemma 4 31B Instruct256K tokensYes$0.12 in · $0.36 out / 1MLightweight open instruct

Ultra-cheap open-weights MoE with tool use, reasoning, and web search.

What is it good for?

  • High-volume agentic coding and tool-calling workflows where per-token cost dominates the budget.
  • Long-context document analysis, summarization, and extraction within the 200K window.
  • Private, open-weights inference for teams that want sovereign control over model weights and data.
  • Rapid prototyping with structured output and reasoning for form filling, validation, and multi-step logic.
  • Creative writing and roleplay that benefit from lower refusal rates and more permissive generation.

Prompting tips

  • Provide explicit JSON schemas when requesting structured output to maximize parseable accuracy.
  • Enable reasoning mode for debugging or multi-step logic, but expect slightly higher latency.
  • For tool use, define function signatures clearly — the model handles parallel function calling well.
  • Long-context prompts benefit from context caching; cached input is billed at $0.04 per 1M tokens.
  • If the model over-thinks on simple queries, set a lower max output length or disable reasoning.

Version history

GLM-4.7-Flash
2026-01

Z.AI base model.

GLM 4.7 Flash Heretic
2026-02

CurrentCommunity abliterated variant released on Hugging Face.

Frequently asked questions

GLM 4.7 Flash Heretic is a community-modified text model derived from Z.AI's GLM-4.7-Flash. It is an open-weights, mixture-of-experts model optimized for fast reasoning, function calling, and web search, with abliteration applied to reduce refusal behavior.

Venice charges $0.07 per 1M input tokens and $0.40 per 1M output tokens. Cached input is $0.04 per 1M tokens. There is no subscription required; you pay only for what you use.

It is not free, but it is extremely inexpensive. At $0.07/$0.40 per 1M tokens, it is one of the cheapest capable reasoning models on Venice. You only need Venice credits to run it.

Yes. The weights are released under an MIT license on Hugging Face and can be downloaded for self-hosting. On Venice it is served as an open-weights, FP8-quantized endpoint.

Yes. Venice's endpoint supports tool use and function calling, reasoning, web search, and structured JSON output. These capabilities make it suitable for agentic workflows.

GLM 4.7 Flash Heretic is far cheaper and faster for high-volume coding and tool use. GLM 5.1 is the newer flagship with stronger overall performance but costs significantly more. Choose the Heretic variant for cost-efficient agents and GLM 5.1 for maximum quality.

It runs under Venice's private, zero-retention tier. Your prompts are not stored, profiled, or used for training. Venice does not build a conversation history from your requests.

'Heretic' refers to the Heretic abliteration method used by the community contributor Olafangensan. It is an open-source technique designed to strip automated refusal behavior while aiming to preserve the base model's reasoning and coding capabilities.

Related models

Run GLM 4.7 Flash Heretic privately.

No prompt logging. No data used for training. Free to start — no credit card.

Room