Qwen3.6-35B-A3B-4bit-Native This repository provides the 4-bit (NF4) quantized weights for the Qwen3.6-35B-A3B Mixture-of-Experts model. These weights were generated using the bitsandbytes library with double quantization enabled to ensure maximum precision at a reduced memory footprint.

Model Details

Base Model: Qwen3.6-35B-A3B Quantization: 4-bit NormalFloat (NF4) Framework: Hugging Face Transformers Total Parameters: ~35B Expert Architecture: 256 Experts per Layer

Key Features

Native Compatibility: Designed to work seamlessly with the transformers library without additional conversion layers. Memory Efficiency: Optimized to fit within ~20GB of memory (VRAM/RAM combined), making it accessible for mid-range hardware environments. Precision: Uses Double Quantization to minimize perplexity degradation compared to the original BF16 weights.

Downloads last month
27
Safetensors
Model size
35B params
Tensor type
F32
BF16
U8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for bombman/Qwen3.6-35B-A3B-4bit-Native

Quantized
(757)
this model