JThomas-CoE/coe-gemma4-math-mmlu_pro-14b-a4b-q4

https://huggingface.co/JThomas-CoE/coe-gemma4-math-mmlu_pro-14b-a4b-q4
Activeby JThomas-CoE3061updated 4 months ago

Base model: google/gemma-4-26b-it Architecture: MoE — 26B total / ≈4B active parameters (1 shared expert + 8 routed from a pool of 128 per MoE layer, 30 MoE layers) Method: Activation-directed expert surgery — 128 → 64 experts per layer (50% reduction) Quantization: Q4KM (≈9.7 GB on disk) Tags:…

Sourced from

  • HuggingFace — JThomas-CoE/coe-gemma4-math-mmlu_pro-14b-a4b-q4

Related resources

Base model: google/gemma-4-26b-it Architecture: MoE — 26B total / ≈4B active parameters (1 shared expert + 8 routed from a pool of 128 per MoE layer, 30 MoE layers) Method: Activation-directed expert surgery — 128 → 64 experts per layer (50% reduction) Quantization: Q4KM (≈9.7 GB on disk) Tags:…

Active1904 months ago

Base model: google/gemma-4-26b-it Architecture: MoE — 26B total / ≈4B active parameters (1 shared expert + 8 routed from a pool of 128 per MoE layer, 30 MoE layers) Method: Activation-directed expert surgery — 128 → 64 experts per layer (50% reduction) Quantization: Q4KM (≈9.7 GB on disk) Tags:…

Active674 months ago

Gravity-bio-16B-A3B is a biology-focus midtrained model derived from Gravity-16B-A3B-Base. It uses the same sparse Mixture-of-Experts (MoE) architecture and tokenizer as Gravity-16B-A3B-Base, with additional midtraining for biological understanding on TheBioCollection corpus.

Active2353 months ago
Python

Using llama.cpp release b2440 for quantization.

Stale6152 years ago
Python

Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants.

Idle5.7K1 year ago
Python