JayhC
/

L3_SnowStorm_4x8B-6.5bpw-h8-exl2

Text Generation

Mixture of Experts

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

Configuration Parsing Warning: In config.json: "quantization_config.bits" must be an integer

6.5bpw/h8 exl2 quantization of xxx777xxxASD/L3_SnowStorm_4x8B using exllamav2 0.0.21 and default calibration dataset.

ORIGINAL CARD:

(Maybe i'll change the waifu picture later)

GGUF quants

Experimental RP-oriented MoE, the idea was to get a model that would be equal to or better than Mixtral 8x7B and it's finetunes in RP/ERP tasks.

Llama 3 SnowStorm 4x8B

base_model: NeverSleep_Llama-3-Lumimaid-8B-v0.1-OAS
gate_mode: random
dtype: bfloat16
experts_per_token: 2
experts:
  - source_model: ChaoticNeutrals_Poppy_Porpoise-v0.7-L3-8B
  - source_model: NeverSleep_Llama-3-Lumimaid-8B-v0.1-OAS
  - source_model: openlynn_Llama-3-Soliloquy-8B-v2
  - source_model: Sao10K_L3-8B-Stheno-v3.1

Models used

Difference(from ChaoticSoliloquy v1.5)

Update from NeverSleep/Llama-3-Lumimaid-8B-v0.1 to NeverSleep/Llama-3-Lumimaid-8B-v0.1-OAS
Update from openlynn/Llama-3-Soliloquy-8B-v1 to openlynn/Llama-3-Soliloquy-8B-v2
Update from Sao10K/L3-Solana-8B-v1 to Sao10K/L3-8B-Stheno-v3.1

Vision

Prompt format: Llama 3

Downloads last month: 6

Inference Examples

Text Generation

This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social visibility and check back later, or deploy to Inference Endpoints (dedicated) instead.