README.md · anthracite-org/magnum-v2-4b-exl2 at 6b8be2fb6d19e8a0828bac1988399cde4da2e696

metadata

license: apache-2.0
language:
  - en
pipeline_tag: text-generation
tags:
  - chat

This repo contains EXL2 quants of the model. If you need the original weights, please find them here.

Base repo only contains the measurement file, see revisions for your quant of choice.

This is the eighth in a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet and Opus. This model is fine-tuned on top of IntervitensInc/Llama-3.1-Minitron-4B-Width-Base-chatml.

Prompting

Model has been Instruct tuned with the ChatML formatting. A typical input would look like this:

"""<|im_start|>system
system prompt<|im_end|>
<|im_start|>user
Hi there!<|im_end|>
<|im_start|>assistant
Nice to meet you!<|im_end|>
<|im_start|>user
Can I ask a question?<|im_end|>
<|im_start|>assistant
"""

Support

To run inference on this model, you'll need to use Aphrodite or vLLM or EXL2/tabbyAPI, as llama.cpp hasn't yet merged the required pull request to fix the llama3.1 rope_freqs issue with custom head dimensions.

However, you can work around this by quantizing the model yourself to create a functional GGUF file. Note that until this PR is merged, the context will be limited to 8k tokens.

To create a working GGUF file, make the following adjustments:

Remove the "rope_scaling": {} entry from config.json
Change "max_position_embeddings" to 8192 in config.json

These modifications should allow you to use the model with llama.cpp, albeit with the mentioned context limitation.

Credits

This model has been a team effort, and the credits goes to all members of Anthracite.

Training

The training was done for 2 epochs. We used 2 x RTX 6000s GPUs graciously provided by Kubernetes_Bad for the full-parameter fine-tuning of the model.

Safety

...