Triangle104 commited on
Commit
fe01368
·
verified ·
1 Parent(s): 0351d0b

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +75 -0
README.md CHANGED
@@ -12,6 +12,81 @@ tags:
12
  This model was converted to GGUF format from [`PygmalionAI/Pygmalion-3-12B`](https://huggingface.co/PygmalionAI/Pygmalion-3-12B) using llama.cpp via the ggml.ai's [GGUF-my-repo](https://huggingface.co/spaces/ggml-org/gguf-my-repo) space.
13
  Refer to the [original model card](https://huggingface.co/PygmalionAI/Pygmalion-3-12B) for more details on the model.
14
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
15
  ## Use with llama.cpp
16
  Install llama.cpp through brew (works on Mac and Linux)
17
 
 
12
  This model was converted to GGUF format from [`PygmalionAI/Pygmalion-3-12B`](https://huggingface.co/PygmalionAI/Pygmalion-3-12B) using llama.cpp via the ggml.ai's [GGUF-my-repo](https://huggingface.co/spaces/ggml-org/gguf-my-repo) space.
13
  Refer to the [original model card](https://huggingface.co/PygmalionAI/Pygmalion-3-12B) for more details on the model.
14
 
15
+ ---
16
+
17
+
18
+
19
+
20
+
21
+
22
+
23
+
24
+ Dataset
25
+ -
26
+
27
+
28
+
29
+ We've gathered a large collection of instructions and roleplaying totaling hundreds of millions of tokens, including our PIPPA dataset and roleplaying forums.
30
+
31
+
32
+
33
+
34
+
35
+
36
+
37
+ Limitations and biases
38
+ -
39
+
40
+
41
+
42
+ The intended use-case for this model is fictional writing for entertainment purposes. Any other sort of usage is out of scope.
43
+
44
+
45
+ As such, it was not fine-tuned to be safe and
46
+ harmless: the base model and this fine-tune have been trained on data
47
+ known to contain profanity and texts that are lewd or otherwise
48
+ offensive. It may produce socially unacceptable or undesirable text,
49
+ even if the prompt itself does not include anything explicitly
50
+ offensive. Outputs might often be factually wrong or misleading.
51
+
52
+
53
+
54
+
55
+
56
+
57
+
58
+ Training Specifications
59
+ -
60
+
61
+ We trained our model as a rank-32 LoRA adapter with one epoch over
62
+ our data using 8x NVIDIA A40 GPUs. For this run, we employed a learning
63
+ rate of 2e-4 and a total batch size across all GPUs of 24. A cosine
64
+ learning rate scheduler was used with a 100 step warmup. DeepSpeed ZeRO
65
+ was used to successfully get memory usage down.
66
+
67
+
68
+
69
+
70
+
71
+
72
+
73
+ Acknowledgements
74
+ -
75
+
76
+
77
+
78
+ This project could not have been done without the compute support of Hive Digital Technologies and the Axolotl training software.
79
+
80
+
81
+ We'd like to extensively thank lemonilia for their wonderful help in compiling roleplay forum data.
82
+
83
+
84
+ And most of all, we dedicate this model to our great community,
85
+ who've stuck with us through everything until now. Sincerely, thank you
86
+ so much. We hope you enjoy our work to the fullest and we promise more
87
+ is on the way soon.
88
+
89
+ ---
90
  ## Use with llama.cpp
91
  Install llama.cpp through brew (works on Mac and Linux)
92