Cherrytest commited on
Commit
f0cf451
·
1 Parent(s): e00d68b

update file

Browse files
.gitattributes CHANGED
@@ -33,7 +33,3 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
- *.md filter=lfs diff=lfs merge=lfs -text
37
- *.png filter=lfs diff=lfs merge=lfs -text
38
- *.json filter=lfs diff=lfs merge=lfs -text
39
- *.py filter=lfs diff=lfs merge=lfs -text
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
LOGO.png CHANGED

Git LFS Details

  • SHA256: f11ce0268d9a5292f16388a70f5ea437e3e8a0b00e1bb73e64a970d24cc8340c
  • Pointer size: 130 Bytes
  • Size of remote file: 21.8 kB
MODEL_LICENSE.md CHANGED
@@ -1,3 +1,47 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:58043a7f560e11559402d57e7fcb49633b55b76f056b3fc24bbf24c3a8b15f96
3
- size 6998
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # CodeFuse COMMUNITY LICENSE AGREEMENT
2
+ CodeFuse Release Date: September 8, 2023
3
+
4
+ By clicking to agree or by using or distributing any portion or element of the Materials, you will be deemed to have recognized and accepted the content of this Agreement, which is effective immediately.
5
+
6
+ 1. Definitions.
7
+ a. This CodeFuse COMMUNITY LICENSE AGREEMENT (this "Agreement") shall mean the terms and conditions for use, reproduction, distribution and modification of the Materials as defined by this Agreement.
8
+ b. "Ant" or "We" (or "Us") shall mean Ant Group.
9
+ c. "CodeFuse" shall mean the large language models (including CodeFuse-13B and CodeFuse-CodeLlaMa-34B), and software and algorithms, consisting of trained model weights, parameters (including optimizer states), machine-learning model code, and other elements of the foregoing distributed by Us.
10
+ d. "Documentation" shall mean the specifications, manuals and documentation accompanying CodeFuse distributed by Us.
11
+ e. "Materials" shall mean, collectively, Ant's proprietary CodeFuse and Documentation (and any portion thereof) made available under this Agreement.
12
+ f. "Object" form shall mean any form resulting from mechanical transformation or translation of a Source form, including but not limited to compiled object code, generated documentation, and conversions to other media types.
13
+ g. "Source" form shall mean the preferred form for making modifications, including but not limited to model source code, documentation source, and configuration files.
14
+ h. "Third Parties" (or "Third Party") shall mean individuals or legal entities that are not controlling, controlled by Us or You, or under common control with Us or You.
15
+ i. "You" (or "Your") shall mean a natural person or legal entity exercising the rights granted by this Agreement and/or using the Materials for any purpose and in any field of use.
16
+
17
+ 2. Grant of Rights.
18
+ You are granted a non-exclusive, worldwide, non-transferable and royalty-free limited license under Ant's intellectual property or other rights owned by Ant embodied in the Materials to use, reproduce, distribute, copy, create derivative works of, and make modifications to the Materials.
19
+
20
+ 3. Redistribution.
21
+ You may distribute or make the Materials or derivative works thereof available to a Third Party in any medium, with or without modifications, and in Source or Object form, provided that You meet the following conditions:
22
+ a. You shall provide a copy of this Agreement to such Third Party;
23
+ b. if You modify the CodeFuse model, You shall provide a prominent notice, stating how You have modified the CodeFuse model, to such Third Party; and
24
+ c. You shall retain in all copies of the Materials that You distribute the following attribution notices within a "Notice" text file distributed as a part of such copies: "CodeFuse is licensed under the CodeFuse COMMUNITY LICENSE AGREEMENT, Copyright (c) Ant Group. All Rights Reserved."
25
+ You may add Your own copyright statement to Your modifications and may provide additional or different license terms and conditions for use, reproduction, or distribution of Your modifications, or for any such derivative works as a whole, provided Your use, reproduction, and distribution of the work otherwise complies with the terms and conditions of this Agreement.
26
+
27
+ 4. Rules of Use.
28
+ You shall comply with applicable laws and regulations (including without limitation export controls or restrictions) in Your use of the Materials.
29
+
30
+ 5. Intellectual Property.
31
+ a. Ant retains ownership of all intellectual property rights in and to the Materials and derivatives made by or for Ant. Conditioned upon compliance with the terms and conditions of this Agreement, with respect to any derivative works and modifications of the Materials that are made by You, You are and will be the owner of such derivative works and modifications.
32
+ b. No trademark license is granted to use the trade names, trademarks, service marks, or product names of Ant, except as required to fulfill notice requirements under this Agreement or as required for reasonable and customary use in describing and redistributing the Materials.
33
+ c. If You commence a lawsuit or other proceedings (including a cross-claim or counterclaim in a lawsuit) against Ant or any entity alleging that the Materials or any output therefrom, or any part of the foregoing, infringe any intellectual property or other right owned or licensable by You, then all licences granted to You under this Agreement shall terminate as of the date such lawsuit or other proceeding is commenced or brought.
34
+
35
+ 6. Disclaimer of Warranty and Limitation of Liability.
36
+ a. Ant is not obligated to support, update, provide training for, or develop any further version of the Materials or to grant any license thereto.
37
+ b. THE MATERIALS ARE PROVIDED "AS IS" WITHOUT ANY EXPRESS OR IMPLIED WARRANTY OF ANY KIND INCLUDING WARRANTIES OF TITLE, MERCHANTABILITY, NONINFRINGEMENT, OR FITNESS FOR A PARTICULAR PURPOSE. WE MAKE NO WARRANTY AND ASSUME NO RESPONSIBILITY FOR THE SAFETY OR STABILITY OF THE MATERIALS AND ANY OUTPUT THEREFROM.
38
+ c. YOU ARE SOLELY RESPONSIBLE FOR DETERMINING THE APPROPRIATENESS OF USING OR REDISTRIBUTING THE MATERIALS AND ASSUME ANY RISKS ASSOCIATED WITH YOUR USE OF THE MATERIALS AND ANY OUTPUT AND RESULTS. IN NO EVENT SHALL WE BE LIABLE TO YOU FOR ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, TORT, NEGLIGENCE, PRODUCTS LIABILITY, OR OTHERWISE, ARISING OUT OF THIS AGREEMENT OR ARISING FROM YOUR USE OR INABILITY TO USE THE MATERIALS OR ANY OUTPUT OF IT, FOR ANY DIRECT, OR INDIRECT, SPECIAL, CONSEQUENTIAL, INCIDENTAL, EXEMPLARY OR PUNITIVE DAMAGES, NO MATTER HOW IT'S CAUSED OR EVEN IF ANT OR ITS AFFILIATES HAVE BEEN ADVISED OF THE POSSIBILITY OF ANY OF THE FOREGOING.
39
+ d. You will defend, indemnify and hold harmless Ant from and against any claim by any Third Party arising out of or related to Your use or distribution of the Materials.
40
+
41
+ 7. Survival and Termination.
42
+ a. The term of this Agreement shall commence upon Your acceptance of this Agreement or access to the Materials and will continue in full force and effect until terminated in accordance with the terms and conditions herein.
43
+ b. We may terminate this Agreement if You breach any of the terms or conditions of this Agreement. Upon termination of this Agreement, You must delete and cease use of the Materials. Sections 6 and 8 shall survive the termination of this Agreement.
44
+
45
+ 8. Governing Law and Jurisdiction.
46
+ a. This Agreement and any dispute arising out of or relating to it, whether in contract, tort, negligence, products liability, or otherwise, will be governed by the laws of China, without regard to conflict of law principles, and the UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement.
47
+ b. The People's Courts in Hangzhou City shall have exclusive jurisdiction over any dispute arising out of this Agreement.
config.json CHANGED
@@ -1,3 +1,29 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:214855c09d7b611d331172926256fed1d195c7976a71845b14e890200e79dec9
3
- size 724
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_name_or_path": "/CodeFuse-13B",
3
+ "architectures": [
4
+ "GPTNeoXForCausalLM"
5
+ ],
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 100256,
8
+ "classifier_dropout": 0.1,
9
+ "eos_token_id": 100256,
10
+ "hidden_act": "gelu",
11
+ "hidden_dropout": 0.0,
12
+ "hidden_size": 5120,
13
+ "initializer_range": 0.02,
14
+ "intermediate_size": 20480,
15
+ "layer_norm_eps": 1e-05,
16
+ "max_position_embeddings": 4096,
17
+ "model_type": "gpt_neox",
18
+ "num_attention_heads": 40,
19
+ "num_hidden_layers": 40,
20
+ "rope_scaling": null,
21
+ "rotary_emb_base": 10000,
22
+ "rotary_pct": 1.0,
23
+ "tie_word_embeddings": false,
24
+ "torch_dtype": "float16",
25
+ "transformers_version": "4.31.0",
26
+ "use_cache": false,
27
+ "use_parallel_residual": true,
28
+ "vocab_size": 100831
29
+ }
generation_config.json CHANGED
@@ -1,3 +1,6 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:0409601bbbd56f08fac8d63b7d5b7c345e7abcee0fa3e1308638c3422d2c2ec5
3
- size 121
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 100256,
4
+ "eos_token_id": 100256,
5
+ "transformers_version": "4.31.0"
6
+ }
pytorch_model.bin.index.json CHANGED
@@ -1,3 +1,531 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:6afb71f636fec1bd4697f0b6228fa2864161dded61247fd42cdc30e1c3d65f3c
3
- size 46064
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "metadata": {
3
+ "total_size": 54472386560
4
+ },
5
+ "weight_map": {
6
+ "embed_out.weight": "pytorch_model-00006-of-00006.bin",
7
+ "gpt_neox.embed_in.weight": "pytorch_model-00001-of-00006.bin",
8
+ "gpt_neox.final_layer_norm.bias": "pytorch_model-00006-of-00006.bin",
9
+ "gpt_neox.final_layer_norm.weight": "pytorch_model-00006-of-00006.bin",
10
+ "gpt_neox.layers.0.attention.dense.bias": "pytorch_model-00001-of-00006.bin",
11
+ "gpt_neox.layers.0.attention.dense.weight": "pytorch_model-00001-of-00006.bin",
12
+ "gpt_neox.layers.0.attention.query_key_value.bias": "pytorch_model-00001-of-00006.bin",
13
+ "gpt_neox.layers.0.attention.query_key_value.weight": "pytorch_model-00001-of-00006.bin",
14
+ "gpt_neox.layers.0.attention.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
15
+ "gpt_neox.layers.0.input_layernorm.bias": "pytorch_model-00001-of-00006.bin",
16
+ "gpt_neox.layers.0.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
17
+ "gpt_neox.layers.0.mlp.dense_4h_to_h.bias": "pytorch_model-00001-of-00006.bin",
18
+ "gpt_neox.layers.0.mlp.dense_4h_to_h.weight": "pytorch_model-00001-of-00006.bin",
19
+ "gpt_neox.layers.0.mlp.dense_h_to_4h.bias": "pytorch_model-00001-of-00006.bin",
20
+ "gpt_neox.layers.0.mlp.dense_h_to_4h.weight": "pytorch_model-00001-of-00006.bin",
21
+ "gpt_neox.layers.0.post_attention_layernorm.bias": "pytorch_model-00001-of-00006.bin",
22
+ "gpt_neox.layers.0.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
23
+ "gpt_neox.layers.1.attention.dense.bias": "pytorch_model-00001-of-00006.bin",
24
+ "gpt_neox.layers.1.attention.dense.weight": "pytorch_model-00001-of-00006.bin",
25
+ "gpt_neox.layers.1.attention.query_key_value.bias": "pytorch_model-00001-of-00006.bin",
26
+ "gpt_neox.layers.1.attention.query_key_value.weight": "pytorch_model-00001-of-00006.bin",
27
+ "gpt_neox.layers.1.attention.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
28
+ "gpt_neox.layers.1.input_layernorm.bias": "pytorch_model-00001-of-00006.bin",
29
+ "gpt_neox.layers.1.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
30
+ "gpt_neox.layers.1.mlp.dense_4h_to_h.bias": "pytorch_model-00001-of-00006.bin",
31
+ "gpt_neox.layers.1.mlp.dense_4h_to_h.weight": "pytorch_model-00001-of-00006.bin",
32
+ "gpt_neox.layers.1.mlp.dense_h_to_4h.bias": "pytorch_model-00001-of-00006.bin",
33
+ "gpt_neox.layers.1.mlp.dense_h_to_4h.weight": "pytorch_model-00001-of-00006.bin",
34
+ "gpt_neox.layers.1.post_attention_layernorm.bias": "pytorch_model-00001-of-00006.bin",
35
+ "gpt_neox.layers.1.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
36
+ "gpt_neox.layers.10.attention.dense.bias": "pytorch_model-00002-of-00006.bin",
37
+ "gpt_neox.layers.10.attention.dense.weight": "pytorch_model-00002-of-00006.bin",
38
+ "gpt_neox.layers.10.attention.query_key_value.bias": "pytorch_model-00002-of-00006.bin",
39
+ "gpt_neox.layers.10.attention.query_key_value.weight": "pytorch_model-00002-of-00006.bin",
40
+ "gpt_neox.layers.10.attention.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
41
+ "gpt_neox.layers.10.input_layernorm.bias": "pytorch_model-00002-of-00006.bin",
42
+ "gpt_neox.layers.10.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
43
+ "gpt_neox.layers.10.mlp.dense_4h_to_h.bias": "pytorch_model-00002-of-00006.bin",
44
+ "gpt_neox.layers.10.mlp.dense_4h_to_h.weight": "pytorch_model-00002-of-00006.bin",
45
+ "gpt_neox.layers.10.mlp.dense_h_to_4h.bias": "pytorch_model-00002-of-00006.bin",
46
+ "gpt_neox.layers.10.mlp.dense_h_to_4h.weight": "pytorch_model-00002-of-00006.bin",
47
+ "gpt_neox.layers.10.post_attention_layernorm.bias": "pytorch_model-00002-of-00006.bin",
48
+ "gpt_neox.layers.10.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
49
+ "gpt_neox.layers.11.attention.dense.bias": "pytorch_model-00002-of-00006.bin",
50
+ "gpt_neox.layers.11.attention.dense.weight": "pytorch_model-00002-of-00006.bin",
51
+ "gpt_neox.layers.11.attention.query_key_value.bias": "pytorch_model-00002-of-00006.bin",
52
+ "gpt_neox.layers.11.attention.query_key_value.weight": "pytorch_model-00002-of-00006.bin",
53
+ "gpt_neox.layers.11.attention.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
54
+ "gpt_neox.layers.11.input_layernorm.bias": "pytorch_model-00002-of-00006.bin",
55
+ "gpt_neox.layers.11.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
56
+ "gpt_neox.layers.11.mlp.dense_4h_to_h.bias": "pytorch_model-00002-of-00006.bin",
57
+ "gpt_neox.layers.11.mlp.dense_4h_to_h.weight": "pytorch_model-00002-of-00006.bin",
58
+ "gpt_neox.layers.11.mlp.dense_h_to_4h.bias": "pytorch_model-00002-of-00006.bin",
59
+ "gpt_neox.layers.11.mlp.dense_h_to_4h.weight": "pytorch_model-00002-of-00006.bin",
60
+ "gpt_neox.layers.11.post_attention_layernorm.bias": "pytorch_model-00002-of-00006.bin",
61
+ "gpt_neox.layers.11.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
62
+ "gpt_neox.layers.12.attention.dense.bias": "pytorch_model-00002-of-00006.bin",
63
+ "gpt_neox.layers.12.attention.dense.weight": "pytorch_model-00002-of-00006.bin",
64
+ "gpt_neox.layers.12.attention.query_key_value.bias": "pytorch_model-00002-of-00006.bin",
65
+ "gpt_neox.layers.12.attention.query_key_value.weight": "pytorch_model-00002-of-00006.bin",
66
+ "gpt_neox.layers.12.attention.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
67
+ "gpt_neox.layers.12.input_layernorm.bias": "pytorch_model-00002-of-00006.bin",
68
+ "gpt_neox.layers.12.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
69
+ "gpt_neox.layers.12.mlp.dense_4h_to_h.bias": "pytorch_model-00002-of-00006.bin",
70
+ "gpt_neox.layers.12.mlp.dense_4h_to_h.weight": "pytorch_model-00002-of-00006.bin",
71
+ "gpt_neox.layers.12.mlp.dense_h_to_4h.bias": "pytorch_model-00002-of-00006.bin",
72
+ "gpt_neox.layers.12.mlp.dense_h_to_4h.weight": "pytorch_model-00002-of-00006.bin",
73
+ "gpt_neox.layers.12.post_attention_layernorm.bias": "pytorch_model-00002-of-00006.bin",
74
+ "gpt_neox.layers.12.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
75
+ "gpt_neox.layers.13.attention.dense.bias": "pytorch_model-00002-of-00006.bin",
76
+ "gpt_neox.layers.13.attention.dense.weight": "pytorch_model-00002-of-00006.bin",
77
+ "gpt_neox.layers.13.attention.query_key_value.bias": "pytorch_model-00002-of-00006.bin",
78
+ "gpt_neox.layers.13.attention.query_key_value.weight": "pytorch_model-00002-of-00006.bin",
79
+ "gpt_neox.layers.13.attention.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
80
+ "gpt_neox.layers.13.input_layernorm.bias": "pytorch_model-00002-of-00006.bin",
81
+ "gpt_neox.layers.13.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
82
+ "gpt_neox.layers.13.mlp.dense_4h_to_h.bias": "pytorch_model-00002-of-00006.bin",
83
+ "gpt_neox.layers.13.mlp.dense_4h_to_h.weight": "pytorch_model-00002-of-00006.bin",
84
+ "gpt_neox.layers.13.mlp.dense_h_to_4h.bias": "pytorch_model-00002-of-00006.bin",
85
+ "gpt_neox.layers.13.mlp.dense_h_to_4h.weight": "pytorch_model-00002-of-00006.bin",
86
+ "gpt_neox.layers.13.post_attention_layernorm.bias": "pytorch_model-00002-of-00006.bin",
87
+ "gpt_neox.layers.13.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
88
+ "gpt_neox.layers.14.attention.dense.bias": "pytorch_model-00003-of-00006.bin",
89
+ "gpt_neox.layers.14.attention.dense.weight": "pytorch_model-00003-of-00006.bin",
90
+ "gpt_neox.layers.14.attention.query_key_value.bias": "pytorch_model-00003-of-00006.bin",
91
+ "gpt_neox.layers.14.attention.query_key_value.weight": "pytorch_model-00003-of-00006.bin",
92
+ "gpt_neox.layers.14.attention.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
93
+ "gpt_neox.layers.14.input_layernorm.bias": "pytorch_model-00002-of-00006.bin",
94
+ "gpt_neox.layers.14.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
95
+ "gpt_neox.layers.14.mlp.dense_4h_to_h.bias": "pytorch_model-00003-of-00006.bin",
96
+ "gpt_neox.layers.14.mlp.dense_4h_to_h.weight": "pytorch_model-00003-of-00006.bin",
97
+ "gpt_neox.layers.14.mlp.dense_h_to_4h.bias": "pytorch_model-00003-of-00006.bin",
98
+ "gpt_neox.layers.14.mlp.dense_h_to_4h.weight": "pytorch_model-00003-of-00006.bin",
99
+ "gpt_neox.layers.14.post_attention_layernorm.bias": "pytorch_model-00002-of-00006.bin",
100
+ "gpt_neox.layers.14.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
101
+ "gpt_neox.layers.15.attention.dense.bias": "pytorch_model-00003-of-00006.bin",
102
+ "gpt_neox.layers.15.attention.dense.weight": "pytorch_model-00003-of-00006.bin",
103
+ "gpt_neox.layers.15.attention.query_key_value.bias": "pytorch_model-00003-of-00006.bin",
104
+ "gpt_neox.layers.15.attention.query_key_value.weight": "pytorch_model-00003-of-00006.bin",
105
+ "gpt_neox.layers.15.attention.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
106
+ "gpt_neox.layers.15.input_layernorm.bias": "pytorch_model-00003-of-00006.bin",
107
+ "gpt_neox.layers.15.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
108
+ "gpt_neox.layers.15.mlp.dense_4h_to_h.bias": "pytorch_model-00003-of-00006.bin",
109
+ "gpt_neox.layers.15.mlp.dense_4h_to_h.weight": "pytorch_model-00003-of-00006.bin",
110
+ "gpt_neox.layers.15.mlp.dense_h_to_4h.bias": "pytorch_model-00003-of-00006.bin",
111
+ "gpt_neox.layers.15.mlp.dense_h_to_4h.weight": "pytorch_model-00003-of-00006.bin",
112
+ "gpt_neox.layers.15.post_attention_layernorm.bias": "pytorch_model-00003-of-00006.bin",
113
+ "gpt_neox.layers.15.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
114
+ "gpt_neox.layers.16.attention.dense.bias": "pytorch_model-00003-of-00006.bin",
115
+ "gpt_neox.layers.16.attention.dense.weight": "pytorch_model-00003-of-00006.bin",
116
+ "gpt_neox.layers.16.attention.query_key_value.bias": "pytorch_model-00003-of-00006.bin",
117
+ "gpt_neox.layers.16.attention.query_key_value.weight": "pytorch_model-00003-of-00006.bin",
118
+ "gpt_neox.layers.16.attention.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
119
+ "gpt_neox.layers.16.input_layernorm.bias": "pytorch_model-00003-of-00006.bin",
120
+ "gpt_neox.layers.16.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
121
+ "gpt_neox.layers.16.mlp.dense_4h_to_h.bias": "pytorch_model-00003-of-00006.bin",
122
+ "gpt_neox.layers.16.mlp.dense_4h_to_h.weight": "pytorch_model-00003-of-00006.bin",
123
+ "gpt_neox.layers.16.mlp.dense_h_to_4h.bias": "pytorch_model-00003-of-00006.bin",
124
+ "gpt_neox.layers.16.mlp.dense_h_to_4h.weight": "pytorch_model-00003-of-00006.bin",
125
+ "gpt_neox.layers.16.post_attention_layernorm.bias": "pytorch_model-00003-of-00006.bin",
126
+ "gpt_neox.layers.16.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
127
+ "gpt_neox.layers.17.attention.dense.bias": "pytorch_model-00003-of-00006.bin",
128
+ "gpt_neox.layers.17.attention.dense.weight": "pytorch_model-00003-of-00006.bin",
129
+ "gpt_neox.layers.17.attention.query_key_value.bias": "pytorch_model-00003-of-00006.bin",
130
+ "gpt_neox.layers.17.attention.query_key_value.weight": "pytorch_model-00003-of-00006.bin",
131
+ "gpt_neox.layers.17.attention.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
132
+ "gpt_neox.layers.17.input_layernorm.bias": "pytorch_model-00003-of-00006.bin",
133
+ "gpt_neox.layers.17.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
134
+ "gpt_neox.layers.17.mlp.dense_4h_to_h.bias": "pytorch_model-00003-of-00006.bin",
135
+ "gpt_neox.layers.17.mlp.dense_4h_to_h.weight": "pytorch_model-00003-of-00006.bin",
136
+ "gpt_neox.layers.17.mlp.dense_h_to_4h.bias": "pytorch_model-00003-of-00006.bin",
137
+ "gpt_neox.layers.17.mlp.dense_h_to_4h.weight": "pytorch_model-00003-of-00006.bin",
138
+ "gpt_neox.layers.17.post_attention_layernorm.bias": "pytorch_model-00003-of-00006.bin",
139
+ "gpt_neox.layers.17.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
140
+ "gpt_neox.layers.18.attention.dense.bias": "pytorch_model-00003-of-00006.bin",
141
+ "gpt_neox.layers.18.attention.dense.weight": "pytorch_model-00003-of-00006.bin",
142
+ "gpt_neox.layers.18.attention.query_key_value.bias": "pytorch_model-00003-of-00006.bin",
143
+ "gpt_neox.layers.18.attention.query_key_value.weight": "pytorch_model-00003-of-00006.bin",
144
+ "gpt_neox.layers.18.attention.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
145
+ "gpt_neox.layers.18.input_layernorm.bias": "pytorch_model-00003-of-00006.bin",
146
+ "gpt_neox.layers.18.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
147
+ "gpt_neox.layers.18.mlp.dense_4h_to_h.bias": "pytorch_model-00003-of-00006.bin",
148
+ "gpt_neox.layers.18.mlp.dense_4h_to_h.weight": "pytorch_model-00003-of-00006.bin",
149
+ "gpt_neox.layers.18.mlp.dense_h_to_4h.bias": "pytorch_model-00003-of-00006.bin",
150
+ "gpt_neox.layers.18.mlp.dense_h_to_4h.weight": "pytorch_model-00003-of-00006.bin",
151
+ "gpt_neox.layers.18.post_attention_layernorm.bias": "pytorch_model-00003-of-00006.bin",
152
+ "gpt_neox.layers.18.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
153
+ "gpt_neox.layers.19.attention.dense.bias": "pytorch_model-00003-of-00006.bin",
154
+ "gpt_neox.layers.19.attention.dense.weight": "pytorch_model-00003-of-00006.bin",
155
+ "gpt_neox.layers.19.attention.query_key_value.bias": "pytorch_model-00003-of-00006.bin",
156
+ "gpt_neox.layers.19.attention.query_key_value.weight": "pytorch_model-00003-of-00006.bin",
157
+ "gpt_neox.layers.19.attention.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
158
+ "gpt_neox.layers.19.input_layernorm.bias": "pytorch_model-00003-of-00006.bin",
159
+ "gpt_neox.layers.19.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
160
+ "gpt_neox.layers.19.mlp.dense_4h_to_h.bias": "pytorch_model-00003-of-00006.bin",
161
+ "gpt_neox.layers.19.mlp.dense_4h_to_h.weight": "pytorch_model-00003-of-00006.bin",
162
+ "gpt_neox.layers.19.mlp.dense_h_to_4h.bias": "pytorch_model-00003-of-00006.bin",
163
+ "gpt_neox.layers.19.mlp.dense_h_to_4h.weight": "pytorch_model-00003-of-00006.bin",
164
+ "gpt_neox.layers.19.post_attention_layernorm.bias": "pytorch_model-00003-of-00006.bin",
165
+ "gpt_neox.layers.19.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
166
+ "gpt_neox.layers.2.attention.dense.bias": "pytorch_model-00001-of-00006.bin",
167
+ "gpt_neox.layers.2.attention.dense.weight": "pytorch_model-00001-of-00006.bin",
168
+ "gpt_neox.layers.2.attention.query_key_value.bias": "pytorch_model-00001-of-00006.bin",
169
+ "gpt_neox.layers.2.attention.query_key_value.weight": "pytorch_model-00001-of-00006.bin",
170
+ "gpt_neox.layers.2.attention.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
171
+ "gpt_neox.layers.2.input_layernorm.bias": "pytorch_model-00001-of-00006.bin",
172
+ "gpt_neox.layers.2.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
173
+ "gpt_neox.layers.2.mlp.dense_4h_to_h.bias": "pytorch_model-00001-of-00006.bin",
174
+ "gpt_neox.layers.2.mlp.dense_4h_to_h.weight": "pytorch_model-00001-of-00006.bin",
175
+ "gpt_neox.layers.2.mlp.dense_h_to_4h.bias": "pytorch_model-00001-of-00006.bin",
176
+ "gpt_neox.layers.2.mlp.dense_h_to_4h.weight": "pytorch_model-00001-of-00006.bin",
177
+ "gpt_neox.layers.2.post_attention_layernorm.bias": "pytorch_model-00001-of-00006.bin",
178
+ "gpt_neox.layers.2.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
179
+ "gpt_neox.layers.20.attention.dense.bias": "pytorch_model-00003-of-00006.bin",
180
+ "gpt_neox.layers.20.attention.dense.weight": "pytorch_model-00003-of-00006.bin",
181
+ "gpt_neox.layers.20.attention.query_key_value.bias": "pytorch_model-00003-of-00006.bin",
182
+ "gpt_neox.layers.20.attention.query_key_value.weight": "pytorch_model-00003-of-00006.bin",
183
+ "gpt_neox.layers.20.attention.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
184
+ "gpt_neox.layers.20.input_layernorm.bias": "pytorch_model-00003-of-00006.bin",
185
+ "gpt_neox.layers.20.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
186
+ "gpt_neox.layers.20.mlp.dense_4h_to_h.bias": "pytorch_model-00003-of-00006.bin",
187
+ "gpt_neox.layers.20.mlp.dense_4h_to_h.weight": "pytorch_model-00003-of-00006.bin",
188
+ "gpt_neox.layers.20.mlp.dense_h_to_4h.bias": "pytorch_model-00003-of-00006.bin",
189
+ "gpt_neox.layers.20.mlp.dense_h_to_4h.weight": "pytorch_model-00003-of-00006.bin",
190
+ "gpt_neox.layers.20.post_attention_layernorm.bias": "pytorch_model-00003-of-00006.bin",
191
+ "gpt_neox.layers.20.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
192
+ "gpt_neox.layers.21.attention.dense.bias": "pytorch_model-00003-of-00006.bin",
193
+ "gpt_neox.layers.21.attention.dense.weight": "pytorch_model-00003-of-00006.bin",
194
+ "gpt_neox.layers.21.attention.query_key_value.bias": "pytorch_model-00003-of-00006.bin",
195
+ "gpt_neox.layers.21.attention.query_key_value.weight": "pytorch_model-00003-of-00006.bin",
196
+ "gpt_neox.layers.21.attention.rotary_emb.inv_freq": "pytorch_model-00003-of-00006.bin",
197
+ "gpt_neox.layers.21.input_layernorm.bias": "pytorch_model-00003-of-00006.bin",
198
+ "gpt_neox.layers.21.input_layernorm.weight": "pytorch_model-00003-of-00006.bin",
199
+ "gpt_neox.layers.21.mlp.dense_4h_to_h.bias": "pytorch_model-00004-of-00006.bin",
200
+ "gpt_neox.layers.21.mlp.dense_4h_to_h.weight": "pytorch_model-00004-of-00006.bin",
201
+ "gpt_neox.layers.21.mlp.dense_h_to_4h.bias": "pytorch_model-00003-of-00006.bin",
202
+ "gpt_neox.layers.21.mlp.dense_h_to_4h.weight": "pytorch_model-00003-of-00006.bin",
203
+ "gpt_neox.layers.21.post_attention_layernorm.bias": "pytorch_model-00003-of-00006.bin",
204
+ "gpt_neox.layers.21.post_attention_layernorm.weight": "pytorch_model-00003-of-00006.bin",
205
+ "gpt_neox.layers.22.attention.dense.bias": "pytorch_model-00004-of-00006.bin",
206
+ "gpt_neox.layers.22.attention.dense.weight": "pytorch_model-00004-of-00006.bin",
207
+ "gpt_neox.layers.22.attention.query_key_value.bias": "pytorch_model-00004-of-00006.bin",
208
+ "gpt_neox.layers.22.attention.query_key_value.weight": "pytorch_model-00004-of-00006.bin",
209
+ "gpt_neox.layers.22.attention.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
210
+ "gpt_neox.layers.22.input_layernorm.bias": "pytorch_model-00004-of-00006.bin",
211
+ "gpt_neox.layers.22.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
212
+ "gpt_neox.layers.22.mlp.dense_4h_to_h.bias": "pytorch_model-00004-of-00006.bin",
213
+ "gpt_neox.layers.22.mlp.dense_4h_to_h.weight": "pytorch_model-00004-of-00006.bin",
214
+ "gpt_neox.layers.22.mlp.dense_h_to_4h.bias": "pytorch_model-00004-of-00006.bin",
215
+ "gpt_neox.layers.22.mlp.dense_h_to_4h.weight": "pytorch_model-00004-of-00006.bin",
216
+ "gpt_neox.layers.22.post_attention_layernorm.bias": "pytorch_model-00004-of-00006.bin",
217
+ "gpt_neox.layers.22.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
218
+ "gpt_neox.layers.23.attention.dense.bias": "pytorch_model-00004-of-00006.bin",
219
+ "gpt_neox.layers.23.attention.dense.weight": "pytorch_model-00004-of-00006.bin",
220
+ "gpt_neox.layers.23.attention.query_key_value.bias": "pytorch_model-00004-of-00006.bin",
221
+ "gpt_neox.layers.23.attention.query_key_value.weight": "pytorch_model-00004-of-00006.bin",
222
+ "gpt_neox.layers.23.attention.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
223
+ "gpt_neox.layers.23.input_layernorm.bias": "pytorch_model-00004-of-00006.bin",
224
+ "gpt_neox.layers.23.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
225
+ "gpt_neox.layers.23.mlp.dense_4h_to_h.bias": "pytorch_model-00004-of-00006.bin",
226
+ "gpt_neox.layers.23.mlp.dense_4h_to_h.weight": "pytorch_model-00004-of-00006.bin",
227
+ "gpt_neox.layers.23.mlp.dense_h_to_4h.bias": "pytorch_model-00004-of-00006.bin",
228
+ "gpt_neox.layers.23.mlp.dense_h_to_4h.weight": "pytorch_model-00004-of-00006.bin",
229
+ "gpt_neox.layers.23.post_attention_layernorm.bias": "pytorch_model-00004-of-00006.bin",
230
+ "gpt_neox.layers.23.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
231
+ "gpt_neox.layers.24.attention.dense.bias": "pytorch_model-00004-of-00006.bin",
232
+ "gpt_neox.layers.24.attention.dense.weight": "pytorch_model-00004-of-00006.bin",
233
+ "gpt_neox.layers.24.attention.query_key_value.bias": "pytorch_model-00004-of-00006.bin",
234
+ "gpt_neox.layers.24.attention.query_key_value.weight": "pytorch_model-00004-of-00006.bin",
235
+ "gpt_neox.layers.24.attention.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
236
+ "gpt_neox.layers.24.input_layernorm.bias": "pytorch_model-00004-of-00006.bin",
237
+ "gpt_neox.layers.24.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
238
+ "gpt_neox.layers.24.mlp.dense_4h_to_h.bias": "pytorch_model-00004-of-00006.bin",
239
+ "gpt_neox.layers.24.mlp.dense_4h_to_h.weight": "pytorch_model-00004-of-00006.bin",
240
+ "gpt_neox.layers.24.mlp.dense_h_to_4h.bias": "pytorch_model-00004-of-00006.bin",
241
+ "gpt_neox.layers.24.mlp.dense_h_to_4h.weight": "pytorch_model-00004-of-00006.bin",
242
+ "gpt_neox.layers.24.post_attention_layernorm.bias": "pytorch_model-00004-of-00006.bin",
243
+ "gpt_neox.layers.24.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
244
+ "gpt_neox.layers.25.attention.dense.bias": "pytorch_model-00004-of-00006.bin",
245
+ "gpt_neox.layers.25.attention.dense.weight": "pytorch_model-00004-of-00006.bin",
246
+ "gpt_neox.layers.25.attention.query_key_value.bias": "pytorch_model-00004-of-00006.bin",
247
+ "gpt_neox.layers.25.attention.query_key_value.weight": "pytorch_model-00004-of-00006.bin",
248
+ "gpt_neox.layers.25.attention.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
249
+ "gpt_neox.layers.25.input_layernorm.bias": "pytorch_model-00004-of-00006.bin",
250
+ "gpt_neox.layers.25.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
251
+ "gpt_neox.layers.25.mlp.dense_4h_to_h.bias": "pytorch_model-00004-of-00006.bin",
252
+ "gpt_neox.layers.25.mlp.dense_4h_to_h.weight": "pytorch_model-00004-of-00006.bin",
253
+ "gpt_neox.layers.25.mlp.dense_h_to_4h.bias": "pytorch_model-00004-of-00006.bin",
254
+ "gpt_neox.layers.25.mlp.dense_h_to_4h.weight": "pytorch_model-00004-of-00006.bin",
255
+ "gpt_neox.layers.25.post_attention_layernorm.bias": "pytorch_model-00004-of-00006.bin",
256
+ "gpt_neox.layers.25.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
257
+ "gpt_neox.layers.26.attention.dense.bias": "pytorch_model-00004-of-00006.bin",
258
+ "gpt_neox.layers.26.attention.dense.weight": "pytorch_model-00004-of-00006.bin",
259
+ "gpt_neox.layers.26.attention.query_key_value.bias": "pytorch_model-00004-of-00006.bin",
260
+ "gpt_neox.layers.26.attention.query_key_value.weight": "pytorch_model-00004-of-00006.bin",
261
+ "gpt_neox.layers.26.attention.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
262
+ "gpt_neox.layers.26.input_layernorm.bias": "pytorch_model-00004-of-00006.bin",
263
+ "gpt_neox.layers.26.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
264
+ "gpt_neox.layers.26.mlp.dense_4h_to_h.bias": "pytorch_model-00004-of-00006.bin",
265
+ "gpt_neox.layers.26.mlp.dense_4h_to_h.weight": "pytorch_model-00004-of-00006.bin",
266
+ "gpt_neox.layers.26.mlp.dense_h_to_4h.bias": "pytorch_model-00004-of-00006.bin",
267
+ "gpt_neox.layers.26.mlp.dense_h_to_4h.weight": "pytorch_model-00004-of-00006.bin",
268
+ "gpt_neox.layers.26.post_attention_layernorm.bias": "pytorch_model-00004-of-00006.bin",
269
+ "gpt_neox.layers.26.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
270
+ "gpt_neox.layers.27.attention.dense.bias": "pytorch_model-00004-of-00006.bin",
271
+ "gpt_neox.layers.27.attention.dense.weight": "pytorch_model-00004-of-00006.bin",
272
+ "gpt_neox.layers.27.attention.query_key_value.bias": "pytorch_model-00004-of-00006.bin",
273
+ "gpt_neox.layers.27.attention.query_key_value.weight": "pytorch_model-00004-of-00006.bin",
274
+ "gpt_neox.layers.27.attention.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
275
+ "gpt_neox.layers.27.input_layernorm.bias": "pytorch_model-00004-of-00006.bin",
276
+ "gpt_neox.layers.27.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
277
+ "gpt_neox.layers.27.mlp.dense_4h_to_h.bias": "pytorch_model-00004-of-00006.bin",
278
+ "gpt_neox.layers.27.mlp.dense_4h_to_h.weight": "pytorch_model-00004-of-00006.bin",
279
+ "gpt_neox.layers.27.mlp.dense_h_to_4h.bias": "pytorch_model-00004-of-00006.bin",
280
+ "gpt_neox.layers.27.mlp.dense_h_to_4h.weight": "pytorch_model-00004-of-00006.bin",
281
+ "gpt_neox.layers.27.post_attention_layernorm.bias": "pytorch_model-00004-of-00006.bin",
282
+ "gpt_neox.layers.27.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
283
+ "gpt_neox.layers.28.attention.dense.bias": "pytorch_model-00004-of-00006.bin",
284
+ "gpt_neox.layers.28.attention.dense.weight": "pytorch_model-00004-of-00006.bin",
285
+ "gpt_neox.layers.28.attention.query_key_value.bias": "pytorch_model-00004-of-00006.bin",
286
+ "gpt_neox.layers.28.attention.query_key_value.weight": "pytorch_model-00004-of-00006.bin",
287
+ "gpt_neox.layers.28.attention.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
288
+ "gpt_neox.layers.28.input_layernorm.bias": "pytorch_model-00004-of-00006.bin",
289
+ "gpt_neox.layers.28.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
290
+ "gpt_neox.layers.28.mlp.dense_4h_to_h.bias": "pytorch_model-00004-of-00006.bin",
291
+ "gpt_neox.layers.28.mlp.dense_4h_to_h.weight": "pytorch_model-00004-of-00006.bin",
292
+ "gpt_neox.layers.28.mlp.dense_h_to_4h.bias": "pytorch_model-00004-of-00006.bin",
293
+ "gpt_neox.layers.28.mlp.dense_h_to_4h.weight": "pytorch_model-00004-of-00006.bin",
294
+ "gpt_neox.layers.28.post_attention_layernorm.bias": "pytorch_model-00004-of-00006.bin",
295
+ "gpt_neox.layers.28.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
296
+ "gpt_neox.layers.29.attention.dense.bias": "pytorch_model-00004-of-00006.bin",
297
+ "gpt_neox.layers.29.attention.dense.weight": "pytorch_model-00004-of-00006.bin",
298
+ "gpt_neox.layers.29.attention.query_key_value.bias": "pytorch_model-00004-of-00006.bin",
299
+ "gpt_neox.layers.29.attention.query_key_value.weight": "pytorch_model-00004-of-00006.bin",
300
+ "gpt_neox.layers.29.attention.rotary_emb.inv_freq": "pytorch_model-00004-of-00006.bin",
301
+ "gpt_neox.layers.29.input_layernorm.bias": "pytorch_model-00004-of-00006.bin",
302
+ "gpt_neox.layers.29.input_layernorm.weight": "pytorch_model-00004-of-00006.bin",
303
+ "gpt_neox.layers.29.mlp.dense_4h_to_h.bias": "pytorch_model-00005-of-00006.bin",
304
+ "gpt_neox.layers.29.mlp.dense_4h_to_h.weight": "pytorch_model-00005-of-00006.bin",
305
+ "gpt_neox.layers.29.mlp.dense_h_to_4h.bias": "pytorch_model-00005-of-00006.bin",
306
+ "gpt_neox.layers.29.mlp.dense_h_to_4h.weight": "pytorch_model-00005-of-00006.bin",
307
+ "gpt_neox.layers.29.post_attention_layernorm.bias": "pytorch_model-00004-of-00006.bin",
308
+ "gpt_neox.layers.29.post_attention_layernorm.weight": "pytorch_model-00004-of-00006.bin",
309
+ "gpt_neox.layers.3.attention.dense.bias": "pytorch_model-00001-of-00006.bin",
310
+ "gpt_neox.layers.3.attention.dense.weight": "pytorch_model-00001-of-00006.bin",
311
+ "gpt_neox.layers.3.attention.query_key_value.bias": "pytorch_model-00001-of-00006.bin",
312
+ "gpt_neox.layers.3.attention.query_key_value.weight": "pytorch_model-00001-of-00006.bin",
313
+ "gpt_neox.layers.3.attention.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
314
+ "gpt_neox.layers.3.input_layernorm.bias": "pytorch_model-00001-of-00006.bin",
315
+ "gpt_neox.layers.3.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
316
+ "gpt_neox.layers.3.mlp.dense_4h_to_h.bias": "pytorch_model-00001-of-00006.bin",
317
+ "gpt_neox.layers.3.mlp.dense_4h_to_h.weight": "pytorch_model-00001-of-00006.bin",
318
+ "gpt_neox.layers.3.mlp.dense_h_to_4h.bias": "pytorch_model-00001-of-00006.bin",
319
+ "gpt_neox.layers.3.mlp.dense_h_to_4h.weight": "pytorch_model-00001-of-00006.bin",
320
+ "gpt_neox.layers.3.post_attention_layernorm.bias": "pytorch_model-00001-of-00006.bin",
321
+ "gpt_neox.layers.3.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
322
+ "gpt_neox.layers.30.attention.dense.bias": "pytorch_model-00005-of-00006.bin",
323
+ "gpt_neox.layers.30.attention.dense.weight": "pytorch_model-00005-of-00006.bin",
324
+ "gpt_neox.layers.30.attention.query_key_value.bias": "pytorch_model-00005-of-00006.bin",
325
+ "gpt_neox.layers.30.attention.query_key_value.weight": "pytorch_model-00005-of-00006.bin",
326
+ "gpt_neox.layers.30.attention.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
327
+ "gpt_neox.layers.30.input_layernorm.bias": "pytorch_model-00005-of-00006.bin",
328
+ "gpt_neox.layers.30.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
329
+ "gpt_neox.layers.30.mlp.dense_4h_to_h.bias": "pytorch_model-00005-of-00006.bin",
330
+ "gpt_neox.layers.30.mlp.dense_4h_to_h.weight": "pytorch_model-00005-of-00006.bin",
331
+ "gpt_neox.layers.30.mlp.dense_h_to_4h.bias": "pytorch_model-00005-of-00006.bin",
332
+ "gpt_neox.layers.30.mlp.dense_h_to_4h.weight": "pytorch_model-00005-of-00006.bin",
333
+ "gpt_neox.layers.30.post_attention_layernorm.bias": "pytorch_model-00005-of-00006.bin",
334
+ "gpt_neox.layers.30.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
335
+ "gpt_neox.layers.31.attention.dense.bias": "pytorch_model-00005-of-00006.bin",
336
+ "gpt_neox.layers.31.attention.dense.weight": "pytorch_model-00005-of-00006.bin",
337
+ "gpt_neox.layers.31.attention.query_key_value.bias": "pytorch_model-00005-of-00006.bin",
338
+ "gpt_neox.layers.31.attention.query_key_value.weight": "pytorch_model-00005-of-00006.bin",
339
+ "gpt_neox.layers.31.attention.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
340
+ "gpt_neox.layers.31.input_layernorm.bias": "pytorch_model-00005-of-00006.bin",
341
+ "gpt_neox.layers.31.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
342
+ "gpt_neox.layers.31.mlp.dense_4h_to_h.bias": "pytorch_model-00005-of-00006.bin",
343
+ "gpt_neox.layers.31.mlp.dense_4h_to_h.weight": "pytorch_model-00005-of-00006.bin",
344
+ "gpt_neox.layers.31.mlp.dense_h_to_4h.bias": "pytorch_model-00005-of-00006.bin",
345
+ "gpt_neox.layers.31.mlp.dense_h_to_4h.weight": "pytorch_model-00005-of-00006.bin",
346
+ "gpt_neox.layers.31.post_attention_layernorm.bias": "pytorch_model-00005-of-00006.bin",
347
+ "gpt_neox.layers.31.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
348
+ "gpt_neox.layers.32.attention.dense.bias": "pytorch_model-00005-of-00006.bin",
349
+ "gpt_neox.layers.32.attention.dense.weight": "pytorch_model-00005-of-00006.bin",
350
+ "gpt_neox.layers.32.attention.query_key_value.bias": "pytorch_model-00005-of-00006.bin",
351
+ "gpt_neox.layers.32.attention.query_key_value.weight": "pytorch_model-00005-of-00006.bin",
352
+ "gpt_neox.layers.32.attention.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
353
+ "gpt_neox.layers.32.input_layernorm.bias": "pytorch_model-00005-of-00006.bin",
354
+ "gpt_neox.layers.32.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
355
+ "gpt_neox.layers.32.mlp.dense_4h_to_h.bias": "pytorch_model-00005-of-00006.bin",
356
+ "gpt_neox.layers.32.mlp.dense_4h_to_h.weight": "pytorch_model-00005-of-00006.bin",
357
+ "gpt_neox.layers.32.mlp.dense_h_to_4h.bias": "pytorch_model-00005-of-00006.bin",
358
+ "gpt_neox.layers.32.mlp.dense_h_to_4h.weight": "pytorch_model-00005-of-00006.bin",
359
+ "gpt_neox.layers.32.post_attention_layernorm.bias": "pytorch_model-00005-of-00006.bin",
360
+ "gpt_neox.layers.32.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
361
+ "gpt_neox.layers.33.attention.dense.bias": "pytorch_model-00005-of-00006.bin",
362
+ "gpt_neox.layers.33.attention.dense.weight": "pytorch_model-00005-of-00006.bin",
363
+ "gpt_neox.layers.33.attention.query_key_value.bias": "pytorch_model-00005-of-00006.bin",
364
+ "gpt_neox.layers.33.attention.query_key_value.weight": "pytorch_model-00005-of-00006.bin",
365
+ "gpt_neox.layers.33.attention.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
366
+ "gpt_neox.layers.33.input_layernorm.bias": "pytorch_model-00005-of-00006.bin",
367
+ "gpt_neox.layers.33.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
368
+ "gpt_neox.layers.33.mlp.dense_4h_to_h.bias": "pytorch_model-00005-of-00006.bin",
369
+ "gpt_neox.layers.33.mlp.dense_4h_to_h.weight": "pytorch_model-00005-of-00006.bin",
370
+ "gpt_neox.layers.33.mlp.dense_h_to_4h.bias": "pytorch_model-00005-of-00006.bin",
371
+ "gpt_neox.layers.33.mlp.dense_h_to_4h.weight": "pytorch_model-00005-of-00006.bin",
372
+ "gpt_neox.layers.33.post_attention_layernorm.bias": "pytorch_model-00005-of-00006.bin",
373
+ "gpt_neox.layers.33.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
374
+ "gpt_neox.layers.34.attention.dense.bias": "pytorch_model-00005-of-00006.bin",
375
+ "gpt_neox.layers.34.attention.dense.weight": "pytorch_model-00005-of-00006.bin",
376
+ "gpt_neox.layers.34.attention.query_key_value.bias": "pytorch_model-00005-of-00006.bin",
377
+ "gpt_neox.layers.34.attention.query_key_value.weight": "pytorch_model-00005-of-00006.bin",
378
+ "gpt_neox.layers.34.attention.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
379
+ "gpt_neox.layers.34.input_layernorm.bias": "pytorch_model-00005-of-00006.bin",
380
+ "gpt_neox.layers.34.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
381
+ "gpt_neox.layers.34.mlp.dense_4h_to_h.bias": "pytorch_model-00005-of-00006.bin",
382
+ "gpt_neox.layers.34.mlp.dense_4h_to_h.weight": "pytorch_model-00005-of-00006.bin",
383
+ "gpt_neox.layers.34.mlp.dense_h_to_4h.bias": "pytorch_model-00005-of-00006.bin",
384
+ "gpt_neox.layers.34.mlp.dense_h_to_4h.weight": "pytorch_model-00005-of-00006.bin",
385
+ "gpt_neox.layers.34.post_attention_layernorm.bias": "pytorch_model-00005-of-00006.bin",
386
+ "gpt_neox.layers.34.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
387
+ "gpt_neox.layers.35.attention.dense.bias": "pytorch_model-00005-of-00006.bin",
388
+ "gpt_neox.layers.35.attention.dense.weight": "pytorch_model-00005-of-00006.bin",
389
+ "gpt_neox.layers.35.attention.query_key_value.bias": "pytorch_model-00005-of-00006.bin",
390
+ "gpt_neox.layers.35.attention.query_key_value.weight": "pytorch_model-00005-of-00006.bin",
391
+ "gpt_neox.layers.35.attention.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
392
+ "gpt_neox.layers.35.input_layernorm.bias": "pytorch_model-00005-of-00006.bin",
393
+ "gpt_neox.layers.35.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
394
+ "gpt_neox.layers.35.mlp.dense_4h_to_h.bias": "pytorch_model-00005-of-00006.bin",
395
+ "gpt_neox.layers.35.mlp.dense_4h_to_h.weight": "pytorch_model-00005-of-00006.bin",
396
+ "gpt_neox.layers.35.mlp.dense_h_to_4h.bias": "pytorch_model-00005-of-00006.bin",
397
+ "gpt_neox.layers.35.mlp.dense_h_to_4h.weight": "pytorch_model-00005-of-00006.bin",
398
+ "gpt_neox.layers.35.post_attention_layernorm.bias": "pytorch_model-00005-of-00006.bin",
399
+ "gpt_neox.layers.35.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
400
+ "gpt_neox.layers.36.attention.dense.bias": "pytorch_model-00005-of-00006.bin",
401
+ "gpt_neox.layers.36.attention.dense.weight": "pytorch_model-00005-of-00006.bin",
402
+ "gpt_neox.layers.36.attention.query_key_value.bias": "pytorch_model-00005-of-00006.bin",
403
+ "gpt_neox.layers.36.attention.query_key_value.weight": "pytorch_model-00005-of-00006.bin",
404
+ "gpt_neox.layers.36.attention.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
405
+ "gpt_neox.layers.36.input_layernorm.bias": "pytorch_model-00005-of-00006.bin",
406
+ "gpt_neox.layers.36.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
407
+ "gpt_neox.layers.36.mlp.dense_4h_to_h.bias": "pytorch_model-00005-of-00006.bin",
408
+ "gpt_neox.layers.36.mlp.dense_4h_to_h.weight": "pytorch_model-00005-of-00006.bin",
409
+ "gpt_neox.layers.36.mlp.dense_h_to_4h.bias": "pytorch_model-00005-of-00006.bin",
410
+ "gpt_neox.layers.36.mlp.dense_h_to_4h.weight": "pytorch_model-00005-of-00006.bin",
411
+ "gpt_neox.layers.36.post_attention_layernorm.bias": "pytorch_model-00005-of-00006.bin",
412
+ "gpt_neox.layers.36.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
413
+ "gpt_neox.layers.37.attention.dense.bias": "pytorch_model-00006-of-00006.bin",
414
+ "gpt_neox.layers.37.attention.dense.weight": "pytorch_model-00006-of-00006.bin",
415
+ "gpt_neox.layers.37.attention.query_key_value.bias": "pytorch_model-00005-of-00006.bin",
416
+ "gpt_neox.layers.37.attention.query_key_value.weight": "pytorch_model-00005-of-00006.bin",
417
+ "gpt_neox.layers.37.attention.rotary_emb.inv_freq": "pytorch_model-00005-of-00006.bin",
418
+ "gpt_neox.layers.37.input_layernorm.bias": "pytorch_model-00005-of-00006.bin",
419
+ "gpt_neox.layers.37.input_layernorm.weight": "pytorch_model-00005-of-00006.bin",
420
+ "gpt_neox.layers.37.mlp.dense_4h_to_h.bias": "pytorch_model-00006-of-00006.bin",
421
+ "gpt_neox.layers.37.mlp.dense_4h_to_h.weight": "pytorch_model-00006-of-00006.bin",
422
+ "gpt_neox.layers.37.mlp.dense_h_to_4h.bias": "pytorch_model-00006-of-00006.bin",
423
+ "gpt_neox.layers.37.mlp.dense_h_to_4h.weight": "pytorch_model-00006-of-00006.bin",
424
+ "gpt_neox.layers.37.post_attention_layernorm.bias": "pytorch_model-00005-of-00006.bin",
425
+ "gpt_neox.layers.37.post_attention_layernorm.weight": "pytorch_model-00005-of-00006.bin",
426
+ "gpt_neox.layers.38.attention.dense.bias": "pytorch_model-00006-of-00006.bin",
427
+ "gpt_neox.layers.38.attention.dense.weight": "pytorch_model-00006-of-00006.bin",
428
+ "gpt_neox.layers.38.attention.query_key_value.bias": "pytorch_model-00006-of-00006.bin",
429
+ "gpt_neox.layers.38.attention.query_key_value.weight": "pytorch_model-00006-of-00006.bin",
430
+ "gpt_neox.layers.38.attention.rotary_emb.inv_freq": "pytorch_model-00006-of-00006.bin",
431
+ "gpt_neox.layers.38.input_layernorm.bias": "pytorch_model-00006-of-00006.bin",
432
+ "gpt_neox.layers.38.input_layernorm.weight": "pytorch_model-00006-of-00006.bin",
433
+ "gpt_neox.layers.38.mlp.dense_4h_to_h.bias": "pytorch_model-00006-of-00006.bin",
434
+ "gpt_neox.layers.38.mlp.dense_4h_to_h.weight": "pytorch_model-00006-of-00006.bin",
435
+ "gpt_neox.layers.38.mlp.dense_h_to_4h.bias": "pytorch_model-00006-of-00006.bin",
436
+ "gpt_neox.layers.38.mlp.dense_h_to_4h.weight": "pytorch_model-00006-of-00006.bin",
437
+ "gpt_neox.layers.38.post_attention_layernorm.bias": "pytorch_model-00006-of-00006.bin",
438
+ "gpt_neox.layers.38.post_attention_layernorm.weight": "pytorch_model-00006-of-00006.bin",
439
+ "gpt_neox.layers.39.attention.dense.bias": "pytorch_model-00006-of-00006.bin",
440
+ "gpt_neox.layers.39.attention.dense.weight": "pytorch_model-00006-of-00006.bin",
441
+ "gpt_neox.layers.39.attention.query_key_value.bias": "pytorch_model-00006-of-00006.bin",
442
+ "gpt_neox.layers.39.attention.query_key_value.weight": "pytorch_model-00006-of-00006.bin",
443
+ "gpt_neox.layers.39.attention.rotary_emb.inv_freq": "pytorch_model-00006-of-00006.bin",
444
+ "gpt_neox.layers.39.input_layernorm.bias": "pytorch_model-00006-of-00006.bin",
445
+ "gpt_neox.layers.39.input_layernorm.weight": "pytorch_model-00006-of-00006.bin",
446
+ "gpt_neox.layers.39.mlp.dense_4h_to_h.bias": "pytorch_model-00006-of-00006.bin",
447
+ "gpt_neox.layers.39.mlp.dense_4h_to_h.weight": "pytorch_model-00006-of-00006.bin",
448
+ "gpt_neox.layers.39.mlp.dense_h_to_4h.bias": "pytorch_model-00006-of-00006.bin",
449
+ "gpt_neox.layers.39.mlp.dense_h_to_4h.weight": "pytorch_model-00006-of-00006.bin",
450
+ "gpt_neox.layers.39.post_attention_layernorm.bias": "pytorch_model-00006-of-00006.bin",
451
+ "gpt_neox.layers.39.post_attention_layernorm.weight": "pytorch_model-00006-of-00006.bin",
452
+ "gpt_neox.layers.4.attention.dense.bias": "pytorch_model-00001-of-00006.bin",
453
+ "gpt_neox.layers.4.attention.dense.weight": "pytorch_model-00001-of-00006.bin",
454
+ "gpt_neox.layers.4.attention.query_key_value.bias": "pytorch_model-00001-of-00006.bin",
455
+ "gpt_neox.layers.4.attention.query_key_value.weight": "pytorch_model-00001-of-00006.bin",
456
+ "gpt_neox.layers.4.attention.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
457
+ "gpt_neox.layers.4.input_layernorm.bias": "pytorch_model-00001-of-00006.bin",
458
+ "gpt_neox.layers.4.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
459
+ "gpt_neox.layers.4.mlp.dense_4h_to_h.bias": "pytorch_model-00001-of-00006.bin",
460
+ "gpt_neox.layers.4.mlp.dense_4h_to_h.weight": "pytorch_model-00001-of-00006.bin",
461
+ "gpt_neox.layers.4.mlp.dense_h_to_4h.bias": "pytorch_model-00001-of-00006.bin",
462
+ "gpt_neox.layers.4.mlp.dense_h_to_4h.weight": "pytorch_model-00001-of-00006.bin",
463
+ "gpt_neox.layers.4.post_attention_layernorm.bias": "pytorch_model-00001-of-00006.bin",
464
+ "gpt_neox.layers.4.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
465
+ "gpt_neox.layers.5.attention.dense.bias": "pytorch_model-00001-of-00006.bin",
466
+ "gpt_neox.layers.5.attention.dense.weight": "pytorch_model-00001-of-00006.bin",
467
+ "gpt_neox.layers.5.attention.query_key_value.bias": "pytorch_model-00001-of-00006.bin",
468
+ "gpt_neox.layers.5.attention.query_key_value.weight": "pytorch_model-00001-of-00006.bin",
469
+ "gpt_neox.layers.5.attention.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
470
+ "gpt_neox.layers.5.input_layernorm.bias": "pytorch_model-00001-of-00006.bin",
471
+ "gpt_neox.layers.5.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
472
+ "gpt_neox.layers.5.mlp.dense_4h_to_h.bias": "pytorch_model-00001-of-00006.bin",
473
+ "gpt_neox.layers.5.mlp.dense_4h_to_h.weight": "pytorch_model-00001-of-00006.bin",
474
+ "gpt_neox.layers.5.mlp.dense_h_to_4h.bias": "pytorch_model-00001-of-00006.bin",
475
+ "gpt_neox.layers.5.mlp.dense_h_to_4h.weight": "pytorch_model-00001-of-00006.bin",
476
+ "gpt_neox.layers.5.post_attention_layernorm.bias": "pytorch_model-00001-of-00006.bin",
477
+ "gpt_neox.layers.5.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
478
+ "gpt_neox.layers.6.attention.dense.bias": "pytorch_model-00002-of-00006.bin",
479
+ "gpt_neox.layers.6.attention.dense.weight": "pytorch_model-00002-of-00006.bin",
480
+ "gpt_neox.layers.6.attention.query_key_value.bias": "pytorch_model-00001-of-00006.bin",
481
+ "gpt_neox.layers.6.attention.query_key_value.weight": "pytorch_model-00001-of-00006.bin",
482
+ "gpt_neox.layers.6.attention.rotary_emb.inv_freq": "pytorch_model-00001-of-00006.bin",
483
+ "gpt_neox.layers.6.input_layernorm.bias": "pytorch_model-00001-of-00006.bin",
484
+ "gpt_neox.layers.6.input_layernorm.weight": "pytorch_model-00001-of-00006.bin",
485
+ "gpt_neox.layers.6.mlp.dense_4h_to_h.bias": "pytorch_model-00002-of-00006.bin",
486
+ "gpt_neox.layers.6.mlp.dense_4h_to_h.weight": "pytorch_model-00002-of-00006.bin",
487
+ "gpt_neox.layers.6.mlp.dense_h_to_4h.bias": "pytorch_model-00002-of-00006.bin",
488
+ "gpt_neox.layers.6.mlp.dense_h_to_4h.weight": "pytorch_model-00002-of-00006.bin",
489
+ "gpt_neox.layers.6.post_attention_layernorm.bias": "pytorch_model-00001-of-00006.bin",
490
+ "gpt_neox.layers.6.post_attention_layernorm.weight": "pytorch_model-00001-of-00006.bin",
491
+ "gpt_neox.layers.7.attention.dense.bias": "pytorch_model-00002-of-00006.bin",
492
+ "gpt_neox.layers.7.attention.dense.weight": "pytorch_model-00002-of-00006.bin",
493
+ "gpt_neox.layers.7.attention.query_key_value.bias": "pytorch_model-00002-of-00006.bin",
494
+ "gpt_neox.layers.7.attention.query_key_value.weight": "pytorch_model-00002-of-00006.bin",
495
+ "gpt_neox.layers.7.attention.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
496
+ "gpt_neox.layers.7.input_layernorm.bias": "pytorch_model-00002-of-00006.bin",
497
+ "gpt_neox.layers.7.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
498
+ "gpt_neox.layers.7.mlp.dense_4h_to_h.bias": "pytorch_model-00002-of-00006.bin",
499
+ "gpt_neox.layers.7.mlp.dense_4h_to_h.weight": "pytorch_model-00002-of-00006.bin",
500
+ "gpt_neox.layers.7.mlp.dense_h_to_4h.bias": "pytorch_model-00002-of-00006.bin",
501
+ "gpt_neox.layers.7.mlp.dense_h_to_4h.weight": "pytorch_model-00002-of-00006.bin",
502
+ "gpt_neox.layers.7.post_attention_layernorm.bias": "pytorch_model-00002-of-00006.bin",
503
+ "gpt_neox.layers.7.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
504
+ "gpt_neox.layers.8.attention.dense.bias": "pytorch_model-00002-of-00006.bin",
505
+ "gpt_neox.layers.8.attention.dense.weight": "pytorch_model-00002-of-00006.bin",
506
+ "gpt_neox.layers.8.attention.query_key_value.bias": "pytorch_model-00002-of-00006.bin",
507
+ "gpt_neox.layers.8.attention.query_key_value.weight": "pytorch_model-00002-of-00006.bin",
508
+ "gpt_neox.layers.8.attention.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
509
+ "gpt_neox.layers.8.input_layernorm.bias": "pytorch_model-00002-of-00006.bin",
510
+ "gpt_neox.layers.8.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
511
+ "gpt_neox.layers.8.mlp.dense_4h_to_h.bias": "pytorch_model-00002-of-00006.bin",
512
+ "gpt_neox.layers.8.mlp.dense_4h_to_h.weight": "pytorch_model-00002-of-00006.bin",
513
+ "gpt_neox.layers.8.mlp.dense_h_to_4h.bias": "pytorch_model-00002-of-00006.bin",
514
+ "gpt_neox.layers.8.mlp.dense_h_to_4h.weight": "pytorch_model-00002-of-00006.bin",
515
+ "gpt_neox.layers.8.post_attention_layernorm.bias": "pytorch_model-00002-of-00006.bin",
516
+ "gpt_neox.layers.8.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin",
517
+ "gpt_neox.layers.9.attention.dense.bias": "pytorch_model-00002-of-00006.bin",
518
+ "gpt_neox.layers.9.attention.dense.weight": "pytorch_model-00002-of-00006.bin",
519
+ "gpt_neox.layers.9.attention.query_key_value.bias": "pytorch_model-00002-of-00006.bin",
520
+ "gpt_neox.layers.9.attention.query_key_value.weight": "pytorch_model-00002-of-00006.bin",
521
+ "gpt_neox.layers.9.attention.rotary_emb.inv_freq": "pytorch_model-00002-of-00006.bin",
522
+ "gpt_neox.layers.9.input_layernorm.bias": "pytorch_model-00002-of-00006.bin",
523
+ "gpt_neox.layers.9.input_layernorm.weight": "pytorch_model-00002-of-00006.bin",
524
+ "gpt_neox.layers.9.mlp.dense_4h_to_h.bias": "pytorch_model-00002-of-00006.bin",
525
+ "gpt_neox.layers.9.mlp.dense_4h_to_h.weight": "pytorch_model-00002-of-00006.bin",
526
+ "gpt_neox.layers.9.mlp.dense_h_to_4h.bias": "pytorch_model-00002-of-00006.bin",
527
+ "gpt_neox.layers.9.mlp.dense_h_to_4h.weight": "pytorch_model-00002-of-00006.bin",
528
+ "gpt_neox.layers.9.post_attention_layernorm.bias": "pytorch_model-00002-of-00006.bin",
529
+ "gpt_neox.layers.9.post_attention_layernorm.weight": "pytorch_model-00002-of-00006.bin"
530
+ }
531
+ }
special_tokens_map.json CHANGED
@@ -1,3 +1,6 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:8bcf4e58d0970bbcb3b111a4c1e297c082fbf61253b14b3c0c8f2972ccb61ba4
3
- size 131
 
 
 
 
1
+ {
2
+ "bos_token": "<|endoftext|>",
3
+ "eos_token": "<|endoftext|>",
4
+ "pad_token": "<|endoftext|>",
5
+ "unk_token": "<|endoftext|>"
6
+ }
tokenizer_config.json CHANGED
@@ -1,3 +1,6 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:aa93de105cdad2fe3162a41a6e0b8f188eed2f4f0ab855c0f0e0abcfe97b2ec8
3
- size 146
 
 
 
 
1
+ {
2
+ "clean_up_tokenization_spaces": true,
3
+ "model_max_length": 1024,
4
+ "padding_side": "right",
5
+ "tokenizer_class": "PreTrainedTokenizerFast"
6
+ }
trainer_state.json CHANGED
The diff for this file is too large to render. See raw diff
 
zero_to_fp32.py CHANGED
@@ -1,3 +1,578 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:061bdb6df821b30589945691c45ddf5731de7baf150f76b0555f850187e4062f
3
- size 23610
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python
2
+
3
+ # Copyright (c) Microsoft Corporation.
4
+ # SPDX-License-Identifier: Apache-2.0
5
+
6
+ # DeepSpeed Team
7
+
8
+ # This script extracts fp32 consolidated weights from a zero 1, 2 and 3 DeepSpeed checkpoints. It gets
9
+ # copied into the top level checkpoint dir, so the user can easily do the conversion at any point in
10
+ # the future. Once extracted, the weights don't require DeepSpeed and can be used in any
11
+ # application.
12
+ #
13
+ # example: python zero_to_fp32.py . pytorch_model.bin
14
+
15
+ import argparse
16
+ import torch
17
+ import glob
18
+ import math
19
+ import os
20
+ import re
21
+ from collections import OrderedDict
22
+ from dataclasses import dataclass
23
+
24
+ # while this script doesn't use deepspeed to recover data, since the checkpoints are pickled with
25
+ # DeepSpeed data structures it has to be available in the current python environment.
26
+ from deepspeed.utils import logger
27
+ from deepspeed.checkpoint.constants import (DS_VERSION, OPTIMIZER_STATE_DICT, SINGLE_PARTITION_OF_FP32_GROUPS,
28
+ FP32_FLAT_GROUPS, ZERO_STAGE, PARTITION_COUNT, PARAM_SHAPES, BUFFER_NAMES,
29
+ FROZEN_PARAM_SHAPES, FROZEN_PARAM_FRAGMENTS)
30
+
31
+
32
+ @dataclass
33
+ class zero_model_state:
34
+ buffers: dict()
35
+ param_shapes: dict()
36
+ shared_params: list
37
+ ds_version: int
38
+ frozen_param_shapes: dict()
39
+ frozen_param_fragments: dict()
40
+
41
+
42
+ debug = 0
43
+
44
+ # load to cpu
45
+ device = torch.device('cpu')
46
+
47
+
48
+ def atoi(text):
49
+ return int(text) if text.isdigit() else text
50
+
51
+
52
+ def natural_keys(text):
53
+ '''
54
+ alist.sort(key=natural_keys) sorts in human order
55
+ http://nedbatchelder.com/blog/200712/human_sorting.html
56
+ (See Toothy's implementation in the comments)
57
+ '''
58
+ return [atoi(c) for c in re.split(r'(\d+)', text)]
59
+
60
+
61
+ def get_model_state_file(checkpoint_dir, zero_stage):
62
+ if not os.path.isdir(checkpoint_dir):
63
+ raise FileNotFoundError(f"Directory '{checkpoint_dir}' doesn't exist")
64
+
65
+ # there should be only one file
66
+ if zero_stage <= 2:
67
+ file = os.path.join(checkpoint_dir, "mp_rank_00_model_states.pt")
68
+ elif zero_stage == 3:
69
+ file = os.path.join(checkpoint_dir, "zero_pp_rank_0_mp_rank_00_model_states.pt")
70
+
71
+ if not os.path.exists(file):
72
+ raise FileNotFoundError(f"can't find model states file at '{file}'")
73
+
74
+ return file
75
+
76
+
77
+ def get_checkpoint_files(checkpoint_dir, glob_pattern):
78
+ # XXX: need to test that this simple glob rule works for multi-node setup too
79
+ ckpt_files = sorted(glob.glob(os.path.join(checkpoint_dir, glob_pattern)), key=natural_keys)
80
+
81
+ if len(ckpt_files) == 0:
82
+ raise FileNotFoundError(f"can't find {glob_pattern} files in directory '{checkpoint_dir}'")
83
+
84
+ return ckpt_files
85
+
86
+
87
+ def get_optim_files(checkpoint_dir):
88
+ return get_checkpoint_files(checkpoint_dir, "*_optim_states.pt")
89
+
90
+
91
+ def get_model_state_files(checkpoint_dir):
92
+ return get_checkpoint_files(checkpoint_dir, "*_model_states.pt")
93
+
94
+
95
+ def parse_model_states(files):
96
+ zero_model_states = []
97
+ for file in files:
98
+ state_dict = torch.load(file, map_location=device)
99
+
100
+ if BUFFER_NAMES not in state_dict:
101
+ raise ValueError(f"{file} is not a model state checkpoint")
102
+ buffer_names = state_dict[BUFFER_NAMES]
103
+ if debug:
104
+ print("Found buffers:", buffer_names)
105
+
106
+ # recover just the buffers while restoring them to fp32 if they were saved in fp16
107
+ buffers = {k: v.float() for k, v in state_dict["module"].items() if k in buffer_names}
108
+ param_shapes = state_dict[PARAM_SHAPES]
109
+
110
+ # collect parameters that are included in param_shapes
111
+ param_names = []
112
+ for s in param_shapes:
113
+ for name in s.keys():
114
+ param_names.append(name)
115
+
116
+ # update with frozen parameters
117
+ frozen_param_shapes = state_dict.get(FROZEN_PARAM_SHAPES, None)
118
+ if frozen_param_shapes is not None:
119
+ if debug:
120
+ print(f"Found frozen_param_shapes: {frozen_param_shapes}")
121
+ param_names += list(frozen_param_shapes.keys())
122
+
123
+ # handle shared params
124
+ shared_params = [[k, v] for k, v in state_dict["shared_params"].items()]
125
+
126
+ ds_version = state_dict.get(DS_VERSION, None)
127
+
128
+ frozen_param_fragments = state_dict.get(FROZEN_PARAM_FRAGMENTS, None)
129
+
130
+ z_model_state = zero_model_state(buffers=buffers,
131
+ param_shapes=param_shapes,
132
+ shared_params=shared_params,
133
+ ds_version=ds_version,
134
+ frozen_param_shapes=frozen_param_shapes,
135
+ frozen_param_fragments=frozen_param_fragments)
136
+ zero_model_states.append(z_model_state)
137
+
138
+ return zero_model_states
139
+
140
+
141
+ def parse_optim_states(files, ds_checkpoint_dir):
142
+
143
+ total_files = len(files)
144
+ state_dicts = []
145
+ for f in files:
146
+ state_dicts.append(torch.load(f, map_location=device))
147
+
148
+ if not ZERO_STAGE in state_dicts[0][OPTIMIZER_STATE_DICT]:
149
+ raise ValueError(f"{files[0]} is not a zero checkpoint")
150
+ zero_stage = state_dicts[0][OPTIMIZER_STATE_DICT][ZERO_STAGE]
151
+ world_size = state_dicts[0][OPTIMIZER_STATE_DICT][PARTITION_COUNT]
152
+
153
+ # For ZeRO-2 each param group can have different partition_count as data parallelism for expert
154
+ # parameters can be different from data parallelism for non-expert parameters. So we can just
155
+ # use the max of the partition_count to get the dp world_size.
156
+
157
+ if type(world_size) is list:
158
+ world_size = max(world_size)
159
+
160
+ if world_size != total_files:
161
+ raise ValueError(
162
+ f"Expected {world_size} of '*_optim_states.pt' under '{ds_checkpoint_dir}' but found {total_files} files. "
163
+ "Possibly due to an overwrite of an old checkpoint, or a checkpoint didn't get saved by one or more processes."
164
+ )
165
+
166
+ # the groups are named differently in each stage
167
+ if zero_stage <= 2:
168
+ fp32_groups_key = SINGLE_PARTITION_OF_FP32_GROUPS
169
+ elif zero_stage == 3:
170
+ fp32_groups_key = FP32_FLAT_GROUPS
171
+ else:
172
+ raise ValueError(f"unknown zero stage {zero_stage}")
173
+
174
+ if zero_stage <= 2:
175
+ fp32_flat_groups = [state_dicts[i][OPTIMIZER_STATE_DICT][fp32_groups_key] for i in range(len(state_dicts))]
176
+ elif zero_stage == 3:
177
+ # if there is more than one param group, there will be multiple flattened tensors - one
178
+ # flattened tensor per group - for simplicity merge them into a single tensor
179
+ #
180
+ # XXX: could make the script more memory efficient for when there are multiple groups - it
181
+ # will require matching the sub-lists of param_shapes for each param group flattened tensor
182
+
183
+ fp32_flat_groups = [
184
+ torch.cat(state_dicts[i][OPTIMIZER_STATE_DICT][fp32_groups_key], 0) for i in range(len(state_dicts))
185
+ ]
186
+
187
+ return zero_stage, world_size, fp32_flat_groups
188
+
189
+
190
+ def _get_fp32_state_dict_from_zero_checkpoint(ds_checkpoint_dir):
191
+ """
192
+ Returns fp32 state_dict reconstructed from ds checkpoint
193
+
194
+ Args:
195
+ - ``ds_checkpoint_dir``: path to the deepspeed checkpoint folder (where the optimizer files are)
196
+
197
+ """
198
+ print(f"Processing zero checkpoint '{ds_checkpoint_dir}'")
199
+
200
+ optim_files = get_optim_files(ds_checkpoint_dir)
201
+ zero_stage, world_size, fp32_flat_groups = parse_optim_states(optim_files, ds_checkpoint_dir)
202
+ print(f"Detected checkpoint of type zero stage {zero_stage}, world_size: {world_size}")
203
+
204
+ model_files = get_model_state_files(ds_checkpoint_dir)
205
+
206
+ zero_model_states = parse_model_states(model_files)
207
+ print(f'Parsing checkpoint created by deepspeed=={zero_model_states[0].ds_version}')
208
+
209
+ if zero_stage <= 2:
210
+ return _get_fp32_state_dict_from_zero2_checkpoint(world_size, fp32_flat_groups, zero_model_states)
211
+ elif zero_stage == 3:
212
+ return _get_fp32_state_dict_from_zero3_checkpoint(world_size, fp32_flat_groups, zero_model_states)
213
+
214
+
215
+ def _zero2_merge_frozen_params(state_dict, zero_model_states):
216
+ if zero_model_states[0].frozen_param_shapes is None or len(zero_model_states[0].frozen_param_shapes) == 0:
217
+ return
218
+
219
+ frozen_param_shapes = zero_model_states[0].frozen_param_shapes
220
+ frozen_param_fragments = zero_model_states[0].frozen_param_fragments
221
+
222
+ if debug:
223
+ num_elem = sum(s.numel() for s in frozen_param_shapes.values())
224
+ print(f'rank 0: {FROZEN_PARAM_SHAPES}.numel = {num_elem}')
225
+
226
+ wanted_params = len(frozen_param_shapes)
227
+ wanted_numel = sum(s.numel() for s in frozen_param_shapes.values())
228
+ avail_numel = sum([p.numel() for p in frozen_param_fragments.values()])
229
+ print(f'Frozen params: Have {avail_numel} numels to process.')
230
+ print(f'Frozen params: Need {wanted_numel} numels in {wanted_params} params')
231
+
232
+ total_params = 0
233
+ total_numel = 0
234
+ for name, shape in frozen_param_shapes.items():
235
+ total_params += 1
236
+ unpartitioned_numel = shape.numel()
237
+ total_numel += unpartitioned_numel
238
+
239
+ state_dict[name] = frozen_param_fragments[name]
240
+
241
+ if debug:
242
+ print(f"{name} full shape: {shape} unpartitioned numel {unpartitioned_numel} ")
243
+
244
+ print(f"Reconstructed Frozen fp32 state dict with {total_params} params {total_numel} elements")
245
+
246
+
247
+ def _zero2_merge_trainable_params(state_dict, world_size, fp32_flat_groups, zero_model_states):
248
+ param_shapes = zero_model_states[0].param_shapes
249
+
250
+ # Reconstruction protocol:
251
+ #
252
+ # XXX: document this
253
+
254
+ if debug:
255
+ for i in range(world_size):
256
+ for j in range(len(fp32_flat_groups[0])):
257
+ print(f"{FP32_FLAT_GROUPS}[{i}][{j}].shape={fp32_flat_groups[i][j].shape}")
258
+
259
+ # XXX: memory usage doubles here (zero2)
260
+ num_param_groups = len(fp32_flat_groups[0])
261
+ merged_single_partition_of_fp32_groups = []
262
+ for i in range(num_param_groups):
263
+ merged_partitions = [sd[i] for sd in fp32_flat_groups]
264
+ full_single_fp32_vector = torch.cat(merged_partitions, 0)
265
+ merged_single_partition_of_fp32_groups.append(full_single_fp32_vector)
266
+ avail_numel = sum(
267
+ [full_single_fp32_vector.numel() for full_single_fp32_vector in merged_single_partition_of_fp32_groups])
268
+
269
+ if debug:
270
+ wanted_params = sum([len(shapes) for shapes in param_shapes])
271
+ wanted_numel = sum([sum(shape.numel() for shape in shapes.values()) for shapes in param_shapes])
272
+ # not asserting if there is a mismatch due to possible padding
273
+ print(f"Have {avail_numel} numels to process.")
274
+ print(f"Need {wanted_numel} numels in {wanted_params} params.")
275
+
276
+ # params
277
+ # XXX: for huge models that can't fit into the host's RAM we will have to recode this to support
278
+ # out-of-core computing solution
279
+ total_numel = 0
280
+ total_params = 0
281
+ for shapes, full_single_fp32_vector in zip(param_shapes, merged_single_partition_of_fp32_groups):
282
+ offset = 0
283
+ avail_numel = full_single_fp32_vector.numel()
284
+ for name, shape in shapes.items():
285
+
286
+ unpartitioned_numel = shape.numel()
287
+ total_numel += unpartitioned_numel
288
+ total_params += 1
289
+
290
+ if debug:
291
+ print(f"{name} full shape: {shape} unpartitioned numel {unpartitioned_numel} ")
292
+ state_dict[name] = full_single_fp32_vector.narrow(0, offset, unpartitioned_numel).view(shape)
293
+ offset += unpartitioned_numel
294
+
295
+ # Z2 started to align to 2*world_size to improve nccl performance. Therefore both offset and
296
+ # avail_numel can differ by anywhere between 0..2*world_size. Due to two unrelated complex
297
+ # paddings performed in the code it's almost impossible to predict the exact numbers w/o the
298
+ # live optimizer object, so we are checking that the numbers are within the right range
299
+ align_to = 2 * world_size
300
+
301
+ def zero2_align(x):
302
+ return align_to * math.ceil(x / align_to)
303
+
304
+ if debug:
305
+ print(f"original offset={offset}, avail_numel={avail_numel}")
306
+
307
+ offset = zero2_align(offset)
308
+ avail_numel = zero2_align(avail_numel)
309
+
310
+ if debug:
311
+ print(f"aligned offset={offset}, avail_numel={avail_numel}")
312
+
313
+ # Sanity check
314
+ if offset != avail_numel:
315
+ raise ValueError(f"consumed {offset} numels out of {avail_numel} - something is wrong")
316
+
317
+ print(f"Reconstructed fp32 state dict with {total_params} params {total_numel} elements")
318
+
319
+
320
+ def _get_fp32_state_dict_from_zero2_checkpoint(world_size, fp32_flat_groups, zero_model_states):
321
+ state_dict = OrderedDict()
322
+
323
+ # buffers
324
+ buffers = zero_model_states[0].buffers
325
+ state_dict.update(buffers)
326
+ if debug:
327
+ print(f"added {len(buffers)} buffers")
328
+
329
+ _zero2_merge_frozen_params(state_dict, zero_model_states)
330
+
331
+ _zero2_merge_trainable_params(state_dict, world_size, fp32_flat_groups, zero_model_states)
332
+
333
+ # recover shared parameters
334
+ for pair in zero_model_states[0].shared_params:
335
+ if pair[1] in state_dict:
336
+ state_dict[pair[0]] = state_dict[pair[1]]
337
+
338
+ return state_dict
339
+
340
+
341
+ def zero3_partitioned_param_info(unpartitioned_numel, world_size):
342
+ remainder = unpartitioned_numel % world_size
343
+ padding_numel = (world_size - remainder) if remainder else 0
344
+ partitioned_numel = math.ceil(unpartitioned_numel / world_size)
345
+ return partitioned_numel, padding_numel
346
+
347
+
348
+ def _zero3_merge_frozen_params(state_dict, world_size, zero_model_states):
349
+ if zero_model_states[0].frozen_param_shapes is None or len(zero_model_states[0].frozen_param_shapes) == 0:
350
+ return
351
+
352
+ if debug:
353
+ for i in range(world_size):
354
+ num_elem = sum(s.numel() for s in zero_model_states[i].frozen_param_fragments.values())
355
+ print(f'rank {i}: {FROZEN_PARAM_SHAPES}.numel = {num_elem}')
356
+
357
+ frozen_param_shapes = zero_model_states[0].frozen_param_shapes
358
+ wanted_params = len(frozen_param_shapes)
359
+ wanted_numel = sum(s.numel() for s in frozen_param_shapes.values())
360
+ avail_numel = sum([p.numel() for p in zero_model_states[0].frozen_param_fragments.values()]) * world_size
361
+ print(f'Frozen params: Have {avail_numel} numels to process.')
362
+ print(f'Frozen params: Need {wanted_numel} numels in {wanted_params} params')
363
+
364
+ total_params = 0
365
+ total_numel = 0
366
+ for name, shape in zero_model_states[0].frozen_param_shapes.items():
367
+ total_params += 1
368
+ unpartitioned_numel = shape.numel()
369
+ total_numel += unpartitioned_numel
370
+
371
+ param_frags = tuple(model_state.frozen_param_fragments[name] for model_state in zero_model_states)
372
+ state_dict[name] = torch.cat(param_frags, 0).narrow(0, 0, unpartitioned_numel).view(shape)
373
+
374
+ partitioned_numel, partitioned_padding_numel = zero3_partitioned_param_info(unpartitioned_numel, world_size)
375
+
376
+ if debug:
377
+ print(
378
+ f"Frozen params: {total_params} {name} full shape: {shape} partition0 numel={partitioned_numel} partitioned_padding_numel={partitioned_padding_numel}"
379
+ )
380
+
381
+ print(f"Reconstructed Frozen fp32 state dict with {total_params} params {total_numel} elements")
382
+
383
+
384
+ def _zero3_merge_trainable_params(state_dict, world_size, fp32_flat_groups, zero_model_states):
385
+ param_shapes = zero_model_states[0].param_shapes
386
+ avail_numel = fp32_flat_groups[0].numel() * world_size
387
+ # Reconstruction protocol: For zero3 we need to zip the partitions together at boundary of each
388
+ # param, re-consolidating each param, while dealing with padding if any
389
+
390
+ # merge list of dicts, preserving order
391
+ param_shapes = {k: v for d in param_shapes for k, v in d.items()}
392
+
393
+ if debug:
394
+ for i in range(world_size):
395
+ print(f"{FP32_FLAT_GROUPS}[{i}].shape={fp32_flat_groups[i].shape}")
396
+
397
+ wanted_params = len(param_shapes)
398
+ wanted_numel = sum(shape.numel() for shape in param_shapes.values())
399
+ # not asserting if there is a mismatch due to possible padding
400
+ avail_numel = fp32_flat_groups[0].numel() * world_size
401
+ print(f"Trainable params: Have {avail_numel} numels to process.")
402
+ print(f"Trainable params: Need {wanted_numel} numels in {wanted_params} params.")
403
+
404
+ # params
405
+ # XXX: for huge models that can't fit into the host's RAM we will have to recode this to support
406
+ # out-of-core computing solution
407
+ offset = 0
408
+ total_numel = 0
409
+ total_params = 0
410
+ for name, shape in param_shapes.items():
411
+
412
+ unpartitioned_numel = shape.numel()
413
+ total_numel += unpartitioned_numel
414
+ total_params += 1
415
+
416
+ partitioned_numel, partitioned_padding_numel = zero3_partitioned_param_info(unpartitioned_numel, world_size)
417
+
418
+ if debug:
419
+ print(
420
+ f"Trainable params: {total_params} {name} full shape: {shape} partition0 numel={partitioned_numel} partitioned_padding_numel={partitioned_padding_numel}"
421
+ )
422
+
423
+ # XXX: memory usage doubles here
424
+ state_dict[name] = torch.cat(
425
+ tuple(fp32_flat_groups[i].narrow(0, offset, partitioned_numel) for i in range(world_size)),
426
+ 0).narrow(0, 0, unpartitioned_numel).view(shape)
427
+ offset += partitioned_numel
428
+
429
+ offset *= world_size
430
+
431
+ # Sanity check
432
+ if offset != avail_numel:
433
+ raise ValueError(f"consumed {offset} numels out of {avail_numel} - something is wrong")
434
+
435
+ print(f"Reconstructed Trainable fp32 state dict with {total_params} params {total_numel} elements")
436
+
437
+
438
+ def _get_fp32_state_dict_from_zero3_checkpoint(world_size, fp32_flat_groups, zero_model_states):
439
+ state_dict = OrderedDict()
440
+
441
+ # buffers
442
+ buffers = zero_model_states[0].buffers
443
+ state_dict.update(buffers)
444
+ if debug:
445
+ print(f"added {len(buffers)} buffers")
446
+
447
+ _zero3_merge_frozen_params(state_dict, world_size, zero_model_states)
448
+
449
+ _zero3_merge_trainable_params(state_dict, world_size, fp32_flat_groups, zero_model_states)
450
+
451
+ # recover shared parameters
452
+ for pair in zero_model_states[0].shared_params:
453
+ if pair[1] in state_dict:
454
+ state_dict[pair[0]] = state_dict[pair[1]]
455
+
456
+ return state_dict
457
+
458
+
459
+ def get_fp32_state_dict_from_zero_checkpoint(checkpoint_dir, tag=None):
460
+ """
461
+ Convert ZeRO 2 or 3 checkpoint into a single fp32 consolidated state_dict that can be loaded with
462
+ ``load_state_dict()`` and used for training without DeepSpeed or shared with others, for example
463
+ via a model hub.
464
+
465
+ Args:
466
+ - ``checkpoint_dir``: path to the desired checkpoint folder
467
+ - ``tag``: checkpoint tag used as a unique identifier for checkpoint. If not provided will attempt to load tag in 'latest' file. e.g., ``global_step14``
468
+
469
+ Returns:
470
+ - pytorch ``state_dict``
471
+
472
+ Note: this approach may not work if your application doesn't have sufficient free CPU memory and
473
+ you may need to use the offline approach using the ``zero_to_fp32.py`` script that is saved with
474
+ the checkpoint.
475
+
476
+ A typical usage might be ::
477
+
478
+ from deepspeed.utils.zero_to_fp32 import get_fp32_state_dict_from_zero_checkpoint
479
+ # do the training and checkpoint saving
480
+ state_dict = get_fp32_state_dict_from_zero_checkpoint(checkpoint_dir) # already on cpu
481
+ model = model.cpu() # move to cpu
482
+ model.load_state_dict(state_dict)
483
+ # submit to model hub or save the model to share with others
484
+
485
+ In this example the ``model`` will no longer be usable in the deepspeed context of the same
486
+ application. i.e. you will need to re-initialize the deepspeed engine, since
487
+ ``model.load_state_dict(state_dict)`` will remove all the deepspeed magic from it.
488
+
489
+ If you want it all done for you, use ``load_state_dict_from_zero_checkpoint`` instead.
490
+
491
+ """
492
+ if tag is None:
493
+ latest_path = os.path.join(checkpoint_dir, 'latest')
494
+ if os.path.isfile(latest_path):
495
+ with open(latest_path, 'r') as fd:
496
+ tag = fd.read().strip()
497
+ else:
498
+ raise ValueError(f"Unable to find 'latest' file at {latest_path}")
499
+
500
+ ds_checkpoint_dir = os.path.join(checkpoint_dir, tag)
501
+
502
+ if not os.path.isdir(ds_checkpoint_dir):
503
+ raise FileNotFoundError(f"Directory '{ds_checkpoint_dir}' doesn't exist")
504
+
505
+ return _get_fp32_state_dict_from_zero_checkpoint(ds_checkpoint_dir)
506
+
507
+
508
+ def convert_zero_checkpoint_to_fp32_state_dict(checkpoint_dir, output_file, tag=None):
509
+ """
510
+ Convert ZeRO 2 or 3 checkpoint into a single fp32 consolidated ``state_dict`` file that can be
511
+ loaded with ``torch.load(file)`` + ``load_state_dict()`` and used for training without DeepSpeed.
512
+
513
+ Args:
514
+ - ``checkpoint_dir``: path to the desired checkpoint folder. (one that contains the tag-folder, like ``global_step14``)
515
+ - ``output_file``: path to the pytorch fp32 state_dict output file (e.g. path/pytorch_model.bin)
516
+ - ``tag``: checkpoint tag used as a unique identifier for checkpoint. If not provided will attempt to load tag in the file named ``latest`` in the checkpoint folder, e.g., ``global_step14``
517
+ """
518
+
519
+ state_dict = get_fp32_state_dict_from_zero_checkpoint(checkpoint_dir, tag)
520
+ print(f"Saving fp32 state dict to {output_file}")
521
+ torch.save(state_dict, output_file)
522
+
523
+
524
+ def load_state_dict_from_zero_checkpoint(model, checkpoint_dir, tag=None):
525
+ """
526
+ 1. Put the provided model to cpu
527
+ 2. Convert ZeRO 2 or 3 checkpoint into a single fp32 consolidated ``state_dict``
528
+ 3. Load it into the provided model
529
+
530
+ Args:
531
+ - ``model``: the model object to update
532
+ - ``checkpoint_dir``: path to the desired checkpoint folder. (one that contains the tag-folder, like ``global_step14``)
533
+ - ``tag``: checkpoint tag used as a unique identifier for checkpoint. If not provided will attempt to load tag in the file named ``latest`` in the checkpoint folder, e.g., ``global_step14``
534
+
535
+ Returns:
536
+ - ``model`: modified model
537
+
538
+ Make sure you have plenty of CPU memory available before you call this function. If you don't
539
+ have enough use the ``zero_to_fp32.py`` utility to do the conversion. You will find it
540
+ conveniently placed for you in the checkpoint folder.
541
+
542
+ A typical usage might be ::
543
+
544
+ from deepspeed.utils.zero_to_fp32 import load_state_dict_from_zero_checkpoint
545
+ model = load_state_dict_from_zero_checkpoint(trainer.model, checkpoint_dir)
546
+ # submit to model hub or save the model to share with others
547
+
548
+ Note, that once this was run, the ``model`` will no longer be usable in the deepspeed context
549
+ of the same application. i.e. you will need to re-initialize the deepspeed engine, since
550
+ ``model.load_state_dict(state_dict)`` will remove all the deepspeed magic from it.
551
+
552
+ """
553
+ logger.info(f"Extracting fp32 weights")
554
+ state_dict = get_fp32_state_dict_from_zero_checkpoint(checkpoint_dir, tag)
555
+
556
+ logger.info(f"Overwriting model with fp32 weights")
557
+ model = model.cpu()
558
+ model.load_state_dict(state_dict, strict=False)
559
+
560
+ return model
561
+
562
+
563
+ if __name__ == "__main__":
564
+
565
+ parser = argparse.ArgumentParser()
566
+ parser.add_argument("checkpoint_dir",
567
+ type=str,
568
+ help="path to the desired checkpoint folder, e.g., path/checkpoint-12")
569
+ parser.add_argument(
570
+ "output_file",
571
+ type=str,
572
+ help="path to the pytorch fp32 state_dict output file (e.g. path/checkpoint-12/pytorch_model.bin)")
573
+ parser.add_argument("-d", "--debug", action='store_true', help="enable debug")
574
+ args = parser.parse_args()
575
+
576
+ debug = args.debug
577
+
578
+ convert_zero_checkpoint_to_fp32_state_dict(args.checkpoint_dir, args.output_file)