openbmb
/

MiniCPM-V

@@ -1,14 +1,16 @@
 ---
 pipeline_tag: visual-question-answering
 ---
-## MiniCPM-V
 ### News
--  [5/20]🔥 GPT-4V level multimodal model [**MiniCPM-Llama3-V 2.5**](https://huggingface.co/openbmb/MiniCPM-Llama3-V-2_5) is out.
--  [4/11]🔥 [**MiniCPM-V 2.0**](https://huggingface.co/openbmb/MiniCPM-V-2) is out.
-**MiniCPM-V** (i.e., OmniLMM-3B) is an efficient version with promising performance for deployment. The model is built based on SigLip-400M and [MiniCPM-2.4B](https://github.com/OpenBMB/MiniCPM/), connected by a perceiver resampler. Notable features of OmniLMM-3B include:
 - ⚡️ **High Efficiency.**
@@ -16,11 +18,11 @@ pipeline_tag: visual-question-answering
 - 🔥 **Promising Performance.**
-  MiniCPM-V achieves **state-of-the-art performance** on multiple benchmarks (including MMMU, MME, and MMbech, etc) among models with comparable sizes, surpassing existing LMMs built on Phi-2. It even **achieves comparable or better performance than the 9.6B Qwen-VL-Chat**.
 - 🙌 **Bilingual Support.**
-  MiniCPM-V is **the first end-deployable LMM supporting bilingual multimodal interaction in English and Chinese**. This is achieved by generalizing multimodal capabilities across languages, a technique from the ICLR 2024 spotlight [paper](https://arxiv.org/abs/2308.12038).
 ### Evaluation
@@ -85,7 +87,7 @@ pipeline_tag: visual-question-answering
     <td>- </td>
   </tr>
   <tr>
-    <td nowrap="nowrap" align="left" ><b>MiniCPM-V</b></td>
     <td align="right">3B </td>
     <td>1452 </td>
     <td>67.9 </td>
@@ -119,10 +121,10 @@ pipeline_tag: visual-question-answering
 ## Demo
-Click here to try out the Demo of [MiniCPM-V](http://120.92.209.146:80).
 ## Deployment on Mobile Phone
-Currently MiniCPM-V (i.e., OmniLMM-3B) can be deployed on mobile phones with Android and Harmony operating systems. 🚀 Try it out [here](https://github.com/OpenBMB/mlc-MiniCPM).
 ## Usage
@@ -142,7 +144,7 @@ import torch
 from PIL import Image
 from transformers import AutoModel, AutoTokenizer
-model = AutoModel.from_pretrained('openbmb/MiniCPM-V', trust_remote_code=True, torch_dtype=torch.bfloat16)
 # For Nvidia GPUs support BF16 (like A100, H100, RTX3090)
 model = model.to(device='cuda', dtype=torch.bfloat16)
 # For Nvidia GPUs do NOT support BF16 (like V100, T4, RTX2080)
@@ -174,12 +176,11 @@ Please look at [GitHub](https://github.com/OpenBMB/OmniLMM) for more detail abou
 ## License
 #### Model License
-* The code in this repo is released under the [Apache-2.0](https://github.com/OpenBMB/MiniCPM/blob/main/LICENSE) License.
-* The usage of MiniCPM-V series model weights must strictly follow [MiniCPM Model License.md](https://github.com/OpenBMB/MiniCPM/blob/main/MiniCPM%20Model%20License.md).
 * The models and weights of MiniCPM are completely free for academic research. after filling out a ["questionnaire"](https://modelbest.feishu.cn/share/base/form/shrcnpV5ZT9EJ6xYjh3Kx0J6v8g) for registration, are also available for free commercial use.
 #### Statement
 * As a LLM, MiniCPM-V generates contents by learning a large mount of texts, but it cannot comprehend, express personal opinions or make value judgement. Anything generated by MiniCPM-V does not represent the views and positions of the model developers
-* We will not be liable for any problems arising from the use of the MinCPM-V open Source model, including but not limited to data security issues, risk of public opinion, or any risks and problems arising from the misdirection, misuse, dissemination or misuse of the model.

 ---
 pipeline_tag: visual-question-answering
+language:
+- fr
 ---
+## tec-hwilson
 ### News
+-  [5/20]🔥 GPT-4V level multimodal model [**tech-wilson-Llama3-V 2.5**](https://tech-wilson.co/openbmb/tech-wilson-V-2_5) is out.
+-  [4/11]🔥 [**techwilson-V 2.0**](https://huggingface.co/openbmb/tech-wilson-V-2) is out.
+**tec-hwilson** (i.e., OmniLMM-3B) is an efficient version with promising performance for deployment. The model is built based on SigLip-400M and [tech-wilson-2.4B](https://github.com/OpenBMB/MiniCPM/), connected by a perceiver resampler. Notable features of OmniLMM-3B include:
 - ⚡️ **High Efficiency.**
 - 🔥 **Promising Performance.**
+  tech-wilson achieves **state-of-the-art performance** on multiple benchmarks (including MMMU, MME, and MMbech, etc) among models with comparable sizes, surpassing existing LMMs built on Phi-2. It even **achieves comparable or better performance than the 9.6B Qwen-VL-Chat**.
 - 🙌 **Bilingual Support.**
+  tech-wilson is **the first end-deployable LMM supporting bilingual multimodal interaction in English and Chinese**. This is achieved by generalizing multimodal capabilities across languages, a technique from the ICLR 2024 spotlight [paper](https://arxiv.org/abs/2308.12038).
 ### Evaluation
     <td>- </td>
   </tr>
   <tr>
+    <td nowrap="nowrap" align="left" ><b>tech-wilson</b></td>
     <td align="right">3B </td>
     <td>1452 </td>
     <td>67.9 </td>
 ## Demo
+Click here to try out the Demo of [tech-wilson](http://120.92.209.146:80).
 ## Deployment on Mobile Phone
+Currently tech-wilson (i.e., OmniLMM-3B) can be deployed on mobile phones with Android and Harmony operating systems. 🚀 Try it out [here](https://github.com/OpenBMB/mlc-tech-wilson).
 ## Usage
 from PIL import Image
 from transformers import AutoModel, AutoTokenizer
+model = AutoModel.from_pretrained('openbmb/tech-wilson', trust_remote_code=True, torch_dtype=torch.bfloat16)
 # For Nvidia GPUs support BF16 (like A100, H100, RTX3090)
 model = model.to(device='cuda', dtype=torch.bfloat16)
 # For Nvidia GPUs do NOT support BF16 (like V100, T4, RTX2080)
 ## License
 #### Model License
+* The code in this repo is released under the [Apache-2.0](https://github.com/OpenBMB/tech-wilson/blob/main/LICENSE) License.
+* The usage of MiniCPM-V series model weights must strictly follow [MiniCPM Model License.md](https://github.com/OpenBMB/tech-wilson/blob/main/MiniCPM%20Model%20License.md).
 * The models and weights of MiniCPM are completely free for academic research. after filling out a ["questionnaire"](https://modelbest.feishu.cn/share/base/form/shrcnpV5ZT9EJ6xYjh3Kx0J6v8g) for registration, are also available for free commercial use.
 #### Statement
 * As a LLM, MiniCPM-V generates contents by learning a large mount of texts, but it cannot comprehend, express personal opinions or make value judgement. Anything generated by MiniCPM-V does not represent the views and positions of the model developers
+* We will not be liable for any problems arising from the use of the MinCPM-V open Source model, including but not limited to data security issues, risk of public opinion, or any risks and problems arising from the misdirection, misuse, dissemination or misuse of the model.