metadata

language:
  - ko
license: apache-2.0
library_name: transformers
tags:
  - text2text-generation
datasets:
  - aihub
metrics:
  - bleu
  - rouge
model-index:
  - name: ko-TextNumbarT
    results:
      - task:
          type: text2text-generation
          name: text2text-generation
        metrics:
          - type: bleu
            value: 0.958234790096092
            name: eval_bleu
            verified: false
          - type: rouge1
            value: 0.9735361877162854
            name: eval_rouge1
            verified: false
          - type: rouge2
            value: 0.9493975212378124
            name: eval_rouge2
            verified: false
          - type: rougeL
            value: 0.9734558938864928
            name: eval_rougeL
            verified: false
          - type: rougeLsum
            value: 0.9734350757552404
            name: eval_rougeLsum
            verified: false

ko-TextNumbarT(TNT Model🧨): Try Korean Reading To Number(한글을 숫자로 바꾸는 모델)

ko-TextNumbarT(TNT Model🧨): Try Korean Reading To Number(한글을 숫자로 바꾸는 모델)

Model Details

Model Description: 뭔가 찾아봐도 모델이나 알고리즘이 딱히 없어서 만들어본 모델입니다.
BartForConditionalGeneration Fine-Tuning Model For Korean To Number
BartForConditionalGeneration으로 파인튜닝한, 한글을 숫자로 변환하는 Task 입니다.
Dataset use Korea aihub
I can't open my fine-tuning datasets for my private issue
데이터셋은 Korea aihub에서 받아서 사용하였으며, 파인튜닝에 사용된 모든 데이터를 사정상 공개해드릴 수는 없습니다.
Korea aihub data is ONLY permit to Korean!!!!!!!
aihub에서 데이터를 받으실 분은 한국인일 것이므로, 한글로만 작성합니다.
정확히는 철자전사를 음성전사로 번역하는 형태로 학습된 모델입니다. (ETRI 전사기준)
In case, ten million, some people use 10 million or some people use 10000000, so this model is crucial for training datasets
천만을 1000만 혹은 10000000으로 쓸 수도 있기에, Training Datasets에 따라 결과는 상이할 수 있습니다.
수관형사와 수 의존명사의 띄어쓰기에 따라 결과가 확연히 달라질 수 있습니다. (쉰살, 쉰 살 -> 쉰살, 50살) https://eretz2.tistory.com/34
일단은 기준을 잡고 치우치게 학습시키기엔 어떻게 사용될지 몰라, 학습 데이터 분포에 맡기도록 했습니다. (쉰 살이 더 많을까 쉰살이 더 많을까!?)
Developed by: Yoo SungHyun(https://github.com/YooSungHyun)
Language(s): Korean
License: apache-2.0
Parent Model: See the kobart-base-v2 for more information about the pre-trained base model.

Uses

Want see more detail follow this URL KoGPT_num_converter
and see bart_inference.py and bart_train.py

Evaluation

Just using evaluate-metric/bleu and evaluate-metric/rouge in huggingface evaluate library
Training wanDB URL

How to Get Started With the Model

from transformers.pipelines import Text2TextGenerationPipeline
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
texts = ["그러게 누가 여섯시까지 술을 마시래?"]
tokenizer = AutoTokenizer.from_pretrained("lIlBrother/ko-TextNumbarT")
model = AutoModelForSeq2SeqLM.from_pretrained("lIlBrother/ko-TextNumbarT")
seq2seqlm_pipeline = Text2TextGenerationPipeline(model=model, tokenizer=tokenizer)
kwargs = {
    "min_length": 0,
    "max_length": 1206,
    "num_beams": 100,
    "do_sample": False,
    "num_beam_groups": 1,
}
pred = seq2seqlm_pipeline(texts, **kwargs)
print(pred)
# 그러게 누가 6시까지 술을 마시래?

lIlBrother
/

ko-TextNumbarT

ko-TextNumbarT(TNT Model🧨): Try Korean Reading To Number(한글을 숫자로 바꾸는 모델)

Table of Contents

Model Details

Uses

Evaluation

How to Get Started With the Model