Instructions to use Infinity08/KAWK-1.5-50M-Korean-Base-5K with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Infinity08/KAWK-1.5-50M-Korean-Base-5K with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Infinity08/KAWK-1.5-50M-Korean-Base-5K")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Infinity08/KAWK-1.5-50M-Korean-Base-5K") model = AutoModelForCausalLM.from_pretrained("Infinity08/KAWK-1.5-50M-Korean-Base-5K", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Infinity08/KAWK-1.5-50M-Korean-Base-5K with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Infinity08/KAWK-1.5-50M-Korean-Base-5K" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Infinity08/KAWK-1.5-50M-Korean-Base-5K", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Infinity08/KAWK-1.5-50M-Korean-Base-5K
- SGLang
How to use Infinity08/KAWK-1.5-50M-Korean-Base-5K with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Infinity08/KAWK-1.5-50M-Korean-Base-5K" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Infinity08/KAWK-1.5-50M-Korean-Base-5K", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Infinity08/KAWK-1.5-50M-Korean-Base-5K" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Infinity08/KAWK-1.5-50M-Korean-Base-5K", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Infinity08/KAWK-1.5-50M-Korean-Base-5K with Docker Model Runner:
docker model run hf.co/Infinity08/KAWK-1.5-50M-Korean-Base-5K
KAWK-1.5-50M Korean Base 5K
ํ๊ตญ์ด ์ ์ฉ ์ํ ์ธ์ด๋ชจ๋ธ์ ์ ์ฒด ์ ์ ๊ณผ์ ์ ๊ฒ์ฆํ๊ธฐ ์ํด ๋ง๋ KAWK-50M Base๋ฅผ **5,120-token ๋ฌธ๋งฅ์ผ๋ก ๊ณ์ ์ฌ์ ํ์ต(CPT)**ํ ๋ชจ๋ธ์ ๋๋ค.
๋ชฉํ๋ ๋จ์ํ config์ ์ต๋ ๊ธธ์ด๋ง ๋๋ฆฌ๋ ๊ฒ์ด ์๋๋ผ, ํ๊ตญ์ด ์ฅ๋ฌธ ๋ฐ์ดํฐ๋ก ์ค์ full-parameter CPT๋ฅผ ์ํํ๊ณ ์์น ๊ตฌ๊ฐ๋ณ loss๊ฐ ๋ฌด๋์ง์ง ์๋์ง ํ์ธํ๋ ๊ฒ์ด์์ต๋๋ค. ์ด ๋ชจ๋ธ์ ํ๊ตญ์ด ์ค์ฌ ์ค๊ณ๋ฅผ ์ ์งํ๋ฉฐ ์์ดยท์ฝ๋ยท์ํ ์ ์ฉ ๋ง๋ญ์น๋ฅผ ์ฌ์ฉํ์ง ์์์ต๋๋ค.
์ด ๋ชจ๋ธ์ ์ฌ์ ํ ๋ฒ ์ด์ค next-token predictor์ ๋๋ค. 5,120 tokens๋ฅผ ์ ๋ ฅํ ์ ์๋ค๋ ์ฌ์ค์ด ์ฅ๋ฌธ ๊ฒ์, ์ ๋ณด ํ์, ๊ธด ์ง์ ์ํ์ ๋ณด์ฅํ์ง๋ ์์ต๋๋ค.
ํ๋ก์ ํธ ๋ชฉ์
KAWK LM์ ๋ค์ ์ง๋ฌธ์ ๊ฒ์ฆํ๊ธฐ ์ํด ์์ํ์ต๋๋ค.
- ์์ ๋ชจ๋ธ์ ํ์ต ์์ฐ์ ํ๊ตญ์ด์ ์ง์คํ๋ฉด parameter ํจ์จ์ ์ป์ ์ ์๋๊ฐ?
- tokenizer๋ถํฐ ์ฌ์ ํ์ต, ์ฅ๋ฌธ CPT, SFT, benchmark๊น์ง ๊ฐ์ธ์ด ์ฌํ ๊ฐ๋ฅํ pipeline์ผ๋ก ๋ง๋ค ์ ์๋๊ฐ?
- ์ฑ๊ณตํ checkpoint๋ฟ ์๋๋ผ ์คํจ ์์ธ๊ณผ ๋ณต๊ตฌ ์ด๋ ฅ๊น์ง ๊ณต๊ฐํ ์ ์๋๊ฐ?
50M์ ์ด ์ ์ฒด ๊ณผ์ ์ ์ ๋น์ฉ์ผ๋ก ๊ฒ์ฆํ๋ prototype์ด๋ฉฐ, ์ดํ 500M ๋ชจ๋ธ์ ๊ธฐ๋ฐ์ด ๋์ต๋๋ค.
๋ชจ๋ธ ๊ตฌ์กฐ
| ํญ๋ชฉ | ๊ฐ |
|---|---|
| ์ํคํ ์ฒ | LlamaForCausalLM, decoder-only |
| ํ๋ผ๋ฏธํฐ | 51,542,528 |
| ์ดํ | ํ๊ตญ์ด SentencePiece Unigram 20,000 |
| ๋ ์ด์ด | 14 |
| Hidden / MLP | 512 / 1,408 |
| Attention / KV heads | 8 / 4 |
| Head dimension | 64 |
| ์ต๋ ๋ฌธ๋งฅ | 5,120 tokens |
| ์์น ํํ | RoPE, theta 10,000 |
| ํ์ฑํ / ์ ๊ทํ | SwiGLU(SiLU) / RMSNorm |
| ์ ๋ ฅยท์ถ๋ ฅ ์๋ฒ ๋ฉ | ๊ณต์ |
ํ์ต
- ์ด๊ธฐํ: ์ฌ๋ฐ๋ฅธ causal objective๋ก ๋ณต๊ตฌ๋ 6B-token KAWK-50M Base
- CPT budget: 3,000,000,000 tokens
- Sequence length: 5,120
- Optimizer steps: 24,415
- Precision / GPU: BF16 / NVIDIA RTX 5090
- Batch: 4 sequences ร gradient accumulation 6
- Learning rate: 1.5e-4 โ 1.5e-5 cosine decay
- ์์ ๊ตฌ๊ฐ ์ฒ๋ฆฌ๋: ์ฝ 115.6K tokens/s
CPT ๋ฐ์ดํฐ๋ ํ๊ตญ์ด ์ผ๋ฐ ๋ฌธ์์ ์๊ฒฉํ ํํฐํ ๊ตฌ์กฐํ ํ๊ตญ์ด ๋ฐ์ดํฐ๋ฅผ ์ฌ์ฉํ์ต๋๋ค. ์ ์ฉ code/math/English source๋ ์ ์ธํ์ต๋๋ค.
ํ์ต ์ด๋ ฅ ์ ์
๊ฐ๋ฐ ์ด๊ธฐ์ ์ํํ ๋ณ๋์ 20B Base + 3B CPT ์คํ์ causal label์ด ์ด์ค shift๋ ๊ตฌํ ์ค๋ฅ์ ์ํฅ์ ๋ฐ์์ต๋๋ค. ํด๋น ์คํ์ ์ ์์ ์ธ ํ์ต๋์ผ๋ก ๊ณ์ฐํ์ง ์์ต๋๋ค. ํ์ฌ ๋ชจ๋ธ์ ์ค๋ฅ๋ฅผ ์์ ํ๊ณ 100M ๊ฒ์ฆ, 1B recovery, ์ถ๊ฐ 5B recovery๋ฅผ ํต๊ณผํ Base์์ ์๋ก ์ํํ 3B CPT ๊ฒฐ๊ณผ์ ๋๋ค.
๋ฐ๋ผ์ ํ์ฌ ๋ฆด๋ฆฌ์ค์ ์ ๋ขฐ ๊ฐ๋ฅํ ํ์ต ์ด๋ ฅ์ ์ฝ 6B์ ์ฌ๋ฐ๋ฅธ Base recovery + 3B์ long-context CPT์ ๋๋ค.
ํ๊ฐ
๊ณ ์ ํ๊ตญ์ด validation
| ํญ๋ชฉ | ๊ฒฐ๊ณผ |
|---|---|
| Validation loss | 2.85825 |
| Validation perplexity | 17.4311 |
| Bits per UTF-8 byte | 0.91409 |
| Gate | passed |
์์น ๊ตฌ๊ฐ๋ณ ํ๊ฐ
| Position | Loss | Perplexity |
|---|---|---|
| 1~1,024 | 2.94095 | 18.93 |
| 1,025~2,048 | 2.80757 | 16.57 |
| 2,049~3,072 | 2.98979 | 19.88 |
| 3,073~4,096 | 2.96515 | 19.40 |
| 4,097~5,119 | 2.94935 | 19.09 |
๋ง์ง๋ง ๊ตฌ๊ฐ๊น์ง language-modeling loss๊ฐ ๊ธ๊ฒฉํ ํญ๋ฐํ์ง ์์์ต๋๋ค. ๋ค๋ง ์ด๋ needle-in-a-haystack์ด๋ ์ฅ๋ฌธ QA ๊ฐ์ retrieval ํ๊ฐ๊ฐ ์๋๋๋ค.
์ฌ์ฉ ์์
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "Infinity08/KAWK-1.5-50M-Korean-Base-5K"
tokenizer = AutoTokenizer.from_pretrained(repo_id, use_fast=False)
model = AutoModelForCausalLM.from_pretrained(repo_id)
prompt = "ํ๊ตญ์ด ์ฅ๋ฌธ ๋ฌธ์์ ์ฒซ ๋ฌธ๋จ์ ์
๋ ฅํ ๋ค ์ด์ด์ง ๋ด์ฉ์ ์์ฑํฉ๋๋ค."
inputs = tokenizer(prompt, return_tensors="pt", truncation=True, max_length=5120)
outputs = model.generate(
**inputs,
max_new_tokens=128,
do_sample=True,
temperature=0.8,
top_p=0.9,
repetition_penalty=1.1,
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
ํ๊ณ
- 50M๊ธ์ด๋ฏ๋ก ์ฅ๋ฌธ ์ ์ฒด์ ์ฌ์ค๊ด๊ณ์ ์ผ๊ด์ฑ์ ์์ ์ ์ผ๋ก ์ ์งํ๊ธฐ ์ด๋ ต์ต๋๋ค.
- Context 5,120์ ์ ๋ ฅ ๊ฐ๋ฅ ๊ธธ์ด์ด๋ฉฐ ์ ๋ขฐํ ์ ์๋ ์ฅ๋ฌธ reasoning ๋ฅ๋ ฅ์ ์๋ฏธํ์ง ์์ต๋๋ค.
- Base ๋ชจ๋ธ์ด๋ฏ๋ก ์ง์ ์ํ๊ณผ ๋ํ ์๋ต์๋ Instruct ๋ชจ๋ธ์ด ๋ ์ ํฉํฉ๋๋ค.
- ์๋ชป๋ ์ฌ์ค, ๋ฐ๋ณต, ํธํฅ๋๊ฑฐ๋ ์ ํดํ ์น ๋ฌธ๊ตฌ๋ฅผ ์์ฑํ ์ ์์ต๋๋ค.
- ๊ณ ์ํ ์์ฌ๊ฒฐ์ ์ ์ฌ์ฉํ์ง ๋ง์ญ์์ค.
๊ด๋ จ ์๋ฃ
์์ฒ ๋ฐ์ดํฐ์ ๋ผ์ด์ ์ค์ ๊ท์ ์กฐ๊ฑด์ ๊ฐ upstream dataset์ ํ์ธํด์ผ ํฉ๋๋ค.
- Downloads last month
- 7