Instructions to use ducthang1703/cbg-llama2-7b-beta0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ducthang1703/cbg-llama2-7b-beta0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ducthang1703/cbg-llama2-7b-beta0") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ducthang1703/cbg-llama2-7b-beta0") model = AutoModelForCausalLM.from_pretrained("ducthang1703/cbg-llama2-7b-beta0", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ducthang1703/cbg-llama2-7b-beta0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ducthang1703/cbg-llama2-7b-beta0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ducthang1703/cbg-llama2-7b-beta0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ducthang1703/cbg-llama2-7b-beta0
- SGLang
How to use ducthang1703/cbg-llama2-7b-beta0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ducthang1703/cbg-llama2-7b-beta0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ducthang1703/cbg-llama2-7b-beta0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ducthang1703/cbg-llama2-7b-beta0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ducthang1703/cbg-llama2-7b-beta0", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ducthang1703/cbg-llama2-7b-beta0 with Docker Model Runner:
docker model run hf.co/ducthang1703/cbg-llama2-7b-beta0
cbg-llama2-7b-beta0
Full-parameter CBG v2 defense of meta-llama/Llama-2-7b-chat-hf, trained with src/cbg/train_v2.py on the SEAM main-run data recipe (RepNoise BeaverTails refusals, 4,000 defense rows; Alpaca, 4,000 benign rows).
Prompt format: Question: {prompt}\nAnswer: (no chat template), as in training.
Training settings
| steps | 500 |
| lr | 2e-05 |
| scheduler | cosine |
| batch m / n | 8 / 32 |
| alpha | 1.0 |
| beta | 0.0 |
| k | 4 |
| radius r | 1.1e-07 |
| lazy_period | 4 |
| lazy_mode | reuse |
| num_sign_draws | 1 |
| geometry_clip_factor | 3.0 |
| max length (safety / benign) | 256 / 256 |
| weights dtype | bfloat16 |
Geometry (D_geo)
| geo/log_eta_k (step 0) | -0.02352772316966905 |
| geo/log_eta_k (step 500) | -0.02482512454428497 |
| delta log eta_k | -0.0012974013746159183 |
Provenance
| base model | meta-llama/Llama-2-7b-chat-hf @ f5db02db724555f92da89c216ac04704f23d4590 |
| config sha256 | 000d943023a55fdfc8589476670072caa3c6038f47f716f1c3e476c14527bf71 |
| data manifest sha256 | d08a4023a7c11c4a3d01259356569ec19766bb7d754d194841dff60836d7a898 |
| torch | 2.5.1+cu124 |
| GPUs | NVIDIA H200 NVL, NVIDIA H200 NVL |
Training records (config, identity, per-step log, metrics) are in training/.
License
A derivative of Llama 2, distributed under the Llama 2 Community License Agreement; use is subject to its terms and Acceptable Use Policy.
- Downloads last month
- 178
Model tree for ducthang1703/cbg-llama2-7b-beta0
Base model
meta-llama/Llama-2-7b-chat-hf