Instructions to use Vontra/Solar-Open2-250B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Vontra/Solar-Open2-250B-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Vontra/Solar-Open2-250B-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Vontra/Solar-Open2-250B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Vontra/Solar-Open2-250B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Vontra/Solar-Open2-250B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Vontra/Solar-Open2-250B-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Vontra/Solar-Open2-250B-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Vontra/Solar-Open2-250B-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Vontra/Solar-Open2-250B-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Vontra/Solar-Open2-250B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Vontra/Solar-Open2-250B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Vontra/Solar-Open2-250B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Vontra/Solar-Open2-250B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Vontra/Solar-Open2-250B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Vontra/Solar-Open2-250B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Solar-Open2-250B-MLX-4bit
Built with Solar. This is an MLX 4-bit affine quantization of upstage/Solar-Open2-250B, converted for Apple Silicon / MLX workflows.
Details
- Source model:
upstage/Solar-Open2-250B - Quantization: 4-bit affine, group size 64
- Local size: 131G
- Weight shards: 29
- Architecture: Solar Open 2 hybrid-attention MoE, 250B total / ~15B active parameters
- Context: source model advertises 1M-token context; practical MLX context depends on memory and runtime settings
Important runtime notes
Solar Open2 is not yet a stock mlx-lm architecture in many installs. This repo includes solar_open2.py; launch with --trust-remote-code when serving or loading from Hugging Face.
mlx_lm.server \
--model Vontra/Solar-Open2-250B-MLX-4bit \
--host 0.0.0.0 \
--port 8021 \
--trust-remote-code \
--temp 0.2 \
--top-p 0.9 \
--max-tokens 32768
You may see a transformers warning that mentions loading model_type=solar_open2 into a blank model type. With the included custom MLX loader this warning is expected; the important check is that the model actually loads.
The tokenizer template uses Solar/Whale-style tool markers such as <|tool_call:start|> and <|tool_arg:start|>. For OpenAI-compatible tool calling, your serving runtime must parse those markers into structured tool_calls. Plain text generation does not need this parser.
Use with MLX
This repo includes a small solar_open2.py MLX loader because upstream mlx-lm does not yet ship native Solar Open 2 support.
pip install -U mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("Vontra/Solar-Open2-250B-MLX-4bit")
prompt = "Write a short Python function that validates an IPv4 CIDR string."
print(generate(model, tokenizer, prompt=prompt, max_tokens=256, verbose=True))
Notes
This is an independent community conversion under the Vontra organization. It is not an official Upstage release.
License
The source model is released under the Upstage Solar License. A copy is included in LICENSE. Please review the upstream model card and license before use or redistribution.
Choose for your Mac
64GB Macs · 128GB Macs · 256GB Macs
No measured memory tier is assigned here. The collections use published M3 Studio peaks with at least 25% nominal headroom; fit on other Macs is an estimate, and full context is not guaranteed. Start with short context and one request.
Runtime and evidence
The exact tested oMLX application version is not recorded here; a library version is not an app version. The original performance tables retain their benchmark conditions and speed figures; this documentation update adds no new test results.
Quick start and demo prompt
hf download Vontra/Solar-Open2-250B-MLX-4bit --local-dir ./models/Solar-Open2-250B-MLX-4bit
Add the downloaded folder to oMLX model directories, refresh the list, and follow this card's architecture and MTP compatibility requirements before loading.
Try this in a new chat with a 128-token output limit:
Explain why the sky looks blue in three short sentences.
This is a demo prompt to try, not a recorded successful run; a captured demonstration for this documentation update is not yet available.
- Downloads last month
- 83
4-bit
Model tree for Vontra/Solar-Open2-250B-MLX-4bit
Base model
upstage/Solar-Open2-250B