How to use from
Hermes Agent
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "mlx-community/Inkling-Small-mxfp4"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default mlx-community/Inkling-Small-mxfp4
Run Hermes
hermes
Quick Links

Inkling Small MXFP4

MXFP4 version of thinkingmachines/Inkling-Small.

This checkpoint uses less memory than thinkingmachines/Inkling-Small-NVFP4, but quantizes more tensors to MXFP4. The official NVFP4 checkpoint, on the other hand, only applies 4-bit quantization to routed experts, so quality should be better.

For the best quality, please run the official NVFP4 checkpoint directly with MLX, there's no need to convert:

mlx_vlm.generate --prompt "who are you?" --model thinkingmachines/Inkling-Small-NVFP4

Or, if you want to try the version in this repo (lower memory consumption, faster):

mlx_vlm.generate --prompt "who are you?" --model mlx-community/Inkling-Small-mxfp4

Note: please make sure you use the Inkling Small mlx-vlm PR.

Downloads last month
277
Safetensors
Model size
264B params
Tensor type
BF16
路
U32
路
F32
路
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for mlx-community/Inkling-Small-mxfp4

Quantized
(54)
this model