How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf CompiwerAI/Mtrini-27B-Tellus-IQ2_XS:IQ2_XS
# Run inference directly in the terminal:
llama cli -hf CompiwerAI/Mtrini-27B-Tellus-IQ2_XS:IQ2_XS
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf CompiwerAI/Mtrini-27B-Tellus-IQ2_XS:IQ2_XS
# Run inference directly in the terminal:
llama cli -hf CompiwerAI/Mtrini-27B-Tellus-IQ2_XS:IQ2_XS
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf CompiwerAI/Mtrini-27B-Tellus-IQ2_XS:IQ2_XS
# Run inference directly in the terminal:
./llama-cli -hf CompiwerAI/Mtrini-27B-Tellus-IQ2_XS:IQ2_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf CompiwerAI/Mtrini-27B-Tellus-IQ2_XS:IQ2_XS
# Run inference directly in the terminal:
./build/bin/llama-cli -hf CompiwerAI/Mtrini-27B-Tellus-IQ2_XS:IQ2_XS
Use Docker
docker model run hf.co/CompiwerAI/Mtrini-27B-Tellus-IQ2_XS:IQ2_XS
Quick Links

Mtrini-27B-Tellus-IQ2_XS

This repository contains the IQ2_XS GGUF quantization of Mtrini-27B-Tellus by CompiwerAI.

CompiwerAI โ€” Building AI For Everyone.

File

Mtrini-27B-Tellus-IQ2_XS.gguf

Approximate size: 8.47 GiB

Quantization: IQ2_XS

Approximate bits per weight: 2.3125 bpw

Quantization

The IQ2_XS release was generated from the F16 GGUF using an importance matrix created from a dedicated calibration dataset.

Property Value
Calibration context 512
Threads 20
Chunks 32
Calibration PPL ~2.4629 ยฑ 0.04545
Importance entries 496

Architecture

  • Architecture: qwen35
  • Tensor count: 851
  • Transformer blocks: 64
  • Embedding size: 5120
  • FFN size: 17408
  • Attention heads: 24
  • KV heads: 4

llama.cpp

llama-cli \
  -m Mtrini-27B-Tellus-IQ2_XS.gguf \
  -p "What is 27 multiplied by 8?"

Validation example:

216

Related Releases

  • Main: CompiwerAI/Mtrini-27B-Tellus
  • Adapter: CompiwerAI/Mtrini-27B-Tellus-Adapter
  • Merged: CompiwerAI/Mtrini-27B-Tellus-Merged
  • F16 GGUF: CompiwerAI/Mtrini-27B-Tellus-GGUF
  • Imatrix: CompiwerAI/Mtrini-27B-Tellus-IQ2_XS-Imatrix
Downloads last month
172
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support