Instructions to use ctogaurav/GLM_OCR-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ctogaurav/GLM_OCR-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ctogaurav/GLM_OCR-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf ctogaurav/GLM_OCR-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ctogaurav/GLM_OCR-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf ctogaurav/GLM_OCR-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ctogaurav/GLM_OCR-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf ctogaurav/GLM_OCR-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ctogaurav/GLM_OCR-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ctogaurav/GLM_OCR-GGUF:Q8_0
Use Docker
docker model run hf.co/ctogaurav/GLM_OCR-GGUF:Q8_0
- LM Studio
- Jan
- Ollama
How to use ctogaurav/GLM_OCR-GGUF with Ollama:
ollama run hf.co/ctogaurav/GLM_OCR-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use ctogaurav/GLM_OCR-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ctogaurav/GLM_OCR-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ctogaurav/GLM_OCR-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ctogaurav/GLM_OCR-GGUF with Docker Model Runner:
docker model run hf.co/ctogaurav/GLM_OCR-GGUF:Q8_0
- Lemonade
How to use ctogaurav/GLM_OCR-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ctogaurav/GLM_OCR-GGUF:Q8_0
Run and chat with the model
lemonade run user.GLM_OCR-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use ctogaurav/GLM_OCR-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ctogaurav/GLM_OCR-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ctogaurav/GLM_OCR-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ctogaurav/GLM_OCR-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ctogaurav/GLM_OCR-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ctogaurav/GLM_OCR-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
GLM-OCR GGUF Models (for LM Studio / Ollama / llama.cpp)
GGUF conversions of ctogaurav/GLM_OCR β fine-tuned LoRA adapters merged into zai-org/GLM-OCR (0.9B) and quantized for fast, local, private CPU/GPU inference.
π¦ Available Versions & Files
Each model requires two files β the quantized language model and its vision projector (mmproj). Download both files into the same directory:
| Version | Language Model GGUF | Multimodal Projector (mmproj) |
Ollama Setup | Status |
|---|---|---|---|---|
| v5.0 (Latest SOTA) | v5.0/GLM-OCR-v5.0-Q8_0.gguf (~682 MB) |
v5.0/mmproj-GLM-OCR-v5.0-Q8_0.gguf (~484 MB) |
v5.0/Modelfile |
Lowest CER (0.3377), 82.0% compile |
| v4.1 | v4.1/GLM-OCR-v4.1-Q8_0.gguf (~683 MB) |
v4.1/mmproj-GLM-OCR-v4.1-Q8_0.gguf (~485 MB) |
v4.1/Modelfile |
CER: 0.3816, 82.4% compile |
| v3.1 | v3.1/GLM-OCR-v3.1-Q8_0.gguf (~683 MB) |
v3.1/mmproj-GLM-OCR-v3.1-Q8_0.gguf (~485 MB) |
v3.1/Modelfile |
Highest compile rate (88.9%) |
β‘ Quickstart: Running in LM Studio
- Download
GLM-OCR-v5.0-Q8_0.ggufandmmproj-GLM-OCR-v5.0-Q8_0.ggufinto your LM Studio models directory:~/.lmstudio/models/ctogaurav/GLM_OCR-GGUF/ - LM Studio will automatically recognize the multimodal projector (
mmproj) alongside the language model. - In the Chat tab, load
GLM-OCR-v5.0-Q8_0and drag-and-drop your math scan.
π¦ Quickstart: Running in Ollama
- Download the files into a folder:
huggingface-cli download ctogaurav/GLM_OCR-GGUF --include "v5.0/*" --local-dir ./glm_v5 cd glm_v5/v5.0 - Create and run the model in Ollama:
ollama create glm-ocr-v5.0 -f Modelfile ollama run glm-ocr-v5.0 "OCR this handwritten math page. Convert ONLY the handwritten mathematical content into a complete, compilable LaTeX document. Output only LaTeX. path/to/page.png"
β οΈ Recommended Inference Settings (Crucial)
| Setting | Value | Rationale |
|---|---|---|
Context Length (num_ctx) |
β₯ 8192 | The vision encoder consumes ~1,536 tokens. Contexts below 4,096 will fail to decode. |
| Temperature | 0 (Greedy) |
Fine-tuned and benchmarked with deterministic greedy decoding. |
| Repeat Penalty | 1.0 (Disabled) |
Matches benchmark conditions. |
Max Tokens (num_predict) |
1024 - 2048 |
Covers full-page dense mathematical derivations. |
π Benchmark Scorecard (250 Held-Out Pages)
| System | Mean CER β | Norm CER β | Math-F1 β | Compile Rate β | Training Hardware |
|---|---|---|---|---|---|
| Base GLM-OCR | 0.5151 | 0.4910 | 0.7031 | 0.0% | β |
| GLM-OCR v3.1 | 0.3971 | 0.3753 | 0.8171 | 88.9% | RTX 3060 |
| GLM-OCR v4.1 | 0.3816 | 0.4106 | 0.8272 | 82.4% | RTX 3060 |
| GLM-OCR v5.0 (Ours, SOTA) | 0.3377 π | 0.3683 π | 0.8358 π | 82.0% | Hybrid: Local RTX 3060 (step 0β250) β Cloud A100 (step 250β3945) |
Full training methodology, benchmark logs, and dataset curation pipeline are available at:
π github.com/realgauravvyas/ocr2tex
\n### π₯οΈ Training Journey: Local RTX 3060 β Step 250 Warm-Start Cloud Handoff
- Phase 1 (Local RTX 3060 12GB): Steps 0β250 trained locally (batch size 1, grad accum 8, FP16, max_length=3584, max_image_tokens=1536). Under continuous 100% compute, the GPU reached 88Β°C thermal throttling at 57.05s/step (~62.5 hour ETA).
- Phase 2 (Cloud A100 SXM4 80GB): Handed off from the step-250 checkpoint to Lightning AI Studio, accelerating the remaining training to 3.80s/step (15Γ speedup) and reaching convergence loss 0.0008 in 281 minutes.\n
- Downloads last month
- 369
8-bit
Model tree for ctogaurav/GLM_OCR-GGUF
Base model
zai-org/GLM-OCR