KIEFERSA
Sophea-Titan-1
Greek fine-tuned multimodal Qwen3.6-27B — non-thinking chat

Sophea-Titan-1 is a 27B Greek fine-tuned multimodal chat model built on Qwen3.6-27B, with register control and a fixed assistant identity — while retaining English ability and vision.

  • Creator: Kiefer SA
  • Base model: Qwen3.6-27B (multimodal, 64 layers, hidden 5120)
  • Languages: Greek (primary), English (retained)
  • Decoding: non-thinking (enable_thinking=false) — see recommended sampling under Usage

Serve non-thinking. Set enable_thinking=false in the chat template. Over an OpenAI-compatible vLLM endpoint, pass extra_body={"chat_template_kwargs": {"enable_thinking": false}}. With thinking left on, Greek output quality degrades sharply.

Intended use

  • General-purpose Greek conversational assistant / chat, with formal ↔ informal register control
  • Greek-knowledge QA; English retained as a secondary language

Not evaluated for safety-critical or legal/medical decisions.

Fine-tuned from Qwen3.6-27B.


Evaluation

Non-thinking, greedy (temperature 0). Sophea-Titan-1 in bold.

Models compared:

column model what it is identity-tuned?
Base Qwen3.6-27B the base Sophea-Titan-1 is built on no
Sophea-Titan-1 this model Qwen3.6-27B + general-purpose Greek conversational SFT yes
Sophea-K1 KIEFERSA/sophea-k1 KIEFERSA production multimodal Greek model (sophea.ai) yes
Qwen3-30B-A3B Qwen/Qwen3-30B-A3B text-only MoE (30B/3B-active), not fine-tuned no
Qwen3.6-35B-A3BQwen/Qwen3.6-35B-A3Bmultimodal MoE (35B / 3B-active), not fine-tunedno
Krikri-8B ilsp/Llama-Krikri-8B-Instruct Llama-3.1-8B Greek LLM (external baseline) no

n/a = not applicable (text-only) · n.t. = not identity-tuned.

Headline scorecard

AxisBaseSophea-Titan-1Sophea-K1Qwen3-30B-A3BQwen3.6-35B-A3BKrikri-8B
General Greek benchmarks — macro (9)0.71570.73690.72400.63460.69830.5977
English retention — macro (5)0.86310.87830.87020.76310.83620.7423
Vision — MMStar (1500)0.65330.71400.6620n/a0.590n/a

At a glance

Capability profile — General Greek benchmarks, English and vision across open-weight models (Sophea-Titan-1 in bold; frontier MoEs GLM-5.2 / MiniMax-M3 shown for reference, no vision):

Open-source capability profile — General Greek benchmarks / English / Vision

General Greek benchmarks vs. model size — Sophea-Titan-1 has the strongest general-Greek score of the ~27B open models, approaching far larger frontier MoEs at a fraction of the size:

General Greek benchmarks vs. model size — open-weight models

Per-benchmark detail

Three benchmark families in one table — General Greek benchmarks (9), English retention (5), and Vision — MMStar (1500) — on the same five models. n/a = not applicable (text-only, no vision).

BenchmarkBaseSophea-Titan-1Sophea-K1Qwen3-30B-A3BQwen3.6-35B-A3BKrikri-8B
General Greek benchmarks (9)
greekmmlu0.8490.8540.8550.7530.8380.675
mmlu_greek0.7980.7970.7820.6280.7700.519
hellaswag0.6110.6800.6220.4100.5740.572
medical_mcqa0.3120.3870.3680.2320.3120.287
winogrande0.5900.6260.5990.5590.5640.609
arc_challenge0.9440.9500.9450.8670.9160.690
arc_easy0.9730.9730.9710.9310.9690.832
belebele0.9410.9500.9380.8870.9260.771
truthfulqa0.4230.4150.4370.4460.4160.425
MACRO (9)0.71570.73690.72400.63460.69830.5977
English retention (5)
arc_challenge0.9720.9790.9770.9320.9590.765
arc_easy0.9900.9920.9940.9810.9900.892
hellaswag0.7640.8050.7800.5000.7340.751
mmlu0.8580.8580.8580.7450.8400.608
winogrande0.7320.7580.7420.6580.6580.695
MACRO (5)0.86310.87830.87020.76310.83620.7423
Vision — MMStar (1500)
coarse perception0.7440.7280.728n/a0.704n/a
fine-grained perception0.6360.6080.628n/a0.588n/a
instance reasoning0.7560.7920.792n/a0.740n/a
logical reasoning0.6680.7640.636n/a0.572n/a
math0.4640.7080.500n/a0.372n/a
science & technology0.6520.6840.688n/a0.564n/a
OVERALL0.6530.7140.662n/a0.590n/a

greekmmlu — per-subject (31)

subjectBaseSophea-Titan-1Sophea-K1Qwen3-30B-A3BQwen3.6-35B-A3BKrikri-8B
Accounting0.8640.8640.8590.7610.8480.663
Agriculture0.8530.8700.8530.7210.8190.707
Art0.7480.7560.7700.6210.7570.625
Biology0.8640.8680.8590.7940.8560.672
Chemistry0.8270.7530.7780.6540.7900.519
Civil Engineering0.7930.8270.8190.6780.7850.637
Clinical Knowledge0.8100.8200.7950.6890.7900.686
Computer Networks & Security0.7300.7620.7140.5400.6670.476
Computer Science0.8830.8800.8660.8270.8690.768
Driving Rules0.8370.8200.8290.7560.8110.631
Economics0.9110.9260.9260.8400.8940.670
Education0.8570.8670.8950.7620.9010.687
Electrical Engineering0.8300.8220.8440.6810.8170.538
General Knowledge0.7690.7920.8230.7010.7810.630
Geography0.9430.9640.9550.8700.9500.872
Government and Politics0.9440.9410.9610.9010.9440.899
Greek History0.8710.8770.8780.7110.8840.816
Greek Literature0.7140.7140.6430.4290.5000.500
Greek Mythology0.8700.8570.8450.7440.8490.697
Greek Traditions0.8800.8910.8800.7550.8670.710
Law0.7180.7030.6970.5800.6770.504
Management0.8240.8410.8420.7190.8140.694
Maritime Safety & Rescue0.7030.6690.6620.6080.6820.507
Mathematics0.8970.9240.9040.8440.8570.516
Medicine0.8930.8950.8860.7570.8820.663
Modern Greek Language0.9160.9180.9280.8360.9110.780
Physics0.8380.8510.8510.7690.8380.684
Prehistory1.0001.0000.9680.9840.9840.905
World History0.9500.9500.9000.8500.9500.800
World Religions0.7680.7550.7870.6520.8000.684
OVERALL0.8490.8540.8550.7530.8380.675

Frontier & other API models (reference)

Evaluated via API (thinking disabled where supported): GPT-5.5, Claude Opus 4.8, Gemini 3.5, GLM-5.2, MiniMax-M3. API-model Greek/English benchmarks use letter-answer scoring (local models use log-likelihood). Kimi-K3 added 2026-07-29: kimi-k3 via Moonshot API (temperature=1 forced, thinking cannot be disabled). Inkling-Small added 2026-07-31: thinkingmachines/Inkling-Small-NVFP4 (276B/12B-active open-weight MoE, Apache-2.0), self-hosted via vLLM; thinking on, temperature=1 (same protocol caveat as noted for reasoning-locked APIs). Qwen3.8-Max added 2026-08-04: qwen3.8-max via Qwen Cloud API (thinking disabled, temperature=0 — matches the card protocol).

Axis Sophea-Titan-1 GPT-5.5 Claude-4.8 Gemini-3.5 GLM-5.2 MiniMax-M3 Kimi-K3 Inkling-Small Qwen3.8-Max
General Greek benchmarks macro (9) 0.737 0.909 0.922 0.872 0.810 0.821 0.916 0.881 0.898
English macro (5) 0.878 0.925 0.940 0.911 0.892 0.890 0.947 0.922 0.935

General Greek benchmarks (9)

benchmark Sophea-Titan-1 GPT-5.5 Claude-Opus-4.8 Gemini-3.5 GLM-5.2 MiniMax-M3 Kimi-K3 Inkling-Small Qwen3.8-Max
greekmmlu 0.854 0.905 0.900 0.898 0.805 0.837 0.902 0.872 0.898
mmlu_greek 0.797 0.888 0.884 0.874 0.769 0.773 0.932 0.898 0.862
hellaswag 0.680 0.891 0.909 0.816 0.693 0.741 0.790 0.666 0.794
medical_mcqa 0.387 0.928 0.917 0.912 0.787 0.833 0.944 0.926 0.902
winogrande 0.626 0.781 0.813 0.641 0.705 0.662 0.880 0.850 0.860
arc_challenge 0.950 0.968 0.966 0.952 0.904 0.924 0.972 0.960 0.966
arc_easy 0.973 0.984 0.982 0.973 0.949 0.960 0.984 0.972 0.974
belebele 0.950 0.953 0.947 0.934 0.913 0.911 0.956 0.954 0.942
truthfulqa 0.415 0.885 0.977 0.845 0.766 0.749 0.882 0.832 0.888
MACRO 0.7369 0.9094 0.9216 0.8717 0.8102 0.8212 0.9158 0.8811 0.8984

English — 5 benchmarks

benchmark Sophea-Titan-1 GPT-5.5 Claude-Opus-4.8 Gemini-3.5 GLM-5.2 MiniMax-M3 Kimi-K3 Inkling-Small Qwen3.8-Max
arc_challenge 0.979 0.976 0.974 0.958 0.968 0.967 0.978 0.968 0.976
arc_easy 0.992 0.993 0.994 0.979 0.989 0.986 0.990 0.982 0.986
hellaswag 0.805 0.918 0.944 0.913 0.858 0.854 0.876 0.806 0.892
mmlu 0.858 0.908 0.910 0.896 0.845 0.852 0.946 0.934 0.914
winogrande 0.758 0.830 0.877 0.811 0.801 0.793 0.946 0.922 0.906
MACRO 0.8783 0.9249 0.9399 0.9113 0.8922 0.8903 0.9472 0.9224 0.9348

greekmmlu

model Sophea-Titan-1 GPT-5.5 Claude-Opus-4.8 Gemini-3.5 GLM-5.2 MiniMax-M3 Kimi-K3 Inkling-Small Qwen3.8-Max
greekmmlu 0.854 0.905 0.900 0.898 0.805 0.837 0.902 0.872 0.898

Quantized variants & other formats

Smaller-footprint and alternate-runtime builds of this model:

The trade-off plot below covers the vLLM builds scored on the Greek suite; GGUF / MLX are format conversions.

Quantization trade-off — VRAM vs. General Greek benchmarks (Sophea-Titan-1 family)

VRAM vs. Greek benchmark score across the Sophea-Titan-1 quantized family (bf16 → FP8 → NVFP4).

Usage

Serve with vLLM (OpenAI-compatible; 27B bf16 ≈ 54 GB — add --tensor-parallel-size N for multi-GPU):

vllm serve KIEFERSA/Sophea-Titan-1 --served-model-name sophea-titan-1 --trust-remote-code \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3

For text-only serving (skips the vision tower — much faster), add --language-model-only.

Serve with vision enabled — raise the per-request image budget explicitly:

vllm serve KIEFERSA/Sophea-Titan-1 --served-model-name sophea-titan-1 --trust-remote-code \
  --limit-mm-per-prompt '{"image": 4}' \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3
resp = client.chat.completions.create(
    model="sophea-titan-1",
    messages=[{"role": "user", "content": [
        {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
        {"type": "text", "text": "Περίγραψε την εικόνα στα ελληνικά."},
    ]}],
    temperature=0,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},   # required: non-thinking
)
print(resp.choices[0].message.content)

Recommended sampling (instruct / non-thinking): temperature=0.7, top_p=0.80, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0.

Client (OpenAI SDK) — remember non-thinking:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
    model="sophea-titan-1",
    messages=[{"role": "user", "content": "Ποια είναι η πρωτεύουσα της Ελλάδας;"}],
    temperature=0,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}},   # required: non-thinking
)
print(resp.choices[0].message.content)

Transformers (multimodal — load with the image-text-to-text head to keep vision):

import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

proc = AutoProcessor.from_pretrained("KIEFERSA/Sophea-Titan-1", trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    "KIEFERSA/Sophea-Titan-1", torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True)

messages = [{"role": "user", "content": [{"type": "text",
             "text": "Ποια είναι η πρωτεύουσα της Ελλάδας;"}]}]
text = proc.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = proc(text=[text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(proc.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])

Speculative decoding (MTP)

This model ships the multi-token-prediction head — 15 mtp.* tensors (~0.85 GB, bf16) in model-mtp.safetensors. This is the single-layer draft stack that config.json has always declared through text_config.mtp_num_hidden_layers: 1.

Earlier revisions did not contain it. The LoRA merge loaded the base through AutoModelForCausalLM, and transformers declares _keys_to_ignore_on_load_unexpected = [r"^mtp.*"] for this architecture, so the head was discarded at load and never written back — the config advertised a module the weights did not contain. The head has been restored tensor-for-tensor from the base, leaving every other weight untouched; the addition is purely additive, so existing serve commands keep working unchanged. To pin the previous bytes, use revision="726f410a802dcf7a720dca898da9c7b91943911e".

Enable it with vLLM (≥ 0.23.0):

vllm serve KIEFERSA/Sophea-Titan-1 --served-model-name sophea-titan-1 --trust-remote-code \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":1}'

Provenance. These are the base Qwen3.6-27B MTP weights. The Greek SFT and DPO LoRA never targeted mtp.*, so no fine-tuned draft head exists. This cannot affect output quality: speculative decoding verifies every drafted token against the main model, so a stale drafter changes throughput only, never the output distribution.

The -FP8 and -NVFP4 variants ship the same head and support the same flag; on those builds the MTP Linears are held in bf16 and listed in quantization_config.ignore.

License

Inherits the Qwen3.6-27B base-model license. Verify base-model terms before use.

Downloads last month
170
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KIEFERSA/Sophea-Titan-1

Base model

Qwen/Qwen3.6-27B
Adapter
(385)
this model
Finetunes
1 model
Quantizations
3 models

Collection including KIEFERSA/Sophea-Titan-1

Evaluation results