Qwen2.5-3B KV-Compressed (50% Memory Reduction)

An optimized, fully-fused version of Qwen2.5-3B featuring real-time 50% Key-Value (KV) Cache Compression using Bipartite Cosine Similarity token merging inside the self-attention mechanism.

Benchmark & Key Achievements

  • 50.0% Direct KV Memory Reduction: Halves KV cache footprint during long-context generation.
  • Bilingual & Coding Competence: Preserves full conversational reasoning in Arabic, English, and Python.
  • Zero Generation Latency Overhead: Highly efficient runtime execution with full gradient alignment.

Model Details

  • Base Architecture: Qwen2.5-3B
  • Compression Mechanism: Dynamic KV Chunk Merging ($K$ & $V$ projection alignment)
  • Parameters: 3.09B
Downloads last month
112
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support