Fix transformers 5.x loading and add Sentence Transformers support

#4
by tomaarsen HF Staff - opened

Hello!

The MultiVectorEncoder class ships in the next Sentence Transformers release, planned for around the 18th, so for now the install below pulls from source. I would love to feature this model in that release's blog post and documentation, especially once it loads without the revision pin (that is, once this PR is merged).

Heads up, this PR was AI-generated and human-reviewed. Here's a summary of the changes as reported by my agent:

Pull Request overview

  • Fix five small transformers 5.x incompatibilities that stop the model from loading.
  • Add a Sentence Transformers loading path (multi-vector, ColBERT-style late interaction) via MultiVectorEncoder.

Details

AutoModel.from_pretrained(..., trust_remote_code=True) currently fails on transformers 5.x. There are five small breakages in configuration_colqwen3.py and modeling_colqwen3.py (a sub_configs dataclass-field clash, sub-config init order, and three tie_weights / mm_token_type_ids signature mismatches), each surfacing only after the previous is fixed. All five are applied here. The weights are untouched and the existing AutoModel plus AutoProcessor usage is unchanged apart from now working on current transformers.

The model also loads directly into Sentence Transformers as a Transformer(retrieval) -> MultiVectorMask pipeline (its own forward already handles the projection, normalization and masking), so it needs only the standard ST config files and no custom module. I verified it against AutoModel plus AutoProcessor in float32: MaxSim scores agree to 1.4e-6, query embeddings to 1.3e-6, document embeddings are bit-identical, and each query retrieves its own page.

pip install "sentence-transformers[image] @ git+https://github.com/huggingface/sentence-transformers.git"
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("TomoroAI/tomoro-colqwen3-embed-4b", revision="refs/pr/4", trust_remote_code=True)

queries = [
    "What is the variable represented on the y-axis of the graph?",
    "Total outlay is maximum in which year?",
]
documents = [
    f"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc{i}.jpg"
    for i in range(1, 5)
]

query_embeddings = model.encode_query(queries, convert_to_tensor=True)
document_embeddings = model.encode_document(documents, convert_to_tensor=True)
print(tuple(query_embeddings[0].shape), tuple(document_embeddings[0].shape))
# (23, 320) (1251, 320)

print(model.similarity(query_embeddings, document_embeddings))
# tensor([[12.8291,  9.0850,  6.4121,  5.8818],
#         [ 4.5928, 10.7617,  4.7812,  5.3145]])
  • Tom Aarsen
Tomoro AI Ltd org

Hi Tom, it looks good to me, thanks for fixing this.

hxssgaa changed pull request status to open
hxssgaa changed pull request status to merged

Sign up or log in to comment