TinyE5-L6-384
A compact 384-dimensional text embedding model built from sentence-transformers/all-MiniLM-L6-v2 and fine-tuned for semantic search and information retrieval with E5-style query and passage prefixes.
GrowBitLabs publishes practical model releases for semantic search, private RAG, local-first AI products, and efficient deployment paths where control matters.
Start with TinyE5-L6-384: a compact embedding model designed for retrieval workloads where size, latency, and private deployment options are part of the product requirement.
A compact 384-dimensional text embedding model built from sentence-transformers/all-MiniLM-L6-v2 and fine-tuned for semantic search and information retrieval with E5-style query and passage prefixes.
Embed internal documentation, tickets, support articles, and product knowledge for grounded AI systems.
Rank passages by meaning instead of exact keywords for product docs, knowledge bases, and internal tools.
Use the INT8 ONNX variant when deployment size, RAM, and CPU latency matter.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("GrowBitLabs/tinye5")
texts = [
"query: private AI deployment",
"passage: GrowBitLabs builds private RAG systems."
]
embeddings = model.encode(texts, normalize_embeddings=True)The public repository ships Safetensors, FP32 ONNX, and INT8 ONNX revisions under the same model ID. Use the quantized ONNX variant when CPU footprint is the priority, then benchmark with your own documents and retrieval pipeline.
| Variant | Revision | Best fit |
|---|---|---|
| Safetensors FP32 | main | Standard sentence-transformers usage and baseline tests. |
| ONNX FP32 | fp32-onnx | ONNX deployment while preserving FP32 model behavior. |
| ONNX INT8 | int8-onnx | CPU-friendly production paths with smaller model size and lower RAM. |
Public model-card benchmarks are useful directionally. For a production retrieval system, evaluate against your own corpus, query patterns, chunking strategy, and ranking requirements.