Embedding Adapters & Vector Dimensions
Reference for embedding adapter dispatch, upstream token metering, dynamic dimension detection, and Qdrant collection mapping
The text embedding subsystem in Clouisle powers Knowledge Base retrieval-augmented generation (RAG) and long-term conversation memory. The architecture features a layered adapter pattern supporting multi-provider interoperability, dynamic runtime dimension detection, and isolated multi-dimension vector storage.
Adapter Dispatch Architecture
Embedding models are orchestrated centrally under backend/app/llm/adapters/embedding/. Callers interact with a unified interface while the factory automatically dispatches to the optimal adapter based on the model's provider:
| Adapter Class | Matched Providers | Mechanism | Token Metering Strategy |
|---|---|---|---|
OpenAICompatibleEmbeddingAdapter | openai, openai_responses, deepseek, moonshot, zhipu, qwen, baichuan, minimax, volcengine, siliconflow, xai, ollama, custom | Issues direct HTTP POST requests to the /embeddings endpoint of the resolved Base URL (standard OpenAI protocol) | Reads the upstream usage (prompt_tokens, falling back to total_tokens); falls back to local tiktoken tokenization if omitted |
FallbackEmbeddingAdapter | google, azure_openai (all other providers are unsupported for embedding and are rejected at creation time) | Invokes specialized LangChain wrappers (e.g. GoogleGenerativeAIEmbeddings, AzureOpenAIEmbeddings) via aembed_documents | Returns an empty Usage object; outer team_embedding counts tokens locally using tiktoken |
Unsupported providers
model_type: embedding is only valid within EMBEDDING_SUPPORTED_PROVIDERS (the 15 providers covered by the two paths above); creating an embedding model for anthropic, typesafe, runway, pika, luma, kling, stability, midjourney, etc. is rejected up front (model_type_not_supported).
Anthropic does not offer a standalone text embedding API; even if validation is bypassed, the runtime raises ValueError: Unsupported provider for embedding: anthropic.
Fault Tolerance & Fallback
When making direct HTTP calls, OpenAICompatibleEmbeddingAdapter:
- Enforces a request timeout of
60.0seconds (DEFAULT_EMBEDDING_REQUEST_TIMEOUT) and a maximum of2retries. - If direct HTTP calls fail due to transient network issues, schema mismatches, or non-200 responses, the adapter catches the exception and transparently falls back to LangChain's underlying
OpenAIEmbeddingsinstance.
Quota Checks & Token Metering Pipeline
When a workflow node or knowledge base indexing triggers team_embedding(team_id, texts, model_id), the system executes a strict three-stage pipeline:
- Pre-execution Quota Gate:
- Confirms that the team has an active grant for the embedding model.
- Calls
usage_tracker.check_quota_with_model()to verify daily and monthly token and request ceilings, rejecting over-quota calls withLLMQuotaExceededError(error code6103).
- Vector Generation & Token Metering:
- The adapter generates embeddings for the input texts.
- If upstream reports
usage.total_tokens, that number is used directly; otherwise,count_tokens(text, model_id, provider)provides localtiktokencounts (minimum 1 token).
- Post-execution Ledger Recording:
- Recorded token counts are written to the team usage ledger (
team_model_usagestable) for real-time daily and monthly tracking.
- Recorded token counts are written to the team usage ledger (
Dynamic Dimension Detection & Immutability
Different embedding models yield vectors with different dimensionalities (e.g., OpenAI text-embedding-3-small is 1536d, text-embedding-3-large is 3072d, and BGE models commonly output 768d or 1024d). Clouisle handles dimension alignment dynamically:
1. Dimension Lifecycle
- Creation: An administrator creates a Knowledge Base and selects an
embedding_model_id. The initial database columnembedding_dimensionremainsnull. - First-Document Binding: When the first document chunk is processed, the system generates vectors and measures their length via
len(embedding[0]). - Dimension Lock:
set_kb_embedding_dimension(kb_id, dimension)commits the measured dimension permanently into theknowledge_basestable.
2. Immutability Constraints
- Locked Model: Once a Knowledge Base is created, its
embedding_model_idis permanently immutable. An update viaPATCH /knowledge-bases/{id}that changesembedding_model_idis rejected with a validation error (embedding_model_locked_after_kb_creation). - Dimension Enforcement: Any subsequent chunks undergo validation via
_validate_embeddings. If the returned dimension disagrees with the lockedembedding_dimension, the system aborts indexing with aDimensionMismatchError.
Qdrant Multi-Dimension Collection Isolation
Clouisle uses Qdrant as its default vector database. To support heterogeneous models concurrently, collections are partitioned by dimension:
Qdrant Cluster (http://localhost:6333)
├── kb_dim_768 (e.g., BGE-base)
├── kb_dim_1024 (e.g., BGE-large)
├── kb_dim_1536 (e.g., text-embedding-3-small)
└── kb_dim_3072 (e.g., text-embedding-3-large)Collection & Indexing Rules
- Naming Convention:
${QDRANT_COLLECTION_PREFIX}_${dimension}(default prefix iskb_dim, producingkb_dim_1536). - Distance Metric: Defaults to
Cosine(configurable via environment variables toDotorEuclid). - Automatic Provisioning: When vectors of a new dimension are first stored, the system provisions the Qdrant collection and creates payload indices on
kb_idanddocument_idfor fast scoped filtering.
Batching Specifications & Timeouts
| Configuration Parameter | Setting | Description |
|---|---|---|
| Indexing Chunk Batch Size | 25 (default) | store_chunks_with_progress embeds, stores, and updates progress in 25-chunk batches |
| Adapter Request Timeout | 60.0s | DEFAULT_EMBEDDING_REQUEST_TIMEOUT used across HTTPX and LangChain clients |
| Vector Store Operation Timeout | 120.0s | EMBEDDING_REQUEST_TIMEOUT_SECONDS governing end-to-end vector queries and bulk inserts |
| Adapter Maximum Retries | 2 | DEFAULT_EMBEDDING_MAX_RETRIES |
How is this guide?