लागू करने के लिए RAG: NNX और क्यूडमेंट के साथ सीपीयू-क्वेंटिक खोज (हिन्दी (Hindi))

लागू करने के लिए RAG: NNX और क्यूडमेंट के साथ सीपीयू-क्वेंटिक खोज

Tuesday, 25 November 2025

//

28 minute read

परिचय

RAG वस्तुओं का सबसे बड़ा भाग: यह 4a - कोर कार्यान्वयन है:

भाग 1-3 समझाता है क्यों खोज का काम करता है. यह लेख दिखाता है कैसे नींव का निर्माण करने के लिए शून्य आधार, सीपीयू- कुशल कार्यान्वयन परनिक्स रन टाइम और क्यूडमेंट का उपयोग. पार्ट 4ख खोज यूआई तथा ओपन- पीजीपी खोज कार्यान्वयन, और पार्ट 5 उत्पादन स्वचालित कवरिंग कवर.

चुनौती: अधिकांश खोज समाधानों के लिए महँगी जीयूपी को या महँगी सेवा की आवश्यकता होती है. क्या आप एक साधारण वीपीपी पर एक ब्लॉग चलाते हुए एक ब्लॉग चला रहे हैं?

समाधानः एक पूरी तरह से कार्यात्मक खोज प्रणाली जो सीपीयू पर पूरी तरह चलता है, मुक्त स्रोत औज़ारों का उपयोग करके. यह इस ब्लॉग पर पूर्ण रूप से सेटअप है मौजूदा होस्ट के पार.

कोरलेट्स

इन धारणाओं को गहराई में छिपाया जाता है RAG श्रृंखला, लेकिन यहाँ आप इस कार्यान्वयन के लिए क्या जानने की जरूरत है:

एम्बेडिंग: पाठ संख्या की तरह

एम्बेडिंग सदिश हैं जो कि कैप्चर करते हैं अर्थ पाठ का. समान अर्थ ऐसे सदिश उत्पन्‍न करते हैं - कि जादू है.

graph TD
    A["Text: 'The cat sat on the mat'"] --> B[Embedding Model]
    B --> C["Vector: [0.25, -0.18, 0.91, ... 384 more numbers]"]
    D["Text: 'A feline rested on the carpet'"] --> B
    B --> E["Vector: [0.27, -0.16, 0.89, ... similar numbers!]"]

    C -.Similar vectors = similar meaning.-> E

    style A stroke:#10b981,stroke-width:2px
    style D stroke:#10b981,stroke-width:2px
    style B stroke:#6366f1,stroke-width:3px
    style C stroke:#f59e0b,stroke-width:2px
    style E stroke:#f59e0b,stroke-width:2px

मुख्य अन्तर्दृष्टि: समान अर्थों के साथ पाठ समान सदिश (एमबिंग) होगा. यह हम कैसे पा सकते हैं "से जुड़े" सामग्री - हम वास्तव में अर्थों के बीच दूरी माप रहे हैं!

बदकारी को समझना

कोसाइन दो सदिशों के बीच कोण मापो - यदि वे समान दिशाओं में बात करते हैं, वे समान रूप से समान हैं:

flowchart LR
    subgraph "Vector Space (simplified to 2D)"
        direction TB
        A["'Docker tutorial'"] -.-> B((0.85))
        C["'Container deployment'"] -.-> B
        D["'Cooking recipes'"] -.-> E((0.12))
        A -.-> E
    end

    B --> F["High Similarity<br/>Related content!"]
    E --> G["Low Similarity<br/>Different topics"]

    style A stroke:#10b981,stroke-width:2px
    style C stroke:#10b981,stroke-width:2px
    style D stroke:#f59e0b,stroke-width:2px
    style B stroke:#22c55e,stroke-width:3px
    style E stroke:#ef4444,stroke-width:3px
    style F stroke:#22c55e,stroke-width:2px
    style G stroke:#ef4444,stroke-width:2px

सूत्र similarity = (A · B) / (||A|| × ||B||) - लेकिन जब से हम L2-सामान्य हमारे सदिश का पालन करते हैं, यह सिर्फ बिंदु उत्पाद के लिए सरल है!

क्या पर है?

ऑननेटएक्स (ओस्टल नेटवर्क एक्सचेंज) मशीन सीखने वाले मॉडलों के लिए एक खुला मानक फॉर्मेट है जो उन्हें विभिन्न मंचों पर आसानी से चलाने की अनुमति देता है. इसके बारे में सोचो एआई मॉडलों के लिए एक विश्वव्यापी अनुवादक. ऑन NNX रन समय Microsoft की उच्च क्षमता इंजन है कि इन मॉडलों को चलाने के लिए।

हमारे मामले में क्यों नैट:

flowchart LR
    subgraph "ONNX Inference Pipeline"
        A[Raw Text] --> B[Tokenizer]
        B --> C["Tokens: [CLS] the cat sat [SEP]"]
        C --> D[Token IDs: 101 1996 4937 2068 102]
        D --> E[ONNX Runtime]
        E --> F[384-dim Vector]
        F --> G[L2 Normalize]
        G --> H[Final Embedding]
    end

    style A stroke:#10b981,stroke-width:2px
    style B stroke:#f59e0b,stroke-width:2px
    style C stroke:#f59e0b,stroke-width:2px
    style D stroke:#f59e0b,stroke-width:2px
    style E stroke:#6366f1,stroke-width:3px
    style F stroke:#8b5cf6,stroke-width:2px
    style G stroke:#8b5cf6,stroke-width:2px
    style H stroke:#ef4444,stroke-width:2px

क्वीप क्या है?

स्केल्स एक ओपन-source वेक्टर डाटाबेस - मूल रूप से एक डाटाबेस भंडारित किया गया है इन एम्बेडेड सदिशों को भंडारित करने और खोज करने के लिए. एक गहरी बिट के लिए, विन्यास, और C# एकीकरण, देखें. क्यूवर के साथ स्व-Hed सदिश डाटाबेस. जबकि आप कर सकता है एसक्यूएल में वेक्टर संग्रह इस तथा प्रस्तुत करने के लिए उद्देश्य है:

flowchart TB
    subgraph "Qdrant Vector Storage"
        direction TB
        A[Collection: blog_posts] --> B[Point 1]
        A --> C[Point 2]
        A --> D[Point N...]

        B --> B1["Vector: [0.12, -0.08, ...]"]
        B --> B2["Payload: {slug, title, language}"]

        C --> C1["Vector: [0.25, 0.14, ...]"]
        C --> C2["Payload: {slug, title, language}"]
    end

    subgraph "Vector Search"
        E[Query Vector] --> F[HNSW Index]
        F --> G[Cosine Similarity]
        G --> H[Top K Results]
    end

    style A stroke:#ef4444,stroke-width:3px
    style B stroke:#8b5cf6,stroke-width:2px
    style C stroke:#8b5cf6,stroke-width:2px
    style D stroke:#8b5cf6,stroke-width:2px
    style B1 stroke:#f59e0b,stroke-width:2px
    style B2 stroke:#10b981,stroke-width:2px
    style C1 stroke:#f59e0b,stroke-width:2px
    style C2 stroke:#10b981,stroke-width:2px
    style E stroke:#6366f1,stroke-width:2px
    style F stroke:#ec4899,stroke-width:3px
    style G stroke:#ec4899,stroke-width:2px
    style H stroke:#10b981,stroke-width:2px

ओवरव्यू

यहाँ हमारा खोज प्रणाली एक साथ फिट कैसे है:

flowchart TB
    subgraph "Content Ingestion"
        A[Blog Post Markdown] --> B[Extract Plain Text]
        B --> C[ONNX Embedding Service]
        C --> D[Generate 384-dim Vector]
        D --> E[Qdrant Vector Store]
    end

    subgraph "Search Flow"
        F[User Query] --> G[ONNX Embedding Service]
        G --> H[Generate Query Vector]
        H --> I[Qdrant Search]
        E -.Vector Similarity.-> I
        I --> J[Ranked Results]
    end

    subgraph "Related Posts"
        K[Current Blog Post] --> L[Get Post Vector from Qdrant]
        L --> M[Find Similar Vectors]
        E -.->M
        M --> N[Top 5 Related Posts]
    end

    style A stroke:#10b981,stroke-width:2px
    style B stroke:#10b981,stroke-width:2px
    style C stroke:#6366f1,stroke-width:3px
    style D stroke:#f59e0b,stroke-width:2px
    style E stroke:#ef4444,stroke-width:3px
    style F stroke:#10b981,stroke-width:2px
    style G stroke:#6366f1,stroke-width:3px
    style H stroke:#f59e0b,stroke-width:2px
    style I stroke:#ef4444,stroke-width:2px
    style J stroke:#8b5cf6,stroke-width:2px
    style K stroke:#10b981,stroke-width:2px
    style L stroke:#ef4444,stroke-width:2px
    style M stroke:#ef4444,stroke-width:2px
    style N stroke:#8b5cf6,stroke-width:2px

सादा अंग्रेज़ी में प्रवाह:

  1. सूचीकरण: जब आप ब्लॉग पोस्ट लिखते हैं, तो हम इसे एक सदिश में बदल देते हैं और इसे क्यूवर में जमा करते हैं
  2. खोज रहा है: जब कोई जाँच करता है, हम उनके प्रश्न को एक सदिश में बदल देते हैं और क्यूवर में समान सदिश खोज पाते हैं
  3. संबंधित पोस्ट: किसी भी ब्लॉग पोस्ट के लिए, हम समान सदिशों के साथ अन्य पोस्टों को पा सकते हैं

परियोजना स्ट्रक्चर

हमने एक साफ, निचली संरचना बनाई है:

Mostlylucid.SemanticSearch/
├── Config/
│   └── SemanticSearchConfig.cs      # Configuration settings
├── Models/
│   ├── BlogPostDocument.cs          # Document model for indexing
│   └── SearchResult.cs               # Search result model
├── Services/
│   ├── IEmbeddingService.cs         # Embedding interface
│   ├── OnnxEmbeddingService.cs      # ONNX-based embeddings
│   ├── IVectorStoreService.cs       # Vector store interface
│   ├── QdrantVectorStoreService.cs  # Qdrant implementation
│   ├── ISemanticSearchService.cs    # High-level search interface
│   └── SemanticSearchService.cs     # Orchestration service
├── Extensions/
│   └── ServiceCollectionExtensions.cs  # DI registration
├── download-models.sh               # Model download script
└── README.md

कार्यान्वयन

चरण 1: परियोजना ऊपर विन्यास

प्रथम, नई क्लास लाइब्रेरी बनाएँ:

dotnet new classlib -n Mostlylucid.SemanticSearch -f net9.0
dotnet sln add Mostlylucid.SemanticSearch

आवश्यक निग पैकेज जोड़ें:

cd Mostlylucid.SemanticSearch
dotnet add package Microsoft.Extensions.Logging.Abstractions
dotnet add package Microsoft.ML.OnnxRuntime --version 1.21.1
dotnet add package Qdrant.Client --version 1.14.0
dotnet add reference ../Mostlylucid.Shared/Mostlylucid.Shared.csproj

चरण 2: कॉन्फ़िगरेशन

हम उपयोग कर रहे हैं अपने विन्यास क्लास सेट. IConfigSection पैटर्न है कि सबसे अधिक से कम इस्तेमाल किया गया है:

using Mostlylucid.Shared.Config;

namespace Mostlylucid.SemanticSearch.Config;

/// <summary>
/// Configuration for semantic search functionality
/// </summary>

public class SemanticSearchConfig : IConfigSection
{
    public static string Section => "SemanticSearch";

    /// <summary>
    /// Enable or disable semantic search
    /// </summary>

    public bool Enabled { get; set; } = true;

    /// <summary>
    /// Qdrant server URL (e.g., http://localhost:6333)
    /// </summary>

    public string QdrantUrl { get; set; } = "http://localhost:6333";

    /// <summary>
    /// Optional read-only API key for Qdrant (used for search operations)
    /// </summary>

    public string? ReadApiKey { get; set; }

    /// <summary>
    /// Optional read-write API key for Qdrant (used for indexing operations)
    /// </summary>

    public string? WriteApiKey { get; set; }

    /// <summary>
    /// Collection name in Qdrant for blog posts
    /// </summary>

    public string CollectionName { get; set; } = "blog_posts";

    /// <summary>
    /// Path to the ONNX embedding model file
    /// </summary>

    public string EmbeddingModelPath { get; set; } = "models/all-MiniLM-L6-v2.onnx";

    /// <summary>
    /// Path to the tokenizer vocabulary file
    /// </summary>

    public string VocabPath { get; set; } = "models/vocab.txt";

    /// <summary>
    /// Embedding vector size (384 for all-MiniLM-L6-v2)
    /// </summary>

    public int VectorSize { get; set; } = 384;

    /// <summary>
    /// Number of related posts to return
    /// </summary>

    public int RelatedPostsCount { get; set; } = 5;

    /// <summary>
    /// Minimum similarity score (0-1) for related posts
    /// </summary>

    public float MinimumSimilarityScore { get; set; } = 0.5f;

    /// <summary>
    /// Number of search results to return
    /// </summary>

    public int SearchResultsCount { get; set; } = 10;
}

क्यों अलग हाथों की कुंजियाँ? सुरक्षा! आपकी पढ़ने की कुंजी सार्वजनिक खोज अंत बिन्दुओं में उपयोग में ली जा सकती है, जबकि आपका लेखन कुंजी केवल प्रशासक ऑपरेशन के लिए सर्वर- साइड पर आधारित है.

इसे आपके पास जोड़ें appsettings.json:

{
  "SemanticSearch": {
    "Enabled": false,
    "QdrantUrl": "http://localhost:6333",
    "ReadApiKey": "",
    "WriteApiKey": "",
    "CollectionName": "blog_posts",
    "EmbeddingModelPath": "models/all-MiniLM-L6-v2.onnx",
    "VocabPath": "models/vocab.txt",
    "VectorSize": 384,
    "RelatedPostsCount": 5,
    "MinimumSimilarityScore": 0.5,
    "SearchResultsCount": 10
  }
}

कदम 3: ऊपर दिया गया सेवा

यह जादू जहां होता है। हम सभी एमीनीएम-L6-v2 मॉडल का उपयोग कर रहे हैं, जो विशेष रूप से इमानिक समानता कार्य के लिए तैयार किया जाता है और सीपीयू पर तेजी से चलता है।

यह आदर्श क्यों?

  • छोटा आकार (0-290MBMB)
  • सीपीयू पर तेज अंतराल (0-250-100ms प्रति एम्बेडिंग)
  • अच्छा विशेषता एम्बेडिंग (384 आयाम)
  • 1 अरब से भी ज़्यादा आदमियों को तालीम दी गयी

यहाँ पूरा कार्यान्वयन है:

using Microsoft.Extensions.Logging;
using Microsoft.ML.OnnxRuntime;
using Microsoft.ML.OnnxRuntime.Tensors;
using Mostlylucid.SemanticSearch.Config;
using System.Text.RegularExpressions;

namespace Mostlylucid.SemanticSearch.Services;

public class OnnxEmbeddingService : IEmbeddingService, IDisposable
{
    private readonly ILogger<OnnxEmbeddingService> _logger;
    private readonly SemanticSearchConfig _config;
    private readonly InferenceSession? _session;
    private readonly Dictionary<string, int> _vocabulary;
    private readonly SemaphoreSlim _semaphore = new(1, 1);
    private bool _disposed;

    private const int MaxSequenceLength = 256;
    private const string PadToken = "[PAD]";
    private const string UnkToken = "[UNK]";
    private const string ClsToken = "[CLS]";
    private const string SepToken = "[SEP]";

    public OnnxEmbeddingService(
        ILogger<OnnxEmbeddingService> logger,
        SemanticSearchConfig config)
    {
        _logger = logger;
        _config = config;
        _vocabulary = new Dictionary<string, int>();

        if (!_config.Enabled)
        {
            _logger.LogInformation("Semantic search is disabled");
            return;
        }

        try
        {
            // Check if model file exists
            if (!File.Exists(_config.EmbeddingModelPath))
            {
                _logger.LogWarning("Embedding model not found at {Path}. Semantic search will be disabled.",
                    _config.EmbeddingModelPath);
                return;
            }

            // Load vocabulary if it exists
            if (File.Exists(_config.VocabPath))
            {
                LoadVocabulary(_config.VocabPath);
            }

            // Create ONNX session with CPU execution provider
            var sessionOptions = new SessionOptions
            {
                ExecutionMode = ExecutionMode.ORT_SEQUENTIAL,
                GraphOptimizationLevel = GraphOptimizationLevel.ORT_ENABLE_ALL
            };

            _session = new InferenceSession(_config.EmbeddingModelPath, sessionOptions);
            _logger.LogInformation("ONNX embedding model loaded successfully from {Path}",
                _config.EmbeddingModelPath);
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Failed to initialize ONNX embedding service");
        }
    }

    private void LoadVocabulary(string vocabPath)
    {
        var lines = File.ReadAllLines(vocabPath);
        for (int i = 0; i < lines.Length; i++)
        {
            var token = lines[i].Trim();
            if (!string.IsNullOrEmpty(token))
            {
                _vocabulary[token] = i;
            }
        }
        _logger.LogInformation("Loaded vocabulary with {Count} tokens", _vocabulary.Count);
    }

    public async Task<float[]> GenerateEmbeddingAsync(string text, CancellationToken cancellationToken = default)
    {
        if (_session == null || !_config.Enabled)
        {
            return new float[_config.VectorSize];
        }

        if (string.IsNullOrWhiteSpace(text))
        {
            return new float[_config.VectorSize];
        }

        // Use semaphore to prevent concurrent ONNX inference (not thread-safe)
        await _semaphore.WaitAsync(cancellationToken);
        try
        {
            return await Task.Run(() => GenerateEmbedding(text), cancellationToken);
        }
        finally
        {
            _semaphore.Release();
        }
    }

    private float[] GenerateEmbedding(string text)
    {
        try
        {
            // Tokenize the input text
            var tokens = Tokenize(text);

            // Create input tensors for ONNX model
            var inputIds = CreateInputTensor(tokens, "input_ids");
            var attentionMask = CreateAttentionMaskTensor(tokens.Length);
            var tokenTypeIds = CreateTokenTypeIdsTensor(tokens.Length);

            // Run inference
            var inputs = new List<NamedOnnxValue>
            {
                NamedOnnxValue.CreateFromTensor("input_ids", inputIds),
                NamedOnnxValue.CreateFromTensor("attention_mask", attentionMask),
                NamedOnnxValue.CreateFromTensor("token_type_ids", tokenTypeIds)
            };

            using var results = _session!.Run(inputs);

            // Extract the output tensor (sentence embedding)
            var output = results.First().AsTensor<float>();
            var embedding = output.ToArray();

            // Normalize the vector (L2 normalization)
            return NormalizeVector(embedding);
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Error generating embedding for text: {Text}",
                text[..Math.Min(100, text.Length)]);
            return new float[_config.VectorSize];
        }
    }

    private List<int> Tokenize(string text)
    {
        // Simple whitespace + punctuation tokenization
        var tokens = new List<int>();

        // Add [CLS] token at the start
        if (_vocabulary.TryGetValue(ClsToken, out var clsId))
            tokens.Add(clsId);

        // Tokenize the text
        var words = Regex.Split(text.ToLowerInvariant(), @"(\W+)")
            .Where(w => !string.IsNullOrWhiteSpace(w))
            .Take(MaxSequenceLength - 2); // Leave room for [CLS] and [SEP]

        foreach (var word in words)
        {
            if (_vocabulary.Count > 0)
            {
                if (_vocabulary.TryGetValue(word, out var tokenId))
                    tokens.Add(tokenId);
                else if (_vocabulary.TryGetValue(UnkToken, out var unkId))
                    tokens.Add(unkId);
            }
            else
            {
                // Fallback: use hash code as token ID
                tokens.Add(Math.Abs(word.GetHashCode()) % 30000);
            }
        }

        // Add [SEP] token at the end
        if (_vocabulary.TryGetValue(SepToken, out var sepId))
            tokens.Add(sepId);

        return tokens;
    }

    private Tensor<long> CreateInputTensor(List<int> tokens, string name)
    {
        var length = Math.Min(tokens.Count, MaxSequenceLength);
        var tensorData = new long[1, MaxSequenceLength];

        for (int i = 0; i < length; i++)
        {
            tensorData[0, i] = tokens[i];
        }

        // Pad the rest
        var padId = _vocabulary.TryGetValue(PadToken, out var id) ? id : 0;
        for (int i = length; i < MaxSequenceLength; i++)
        {
            tensorData[0, i] = padId;
        }

        return new DenseTensor<long>(tensorData, new[] { 1, MaxSequenceLength });
    }

    private Tensor<long> CreateAttentionMaskTensor(int actualLength)
    {
        var length = Math.Min(actualLength, MaxSequenceLength);
        var tensorData = new long[1, MaxSequenceLength];

        for (int i = 0; i < length; i++)
        {
            tensorData[0, i] = 1; // Attend to actual tokens
        }

        return new DenseTensor<long>(tensorData, new[] { 1, MaxSequenceLength });
    }

    private Tensor<long> CreateTokenTypeIdsTensor(int actualLength)
    {
        var tensorData = new long[1, MaxSequenceLength];
        // All zeros for single sentence
        return new DenseTensor<long>(tensorData, new[] { 1, MaxSequenceLength });
    }

    private float[] NormalizeVector(float[] vector)
    {
        // L2 normalization
        var sumOfSquares = vector.Sum(v => v * v);
        var magnitude = MathF.Sqrt(sumOfSquares);

        if (magnitude > 0)
        {
            for (int i = 0; i < vector.Length; i++)
            {
                vector[i] /= magnitude;
            }
        }

        return vector;
    }

    public void Dispose()
    {
        if (_disposed) return;

        _session?.Dispose();
        _semaphore?.Dispose();
        _disposed = true;

        GC.SuppressFinalize(this);
    }
}

जुटेसी डेव्स के लिए कुंजी पाइंट्स:

  1. टोकनीकरण: हम पाठ को छोटे टुकड़ों में तोड़ रहे हैं (kss) मॉडल समझ सकता है कि
  2. टेर्सर्स: ये बहुत-से बहु-विष्टि हैं कि NNX मॉडलों के साथ काम कर रहे हैं
  3. ध्यान मास्क: पैटर्न बताता है कि इनपुट के कौन से भाग वास्तविक सामग्री v हैं. पैडिंग
  4. एल2 सामान्य करें: सभी सदिशों को एक ही "सेट" है, तो हम उन्हें पूरी तरह से तुलना कर सकते हैं
  5. केरोफ़ॉयर: सुनिश्चित करें कि थ्रेड सुरक्षा (एनएनएक्स तयशुदा रूप से सुरक्षित नहीं है)

चरण 4: प्रवर वेक्टर भंडार

अब चलो सदिश भंडारण और खोज को लागू करें:

using Microsoft.Extensions.Logging;
using Mostlylucid.SemanticSearch.Config;
using Mostlylucid.SemanticSearch.Models;
using Qdrant.Client;
using Qdrant.Client.Grpc;

namespace Mostlylucid.SemanticSearch.Services;

public class QdrantVectorStoreService : IVectorStoreService
{
    private readonly ILogger<QdrantVectorStoreService> _logger;
    private readonly SemanticSearchConfig _config;
    private readonly QdrantClient? _client;
    private bool _collectionInitialized;

    public QdrantVectorStoreService(
        ILogger<QdrantVectorStoreService> logger,
        SemanticSearchConfig config)
    {
        _logger = logger;
        _config = config;

        if (!_config.Enabled)
        {
            _logger.LogInformation("Semantic search is disabled");
            return;
        }

        try
        {
            var uri = new Uri(_config.QdrantUrl);
            var host = uri.Host;
            var port = uri.Port > 0 ? uri.Port : 6334; // Default gRPC port

            _client = new QdrantClient(host, port, https: uri.Scheme == "https");
            _logger.LogInformation("Connected to Qdrant at {Host}:{Port}", host, port);
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Failed to connect to Qdrant at {Url}", _config.QdrantUrl);
        }
    }

    public async Task InitializeCollectionAsync(CancellationToken cancellationToken = default)
    {
        if (_client == null || !_config.Enabled || _collectionInitialized)
            return;

        try
        {
            var collections = await _client.ListCollectionsAsync(cancellationToken);
            var collectionExists = collections.Any(c => c.Name == _config.CollectionName);

            if (!collectionExists)
            {
                _logger.LogInformation("Creating collection {CollectionName}", _config.CollectionName);

                await _client.CreateCollectionAsync(
                    collectionName: _config.CollectionName,
                    vectorsConfig: new VectorParams
                    {
                        Size = (ulong)_config.VectorSize,
                        Distance = Distance.Cosine // Cosine similarity for semantic search
                    },
                    cancellationToken: cancellationToken
                );

                _logger.LogInformation("Collection {CollectionName} created successfully", _config.CollectionName);
            }

            _collectionInitialized = true;
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Failed to initialize collection {CollectionName}", _config.CollectionName);
            throw;
        }
    }

    public async Task<List<SearchResult>> FindRelatedPostsAsync(
        string slug,
        string language,
        int limit = 5,
        CancellationToken cancellationToken = default)
    {
        if (_client == null || !_config.Enabled)
            return new List<SearchResult>();

        try
        {
            // Find the document by slug and language
            var scrollResults = await _client.ScrollAsync(
                collectionName: _config.CollectionName,
                filter: new Filter
                {
                    Must =
                    {
                        new Condition
                        {
                            Field = new FieldCondition
                            {
                                Key = "slug",
                                Match = new Match { Keyword = slug }
                            }
                        },
                        new Condition
                        {
                            Field = new FieldCondition
                            {
                                Key = "language",
                                Match = new Match { Keyword = language }
                            }
                        }
                    }
                },
                limit: 1,
                cancellationToken: cancellationToken
            );

            var point = scrollResults.FirstOrDefault();
            if (point == null)
            {
                _logger.LogWarning("Post {Slug} ({Language}) not found in vector store", slug, language);
                return new List<SearchResult>();
            }

            // Use the document's vector to find similar posts
            var searchResults = await _client.SearchAsync(
                collectionName: _config.CollectionName,
                vector: point.Vectors.Vector.Data.ToArray(),
                limit: (ulong)(limit + 1), // +1 because the first result will be the post itself
                scoreThreshold: _config.MinimumSimilarityScore,
                cancellationToken: cancellationToken
            );

            // Filter out the original post and return top N similar posts
            return searchResults
                .Where(r => r.Payload["slug"].StringValue != slug || r.Payload["language"].StringValue != language)
                .Take(limit)
                .Select(result => new SearchResult
                {
                    Slug = result.Payload["slug"].StringValue,
                    Title = result.Payload["title"].StringValue,
                    Language = result.Payload["language"].StringValue,
                    Categories = result.Payload.TryGetValue("categories", out var cats)
                        ? cats.ListValue.Values.Select(v => v.StringValue).ToList()
                        : new List<string>(),
                    Score = result.Score,
                    PublishedDate = DateTime.Parse(result.Payload["published_date"].StringValue)
                })
                .ToList();
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Failed to find related posts for {Slug} ({Language})", slug, language);
            return new List<SearchResult>();
        }
    }

    // ... Additional methods for IndexDocument, Search, Delete, etc.
}

यहाँ क्या हो रहा है:

  1. कोसाइन दूरी: हम कोज्या के उपयोग से कर रहे हैं, जो सामान्य सदिशों की तुलना करने के लिए सही है
  2. मेटाडाटा भंडारण: कवरेज से हमें अतिरिक्त डाटा भंडारित करने देता है (लोडिंग) साथ सदिश के साथ
  3. फ़िल्टर किया जा रहा है: सदिशों की तुलना करने से पहले हम मेटाडाटा द्वारा परिणाम फ़िल्टर कर सकते हैं
  4. स्कोर दहलीज: केवल वापसी परिणाम किसी विशेष समानता अंक के ऊपर

कदम 5: बदलते हालात के मुताबिक सेवा

यह उच्च स्तर सेवा सेवा सभी एक साथ संलग्न है:

using Microsoft.Extensions.Logging;
using Mostlylucid.SemanticSearch.Config;
using Mostlylucid.SemanticSearch.Models;
using System.Security.Cryptography;
using System.Text;

namespace Mostlylucid.SemanticSearch.Services;

public class SemanticSearchService : ISemanticSearchService
{
    private readonly ILogger<SemanticSearchService> _logger;
    private readonly SemanticSearchConfig _config;
    private readonly IEmbeddingService _embeddingService;
    private readonly IVectorStoreService _vectorStoreService;

    public SemanticSearchService(
        ILogger<SemanticSearchService> logger,
        SemanticSearchConfig config,
        IEmbeddingService embeddingService,
        IVectorStoreService vectorStoreService)
    {
        _logger = logger;
        _config = config;
        _embeddingService = embeddingService;
        _vectorStoreService = vectorStoreService;
    }

    public async Task IndexPostAsync(BlogPostDocument document, CancellationToken cancellationToken = default)
    {
        if (!_config.Enabled)
            return;

        try
        {
            // Prepare text for embedding: combine title and content
            // We give more weight to the title by including it twice
            var textToEmbed = $"{document.Title}. {document.Title}. {document.Content}";

            // Truncate to reasonable length (embedding models have token limits)
            const int maxLength = 2000;
            if (textToEmbed.Length > maxLength)
            {
                textToEmbed = textToEmbed[..maxLength];
            }

            // Generate embedding
            var embedding = await _embeddingService.GenerateEmbeddingAsync(textToEmbed, cancellationToken);

            // Compute content hash if not provided
            if (string.IsNullOrEmpty(document.ContentHash))
            {
                document.ContentHash = ComputeContentHash(document.Content);
            }

            // Store in vector database
            await _vectorStoreService.IndexDocumentAsync(document, embedding, cancellationToken);

            _logger.LogInformation("Indexed post {Slug} ({Language})", document.Slug, document.Language);
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Failed to index post {Slug} ({Language})", document.Slug, document.Language);
        }
    }

    public async Task<List<SearchResult>> SearchAsync(
        string query,
        int limit = 10,
        CancellationToken cancellationToken = default)
    {
        if (!_config.Enabled || string.IsNullOrWhiteSpace(query))
            return new List<SearchResult>();

        try
        {
            // Generate embedding for the search query
            var queryEmbedding = await _embeddingService.GenerateEmbeddingAsync(query, cancellationToken);

            // Search in vector store
            var results = await _vectorStoreService.SearchAsync(
                queryEmbedding,
                Math.Min(limit, _conken);

            _logger.LogDebug("Search for '{Query}' returned {Count} results", query, results.Count);

            return results;
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Search failed for query '{Query}'", query);
            return new List<SearchResult>();
        }
    }

    public async Task<List<SearchResult>> GetRelatedPostsAsync(
        string slug,
        string language,
        int limit = 5,
        CancellationToken cancellationToken = default)
    {
        if (!_config.Enabled)
            return new List<SearchResult>();

        try
        {
            var results = await _vectorStoreService.FindRelatedPostsAsync(
                slug,
                language,
                Math.Min(limit, _config.RelatedPostsCount),
                cancellationToken);

            _logger.LogDebug("Found {Count} related posts for {Slug} ({Language})",
                results.Count, slug, language);

            return results;
        }
        catch (Exception ex)
        {
            _logger.LogError(ex, "Failed to get related posts for {Slug} ({Language})", slug, language);
            return new List<SearchResult>();
        }
    }

    private string ComputeContentHash(string content)
    {
        using var sha256 = SHA256.Create();
        var bytes = Encoding.UTF8.GetBytes(content);
        var hashBytes = sha256.ComputeHash(bytes);
        return Convert.ToBase64String(hashBytes);
    }
}

चरण 6: डिपेंडेंसी इनसन सेटअप

थण्डॉर्न में सब कुछ पंजीकृत करें:

using Microsoft.Extensions.Configuration;
using Microsoft.Extensions.DependencyInjection;
using Mostlylucid.SemanticSearch.Config;
using Mostlylucid.SemanticSearch.Services;
using Mostlylucid.Shared.Config;

namespace Mostlylucid.SemanticSearch.Extensions;

public static class ServiceCollectionExtensions
{
    public static void AddSemanticSearch(
        this IServiceCollection services,
        IConfiguration configuration)
    {
        // Bind configuration using POCO pattern
        services.ConfigurePOCO<SemanticSearchConfig>(
            configuration.GetSection(SemanticSearchConfig.Section));

        // Register services as singletons for efficiency
        services.AddSingleton<IEmbeddingService, OnnxEmbeddingService>();
        services.AddSingleton<IVectorStoreService, QdrantVectorStoreService>();
        services.AddSingleton<ISemanticSearchService, SemanticSearchService>();
    }
}

अपने में Program.cs:

using Mostlylucid.SemanticSearch.Extensions;
using Mostlylucid.SemanticSearch.Services;

// Add services
services.AddSemanticSearch(config);

// Initialize after building the app
using (var scope = app.Services.CreateScope())
{
    var semanticSearch = scope.ServiceProvider.GetRequiredService<ISemanticSearchService>();
    await semanticSearch.InitializeAsync();
}

आई. वी.

क्यू चीज़ के लिए डॉकर बनाएं

खोज सेवाों के लिए अलग- अलग डॉकe- क़िस्म की खोज सेवाएँ बनाएं:

version: '3.8'

services:
  qdrant:
    image: qdrant/qdrant:latest
    container_name: mostlylucid-qdrant
    restart: unless-stopped
    ports:
      - "6333:6333"  # HTTP API
      - "6334:6334"  # gRPC API
    volumes:
      - qdrant_storage:/qdrant/storage
    environment:
      - QDRANT__SERVICE__HTTP_PORT=6333
      - QDRANT__SERVICE__GRPC_PORT=6334
    networks:
      - mostlylucid_network
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:6333/health"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 40s

volumes:
  qdrant_storage:
    driver: local

networks:
  mostlylucid_network:
    name: mostlylucidweb_app_network
    external: true

इसे इसके साथ प्रारंभ करें:

docker-compose -f semantic-search-docker-compose.yml up -d

एम्बेडिंग मॉडल डाउनलोड करें

हम इस्तेमाल कर रहे हैं सभी मारनीLM-L6-v2 एचडिफ्ट मुख्स से मॉडल वाक्य रूपांतरण लाइब्रेरी. इस मॉडल को विशेष रूप से प्रोटेस्टेंट समानता कार्यों पर प्रशिक्षित किया जाता है और 384-विधन पैदा करता है.

स्वचालित डाउनलोड (पुष्टिकारी)

सेवा स्वतः डाउनलोड करने के पहले चालू स्क्रीन से मॉडल को डाउनलोड कर रहा है यदि यह मौजूद नहीं है:

// In OnnxEmbeddingService.cs
private const string ModelUrl = "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/onnx/model.onnx";
private const string VocabUrl = "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/vocab.txt";

public async Task EnsureInitializedAsync(CancellationToken cancellationToken = default)
{
    if (_initialized || !_config.Enabled) return;

    // Download model if not exists
    if (!File.Exists(_config.EmbeddingModelPath))
    {
        _logger.LogInformation("Downloading ONNX embedding model to {Path}...", _config.EmbeddingModelPath);
        await DownloadFileAsync(ModelUrl, _config.EmbeddingModelPath, cancellationToken);
    }

    // Download vocab if not exists
    if (!File.Exists(_config.VocabPath))
    {
        _logger.LogInformation("Downloading vocabulary file to {Path}...", _config.VocabPath);
        await DownloadFileAsync(VocabUrl, _config.VocabPath, cancellationToken);
    }

    // Initialize ONNX session...
}

यह विशेष रूप से उपयोगी है जब डॉकयर के साथ तैनात किया जा रहा है - आप मॉडल डिरेक्ट्री के लिए एक वॉल्यूम मैप कर सकते हैं:

volumes:
  - ./mlmodels:/app/mlmodels  # Model persists across container restarts

हस्तचालित डाउनलोड

वैकल्पिक रूप से, आप दस्ती रूप से डाउनलोड कर सकते हैं:

chmod +x Mostlylucid.SemanticSearch/download-models.sh
./Mostlylucid.SemanticSearch/download-models.sh

या सीधे एचडिस्क चेहरे से:

mkdir -p mlmodels
curl -L https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/onnx/model.onnx -o mlmodels/all-MiniLM-L6-v2.onnx
curl -L https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/vocab.txt -o mlmodels/vocab.txt

इस डाउनलोड्स

परफ़ॉर्मेंस पर ध्यान दें

अंतर्निर्मित पीढ़ी

  • सीपीयू परफ़ॉर्मेंसComment: ~50-100 ms प्रति आधुनिक सीपीयू पर एम्बेड किया जा रहा है
  • ऑप्टीमाइज़ेशन: हम एक मौजूदा स्तर पर रोक लगाने के लिए एक समोग्रोक्स का उपयोग करते हैं
  • बैचिंग: बड़ी सूची के लिए, प्रक्रिया पोस्टिंग 10-20 के बैचेट में

सदिश खोज

  • ढूंढें गति: < 10ms के लिए 100के वेक्टर में संग्रह के लिए
  • मेमोरी उपयोग: ~1के प्रति सदिश ( मेटाडाटा के साथ)
  • स्केल्स: प्रवरेज, साधारण हार्डवेयर पर लाखों सदिशों को संभाल सकता है

कैश प्लान किया जा रहा है

हम अप्रयोगक का उपयोग करते हैं.

[OutputCache(Duration = 7200, VaryByRouteValueNames = new[] {"slug", "language"})]

इस कैश्स में 2 घंटे से संबंधित पोस्टों के बारे में जानकारी दी गयी है, लोड बहुत ही कम किया जा रहा है.

हमने क्या बनाया है

इस बिंदु पर आप एक पूरा है, खोज आधार काम कर रहा है:

  • ✅ ऑन- लाइनिंग - स्विंगिंग मुख से सीपीयू मित्रता
  • ✅ बैच वेक्टर भंडारण - मेटाडाटा फिल्टरों के साथ तेज सी स्थिति खोज
  • ✅ सम्बन्धित पोस्ट्स - ठीक इसी तरह की सामग्री ढूंढें
  • ✅ ढूंढें एपीआई - प्राकृतिक भाषा स्वाभाविक है
  • ✅ विषयवस्तु बॉक्सिंग - ब्लॉग पोस्ट को सदिश के रूप में भंडारित करें

यह इस ब्लॉग पर पूर्ण सेटअप है - शून्य जीपी, शून्य अतिरिक्त लागत.

अगला: कार्य में टिकिक खोज

में भाग ४ख: कार्य में सेप्टिक खोज, हम कवर:

  • टाइप हेड खोज - खोज के रूप में कैसे काम करता है आप-प्रकार As.js के साथ काम करता है
  • एचब्रिएड खोजQuery - उपभोग किया जा रहा है
  • ढूंढें एपीआई - फिल्टरों के साथ पूर्ण एपीआई दस्तावेज
  • सम्बन्धित पोस्ट यूआई - HMASX आलसी लोड के साथ घटक
  • विस्तृत फ़िल्टर्स - भाषा और तिथि सीमा फिल्टरिंग

यहाँ जारी रखें पार्ट 4ख खोज यूआई और ओपन- पीजीपी खोज कार्यान्वयन के लिए.

तब पार्ट 5: Hybd खोज स्वचलित किया जा रहा है बनाने का तरीका अलग - अलग होता है ।

संसाधन

ऑननेटएक्स दस्तावेज़ीकरण

स्ट्रिगी दस्तावेज़ीकरण

मॉडलों को एम्बेड करना

कोड पूरा करें

सभी कोड इस पर उपलब्ध: gdyb.com/sek/ allyptib

  • Mostlylucid.SemanticSearch/ - कोरलिक खोज लाइब्रेरी
Finding related posts...
logo

© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.