Back to "DocSummarizer भाग 4 - Building RAG Pipelines"

This is a viewer only at the moment see the article on how this works.

To update the preview hit Ctrl-Alt-R (or ⌘-Alt-R on Mac) or Enter to refresh. The Save icon lets you save the markdown file to disk

This is a preview from the server running through my markdig pipeline

AI C# Embeddings LLM ONNX RAG Semantic Search

DocSummarizer भाग 4 - Building RAG Pipelines

Tuesday, 30 December 2025

न्यूगेटGenericName एनबीएम .NET नोड

यह है भाग 4 DocSummarizer श्रृंखला के. देखें भाग 1 वास्तुकला के लिए भाग 2 के लिए CLI उपकरण, या भाग 3 embeddings पर गहरी डुबो के लिए

RAG का हार्ड भाग LLM के पहले सब कुछ है

आपने शायद इस पैटर्न को देखा है।

  • प्रत्येक प्रारूप के लिए दस्तावेज़ पदवर्णक
  • अर्थात्मक सीमाओं को मानने वाला तर्क
  • टोकनीकरण जो आपके एम्बेडिंग मॉडल से मिल जाता है
  • प्रभावी एम्बेडिंग उत्पादन के लिए बैचिंग
  • सलिंस स्कोरिंग इसलिए सभी टुकड़े समान नहीं किया जाता है
  • उद्धरण ट्रैकिंग ताकि आप जानते हों कि जवाब कहाँ से आये थे

आप अनुप्रयोग कोड के एक ही लाइन लिखने से पहले यह एक बहुत बुनियादी ढांचा है

DocSummarizer.Core यह सब एक ही पैकेज में संभालता है दोनों के लिए उपलब्ध .NET और नोड

DocSummarizer.Core अनिवार्य रूप से एक है दस्तावेज़ आसूचना परतयह सूचनाओं को हल करता है

इंजेक्शन पाइपलाइन

यहाँ DocSummarizer क्या करता है - और महत्वपूर्ण रूप से , यह क्या नहीं करता do:

flowchart TB
    subgraph INPUT["Input (Your Document)"]
        DOC[/"PDF / DOCX / Markdown / HTML / URL"/]
    end
    
    subgraph DOCSUMMARIZER["DocSummarizer.Core (Deterministic)"]
        direction TB
        PARSE["Parse & Structure"]
        SEGMENT["Segment by Semantics"]
        EMBED["Generate Embeddings<br/>(ONNX - Local)"]
        SCORE["Compute Salience"]
        CITE["Assign Citation IDs"]
        
        PARSE --> SEGMENT
        SEGMENT --> EMBED
        EMBED --> SCORE
        SCORE --> CITE
    end
    
    subgraph OUTPUT["Output (ExtractionResult)"]
        SEGMENTS[/"Segments[]<br/>• Original text (verbatim)<br/>• float[384] embedding<br/>• Salience score<br/>• StartChar / EndChar<br/>• Section context"/]
    end
    
    subgraph YOURS["Your Code"]
        STORE[("Vector Store<br/>(Qdrant / pgvector / etc)")]
    end
    
    subgraph QUERY["Query Time (Later)"]
        Q["User Question"]
        RETRIEVE["Retrieve Top-K"]
        LLM["LLM Synthesis"]
        ANS["Answer + Citations"]
        
        Q --> RETRIEVE
        RETRIEVE --> LLM
        LLM --> ANS
    end
    
    DOC --> PARSE
    CITE --> SEGMENTS
    SEGMENTS --> STORE
    STORE --> RETRIEVE
    
    style DOCSUMMARIZER stroke:#27ae60,stroke-width:3px
    style YOURS stroke:#3498db,stroke-width:2px
    style QUERY stroke:#9b59b6,stroke-width:2px
    style LLM stroke:#e74c3c,stroke-width:2px

क्या DocSummarizer करता है (green box):

  • संरचना को सुरक्षित रखने वाले दस्तावेजों का विश्लेषण करता है
  • सेमेटिक खण्डों में विभाजित करता है
  • एनएनक्स के साथ स्थानीय रूप से एम्बेडिंग उत्पन्न करता है
  • सुस्पष्टता प्राप्तियों को संगणित करता है
  • क्यारेक्टर पोजीशनों के साथ उद्धरण आईडी माना जाता है

क्या आप कर रहे हैं (blue box

  • अपने भेक्टर डाटाबेस में खण्डों को भंडारित करें
  • अपने पुनःप्राप्ति तर्क बनाएँ

क्वेरी के समय क्या होता है (purple box

  • प्रश्न को अंतःस्थापित करें (same model)
  • समान खंडों को पुनः प्राप्त करें
  • संश्लेषण के लिए एलएलएम को भेजें

प्रमुख अंतर्दृष्टि LLM (रेड किनारा) केवल क्वेरी समय में शामिल होता है।

पुनरुत्पादकता बोनस निर्णायक अंतरण का अर्थ होता है, आप फिर से आरएजी पाइपलाइन को डिबग कर सकते हैं जैसे किसी भी अन्य निर्माण अवयव की तरह।

इंजेक्शन के लिए कोई एलएलएम क्यों नहीं

के लिए RAG, आप चाहते हैं वास्तविक वाक्य अपने दस्तावेजों से - LLM नहीं है

LLM बाद में लाया जाता है। क्वेरी समय पर। पुस्तक प्राप्त टुकडों से एक जवाब बनाने के लिए।

DocSummarizer ExtractSegmentsAsync आपको ठीक से यह देता है: embeddings के साथ मूल पाठ सेगमेंट्स

मूल्य प्रस्ताव

यहाँ-' है क्या आप एक के साथ मिलता है dotnet add package:

dotnet add package Mostlylucid.DocSummarizer
  • स्मार्ट खण्डन - शीर्षकों पर विभाजन
  • ऑनिक्स सम्मिलन कोई API कुंजी नहीं
  • नम्रता अंकन प्रत्येक खंड के लिए संगणित महत्व
  • उद्धरण ट्रैकिंग - प्रत्येक खंड एक विशिष्ट आईडी है
  • बहुविध प्रारूप मार्कडाउन
  • सदिश भंडार विकल्प - InMemory

कोई पाइथोन नहीं . कोई बाहरी APIs नहीं है | . | कोई जटिल सेटअप नहीं है & #44; . | पहली मॉडल डाउनलोड के बाद ऑफ़लाइन काम करता है

सेगमेंट्स निकालें: कोर API

सरलतम उपयोग केस - अपने भेक्टर भंडार के लिए तैयार एम्बेडिंग सहित सेगमेंट्स निकालें

using Microsoft.Extensions.DependencyInjection;
using Mostlylucid.DocSummarizer;

// Setup DI
var services = new ServiceCollection();
services.AddDocSummarizer();
var provider = services.BuildServiceProvider();

var summarizer = provider.GetRequiredService<IDocumentSummarizer>();

// Extract segments with embeddings
string markdown = File.ReadAllText("document.md");
var extraction = await summarizer.ExtractSegmentsAsync(markdown);

foreach (var segment in extraction.AllSegments)
{
    Console.WriteLine($"[{segment.Type}] {segment.SectionTitle}");
    Console.WriteLine($"  ID: {segment.Id}");
    Console.WriteLine($"  Salience: {segment.SalienceScore:F2}");
    Console.WriteLine($"  Embedding: float[{segment.Embedding?.Length}]");
    Console.WriteLine($"  Text: {segment.Text[..Math.Min(80, segment.Text.Length)]}...");
}

आउटपुट

[Heading] Introduction
  ID: a1b2c3d4e5f6g7h8_h_0
  Salience: 0.85
  Embedding: float[384]
  Text: This document describes the architecture of our new microservices platform...

[Sentence] Introduction  
  ID: a1b2c3d4e5f6g7h8_s_1
  Salience: 0.72
  Embedding: float[384]
  Text: The system is designed to handle 10,000 requests per second with sub-100ms...

That's itM SK1 No orchestrationMSC2 no prompts, no opinions MSP4 just segments with embeddings and provenanceMST5 Ready for your vector databaseM ST6

दस्तावेज़ आईडी

प्रत्येक खंड Id एक दस्तावेज आईडी प्लस प्रकार और अनुक्रमणिका से बना है {docId}_{type}_{index}.

आप अपना दस्तावेज आईडी प्रदान कर सकते हैं

// Option 1: Provide your own ID (useful for tracking documents in your system)
var extraction = await summarizer.ExtractSegmentsAsync(markdown, documentId: "contract-2024-001");
// Segments get IDs like: contract_2024_001_s_0, contract_2024_001_h_1, ...

// Option 2: Auto-generated from content hash (default)
var extraction = await summarizer.ExtractSegmentsAsync(markdown);
// Segments get IDs like: a1b2c3d4e5f6g7h8_s_0, a1b2c3d4e5f6g7h8_h_1, ...
// Same document = same hash = same IDs (deterministic)

RAG के लिए यह क्यों महत्वपूर्ण है

  • अनुकूल आईडी आपको अपने दस्तावेज प्रबंधन प्रणाली के साथ खंडों को जोड़ने देता है
  • सामग्री-hash आईडी यह सुनिश्चित करती है कि एक ही दस्तावेज़ को पुनः indexing करने से समान खंड आईडी उत्पन्न होती है
  • दोनों दृष्टिकोण स्थिर उद्धरणों को समर्थन देते हैं - [s42] हमेशा एक ही स्रोत पाठ में हल करता है

एक खंड में क्या है

प्रत्येक निकाले गए खंड में सब कुछ है जो आपको RAG के लिए आवश्यक है

public class Segment
{
    string Id;              // Unique ID: "mydoc_s_42" (for citations)
    string Text;            // The actual content
    SegmentType Type;       // Sentence, Heading, ListItem, CodeBlock, Quote, TableRow
    int Index;              // 0-based order in document
    
    // Source location tracking
    int StartChar;          // Character offset where segment starts
    int EndChar;            // Character offset where segment ends
    int? PageNumber;        // Page number (for PDFs)
    int? LineNumber;        // Line number (for text/markdown)
    
    // Section context
    string SectionTitle;    // "Introduction" - immediate heading
    string HeadingPath;     // "Chapter 1 > Introduction > Overview"
    int HeadingLevel;       // 1-6 (heading depth)
    
    // Computed during extraction
    float[] Embedding;      // 384-dim vector (default model)
    double SalienceScore;   // 0-1 importance score
    string ContentHash;     // Stable hash for citation tracking across re-indexing
    
    // For retrieval (set during query)
    double QuerySimilarity; // Similarity to the query
    double RetrievalScore;  // Combined score: similarity + salience
    
    string Citation { get; } // Auto-generated: "[s42]", "[h3]", etc.
}

Id उद्धरण ट्रैकिंग के लिए कुंजी है. जब आपका LLM आउटपुट [s42], आप इसे सही स्रोत स्थान पर उपयोग कर वापस हल कर सकते हैं StartChar/EndChar.

बोनस ExtractionResult उद्धरण रिजोल्यूशन के लिए सहायक विधि शामिल हैं:

var extraction = await summarizer.ExtractSegmentsAsync(markdown);

// Fast O(1) lookups
var segment = extraction.GetSegment("mydoc_s_42");
var segmentByIdx = extraction.GetSegmentByIndex(42);

// Find segment at a character position
var segmentAtPos = extraction.GetSegmentAtPosition(5432);

// Get all segments on page 5 (for PDFs)
var pageSegments = extraction.GetSegmentsOnPage(5);

// Get source location for highlighting
var location = extraction.GetSourceLocation("mydoc_s_42");
// Returns: StartChar, EndChar, LineNumber, PageNumber, SectionTitle, HeadingPath

// Extract highlighted text with context
var highlight = extraction.GetHighlightedText(originalMarkdown, "mydoc_s_42", contextChars: 50);
Console.WriteLine(highlight.ToHtml());  // <span class="highlight">...</span>
Console.WriteLine(highlight.ToMarkdown()); // **...**

वेक्टर भंडार में प्लगिंग

DocSummarizer आपको एम्बेडिंग देता है

क्वारंटCity name (optional, probably does not need a translation)

var points = extraction.AllSegments.Select((s, i) => new PointStruct
{
    Id = (ulong)i,
    Vectors = s.Embedding,
    Payload = 
    {
        ["text"] = s.Text,
        ["section"] = s.SectionTitle,
        ["salience"] = s.SalienceScore,
        ["segment_id"] = s.Id,
        ["start_char"] = s.StartChar,
        ["end_char"] = s.EndChar
    }
}).ToList();

await qdrantClient.UpsertAsync("documents", points);

PostgreSQL + pgvector

foreach (var segment in extraction.Segments)
{
    await connection.ExecuteAsync(
        @"INSERT INTO documents (segment_id, text, heading, salience, embedding) 
          VALUES (@id, @text, @heading, @salience, @embedding::vector)",
        new { 
            id = segment.Id,
            text = segment.Text, 
            heading = segment.SectionTitle,
            salience = segment.SalienceScore,
            // NOTE: String interpolation is for demo simplicity only.
            // For production, use NpgsqlParameter with Vector type for better
            // performance and to avoid culture-dependent decimal separators.
            embedding = $"[{string.Join(",", segment.Embedding)}]"
        });
}

या निर्मित स्टोर्स का उपयोग करें

एक अलग डाटाबेस का प्रबंधन नहीं करना चाहते हैं

services.AddDocSummarizer(options =>
{
    // In-memory (fastest, no persistence)
    options.BertRag.VectorStore = VectorStoreBackend.InMemory;
    
    // DuckDB (embedded file-based, default)
    options.BertRag.VectorStore = VectorStoreBackend.DuckDB;
    
    // Qdrant (external server)
    options.BertRag.VectorStore = VectorStoreBackend.Qdrant;
    options.Qdrant.Host = "localhost";
    options.Qdrant.Port = 6334;
});

सलिन्स का क्या महत्व है?

अधिकतर RAG तंत्र असफल नहीं है क्योंकि एम्बेडिंग खराब हैं, लेकिन इसलिए कि सभी टुकड़े समान रूप से महत्वपूर्ण माना जाता है

flowchart LR
    subgraph DOC["Document"]
        H1["# Title"]
        P1["First paragraph<br/>(intro)"]
        H2["## Methods"]
        P2["Technical details..."]
        P3["More details..."]
        H3["## Results"]
        P4["Key findings here"]
        H4["## Appendix"]
        P5["Reference data..."]
    end
    
    subgraph SCORES["Salience Scores"]
        S1["0.95"]
        S2["0.85"]
        S3["0.70"]
        S4["0.65"]
        S5["0.60"]
        S6["0.80"]
        S7["0.30"]
    end
    
    H1 --> S1
    P1 --> S2
    H2 --> S3
    P2 --> S4
    P3 --> S5
    P4 --> S6
    P5 --> S7
    
    style S1 stroke:#27ae60,stroke-width:3px
    style S2 stroke:#27ae60,stroke-width:2px
    style S6 stroke:#27ae60,stroke-width:2px
    style S7 stroke:#e74c3c,stroke-width:2px

सारांश में एक वाक्य appendix में से अधिक महत्वपूर्ण है

// Get the top 20% most salient segments
var topSegments = extraction.Segments
    .OrderByDescending(s => s.SalienceScore)
    .Take((int)(extraction.Segments.Count * 0.2));

सलायन कारक

गुणक |--------|--------| | स्थिति शीर्षक समीपता | शीर्ष के बाद पहले वाक्य विषय वाक्य हैं | लंबाई अनुच्छेद प्रकार सामग्री प्रकार | कोड ब्लॉक्स

इसका मतलब है कि आपके पुनर्प्राप्ति द्वारा वजन कर सकते हैं (similarity * salience) सिर्फ समानता के बजाय

दस्तावेज़ वर्गीकरण

DocSummarizer auto-Content पर हियूरिसिक्स का उपयोग करके दस्तावेज़ प्रकार पता लगाता है

var extraction = await summarizer.ExtractSegmentsAsync(markdown);

// Document type detected from content
Console.WriteLine($"Type: {extraction.DocumentType}");     // Technical, Narrative, Legal, etc.
Console.WriteLine($"Confidence: {extraction.Confidence}"); // High, Medium, Low

वर्गीकरण पुनर्प्राप्ति पर प्रभाव डालता हैदस्तावेज़ एन्ट्रोपी के साथ रिट्रीव गहराई स्केल,Not a hardcoded TopK. Narrative documents (Fiction,History) get a 1.5x boost to retrieval count because they need more context.Technical documents with clear structure Need less

शल्यचिकित्सा की दृष्टि से

  • कोड ब्लॉक आवृत्ति (तकनीकी
  • शीर्षक संरचना (
  • संवाद पैटर्न
  • विधिक शब्दावली (संविदाएं
  • सूची घनत्व

यदि वेuristics अनिश्चित हैं, तो DocSummarizer एक "sentinel" मॉडल के प्रयोग से एक तेज LLM वर्गीकरण में वापस आ सकता है tinyllamaविन्यास के माध्यम से इसे सक्षम करें

services.AddDocSummarizer(options =>
{
    options.Ollama.BaseUrl = "http://localhost:11434";
    options.Ollama.Model = "tinyllama";
});

// Then use with LLM fallback enabled
var extraction = await summarizer.ExtractSegmentsAsync(markdown, useLlmFallback: true);

अधिकतर दस्तावेज़ों के लिए ह्य,हृ केवल हियूरिसिक्स काफी सटीक है।

पूर्ण आरएजी पाइपलाइन उदाहरण

यहाँ' एक पूर्ण सूचकांक- और - प्रश्न पाइपलाइन. हम' यह तीन चरणों में बना देंगे

sequenceDiagram
    participant User
    participant App as Your App
    participant DS as DocSummarizer
    participant VS as Vector Store
    participant LLM
    
    Note over DS: INGESTION (No LLM)
    App->>DS: ExtractSegmentsAsync(markdown)
    DS->>DS: Parse structure
    DS->>DS: Split into segments
    DS->>DS: Generate embeddings (ONNX)
    DS->>DS: Compute salience
    DS-->>App: ExtractionResult
    App->>VS: Store segments + vectors
    
    Note over LLM: QUERY TIME (LLM involved)
    User->>App: "What about X?"
    App->>DS: EmbedAsync(question)
    DS-->>App: float[384]
    App->>VS: Search(vector, topK=5)
    VS-->>App: Top segments
    App->>LLM: Question + Context
    LLM-->>App: Answer with [citations]
    App-->>User: Answer

चरण 1: डेटा संरचना

public class SimpleRagService
{
    private readonly IDocumentSummarizer _summarizer;
    
    // In-memory segment store - maps "docId:segmentId" to the full segment
    private readonly Dictionary<string, ExtractedSegment> _segments = new();
    
    // In-memory vector index - pairs of (id, embedding vector)
    private readonly List<(string Id, float[] Vector)> _index = new();

उत्पादन में आप एक वास्तविक भेक्टर डाटाबेस का उपयोग करते हैं

चरण 2: अनुक्रमण दस्तावेज

    public async Task IndexAsync(string markdown, string docId)
    {
        // Extract segments with embeddings - this is where DocSummarizer does the work
        var extraction = await _summarizer.ExtractSegmentsAsync(markdown);
        
        // Store each segment and its vector
        foreach (var segment in extraction.Segments)
        {
            // Composite key: document + segment for citation tracking
            var id = $"{docId}:{segment.SegmentId}";
            
            // Keep the full segment for retrieval
            _segments[id] = segment;
            
            // Add to vector index for similarity search
            _index.Add((id, segment.Embedding));
        }
    }

नोट: कोई LLM शामिल नहीं है वास्तविक दस्तावेज़ पाठ, सारांश नहीं

चरण 3: क्वेरी

    public async Task<string> QueryAsync(string question, int topK = 5)
    {
        // Embed the question using the same model as documents
        // This ensures vectors are in the same space
        var embedding = await _summarizer.EmbedAsync(question);
        
        // Find top-K most similar segments
        var results = _index
            .Select(x => (x.Id, Similarity: CosineSimilarity(embedding, x.Vector)))
            .OrderByDescending(x => x.Similarity)
            .Take(topK)
            .Select(x => _segments[x.Id])
            .ToList();
        
        // Build context with citation markers
        // The LLM can reference [chunk-3] and we can trace it back
        var context = string.Join("\n\n", results.Select(s => 
            $"[{s.SegmentId}] {s.Text}"));
        
        return context; // Send this + the question to your LLM
    }

लौटाए गए संदर्भ में है वास्तविक खण्ड आईडी के साथ दस्तावेज़ पाठ

Answer the question based on the following context.
Cite sources using the [chunk-N] markers.

Context:
{context}

Question: {question}

गणित (मानक कोसाइन समानता

    private static float CosineSimilarity(float[] a, float[] b)
    {
        float dot = 0, normA = 0, normB = 0;
        for (int i = 0; i < a.Length; i++)
        {
            dot += a[i] * b[i];
            normA += a[i] * a[i];
            normB += b[i] * b[i];
        }
        return dot / (MathF.Sqrt(normA) * MathF.Sqrt(normB));
    }
}

DocSummarizer में शामिल है VectorMath.CosineSimilarity() यदि आप यह स्वयं लिखना नहीं चाहते हैं

DocSummarizer भी प्रदर्शित करता है IEmbeddingService directly if you need to embed queries separately from the full summarization pipeline

सम्मिलित मॉडल

डिफ़ॉल्ट है AllMiniLmL6V2 - तेजी से,छोटीM SK2सबसे अच्छी गुणवत्ताMSC3अपनी आवश्यकताओं पर आधारित चुनें

services.AddDocSummarizer(options =>
{
    options.Onnx.EmbeddingModel = OnnxEmbeddingModel.BgeBaseEnV15;
});
मॉडल Dims
AllMiniLmL6V2 डिफ़ॉल्ट
BgeSmallEnV15 सर्वोत्तम गुणवत्ता
BgeBaseEnV15 उत्पादन गुणवत्ता
JinaEmbeddingsV2BaseEn

मॉडल स्वचालित-हेगिंगफेस से पहली बार उपयोग पर डाउनलोड करें

ओपन टेलीमीटरी

अपने RAG पाइपलाइन को उत्पादन में मॉनीटर करें

services.AddOpenTelemetry()
    .WithTracing(tracing => tracing
        .AddSource("Mostlylucid.DocSummarizer")
        .AddSource("Mostlylucid.DocSummarizer.Ollama")
        .AddSource("Mostlylucid.DocSummarizer.WebFetcher")
        .AddOtlpExporter())
    .WithMetrics(metrics => metrics
        .AddMeter("Mostlylucid.DocSummarizer")
        .AddMeter("Mostlylucid.DocSummarizer.Ollama")
        .AddMeter("Mostlylucid.DocSummarizer.WebFetcher")
        .AddPrometheusExporter());

कुंजी मापन

  • docsummarizer.summarizations - अनुरोध गणना
  • docsummarizer.summarization.duration - प्रक्रिया समय ms में
  • docsummarizer.document.size - दस्तावेज़ आकार
  • docsummarizer.ollama.embed.requests - API कॉल सम्मिलित करना

प्रारूप समर्थन

DocSummarizer कई दस्तावेज़ प्रारूपों को स्मार्ट पता लगाने और प्रक्रमण के साथ नियंत्रित करता है

प्रत्यक्ष संसाधन (अंतरिक निर्भरता नहीं

इन प्रारूपों को मूल रूप से प्रसंस्करण किया जाता है - कोई डॉकलिंग या अन्य सेवाओं की आवश्यकता नहीं

प्रारूप | एक्सटेंशन |--------|-----------|------------| मार्कडाउन .md, .markdown | मार्कडिग के साथ वर्णित सादा पाठ .txt, .text पैराग्राफों द्वारा विभाजित एचटीएमएल .html, .htm | Sanitized, Markdown में परिवर्तित किया गया ZIP अभिलेख .zip पाठ फ़ाइलें निकालता है

सादा पाठ स्मार्ट हैंडलिंग प्राप्त करता हैजब कोई मार्कडाउन शीर्षक नहीं है तो खंडक अनुच्छेद में स्विच करता है।

// Plain text works the same way
var plainText = File.ReadAllText("notes.txt");
var extraction = await summarizer.ExtractSegmentsAsync(plainText);
// Chunks split by paragraphs, embeddings generated

डॉकलिंग के साथ समृद्ध दस्तावेज

पीडीएफ के लिए

docker run -d -p 5001:5001 quay.io/docling-project/docling-serve
services.AddDocSummarizer(options =>
{
    options.Docling.BaseUrl = "http://localhost:5001";
});

// PDF, DOCX, PPTX, images all work
var pdfBytes = await File.ReadAllBytesAsync("document.pdf");
var extraction = await summarizer.ExtractSegmentsAsync(pdfBytes, "document.pdf");

डॉकलिंग दस्तावेज़ संरचना को सुरक्षित रखता है - शीर्षकों, तालिकाओंM SK2 सूची उचित मार्कडाउन के रूप में आते हैं

प्रारूप | एक्सटेंशन |--------|-----------|-------| पीडीएफ .pdf पाठ + सजावट सुरक्षित है शब्द .docx पूर्ण ढाँचाबद्धता PowerPoint .pptx स्लाइड खंड बन जाते हैं Excel .xlsx निकाले गए तालिकाएं छवियाँ .png, .jpg, .tiff डॉकलिंग के माध्यम से ओसीआर |

देखें भाग 1 डॉकलिंग इंटीग्रेशन पर अधिक जानकारी के लिए बहु--Format दस्तावेज़ रूपांतरण गहराई में डुबोने के लिए

हार्ड पार्ट्स आप छोड़ें

समस्या |---------|----------------------| | सेमेटिक सीमाओं पर वर्णन | | | शीर्षकों में विभाजित |, | समूहों से संबंधित सामग्री प्रति मॉडल | टोकनाइजेशन एम्बेडिंग बैचिंग | उद्धरण ट्रैकिंग SegmentId | एकल API के माध्यम से | ढाँचा रूपांतरण | मॉडल डाउनलोड ~/.docsummarizer |

सारांश

RAG पाइपलाइनों के रोचक भाग से पहले अवसंरचना की आवश्यकता है

  1. सम्मिलित सेगमेंट्स - किसी भी वेक्टर भंडार के लिए तैयार
  2. सलिन्स स्कोर्स - सभी टुकड़े समान नहीं हैं
  3. उद्धरण ट्रैकिंग - पता कहाँ से उत्तर आते हैं
  4. स्थानीय-पहले कोई API कुंजी नहीं
  5. उत्पादन टेलीमीटरी - ओपन टेलीमीटरी में निर्मित
dotnet add package Mostlylucid.DocSummarizer

पाईपिंग किया गया है. अपने RAG अनुप्रयोग का निर्माण करें

संबंधित लेख

डॉक-सममैजर श्रृङ्खला

आर. ए. जी. डी. डायव्स

ग्राफ आर ए जी

संबंधित उपकरण

लिंक्स

logo

© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.