# "Lakimies GPT:n" rakentaminen blogiasi varten - Osa 9: Dokumentin nieleminen Doclingilla

<!--category-- AI, LLM, Docling, RAG, C#, AI-Article, mostlylucid.blogllm -->
<datetime class="hidden">2025-12-15T22:45</datetime>

Tervetuloa 9. osaan! Edellisissä osissa olemme rakentaneet vankan RAG-järjestelmän, joka käsittelee markedown-blogikirjoituksia ja tekee niistä hakukelpoisia semanttisten upotusten kautta. Nyt on aika laajentaa kykyjämme **Dokulaatio** - tehokas dokumenttien käsittelykirjasto, joka voi nauttia pdf-, DOCX- ja muita formaatteja, mikä tekee asianajajamme GPT:stä todella kattavan.

> HUOMAUTUS: Tämä on osa kokeilujani tekoälyllä (avustettu piirtäminen) + omalla muokkauksellani. Sama ääni, sama käytännönläheisyys, vain nopeammat sormet.

Ei sillä, että tarvitsisin tätä blogiin (se on kaikki marketdown), mutta loppuun asti ajattelin näyttää, kuinka helposti lisään dokumentin nielemiskykyä RAG-putkeen. Tämä on erityisen hyödyllistä, jos rakennat järjestelmää, joka käsittelee oikeudellisia asiakirjoja, sopimuksia tai muita liiketoiminta-asiakirjoja.

[TOC]

## Miksi asiakirjasyöminen on tärkeää lailliselle tekoälylle

Nykyaikaiset oikeuskäytännöt käsittelevät erilaisia asiakirjaformaatteja pelkän tekstin lisäksi. Lakimiesten tulee viitata:

- **PDF-sopimukset** ja oikeudelliset sopimukset
- **DOCX-kirjaimet** ja esityksiä
- **Sähköpostilangat** ja kirjeenvaihto
- **Tutkitut asiakirjat** (OCR:n kautta)
- **Taulukkotaulukot** tapaustietojen kanssa

Dokumentoinnin avulla asianajajamme GPT voi käsitellä näitä formaatteja ja tehdä niistä hakukelpoisia, luoden todella kattavan tietopohjan, joka heijastaa sitä, miten nykyaikaiset lakifirmat käyttävät tekoälyä.

## Mikä Docling on?

[Dokulaatio](https://github.com/docling-project/docling) on avoimen lähdekoodin asiakirjojen käsittelytyökalupaketti IBM:ltä, joka

- Muuntaa PDF:t, DOCX:n, PPPTX:n, HTML:n ja muut formaatit strukturoiduiksi markdowniksi tai JSON:iksi
- Arkistojen muotoilu, taulukot ja asiakirjarakenne
- Tukee läpivalaisuasiakirjojen OCR:ää
- Toimii Docker-palveluna [Valmennuspalvelu](https://github.com/docling-project/docling-serve)
- REST-rajapinta on helppo integroida mihin tahansa kieleen

## Opettelun alustaminen

### Docker-palvelun käyttöönotto

Docling Serve on helpoin tapa hoitaa Doclingia palveluna. Perustetaan se:

```bash
# Using the official container image from Quay.io
docker run -p 5001:5001 quay.io/docling-project/docling-serve

# Or with the UI enabled for testing
docker run -p 5001:5001 -e DOCLING_SERVE_ENABLE_UI=1 quay.io/docling-project/docling-serve
```

**Saatavilla olevat konttikuvat:**

Kuvan kuvaus Koko
|-------|-------------|------|
| `quay.io/docling-project/docling-serve` Peruskuva (PyPI-paketit) ~8,7GB (amd64)
| `quay.io/docling-project/docling-serve-cpu` Vain CPU-muunnos ~4.4GB
| `quay.io/docling-project/docling-serve-cu126` CUDA 12,6 GPU:lle ~10GB:lle
| `quay.io/docling-project/docling-serve-cu128` CUDA 12.8 GPU:n osalta ~11,4GB:n osalta

**Loppupäätelmät:**

- API: `http://localhost:5001`
- API-dokumentaatio (Swagger): `http://localhost:5001/docs`
- UI Playground: `http://localhost:5001/ui` (kun käytössä)

### Dockerin sävellysasetukset

Tuotantoa varten lisää Docling nykyiseen `docker-compose.yml`:

```yaml
services:
  docling:
    image: quay.io/docling-project/docling-serve:latest
    ports:
      - "5001:5001"
    environment:
      - DOCLING_SERVE_ENABLE_UI=0
      - DOCLING_SERVE_MAX_WORKERS=4
    volumes:
      - docling_cache:/root/.cache
    restart: unless-stopped
    
volumes:
  docling_cache:
```

## Testaaminen Dokulaatio kiharalla

Ennen integroitumista C#:n kanssa testataan API:

```bash
# Convert a PDF from URL
curl -X 'POST' \
  'http://localhost:5001/v1/convert/source' \
  -H 'accept: application/json' \
  -H 'Content-Type: application/json' \
  -d '{
    "sources": [{"kind": "http", "url": "https://arxiv.org/pdf/2501.17887"}],
    "options": {
      "to_formats": ["md"]
    }
  }'

# Convert a local file (upload)
curl -X 'POST' \
  'http://localhost:5001/v1/convert/file' \
  -H 'accept: application/json' \
  -F 'files=@contract.pdf'
```

## C# Integrointi

Koska Docling on Python-palvelu, integroitumme HTTP:n kautta. Tehdään vahva C#-asiakas.

### API-mallien muokkaus

```csharp
using System.Text.Json.Serialization;

namespace Mostlylucid.BlogLLM.Core.Models
{
    /// <summary>
    /// Request to convert documents from URLs or base64 content
    /// </summary>

    public class DoclingConvertRequest
    {
        [JsonPropertyName("sources")]
        public List<DoclingSource> Sources { get; set; } = new();

        [JsonPropertyName("options")]
        public DoclingOptions? Options { get; set; }
    }

    public class DoclingSource
    {
        [JsonPropertyName("kind")]
        public string Kind { get; set; } = "http"; // "http", "base64", "file"

        [JsonPropertyName("url")]
        public string? Url { get; set; }

        [JsonPropertyName("base64")]
        public string? Base64Content { get; set; }

        [JsonPropertyName("filename")]
        public string? Filename { get; set; }
    }

    public class DoclingOptions
    {
        [JsonPropertyName("to_formats")]
        public List<string> ToFormats { get; set; } = new() { "md" }; // "md", "json", "text"

        [JsonPropertyName("ocr")]
        public bool Ocr { get; set; } = true;

        [JsonPropertyName("table_mode")]
        public string TableMode { get; set; } = "accurate"; // "fast", "accurate"
    }

    /// <summary>
    /// Response from Docling conversion
    /// </summary>

    public class DoclingConvertResponse
    {
        [JsonPropertyName("document")]
        public DoclingDocument? Document { get; set; }

        [JsonPropertyName("status")]
        public string Status { get; set; } = string.Empty;

        [JsonPropertyName("errors")]
        public List<string>? Errors { get; set; }
    }

    public class DoclingDocument
    {
        [JsonPropertyName("md_content")]
        public string? MarkdownContent { get; set; }

        [JsonPropertyName("json_content")]
        public object? JsonContent { get; set; }

        [JsonPropertyName("text_content")]
        public string? TextContent { get; set; }

        [JsonPropertyName("metadata")]
        public DoclingMetadata? Metadata { get; set; }
    }

    public class DoclingMetadata
    {
        [JsonPropertyName("filename")]
        public string? Filename { get; set; }

        [JsonPropertyName("page_count")]
        public int? PageCount { get; set; }

        [JsonPropertyName("file_type")]
        public string? FileType { get; set; }
    }

    /// <summary>
    /// Our internal document model
    /// </summary>

    public class ProcessedDocument
    {
        public string DocumentId { get; set; } = Guid.NewGuid().ToString();
        public string FileName { get; set; } = string.Empty;
        public string OriginalFormat { get; set; } = string.Empty;
        public string MarkdownContent { get; set; } = string.Empty;
        public string? TextContent { get; set; }
        public DateTime ProcessedDate { get; set; } = DateTime.UtcNow;
        public string[] Categories { get; set; } = Array.Empty<string>();
        public int? PageCount { get; set; }
        public string ContentHash { get; set; } = string.Empty;
    }
}
```

### Customer Servicen opettelu

```csharp
using Microsoft.Extensions.Logging;
using Microsoft.Extensions.Options;
using Mostlylucid.BlogLLM.Core.Models;
using System.Net.Http.Json;
using System.Security.Cryptography;
using System.Text;
using System.Text.Json;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public class DoclingClientOptions
    {
        public string BaseUrl { get; set; } = "http://localhost:5001";
        public int TimeoutSeconds { get; set; } = 300; // 5 minutes for large documents
        public bool EnableOcr { get; set; } = true;
        public string TableMode { get; set; } = "accurate";
    }

    public class DoclingClient : IDisposable
    {
        private readonly HttpClient _httpClient;
        private readonly ILogger<DoclingClient> _logger;
        private readonly DoclingClientOptions _options;

        public DoclingClient(
            HttpClient httpClient,
            ILogger<DoclingClient> logger,
            IOptions<DoclingClientOptions> options)
        {
            _httpClient = httpClient;
            _logger = logger;
            _options = options.Value;

            _httpClient.BaseAddress = new Uri(_options.BaseUrl);
            _httpClient.Timeout = TimeSpan.FromSeconds(_options.TimeoutSeconds);
        }

        /// <summary>
        /// Convert a document from a URL
        /// </summary>

        public async Task<ProcessedDocument?> ConvertFromUrlAsync(
            string url,
            string[]? categories = null,
            CancellationToken cancellationToken = default)
        {
            _logger.LogInformation("Converting document from URL: {Url}", url);

            var request = new DoclingConvertRequest
            {
                Sources = new List<DoclingSource>
                {
                    new() { Kind = "http", Url = url }
                },
                Options = new DoclingOptions
                {
                    ToFormats = new List<string> { "md", "text" },
                    Ocr = _options.EnableOcr,
                    TableMode = _options.TableMode
                }
            };

            return await SendConversionRequestAsync(request, categories, cancellationToken);
        }

        /// <summary>
        /// Convert a local file
        /// </summary>

        public async Task<ProcessedDocument?> ConvertFileAsync(
            string filePath,
            string[]? categories = null,
            CancellationToken cancellationToken = default)
        {
            if (!File.Exists(filePath))
            {
                _logger.LogError("File not found: {FilePath}", filePath);
                return null;
            }

            _logger.LogInformation("Converting local file: {FilePath}", filePath);

            // Read file and convert to base64
            var fileBytes = await File.ReadAllBytesAsync(filePath, cancellationToken);
            var base64Content = Convert.ToBase64String(fileBytes);
            var fileName = Path.GetFileName(filePath);

            var request = new DoclingConvertRequest
            {
                Sources = new List<DoclingSource>
                {
                    new()
                    {
                        Kind = "base64",
                        Base64Content = base64Content,
                        Filename = fileName
                    }
                },
                Options = new DoclingOptions
                {
                    ToFormats = new List<string> { "md", "text" },
                    Ocr = _options.EnableOcr,
                    TableMode = _options.TableMode
                }
            };

            return await SendConversionRequestAsync(request, categories, cancellationToken);
        }

        /// <summary>
        /// Convert a file using multipart form upload (more efficient for large files)
        /// </summary>

        public async Task<ProcessedDocument?> ConvertFileUploadAsync(
            string filePath,
            string[]? categories = null,
            CancellationToken cancellationToken = default)
        {
            if (!File.Exists(filePath))
            {
                _logger.LogError("File not found: {FilePath}", filePath);
                return null;
            }

            _logger.LogInformation("Uploading and converting file: {FilePath}", filePath);

            try
            {
                using var fileStream = File.OpenRead(filePath);
                using var content = new MultipartFormDataContent();
                using var streamContent = new StreamContent(fileStream);

                var fileName = Path.GetFileName(filePath);
                content.Add(streamContent, "files", fileName);

                var response = await _httpClient.PostAsync(
                    "/v1/convert/file",
                    content,
                    cancellationToken);

                if (!response.IsSuccessStatusCode)
                {
                    var errorContent = await response.Content.ReadAsStringAsync(cancellationToken);
                    _logger.LogError("Docling conversion failed: {StatusCode} - {Error}",
                        response.StatusCode, errorContent);
                    return null;
                }

                var result = await response.Content.ReadFromJsonAsync<DoclingConvertResponse>(
                    cancellationToken: cancellationToken);

                return MapToProcessedDocument(result, fileName, categories);
            }
            catch (Exception ex)
            {
                _logger.LogError(ex, "Error uploading file to Docling: {FilePath}", filePath);
                return null;
            }
        }

        private async Task<ProcessedDocument?> SendConversionRequestAsync(
            DoclingConvertRequest request,
            string[]? categories,
            CancellationToken cancellationToken)
        {
            try
            {
                var response = await _httpClient.PostAsJsonAsync(
                    "/v1/convert/source",
                    request,
                    cancellationToken);

                if (!response.IsSuccessStatusCode)
                {
                    var errorContent = await response.Content.ReadAsStringAsync(cancellationToken);
                    _logger.LogError("Docling conversion failed: {StatusCode} - {Error}",
                        response.StatusCode, errorContent);
                    return null;
                }

                var result = await response.Content.ReadFromJsonAsync<DoclingConvertResponse>(
                    cancellationToken: cancellationToken);

                var filename = request.Sources.FirstOrDefault()?.Filename
                    ?? request.Sources.FirstOrDefault()?.Url
                    ?? "unknown";

                return MapToProcessedDocument(result, filename, categories);
            }
            catch (HttpRequestException ex)
            {
                _logger.LogError(ex, "HTTP error calling Docling API");
                return null;
            }
            catch (TaskCanceledException ex)
            {
                _logger.LogError(ex, "Docling conversion timed out");
                return null;
            }
            catch (Exception ex)
            {
                _logger.LogError(ex, "Unexpected error calling Docling API");
                return null;
            }
        }

        private ProcessedDocument? MapToProcessedDocument(
            DoclingConvertResponse? response,
            string filename,
            string[]? categories)
        {
            if (response?.Document == null)
            {
                _logger.LogWarning("Docling returned empty document");
                return null;
            }

            var markdownContent = response.Document.MarkdownContent ?? string.Empty;

            return new ProcessedDocument
            {
                DocumentId = Guid.NewGuid().ToString(),
                FileName = Path.GetFileName(filename),
                OriginalFormat = Path.GetExtension(filename).TrimStart('.').ToLower(),
                MarkdownContent = markdownContent,
                TextContent = response.Document.TextContent,
                ProcessedDate = DateTime.UtcNow,
                Categories = categories ?? Array.Empty<string>(),
                PageCount = response.Document.Metadata?.PageCount,
                ContentHash = ComputeHash(markdownContent)
            };
        }

        private static string ComputeHash(string content)
        {
            using var sha256 = SHA256.Create();
            var bytes = sha256.ComputeHash(Encoding.UTF8.GetBytes(content));
            return Convert.ToBase64String(bytes);
        }

        public void Dispose()
        {
            _httpClient?.Dispose();
        }
    }
}
```

### Dokumentin nielemispalvelu

Nyt yhdistetään Docling nykyiseen RAG-putkistoomme:

```csharp
using Microsoft.Extensions.Logging;
using Mostlylucid.BlogLLM.Core.Models;

namespace Mostlylucid.BlogLLM.Core.Services
{
    public class DocumentIngestionService
    {
        private readonly ILogger<DocumentIngestionService> _logger;
        private readonly DoclingClient _doclingClient;
        private readonly MarkdownParserService _markdownParser;
        private readonly ChunkingService _chunker;
        private readonly BatchEmbeddingService _embedder;
        private readonly QdrantVectorStore _vectorStore;

        public DocumentIngestionService(
            ILogger<DocumentIngestionService> logger,
            DoclingClient doclingClient,
            MarkdownParserService markdownParser,
            ChunkingService chunker,
            BatchEmbeddingService embedder,
            QdrantVectorStore vectorStore)
        {
            _logger = logger;
            _doclingClient = doclingClient;
            _markdownParser = markdownParser;
            _chunker = chunker;
            _embedder = embedder;
            _vectorStore = vectorStore;
        }

        /// <summary>
        /// Process a document file and add to vector store
        /// </summary>

        public async Task<bool> ProcessDocumentAsync(
            string filePath,
            string[]? categories = null,
            CancellationToken cancellationToken = default)
        {
            try
            {
                _logger.LogInformation("Processing document: {FilePath}", filePath);

                // Step 1: Convert with Docling
                var document = await _doclingClient.ConvertFileUploadAsync(
                    filePath, categories, cancellationToken);

                if (document == null)
                {
                    _logger.LogError("Failed to convert document: {FilePath}", filePath);
                    return false;
                }

                return await ProcessConvertedDocumentAsync(document, cancellationToken);
            }
            catch (Exception ex)
            {
                _logger.LogError(ex, "Error processing document {FilePath}", filePath);
                return false;
            }
        }

        /// <summary>
        /// Process a document from URL and add to vector store
        /// </summary>

        public async Task<bool> ProcessDocumentFromUrlAsync(
            string url,
            string[]? categories = null,
            CancellationToken cancellationToken = default)
        {
            try
            {
                _logger.LogInformation("Processing document from URL: {Url}", url);

                // Step 1: Convert with Docling
                var document = await _doclingClient.ConvertFromUrlAsync(
                    url, categories, cancellationToken);

                if (document == null)
                {
                    _logger.LogError("Failed to convert document from URL: {Url}", url);
                    return false;
                }

                return await ProcessConvertedDocumentAsync(document, cancellationToken);
            }
            catch (Exception ex)
            {
                _logger.LogError(ex, "Error processing document from URL {Url}", url);
                return false;
            }
        }

        private async Task<bool> ProcessConvertedDocumentAsync(
            ProcessedDocument document,
            CancellationToken cancellationToken)
        {
            _logger.LogInformation("Document converted: {FileName} ({PageCount} pages)",
                document.FileName, document.PageCount ?? 0);

            // Step 2: Parse the markdown content
            var post = _markdownParser.ParseMarkdownFromContent(
                document.MarkdownContent,
                document.FileName);
            post.Categories = document.Categories;

            // Step 3: Chunk the content
            var chunks = _chunker.ChunkBlogPost(post);
            _logger.LogInformation("Created {ChunkCount} chunks from {FileName}",
                chunks.Count, document.FileName);

            if (chunks.Count == 0)
            {
                _logger.LogWarning("No chunks created for document: {FileName}", document.FileName);
                return false;
            }

            // Step 4: Generate embeddings
            var progress = new Progress<int>(processed =>
            {
                _logger.LogDebug("Embedded {Processed}/{Total} chunks",
                    processed, chunks.Count);
            });

            await _embedder.GenerateEmbeddingsAsync(chunks, progress, cancellationToken);

            // Step 5: Store in vector database
            await _vectorStore.UpsertChunksAsync(chunks);

            _logger.LogInformation("Successfully processed document: {FileName} ({ChunkCount} chunks)",
                document.FileName, chunks.Count);

            return true;
        }

        /// <summary>
        /// Batch process multiple documents
        /// </summary>

        public async Task<(int success, int failed)> ProcessDocumentBatchAsync(
            IEnumerable<string> filePaths,
            string[]? categories = null,
            CancellationToken cancellationToken = default)
        {
            int success = 0;
            int failed = 0;

            foreach (var filePath in filePaths)
            {
                if (cancellationToken.IsCancellationRequested)
                    break;

                var result = await ProcessDocumentAsync(filePath, categories, cancellationToken);
                if (result)
                    success++;
                else
                    failed++;
            }

            _logger.LogInformation("Batch processing complete: {Success} succeeded, {Failed} failed",
                success, failed);

            return (success, failed);
        }
    }
}
```

### Riippuvuusinjektion asettelu

```csharp
using Microsoft.Extensions.DependencyInjection;
using Mostlylucid.BlogLLM.Core.Services;

public static class ServiceCollectionExtensions
{
    public static IServiceCollection AddDoclingServices(
        this IServiceCollection services,
        Action<DoclingClientOptions>? configureOptions = null)
    {
        // Configure options
        if (configureOptions != null)
        {
            services.Configure(configureOptions);
        }
        else
        {
            services.Configure<DoclingClientOptions>(options =>
            {
                options.BaseUrl = "http://localhost:5001";
                options.TimeoutSeconds = 300;
                options.EnableOcr = true;
            });
        }

        // Register HttpClient with configuration
        services.AddHttpClient<DoclingClient>((serviceProvider, client) =>
        {
            var options = serviceProvider
                .GetRequiredService<IOptions<DoclingClientOptions>>().Value;
            client.BaseAddress = new Uri(options.BaseUrl);
            client.Timeout = TimeSpan.FromSeconds(options.TimeoutSeconds);
        });

        // Register services
        services.AddScoped<DocumentIngestionService>();

        return services;
    }
}
```

### Asetukset

Lisää ruutuusi `appsettings.json`:

```json
{
  "Docling": {
    "BaseUrl": "http://localhost:5001",
    "TimeoutSeconds": 300,
    "EnableOcr": true,
    "TableMode": "accurate"
  }
}
```

## Käytä esimerkkejä

### Eri asiakirjatyyppien käsittely

```csharp
// PDF Processing
await documentIngestionService.ProcessDocumentAsync(
    "C:\\documents\\client_contract.pdf",
    new[] { "legal", "contracts", "client-agreements" });

// DOCX Processing
await documentIngestionService.ProcessDocumentAsync(
    "C:\\documents\\motion_to_dismiss.docx",
    new[] { "legal", "briefs", "motions" });

// URL Processing (great for public documents)
await documentIngestionService.ProcessDocumentFromUrlAsync(
    "https://arxiv.org/pdf/2501.17887",
    new[] { "research", "ai", "docling" });

// Batch Processing
var files = Directory.GetFiles("C:\\documents\\legal", "*.pdf");
var (success, failed) = await documentIngestionService.ProcessDocumentBatchAsync(
    files,
    new[] { "legal", "batch-import" });
Console.WriteLine($"Processed {success} files, {failed} failures");
```

### Tiedostontarkkailija automaattiselle syönnille

```csharp
public class DocumentWatcherService : BackgroundService
{
    private readonly IServiceProvider _serviceProvider;
    private readonly ILogger<DocumentWatcherService> _logger;
    private FileSystemWatcher? _watcher;
    private readonly string _watchPath;

    public DocumentWatcherService(
        IServiceProvider serviceProvider,
        ILogger<DocumentWatcherService> logger,
        IConfiguration configuration)
    {
        _serviceProvider = serviceProvider;
        _logger = logger;
        _watchPath = configuration["DocumentWatch:Path"] ?? "C:\\documents\\incoming";
    }

    protected override Task ExecuteAsync(CancellationToken stoppingToken)
    {
        if (!Directory.Exists(_watchPath))
        {
            Directory.CreateDirectory(_watchPath);
        }

        _watcher = new FileSystemWatcher(_watchPath)
        {
            NotifyFilter = NotifyFilters.FileName | NotifyFilters.LastWrite,
            IncludeSubdirectories = false
        };

        // Watch for common document types
        _watcher.Filters.Add("*.pdf");
        _watcher.Filters.Add("*.docx");
        _watcher.Filters.Add("*.doc");
        _watcher.Filters.Add("*.pptx");

        _watcher.Created += OnFileCreated;
        _watcher.EnableRaisingEvents = true;

        _logger.LogInformation("Watching for documents in: {Path}", _watchPath);

        return Task.CompletedTask;
    }

    private async void OnFileCreated(object sender, FileSystemEventArgs e)
    {
        _logger.LogInformation("New document detected: {FileName}", e.Name);

        // Wait for file to be fully written
        await Task.Delay(1000);

        using var scope = _serviceProvider.CreateScope();
        var ingestionService = scope.ServiceProvider
            .GetRequiredService<DocumentIngestionService>();

        await ingestionService.ProcessDocumentAsync(e.FullPath);
    }

    public override void Dispose()
    {
        _watcher?.Dispose();
        base.Dispose();
    }
}
```

## Suorituskykyä koskevia huomioita

### Asiakirjan kokovaikutus

Asiakirjan tyyppi Tyypillinen koko Dokumentin käsittely Upotusaika
|--------------|-------------|-------------------|----------------|
Yksinkertainen PDF (1-5 sivua) 100KB-500KB 2-5 sekuntia 0,5-1 sekuntia
Complex PDF (20+ sivua) 1-5MB 10-30 sekuntia 2-5 sekuntia
Skannattu PDF-muodossa 1-10MB 30-120 sekuntia 2-5 sekuntia
DOCX-asiakirja 50KB-500KB 1-3 sekuntia 0,3-0,8 sekuntia
PPTX Esitys 1-20MB 5-30 sekuntia 1-3 sekuntia

### Optimointivinkkejä

1. **Käytä GPU-kuvia** OCR:n raskaiden työmäärien osalta (`docling-serve-cu126` tai `docling-serve-cu128`)
2. **Poista OCR käytöstä** digitaalisten PDF-levyjen käsittely nopeutuu
3. **Käyttö `fast` Pöytätila** kun taulukon tarkkuus ei ole kriittinen
4. **Eräprosessi** suurten dokumenttisarjojen puheajan ulkopuolella
5. **Välimuistin tulokset** - säilytä sisältö hash, jotta et käsittelisi asiakirjoja muuttumattomina

## Yhteenveto

Olemme onnistuneesti integroineet Doclingin Lakimies GPT -järjestelmään, mikä mahdollistaa:

1. PDF-, DOCX-, PPTX- ja muiden asiakirjamuotojen käsittely
2. OCR-tuki skannatuille dokumenteille
3. Puhdas integroituminen nykyiseen RAG-putkistoomme
4. Dockerin käyttöönotto skaalautuvuuteen
5. Erän käsittelyominaisuudet
6. Automaattinen tiedoston tarkkailu näppäilyn varalta

Tämä laajentaa tietopohjaamme blogikirjoitusten lisäksi myös oikeudellisiin asiakirjoihin, sopimuksiin, tutkimuspapereihin ja muihin tärkeisiin materiaaleihin.

## Sarjanavigointi

- [Osa 1: Johdanto ja arkkitehtuuri](/blog/building-a-lawyer-gpt-for-your-blog-part1)
- [Osa 2: GPU:n asetukset ja CUDA C#-muodossa](/blog/building-a-lawyer-gpt-for-your-blog-part2)
- [Osa 3: Understanding Upbeddings & Vector Databases](/blog/building-a-lawyer-gpt-for-your-blog-part3)
- [Osa 4: Ruoansulatusputken rakentaminen](/blog/building-a-lawyer-gpt-for-your-blog-part4)
- [Osa 5: Windows-asiakas](/blog/building-a-lawyer-gpt-for-your-blog-part5)
- [Osa 6: Paikallinen LLM-integraatio](/blog/building-a-lawyer-gpt-for-your-blog-part6)
- [Osa 7: Content Generation & Prompt Engineering](/blog/building-a-lawyer-gpt-for-your-blog-part7)
- [Osa 8: Kehittyneet ominaisuudet ja tuotannon käyttöönotto](/blog/building-a-lawyer-gpt-for-your-blog-part8)
- **Osa 9: Dokumenttien nauttiminen** (tässä virassa)

## Resurssit

- [GitHub-varastojen operointi](https://github.com/docling-project/docling)
- [Dokumentointi GitHub-varastojen tarjoilu](https://github.com/docling-project/docling-serve)
- [Dokumentaation opettelu](https://github.com/docling-project/docling-serve/blob/main/docs/README.md)
- [Docling Serve API -dokumentaatio](http://localhost:5001/docs) (toimittaessa paikallisesti)
- [ArXiv-paperin muokkaus](https://arxiv.org/abs/2501.17887)