Back to "वीडियोसम्मैजर: वीडियो के लिए कम RAG (Shots | → | दृश्य |→ | Evidence"

This is a viewer only at the moment see the article on how this works.

To update the preview hit Ctrl-Alt-R (or ⌘-Alt-R on Mac) or Enter to refresh. The Save icon lets you save the markdown file to disk

This is a preview from the server running through my markdig pipeline

AI Architecture CLIP LLM NER ONNX Patterns Video

वीडियोसम्मैजर: वीडियो के लिए कम RAG (Shots | → | दृश्य |→ | Evidence

Thursday, 15 January 2026

स्थिति: विकास के एक भाग के रूप में सुस्पष्टRAG. स्रोत: github.comM SK1scottgal/lucidrag

जहाँ यह ठीक है: वीडियोसम्मैजर है वाद्यकार का सुस्पष्टRAG परिवार, तीन पाइपलाइनों को एक एकीकृत वीडियो विश्लेषण इंजन में सम्मिलित करने के लिए

सभी एक ही अनुसरण करें कम RAG पैटर्न: एक बार सिग्नलों को निकालने के लिए , साक्ष्य भंडारित करें


एक दो घण्टे फिल्म फ्रेम को प्रसंस्करण करने के लिए CLIP अंतःकरणों के साथ फ्रेम ले जाएगा और गणना में सैकड़ों डॉलर लगेंगे।

VideoSummarizer इस समस्या को तीन प्रमुख optimizations के साथ हल करता है

  1. बोधात्मक हैश द्वैतीकरण महंगे ML के पहले दृश्य रूप से समान फ्रेमें छोड़ दें
  2. बैच क्लिप सम्मिलन प्रक्रिया एक के स्थान पर एक GPU पास प्रति छवि 8 speedup
  3. पाइपलाइन संरचना - चेन इमेजके लिए कुंजी फ्रेम्सका स्ममराइजर

परिणाम : एक 2- घंटे फिल्म प्रक्रियाओं में |~10-15 | मिनटों | , | घंटों नहीं |M. के साथ ही वास्तुकला सिद्धांत ImageSummarizerComment और ऑडियोसममारीजरNameलेकिन एक एकीकृत वीडियो विश्लेषण पाइपलाइन में सम्मिलित

कोर अंतर्दृष्टि वीडियो छवियाँ है | + | ऑडियो |+ | पाठ & #44; . | प्रत्येक डोमेन को विशिष्ट उपकरणों के साथ संसाधित करें

  • प्रक्रिया संरचना पहले (cuts
  • एक बार क्रॉस-मोडल संकेत निकालें (embeddings

शब्दावली

  • गोलीबारी कटों के बीच कैमरा ले (संरचनात्मक,FFmpeg दृश्य पता लगाने से
  • दृश्य एक सुसंगत इकाई (
  • प्रमाण (start_time, end_time) संकेत + संकेतक

कुंजी एमएल मॉडलों का उपयोग किया गया

  • क्लिप OpenAI's कंट्रास्टिव भाषा
  • फुसफुसाएँ OpenAI ध्वनि पहचान मॉडल ' ध्वनी को समयtamps के साथ पाठ में स्थानांतरित करता है
  • BERT-NER नामित अस्तित्व पहचान
  • ऑनिक्स रनटाइम क्रॉसः- platform ML inference, ; CPU पर मॉडल चलाता है.

इस लेख में शामिल है

  • VideoSummarizer तीन पाइपलाइनों को कैसे orchestrates?
  • तरंग वास्तुकला: 16 मानकीकरण से प्रमाण उत्पादन तक तरंग
  • क्षमता प्रणाली: लचीला मॉडल डाउनलोड
  • कुंजीफ्रेम एम्बेडिंग के लिए बैच सीलिप अनुकूलन (3-5x तेज
  • संवेदनात्मक हैश ह्रास (40% फ्रेम ह्रान )
  • बहुविध ध्वनि दृश्य समूहन | ( | सम्मिलित |+ | ट्रांस्क्रिप्ट | МSK3 | काटने का प्रकार |
  • ट्रांस्क्रिप्टों से इकाई निष्कर्षण के लिए एनईआर एकीकरण
  • अल्पावधि परमाणु दर सीमित करने के लिए, समय आकलन, और बैकप्रेस
  • आउटपुट

संबंधित लेख:


समस्या: वीडियो किफायती है

मानदंडएमडीएएम पर मापा गया निम्न संख्याएँ 9950X | (16- | Core | | 3 | 4 | NVIDIA A | 5 | 6 | GB | 7 | 8 | 9 |GB RAM | 10 | NVMe | 11 | 12 | p H | 13 | Whisper base | 14 | आपका मील बदलेगा | 15

एक सामान्य फिल्म में शामिल है:

  • फ्रेम घंटे पर 24fps
  • ~2 श्रव्य घण्टे स्पीच
  • बहुविध पाठ परतें उपशीर्षक

स्ट्रोमन दृष्टिकोण (कहीं ऐसा नहीं करता है

  • प्रति फ्रेम में सीलिप सम्मिलन घंटे
  • प्रति फ्रेम दृश्य एलएलएम घंटे (cloud Vision API गेंद पार्क

कुंजी फ्रेम एक्सटेक्शन के साथ भी (say, 500-1000 फ्रेमों से), कि' अभी भी serial CLIP inference के सेकेंड

पारंपरिक दृष्टिकोणएक्सट्रेक्ट कुंजी फ्रेम्स

समस्या: यह redundant फ्रेमों पर संगणना बर्न करता है.

समाधान: बहुआयामी,- चरण फिल्टरिंग,, बैच प्रसंस्करण और पाइपलाइन संरचना


वीडियोसममैजर वास्तुकला

वीडियोसममैजर एप्लेट कम RAG तीन चरण कम करने के साथ वीडियो के लिए

flowchart TB
    subgraph Input["Video File (.mp4, .mkv, etc.)"]
        V[Video Stream]
        A[Audio Stream]
    end

    subgraph Stage1["Stage 1: Structural Analysis"]
        N[NormalizeWave<br/>FFprobe metadata]
        SD[ShotDetectionWave<br/>Scene cuts via FFmpeg]
        KE[KeyframeExtractionWave<br/>I-frame + dedup]
    end

    subgraph Stage2["Stage 2: Content Extraction"]
        IS[ImageSummarizer<br/>CLIP, OCR, Vision]
        AS[AudioSummarizer<br/>Whisper, Diarization]
        NER[NER Service<br/>Entity extraction]
    end

    subgraph Stage3["Stage 3: Scene Assembly"]
        SC[SceneClusteringWave<br/>CLIP similarity]
        EV[EvidenceGenerationWave<br/>RAG chunks]
    end

    V --> N --> SD --> KE
    KE --> IS
    A --> AS
    AS --> NER
    IS --> SC
    NER --> SC
    SC --> EV

    style Stage1 stroke:#22c55e,stroke-width:2px
    style Stage2 stroke:#3b82f6,stroke-width:2px
    style Stage3 stroke:#8b5cf6,stroke-width:2px

प्रमाण-सामग्री तैयार की गई

क्रियान्वयन में घुसने से पहले

आकृति कुंजी फ़ील्ड्स स्रोत
दृश्य id, start_time, end_time, key_terms[], speaker_ids[], embedding[512] SceneClusteringWave
गोलीबारी id, start_time, end_time, cut_type, keyframe_path ShotDetectionWave
अभिव्यक्ति id, text, start_time, end_time, speaker_id, confidence ट्रांस्क्रिप्शन वेव
पाठ ट्रैक id, text, start_time, text_type उपशीर्षकExtractionWave
कुंजीफ्रेम id, timestamp, frame_path, dhash, clip_embedding[512] कुंजीफ्रेम एक्सट्रैक्शन वेव

प्रत्येक वास्तु में शामिल है उद्गमस्रोत तरंग,processing timestamp,confidence score.This is the "evidence ledger"that downstream RAG queries operate on

संकेत---अभिज्ञ तरंग पाइपलाइन

वीडियोसममैजर एक का उपयोग करता है संकेत---आधारित तरंग संरचना जहाँ प्रत्येक तरंग अपने संकेत संविदाओं को स्पष्ट रूप से घोषित करता है

public interface ISignalAwareVideoWave
{
    /// <summary>Signals this wave requires before it can run.</summary>

    IReadOnlyList<string> RequiredSignals { get; }

    /// <summary>Signals this wave can optionally use if available.</summary>

    IReadOnlyList<string> OptionalSignals { get; }

    /// <summary>Signals this wave emits on successful completion.</summary>

    IReadOnlyList<string> EmittedSignals { get; }

    /// <summary>Cache keys this wave produces for downstream waves.</summary>

    IReadOnlyList<string> CacheEmits { get; }

    /// <summary>Cache keys this wave consumes from upstream waves.</summary>

    IReadOnlyList<string> CacheUses { get; }
}

यह सक्षम करता है गतिमान तरंग समन्वय:

  • यदि आवश्यक संकेतों की कमी है तो वेव्स स्वचालित रूप से छोड़ें
  • रनटाइम पर निर्भरताओं को हल करता है (नहीं हार्डकोड क्रमादेश
  • आंशिक पुनरावृत्ति संसूचक कुंजी द्वारा कैश (पुनरीकृत होती है
  • UI प्रगति कटीकता: प्रत्येक तरंग स्वतंत्र रूप से प्रगति उत्सर्जित करता है

16-वेव पाइपलाइन

कुंजीफ्रेम निष्कर्षण बेहतर समांतरता और कैश दक्षता के लिए 7 कणिका तरंगों के रूप में कार्यान्वित किया जाता है

तरंग प्राथमिकता आवश्यकताएं МSK3 उत्प्रेषण एमSK4 समय एमSK5
सामान्यizeWave 1000 - video.duration, video.fps, video.normalized
FFmpegShotDetectionWave 900 video.normalized shots.detected, shots.count
आईएफआरएम डिटेक्शन वेव 850 video.normalized keyframes.iframes_detected, keyframes.iframes_count
कुंजीफ्रेम चयन तरंग 840 shots.detected, keyframes.iframes_detected keyframes.selected, keyframes.selected_count
थम्बनेल एक्सट्रैक्शन वेव 830 keyframes.selected keyframes.thumbnails_extracted
कुंजीफ्रेम डेड्यूप्लिकेशन वेव 820 keyframes.thumbnails_extracted keyframes.deduplicated, keyframes.duplicates_skipped
कुंजीफ्रेमFullResExtractionWave 810 keyframes.deduplicated keyframes.extracted, keyframes.count
क्लिप एम्बेडिंग वेव 800 keyframes.extracted clip.embeddings_ready, clip.embeddings_count
छवि विश्लेषण तरंग 790 keyframes.deduplicated keyframes.analyzed, ocr.extracted
TitleCreditsDetectionWave 750 shots.detected title.detected, credits.detected
ऑडियो एक्सट्रैक्शन वेव 650 video.normalized audio.extracted, audio.path
ट्रांस्क्रिप्शन वेव 600 audio.extracted transcription.complete, transcription.utterance_count
उपशीर्षक एक्सट्रैक्शन वेव 550 video.normalized subtitles.extracted
अध्याय एक्सट्रैक्शन वेव 500 video.normalized chapters.extracted
दृश्य क्लस्टरिंग वेव 400 shots.detected scenes.detected, scene.count
प्रमाणGenerationWave 100 scenes.detected evidence.generated

नोट्स

  • ImageAnalysisWave प्रयोग करता है keyframes.deduplicated OCR थम्बनेल पर चलता है
  • तरंग 3-7 कैश दक्षता के लिए कुंजी फ्रेम एक्सटेक्शन है

घंटे के लिए कुल 2- चलचित्र: ~10-15 मिनट (vs. optimization के बिना घंटों

Well-Known संकेत कुंजी

संकेतों को संगतता के लिए स्थिरांक के रूप में परिभाषित किया जाता है

public static class VideoSignals
{
    // NormalizeWave signals
    public const string VideoDuration = "video.duration";
    public const string VideoFps = "video.fps";
    public const string VideoNormalized = "video.normalized";

    // Shot detection signals
    public const string ShotsDetected = "shots.detected";
    public const string ShotsCount = "shots.count";

    // Keyframe signals
    public const string IframesDetected = "keyframes.iframes_detected";
    public const string KeyframesSelected = "keyframes.selected";
    public const string KeyframesDeduplicated = "keyframes.deduplicated";
    public const string KeyframesExtracted = "keyframes.extracted";

    // CLIP embedding signals
    public const string ClipEmbeddingsReady = "clip.embeddings_ready";

    // Scene clustering signals
    public const string ScenesDetected = "scenes.detected";
    public const string SceneCount = "scene.count";

    // Transcription signals
    public const string TranscriptionComplete = "transcription.complete";
}

क्षमता प्रणाली: लचीला मॉडल & रूटिंग

वीडियोसममैजर एक का उपयोग करता है क्षमता-आधारित वास्तुकलाआरंभ पर एक बार GPU पता लगाएँ

मॉडल मैनिफेस्ट (YAML

मॉडलों में परिभाषित किए जाते हैं models.yamlकोड में कोई जादू स्ट्रिंग नहीं

# models.yaml (excerpt)
models:
  clip-vit-b32:
    name: "CLIP ViT-B/32"
    download_url: "https://huggingface.co/openai/clip-vit-base-patch32/resolve/main/onnx/visual_model.onnx"
    preferred_providers: [CUDAExecutionProvider, DmlExecutionProvider, CPUExecutionProvider]

components:
  ClipEmbeddingWave:
    models: [clip-vit-b32]
    fallback_chain: [ImageAnalysisWave]
// Type-safe constants (no raw strings)
await coordinator.EnsureModelAsync(ModelIds.ClipVitB32);
await coordinator.ActivateWaveAsync(ComponentIds.TranscriptionWave);

// Route with fallback
var route = await coordinator.RouteWorkAsync(new[]
{
    ComponentIds.ClipEmbeddingWave,    // Primary (GPU)
    ComponentIds.ImageAnalysisWave     // Fallback (CPU)
});

पाइपलाइन दक्षता परमाणु

दर सीमान, समय अनुमान, और अनुकूली बैकप्रिश्रण UI प्रतिक्रियात्मक बनाए रखते हैं जबकि अधिकतम स्ट्राइपट

// Time estimation from actual data
var estimator = CapabilityAtoms.CreateTimeEstimator();
using (estimator.Time("clip_embedding")) { await ProcessAsync(); }
var eta = estimator.GetEstimate("clip_embedding", remaining: 50);
// eta.Estimated, eta.Optimistic, eta.Pessimistic, eta.Confidence

पूर्ण क्षमता प्रणाली डक: देखें Mostlylucid.Summarizer.Core/Capabilities/ GPU पता लगाने के लिए, सिग्नल pub/subM SK2 बैकप्रेस कंट्रोलर्सMSC3 और मेश टोपोलॉजी डिजाइन


कुंजी अनुकूलन 1: परिकल्पित हैश डुप्लिकेटेशन

expensive CLIP embeddings चलाने से पहले, VideoSummarizer का उपयोग कर दृश्य रूप में समान फ्रेमों को फिल्टर करता है अंतर हैश (dHash).

dHash कैसे काम करता है

public class KeyframeDeduplicationService
{
    // dHash parameters: 9x8 grayscale = 64 bits
    private const int HashWidth = 9;
    private const int HashHeight = 8;
    private const int DefaultHammingThreshold = 10;

    public async Task<ulong> ComputeDHashAsync(string imagePath, CancellationToken ct)
    {
        using var image = Image.Load<Rgba32>(imagePath);

        // Resize to 9x8 (one extra column for gradient comparison)
        image.Mutate(x => x
            .Resize(HashWidth, HashHeight)
            .Grayscale());

        ulong hash = 0;
        int bit = 0;

        // Compare adjacent pixels horizontally
        for (int y = 0; y < HashHeight; y++)
        {
            for (int x = 0; x < HashWidth - 1; x++)
            {
                var left = image[x, y].R;
                var right = image[x + 1, y].R;

                // Set bit if left pixel is brighter than right
                if (left > right)
                {
                    hash |= (1UL << bit);
                }
                bit++;
            }
        }

        return hash;
    }

    public static int HammingDistance(ulong a, ulong b) =>
        BitOperations.PopCount(a ^ b);
}

उदाहरण आउटपुट

Input: 50 keyframe candidates (from codec I-frames)

Deduplication (Hamming threshold 10):
  Frame 0: hash=0x8f3a2c1d → KEEP (first frame)
  Frame 1: hash=0x8f3a2c1e → SKIP (distance=1 from frame 0)
  Frame 2: hash=0x8f3a2c1f → SKIP (distance=2 from frame 0)
  Frame 3: hash=0xc7e1b4a2 → KEEP (distance=28 from frame 0)
  ...

Result: 50 → 30 frames (40% reduction)
Processing saved: ~8 seconds of CLIP inference

यह क्यों महत्वपूर्ण है

  • ~40% फ्रेम घटाना विशिष्ट सामग्री पर
  • प्रति फ्रेम <1ms हैश गणना के लिए (vs.
  • अतिरिक्त फ्रेम फ़िल्टर करता है पहले expensive GPU operations

कुंजी अनुकूलन 2: बैच क्लिप सम्मिलन

एक समय में एक छवि को प्रसंस्करण करने के बजाय

बैच प्रोसेसिंग वास्तुशिल्प

public class BatchClipEmbeddingService
{
    private const int ClipImageSize = 224;
    private const int DefaultBatchSize = 8; // 8 images per GPU pass

    public async Task<Dictionary<int, float[]>> GenerateBatchEmbeddingsAsync(
        Dictionary<int, string> framePaths,
        int batchSize = DefaultBatchSize,
        CancellationToken ct = default)
    {
        var session = await GetOrLoadClipModelAsync(ct);
        var results = new Dictionary<int, float[]>();

        // Pre-index batch for O(1) lookup (not batch.IndexOf!)
        var batches = framePaths
            .Select((kvp, idx) => (idx, kvp.Key, kvp.Value))
            .Chunk(batchSize);

        foreach (var batch in batches)
        {
            // Create batch tensor [batchSize, 3, 224, 224]
            var tensor = new DenseTensor<float>(new[] { batch.Length, 3, ClipImageSize, ClipImageSize });

            // Preprocess images in parallel (simplified; production uses vectorised span copy)
            Parallel.ForEach(batch, item =>
            {
                var (batchIdx, frameIndex, path) = item;
                var localIdx = batchIdx % batchSize;
                PreprocessImageToTensor(path, tensor, localIdx); // ImageSharp pixel buffers
            });

            // Single GPU pass for entire batch
            var inputs = new List<NamedOnnxValue>
            {
                NamedOnnxValue.CreateFromTensor("input", tensor)
            };

            using var outputResults = session.Run(inputs);
            // Extract embeddings from batch output...
        }

        return results;
    }
}

निष्पादन तुलना

Input: 30 keyframes (after deduplication)

Serial processing (1 frame at a time):
  30 × 200ms = 6,000ms (6.0 seconds)

Batch processing (8 frames per pass):
  4 batches × 350ms = 1,400ms (1.4 seconds)

Speedup: 4.3x

बैच प्रसंस्करण क्यों काम करता है

  • GPU समांतरता एकल-छवि निष्कर्ष के साथ कम उपयोग में है
  • बैच टेन्सर [8, 3, 224, 224] एकल छवि के रूप में एक ही GPU स्मृति का उपयोग करता है (अधिकतम)
  • ऑनिक्स रनटाइम आंतरिक रूप से बैच परिचालनों को अनुकूलित करता है

कुंजी अनुकूलन 3: पाइपलाइन संरचना

VideoSummarizer doesn't reinvent ImageSum marizer or AudioSommarizerit शंखियाँ them.

कुंजीफ्रेम उप-Pipeline: ImageSummarizer इंटीग्रेशन

कुंजी फ्रेम निष्कर्षण में विभाजित किया जाता है 7 ग्रेनोलर तरंगों को।

// IFrameDetectionWave → KeyframeSelectionWave → ThumbnailExtractionWave
// → KeyframeDeduplicationWave → KeyframeFullResExtractionWave → ClipEmbeddingWave

// ClipEmbeddingWave coordinates with ImageSummarizer
public class ClipEmbeddingWave : IVideoWave, ISignalAwareVideoWave
{
    public IReadOnlyList<string> RequiredSignals => [VideoSignals.KeyframesExtracted];
    public IReadOnlyList<string> EmittedSignals => [VideoSignals.ClipEmbeddingsReady];

    public async Task ProcessAsync(VideoContext context, CancellationToken ct)
    {
        var keyframes = context.GetCached<Dictionary<int, string>>("keyframes.paths");

        // Batch CLIP embedding (3-5x faster than serial)
        var embeddings = await _batchClipService.GenerateBatchEmbeddingsAsync(
            keyframes, batchSize: 8, ct);

        foreach (var (frameIndex, embedding) in embeddings)
            context.KeyframeEmbeddings[frameIndex] = embedding;
    }
}

// ImageAnalysisWave runs ImageSummarizer on deduplicated frames
public class ImageAnalysisWave : IVideoWave, ISignalAwareVideoWave
{
    public IReadOnlyList<string> RequiredSignals => [VideoSignals.KeyframesDeduplicated];

    public async Task ProcessAsync(VideoContext context, CancellationToken ct)
    {
        var keyframePaths = context.GetCached<List<string>>("keyframes.deduplicated_paths");

        foreach (var path in keyframePaths)
        {
            // Run ImageSummarizer for OCR, vision, captions
            var result = await _imageOrchestrator.AnalyzeAsync(path, ct);
            context.SetCached($"image_analysis.{Path.GetFileName(path)}", result);
        }
    }
}

ट्रांस्क्रिप्शन वेव

ऑडियो निष्कर्षण और ट्रांसक्रिप्शन अब अलग संकेत हैं

// AudioExtractionWave runs first (extracts audio track from video)
public class AudioExtractionWave : IVideoWave, ISignalAwareVideoWave
{
    public IReadOnlyList<string> RequiredSignals => [VideoSignals.VideoNormalized];
    public IReadOnlyList<string> EmittedSignals => ["audio.extracted", "audio.path"];

    public async Task ProcessAsync(VideoContext context, CancellationToken ct)
    {
        var audioPath = await _ffmpegService.ExtractAudioAsync(
            context.VideoPath, context.WorkingDirectory, ct);
        context.SetCached("audio.path", audioPath);
    }
}

// TranscriptionWave depends on audio.extracted signal
public class TranscriptionWave : IVideoWave, ISignalAwareVideoWave
{
    public IReadOnlyList<string> RequiredSignals => ["audio.extracted"];
    public IReadOnlyList<string> EmittedSignals => [
        VideoSignals.TranscriptionComplete,
        "transcription.utterance_count"
    ];

    public async Task ProcessAsync(VideoContext context, CancellationToken ct)
    {
        var audioPath = context.GetCached<string>("audio.path");

        // Run AudioSummarizer pipeline (Whisper + diarization)
        var audioProfile = await _audioOrchestrator.AnalyzeAsync(audioPath, ct);

        // Extract utterances with speaker info
        var turns = audioProfile.GetValue<List<SpeakerTurn>>("speaker.turns");
        foreach (var turn in turns ?? [])
        {
            context.Utterances.Add(new Utterance
            {
                Id = Guid.NewGuid(),
                Text = turn.Text,
                StartTime = turn.StartSeconds,
                EndTime = turn.EndSeconds,
                SpeakerId = turn.SpeakerId,
                Confidence = turn.Confidence
            });
        }

        // Run NER on full transcript for entity extraction
        var transcript = audioProfile.GetValue<string>("transcription.full_text");
        if (!string.IsNullOrEmpty(transcript))
        {
            var entities = await _nerService.ExtractEntitiesAsync(transcript, ct);
            context.SetCached("transcript_entities", entities);

            // Emit entity signals by type (PER, ORG, LOC, MISC)
            foreach (var group in entities.GroupBy(e => e.Type))
            {
                context.AddSignal($"transcript.entities.{group.Key.ToLowerInvariant()}",
                    group.Select(e => e.Text).Distinct().ToList());
            }
        }
    }
}

एनआर इंटीग्रेशनएमएसके0 नामित इकाई पहचान

VideoSummarizer BERT-based NER के प्रयोग से नामित इकाइयों को ट्रांस्क्रिप्टों से निकालता है

OnnxNerService

public class OnnxNerService
{
    // Model: dslim/bert-base-NER (ONNX exported)
    // Entities: PER (Person), ORG (Organization), LOC (Location), MISC (Miscellaneous)

    public async Task<List<EntitySpan>> ExtractEntitiesAsync(string text, CancellationToken ct)
    {
        var entities = new List<EntitySpan>();

        // Chunk long text (BERT max 512 tokens)
        foreach (var chunk in ChunkText(text, maxTokens: 400, overlap: 50))
        {
            // Tokenize with WordPiece
            var tokens = _tokenizer.Tokenize(chunk);

            // Run ONNX inference
            var inputs = PrepareInputs(tokens);
            using var results = _session.Run(inputs);

            // Decode BIO tags
            var predictions = DecodePredictions(results);
            var chunkEntities = ExtractEntitySpans(tokens, predictions);

            entities.AddRange(chunkEntities);
        }

        // Deduplicate entities
        return entities
            .GroupBy(e => (e.Text.ToLowerInvariant(), e.Type))
            .Select(g => g.First())
            .ToList();
    }
}

उदाहरण आउटपुट

Transcript: "Today we're speaking with John Smith from Microsoft about
their new AI lab in Seattle. The project, codenamed Phoenix, builds
on research from Stanford University."

Entities extracted:
  PER: John Smith
  ORG: Microsoft, Stanford University
  LOC: Seattle
  MISC: Phoenix

Signals emitted:
  transcript.entities.per = ["John Smith"]
  transcript.entities.org = ["Microsoft", "Stanford University"]
  transcript.entities.loc = ["Seattle"]
  transcript.entities.misc = ["Phoenix"]

वीडियो के लिए क्यों एनईआर महत्वपूर्ण है

  • जैसे क्वेरी सक्षम करें " माइक्रोसॉफ्ट का उल्लेख करने वाले वीडियो ढूंढें
  • DocSummarizer अस्तित्व ग्राफ के लिए लिंक
  • LLM निष्कर्ष के बिना संरचनात्मक मेटाडेटा उपलब्ध कराता है

बहु--Signal दृश्य क्लस्टरिंग

दृश्य पता लगाने के लिए केवल सीलिप एम्बेडिंग का उपयोग करने से यह काम नहीं करता।

वीडियोसममैजर एक का उपयोग करता है बहु-- सिग्नल दृष्टिकोण जो मजबूत दृश्य सीमा पता लगाने के लिए 4 भारित संकेतों को जोड़ता है

SceneClusteringWave: Multi-Signal Architecture

public class SceneClusteringWave : IVideoWave, ISignalAwareVideoWave
{
    // Signal weights for boundary scoring
    private const double EmbeddingWeight = 0.4;   // CLIP embedding dissimilarity
    private const double TranscriptWeight = 0.3;  // Semantic shift in transcript
    private const double CutTypeWeight = 0.2;     // Fade/dissolve detection
    private const double TemporalWeight = 0.1;    // Time since last scene

    // Temporal constraints
    private const double MinSceneDuration = 15.0;   // Don't split scenes < 15s
    private const double MaxSceneDuration = 300.0;  // Force split at 5 minutes
    private const double TargetSceneDuration = 90.0; // Prefer ~90s scenes

    public IReadOnlyList<string> RequiredSignals => [VideoSignals.ShotsDetected];
    public IReadOnlyList<string> OptionalSignals => [
        VideoSignals.ClipEmbeddingsReady,
        VideoSignals.TranscriptionComplete,
        VideoSignals.KeyframesDeduplicated
    ];
    public IReadOnlyList<string> EmittedSignals => [
        VideoSignals.ScenesDetected,
        "scene.count",
        "scene.avg_duration",
        "scene.clustering_method"
    ];

    private List<(int shotIndex, double score)> ComputeBoundaryScores(VideoContext context)
    {
        var shots = context.Shots.OrderBy(s => s.StartTime).ToList();
        var scores = new List<(int, double)>();

        // Build embedding map with nearest-neighbor interpolation
        var shotEmbeddings = PropagateEmbeddingsToNearbyShots(context, shots);

        // Build transcript windows for semantic shift detection
        var transcriptWindows = BuildTranscriptWindows(context, shots, windowSeconds: 10);

        for (int i = 0; i < shots.Count - 1; i++)
        {
            double score = 0;
            var currentShot = shots[i];
            var nextShot = shots[i + 1];

            // 1. Embedding dissimilarity (40%)
            if (shotEmbeddings.TryGetValue(i, out var currentEmbed) &&
                shotEmbeddings.TryGetValue(i + 1, out var nextEmbed))
            {
                var similarity = CosineSimilarity(currentEmbed, nextEmbed);
                score += (1.0 - similarity) * EmbeddingWeight;
            }

            // 2. Transcript semantic shift (30%)
            if (transcriptWindows.TryGetValue(i, out var currentWords) &&
                transcriptWindows.TryGetValue(i + 1, out var nextWords))
            {
                var overlap = currentWords.Intersect(nextWords).Count();
                var union = currentWords.Union(nextWords).Count();
                var jaccard = union > 0 ? (double)overlap / union : 0;
                score += (1.0 - jaccard) * TranscriptWeight;
            }

            // 3. Cut type signal (20%) - fades/dissolves suggest scene boundaries
            if (currentShot.CutType is "fade" or "dissolve")
            {
                score += CutTypeWeight;
            }

            // 4. Temporal pressure (10%) - encourage splits near target duration
            var timeSinceLastScene = currentShot.EndTime - GetLastSceneBoundary();
            if (timeSinceLastScene > TargetSceneDuration)
            {
                var pressure = Math.Min(1.0, (timeSinceLastScene - TargetSceneDuration) / 60);
                score += pressure * TemporalWeight;
            }

            scores.Add((i, score));
        }

        return scores;
    }
}

प्रमुख नवान्वेषण

  1. निकटतम- पड़ोसी सम्मिलन प्रसारकेवल शॉटों में ~2% सीधी CLIP एम्बेडिंग होती है

  2. ट्रांस्क्रिप्ट सेमेटिक विंडोज़प्रत्येक शॉट के चारों ओर से एक सेकंड शब्द विंडो बनाता है और Jaccard दूरी के माध्यम से अर्थिक परिवर्तनों को पता लगाता है एक सस्ता अर्थिक ड्रिफ्ट प्रॉक्सी।

  3. काटें प्रकार सचेतनता: फेडM SK1to-black and dissolve transitions strongly indicate scene boundaries

  4. अनुकूलित बन्धनएक नियत थ्रेसहोल्ड के बजाय : शीर्ष से किनारे चुनता है

  5. अस्थायी अवरोधन्यूनतम को लागू करता है 15\s दृश्यों और सीमाओं पर बल देता है

उदाहरण:

Input: 1881 shots from a 2-hour movie
       39 keyframes with CLIP embeddings
       2302 utterances from transcript

Boundary scoring per shot:
  Shot 45-46: embedding=0.15, transcript=0.32, cut=0.0, temporal=0.0 → score=0.156
  Shot 46-47: embedding=0.08, transcript=0.12, cut=0.0, temporal=0.0 → score=0.068
  Shot 47-48: embedding=0.35, transcript=0.41, cut=0.2, temporal=0.05 → score=0.388 ← BOUNDARY
  ...

Adaptive threshold (top 25%): 0.25
Natural boundaries found: 45

Output: 47 scenes (avg 2.6 minutes per scene)
  - Min scene: 15.2s
  - Max scene: 298.4s
  - Total coverage: 100%

Signals:
  scenes.detected = true
  scene.count = 47
  scene.avg_duration = 156.3
  scene.clustering_method = "multi_signal_weighted"

वीडियो संकेत संविदा

वीडियोसममेरिजर इमेजसममेराइजर और ऑडियोसममेरिसर से संकेत संविदा बढ़ाता है

public record VideoSignal
{
    public required string Key { get; init; }      // "scene.count", "transcript.entities.per"
    public object? Value { get; init; }
    public double Confidence { get; init; } = 1.0;
    public required string Source { get; init; }   // "SceneClusteringWave"

    // Video-specific: time range
    public double? StartTime { get; init; }
    public double? EndTime { get; init; }

    public DateTime Timestamp { get; init; }
    public Dictionary<string, object>? Metadata { get; init; }
    public List<string>? Tags { get; init; }       // ["visual", "scene"]
}

public static class VideoSignalTags
{
    public const string Visual = "visual";
    public const string Audio = "audio";
    public const string Speech = "speech";
    public const string Ocr = "ocr";
    public const string Motion = "motion";
    public const string Scene = "scene";
    public const string Shot = "shot";
    public const string Metadata = "metadata";
}

जारी कुंजी संकेत

संकेत स्रोत विवरण
video.duration सामान्यizeWave सेकेंडों में कुल अवधि
video.resolution NormalizeWave WidthM SK2Height
video.fps सामान्यizeWave फ्रेम दर
shots.count ShotDetectionWave
keyframes.count कुंजी फ्रेम एक्सट्रैक्शन वेव
keyframes.duplicates_skipped KeyframeExtractionWave dHash द्वारा फिल्टर किए गए फ्रेम
scene.count SceneClusteringWave
transcript.entities.per ट्रांस्क्रिप्शन वेव एनईआर से व्यक्तियों के नाम
transcript.entities.org ट्रांस्क्रिप्शन वेव संगठन नाम
transcript.word_count ट्रांस्क्रिप्शन वेव ट्रांसक्रिप्सन में कुल शब्द

वीडियोपीपिलेन: RAG आउटपुट

VideoPipeline वीडियो संकेतों में बदलता है ContentChunk RAG सूचकीकरण के लिए:

public class VideoPipeline : PipelineBase
{
    public override string PipelineId => "video";
    public override IReadOnlySet<string> SupportedExtensions => new HashSet<string>
    {
        ".mp4", ".mkv", ".avi", ".mov", ".wmv", ".webm", ".flv", ".m4v", ".mpeg", ".mpg"
    };

    private List<ContentChunk> BuildContentChunks(VideoContext context, string filePath)
    {
        var chunks = new List<ContentChunk>();

        // 1. Scene-based chunks (best for video retrieval)
        foreach (var scene in context.Scenes)
        {
            var sceneText = BuildSceneText(context, scene);
            var embedding = context.GetCached<float[]>($"scene_centroid.{scene.Id}");

            chunks.Add(new ContentChunk
            {
                Text = sceneText,
                ContentType = ContentType.Summary,
                Embedding = embedding,  // Proper vector column, not metadata
                Metadata = new Dictionary<string, object?>
                {
                    ["source"] = "video_scene",
                    ["scene_id"] = scene.Id,
                    ["key_terms"] = scene.KeyTerms,
                    ["speakers"] = scene.SpeakerIds,
                    ["start_time"] = scene.StartTime,
                    ["end_time"] = scene.EndTime
                }
            });
        }

        // 2. Transcript chunks (1-minute windows)
        var transcriptChunks = BuildTranscriptChunks(context, filePath);
        chunks.AddRange(transcriptChunks);

        // 3. Text track chunks (on-screen text/subtitles)
        foreach (var textTrack in context.TextTracks)
        {
            chunks.Add(new ContentChunk
            {
                Text = $"On-screen text: {textTrack.Text}",
                ContentType = ContentType.ImageOcr,
                Metadata = new Dictionary<string, object?>
                {
                    ["source"] = "video_ocr",
                    ["text_type"] = textTrack.TextType.ToString(),
                    ["start_time"] = textTrack.StartTime
                }
            });
        }

        return chunks;
    }

    private string BuildSceneText(VideoContext context, SceneSegment scene)
    {
        var parts = new List<string>();

        if (!string.IsNullOrEmpty(scene.Label))
            parts.Add($"Scene: {scene.Label}");

        parts.Add($"[{FormatTime(scene.StartTime)} - {FormatTime(scene.EndTime)}]");

        if (scene.KeyTerms.Count > 0)
            parts.Add($"Topics: {string.Join(", ", scene.KeyTerms)}");

        // Add utterances in this scene
        var sceneUtterances = context.Utterances
            .Where(u => u.StartTime >= scene.StartTime && u.EndTime <= scene.EndTime)
            .OrderBy(u => u.StartTime);

        if (sceneUtterances.Any())
            parts.Add($"Speech: {string.Join(" ", sceneUtterances.Select(u => u.Text))}");

        return string.Join("\n", parts);
    }
}

फिल्म के लिए उदाहरण आउटपुट:

{
  "chunks": [
    {
      "text": "Scene: Opening montage\n[0:00 - 2:34]\nTopics: city, night, traffic\nSpeech: The year is 2049. The world has changed.",
      "contentType": "Summary",
      "metadata": {
        "source": "video_scene",
        "scene_id": "abc123",
        "key_terms": ["city", "night", "traffic"],
        "start_time": 0.0,
        "end_time": 154.0
      }
    },
    {
      "text": "The detective arrived at the crime scene. Forensics had already processed the area.",
      "contentType": "Transcript",
      "metadata": {
        "source": "video_transcript",
        "time_window": "2:34 - 3:34",
        "utterance_count": 4
      }
    },
    {
      "text": "On-screen text: LOS ANGELES 2049",
      "contentType": "ImageOcr",
      "metadata": {
        "source": "video_ocr",
        "text_type": "Title"
      }
    }
  ]
}

निष्पादन विशेषताएँ

प्रक्रमण समय (2-घण्टे चलचित्रM SK1 1080p)

चरण |-------|------|-------| FFprobe मेटाडेटा Shot detection कुंजी फ्रेम एक्सटेक्शन डीहश डुप्लिकेट बैच सीलिप अंतःकरण | ~60 s पाठ के साथ कुंजीफ्रेम | ImageSummarizer OCR ऑडियो एक्सटेक्शन | हिस्सों की प्रतिलिपि स्पीकर डायराइजेशन ट्रांस्क्रिप्ट पर | दृश्य क्लस्टरिंग प्रमाण सृजन | कुल | मिनट | |

अनुकूलन किए बिना

अनुकूलन |--------------|---------| | dHash डेड्यूप्लिकेशन | | | \ ~40% | फ्रेमों को फिल्टर किया गया |= | बैच क्लिप | पाइपलाइन कम्पोशन | कुल बचत | मिनट |

स्मृति उपयोग

अवयव |-----------|--------| CLIP ViT | फुसफुसा आधार 0 ECAPA 1 TDNN 2 ~100 MB 4 0 BERT 1 NER 2 3 MB 4 | शिखर | एमएसके0जीबी |


lucidRAG के साथ एकीकरण

वीडियोसममैजर एक के रूप में पंजीकृत करता है IPipeline स्वचालित रूटिंग के लिए

// In Program.cs
builder.Services.AddDocSummarizer(builder.Configuration.GetSection("DocSummarizer"));
builder.Services.AddDocSummarizerImages(builder.Configuration.GetSection("Images"));
builder.Services.AddVideoSummarizer();  // NEW
builder.Services.AddPipelineRegistry(); // Must be last

// Auto-routing by extension
var registry = services.GetRequiredService<IPipelineRegistry>();
var pipeline = registry.FindForFile("movie.mp4");  // Returns VideoPipeline
var result = await pipeline.ProcessAsync("movie.mp4");

समर्थित एक्सटेंशन

  • .mp4, .mkv, .avi, .mov, .wmv, .webm, .flv, .m4v, .mpeg, .mpg

क्या आप प्राप्त करते हैं

  • दृश्य-स्तर RAG टुकड़े: ट्रांस्क्रिप्टों के साथ सुसंगत खंड
  • बहु---मोडल साक्ष्यदृश्यात्मक (keyframe embeddings), audio (speaker diarization),text (OCR subtitles
  • नामित इकाई: व्यक्तियोंM SK1 संगठनों, स्थानों से ट्रांस्क्रिप्ट एनईआर
  • लेखा-परीक्षा योग्य उद्गम: प्रत्येक संकेत में स्रोत तरंग है
  • कुशल प्रसंस्करणएक घंटे की फिल्म के लिए मिनट

यह कितना खर्च करता है

  • ~1.5GB GPU स्मृति सभी ऑनिक्स मॉडलों के लिए
  • मिनट संसाधन प्रति 2- घंटे फिल्म
  • डिस्क स्थान मध्यवर्ती फ़ाइलों के लिए (स्वचालित रूप से साफ किया जाता है
  • जटिलता: तीन पाइपलाइनों को व्यवस्थित करने के लिए वेव निर्भरताओं को समझना आवश्यक है

निष्कर्ष

वीडियोसममाराइजर यह प्रदर्शित करता है कि पाइपलाइन संरचना स्केल

  1. विशिष्ट पाइपलाइनों का पुनः प्रयोग: Don' ImageSummarizer या AudioSummarzer उन्हें पुनः खोजना नहीं है
  2. महंगे ऑपरेशनों से पहले फ़िल्टर करें: dHash डुप्लीकेट व्यय |<1 |ms | , |सीलिप निष्कर्षों को बचाता है
  3. बैच जीपीयू ऑपरेशनप्रति पास छवियाँ : 8 speedup
  4. सामग्री से पहले संरचना निकालेंShots → Scenes → Evidence ( No raw frames LLM
  5. लचीला मॉडल प्रबंधन: नमूने को तभी डाउनलोड करें जब आवश्यक हो
  6. प्रतिक्रियात्मक रूटिंग: उपलब्ध घटकों के लिए रूट कार्य

परिणाम यह होता है कि एक घंटे की फिल्म दृश्यों के साथ एक संरचनात्मक संकेत लाईडर बन जाती है।

  • दृश्य ढूंढें जहाँ जॉन स्मिथ माइक्रोसॉफ्ट पर चर्चा करता है
  • फ़िनिक्स परियोजना के बारे में स्क्रीन पाठ के साथ क्लिप दिखाएँ
  • इस दृश्य के समान वीडियो ढूंढें

वीडियो के लिए कम RAG पैटर्न

Ingestion:  Video → 16 waves → Signals + Evidence (scenes, transcripts, entities)
Storage:    Signals (indexed) + Embeddings (CLIP, voice) + Evidence (chunks)
Query:      Filter (SQL) → Search (BM25 + vector) → Synthesize (LLM, ~5 results)

क्षमता प्रणाली

Startup:    Detect GPU → Load ModelManifest (YAML) → Initialize SignalSink
Activation: Component requests model → Lazy download → Signal "ModelAvailable"
Routing:    Route to best provider → Fallback chain → Backpressure control
Atoms:      Rate limiting + Time estimation + Pipeline balancing

यह है अवरोधित अस्पष्टता पैमाने पर

  • संभाव्य घटक संकेत प्रस्तावित करते हैं: एम्बेडिंग्स
  • निर्धारणात्मक अंकन संरचना को संगठित करता है: भारित सीमा प्राप्तियां, थ्रेसहोल्ड चयन
  • LLM (वैकल्पिक) प्रमाणों से सिन्थेसाइज़ करता है: सीमाबद्ध संदर्भ

एलएलएम पूर्व कम्प्यूटरी साक्ष्य नहीं कच्चा वीडियो पर कार्य करता है


संसाधन

lucidRAG दस्तावेज़ीकरण

संबंधित लाइब्रेरी

ऑनिक्स मॉडल

संबंधित लेख

कोर पैटर्न

कम RAG कार्यान्वयन


द सर्जरी

भाग पैटर्न
1 अवरोधित अस्पष्टता एकल घटक
2 प्रतिबंधित अस्पष्ट मोएम बहुआयामी अवयव
3 संदर्भ खींचना समय
4 छवि इंटेलिजेंस तरंग संरचना
4.1 Three-Tier OCR पाइपलाइन OCR
4.2 ऑडियोसममारीजरName फ़ॉरेंसिक ऑडियो
4.3 वीडियोसॉम्मारीजर (इस अनुच्छेद) वीडियो ऑर्केस्ट्रेशन, बैच क्लिप

अगला: बहुल-मोडल ग्राफ RAG lucidRAG सहित सभी चार सारांशकर्ताओं को एक एकीकृत ज्ञान ग्राफ में सम्मिलित करने के साथ क्रॉस

सभी भाग एक ही अपरिवर्ती का अनुसरण करते हैं संभाव्यात्मक घटक प्रस्तावित.

logo

© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.