स्थिति: विकास के एक भाग के रूप में सुस्पष्टRAG. स्रोत: github.comM SK1scottgal/lucidrag
जहाँ यह ठीक है: वीडियोसम्मैजर है वाद्यकार का सुस्पष्टRAG परिवार, तीन पाइपलाइनों को एक एकीकृत वीडियो विश्लेषण इंजन में सम्मिलित करने के लिए
सभी एक ही अनुसरण करें कम RAG पैटर्न: एक बार सिग्नलों को निकालने के लिए , साक्ष्य भंडारित करें
एक दो घण्टे फिल्म फ्रेम को प्रसंस्करण करने के लिए CLIP अंतःकरणों के साथ फ्रेम ले जाएगा और गणना में सैकड़ों डॉलर लगेंगे।
VideoSummarizer इस समस्या को तीन प्रमुख optimizations के साथ हल करता है
परिणाम : एक 2- घंटे फिल्म प्रक्रियाओं में |~10-15 | मिनटों | , | घंटों नहीं |M. के साथ ही वास्तुकला सिद्धांत ImageSummarizerComment और ऑडियोसममारीजरNameलेकिन एक एकीकृत वीडियो विश्लेषण पाइपलाइन में सम्मिलित
कोर अंतर्दृष्टि वीडियो छवियाँ है | + | ऑडियो |+ | पाठ & #44; . | प्रत्येक डोमेन को विशिष्ट उपकरणों के साथ संसाधित करें
- प्रक्रिया संरचना पहले (cuts
- एक बार क्रॉस-मोडल संकेत निकालें (embeddings
शब्दावली
(start_time, end_time) संकेत + संकेतककुंजी एमएल मॉडलों का उपयोग किया गया
इस लेख में शामिल है
संबंधित लेख:
मानदंडएमडीएएम पर मापा गया निम्न संख्याएँ 9950X | (16- | Core | | 3 | 4 | NVIDIA A | 5 | 6 | GB | 7 | 8 | 9 |GB RAM | 10 | NVMe | 11 | 12 | p H | 13 | Whisper base | 14 | आपका मील बदलेगा | 15
एक सामान्य फिल्म में शामिल है:
स्ट्रोमन दृष्टिकोण (कहीं ऐसा नहीं करता है
कुंजी फ्रेम एक्सटेक्शन के साथ भी (say, 500-1000 फ्रेमों से), कि' अभी भी serial CLIP inference के सेकेंड
पारंपरिक दृष्टिकोणएक्सट्रेक्ट कुंजी फ्रेम्स
समस्या: यह redundant फ्रेमों पर संगणना बर्न करता है.
समाधान: बहुआयामी,- चरण फिल्टरिंग,, बैच प्रसंस्करण और पाइपलाइन संरचना
वीडियोसममैजर एप्लेट कम RAG तीन चरण कम करने के साथ वीडियो के लिए
flowchart TB
subgraph Input["Video File (.mp4, .mkv, etc.)"]
V[Video Stream]
A[Audio Stream]
end
subgraph Stage1["Stage 1: Structural Analysis"]
N[NormalizeWave<br/>FFprobe metadata]
SD[ShotDetectionWave<br/>Scene cuts via FFmpeg]
KE[KeyframeExtractionWave<br/>I-frame + dedup]
end
subgraph Stage2["Stage 2: Content Extraction"]
IS[ImageSummarizer<br/>CLIP, OCR, Vision]
AS[AudioSummarizer<br/>Whisper, Diarization]
NER[NER Service<br/>Entity extraction]
end
subgraph Stage3["Stage 3: Scene Assembly"]
SC[SceneClusteringWave<br/>CLIP similarity]
EV[EvidenceGenerationWave<br/>RAG chunks]
end
V --> N --> SD --> KE
KE --> IS
A --> AS
AS --> NER
IS --> SC
NER --> SC
SC --> EV
style Stage1 stroke:#22c55e,stroke-width:2px
style Stage2 stroke:#3b82f6,stroke-width:2px
style Stage3 stroke:#8b5cf6,stroke-width:2px
क्रियान्वयन में घुसने से पहले
| आकृति | कुंजी फ़ील्ड्स | स्रोत | ||
|---|---|---|---|---|
| दृश्य | id, start_time, end_time, key_terms[], speaker_ids[], embedding[512] SceneClusteringWave |
|||
| गोलीबारी | id, start_time, end_time, cut_type, keyframe_path ShotDetectionWave |
|||
| अभिव्यक्ति | id, text, start_time, end_time, speaker_id, confidence ट्रांस्क्रिप्शन वेव |
|||
| पाठ ट्रैक | id, text, start_time, text_type उपशीर्षकExtractionWave |
|||
| कुंजीफ्रेम | id, timestamp, frame_path, dhash, clip_embedding[512] कुंजीफ्रेम एक्सट्रैक्शन वेव |
प्रत्येक वास्तु में शामिल है उद्गमस्रोत तरंग,processing timestamp,confidence score.This is the "evidence ledger"that downstream RAG queries operate on
वीडियोसममैजर एक का उपयोग करता है संकेत---आधारित तरंग संरचना जहाँ प्रत्येक तरंग अपने संकेत संविदाओं को स्पष्ट रूप से घोषित करता है
public interface ISignalAwareVideoWave
{
/// <summary>Signals this wave requires before it can run.</summary>
IReadOnlyList<string> RequiredSignals { get; }
/// <summary>Signals this wave can optionally use if available.</summary>
IReadOnlyList<string> OptionalSignals { get; }
/// <summary>Signals this wave emits on successful completion.</summary>
IReadOnlyList<string> EmittedSignals { get; }
/// <summary>Cache keys this wave produces for downstream waves.</summary>
IReadOnlyList<string> CacheEmits { get; }
/// <summary>Cache keys this wave consumes from upstream waves.</summary>
IReadOnlyList<string> CacheUses { get; }
}
यह सक्षम करता है गतिमान तरंग समन्वय:
कुंजीफ्रेम निष्कर्षण बेहतर समांतरता और कैश दक्षता के लिए 7 कणिका तरंगों के रूप में कार्यान्वित किया जाता है
| तरंग | प्राथमिकता | आवश्यकताएं | МSK3 | उत्प्रेषण | एमSK4 | समय | एमSK5 | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| सामान्यizeWave | 1000 | - | video.duration, video.fps, video.normalized |
||||||||
| FFmpegShotDetectionWave | 900 | video.normalized |
shots.detected, shots.count |
||||||||
| आईएफआरएम डिटेक्शन वेव | 850 | video.normalized |
keyframes.iframes_detected, keyframes.iframes_count |
||||||||
| कुंजीफ्रेम चयन तरंग | 840 | shots.detected, keyframes.iframes_detected |
keyframes.selected, keyframes.selected_count |
||||||||
| थम्बनेल एक्सट्रैक्शन वेव | 830 | keyframes.selected |
keyframes.thumbnails_extracted |
||||||||
| कुंजीफ्रेम डेड्यूप्लिकेशन वेव | 820 | keyframes.thumbnails_extracted |
keyframes.deduplicated, keyframes.duplicates_skipped |
||||||||
| कुंजीफ्रेमFullResExtractionWave | 810 | keyframes.deduplicated |
keyframes.extracted, keyframes.count |
||||||||
| क्लिप एम्बेडिंग वेव | 800 | keyframes.extracted |
clip.embeddings_ready, clip.embeddings_count |
||||||||
| छवि विश्लेषण तरंग | 790 | keyframes.deduplicated |
keyframes.analyzed, ocr.extracted |
||||||||
| TitleCreditsDetectionWave | 750 | shots.detected |
title.detected, credits.detected |
||||||||
| ऑडियो एक्सट्रैक्शन वेव | 650 | video.normalized |
audio.extracted, audio.path |
||||||||
| ट्रांस्क्रिप्शन वेव | 600 | audio.extracted |
transcription.complete, transcription.utterance_count |
||||||||
| उपशीर्षक एक्सट्रैक्शन वेव | 550 | video.normalized |
subtitles.extracted |
||||||||
| अध्याय एक्सट्रैक्शन वेव | 500 | video.normalized |
chapters.extracted |
||||||||
| दृश्य क्लस्टरिंग वेव | 400 | shots.detected |
scenes.detected, scene.count |
||||||||
| प्रमाणGenerationWave | 100 | scenes.detected |
evidence.generated |
नोट्स
keyframes.deduplicated OCR थम्बनेल पर चलता हैघंटे के लिए कुल 2- चलचित्र: ~10-15 मिनट (vs. optimization के बिना घंटों
संकेतों को संगतता के लिए स्थिरांक के रूप में परिभाषित किया जाता है
public static class VideoSignals
{
// NormalizeWave signals
public const string VideoDuration = "video.duration";
public const string VideoFps = "video.fps";
public const string VideoNormalized = "video.normalized";
// Shot detection signals
public const string ShotsDetected = "shots.detected";
public const string ShotsCount = "shots.count";
// Keyframe signals
public const string IframesDetected = "keyframes.iframes_detected";
public const string KeyframesSelected = "keyframes.selected";
public const string KeyframesDeduplicated = "keyframes.deduplicated";
public const string KeyframesExtracted = "keyframes.extracted";
// CLIP embedding signals
public const string ClipEmbeddingsReady = "clip.embeddings_ready";
// Scene clustering signals
public const string ScenesDetected = "scenes.detected";
public const string SceneCount = "scene.count";
// Transcription signals
public const string TranscriptionComplete = "transcription.complete";
}
वीडियोसममैजर एक का उपयोग करता है क्षमता-आधारित वास्तुकलाआरंभ पर एक बार GPU पता लगाएँ
मॉडलों में परिभाषित किए जाते हैं models.yamlकोड में कोई जादू स्ट्रिंग नहीं
# models.yaml (excerpt)
models:
clip-vit-b32:
name: "CLIP ViT-B/32"
download_url: "https://huggingface.co/openai/clip-vit-base-patch32/resolve/main/onnx/visual_model.onnx"
preferred_providers: [CUDAExecutionProvider, DmlExecutionProvider, CPUExecutionProvider]
components:
ClipEmbeddingWave:
models: [clip-vit-b32]
fallback_chain: [ImageAnalysisWave]
// Type-safe constants (no raw strings)
await coordinator.EnsureModelAsync(ModelIds.ClipVitB32);
await coordinator.ActivateWaveAsync(ComponentIds.TranscriptionWave);
// Route with fallback
var route = await coordinator.RouteWorkAsync(new[]
{
ComponentIds.ClipEmbeddingWave, // Primary (GPU)
ComponentIds.ImageAnalysisWave // Fallback (CPU)
});
दर सीमान, समय अनुमान, और अनुकूली बैकप्रिश्रण UI प्रतिक्रियात्मक बनाए रखते हैं जबकि अधिकतम स्ट्राइपट
// Time estimation from actual data
var estimator = CapabilityAtoms.CreateTimeEstimator();
using (estimator.Time("clip_embedding")) { await ProcessAsync(); }
var eta = estimator.GetEstimate("clip_embedding", remaining: 50);
// eta.Estimated, eta.Optimistic, eta.Pessimistic, eta.Confidence
पूर्ण क्षमता प्रणाली डक: देखें
Mostlylucid.Summarizer.Core/Capabilities/GPU पता लगाने के लिए, सिग्नल pub/subM SK2 बैकप्रेस कंट्रोलर्सMSC3 और मेश टोपोलॉजी डिजाइन
expensive CLIP embeddings चलाने से पहले, VideoSummarizer का उपयोग कर दृश्य रूप में समान फ्रेमों को फिल्टर करता है अंतर हैश (dHash).
public class KeyframeDeduplicationService
{
// dHash parameters: 9x8 grayscale = 64 bits
private const int HashWidth = 9;
private const int HashHeight = 8;
private const int DefaultHammingThreshold = 10;
public async Task<ulong> ComputeDHashAsync(string imagePath, CancellationToken ct)
{
using var image = Image.Load<Rgba32>(imagePath);
// Resize to 9x8 (one extra column for gradient comparison)
image.Mutate(x => x
.Resize(HashWidth, HashHeight)
.Grayscale());
ulong hash = 0;
int bit = 0;
// Compare adjacent pixels horizontally
for (int y = 0; y < HashHeight; y++)
{
for (int x = 0; x < HashWidth - 1; x++)
{
var left = image[x, y].R;
var right = image[x + 1, y].R;
// Set bit if left pixel is brighter than right
if (left > right)
{
hash |= (1UL << bit);
}
bit++;
}
}
return hash;
}
public static int HammingDistance(ulong a, ulong b) =>
BitOperations.PopCount(a ^ b);
}
उदाहरण आउटपुट
Input: 50 keyframe candidates (from codec I-frames)
Deduplication (Hamming threshold 10):
Frame 0: hash=0x8f3a2c1d → KEEP (first frame)
Frame 1: hash=0x8f3a2c1e → SKIP (distance=1 from frame 0)
Frame 2: hash=0x8f3a2c1f → SKIP (distance=2 from frame 0)
Frame 3: hash=0xc7e1b4a2 → KEEP (distance=28 from frame 0)
...
Result: 50 → 30 frames (40% reduction)
Processing saved: ~8 seconds of CLIP inference
यह क्यों महत्वपूर्ण है
एक समय में एक छवि को प्रसंस्करण करने के बजाय
public class BatchClipEmbeddingService
{
private const int ClipImageSize = 224;
private const int DefaultBatchSize = 8; // 8 images per GPU pass
public async Task<Dictionary<int, float[]>> GenerateBatchEmbeddingsAsync(
Dictionary<int, string> framePaths,
int batchSize = DefaultBatchSize,
CancellationToken ct = default)
{
var session = await GetOrLoadClipModelAsync(ct);
var results = new Dictionary<int, float[]>();
// Pre-index batch for O(1) lookup (not batch.IndexOf!)
var batches = framePaths
.Select((kvp, idx) => (idx, kvp.Key, kvp.Value))
.Chunk(batchSize);
foreach (var batch in batches)
{
// Create batch tensor [batchSize, 3, 224, 224]
var tensor = new DenseTensor<float>(new[] { batch.Length, 3, ClipImageSize, ClipImageSize });
// Preprocess images in parallel (simplified; production uses vectorised span copy)
Parallel.ForEach(batch, item =>
{
var (batchIdx, frameIndex, path) = item;
var localIdx = batchIdx % batchSize;
PreprocessImageToTensor(path, tensor, localIdx); // ImageSharp pixel buffers
});
// Single GPU pass for entire batch
var inputs = new List<NamedOnnxValue>
{
NamedOnnxValue.CreateFromTensor("input", tensor)
};
using var outputResults = session.Run(inputs);
// Extract embeddings from batch output...
}
return results;
}
}
निष्पादन तुलना
Input: 30 keyframes (after deduplication)
Serial processing (1 frame at a time):
30 × 200ms = 6,000ms (6.0 seconds)
Batch processing (8 frames per pass):
4 batches × 350ms = 1,400ms (1.4 seconds)
Speedup: 4.3x
बैच प्रसंस्करण क्यों काम करता है
[8, 3, 224, 224] एकल छवि के रूप में एक ही GPU स्मृति का उपयोग करता है (अधिकतम)VideoSummarizer doesn't reinvent ImageSum marizer or AudioSommarizerit शंखियाँ them.
कुंजी फ्रेम निष्कर्षण में विभाजित किया जाता है 7 ग्रेनोलर तरंगों को।
// IFrameDetectionWave → KeyframeSelectionWave → ThumbnailExtractionWave
// → KeyframeDeduplicationWave → KeyframeFullResExtractionWave → ClipEmbeddingWave
// ClipEmbeddingWave coordinates with ImageSummarizer
public class ClipEmbeddingWave : IVideoWave, ISignalAwareVideoWave
{
public IReadOnlyList<string> RequiredSignals => [VideoSignals.KeyframesExtracted];
public IReadOnlyList<string> EmittedSignals => [VideoSignals.ClipEmbeddingsReady];
public async Task ProcessAsync(VideoContext context, CancellationToken ct)
{
var keyframes = context.GetCached<Dictionary<int, string>>("keyframes.paths");
// Batch CLIP embedding (3-5x faster than serial)
var embeddings = await _batchClipService.GenerateBatchEmbeddingsAsync(
keyframes, batchSize: 8, ct);
foreach (var (frameIndex, embedding) in embeddings)
context.KeyframeEmbeddings[frameIndex] = embedding;
}
}
// ImageAnalysisWave runs ImageSummarizer on deduplicated frames
public class ImageAnalysisWave : IVideoWave, ISignalAwareVideoWave
{
public IReadOnlyList<string> RequiredSignals => [VideoSignals.KeyframesDeduplicated];
public async Task ProcessAsync(VideoContext context, CancellationToken ct)
{
var keyframePaths = context.GetCached<List<string>>("keyframes.deduplicated_paths");
foreach (var path in keyframePaths)
{
// Run ImageSummarizer for OCR, vision, captions
var result = await _imageOrchestrator.AnalyzeAsync(path, ct);
context.SetCached($"image_analysis.{Path.GetFileName(path)}", result);
}
}
}
ऑडियो निष्कर्षण और ट्रांसक्रिप्शन अब अलग संकेत हैं
// AudioExtractionWave runs first (extracts audio track from video)
public class AudioExtractionWave : IVideoWave, ISignalAwareVideoWave
{
public IReadOnlyList<string> RequiredSignals => [VideoSignals.VideoNormalized];
public IReadOnlyList<string> EmittedSignals => ["audio.extracted", "audio.path"];
public async Task ProcessAsync(VideoContext context, CancellationToken ct)
{
var audioPath = await _ffmpegService.ExtractAudioAsync(
context.VideoPath, context.WorkingDirectory, ct);
context.SetCached("audio.path", audioPath);
}
}
// TranscriptionWave depends on audio.extracted signal
public class TranscriptionWave : IVideoWave, ISignalAwareVideoWave
{
public IReadOnlyList<string> RequiredSignals => ["audio.extracted"];
public IReadOnlyList<string> EmittedSignals => [
VideoSignals.TranscriptionComplete,
"transcription.utterance_count"
];
public async Task ProcessAsync(VideoContext context, CancellationToken ct)
{
var audioPath = context.GetCached<string>("audio.path");
// Run AudioSummarizer pipeline (Whisper + diarization)
var audioProfile = await _audioOrchestrator.AnalyzeAsync(audioPath, ct);
// Extract utterances with speaker info
var turns = audioProfile.GetValue<List<SpeakerTurn>>("speaker.turns");
foreach (var turn in turns ?? [])
{
context.Utterances.Add(new Utterance
{
Id = Guid.NewGuid(),
Text = turn.Text,
StartTime = turn.StartSeconds,
EndTime = turn.EndSeconds,
SpeakerId = turn.SpeakerId,
Confidence = turn.Confidence
});
}
// Run NER on full transcript for entity extraction
var transcript = audioProfile.GetValue<string>("transcription.full_text");
if (!string.IsNullOrEmpty(transcript))
{
var entities = await _nerService.ExtractEntitiesAsync(transcript, ct);
context.SetCached("transcript_entities", entities);
// Emit entity signals by type (PER, ORG, LOC, MISC)
foreach (var group in entities.GroupBy(e => e.Type))
{
context.AddSignal($"transcript.entities.{group.Key.ToLowerInvariant()}",
group.Select(e => e.Text).Distinct().ToList());
}
}
}
}
VideoSummarizer BERT-based NER के प्रयोग से नामित इकाइयों को ट्रांस्क्रिप्टों से निकालता है
public class OnnxNerService
{
// Model: dslim/bert-base-NER (ONNX exported)
// Entities: PER (Person), ORG (Organization), LOC (Location), MISC (Miscellaneous)
public async Task<List<EntitySpan>> ExtractEntitiesAsync(string text, CancellationToken ct)
{
var entities = new List<EntitySpan>();
// Chunk long text (BERT max 512 tokens)
foreach (var chunk in ChunkText(text, maxTokens: 400, overlap: 50))
{
// Tokenize with WordPiece
var tokens = _tokenizer.Tokenize(chunk);
// Run ONNX inference
var inputs = PrepareInputs(tokens);
using var results = _session.Run(inputs);
// Decode BIO tags
var predictions = DecodePredictions(results);
var chunkEntities = ExtractEntitySpans(tokens, predictions);
entities.AddRange(chunkEntities);
}
// Deduplicate entities
return entities
.GroupBy(e => (e.Text.ToLowerInvariant(), e.Type))
.Select(g => g.First())
.ToList();
}
}
उदाहरण आउटपुट
Transcript: "Today we're speaking with John Smith from Microsoft about
their new AI lab in Seattle. The project, codenamed Phoenix, builds
on research from Stanford University."
Entities extracted:
PER: John Smith
ORG: Microsoft, Stanford University
LOC: Seattle
MISC: Phoenix
Signals emitted:
transcript.entities.per = ["John Smith"]
transcript.entities.org = ["Microsoft", "Stanford University"]
transcript.entities.loc = ["Seattle"]
transcript.entities.misc = ["Phoenix"]
वीडियो के लिए क्यों एनईआर महत्वपूर्ण है
दृश्य पता लगाने के लिए केवल सीलिप एम्बेडिंग का उपयोग करने से यह काम नहीं करता।
वीडियोसममैजर एक का उपयोग करता है बहु-- सिग्नल दृष्टिकोण जो मजबूत दृश्य सीमा पता लगाने के लिए 4 भारित संकेतों को जोड़ता है
public class SceneClusteringWave : IVideoWave, ISignalAwareVideoWave
{
// Signal weights for boundary scoring
private const double EmbeddingWeight = 0.4; // CLIP embedding dissimilarity
private const double TranscriptWeight = 0.3; // Semantic shift in transcript
private const double CutTypeWeight = 0.2; // Fade/dissolve detection
private const double TemporalWeight = 0.1; // Time since last scene
// Temporal constraints
private const double MinSceneDuration = 15.0; // Don't split scenes < 15s
private const double MaxSceneDuration = 300.0; // Force split at 5 minutes
private const double TargetSceneDuration = 90.0; // Prefer ~90s scenes
public IReadOnlyList<string> RequiredSignals => [VideoSignals.ShotsDetected];
public IReadOnlyList<string> OptionalSignals => [
VideoSignals.ClipEmbeddingsReady,
VideoSignals.TranscriptionComplete,
VideoSignals.KeyframesDeduplicated
];
public IReadOnlyList<string> EmittedSignals => [
VideoSignals.ScenesDetected,
"scene.count",
"scene.avg_duration",
"scene.clustering_method"
];
private List<(int shotIndex, double score)> ComputeBoundaryScores(VideoContext context)
{
var shots = context.Shots.OrderBy(s => s.StartTime).ToList();
var scores = new List<(int, double)>();
// Build embedding map with nearest-neighbor interpolation
var shotEmbeddings = PropagateEmbeddingsToNearbyShots(context, shots);
// Build transcript windows for semantic shift detection
var transcriptWindows = BuildTranscriptWindows(context, shots, windowSeconds: 10);
for (int i = 0; i < shots.Count - 1; i++)
{
double score = 0;
var currentShot = shots[i];
var nextShot = shots[i + 1];
// 1. Embedding dissimilarity (40%)
if (shotEmbeddings.TryGetValue(i, out var currentEmbed) &&
shotEmbeddings.TryGetValue(i + 1, out var nextEmbed))
{
var similarity = CosineSimilarity(currentEmbed, nextEmbed);
score += (1.0 - similarity) * EmbeddingWeight;
}
// 2. Transcript semantic shift (30%)
if (transcriptWindows.TryGetValue(i, out var currentWords) &&
transcriptWindows.TryGetValue(i + 1, out var nextWords))
{
var overlap = currentWords.Intersect(nextWords).Count();
var union = currentWords.Union(nextWords).Count();
var jaccard = union > 0 ? (double)overlap / union : 0;
score += (1.0 - jaccard) * TranscriptWeight;
}
// 3. Cut type signal (20%) - fades/dissolves suggest scene boundaries
if (currentShot.CutType is "fade" or "dissolve")
{
score += CutTypeWeight;
}
// 4. Temporal pressure (10%) - encourage splits near target duration
var timeSinceLastScene = currentShot.EndTime - GetLastSceneBoundary();
if (timeSinceLastScene > TargetSceneDuration)
{
var pressure = Math.Min(1.0, (timeSinceLastScene - TargetSceneDuration) / 60);
score += pressure * TemporalWeight;
}
scores.Add((i, score));
}
return scores;
}
}
निकटतम- पड़ोसी सम्मिलन प्रसारकेवल शॉटों में ~2% सीधी CLIP एम्बेडिंग होती है
ट्रांस्क्रिप्ट सेमेटिक विंडोज़प्रत्येक शॉट के चारों ओर से एक सेकंड शब्द विंडो बनाता है और Jaccard दूरी के माध्यम से अर्थिक परिवर्तनों को पता लगाता है एक सस्ता अर्थिक ड्रिफ्ट प्रॉक्सी।
काटें प्रकार सचेतनता: फेडM SK1to-black and dissolve transitions strongly indicate scene boundaries
अनुकूलित बन्धनएक नियत थ्रेसहोल्ड के बजाय : शीर्ष से किनारे चुनता है
अस्थायी अवरोधन्यूनतम को लागू करता है 15\s दृश्यों और सीमाओं पर बल देता है
उदाहरण:
Input: 1881 shots from a 2-hour movie
39 keyframes with CLIP embeddings
2302 utterances from transcript
Boundary scoring per shot:
Shot 45-46: embedding=0.15, transcript=0.32, cut=0.0, temporal=0.0 → score=0.156
Shot 46-47: embedding=0.08, transcript=0.12, cut=0.0, temporal=0.0 → score=0.068
Shot 47-48: embedding=0.35, transcript=0.41, cut=0.2, temporal=0.05 → score=0.388 ← BOUNDARY
...
Adaptive threshold (top 25%): 0.25
Natural boundaries found: 45
Output: 47 scenes (avg 2.6 minutes per scene)
- Min scene: 15.2s
- Max scene: 298.4s
- Total coverage: 100%
Signals:
scenes.detected = true
scene.count = 47
scene.avg_duration = 156.3
scene.clustering_method = "multi_signal_weighted"
वीडियोसममेरिजर इमेजसममेराइजर और ऑडियोसममेरिसर से संकेत संविदा बढ़ाता है
public record VideoSignal
{
public required string Key { get; init; } // "scene.count", "transcript.entities.per"
public object? Value { get; init; }
public double Confidence { get; init; } = 1.0;
public required string Source { get; init; } // "SceneClusteringWave"
// Video-specific: time range
public double? StartTime { get; init; }
public double? EndTime { get; init; }
public DateTime Timestamp { get; init; }
public Dictionary<string, object>? Metadata { get; init; }
public List<string>? Tags { get; init; } // ["visual", "scene"]
}
public static class VideoSignalTags
{
public const string Visual = "visual";
public const string Audio = "audio";
public const string Speech = "speech";
public const string Ocr = "ocr";
public const string Motion = "motion";
public const string Scene = "scene";
public const string Shot = "shot";
public const string Metadata = "metadata";
}
जारी कुंजी संकेत
| संकेत | स्रोत | विवरण | ||
|---|---|---|---|---|
video.duration सामान्यizeWave |
सेकेंडों में कुल अवधि | |||
video.resolution NormalizeWave |
WidthM SK2Height | |||
video.fps सामान्यizeWave |
फ्रेम दर | |||
shots.count ShotDetectionWave |
||||
keyframes.count कुंजी फ्रेम एक्सट्रैक्शन वेव |
||||
keyframes.duplicates_skipped KeyframeExtractionWave |
dHash द्वारा फिल्टर किए गए फ्रेम | |||
scene.count SceneClusteringWave |
||||
transcript.entities.per ट्रांस्क्रिप्शन वेव |
एनईआर से व्यक्तियों के नाम | |||
transcript.entities.org ट्रांस्क्रिप्शन वेव |
संगठन नाम | |||
transcript.word_count ट्रांस्क्रिप्शन वेव |
ट्रांसक्रिप्सन में कुल शब्द |
VideoPipeline वीडियो संकेतों में बदलता है ContentChunk RAG सूचकीकरण के लिए:
public class VideoPipeline : PipelineBase
{
public override string PipelineId => "video";
public override IReadOnlySet<string> SupportedExtensions => new HashSet<string>
{
".mp4", ".mkv", ".avi", ".mov", ".wmv", ".webm", ".flv", ".m4v", ".mpeg", ".mpg"
};
private List<ContentChunk> BuildContentChunks(VideoContext context, string filePath)
{
var chunks = new List<ContentChunk>();
// 1. Scene-based chunks (best for video retrieval)
foreach (var scene in context.Scenes)
{
var sceneText = BuildSceneText(context, scene);
var embedding = context.GetCached<float[]>($"scene_centroid.{scene.Id}");
chunks.Add(new ContentChunk
{
Text = sceneText,
ContentType = ContentType.Summary,
Embedding = embedding, // Proper vector column, not metadata
Metadata = new Dictionary<string, object?>
{
["source"] = "video_scene",
["scene_id"] = scene.Id,
["key_terms"] = scene.KeyTerms,
["speakers"] = scene.SpeakerIds,
["start_time"] = scene.StartTime,
["end_time"] = scene.EndTime
}
});
}
// 2. Transcript chunks (1-minute windows)
var transcriptChunks = BuildTranscriptChunks(context, filePath);
chunks.AddRange(transcriptChunks);
// 3. Text track chunks (on-screen text/subtitles)
foreach (var textTrack in context.TextTracks)
{
chunks.Add(new ContentChunk
{
Text = $"On-screen text: {textTrack.Text}",
ContentType = ContentType.ImageOcr,
Metadata = new Dictionary<string, object?>
{
["source"] = "video_ocr",
["text_type"] = textTrack.TextType.ToString(),
["start_time"] = textTrack.StartTime
}
});
}
return chunks;
}
private string BuildSceneText(VideoContext context, SceneSegment scene)
{
var parts = new List<string>();
if (!string.IsNullOrEmpty(scene.Label))
parts.Add($"Scene: {scene.Label}");
parts.Add($"[{FormatTime(scene.StartTime)} - {FormatTime(scene.EndTime)}]");
if (scene.KeyTerms.Count > 0)
parts.Add($"Topics: {string.Join(", ", scene.KeyTerms)}");
// Add utterances in this scene
var sceneUtterances = context.Utterances
.Where(u => u.StartTime >= scene.StartTime && u.EndTime <= scene.EndTime)
.OrderBy(u => u.StartTime);
if (sceneUtterances.Any())
parts.Add($"Speech: {string.Join(" ", sceneUtterances.Select(u => u.Text))}");
return string.Join("\n", parts);
}
}
फिल्म के लिए उदाहरण आउटपुट:
{
"chunks": [
{
"text": "Scene: Opening montage\n[0:00 - 2:34]\nTopics: city, night, traffic\nSpeech: The year is 2049. The world has changed.",
"contentType": "Summary",
"metadata": {
"source": "video_scene",
"scene_id": "abc123",
"key_terms": ["city", "night", "traffic"],
"start_time": 0.0,
"end_time": 154.0
}
},
{
"text": "The detective arrived at the crime scene. Forensics had already processed the area.",
"contentType": "Transcript",
"metadata": {
"source": "video_transcript",
"time_window": "2:34 - 3:34",
"utterance_count": 4
}
},
{
"text": "On-screen text: LOS ANGELES 2049",
"contentType": "ImageOcr",
"metadata": {
"source": "video_ocr",
"text_type": "Title"
}
}
]
}
चरण |-------|------|-------| FFprobe मेटाडेटा Shot detection कुंजी फ्रेम एक्सटेक्शन डीहश डुप्लिकेट बैच सीलिप अंतःकरण | ~60 s पाठ के साथ कुंजीफ्रेम | ImageSummarizer OCR ऑडियो एक्सटेक्शन | हिस्सों की प्रतिलिपि स्पीकर डायराइजेशन ट्रांस्क्रिप्ट पर | दृश्य क्लस्टरिंग प्रमाण सृजन | कुल | मिनट | |
अनुकूलन |--------------|---------| | dHash डेड्यूप्लिकेशन | | | \ ~40% | फ्रेमों को फिल्टर किया गया |= | बैच क्लिप | पाइपलाइन कम्पोशन | कुल बचत | मिनट |
अवयव |-----------|--------| CLIP ViT | फुसफुसा आधार 0 ECAPA 1 TDNN 2 ~100 MB 4 0 BERT 1 NER 2 3 MB 4 | शिखर | एमएसके0जीबी |
वीडियोसममैजर एक के रूप में पंजीकृत करता है IPipeline स्वचालित रूटिंग के लिए
// In Program.cs
builder.Services.AddDocSummarizer(builder.Configuration.GetSection("DocSummarizer"));
builder.Services.AddDocSummarizerImages(builder.Configuration.GetSection("Images"));
builder.Services.AddVideoSummarizer(); // NEW
builder.Services.AddPipelineRegistry(); // Must be last
// Auto-routing by extension
var registry = services.GetRequiredService<IPipelineRegistry>();
var pipeline = registry.FindForFile("movie.mp4"); // Returns VideoPipeline
var result = await pipeline.ProcessAsync("movie.mp4");
समर्थित एक्सटेंशन
.mp4, .mkv, .avi, .mov, .wmv, .webm, .flv, .m4v, .mpeg, .mpgवीडियोसममाराइजर यह प्रदर्शित करता है कि पाइपलाइन संरचना स्केल
परिणाम यह होता है कि एक घंटे की फिल्म दृश्यों के साथ एक संरचनात्मक संकेत लाईडर बन जाती है।
वीडियो के लिए कम RAG पैटर्न
Ingestion: Video → 16 waves → Signals + Evidence (scenes, transcripts, entities)
Storage: Signals (indexed) + Embeddings (CLIP, voice) + Evidence (chunks)
Query: Filter (SQL) → Search (BM25 + vector) → Synthesize (LLM, ~5 results)
क्षमता प्रणाली
Startup: Detect GPU → Load ModelManifest (YAML) → Initialize SignalSink
Activation: Component requests model → Lazy download → Signal "ModelAvailable"
Routing: Route to best provider → Fallback chain → Backpressure control
Atoms: Rate limiting + Time estimation + Pipeline balancing
यह है अवरोधित अस्पष्टता पैमाने पर
एलएलएम पूर्व कम्प्यूटरी साक्ष्य नहीं कच्चा वीडियो पर कार्य करता है
कोर पैटर्न
कम RAG कार्यान्वयन
| भाग | पैटर्न | |
|---|---|---|
| 1 | अवरोधित अस्पष्टता एकल घटक | |
| 2 | प्रतिबंधित अस्पष्ट मोएम बहुआयामी अवयव | |
| 3 | संदर्भ खींचना समय | |
| 4 | छवि इंटेलिजेंस तरंग संरचना | |
| 4.1 | Three-Tier OCR पाइपलाइन OCR | |
| 4.2 | ऑडियोसममारीजरName फ़ॉरेंसिक ऑडियो | |
| 4.3 | वीडियोसॉम्मारीजर (इस अनुच्छेद) | वीडियो ऑर्केस्ट्रेशन, बैच क्लिप |
अगला: बहुल-मोडल ग्राफ RAG lucidRAG सहित सभी चार सारांशकर्ताओं को एक एकीकृत ज्ञान ग्राफ में सम्मिलित करने के साथ क्रॉस
सभी भाग एक ही अपरिवर्ती का अनुसरण करते हैं संभाव्यात्मक घटक प्रस्तावित.
© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.