भाग 4: छवि इंटेलिजेंस ImageSummarizer तरंग वास्तुकला और व्यापक पैटर्नों को प्रस्तुत करता है ओसीआर उपतंत्रपाठ उद्धरण के तीन स्तरों , बौद्धिक रूटिंग | , | और फिल्म ट्रिप ऑप्टिमेशन जो एनीमेशनted GIFs के लिए |30× | टोकन कम करता है
क्यों एक अलग लेख ओसीआर पाइपलाइन ने विज़न एलएलएम फेलबैक के साथ टेसेरेक्ट से विकसित किया।
संबंधित लेख:
वास्तविक पर ओसीआर
पारंपरिक दृष्टिकोण
समस्या: यह या तो शैलीबद्ध पाठ को हराता है।
समाधानमध्य स्तर जोड़ें (Florence
सिस्टम तरंगों को प्राथमिकता क्रम में चलाता है
Wave Priority Order:
40: TextLikelinessWave → Heuristic text detection
50: OcrWave → Tesseract OCR (if text-likely)
51: MlOcrWave → Florence-2 ML OCR (if Tesseract low confidence)
55: Florence2Wave → Florence-2 captions (optional)
80: VisionLlmWave → Vision LLM (escalation)
| Priority | Speed | Cost | Best For | ||||
|---|---|---|---|---|---|---|---|
| 50 मुक्त | साफ पाठ, उच्च कंट्रास्टM SK5 मानक फ़ॉन्ट | स्टिलाइज्ड फ़ॉंटMSC7 कम क्वालिटी |
भेजे गए संकेत
ocr.text - निष्कर्षित पाठocr.confidence - टेसरेक्ट औसत विश्वास प्राप्तांक| Priority | Speed | Cost | Best For | ||||
|---|---|---|---|---|---|---|---|
| 51 Free | Stylized fonts, memes , decorative text |
भेजे गए संकेत
ocr.ml.text - एकलM SK1फ्रेम फ्लोरेंसocr.ml.multiframe_text बहुविध GIF पाठocr.ml.confidence - मॉडल विश्वास प्राप्ताङ्क| प्राथमिकता | गति | लागत | सबसे अच्छा | МSK4 | बन्धन | ||||
|---|---|---|---|---|---|---|---|---|---|
| 80 सभी , विशेष रूप से जटिल दृश्यों को मानना होगा. |
भेजे गए संकेत
ocr.vision.text - दृश्य एलएलएम ओसीआर पाठ निष्कर्षणocr.vision.confidence - एलएलएम विश्वासcaption.text वैकल्पिक वर्णनात्मक शीर्षक OCR से अलगतीन ओसीआर स्तरों में डुबोने से पहले निर्णायक एमएल मॉडल जो सिस्टम को शक्ति प्रदान करता है. सभी मॉडलों ने ONNX Runtime के माध्यम से स्थानीय रूप से चलाया
* गौण चेतावनी: GPU निष्पादन प्रदाताओं को असामान्य फ्लोटिंग ला सकता है-बिन्दु निर्धारितता नहीं ला सकते हैंMSC2 संकेत संविदा M SK3 विश्वास थ्रेसहोल्ड्स
नोट: आकार लगभग हैं और वैकल्पिक के आधार पर भिन्न होते हैं
| मॉडल | लगभग | . | आकार | उद्देश्य | МSK4 | गति | एमSK5 | मॉडल प्रकार | ||
|---|---|---|---|---|---|---|---|---|---|---|
| पूर्व दृश्य पाठ पता लगाना | ||||||||||
| क्रैफ्ट | ~150MB | वर्ण-क्षेत्र पाठ पता लगाना | ||||||||
| फ्लोरेंस OCR + क्याप्सनिंग | ||||||||||
| Real-ESRGAN à¤aà¥à¤°à¥à¤ | ||||||||||
| क्लिप Semantic embeddings |
कुल डिस्क स्थानचुने गए मॉडल वैकल्पिकों पर निर्भर करता है
प्रभावी और सटीक दृश्य पाठ डिटेक्टर प्राकृतिक दृश्यों में पाठ क्षेत्र पाता है
// EAST detects text bounding boxes with confidence scores
var result = await textDetector.RunEastDetectionAsync(imagePath);
// Output: List of BoundingBox with coordinates + confidence
// Example: [BoundingBox(x1:50, y1:100, x2:300, y2:150, confidence:0.92)]
यह कैसे काम करता है:
क्यों निर्णायक
< 0.5 → escalate)तकनीकी विवरण:
// EAST preprocessing (from implementation)
- Input size: 320×320 (must be multiple of 32)
- Format: BGR with mean subtraction [123.68, 116.78, 103.94]
- Output stride: 4 (downsampled 4×)
- Score threshold: 0.5
- NMS IoU threshold: 0.4
उदाहरण आउटपुट:
Input: meme.png (800×600)
EAST detection: 15 text regions found
Region 1: (50, 480, 750, 580) - confidence 0.87 [bottom subtitle area]
Region 2: (100, 50, 300, 90) - confidence 0.62 [top text]
Region 3: ...
Route decision: ANIMATED (subtitle pattern in bottom 30%)
अक्षर-स्तर पाठ पता लगाना - वक्रित में उत्कृष्टता करता है
// CRAFT finds character-level regions, then groups into words
var result = await textDetector.RunCraftDetectionAsync(imagePath);
// Better than EAST for: decorative fonts, curved text, logos
यह कैसे काम करता है:
जब क्रैट का उपयोग किया जाता है:
तकनीकी विवरण:
// CRAFT preprocessing
- Max dimension: 1280px (maintains aspect ratio)
- Format: RGB normalized with ImageNet stats
- Mean: [0.485, 0.456, 0.406]
- Std: [0.229, 0.224, 0.225]
- Output stride: 2 (downsampled 2×)
- Threshold: 0.4 for character regions
ईस्ट vs क्रैट तुलना:
विशेषता |---------|------|-------| खोज स्तर | शब्द | / | पंक्ति || | क्यारेक्टर स्पीड के लिए सर्वोत्तम | मानक पाठ | वक्र पाठ मॉडल साइज
ओसीआर से पहले कम-से-कम गुणवत्ता वाले छवियों को बढ़ाता है - 4× अस्पष्ट के लिए उपस्केलन
// Upscale low-quality image before running OCR
if (quality.Sharpness < 30) // Laplacian variance threshold
{
var upscaled = await esrganService.UpscaleAsync(imagePath, scale: 4);
// Now run OCR on the enhanced image
}
जब यह ' का उपयोग किया जाता है:
उदाहरण:
Input: 100×75 screenshot with tiny text
Laplacian variance: 18 (very blurry)
ESRGAN: Upscale to 400×300 (~500ms)
New Laplacian variance: 87 (sharp)
OCR: Tesseract confidence: 0.92 (vs 0.42 before upscaling)
Text: "Click here to continue" (vs garbled before)
तकनीकी विवरण:
// Real-ESRGAN processing
- Input: Any size (processed in 128×128 tiles if large)
- Output: 4× scaled (200×150 → 800×600)
- Model: x4plus variant (general photos)
- Processing: ~500ms for 800×600 image
- Memory: ~2GB peak (tiles reduce this)
टोकन अर्थशास्त्र:
Scenario: Screenshot with tiny text
Option 1: Send low-res to Vision LLM
Image: 100×75 = ~20 tokens
LLM can't read tiny text → fails
Cost: $0.0002 (wasted)
Option 2: Upscale with ESRGAN, use Tesseract
ESRGAN: Free (local), 500ms
Tesseract: Free (local), 50ms
Success: 92% confidence
Cost: $0
Result: ESRGAN + local OCR beats Vision LLM for low-res images
वैचारिक छवि खोज के लिए बहुमोडल एम्बेडिंग साझा भेक्टर स्थान में छवि और पाठ परियोजना करता है
// Generate embedding for semantic search
var embedding = await clipService.GenerateEmbeddingAsync(imagePath);
// Returns: float[512] vector
// Later: semantic search across thousands of images
var similarImages = await vectorDb.SearchAsync(queryEmbedding, topK: 10);
यह कैसे काम करता है:
केस का उपयोग करें:
तकनीकी विवरण:
// CLIP visual encoder
- Model: ViT-B/32 (Vision Transformer)
- Input: 224×224 RGB (center crop + resize)
- Output: 512-dimensional embedding
- Normalized: L2 norm = 1.0
- Speed: ~100ms per image
उदाहरण:
Input images:
cat_on_couch.jpg → [0.23, -0.51, 0.88, ...]
dog_on_couch.jpg → [0.19, -0.48, 0.91, ...]
car_photo.jpg → [-0.67, 0.33, -0.12, ...]
Query: "animals on furniture"
Text embedding → [0.21, -0.50, 0.89, ...]
Cosine similarity:
cat_on_couch: 0.94 (very similar!)
dog_on_couch: 0.91 (similar)
car_photo: 0.12 (not similar)
Result: Returns cat and dog images
फ्लोरेंस के बारे में अधिक जानकारी के लिए टियर 2 को देखें
सभी मॉडलों को पहली बार उपयोग पर स्वचालित रूप से डाउनलोड किया जाता है
$ imagesummarizer image.png --pipeline auto
[First run]
Downloading EAST scene text detector (~100MB)...
Progress: ████████████████████ 100% (102.4 MB)
Downloading Florence-2 base model (~250MB)...
Progress: ████████████████████ 100% (248.7 MB)
Downloading CLIP ViT-B/32 visual (~350MB)...
Progress: ████████████████████ 100% (347.2 MB)
Models saved to: ~/.mostlylucid/models/
Total disk space: 1.16 GB
[Subsequent runs]
All models cached, analysis starts immediately
अनुग्रही अवनति:
// If ONNX model download fails, system falls back gracefully
EAST unavailable → Try CRAFT → Fall back to Tesseract PSM
Real-ESRGAN unavailable → Skip upscaling, use original image
CLIP unavailable → Skip embeddings, OCR still works
Florence-2 unavailable → Use Tesseract → Vision LLM escalation
हर ONNX मॉडल विफलता के साथ वापसी पथ के साथ लॉग किया जाता है।
मूल्य निर्धारण नोटवास्तविक API लागत प्रदाता और मॉडल के अनुसार भिन्न होती है।
बिना ऑनिक्स मॉडल बेसलाइन
Every image → Send to Vision LLM
Cost: ~$0.005/image (example pricing)
Time: ~2s network + inference
100 images = ~$0.50, ~200s
ऑनिक्स मॉडलों के साथ (local
85 images → EAST + Florence-2 (local)
Cost: $0
Time: ~200ms
10 images → EAST + Tesseract (local)
Cost: $0
Time: ~50ms
5 images → EAST + Vision LLM (escalation)
Cost: ~$0.025 (5 × $0.005)
Time: ~2s each
100 images = ~$0.025, ~30s total
बचतलागत कम करना निर्णायक रूटिंग.
ONNX मॉडल तंत्र को "probabilistic सभी तरह से नीचे " करने के लिए रूपांतरित करता है | "deterministic आधार |+ probabilistic escalation केवल जब आवश्यक हो |
बुनियादी लाइन. Fast, deterministicM SK2 साफ पाठ के लिए बहुत अच्छा काम करता है
public class OcrWave : IAnalysisWave
{
public string Name => "OcrWave";
public int Priority => 60; // After color/identity
public async Task<IEnumerable<Signal>> AnalyzeAsync(
string imagePath,
AnalysisContext context,
CancellationToken ct)
{
var signals = new List<Signal>();
// Get preprocessed image from cache
var image = context.GetCached<Image<Rgba32>>("image");
// Run Tesseract OCR
using var engine = new TesseractEngine(@"./tessdata", "eng", EngineMode.Default);
using var page = engine.Process(image);
var text = page.GetText();
var confidence = page.GetMeanConfidence();
signals.Add(new Signal
{
Key = "ocr.text", // Tesseract OCR result
Value = text,
Confidence = confidence,
Source = Name,
Tags = new List<string> { "ocr", "text" },
Metadata = new Dictionary<string, object>
{
["engine"] = "tesseract",
["mean_confidence"] = confidence,
["word_count"] = text.Split(' ').Length
}
});
signals.Add(new Signal
{
Key = "ocr.confidence",
Value = confidence,
Confidence = 1.0,
Source = Name
});
return signals;
}
}
कुंजी संकेत:
ocr.full_text - निकाला गया पाठocr.early_exit - स्तरीय छोड़ने के लिए संकेतमाइक्रोसॉफ्ट's फ्लोरेंस-2 एक दृश्य है।
public class MlOcrWave : IAnalysisWave
{
private readonly Florence2OnnxModel _model;
public string Name => "MlOcrWave";
public int Priority => 51; // Runs AFTER Tesseract (priority 50)
public async Task<IEnumerable<Signal>> AnalyzeAsync(
string imagePath,
AnalysisContext context,
CancellationToken ct)
{
var signals = new List<Signal>();
// Check if Tesseract already succeeded with high confidence
var tesseractConfidence = context.GetValue<double>("ocr.confidence");
if (tesseractConfidence >= 0.95)
{
signals.Add(new Signal
{
Key = "ocr.ml.skipped", // Consistent namespace: ocr.ml.*
Value = true,
Confidence = 1.0,
Source = Name,
Metadata = new Dictionary<string, object>
{
["reason"] = "tesseract_high_confidence",
["tesseract_confidence"] = tesseractConfidence
}
});
return signals;
}
// Run Florence-2 OCR
var result = await _model.ExtractTextAsync(imagePath, ct);
signals.Add(new Signal
{
Key = "ocr.ml.text", // Florence-2 ML OCR text
Value = result.Text,
Confidence = result.Confidence,
Source = Name,
Tags = new List<string> { "ocr", "text", "ml" },
Metadata = new Dictionary<string, object>
{
["model"] = "florence2-base",
["inference_time_ms"] = result.InferenceTime,
["token_count"] = result.TokenCount
}
});
// For animated GIFs, extract all unique frames
if (context.GetValue<int>("identity.frame_count") > 1)
{
var frameResults = await ExtractMultiFrameTextAsync(
imagePath,
maxFrames: 10,
ct);
signals.Add(new Signal
{
Key = "ocr.ml.multiframe_text",
Value = frameResults.CombinedText,
Confidence = frameResults.AverageConfidence,
Source = Name,
Metadata = new Dictionary<string, object>
{
["frames_processed"] = frameResults.FrameCount,
["unique_text_segments"] = frameResults.UniqueSegments,
["deduplication_method"] = "levenshtein_85"
}
});
}
return signals;
}
}
एनिमित GIFs के लिए parallel में नमूना फ्रेमों तक प्रक्रियाएँ
private async Task<MultiFrameResult> ExtractMultiFrameTextAsync(
string imagePath,
int maxFrames,
CancellationToken ct)
{
// Load GIF and extract frames
using var image = await Image.LoadAsync<Rgba32>(imagePath, ct);
var frames = new List<Image<Rgba32>>();
int frameCount = image.Frames.Count;
int step = Math.Max(1, frameCount / maxFrames);
for (int i = 0; i < frameCount; i += step)
{
frames.Add(image.Frames.CloneFrame(i));
}
// Process all frames in parallel (bounded concurrency to avoid thrashing)
var semaphore = new SemaphoreSlim(4); // Max 4 concurrent inferences
var tasks = frames.Select(async frame =>
{
await semaphore.WaitAsync(ct);
try
{
var result = await _model.ExtractTextAsync(frame, ct);
return result;
}
finally
{
semaphore.Release();
}
});
var results = await Task.WhenAll(tasks);
semaphore.Dispose();
// Deduplicate using Levenshtein distance
var uniqueTexts = DeduplicateByLevenshtein(
results.Select(r => r.Text).ToList(),
threshold: 0.85);
return new MultiFrameResult
{
CombinedText = string.Join("\n", uniqueTexts),
FrameCount = frames.Count,
UniqueSegments = uniqueTexts.Count,
AverageConfidence = results.Average(r => r.Confidence)
};
}
private List<string> DeduplicateByLevenshtein(
List<string> texts,
double threshold)
{
var unique = new List<string>();
foreach (var text in texts)
{
bool isDuplicate = false;
foreach (var existing in unique)
{
var distance = LevenshteinDistance(text, existing);
var maxLen = Math.Max(text.Length, existing.Length);
var similarity = 1.0 - (distance / (double)maxLen);
if (similarity >= threshold)
{
isDuplicate = true;
break;
}
}
if (!isDuplicate)
{
unique.Add(text);
}
}
return unique;
}
उदाहरणफ्रेम GIF
Frame 1-45: "I'm not even mad."
Frame 46-93: "That's amazing."
ओपन सीवी पाठ पता लगाना (~5-20ms) निर्धारित करता है कि कौन सा रास्ता ले जाना है
public class TextDetectionService
{
public TextDetectionResult DetectText(Image<Rgba32> image)
{
// Use OpenCV EAST text detector
var (regions, confidence) = RunEastDetector(image);
return new TextDetectionResult
{
HasText = regions.Count > 0,
RegionCount = regions.Count,
Confidence = confidence,
Route = SelectRoute(regions, confidence, image)
};
}
private ProcessingRoute SelectRoute(
List<TextRegion> regions,
double confidence,
Image<Rgba32> image)
{
// No text detected
if (regions.Count == 0)
return ProcessingRoute.NoOcr;
// Animated GIF with subtitle pattern
if (image.Frames.Count > 1 && HasSubtitlePattern(regions))
return ProcessingRoute.AnimatedFilmstrip;
// High confidence, standard text
if (confidence >= 0.8 && HasStandardTextCharacteristics(regions))
return ProcessingRoute.Fast; // Florence-2 only
// Moderate confidence
if (confidence >= 0.5)
return ProcessingRoute.Balanced; // Florence-2 + Tesseract voting
// Low confidence, complex image
return ProcessingRoute.Quality; // Full pipeline + Vision LLM
}
private bool HasSubtitlePattern(List<TextRegion> regions)
{
// Subtitles are typically in bottom 30% of frame
var bottomRegions = regions.Where(r =>
r.BoundingBox.Y > r.ImageHeight * 0.7);
return bottomRegions.Count() >= regions.Count * 0.5;
}
}
| रूट | ट्रिगर करता है जब |
|---|---|
| तेजी से उच्च विश्वसनीयता (>0.8), मानक पाठ | |
| संतुलित मामूली विश्वास | |
| गुणवत्ता | कम विश्वसनीयता |
| एनिमेट उपशीर्षक पैटर्न के साथ GIF |
GIF उपशीर्षक के लिए त्रुटिपूर्ण अनुकूलन सिर्फ पाठ क्षेत्रपूर्ण फ्रेम नहीं
उपशीर्षक सहित फ्रेम GIF के लिए पारंपरिक दृष्टिकोण
Option 1: Process every frame
93 frames × 300×185 × ~150 tokens/frame = 13,950 tokens
Cost: ~$0.14 @ $0.01/1K tokens
Time: ~27 seconds
Option 2: Sample 10 frames
10 frames × 300×185 × ~150 tokens/frame = 1,500 tokens
Cost: ~$0.015
Time: ~3 seconds
Problem: Might miss subtitle changes
सिर्फ पाठ बंडिंग बॉक्स निकालें
2 text regions × 250×50 × ~25 tokens/region = 50 tokens
Cost: ~$0.0005
Time: ~2 seconds
Token reduction: 30×
public class FilmstripService
{
public async Task<TextOnlyStrip> CreateTextOnlyStripAsync(
string imagePath,
CancellationToken ct)
{
using var gif = await Image.LoadAsync<Rgba32>(imagePath, ct);
// 1. Detect subtitle region (bottom 30% of frames)
var subtitleRegion = DetectSubtitleRegion(gif);
// 2. Extract frames with text changes
var uniqueFrames = ExtractUniqueTextFrames(gif, subtitleRegion);
// 3. Extract tight bounding boxes around text
var textRegions = ExtractTextBoundingBoxes(uniqueFrames);
// 4. Create horizontal strip of text-only regions
var strip = CreateHorizontalStrip(textRegions);
return new TextOnlyStrip
{
Image = strip,
RegionCount = textRegions.Count,
TotalTokens = EstimateTokens(strip),
OriginalTokens = EstimateTokens(gif),
Reduction = CalculateReduction(strip, gif)
};
}
private Rectangle DetectSubtitleRegion(Image<Rgba32> gif)
{
// Analyze bottom 30% of frame for text patterns
int subtitleHeight = (int)(gif.Height * 0.3);
int subtitleY = gif.Height - subtitleHeight;
return new Rectangle(0, subtitleY, gif.Width, subtitleHeight);
}
private List<Image<Rgba32>> ExtractUniqueTextFrames(
Image<Rgba32> gif,
Rectangle subtitleRegion)
{
var uniqueFrames = new List<Image<Rgba32>>();
Image<Rgba32>? previousFrame = null;
for (int i = 0; i < gif.Frames.Count; i++)
{
var frame = gif.Frames.CloneFrame(i);
var subtitleCrop = frame.Clone(ctx =>
ctx.Crop(subtitleRegion));
// Compare with previous frame
if (previousFrame == null ||
HasTextChanged(subtitleCrop, previousFrame, threshold: 0.05))
{
uniqueFrames.Add(subtitleCrop);
previousFrame = subtitleCrop;
}
}
return uniqueFrames;
}
private bool HasTextChanged(
Image<Rgba32> current,
Image<Rgba32> previous,
double threshold)
{
// Threshold bright pixels (white/yellow text on dark background)
var currentBright = CountBrightPixels(current);
var previousBright = CountBrightPixels(previous);
// Calculate Jaccard similarity of bright pixels
var intersection = currentBright.Intersect(previousBright).Count();
var union = currentBright.Union(previousBright).Count();
var similarity = union > 0 ? intersection / (double)union : 1.0;
// Text changed if similarity drops below threshold
return similarity < (1.0 - threshold);
}
// Helper type for bounding box + crop
private record TextCrop
{
public required Image<Rgba32> CroppedImage { get; init; }
public required Rectangle Bounds { get; init; }
}
private List<TextCrop> ExtractTextBoundingBoxes(
List<Image<Rgba32>> frames)
{
var textCrops = new List<TextCrop>();
foreach (var frame in frames)
{
// Threshold to get text mask
var mask = ThresholdBrightPixels(frame, minValue: 200);
// Find connected components (text regions)
var components = FindConnectedComponents(mask);
// Get tight bounding box around all components
var bbox = GetTightBoundingBox(components);
// Add padding
bbox.Inflate(5, 5);
// Clone the region (dispose properly in production!)
var cropped = frame.Clone(ctx => ctx.Crop(bbox));
textCrops.Add(new TextCrop
{
CroppedImage = cropped,
Bounds = bbox
});
}
return textCrops;
}
private Image<Rgba32> CreateHorizontalStrip(
List<TextCrop> textCrops)
{
// Calculate strip dimensions
int totalWidth = textCrops.Sum(c => c.Bounds.Width);
int maxHeight = textCrops.Max(c => c.Bounds.Height);
// Create blank canvas
var strip = new Image<Rgba32>(totalWidth, maxHeight);
// Paste text regions horizontally
int xOffset = 0;
foreach (var crop in textCrops)
{
strip.Mutate(ctx => ctx.DrawImage(
crop.CroppedImage,
new Point(xOffset, 0),
opacity: 1.0f));
xOffset += crop.Bounds.Width;
// Dispose crop after use (important!)
crop.CroppedImage.Dispose();
}
return strip;
}
}
इनपुट: anchorman-not-even-mad.gif (93 frames, 300×185)
प्रक्रमण:
1. Detect subtitle region: bottom 30% (300×55)
2. Extract unique frames: 93 frames → 2 text changes
3. Extract tight bounding boxes:
- Frame 1-45: "I'm not even mad." → 252×49 bbox
- Frame 46-93: "That's amazing." → 198×49 bbox
4. Create horizontal strip: 450×49 total
आउटपुटपाठ

टोकन अर्थशास्त्र:
30× कटौती सब उपशीर्षक पाठ को सुरक्षित रखते हुए
जब दोनों Tesseract और फ्लोरेंस -2 विफल या कम उत्पादन करता है - विश्वास परिणामों को | , | एक दृश्य LLM के लिए escalate
public class OcrQualityWave : IAnalysisWave
{
private readonly SpellChecker _spellChecker;
public string Name => "OcrQualityWave";
public int Priority => 58; // After Florence-2 and Tesseract
public async Task<IEnumerable<Signal>> AnalyzeAsync(
string imagePath,
AnalysisContext context,
CancellationToken ct)
{
var signals = new List<Signal>();
// Get best OCR result from earlier waves (priority order)
string? ocrText =
context.GetValue<string>("ocr.ml.text") ?? // Florence-2 (priority 51)
context.GetValue<string>("ocr.text"); // Tesseract (priority 50)
if (string.IsNullOrWhiteSpace(ocrText))
{
signals.Add(new Signal
{
Key = "ocr.quality.no_text",
Value = true,
Confidence = 1.0,
Source = Name
});
return signals;
}
// Run spell check (deterministic quality assessment)
var spellResult = _spellChecker.CheckTextQuality(ocrText);
// Additional quality signals to avoid false positives
var alphanumRatio = CalculateAlphanumericRatio(ocrText); // Letters/digits vs junk
var avgTokenLength = CalculateAverageTokenLength(ocrText);
signals.Add(new Signal
{
Key = "ocr.quality.spell_check_score",
Value = spellResult.CorrectWordsRatio,
Confidence = 1.0,
Source = Name,
Metadata = new Dictionary<string, object>
{
["total_words"] = spellResult.TotalWords,
["correct_words"] = spellResult.CorrectWords,
["garbled_words"] = spellResult.GarbledWords,
["alphanum_ratio"] = alphanumRatio,
["avg_token_length"] = avgTokenLength
}
});
// Deterministic escalation threshold
// NOTE: Spellcheck alone can false-trigger on proper nouns, memes, brand names.
// Use additional signals (alphanum ratio, token length) to reduce false escalations.
bool isGarbled = spellResult.CorrectWordsRatio < 0.5 &&
alphanumRatio > 0.7; // Mostly valid characters, just not in dictionary
signals.Add(new Signal
{
Key = "ocr.quality.is_garbled",
Value = isGarbled,
Confidence = 1.0,
Source = Name
});
// Signal Vision LLM escalation
if (isGarbled)
{
signals.Add(new Signal
{
Key = "ocr.quality.escalation_required",
Value = true,
Confidence = 1.0,
Source = Name,
Tags = new List<string> { "action_required", "escalation" },
Metadata = new Dictionary<string, object>
{
["reason"] = "spell_check_below_threshold",
["quality_score"] = spellResult.CorrectWordsRatio,
["threshold"] = 0.5,
["target_tier"] = "vision_llm"
}
});
// Cache garbled text for Vision LLM to access
context.SetCached("ocr.garbled_text", ocrText);
}
return signals;
}
}
उत्तेजना निर्णायक है: वर्तनी जाँच प्राप्ताङ्क
जब एनिमेटेड GIFs के लिए एस्केलेशन ट्रिगर किया जाता है
public class VisionLlmWave : IAnalysisWave
{
private readonly IVisionLlmClient _client;
public string Name => "VisionLlmWave";
public int Priority => 50;
public async Task<IEnumerable<Signal>> AnalyzeAsync(
string imagePath,
AnalysisContext context,
CancellationToken ct)
{
var signals = new List<Signal>();
// Check if escalation is required
var escalationRequired = context.GetValue<bool>(
"ocr.quality.escalation_required");
if (!escalationRequired)
{
signals.Add(new Signal
{
Key = "vision.llm.skipped",
Value = true,
Confidence = 1.0,
Source = Name,
Metadata = new Dictionary<string, object>
{
["reason"] = "no_escalation_required"
}
});
return signals;
}
// For animated GIFs, use text-only strip
string imageToProcess = imagePath;
bool usedFilmstrip = false;
if (context.GetValue<int>("identity.frame_count") > 1)
{
var filmstrip = await CreateTextOnlyStripAsync(imagePath, ct);
imageToProcess = filmstrip.Path;
usedFilmstrip = true;
signals.Add(new Signal
{
Key = "vision.filmstrip.created",
Value = true,
Confidence = 1.0,
Source = Name,
Metadata = new Dictionary<string, object>
{
["mode"] = "text_only",
["region_count"] = filmstrip.RegionCount,
["token_reduction"] = filmstrip.Reduction,
["original_tokens"] = filmstrip.OriginalTokens,
["final_tokens"] = filmstrip.TotalTokens
}
});
}
// Build constrained prompt
var prompt = BuildConstrainedPrompt(context);
// Call Vision LLM
var result = await _client.ExtractTextAsync(
imageToProcess,
prompt,
ct);
// Emit OCR text signal (Vision LLM tier)
signals.Add(new Signal
{
Key = "ocr.vision.text", // Vision LLM OCR result
Value = result.Text,
Confidence = 0.95, // High but not 1.0 - still probabilistic
Source = Name,
Tags = new List<string> { "ocr", "vision", "llm" },
Metadata = new Dictionary<string, object>
{
["model"] = result.Model,
["used_filmstrip"] = usedFilmstrip,
["inference_time_ms"] = result.InferenceTime,
["token_count"] = result.TokenCount,
["cost_usd"] = result.Cost
}
});
// Optionally emit caption if requested (separate from OCR)
if (result.Caption != null)
{
signals.Add(new Signal
{
Key = "caption.text", // Descriptive caption, not OCR
Value = result.Caption,
Confidence = 0.90,
Source = Name,
Tags = new List<string> { "caption", "description" }
});
}
return signals;
}
private string BuildConstrainedPrompt(AnalysisContext context)
{
var sb = new StringBuilder();
sb.AppendLine("Extract all text from this image.");
sb.AppendLine();
sb.AppendLine("CONSTRAINTS:");
sb.AppendLine("- Only extract text that is actually visible");
sb.AppendLine("- Preserve formatting and line breaks");
sb.AppendLine("- If no text is present, return empty string");
sb.AppendLine();
// Add context from earlier waves
var garbledText = context.GetCached<string>("ocr.garbled_text");
if (!string.IsNullOrEmpty(garbledText))
{
sb.AppendLine("CONTEXT:");
sb.AppendLine("Traditional OCR detected garbled text:");
sb.AppendLine($" \"{garbledText}\"");
sb.AppendLine("Use this as a hint for stylized or unusual fonts.");
sb.AppendLine();
}
sb.AppendLine("Return only the extracted text, no commentary.");
return sb.ToString();
}
}
जब सभी स्तर पूर्ण हो जाएँ तो अंतिम पाठ चयन एक कठोर प्राथमिकता क्रम का प्रयोग करता है
public static string? GetFinalText(DynamicImageProfile profile)
{
// Priority chain (highest to lowest quality)
// NOTE: This selects ONE source, but the ledger exposes ALL sources
// with confidence scores for downstream inspection
// 1. Vision LLM OCR (best for complex/garbled text)
var visionText = profile.GetValue<string>("ocr.vision.text");
if (!string.IsNullOrEmpty(visionText))
return visionText;
// 2. Florence-2 multi-frame GIF OCR (best for animations)
var florenceMultiText = profile.GetValue<string>("ocr.ml.multiframe_text");
if (!string.IsNullOrEmpty(florenceMultiText))
return florenceMultiText;
// 3. Florence-2 single-frame ML OCR (good for stylized fonts)
var florenceText = profile.GetValue<string>("ocr.ml.text");
if (!string.IsNullOrEmpty(florenceText))
return florenceText;
// 4. Tesseract OCR (reliable for clean standard text)
var tesseractText = profile.GetValue<string>("ocr.text");
if (!string.IsNullOrEmpty(tesseractText))
return tesseractText;
// 5. Fallback (empty)
return string.Empty;
}
प्रत्येक स्तर के ज्ञात लक्षण हैं:
स्रोत | | | संकेत कुंजी || | विश्वसनीयता के लिए सर्वोत्तम
|--------|------------|----------|------------|------|-------|
दृश्य एलएम ओसीआर ocr.vision.text जटिल चार्ट
फ्लोरेंस ocr.ml.multiframe_text उपशीर्षक के साथ एनिमित GIFs
फ्लोरेंस ocr.ml.text शैलीकृत फ़ॉन्ट
Tesseract ocr.text शुद्ध मानक पाठ
छवियाँ
100 images × $0.005/image = $0.50
Total time: 100 × 2s = 200 seconds
रूट वितरण
Cost:
60 × $0 = $0
25 × $0 = $0
10 × $0.005 = $0.05
5 × $0.002 = $0.01
Total: $0.06
Time:
60 × 0.1s = 6s
25 × 0.3s = 7.5s
10 × 2s = 20s
5 × 2.5s = 12.5s
Total: 46 seconds
Savings:
Cost: 88% reduction ($0.50 → $0.06)
Time: 77% reduction (200s → 46s)
मध्य स्तर (Florence-2) शून्य लागत पर छवियों को संभालता है
यहाँ subtitles के साथ एक meme GIF के लिए पूरा प्रवाह है
1. Load image: anchorman-not-even-mad.gif (93 frames)
2. IdentityWave (priority 10):
→ identity.frame_count = 93
→ identity.format = "gif"
→ identity.is_animated = true
3. TextLikelinessWave (priority 40, ~10ms):
→ Heuristic text detection: 15 regions in bottom 30%
→ Subtitle pattern: DETECTED
→ text.likeliness = 0.85
4. OcrWave (priority 50, ~60ms):
→ Run Tesseract OCR on first frame
→ ocr.text = "I'm not emn mad." (garbled)
→ ocr.confidence = 0.62
5. MlOcrWave (priority 51, ~180ms):
→ Tesseract confidence < 0.95, run Florence-2
→ Sample 10 frames (animated GIF)
→ Run Florence-2 on each frame (parallel)
→ Deduplicate: 10 results → 2 unique texts
→ ocr.ml.multiframe_text = "I'm not even mad.\nThat's amazing."
→ ocr.ml.confidence = 0.91
6. OcrQualityWave (priority 58, ~5ms):
→ Check Florence-2 result
→ Spell check: 6/6 words correct (100%)
→ ocr.quality.is_garbled = false
→ ocr.quality.escalation_required = false
7. VisionLlmWave (priority 80, SKIPPED):
→ No escalation required (Florence-2 succeeded)
Final output:
Text: "I'm not even mad.\nThat's amazing."
Source: ocr.ml.multiframe_text
Confidence: 0.91
Cost: $0 (local processing)
Time: ~250ms total (Tesseract + Florence-2)
अगर फ्लोरेंस-2 असफल होता है M SK1अभिमानता < 0.5),प्रवाह जारी रहेगा
6. OcrQualityWave:
→ Spell check: 2/6 words correct (33%)
→ ocr.quality.is_garbled = true
→ ocr.quality.escalation_required = true
7. VisionLlmWave:
→ Create text-only filmstrip (2 regions, 450×49)
→ Send to Vision LLM: "Extract all text from this strip"
→ vision.llm.text = "I'm not even mad.\nThat's amazing."
→ Confidence: 0.95
→ Cost: ~$0.002 (30× token reduction vs full frames)
→ Time: ~2.3s
तीन स्तरीय प्रणाली पूरी तरह से कॉन्फ़िगर की जा सकती है
{
"DocSummarizer": {
"Ocr": {
"Tesseract": {
"Enabled": true,
"DataPath": "/usr/share/tesseract-ocr/4.00/tessdata",
"Languages": ["eng"],
"EarlyExitThreshold": 0.95
},
"Florence2": {
"Enabled": true,
"ModelPath": "models/florence2-base",
"ConfidenceThreshold": 0.85,
"MaxFrames": 10,
"DeduplicationMethod": "levenshtein",
"LevenshteinThreshold": 0.85
},
"Quality": {
"SpellCheckThreshold": 0.5,
"EscalationEnabled": true
}
},
"VisionLlm": {
"Enabled": true,
"Provider": "ollama",
"OllamaUrl": "http://localhost:11434",
"Model": "minicpm-v:8b",
"MaxRetries": 3,
"TimeoutSeconds": 30
},
"Filmstrip": {
"TextOnlyMode": true,
"SubtitleRegionPercent": 0.3,
"BrightPixelThreshold": 200,
"TextChangeThreshold": 0.05
},
"Routing": {
"FastRouteConfidence": 0.8,
"BalancedRouteConfidence": 0.5,
"TextDetectionEnabled": true
}
}
}
विफलता |---------|-----------|----------| | टेसरेक्ट असफल | विश्वसनीयता | फ्लोरेंस विश्वसनीयता | दृश्य एलएलएम समय समाप्ति अनुरोध से अधिक है 30s | | | सर्वोत्तम उपलब्ध ओसीआर परिणाम पर वापस लौटें | सभी स्तर असफल सभी परिणाम खाली या टूटे हुए | निर्दिष्टता के साथ रिक्त स्ट्रिंग लौटाएँ | API लागत सीमा दैनिक बजट से अधिक | विकलांग दृश्य एलएलएम | नमूना उपलब्ध नहीं | फ्लोरेंस-2/विजन एलएलएम ऑफ़लाइन
हर विफलता निर्णायक होती है और पूर्ण प्रमाणन के साथ लॉग की जाती है
For each image:
1. Run Tesseract
2. If looks wrong, manually fix or skip
Problems:
- No middle tier (binary: works or doesn't)
- Manual intervention required
- No cost optimization
For each image:
1. Send to GPT-4o/Claude
2. Pay $0.005-0.01 per image
Problems:
- Expensive (85% of images could be free)
- Slow (network latency)
- Still hallucinates without constraints
For each image:
1. OpenCV text detection (5-20ms, free)
2. Route to appropriate tier
3. Florence-2 handles 85% locally (200ms, free)
4. Vision LLM only for complex cases (2-5s, $0.001-0.01)
Benefits:
- 88% cost reduction
- 77% faster (most images process locally)
- Deterministic escalation (auditable)
- Filmstrip optimization (30× token reduction)
- Constrained by deterministic signals
तीन -tier OCR पाइपलाइन यह साबित करता है कि लागत---सूचित रूटिंग और स्थानीय-प्रथम प्रसंस्करण गुणवत्ता का त्याग किए बिना निष्पादन और आर्थिक दोनों में नाटकीय रूप से सुधार कर सकता है
प्रमुख अंतर्दृष्टि
पैटर्न स्केल स्थानीय निर्णायक विश्लेषण → स्थानीय एमएल मॉडलज्ञात विशेषताओं और लागत व्यापार के साथ प्रत्येक स्तर ,
यह OCR : निर्णायक संकेतों पर लागू अवरोधित अस्पष्टता है।
| भाग | पैटर्न | |
|---|---|---|
| 1 | अवरोधित अस्पष्टता एकल घटक | |
| 2 | प्रतिबंधित अस्पष्ट मोएम बहुआयामी अवयव | |
| 3 | संदर्भ खींचना समय | |
| 4 | छवि इंटेलिजेंस तरंग वास्तुकला | |
| 4.1 | तीनों -Tier OCR पाइपलाइन | OCR, ONNX मॉडलों, फिल्म ट्रिप |
अगलाभाग 5 दिखाएगा कि कैसे ImageSummarizer डॉक-सममैरिजर, और डेटा-सममारीजरName LucidRAG के साथ multi-मोडल ग्राफ RAG में सम्मिलित करें
सभी भाग एक ही अपरिवर्ती का अनुसरण करते हैं संभाव्यात्मक घटक प्रस्तावित.
© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.