LLMs (Large Language Models)सेंसर के रूप में प्रयोग किया जा रहा है श्रेणी त्रुटिएक संभाव्य संश्लेषक का उपयोग करते हुए जहां एक निर्धारात्मक सीमा युक्ति की आवश्यकता होती है
यह नहीं है-'\ t विकासकर्ताओं\ '\ t त्रुटि\ .\ most mainstream examples and tutorials lead with the simplest demo\ :\ "\ just send it to the model\ msc5\ That makes onboarding easy\ m sc6\ but it blurs a crucial boundary\ msec7\ t perception vs synthesis\ mc8\ t Industry incentives don\ mSc9\ t help either\ mSC10\ t token\ mcc11\ tpriced systems naturally reward pipelines that do more work inside the LLM\ mcs12\
इस लेख के बारे में है कम RAG - एक वास्तुकला नमूना जहां निर्णायक पाइपलाइन्स संभाव्यात्मक घटकों को खिलाते हैंसही RAG प्रणाली में
उदाहरण: ऑप्टिकल क्यारेक्टर पहचान में | ( | ऑप्टिक क्यारेटर पहचान |) | पाइपलाइन | МSK3 | विज़न एलएलएम टियर है & #44; 3, | टियर नहीं & #39; 1. | यह केवल पाठ के बाद ही चलता है | - | समानता हियूरिसिक्स और स्थानीय ओसीआर विफल होता है सीमा नियम एक वास्तुकला अवरोध जो perceived सुविधा के बावजूद अनुचित उपकरण उपयोग को रोकता है
यहाँ विभिन्न डोमेनों में समान वास्तुकला असफलताएं हैं
flowchart TD
subgraph Wrong["❌ Common Anti-Pattern"]
Raw[Raw Data<br/>Pixels, Waveforms, Frames] -->|Direct feed| LLM1[Vision/Audio LLM<br/>$$$, variance, hallucination]
LLM1 --> Unreliable[Unreliable Output<br/>High cost, non-deterministic]
end
style Wrong fill:none,stroke:#dc2626,stroke-width:3px
style Raw fill:none,stroke:#6b7280,stroke-width:2px
style LLM1 fill:none,stroke:#dc2626,stroke-width:2px
style Unreliable fill:none,stroke:#dc2626,stroke-width:2px
इन वास्तुकला गलतियों से पूर्वानुमानीय असफलताएं उत्पन्न होती हैं
भ्रमित अनुभूति: LLMs होना चाहिए टोकन जारी करें. अनिश्चितता के तहत वे संभाव्य पूर्णताओं से रिक्तियों को भरते हैं यह मुख्य समस्या है
Non-deterministic failure: तापमान और टोकन बजट ड्राइव भिन्नता
संसाधन अपशिष्ट: क्या आप'एक टोकन के लिए भुगतान कर रहे हैं या स्थानीय मॉडल चला रहे है, या नहीं
मुद्दा यह है-' cost - it ' accuracy स्थानीय रूप से होस्ट किया गया एलएलएम भ्रमित ओसीआर परिणाम उतना ही टूटा है जितना कि एक महंगा API कॉल एक ही काम करता है
श्रेणी त्रुटि मौजूद है क्योंकि एक संवेदक और एक संश्लेषक मूलतः भिन्न उपकरण हैं
एक संवेदक एक सीमांत युक्ति है
एक एलएलएम एक उच्च वर्ण संश्लेषक है
flowchart LR
subgraph Sensor["Sensor (Boundary Device)"]
World[Physical World<br/>∞ dimensions] -->|Reduce| Signal[Structured Signal<br/>Bounded dimensions]
Signal -->|Confidence| Facts[Facts<br/>± certainty]
end
subgraph LLM["LLM (Synthesizer)"]
Input[Structured Input] -->|Synthesize| Prose[Unstructured Output<br/>High entropy]
Prose -->|No confidence| Tokens[Token stream<br/>No 'nothing' state]
end
style Sensor fill:none,stroke:#16a34a,stroke-width:2px
style LLM fill:none,stroke:#b45309,stroke-width:2px
style World fill:none,stroke:#6b7280,stroke-width:2px
style Signal fill:none,stroke:#2563eb,stroke-width:2px
style Facts fill:none,stroke:#059669,stroke-width:2px
style Input fill:none,stroke:#6b7280,stroke-width:2px
style Prose fill:none,stroke:#d97706,stroke-width:2px
style Tokens fill:none,stroke:#dc2626,stroke-width:2px
इसीलिए स्टिलियोफ्लो संकेतों को अपरिवर्तनीय तथ्य के रूप में व्यवहार करता है
मस्तिष्क तर्क से नहीं आरंभ करता है बन्धन.
रीटीना कोर्टिक्स नहीं है
जब संवेदी अवरोध कमजोर होते हैं कम प्रकाश , missing edges | , | ambiguous cues |. | This is not a metaphor | - | it is the same failure mode as LLM hallucinations under uncertainty | l. | Engineers already know this | : | when upstream SNR drops
यह समस्या---संशोधन पैटर्न सभी पशु संवेदन प्रणालियों में प्रकट होता है
बैट प्रतिध्वनि निर्धारण लेखापरीक्षा
मधुमक्खियों की दृष्टि गति पता लगाना
flowchart LR
subgraph Bat["Bat Echolocation"]
Echo[Ultrasonic Echo<br/>∞ waveform data] --> Cochlea[Cochlear Filters<br/>Doppler, delay, amplitude]
Cochlea --> BatBrain[Auditory Cortex<br/>Distance, texture facts]
end
subgraph Bee["Honeybee Vision"]
Motion[Visual Field<br/>Rapid motion] --> Lamina[Lamina<br/>Optical flow computation]
Lamina --> BeeBrain[Central Brain<br/>Motion vectors, not pixels]
end
style Bat fill:none,stroke:#7c3aed,stroke-width:2px
style Bee fill:none,stroke:#d97706,stroke-width:2px
style Echo fill:none,stroke:#6b7280,stroke-width:2px
style Cochlea fill:none,stroke:#a855f7,stroke-width:2px
style BatBrain fill:none,stroke:#6366f1,stroke-width:2px
style Motion fill:none,stroke:#6b7280,stroke-width:2px
style Lamina fill:none,stroke:#f59e0b,stroke-width:2px
style BeeBrain fill:none,stroke:#d97706,stroke-width:2px
सामान्य पैटर्न
यह प्रकृति से प्रेरित नहीं है।
इंजीनियरिंग मैपिंग सटीक है
flowchart TD
subgraph Brain["Biological Vision Pipeline"]
Photons[Photons] --> Retina[Retina<br/>Edge detection, contrast]
Retina --> V1[V1 Cortex<br/>Orientation, motion]
V1 --> IT[Inferotemporal Cortex<br/>Object recognition]
IT --> PFC[Prefrontal Cortex<br/>Reasoning, synthesis]
end
subgraph Engineering["Engineering Vision Pipeline"]
Pixels[Raw Pixels] --> OpenCV[OpenCV + Heuristics<br/>Sharpness, text-likeliness]
OpenCV --> Local[Local Models<br/>Florence-2, EAST/CRAFT OCR]
Local --> Structured[Structured Signals<br/>Bounding boxes, confidence]
Structured --> LLM[LLM Synthesis<br/>Only when needed]
end
Brain -.->|Maps to| Engineering
style Brain fill:none,stroke:#7c3aed,stroke-width:2px
style Engineering fill:none,stroke:#2563eb,stroke-width:2px
style Photons fill:none,stroke:#6b7280,stroke-width:2px
style Retina fill:none,stroke:#16a34a,stroke-width:2px
style V1 fill:none,stroke:#059669,stroke-width:2px
style IT fill:none,stroke:#0891b2,stroke-width:2px
style PFC fill:none,stroke:#6366f1,stroke-width:2px
style Pixels fill:none,stroke:#6b7280,stroke-width:2px
style OpenCV fill:none,stroke:#16a34a,stroke-width:2px
style Local fill:none,stroke:#059669,stroke-width:2px
style Structured fill:none,stroke:#0891b2,stroke-width:2px
style LLM fill:none,stroke:#6366f1,stroke-width:2px
मेंदू नहीं समझता
कार्टेक्स उन संकेतों पर काम करता है जो कम किए गए हैं।
यह एक दार्शनिक दृष्टिकोण नहीं है
उत्तेजना रोकी जानी चाहिए।: एलएलएम टियर हैं | 3, | नहीं टियर |1. | सिर्फ जब सस्ती सेंसर असफल होते हैं बुलाएँ
विश्वास थ्रेसहोल्ड स्पष्ट होना चाहिएविश्वास के साथ पता चला गया पाठ
पथ निर्धारणात्मक होना चाहिए: समान सिग्नल
टोकन अर्थशास्त्र एक अवरोध है: यदि आपके पाइपलाइन को raw data आकार के साथ लागत मापन करता है
तथ्यों को प्रमाण की आवश्यकता है: यदि आप कर सकते हैं, तो यह एक तथ्य नहीं है।
प्रमाण: फिल्म ट्रिप अनुकूलन
में वीडियोसममैजर पाइपलाइनकेवल फिल्म ट्रिप्स द्वारा टोकन लागत को ~30x से कम करते हैं जबकि सुधार ओसीआर विश्वसनीयता. एलएलएम संकेत को देखता है | ( | निकाले गए पाठ क्षेत्रों को |), | दृश्य नहीं
यह एक सुधार नहीं है
यह है कम RAG वास्तुकला पैटर्न पाइपलाइन्स संभाव्यात्मक घटकों को खिसकाते हैं.
घटा RAG मानचित्र है
पारंपरिक RAG यह पीछला पेज होम पेज अगला पेज प्राप्त करता है - यह दस्तावेजों को पुनःप्राप्त करता है और आशा करता है कि LLM तथ्य निकालता है पहले तथ्य निकालता है तब LLM संश्लेषित करता है
पैटर्न सभी मल्टीमोडल प्रणालियों में दोहराता है
flowchart TD
subgraph Map["MAP PHASE (Parallel, Deterministic)"]
Raw[Raw Data<br/>10,000 frames] --> Split{Split}
Split --> S1[Sensors<br/>Frame 1-1000]
Split --> S2[Sensors<br/>Frame 1001-2000]
Split --> S3[Sensors<br/>Frame 2001-3000]
Split --> SDots[...]
S1 --> L1[Local Models<br/>Batch 1]
S2 --> L2[Local Models<br/>Batch 2]
S3 --> L3[Local Models<br/>Batch 3]
SDots --> LDots[...]
L1 --> F1[Facts: 120]
L2 --> F2[Facts: 98]
L3 --> F3[Facts: 156]
LDots --> FDots[...]
F1 --> Collect[Collect Facts]
F2 --> Collect
F3 --> Collect
FDots --> Collect
Collect --> Facts[(Facts Database<br/>500 total facts)]
end
subgraph Reduce["REDUCE PHASE (Sequential, Probabilistic)"]
Query[User Query] --> Retrieve[Retrieve Relevant Facts<br/>Filter: 50 facts]
Facts --> Retrieve
Retrieve --> LLM[LLM Synthesis<br/>Reason over 50 facts]
LLM --> Answer[Grounded Answer]
end
style Map fill:none,stroke:#16a34a,stroke-width:3px
style Reduce fill:none,stroke:#6366f1,stroke-width:3px
style Raw fill:none,stroke:#6b7280,stroke-width:2px
style Split fill:none,stroke:#16a34a,stroke-width:2px
style S1 fill:none,stroke:#16a34a,stroke-width:2px
style S2 fill:none,stroke:#16a34a,stroke-width:2px
style S3 fill:none,stroke:#16a34a,stroke-width:2px
style SDots fill:none,stroke:#16a34a,stroke-width:1px,stroke-dasharray: 5 5
style L1 fill:none,stroke:#059669,stroke-width:2px
style L2 fill:none,stroke:#059669,stroke-width:2px
style L3 fill:none,stroke:#059669,stroke-width:2px
style LDots fill:none,stroke:#059669,stroke-width:1px,stroke-dasharray: 5 5
style F1 fill:none,stroke:#0891b2,stroke-width:2px
style F2 fill:none,stroke:#0891b2,stroke-width:2px
style F3 fill:none,stroke:#0891b2,stroke-width:2px
style FDots fill:none,stroke:#0891b2,stroke-width:1px,stroke-dasharray: 5 5
style Collect fill:none,stroke:#16a34a,stroke-width:2px
style Facts fill:none,stroke:#0891b2,stroke-width:3px
style Query fill:none,stroke:#6b7280,stroke-width:2px
style Retrieve fill:none,stroke:#7c3aed,stroke-width:2px
style LLM fill:none,stroke:#6366f1,stroke-width:2px
style Answer fill:none,stroke:#16a34a,stroke-width:2px
दस्तावेज़-first RAG vsM SK1 Reduced RAG:
| Aspect | Document-first RAG | Reduced RAG | ||||||
|---|---|---|---|---|---|---|---|---|
| पैटर्न | पुनःप्राप्त करें | → | निकालें | → | संश्लेषित करें | | नक्शा | МSK4 | बाहर निकालें |
| निष्कर्षण यथार्थता LLM भ्रम संभव है | ||||||||
| भंडारित डेटा दस्तावेज़ | ||||||||
| एलएलएम भूमिका दो कार्य | ||||||||
| त्रुटिमोचन क्षमता प्रांप्ट ट्रेस निरीक्षण करें | ||||||||
| मापनीयता | अनुक्रमिक LLM bottleneck |
इस पैटर्न को तीन उत्पादन प्रणालियां लागू करती हैं
Here are real numbers from VideoSummarizer on a 10-minute video
दृष्टिकोण |----------|----------|-------------|--------| | फ्रेम | भ्रम , विचलन | | | Non |- | Deterministic | | शॉट्स → कुंजीफ्रेम → LLM यथार्थ | निर्णायक निष्कर्षण | फिल्मस्ट्रिप पाठ निष्कर्षण | सर्वोत्तम ओसीआर विश्वसनीयता
सही वास्तुकला अधिक सटीक है यह भी हो जाता है कि 180x सस्ता - लेकिन यह ,' सही बात करने का एक गौण प्रभाव है , , लक्ष्य नहीं है .
अगर आपका एआई सिस्टम LLM के साथ शुरू होता है, तो आप पहले से ही इसका नियंत्रण खो चुके हैं
बुद्धि का आरंभ तर्क से नहीं होता है
संवेदक अनिश्चितता को कम करते हैं
synthesis अंतिम चरण बनाएँ
© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.