This is a viewer only at the moment see the article on how this works.
To update the preview hit Ctrl-Alt-R (or ⌘-Alt-R on Mac) or Enter to refresh. The Save icon lets you save the markdown file to disk
This is a preview from the server running through my markdig pipeline
Saturday, 08 November 2025
快速实施快速API(https://github.com/UKPLab/EasyNMT)是一个极好但被遗弃的神经机器翻译项目,并直接复制了EasyNMT的API(https://github.com/UKPLab/EasyNMT)。
但我添加了SOMY的好功能 来提高可靠性 并准备在生产系统中使用思考快速翻译自我托管...).
自博客开始以来, 便一直热衷于自动翻译博客文章。
是的,我知道在浏览器... etc... 但是这不是重点。mostlylucid-nmt我想知道该怎么做!
也很高兴能欢迎不读英文的人(即使他们把英语作为第二语言阅读,于是我设计了如何做到这一点, 并分享了如何建立这种系统。哦,它给了我如何在ASP.NET中使用它的想法,以便利用信号和光滑实时更新系统自动确定文本(包括动态文本)的位置。 (ASP.NET)
听好! 听好! 听好基本上; 人类写垃圾文字 超音速的噪音让机器能高效处理
所以很多人都在想 如何解决与EasyNMT(这其实是一个研究项目)有关的问题。
一一写入整个系统http://<server>:<port>/demo与一个叫EasyNMT的惊人项目一起实现这一点。
这是一个简单、快速的方法 获得翻译API 不需要支付某些服务 或者运行一个全尺寸的LLM 来得到翻译( 低调) 。
但随着时间的流逝 裂缝开始显现出来是时候做点更好的事了
**跟往常一样,都是在吉特Hub上 免费使用等等...**Dock 拆船船
Reusing loaded model for en->de (3/10 models in cache)Need to load model for en->fr (3/10 models in cache)主要最新情况(v3.1) - 情报和可见度新建于 v3.1 :
====================================================================================================
🚀 DOWNLOADING MODEL
Model: facebook/mbart-large-50-many-to-many-mmt
Family: mbart50
Direction: en → bn
Device: GPU (cuda:0)
Total Size: 2.46 GB
Files: 6 main files
====================================================================================================
[Progress bars for each file...]
====================================================================================================
✅ MODEL READY
Model: facebook/mbart-large-50-many-to-many-mmt
Translation: en → bn is now available
====================================================================================================
**: 默认缓存大小缩到 10 个模型( 从 6 个) 。**每模型设备记录
[Pivot] Languages reachable from en: 85 languages
[Pivot] Languages that can reach bn: 42 languages
[Pivot] Found 38 possible pivot languages
[Pivot] Selected pivot: en → hi → bn (both legs verified)
**:以横幅显示目标设备(GPU/CPU)**进度栏
Request: en→bn with opus-mt
Trying families: ['opus-mt', 'mbart50', 'm2m100'] ✓ All three!
opus-mt: Failed (model doesn't exist)
mbart50: Success! (auto-fallback worked)
数据驱动智能中枢选择- 不再有盲目的企图:
Loading mbart50 model on GPU (cuda:0)Model loaded on device: cuda:0Successfully loaded... on GPU (cuda:0){\fn方正黑体简体\fs18\b1\bord1\shad1\3cH2F2F2F}Esbn不存在enbn 示例
en->hi, hi->bn[Pivot] Both legs loaded and cached. Ready to translate.:看看为什么选择或跳过每个枢轴4. 4个。
model_family总是尝试后退:不再两次重试同一模式实例流动
**成功的信息包括:**6 . 6 . 6 .
**- 已经在演示中工作 :**演示下调允许选择 opus- mt、 mbart50 或 m2m100
requirements-prod.txt强化演示页面**- 制作即时互动接口:**100vw/100vh) 用于缩入翻译体验的全视图端布局( 100vw/ 100vh)
**: FP16 已启用, BATCH_ SIZE=64, MAX_ INFLight=1 (对单个 GPU 而言最佳)**CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU CPU
**- 较小、更快的图像:**从生产结构中去除的试验依赖性( 热、 热、 热、 热)
综合测试和加载测试- 验证一切:
具有现实交通模式跨平台验证脚本
/discover/opus-mt( Power Shell + Bash + Bash )/discover/mbart50模型下载和主轴翻转回回溯测试/discover/m2m100用于快速验证的自动烟雾测试5 个部署文件
Kubernets 以聚氯乙烯、资源限量、健康检查单列零散集装箱实例
scottgal/mostlylucid-nmt:cpu负载测试指导和监测建议:latest解释货币对流量对流量的权衡取舍scottgal/mostlylucid-nmt:cpu-min6 . 6 . 6 .scottgal/mostlylucid-nmt:gpu三个模范家庭scottgal/mostlylucid-nmt:gpu-min- 选择适合你需要的最佳选择:OPMT 执行-执行MT1200+对,最佳质量(独立模型)
latest, min, gpu, gpu-min:50种语言,单2.4GB模式,2 450对20250108.143022:100种语言,单2.2GB模式,9 900对自动后退- 明智地选择现有最佳模式:
- 所有的MBART50双对- M2M100全对
无预加载模型( 按需下载)
不重建的交换式家庭模式10 点单个单个嵌入器仓库
cpu(或)latest) | scottgal/mostlylucid-nmt:cpu- CPU
| cpu-min | scottgal/mostlylucid-nmt:cpu-min- CPU 最小
| gpu | scottgal/mostlylucid-nmt:gpu- GPU与CUDA 12.6
| gpu-min | scottgal/mostlylucid-nmt:gpu-min- GPU 最小11. 十一、十一、十一、十一、十一、十、十一、十一、十一、十一、十一、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十、十十、十、十、十、十、十、十、十、十、十十、十、十、十、十十、十、十十十、十十十、十、十十、十十、十正确版本
用于跟踪版本、构建日期和Git承诺的完整 OCI 标签
docker run -d \
--name mostlylucid-nmt \
-p 8000:8000 \
scottgal/mostlylucid-nmt
12. 十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二、十二
curl -X POST "http://localhost:8000/translate" \
-H "Content-Type: application/json" \
-d '{
"text": ["Hello, how are you?"],
"target_lang": "de"
}'
最新基础图像
{
"translated": ["Hallo, wie geht es Ihnen?"],
"target_lang": "de",
"source_lang": "en",
"translation_time": 0.34
}
Python 3. 12- lim
docker run -d \
--name mostlylucid-nmt \
--gpus all \
-p 8000:8000 \
-e EASYNMT_MODEL_ARGS='{"torch_dtype":"fp16"}' \
scottgal/mostlylucid-nmt:gpu
CUDA 12.6
GPU 图像 Ubuntu 24.04 (最新 NVIDIA 堆栈)
docker run -d \
--name mostlylucid-nmt \
-p 8000:8000 \
-v $HOME/model-cache:/models \
-e MODEL_CACHE_DIR=/models \
scottgal/mostlylucid-nmt:cpu-min
12.4 使用CUDA火炬
docker run -d `
--name mostlylucid-nmt `
-p 8000:8000 `
-v ${HOME}/model-cache:/models `
-e MODEL_CACHE_DIR=/models `
scottgal/mostlylucid-nmt:cpu-min
(与CUDA 12.6运行时间兼容)
docker run -d ^
--name mostlylucid-nmt ^
-p 8000:8000 ^
-v %USERPROFILE%/model-cache:/models ^
-e MODEL_CACHE_DIR=/models ^
scottgal/mostlylucid-nmt:cpu-min
所有附属关系均更新为最新安全版本
13 号
curl http://localhost:8000/healthz
固定折旧折旧警告
已删除已折旧的 Transformers_CACHE(现在使用高频HOME)与变换器 v5 兼容" 快速启动 " (5分钟)
http://localhost:8000/demo/
这是最简单的最简单的方法,
可用的嵌入夹图像
标记 完整图像名称 大小 说明 使用案例
最小图像
使集装箱尺寸保持小
需要 NVIDIA 嵌入运行时间 :
健康检查:
// Example: Translating a 5000-word article
Input: Long article with multiple paragraphs
Step 1: Split by paragraphs (preserves structure)
→ Paragraph 1 (800 chars)
→ Paragraph 2 (1200 chars)
→ Paragraph 3 (600 chars)
...
Step 2: Group into ~1000 character chunks
→ Chunk 1: Paragraphs 1-2
→ Chunk 2: Paragraph 3-4
→ Chunk 3: Paragraphs 5-6
Step 3: Translate each chunk sequentially
→ Shows progress: "Translating chunk 1/3..."
→ Shows progress: "Translating chunk 2/3..."
→ Shows progress: "Translating chunk 3/3..."
Step 4: Reassemble with paragraph breaks
→ Final output: Complete translated article with preserved formatting
现场服务自动流行的语言下载
自动自动处理任意大小的大文本输入
3 个
高级选项:
句号拆分:
实时统计:
翻译 完成/错误
:100种语言,9 900对/demo/在翻译前,请确切查看哪些语文配对可用
演示工具在客户端执行智能文字块块:为什么用演示?快速测试没有写入代码的测试翻译校验语文配对可用语文
比较翻译质量和不同光束大小的翻译质量
粘贴整个博客文章( 5000+单词)演示自动将它挤成可操作的碎片每个块翻译时显示进度
语言检测以未知语言粘贴文本点击“ 检测语言”
**没有后压或排队。**发送太多的要求, 它只是冷淡的翻过来。MODEL_FAMILY无法观察。
# Opus-MT (default, best quality)
MODEL_FAMILY=opus-mt
# mBART50 (50 languages, single model)
MODEL_FAMILY=mbart50
# M2M100 (100 languages, broadest coverage)
MODEL_FAMILY=m2m100
解决办法:最清晰的NMT所以... 我决定建立一个新的 和改良的容易NMT,现在多数为卢布- nmt
MODEL_FAMILY这就是让事情变得更好的原因:opus-mt多模式家庭支助Opus- OPM(赫尔辛基- NLP) - 默认
# Set primary to Opus-MT (best quality)
MODEL_FAMILY=opus-mt
AUTO_MODEL_FALLBACK=1
MODEL_FALLBACK_ORDER=opus-mt,mbart50,m2m100
# Request Ukrainian → French
# 1. Try Opus-MT first (not available)
# 2. Automatically fall back to mBART50 (available!)
# 3. Translation succeeds with mBART50
覆盖面:
每个方向300-500MB
# Enable auto-fallback (default: enabled)
AUTO_MODEL_FALLBACK=1
# Set fallback priority (default: opus-mt → mbart50 → m2m100)
MODEL_FALLBACK_ORDER="opus-mt,mbart50,m2m100"
# Disable for strict single-family mode
AUTO_MODEL_FALLBACK=0
示例:
-min模型大小 :优点:
易换换的易
最强大的新特征之一是
**如何运作:**您设置了初级
效益:
支持100+语言,不管理多种部署
透明记录 :
多模式家庭支助
# NMT: Fits on a USB stick
du -sh model-cache/
2.5G model-cache/
# LLM: Needs serious storage
du -sh llama-models/
140G llama-models/
模型发现终点
符号?
LRU 模型缓存
EasiNMT 兼容性API
用于较小部署量的量控缓存变量。
为什么NMT会超过LLMs来翻译?
Input: "The API returns a 429 status code when rate limited."
NMT (Opus-MT): "Die API gibt einen 429-Statuscode zurück, wenn sie ratenbegrenzt ist."
(Accurate, preserves technical terms)
LLM (might do): "Die API sendet den Fehlercode 429, wenn zu viele Anfragen gestellt werden."
(Interprets rather than translates, adds context not in original)
以下是基于生产用途的现实检查:
GPT-4:每个请求3-10秒(API延时+一代)
GPT-4 APP**:30-60秒**当地Llama 70B
行动执行-MT(每个方向):300-500MB
mBART50(所有50种语言):2.4GBM2M100(所有100种语言):2.2GB
flowchart LR
A[HTTP Client] --> B[API Gateway]
B --> C[Translation Endpoint]
C --> D{Has Capacity?}
D -->|Yes| E[Translation Service]
D -->|No| F[Queue with 429]
F --> E
E --> G[Process Pipeline]
G --> H[Get Model from Cache]
H --> I[Translate]
I --> J[Return Response]
J --> A
LLM :
Retry-After费用:10-20美元/月**NMT的优势:**专门翻译培训Retry-After一致质量(相同的输入=相同的产出)
不需要即时工程
sequenceDiagram
participant Client
participant API
participant Queue
participant Translator
participant Cache
participant Model
Client->>API: POST /translate
API->>Queue: Acquire slot
alt Queue has space
Queue-->>API: Slot acquired
API->>Translator: Process translation
Translator->>Translator: Sanitize input
Translator->>Translator: Split sentences
Translator->>Translator: Chunk text
Translator->>Translator: Mask symbols
Translator->>Cache: Get model (en→de)
alt Cache hit
Cache-->>Translator: Return cached model
else Cache miss
Cache->>Model: Load from Hugging Face
Model-->>Cache: Pipeline loaded
Cache->>Cache: Evict old if at capacity
Cache-->>Translator: Return model
end
Translator->>Model: Translate batches
Model-->>Translator: Translations
Translator->>Translator: Unmask symbols
Translator->>Translator: Post-process
Translator-->>API: Translations
API->>Queue: Release slot
API-->>Client: 200 OK + translations
else Queue full
Queue-->>API: Overflow error
API-->>Client: 429 Too Many Requests\nRetry-After: X seconds
end
实例设想:
graph LR
A[Raw Input] --> B{Sanitize?}
B -->|Yes| C[Check Noise]
B -->|No| D[Split Sentences]
C -->|Is Noise| Z[Return Placeholder]
C -->|Valid| D
D --> E[Enforce Max Length]
E --> F[Chunk for Batching]
F --> G{Symbol Masking?}
G -->|Yes| H[Mask Digits/Punct/Emoji]
G -->|No| I[Translate]
H --> I
I --> J{Direct Model?}
J -->|Available| K[Direct Translation]
J -->|Not Available| L{Pivot Fallback?}
L -->|Yes| M[src→en→tgt]
L -->|No| Z
K --> N[Unmask Syis robust input handling. Here's what happens:
**Noise Detection:**
- Strips control characters (except \t, \n, \r)
- Checks minimum character count (default: 1)
- Calculates alphanumeric ratio (default: must be ≥20%)
- Rejects pure emoji, pure punctuation, or pure whitespace
**Symbol Masking:**
Why mask symbols? Translation models are trained on text, not emoji or special symbols. These can confuse them or get mangled. So we:
1. Extract all digits, punctuation, and emoji as contiguous runs
2. Replace them with sentinel tokens: `⟪MSK0⟫`, `⟪MSK1⟫`, etc.
3. Translate the masked text
4. Restore the original symbols in their positions
Example:
Input: "Hello 👋 world! Price: $99.99" 何时使用 (👋) (!) (:) ($99.99)
**Post-Processing:**
After translation, we remove "symbol loops" - repeated symbols that weren't in the source:
使用 NMT( 大多为薄荷- nmt) 时 : 您需要规模化的一致、快速翻译 预算事项(自行托管或数量大)
### Sentence Splitting & Chunking
Long texts get split intelligently:
```mermaid
graph TD
A[Long Text] --> B[Split on . ! ? …]
B --> C{Sentence > 500 chars?}
C -->|Yes| D[Split on word boundaries]
C -->|No| E[Keep sentence]
D --> E
E --> F[Group into chunks ≤900 chars]
F --> G[Translate each chunk]
G --> H[Join with space]
你在翻译技术内容 代码 结构化数据
你需要创造性的适应,而不是字面翻译
stateDiagram-v2
[*] --> CheckCache
CheckCache --> CacheHit: Model exists
CheckCache --> CacheMiss: Model not loaded
CacheHit --> MoveToEnd: Update LRU order
MoveToEnd --> ReturnModel
CacheMiss --> CheckCapacity
CheckCapacity --> LoadModel: Space available
CheckCapacity --> EvictOldest: Cache full
EvictOldest --> MoveToCPU: Free VRAM
MoveToCPU --> ClearCUDA: torch.cuda.empty_cache()
ClearCUDA --> LoadModel
LoadModel --> AddToCache
AddToCache --> ReturnModel
ReturnModel --> [*]
环境和文化细微差别比速度重要得多
(我的用法案例)NMT是明显的赢家:
# Semaphore limits concurrent translations
MAX_INFLIGHT = 1 # On GPU, 1 at a time for efficiency
MAX_QUEUE_SIZE = 1000 # Up to 1000 waiting
# When full:
# - Returns 429 Too Many Requests
# - Includes Retry-After header
# - Estimates wait time based on average duration
将100+博客文章翻译为12种语言, 时间为~30分钟( GPU)
avg_duration = 2.5 seconds (tracked with EMA)
waiters = 100
slots = 1
estimated_wait = (100 / 1) * 2.5 = 250 seconds
clamped = min(250, 120) = 120 seconds
Retry-After: 120
所有员额的一贯质量(所有员额)
graph LR
A[Ukrainian Text] --> B{Direct uk→fr?}
B -->|Exists| C[Translate Directly]
B -->|Missing| D[Pivot via English]
D --> E[uk→en]
E --> F[en→fr]
F --> G[French Result]
C --> G
全部设置: 一个嵌入容器
单是速度差异就使得NMT成为生产翻译管道的唯一实际选择。
NMT是专门为翻译而设计的,用微薄的硬件运行,速度比LLMS快10-100x。 如果您需要快速、一致、符合成本效益的大规模翻译,NMT就会得手。
# src/core/cache.py
from collections import OrderedDict
import torch
class LRUPipelineCache:
"""LRU cache that automatically cleans up GPU memory when evicting models."""
def __init__(self, capacity: int):
self.cache = OrderedDict() # Maintains insertion order
self.capacity = capacity
def get(self, key: str):
"""Get model from cache, moves it to end (most recently used)."""
if key not in self.cache:
return None
self.cache.move_to_end(key) # Mark as recently used
return self.cache[key]
def put(self, key: str, value):
"""Add model to cache, evicting oldest if at capacity."""
if key in self.cache:
self.cache.move_to_end(key)
else:
self.cache[key] = value
# If cache is full, evict the oldest model
if len(self.cache) > self.capacity:
oldest_key, oldest_pipeline = self.cache.popitem(last=False)
# MAGIC: Move evicted model to CPU to free GPU memory
try:
oldest_pipeline.model.to("cpu")
if torch.cuda.is_available():
torch.cuda.empty_cache() # Tell GPU to release memory
logger.info(f"Evicted {oldest_key}, freed GPU memory")
except Exception as e:
logger.warning(f"Failed to clean GPU memory: {e}")
建筑结构概览
OrderedDict请求流量直截了当:立即向翻译处提出请求
# src/services/model_manager.py
def get_pipeline(self, src: str, tgt: str):
"""Try to get translation model, with automatic fallback to other providers."""
# Determine which model families support this language pair
families_to_try = []
if config.AUTO_MODEL_FALLBACK:
# Try families in priority order: opus-mt → mbart50 → m2m100
for family in config.MODEL_FALLBACK_ORDER.split(","):
if self._is_pair_supported(src, tgt, family.strip()):
families_to_try.append(family.strip())
# Try each family until one succeeds
last_error = None
for family in families_to_try:
try:
model_name, src_lang, tgt_lang, _ = self._get_model_name_and_langs(src, tgt, family)
if family != config.MODEL_FAMILY:
logger.info(f"Using fallback '{family}' for {src}->{tgt}")
# Load the model from HuggingFace
pipeline = transformers.pipeline(
"translation",
model=model_name,
device=device_manager.device_index,
src_lang=src_lang,
tgt_lang=tgt_lang
)
self.cache.put(f"{src}->{tgt}", pipeline)
return pipeline
except Exception as e:
last_error = e
logger.warning(f"Family '{family}' failed for {src}->{tgt}: {e}")
continue # Try next family
# All families failed
raise ModelLoadError(f"{src}->{tgt}", last_error)
否 无
(LRU) 提供翻译模型:
# src/services/queue_manager.py
import asyncio
from contextlib import asynccontextmanager
class QueueManager:
"""Manages request queuing and backpressure."""
def __init__(self, max_inflight: int, max_queue: int):
self.semaphore = asyncio.Semaphore(max_inflight) # Limit concurrent translations
self.max_queue_size = max_queue
self.waiting_count = 0
self.inflight_count = 0
self.avg_duration_sec = 5.0 # Exponential moving average
@asynccontextmanager
async def acquire_slot(self):
"""Try to get a translation slot, track metrics, handle queueing."""
# Check if queue is too full
if self.waiting_count >= self.max_queue_size:
# Calculate how long client should wait before retrying
retry_after = self._estimate_retry_after()
raise QueueOverflowError(self.waiting_count, retry_after)
self.waiting_count += 1
try:
# Wait for available slot (this is the queue!)
await self.semaphore.acquire()
self.waiting_count -= 1
self.inflight_count += 1
start_time = time.time()
yield # Let the translation happen
# Update average duration for retry-after estimates
duration = time.time() - start_time
alpha = config.RETRY_AFTER_ALPHA # Smoothing factor (0.2)
self.avg_duration_sec = alpha * duration + (1 - alpha) * self.avg_duration_sec
finally:
self.inflight_count -= 1
self.semaphore.release()
def _estimate_retry_after(self) -> int:
"""Smart calculation: how many waiting / how many slots * avg time per request."""
if self.inflight_count == 0:
return config.RETRY_AFTER_MIN_SEC
# If 10 people waiting and 2 slots available, and each takes 5 seconds:
# retry_after = (10 / 2) * 5 = 25 seconds
retry_sec = (self.waiting_count / self.semaphore._value) * self.avg_duration_sec
# Clamp between min and max
return max(
config.RETRY_AFTER_MIN_SEC,
min(int(retry_sec), config.RETRY_AFTER_MAX_SEC)
)
Cache Cache Chash 快速反应
max_inflight)@asynccontextmanager返回客户端输入处理管道
# src/utils/symbol_masking.py
import re
def mask_symbols(text: str) -> tuple[str, dict[str, str]]:
"""Replace special symbols with placeholders before translation."""
originals = {}
masked_text = text
placeholder_counter = 0
# Pattern: Match emojis, symbols, special punctuation
# \U0001F300-\U0001F9FF = emoji range
# [\u2600-\u26FF\u2700-\u27BF] = misc symbols
symbol_pattern = re.compile(
r'[\U0001F300-\U0001F9FF\u2600-\u26FF\u2700-\u27BF'
r'\u00A9\u00AE\u2122\u2139\u3030\u303D\u3297\u3299]+'
)
for match in symbol_pattern.finditer(text):
symbol = match.group()
placeholder = f"__SYMBOL_{placeholder_counter}__"
originals[placeholder] = symbol
masked_text = masked_text.replace(symbol, placeholder, 1)
placeholder_counter += 1
return masked_text, originals
def unmask_symbols(text: str, originals: dict[str, str]) -> str:
"""Restore original symbols after translation."""
for placeholder, original in originals.items():
text = text.replace(placeholder, original)
return text
该服务使用复杂的多阶段管道处理混乱的现实世界文本:
# Before translation:
text = "Hello! 👋 Check out this cool feature 🚀"
# Mask symbols:
masked, originals = mask_symbols(text)
# masked = "Hello! __SYMBOL_0__ Check out this cool feature __SYMBOL_1__"
# originals = {"__SYMBOL_0__": "👋", "__SYMBOL_1__": "🚀"}
# Translate the masked text:
translated = translate(masked, "de") # → "Hallo! __SYMBOL_0__ Schau dir diese coole Funktion an __SYMBOL_1__"
# Unmask symbols:
final = unmask_symbols(translated, originals)
# final = "Hallo! 👋 Schau dir diese coole Funktion an 🚀"
面罩:"你好""世界""世界""世界""世界""世界""世界""世界"世界""价格"价格""价格""价格"MS"K2" MSK3""
👋模型不会因为大量投入而窒息__SYMBOL_0__我们可以高效地分批为什么这很重要:
# src/utils/text_processing.py
def chunk_sentences(sentences: list[str], max_chars: int = 900) -> list[list[str]]:
"""Group sentences into chunks that fit within model's max input length."""
chunks = []
current_chunk = []
current_length = 0
for sentence in sentences:
sentence_len = len(sentence)
# If this sentence alone is too long, it goes in its own chunk
if sentence_len > max_chars:
if current_chunk:
chunks.append(current_chunk)
current_chunk = []
current_length = 0
chunks.append([sentence])
continue
# If adding this sentence exceeds limit, start new chunk
if current_length + sentence_len + 1 > max_chars:
chunks.append(current_chunk)
current_chunk = [sentence]
current_length = sentence_len
else:
current_chunk.append(sentence)
current_length += sentence_len + 1 # +1 for space
# Don't forget the last chunk!
if current_chunk:
chunks.append(current_chunk)
return chunks
def split_sentences(text: str, max_sentence_chars: int = 500) -> list[str]:
"""Split text into sentences, enforcing max length."""
# Split on common sentence terminators
sentences = re.split(r'([.!?…]+\s+)', text)
result = []
for sentence in sentences:
if not sentence or sentence.isspace():
continue
# If sentence is too long, split on word boundaries
if len(sentence) > max_sentence_chars:
words = sentence.split()
current = []
current_len = 0
for word in words:
if current_len + len(word) + 1 > max_sentence_chars:
result.append(' '.join(current))
current = [word]
current_len = len(word)
else:
current.append(word)
current_len += len(word) + 1
if current:
result.append(' '.join(current))
else:
result.append(sentence.strip())
return result
GPU 内存值非常宝贵
.!?…我们让最近6个模特儿保持热辣通过英语发挥枢纽作用:
# src/services/model_discovery.py
import httpx
from datetime import datetime, timedelta
class ModelDiscoveryService:
"""Discovers available translation models with 1-hour cache."""
def __init__(self):
self._cache = {} # Cache results to avoid hammering HuggingFace API
self._cache_ttl = timedelta(hours=1)
self._hf_api_base = "https://huggingface.co/api/models"
async def discover_opus_mt_pairs(self, force_refresh: bool = False):
"""Query HuggingFace for all Helsinki-NLP Opus-MT models."""
cache_key = "opus-mt"
# Check cache first
if not force_refresh and cache_key in self._cache:
cached_data, cached_time = self._cache[cache_key]
if datetime.now() - cached_time < self._cache_ttl:
return cached_data # Cache hit!
# Cache miss - query HuggingFace API
async with httpx.AsyncClient() as client:
response = await client.get(
self._hf_api_base,
params={
"author": "Helsinki-NLP",
"search": "opus-mt",
"limit": 1000
},
timeout=30.0
)
models = response.json()
# Extract language pairs from model names
# Example: "Helsinki-NLP/opus-mt-en-de" → ("en", "de")
pairs = []
for model in models:
model_id = model.get("modelId", "")
if model_id.startswith("Helsinki-NLP/opus-mt-"):
# Extract the language codes after "opus-mt-"
lang_part = model_id.replace("Helsinki-NLP/opus-mt-", "")
if "-" in lang_part:
src, tgt = lang_part.split("-", 1)
pairs.append({"source": src, "target": tgt})
# Cache the results
self._cache[cache_key] = (pairs, datetime.now())
return pairs
这增加了双倍的延迟时间,但确保了所有辅助语文对口的覆盖面。
httpx让我们探索一下代码库中最有趣的部分!en最酷的特征之一是智能模型缓存,知道如何处理 GPU 内存:de这是怎么回事?Helsinki-NLP/opus-mt-en-deGPU 清理
# src/core/device.py
import torch
class DeviceManager:
"""Smart device selection with GPU auto-detection."""
def __init__(self):
self.use_gpu = self._should_use_gpu()
self.device_index = self._resolve_device()
self.device_str = "cpu" if self.device_index < 0 else f"cuda:{self.device_index}"
# Auto-configure parallel translation slots based on device
if self.device_index >= 0:
# GPU: Run translations serially to avoid VRAM fragmentation
self.max_inflight = 1
else:
# CPU: Can handle multiple translations in parallel
self.max_inflight = config.MAX_WORKERS_BACKEND
self._log_device_info()
def _should_use_gpu(self) -> bool:
"""Check if GPU should be used."""
if config.USE_GPU.lower() == "false":
return False
if config.USE_GPU.lower() == "true":
return torch.cuda.is_available()
# "auto" mode: use GPU if available
return torch.cuda.is_available()
def _resolve_device(self) -> int:
"""Returns device index: -1 for CPU, 0+ for CUDA."""
if not self.use_gpu:
return -1
# Check if specific CUDA device requested
if config.DEVICE and config.DEVICE.startswith("cuda:"):
device_num = int(config.DEVICE.split(":")[1])
return device_num
return 0 # Use first GPU
def _log_device_info(self):
"""Log device information at startup."""
if self.device_index >= 0:
gpu_name = torch.cuda.get_device_name(self.device_index)
vram_gb = torch.cuda.get_device_properties(self.device_index).total_memory / 1e9
logger.info(f"Using GPU: {gpu_name} ({vram_gb:.1f}GB VRAM)")
logger.info(f"Max inflight translations: {self.max_inflight} (GPU mode)")
else:
cpu_count = os.cpu_count()
logger.info(f"Using CPU ({cpu_count} cores)")
logger.info(f"Max inflight translations: {self.max_inflight} (CPU mode)")
# Global singleton instance
device_manager = DeviceManager()
: 在驱逐模型时, 我们明确将其移动到 CPU 内存, 并告诉 CPU 释放资源
max_inflight=1这个聪明的功能会自动尝试多个 AI 模式提供者, 如果第一个没有您需要的语言配对 :max_inflight=4这是怎么回事?DEVICE=cuda:1:成功模型用语言对配键缓存
# Snippet from QueueManager showing EMA calculation
def update_avg_duration(self, new_duration: float):
"""Update average duration using exponential moving average."""
# EMA formula: new_avg = α × new_value + (1 - α) × old_avg
# α = smoothing factor (0.0 to 1.0)
# - Higher α = more weight to recent values (faster adaptation)
# - Lower α = more weight to historical values (more stable)
alpha = 0.2 # 20% weight to new value, 80% to historical
self.avg_duration_sec = (
alpha * new_duration +
(1 - alpha) * self.avg_duration_sec
)
3 个
# Initial average: 5.0 seconds
# New request takes: 10.0 seconds
# EMA calculation:
new_avg = 0.2 * 10.0 + 0.8 * 5.0
= 2.0 + 4.0
= 6.0 seconds
# Next request takes: 3.0 seconds
new_avg = 0.2 * 3.0 + 0.8 * 6.0
= 0.6 + 4.8
= 5.4 seconds
请求用后压队列( HTTP 429)
Retry-After指数移动平均数: 平滑出请求期间的峰值
临时临时
翻译模式有时会腐蚀或删除emojis----这完全保护了他们!
# Prefer GPU if available (default)
USE_GPU=auto
# Force GPU
USE_GPU=true
# Force CPU
USE_GPU=false
# Explicit device override
DEVICE=cuda:0
DEVICE=cpu
# Model family selection (NEW in v2.0!)
MODEL_FAMILY=opus-mt # Best quality (default)
MODEL_FAMILY=mbart50 # 50 languages, single model
MODEL_FAMILY=m2m100 # 100 languages, maximum coverage
# Auto-fallback between model families (NEW in v2.0!)
AUTO_MODEL_FALLBACK=1 # Enabled by default
MODEL_FALLBACK_ORDER="opus-mt,mbart50,m2m100" # Priority order
# Volume-mapped model cache (NEW in v2.0!)
MODEL_CACHE_DIR=/models # Persistent cache directory
# Model arguments passed to transformers.pipeline
EASYNMT_MODEL_ARGS='{"torch_dtype":"fp16"}'
EASYNMT_MODEL_ARGS='{"torch_dtype":"bf16","cache_dir":"/models"}'
# Preload models at startup (reduces first-request latency)
PRELOAD_MODELS="en->de,de->en,fr->en"
# LRU cache capacity
MAX_CACHED_MODELS=6
将长长的文字分解为符合模式限制的块块,同时保留句子界限:
**这是怎么回事?**分句判刑
opus-mt:使用 Regex 分割时mbart50保存标点m2m100贪婪区块:将尽可能多的句子包装到每个块块中,但不超过限制字词边界分割
1:如果单句太长,则在空格上拆分,而不是削减中字0为什么重要**:翻译模型有输入限制(通常为512-1024象征性)。**这确保了我们永远不超越它们,同时保持上下文不变。
"opus-mt,mbart50,m2m100"与 Caching 同步模式发现同步模式"m2m100,mbart50,opus-mt"这是怎么回事?Async HTTP 客户端:将非阻塞 HTTP 的请求发送到 Hugging 脸孔
/models:为避免限制费率,1小时的仓储结果-v ./model-cache:/models和
fp16调自bf16为什么重要fp32: Hugging Face 有 1200+ Opus-MT 模型。# Batch size for translation (higher = faster but more VRAM)
EASYNMT_BATCH_SIZE=16 # CPU: 8-16, GPU: 32-64
# Maximum text length per item
EASYNMT_MAX_TEXT_LEN=1000
# Maximum beam size (higher = better quality but slower)
EASYNMT_MAX_BEAM_SIZE=5
# Worker thread pools
MAX_WORKERS_BACKEND=1 # Translation workers
MAX_WORKERS_FRONTEND=2 # Language detection workers
# Enable request queueing (highly recommended)
ENABLE_QUEUE=1
# Max concurrent translations
# Auto: 1 on GPU, MAX_WORKERS_BACKEND on CPU
MAX_INFLIGHT_TRANSLATIONS=1
# Max queued requests before 429
MAX_QUEUE_SIZE=1000
# Per-request timeout (0 = disabled)
TRANSLATE_TIMEOUT_SEC=180
# Retry-After estimation
RETRY_AFTER_MIN_SEC=1 # Floor
RETRY_AFTER_MAX_SEC=120 # Ceiling
RETRY_AFTER_ALPHA=0.2 # EMA smoothing factor
# Enable input filtering
INPUT_SANITIZE=1
# Minimum alphanumeric ratio (0.2 = 20%)
INPUT_MIN_ALNUM_RATIO=0.2
# Minimum character count
INPUT_MIN_CHARS=1
# Language code for undetermined/noise
UNDETERMINED_LANG_CODE=und
# Default sentence splitting behavior
PERFORM_SENTENCE_SPLITTING_DEFAULT=1
# Max chars per sentence before word-boundary split
MAX_SENTENCE_CHARS=500
# Max chars per chunk for batching
MAX_CHUNK_CHARS=900
# Sentence joiner
JOIN_SENTENCES_WITH=" "
# Enable symbol masking
SYMBOL_MASKING=1
# What to mask
MASK_DIGITS=1 # Mask 0-9
MASK_PUNCT=1 # Mask .,!? etc.
MASK_EMOJI=1 # Mask 😀🎉 etc.
# Align response array length to input
ALIGN_RESPONSES=1
# Placeholder for failed items (when aligned)
SANITIZE_PLACEHOLDER=""
# Response format
EASYNMT_RESPONSE_MODE=strings # ["translation1", "translation2"]
EASYNMT_RESPONSE_MODE=objects # [{"text":"translation1"}, ...]
# Enable two-hop translation via pivot
PIVOT_FALLBACK=1
# Pivot language (usually English)
PIVOT_LANG=en
# Log level
LOG_LEVEL=INFO
# Per-request logging (verbose)
REQUEST_LOG=1
# Format
LOG_FORMAT=plain # Human-readable
LOG_FORMAT=json # Structured JSON
# File logging with rotation
LOG_TO_FILE=1
LOG_FILE_PATH=/var/log/marian-translator/app.log
LOG_FILE_MAX_BYTES=10485760 # 10MB
LOG_FILE_BACKUP_COUNT=5
# Include raw text in logs (privacy risk!)
LOG_INCLUDE_TEXT=0
# Periodically clear CUDA cache (seconds, 0=disabled)
CUDA_CACHE_CLEAR_INTERVAL_SEC=0
# Worker count (use 1 for single GPU)
WEB_CONCURRENCY=1
# Request timeout
TIMEOUT=60
# Graceful shutdown timeout
GRACEFUL_TIMEOUT=20
# Keep-alive timeout
KEEP_ALIVE=5
# GET request
curl "http://localhost:8000/translate?target_lang=de&text=Hello%20world&source_lang=en"
# Response
{
"translations": ["Hallo Welt"]
}
# POST request
curl -X POST http://localhost:8000/translate \
-H 'Content-Type: application/json' \
-d '{
"text": [
"Hello world",
"This is a test",
"Machine translation is amazing"
],
"target_lang": "de",
"source_lang": "en",
"beam_size": 1,
"perform_sentence_splitting": true
}'
# Response
{
"target_lang": "de",
"source_lang": "en",
"translated": [
"Hallo Welt",
"Das ist ein Test",
"Maschinenübersetzung ist erstaunlich"
],
"translation_time": 0.342
}
# Omit source_lang for auto-detection
curl -X POST http://localhost:8000/translate \
-H 'Content-Type: application/json' \
-d '{
"text": ["Bonjour le monde"],
"target_lang": "en"
}'
# Response
{
"target_lang": "en",
"source_lang": "fr", # Detected
"translated": ["Hello world"],
"translation_time": 0.156
}
# GET
curl "http://localhost:8000/language_detection?text=Hola%20mundo"
# {"language": "es"}
# POST with batch
curl -X POST http://localhost:8000/language_detection \
-H 'Content-Type: application/json' \
-d '{"text": ["Hello", "Bonjour", "Hola"]}'
# {"languages": ["en", "fr", "es"]}
# Health check
curl http://localhost:8000/healthz
# {"status": "ok"}
# Readiness
curl http://localhost:8000/readyz
# {
# "status": "ready",
# "device": "cuda:0",
# "queue_enabled": true,
# "max_inflight": 1
# }
# Cache status
curl http://localhost:8000/cache
# {
# "capacity": 6,
# "size": 3,
# "keys": ["en->de", "de->en", "fr->en"],
# "device": "cuda:0",
# "inflight": 1,
# "queue_enabled": true
# }
# Model info
curl http://localhost:8000/model_name | jq
# When queue is full, you get 429
curl -X POST http://localhost:8000/translate \
-H 'Content-Type: application/json' \
-d '{"text": ["test"], "target_lang": "de"}'
# Response: 429 Too Many Requests
# Headers: Retry-After: 45
# Body:
{
"message": "Too many requests; queue full",
"retry_after_sec": 45
}
# Proper client behavior:
# 1. Read Retry-After header
# 2. Wait that long + jitter
# 3. Retry request
适应实际请求期限的模拟重试后估计值:
示例:
.\build-all.ps1
这是怎么回事?
chmod +x build-all.sh
./build-all.sh
:像一个更重视最近值的加权平均值平滑系数(α):
latest, min, gpu, gpu-min为什么不是简单平均?20250108.143022:使客户现实适应当前系统负荷的时间
# Always get the latest version
docker pull scottgal/mostlylucid-nmt:cpu
# Or use the :latest alias
docker pull scottgal/mostlylucid-nmt:latest
# Pin to a specific version for reproducibility
docker pull scottgal/mostlylucid-nmt:cpu-20250108.143022
docker pull scottgal/mostlylucid-nmt:cpu-min-20250108.143022
资源管理 资源管理 资源管理
:智能缓存、挤块和平行处理
docker inspect scottgal/mostlylucid-nmt:cpu | jq '.[0].Config.Labels'
可观察性详细伐木和指标跟踪.
# Using pre-built image from Docker Hub (recommended)
docker run -d \
--name translator \
-p 8000:8000 \
-e ENABLE_QUEUE=1 \
-e MAX_QUEUE_SIZE=500 \
-e EASYNMT_BATCH_SIZE=16 \
-e TIMEOUT=180 \
-e LOG_LEVEL=INFO \
-e REQUEST_LOG=0 \
scottgal/mostlylucid-nmt
# Or build locally
docker build -t mostlylucid-nmt .
docker run -d --name translator -p 8000:8000 mostlylucid-nmt
# Check logs
docker logs -f translator
# Using pre-built GPU image from Docker Hub (recommended)
docker run -d \
--name translator-gpu \
--gpus all \
-p 8000:8000 \
-e USE_GPU=true \
-e DEVICE=cuda:0 \
-e PRELOAD_MODELS="en->de,de->en,en->fr,fr->en,en->es,es->en" \
-e EASYNMT_MODEL_ARGS='{"torch_dtype":"fp16"}' \
-e EASYNMT_BATCH_SIZE=64 \
-e MAX_CACHED_MODELS=8 \
-e ENABLE_QUEUE=1 \
-e MAX_QUEUE_SIZE=2000 \
-e WEB_CONCURRENCY=1 \
-e TIMEOUT=180 \
-e GRACEFUL_TIMEOUT=30 \
-e LOG_FORMAT=json \
-e LOG_TO_FILE=1 \
-v /var/log/translator:/var/log/marian-translator \
scottgal/mostlylucid-nmt:gpu
# Or build locally
docker build -f Dockerfile.gpu -t mostlylucid-nmt:gpu .
docker run -d --name translator-gpu --gpus all -p 8000:8000 mostlylucid-nmt:gpu
# Monitor cache and performance
watch -n 5 "curl -s http://localhost:8000/cache | jq"
version: '3.8'
services:
translator:
image: scottgal/mostlylucid-nmt:gpu # Use pre-built image
container_name: translator
restart: unless-stopped
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 1
capabilities: [gpu]
ports:
- "8000:8000"
environment:
USE_GPU: "true"
DEVICE: "cuda:0"
PRELOAD_MODELS: "en->de,de->en,en->fr,fr->en"
EASYNMT_MODEL_ARGS: '{"torch_dtype":"fp16"}'
EASYNMT_BATCH_SIZE: "64"
MAX_CACHED_MODELS: "8"
ENABLE_QUEUE: "1"
MAX_QUEUE_SIZE: "2000"
WEB_CONCURRENCY: "1"
TIMEOUT: "180"
LOG_FORMAT: "json"
LOG_TO_FILE: "1"
volumes:
- translator-logs:/var/log/marian-translator
- translator-cache:/root/.cache/huggingface
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/healthz"]
interval: 30s
timeout: 10s
retries: 3
start_period: 40s
volumes:
translator-logs:
translator-cache:
apiVersion: apps/v1
kind: Deployment
metadata:
name: translator
spec:
replicas: 2 # Scale horizontally for CPU, use 1 per GPU
selector:
matchLabels:
app: translator
template:
metadata:
labels:
app: translator
spec:
containers:
- name: translator
image: scottgal/mostlylucid-nmt:gpu
ports:
- containerPort: 8000
env:
- name: USE_GPU
value: "true"
- name: EASYNMT_MODEL_ARGS
value: '{"torch_dtype":"fp16"}'
- name: PRELOAD_MODELS
value: "en->de,de->en"
- name: ENABLE_QUEUE
value: "1"
- name: MAX_QUEUE_SIZE
value: "2000"
resources:
requests:
memory: "4Gi"
cpu: "2"
nvidia.com/gpu: 1
limits:
memory: "8Gi"
cpu: "4"
nvidia.com/gpu: 1
livenessProbe:
httpGet:
path: /healthz
port: 8000
initialDelaySeconds: 30
periodSeconds: 10
readinessProbe:
httpGet:
path: /readyz
port: 8000
initialDelaySeconds: 20
periodSeconds: 5
---
apiVersion: v1
kind: Service
metadata:
name: translator
spec:
selector:
app: translator
ports:
- port: 80
targetPort: 8000
type: LoadBalancer
模模L_家庭
EASYNMT_MODEL_ARGS='{"torch_dtype":"fp16"}'
:质量良好,100种语言,单一2.2GB模式
# Start high, reduce if you get OOM
EASYNMT_BATCH_SIZE=64 # Try 128 on large GPUs
缩写: 缩写: 缩写: 缩写: 缩写: 缩写: 缩写: 缩写:
PRELOAD_MODELS="en->de,de->en,en->fr,fr->en,en->es,es->en"
:当对无配对时,自动尝试其他家庭
WEB_CONCURRENCY=1
MAX_INFLIGHT_TRANSLATIONS=1
(默认):启用 - 最大覆盖率
MAX_CACHED_MODELS=10 # Keep more models in VRAM
:残疾 -- -- 严格的单一家庭模式
# beam_size=1 is 3-5x faster than beam_size=5
# Quality difference is often minimal
curl -X POST ... -d '{"beam_size": 1, ...}'
: 后退优先排序
EASYNMT_BATCH_SIZE=8
默认 :
MAX_WORKERS_BACKEND=4
MAX_INFLIGHT_TRANSLATIONS=4
WEB_CONCURRENCY=2
(质量第一)
PERFORM_SENTENCE_SPLITTING_DEFAULT=0
(涵盖第一个)
// Bad: 100 separate requests
for (const text of texts) {
await translate(text);
}
// Good: 1 batch request
await translate(texts);
MDEL_ CACHE_DIR 蒙古
async function translateWithRetry(texts) {
try {
return await translate(texts);
} catch (err) {
if (err.status === 429) {
const retryAfter = err.headers['retry-after'];
const jitter = Math.random() * 5;
await sleep((retryAfter + jitter) * 1000);
return translateWithRetry(texts);
}
throw err;
}
}
:通过 docker 批量进行持久性模型存储
// Reuse HTTP connections
const agent = new https.Agent({ keepAlive: true });
设定为
// Bad: mixed language pairs in one request
translate([
{ text: "Hello", sourceLang: "en", targetLang: "de" },
{ text: "Bonjour", sourceLang: "fr", targetLang: "de" }
]);
// Good: group by language pair
translateBatch(enToDe, "en", "de");
translateBatch(frToDe, "fr", "de");
使用率示例
translation_requests_total{lang_pair="en->de",status="success"} 1523
translation_requests_total{lang_pair="en->de",status="error"} 7
translation_duration_seconds{lang_pair="en->de",quantile="0.5"} 0.342
translation_duration_seconds{lang_pair="en->de",quantile="0.95"} 1.234
translation_queue_depth 23
translation_cache_size 6
translation_cache_hits_total 8234
translation_cache_misses_total 142
# Enable JSON logging
LOG_FORMAT=json REQUEST_LOG=1
# Output example
{
"ts": "2025-01-08T15:30:45+0000",
"level": "INFO",
"name": "app",
"message": "translate_post done items=5 dt=0.342s",
"req_id": "a3d2f5b1-c4e6-4f7a-9d8c-1e2f3a4b5c6d",
"endpoint": "/translate",
"src": "en",
"tgt": "de",
"items": 5,
"duration_ms": 342
}
批量翻译(建议)
仅识别语言 |---------|---------|-----------------| | 终点处理后压 | 建筑和改造所有 docker 图像现在都包含正确的版本和元数据以进行跟踪 。 | 快速构建快速构建所有 4 个变量, 并自动设定日期时间版本 : | **窗口 :**Linux/Mac: | 版本战略每个建筑创造 | 两个标记命名标签 | ) - 总是指出最近的版本标签 | (例如,- 不可改变的快照 | **实例:**OCIO 标签 | **每个图像包括元数据:**版本版本版本版本版本 | **: 建筑时间戳(YYYMMDD.HHMMSS)**建立日期 | : ISO 8601 时间戳Git 承诺
: cpu- full, cpu- min, gpu- full, 或 gpu- min检查标签 :
详细的建筑说明和CI/CD集成,见
MAX_QUEUE_SIZEMAX_INFLIGHT_TRANSLATIONSGPU部署绩效优化GPU 优化检查列表
使用 FP16 精确度
ENABLE_QUEUE=1Tune 批量大小预装热热模型
人均平均平均平均平均平均单位的单身工人数
EASYNMT_BATCH_SIZEMAX_CACHED_MODELSEASYNMT_MODEL_ARGS='{"torch_dtype":"fp16"}'WEB_CONCURRENCY=1增加平行主义MAX_INFLIGHT_TRANSLATIONS=1客户最佳做法批次申请
尊重重试后
PRELOAD_MODELS="en->de,de->en"
按语言对对分组 Helsinki-NLP/opus-mt-{src}-{tgt}监测和可观察性
要跟踪的密钥矩阵
PIVOT_FALLBACK=1(请求/sec)curl http://localhost:8000/lang_pairs队列深度(当前等待数)
缓存误存率
MASK_EMOJI=0错误率MASK_PUNCT=0SYMBOL_MASKING=0(如果适用)
public class MostlyLucidNmtClient
{
private readonly HttpClient _httpClient;
private readonly string _baseUrl;
public MostlyLucidNmtClient(HttpClient httpClient, string baseUrl)
{
_httpClient = httpClient;
_baseUrl = baseUrl;
}
public async Task<TranslationResponse> TranslateAsync(
List<string> texts,
string targetLang,
string sourceLang = "",
int beamSize = 1,
bool performSentenceSplitting = true,
CancellationToken cancellationToken = default)
{
var request = new TranslationRequest
{
Text = texts,
TargetLang = targetLang,
SourceLang = sourceLang,
BeamSize = beamSize,
PerformSentenceSplitting = performSentenceSplitting
};
var response = await _httpClient.PostAsJsonAsync(
$"{_baseUrl}/translate",
request,
cancellationToken);
if (response.StatusCode == System.Net.HttpStatusCode.TooManyRequests)
{
// Read Retry-After header
var retryAfter = response.Headers.RetryAfter?.Delta?.TotalSeconds ?? 30;
var jitter = Random.Shared.Next(0, 5);
await Task.Delay(TimeSpan.FromSeconds(retryAfter + jitter), cancellationToken);
// Retry
return await TranslateAsync(texts, targetLang, sourceLang, beamSize,
performSentenceSplitting, cancellationToken);
}
response.EnsureSuccessStatusCode();
return await response.Content.ReadFromJsonAsync<TranslationResponse>(cancellationToken);
}
}
public class TranslationRequest
{
[JsonPropertyName("text")]
public List<string> Text { get; set; }
[JsonPropertyName("target_lang")]
public string TargetLang { get; set; }
[JsonPropertyName("source_lang")]
public string SourceLang { get; set; }
[JsonPropertyName("beam_size")]
public int BeamSize { get; set; }
[JsonPropertyName("perform_sentence_splitting")]
public bool PerformSentenceSplitting { get; set; }
}
public class TranslationResponse
{
[JsonPropertyName("target_lang")]
public string TargetLang { get; set; }
[JsonPropertyName("source_lang")]
public string SourceLang { get; set; }
[JsonPropertyName("translated")]
public List<string> Translated { get; set; }
[JsonPropertyName("translation_time")]
public double TranslationTime { get; set; }
}
内存使用
services.AddHttpClient<MostlyLucidNmtClient>(client =>
{
client.BaseAddress = new Uri("http://translator:8000");
client.Timeout = TimeSpan.FromMinutes(3);
});
Prometheus 示例示例**如果将普罗米修斯(Prometheus)(不是内置的,而是容易添加的)融合在一起,**结构日志示例
比较:容易NMT相对于最粗略的NMT
输入处理
可观察性
手动,不要用自动操作 来缓存 LRU 缓存
配置配置配置限量 40+ env vars 用于微调APPI 兼容性
# Maximum coverage with auto-fallback (recommended!)
docker run -d -p 8000:8000 \
-v ./model-cache:/models \
-e MODEL_CACHE_DIR=/models \
-e AUTO_MODEL_FALLBACK=1 \
-e MODEL_FALLBACK_ORDER="opus-mt,mbart50,m2m100" \
scottgal/mostlylucid-nmt:cpu-min
# GPU with best quality
docker run -d --gpus all -p 8000:8000 \
-e USE_GPU=true \
-e MODEL_FAMILY=opus-mt \
-e EASYNMT_MODEL_ARGS='{"torch_dtype":"fp16"}' \
scottgal/mostlylucid-nmt:gpu
# Test it
curl -X POST http://localhost:8000/translate \
-H 'Content-Type: application/json' \
-d '{"text": ["Hello world"], "target_lang": "de"}'
GPU 上的 OOM( 内存外)
原因:
批量大小太高 或太多模型缓存。
慢速第一请求[原因:
Translation NMT Neural Machine Translation Python FastAPI Docker CUDA PyTorch Transformers Helsinki-NLP Production Microservices API
© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.