Unit 2 / 11

Machine Translation Fundamentals: NMT, LLM, Engine Selection and Pre-Evaluation

Gains:

  • Ability to distinguish the strengths and weaknesses of NMT and LLM-based translation and choose the right engine according to the type of business
  • Ability to pre-classify the suitability of a text for the machine and pre-evaluate it for realistic time and price
  • Ability to quickly measure the quality of raw machine output in five dimensions and flag risks without confusing smoothness with accuracy

The most powerful accelerator you have as a translator is machine translation (MT) — but using a tool consciously means applying it to the right job, in the right language pair, and with the right expectations. In this unit, you will get to know the two major families of MT (NMT and LLM-based translation), learn which one is better for which job, how to quickly measure the raw output quality of a translation before delivery, and how to choose the engine according to the job. The aim is to move away from the idea of ​​"all the same, paste-translate" and become an expert who can foresee which text is suitable for the machine and which requires human-intensive labor.

Two families: NMT and LLM-based translation

NMT (Neural Machine Translation) are systems that have learned from millions of bilingual sentences and translate the sentence as a whole; Google Translate, DeepL, Microsoft Translator are examples of these. Strengths: very fast, consistent, fluent in short sentences. Weakness: often does not see context beyond the sentence boundary (meaning from the whole text); It may be difficult to decide whether the word "bank" means bank or river bank by looking at the paragraph and it does not fit the list of terms provided.

LLM-based translation (Large Language Model) is the use of instruction-driven models such as ChatGPT, Claude, Gemini for translation. Strength: takes instructions — understands instructions such as “use formal tone,” “follow this list of terms,” “target audience is children”; sees the broader context; captures idiom and tone better. Weakness: can sometimes be "too creative" and add to the source (risk of hallucinations), loses coherence in long texts, and costs/speed varies depending on the job.

Rule of thumb: in straight, repetitive, technical texts NMT is usually fast and adequate; LLM is more flexible with texts that require discipline in tone, context, vocabulary, and instruction. Most professionals today use both together: quick draft from NMT, tone/term correction from LLM.

Tip: The machine quality of a language pair (e.g. Turkish→English) may differ from the opposite direction (English→Turkish). Languages ​​with agglutinative, flexible syntax, such as Turkish, are generally more challenging for MT. Judge for yourself which engine is better in your own pair with a few sample texts.

Machine compatibility of text: ask this first

Not every text is equally suitable for the machine. Before starting the translation, pass the text through a "suitability" filter:

  • Well suited to the machine: technical manual, product description, repetitive interface text, standard correspondence, table/list. The language is plain, the terminology is fixed, creativity is scarce.
  • Medium suitable: news, corporate content, educational material. The machine produces good outlines but requires serious post-editing for tone and flow.
  • Not suitable for machine (high risk): legal contract, medical text, advertisement/slogan, literary text, humor, poetry. Here the machine only gives ideas; The job largely falls to people.

If you make this distinction from the beginning, you will both give the customer the right time/price and protect yourself from blindly trusting the raw output.

Quickly measure raw output quality

When you see an MT output you wonder "is it good or bad?" Answer the question with a quick checklist, not intuition. Look at these five dimensions: accuracy (is the meaning faithful to the source?), completeness (is the sentence omitted or added?), terminology (is it correct and consistent?), fluency (is it natural in the target language?), format (are tags, numbers, formatting preserved?). We will deepen these dimensions in the next quality unit; For now, let it be your pre-delivery reflex.

Caution: The most dangerous errors of MT output are the "fluent but incorrect" ones. The sentence may be in perfect Turkish, but the "15% discount" in the source may have been changed to "50% discount". Don't be put off by fluency; Always look at the number, negative, and proper noun separately.

three mini cases

Case 1 — The right engine cut the time in half. One translator chose to translate a 12,000-word software interface with NMT and a marketing booklet of the same volume with LLM. NMT's consistency at the interface made the work faster; LLM's tonal flexibility in the booklet reduced rewriting burden by 40%. Giving both jobs to the same engine would waste time.

Case 2 — Pre-appraisal saved the price. An office was about to offer a legal text because "it can be made by machines, it will be cheap". In the pre-evaluation, a 500-word sample was translated; The machine returned two legal terms with the wrong system's equivalent. The office priced the work as "heavy post-editing" and gave a realistic deadline; He avoided losing money because of the wrong offer.

Case 3 — Language pair difference. A team thought that an engine that worked great in English→German was of the same quality in Turkish→Arabic. The 300-word test showed that the Arabic output required much more correction. The team allocated different engines and different post-editing budgets depending on the language pair.

Four copyable templates

1) Text suitability evaluation:

Your role: senior MTPE specialist. Rate the following text for suitability for machine translation: literal/technical or creative/contextual? Classify it as (1) very machine-friendly, (2) medium, (3) high risk and write your justification. Estimate expected post-editing load as light/medium/heavy.Text example: [paste 200-300 words]

2) Instructed LLM translation (with terminology and tone control):

Translate this text [source]→[target].Audience: [who]. Tone: [e.g. official]. Field: [e.g. finance].List of terms (be sure to use): [A→a, B→b].Rules: do not add what is not in the source, keep numbers/dates/names, replace each sentence with a sentence. Mark where you are not sure with [?]. Text: [...]

3) Two engine comparison:

Below are two different translations of the same source sentence (A and B). Compare each in terms of correctness, terminology and naturalness, and tell me which one will require less correction, with the justification. Source: [...] | Translation A: [...] | Translation B: [...]

4) Raw output quick inspection:

Below is the source and machine translation. Just mark the following:(1) omitted/added information, (2) incorrect number/date/name, (3) negation error, (4) inconsistent term. Don't write any corrections, just list the risky areas.Source: [...] | Translation: [...]

Weak prompt / Strong prompt

Weak: "Translate this text and it will be good." (For the engine, "good" is undefined; no tone, no term, no mass.)

Güçlü: "Translate this product description from English to Turkish. Area: e-commerce. Audience: young consumer. Tone: friendly but reassuring. 'checkout' → 'checkout step', 'wishlist' → 'wish list'. Change the numbers and brand name. Fluent Turkish without any translation odor."

Difference: strong prompt gives the engine a measurable target; the output directly approximates the post-edited draft.

NMT vs LLM comparison chart

Size

NMT (Google, DeepL)

LLM (ChatGPT, Claude)

speed

very fast

medium

Receive instruction/tone

weak

strong

Comply with the term list

limited

Good (with prompt)

Broad context

weak

good

Risk of hallucinations

low

Medium (control required)

most suitable job

Straight/technical/repetitive

Tone/context/creative

Common mistakes

  • Using the same engine for every job. Not choosing an engine according to the type of work causes loss of time and quality.
  • Giving price/time without pre-evaluation. A 200-300 word test reveals the surprises from the beginning.
  • Mistaking fluency for accuracy. A beautiful sentence does not mean a correct sentence.
  • Ignoring the language pair and direction difference. A good engine in one pair may be weak in another.
  • Not noticing LLM's additions. LLM sometimes adds sentences not in the source "to help"; check this out.

Context window and consistency

When translating long text with LLM, it is necessary to know one technical limit: the context window — the amount of text the model can "see" and retain in mind at a time. If the text is larger than this window, you will have to split it into chunks, and the model may "forget" in the next chunk the term decision or tone you made in the previous chunks. So, instead of translating a long book or document in one piece, reminding each piece of the termbase and style summary maintains consistency. This problem does not occur in short texts; However, in long works such as novels and manuals, it is necessary to additionally check the consistency in passages. Practical solution: to prepare a "project memory" (terms + characters + tone decisions) and add it to the beginning of each piece and scan for consistency between the pieces at the end of the translation. In NMT, this problem is experienced in a different way: Since NMT translates sentence by sentence, paragraph length consistency is already weak, so termbase undertakes the term discipline.

In summary

There are two families of machine translation: NMT, which is fast and consistent, and LLM, which takes instruction and sees context and tone better. The master translator grades the suitability of the text for the machine in advance, selects the engine according to the job, quickly measures the quality of the raw output before delivery, and never confuses fluency with accuracy. This pre-evaluation reflex allows you to both give a realistic time/price and protect yourself from blind trust.

Application task

Translate the same 200-word text with both an NMT tool and an LLM. Compare the two outputs on five dimensions (accuracy, completeness, terminology, fluency, format) and write reasons for which one requires less post-editing for this text. Then flag issue/contingency/term risks on both with the “raw output quick check” template.

checklist

  • [ ] I classified the text's suitability for the machine (low/medium/high risk).
  • [ ] Depending on the job type, I decided whether to use NMT or LLM.
  • [ ] I made a preliminary evaluation with a small sample and determined the duration/price accordingly.
  • [ ] I quickly inspected the raw output in five dimensions.
  • [ ] I checked the number, date, name and negatives separately without being fooled by the fluency.