Is AI Translation Good Enough for Your App? An Honest Answer
Yes — conditionally. For high-resource, English-centric pairs, large language models are now competitive with dedicated neural MT on general text under human evaluation, and that output can be a usable first pass on app strings if you supply context, protect placeholders, and run a review. It is not a drop-in replacement for a human localizer on short, ambiguous UI chrome, on low-resource languages, or on legal, medical, or payment copy. Anyone quoting a general MT benchmark as proof that your strings will be fine is overreaching. The honest answer depends on the failure modes of short UI strings, not on a news-domain score.
What the research actually measures
WMT23 and WMT24 evaluate general-domain machine translation: mostly news and web text, not software UI. Systems include traditional neural MT, online MT, and LLM-based systems. Automatic metrics (COMET as primary, with chrF and BLEU) sit alongside human Direct Assessment and MQM.
In WMT23, GPT-4 5-shot ranked at or near the top by human DA for English→German, German→English, and Japanese→English. For English–German, MQM showed the reference near the best, then GPT-4 5-shot ahead of at least one strong online NMT system. The findings note that LLMs exhibit strong performance across the majority of language pairs, based on only two LLM-based submissions. The paper’s own framing — LLMs are here but not quite there yet — is the right summary: state-of-the-art or very close in some high-resource pairs under human evaluation, not uniformly superior, and weakly represented in the task.
WMT24 treated instruction-tuned LLMs as a standard baseline (eight LLMs and four online MT providers). They can be competitive with or better than some commercial MT systems in several high-resource pairs. The advantage is pair- and direction-dependent. Performance drops substantially on low-resource and non-English-centric directions. A TACL analysis on WMT23 data with Llama-2-based models confirms the same pattern: translation quality correlates strongly with resource availability; X→English is relatively stable; low-resource directions degrade. Broader work on Galician shows fine-tuned LLMs outperforming seq2seq on the more distant English↔Galician pair, while bilingual seq2seq remains competitive on closely related Spanish↔Galician.
That is the evidence. It is about news and general prose.
The UI-string benchmark gap
There is no widely recognised, WMT-style, multi-language public benchmark for software UI strings — no shared task that runs MT and LLM systems on button labels, menu items, and format strings with standardised automatic metrics and human evaluation over multiple years. What exists is smaller and narrower: an English→German UI disambiguation test set from a 2012 EAMT paper on context-aware localisation; domain-adaptation work on software localisation corpora; translation memories built from Microsoft Visual Studio UI strings; a 2021 Microsoft Word UI user study; and proprietary or vendor evaluations. That absence is the finding. UI string quality is under-benchmarked compared with news. A COMET score on WMT news does not measure whether “Open” became the right part of speech on a toolbar.
Failure modes specific to UI strings
These are the errors that decide whether machine translation is good enough for an app. They are poorly captured by sentence-level news tests.
Ambiguous short strings without context
UI literals are often a single word. English Open, New, Save, or Clear can be a verb, adjective, or noun depending on the widget. The English→German UI test set was built to stress this: models translating segments in isolation frequently pick the wrong sense. Document-level or widget-level context reduces those errors; it does not eliminate them. If you send a string table with no comments, you are running the exact condition the literature flags as severe.
Placeholder corruption
Localisable files contain variables and markup such as %s and %1$d. MT can move, drop, or rewrite those tokens. Autodesk’s Apertium Spanish–Brazilian Portuguese work treats placeholder integrity as a post-edit and QA requirement, not as something the engine is assumed to preserve. Placeholder mistakes increase edit operations. There is no large public count of how often modern LLM MT corrupts format strings on UI corpora; the failure mode is documented in localisation infrastructure, not scored like WMT.
Text expansion breaking layout
Translated strings need to stay roughly the size of the source. Over-long labels truncate, wrap, or overlap. MT does not optimise for character length or layout. Web-UI internationalisation work treats this as a structural constraint of the medium: strings are short, embedded in code, and the layout was designed around the source. Unconstrained MT will not respect that.
Plural categories
English has singular and plural. Many target languages do not. Slavic and other systems need discrete forms for one, few, many, and related categories. A single English source often must map to multiple target strings by numeric context. Sentence-level MT does not know whether the runtime value is 1, 3, or 21, and generic engines mishandle that mapping. If your resource files do not encode plural categories, the model cannot invent a correct CLDR-style split.
Gender agreement
Many controls refer to an object whose grammatical gender is invisible in English. English→German UI work flags wrong articles and adjective endings as typical. Sentence-level MT also fails on pronouns, gender, and coreference; labels, tooltips, and dialog titles split across segments still have to agree. Without the referent, the model guesses.
Even when the string is “correct enough” to complete a task, unedited MT UI has a UX cost. In a 2021 study, 84 people used Microsoft Word in published human-translated UI, unedited machine-translated UI, or English. Task completion did not differ significantly. Efficiency and satisfaction were significantly higher for the released human translation, especially for less experienced users. Eye-tracking showed higher cognitive load for localised UI versus English, and the highest cognitive cost for unedited MT. Semantic adequacy is not the same as a usable chrome.
What measurably helps
Context and comments. On the English→German UI disambiguation set, context-aware NMT outperforms sentence-level baselines; surrounding strings cut sense errors. Document-level MT research makes the same claim in general: adjacent text improves ambiguity, consistency, and coherence. Instruction-tuned LLMs given document context produce higher-quality translations than independent sentences, even without document-level fine-tuning. For an app, that means translator comments, screen or widget names, neighbouring strings, and the role of the control — not a naked key/value dump. A vendor (non-peer-reviewed) experiment on real UI strings reported further gains from multilingual hints and screenshots; treat that as indicative vendor evidence, not independent science.
Glossaries and domain data. Generic MT underperforms on software UI without domain adaptation. Recurrent-NMT localisation work showed that adaptation on localisation corpora, including UI and terminology glossaries, improves quality versus out-of-domain models. Fine-tuning with UI-derived translation memories (Visual Studio strings, English→Turkish, Spanish→Turkish, English→Catalan) improved localisation-task quality versus generic MT. A product glossary is not decoration; it is the cheapest way to stop the model inventing a second word for “Settings”.
A review pass. In Autodesk software-localisation tests, MT plus post-editing raised average throughput by 74% and saved about 43% of translation time versus from-scratch work, without a significant quality loss in that setting. A later Autodesk test reported productivity gains of around 18% for phrase-based MT and 36% for NMT, with edit distance strongly correlated with the gain. A 2021 professional-translator experiment found post-editing cut time by 63%, keystrokes by 59%, and pauses by 63%, with no significant quality difference. Other NMT studies report on the order of 20% time savings, or 20–60% across the literature when MT is strong. Those gains shrink or reverse when MT is poor, and post-editing skill matters: student PEMT has been measured worse than the same students translating from scratch. Review is not optional QA theatre. It is the step that turns a news-domain-competitive engine into something you can ship on a toolbar.
A decision rule by content type
UI chrome (menus, buttons, toggles, empty states, format strings): highest risk. Short, ambiguous, layout-bound, placeholder-heavy, split across files. Do not ship raw MT. Provide comments and glossary, constrain or check length, lock placeholders, then post-edit. The Word UI study is the relevant user-facing evidence: people can finish tasks and still be slower and less satisfied.
Store listing copy (description, screenshots captions, what’s new): closer to the news and general prose WMT actually measures. High-resource pairs are the setting where LLMs are competitive with dedicated MT under human evaluation. Still run a native review for tone and claims. This is the content type where a strong general engine is least mismatched to the research.
Legal, medical, and payment text (terms, privacy, consent, dosages, prices, refunds, tax): do not use unedited MT as the source of truth. The literature does not give you a UI-specific legal benchmark; it does give you placeholder corruption, agreement errors, and a measured UX penalty even when tasks still complete. Those failure modes are unacceptable when the string is a contract, a health instruction, or a charge. Human translation or specialist post-editing, not a raw model dump.
Strongest and weakest language pairs
Strongest, in the WMT and related evidence: high-resource English-centric pairs — English↔German, English↔French, English↔Spanish, English↔Chinese, English↔Japanese — plus closely related pairs with decent data such as Spanish↔Catalan and well-trained Spanish↔Galician. X→English is generally more stable than the reverse in the Llama-2/WMT23 analysis. English-centric pairs also show higher average COMET and lower variance in a 22-pair LLM failure-mode study.
Weakest: low-resource languages, including many Indic, African, and indigenous languages, where NMT and LLM MT both degrade for lack of parallel and pretraining data. Documented hard LLM directions include Arabic↔Hebrew, English↔Tamil, and Khmer↔English. ChatGPT-style MT has been described as competitive for high-resource languages and not for low-resource ones. Resource availability and English-centricity predict performance. If your next locale is a low-resource or non-English-centric pair, the WMT “LLMs are competitive” headline does not apply; budget for human translation.
Machine translation of UI strings is measurably strong for many high-resource pairs and still fragile next to human translation, especially on short ambiguous strings and low-resource languages. Use the engine where the research matches the text. Do not use a news benchmark as a proxy for a settings screen.
The cost side of the same decision is in the app localization cost guide, and services that sell human translation by the word are compared separately. Games carry far more text than a utility app, which changes the review budget — see what game localization costs.
Frequently asked questions
Can I ship LLM-translated UI strings with no human review?
Not if you care about UX. Unedited MT UI in a Microsoft Word study did not reduce task completion, but it cut efficiency and satisfaction and raised cognitive load versus the released human translation. Review is the quality step, not a nice-to-have.
Does a top WMT score mean my app strings will be fine?
No. WMT23/WMT24 measure news and general prose. There is no widely recognised public benchmark for software UI strings. Quoting a general MT score as proof for buttons and format strings is overreaching.
Are LLMs better than dedicated MT for localisation?
On high-resource general-text pairs, WMT human eval puts some LLMs at or near the top versus specialist NMT. That result is pair-dependent, weakly sampled in WMT23, and not UI-specific. For UI, context, glossary, and post-edit matter more than the engine brand.
What should I put in the string file besides the source text?
Comments on widget role, neighbouring strings or screen name, a glossary of product terms, and intact placeholders. Context-aware UI MT measurably cuts disambiguation errors on cases like Open/New/Save; domain TMs and glossaries beat generic MT on localisation text.
Which languages are safest for AI app translation?
High-resource English-centric pairs such as English↔German, French, Spanish, Chinese, and Japanese, and close pairs like Spanish↔Catalan. Low-resource and non-English-centric directions (e.g. English↔Tamil, Khmer↔English, Arabic↔Hebrew) show clear gaps.
Is post-editing actually faster than translating from scratch?
Yes when MT is strong: Autodesk localisation tests reported large throughput gains (including ~43% time saved in one study; later ~18% PBMT / ~36% NMT); other experiments show ~20–63% time cuts with similar quality. Poor MT can cost more than translating from scratch.
Ship your app in 50 languages by tonight
Upload your localization file, review side by side, download ready-to-import files for every language. One-time credits from $9 — no subscription.
No subscription. Credits never expire.