Skip to content

How to Choose an Enterprise Translation LSP: A Practical Evaluation Scorecard

Enterprise TranslationYou have three proposals on your desk. All three say they use qualified linguists. All three mention ISO somewhere in the second paragraph. Two of them quote within four cents of each other, and the third is cheaper in a way that makes you slightly uneasy.

This is the piece that is missing: Six key areas to evaluate, a 100-point scorecard that helps you compare providers based on actual evidence rather than marketing language, and a practical pilot approach designed to uncover the issues that a sales demo is unlikely to reveal. What criteria should an enterprise use to evaluate an LSP?

Evaluate a language service provider across six core areas: Quality governance, security and privacy, technology integration, domain and locale expertise, program management, and commercial model. Where your specific content or requirements call for it, feel free to add additional evaluation areas, but do not use fewer than these six. Start with scope fit. A provider that does beautiful work on brochures may not be built for continuous software localization, multilingual SEO, or large-scale machine translation post-editing. Ask how they assign linguists by subject, how they manage terminology, what happens when a reviewer disagrees, and who owns the translation memory when the contract ends.

ISO 17100:2015 is the most useful baseline available, because it sets requirements for translation processes, translator qualifications, revision, and project management. You can read the standard summary at ISO. Certification is not the whole story, but a mature provider should be able to walk you through its workflow clause by clause without preparing for a week.

One distinction saves a lot of pain later: Language coverage and locale capability are different claims. Spanish for Spain, Spanish for Mexico, and Spanish for the United States can need different terminology, tone, currency formats, legal disclaimers, and search behaviour. The technical model is language tags, defined in RFC 5646, where es-MX means something specific and “Spanish” does not.

Ask for evidence in each area rather than a yes. Ask what the escalation path looks like when a medical term is disputed. Ask how they stop a false friend such as Spanish “embarazada”, which the Real Academia EspaƱola defines as pregnant rather than embarrassed, from reaching a published page.

 

What goes wrong when the evaluation gets rushed?

Most rushed vendor selections do not fail because of the translation itself. The problems usually surface later, during engineering, legal review, or market launch, when fixing them becomes both costly and highly visible. Four patterns account for most of the damage.

  • A release going live with broken strings. The provider treated a resource file as plain text, so the variables inside it got translated too. The build compiles. The interface displays “{count} archivos” as literal text. A customer in the target market finds it before QA does.
  • A contract obligation changes meaning. A translator without legal qualifications renders a conditional obligation as an absolute one. Nobody notices until the clause is disputed, and by then the translated version is the one the counterparty relied on.
  • An audit finds an unapproved data path. Files containing employee records went through a public machine translation interface, handled by a subcontractor nobody named in the contract. It surfaces months after the content shipped.
  • Localized pages never rank. The translation is accurate. The metadata was translated literally rather than researched, and the hreflang annotations do not reciprocate. The market underperforms, and everyone blames demand.

None of these are exotic. They are the ordinary result of evaluating on price and a sample paragraph. The scorecard below exists to make each one visible before you sign rather than after.

 

How should translation quality be proven?

Through a documented workflow, qualified linguists, revision by a second linguist, terminology control, and metrics that match the content type. Not through a sample.

A sample sentence tells you almost nothing. Enterprise content contains product names, UI labels, legal disclaimers, SEO metadata, and regulated terminology, often in the same file. What you want to know is who translates, who revises, who approves, and how errors get classified when they are found.

If machine translation is part of the workflow, ask for the post-editing method separately. Machine translation post-editing, or MTPE, means the controlled correction of machine output by a qualified editor, and it has its own standard in ISO 18587:2017. This matters commercially as much as technically. MTPE and unreviewed raw output are frequently sold at similar rates and are not remotely the same service.

A strong provider builds a Termbase and style guide before production starts, not after the first complaint. For software and product content, terminology decisions should cover forbidden translations, approved abbreviations, UI length limits, and locale-specific capitalization. Take the word “Checkout”. It might be a button, a noun, or an instruction. Send it without context, and you are paying for a guess.

Review design matters too. A single in-country reviewer will catch tone problems, but that reviewer should not quietly rewrite approved terminology. Ask for a closed loop: Reviewer change, provider response, final decision, asset updated. Without it, the same argument returns every sprint.

Ask for this Because Good looks like
ISO 17100:2015 workflow mapping It shows whether translation, revision, and project management are controlled steps or informal habits. Documented roles, revision rules, qualifications, and handoff points.
ISO 18587:2017 MTPE process It prevents raw machine output from being presented as a reviewed translation. Defined rules for when MT is allowed, how post-editors work, and how output is sampled.
An error taxonomy by content type Legal, UI, SEO, and support content carry different levels of risk. Critical errors named explicitly: mistranslated obligations, broken variables, and incorrect locale conventions.
Terminology governance It prevents the same terminology dispute from recurring across markets and releases. Termbase entries with an owner, date, approved equivalent, and usage note.
A real QA report It shows whether quality is measured or simply asserted. An anonymized report with error counts by severity. Not a template.

 

What security and AI governance should you expect?

Documented information security controls, clear privacy obligations, and written AI usage rules. All three before a single source file moves.

Translation workflows routinely carry unreleased products, employee data, contracts, support tickets, and marketing plans. ISO/IEC 27001:2022 is the current management-system standard for information security, and a certified provider should still be able to explain the practical controls: Access management, subcontractor approval, encryption, retention, incident notification.

Privacy affects selection directly. The EU General Data Protection Regulation has applied since 25 May 2018. If your content includes customer names, employee records, patient narratives, or support exports, the provider should be ready to sign data processing terms and to restrict where content is processed.

 

The AI Act timeline moved to July 2026

This one is worth flagging, because the date most compliance calendars still carry is wrong. The EU AI Act, Regulation (EU) 2024/1689, came into force on 1 August 2024, and its high-risk obligations were originally set for 2 August 2026. Six days before that date, Regulation (EU) 2026/1744 came into force and deferred them.

Where things stand now: General-purpose AI obligations have applied since 2 August 2025 and were not deferred. Article 50 transparency duties applied from 2 August 2026 and were not deferred. High-risk obligations for standalone Annex III systems now apply from 2 December 2027, and for AI embedded in regulated products from 2 August 2028.

For most localization programmes, that means more time, not less scrutiny. It does not change a single question you should be asking. And a provider still quoting the old August 2026 deadline, in either direction, is telling you something useful about how closely they follow this.

Three questions, and get the answers into the contract rather than the sales call. Is our content used to train AI systems? Are public machine translation engines prohibited for confidential material? How is AI involvement disclosed inside the workflow? Then make the answers operational: Named subprocessors, approved engines, retention periods, breach notice timing, and client ownership of translation memory, termbases, and bilingual files.

 

Which technology capabilities actually matter?

Translation management system integration, XLIFF 2.1 file handling, terminology exchange, ICU MessageFormat support, and multilingual SEO validation. In that order of how often they break.

Technology fit decides whether translation scales or degrades into copy and paste. For software, the provider should process resource files without breaking variables, tags or placeholders. XLIFF is the standard bilingual format for moving content between systems, and version 2.1 became an OASIS Standard in 2018. Ask them to show you how inline codes survive a round trip, using your file, not a demo file.

Plural logic breaks more often than anything else. ICU MessageFormat is the syntax for messages that change with a number. A string like {count, plural, one {# file} other {# files}} is functional code, not prose. Languages with more plural categories than English, such as Polish, Russian, and Arabic, are where it fails first.

Websites add search requirements. Google Search Central states that hreflang annotations should reference alternate versions and include return links. The classic enterprise defect is a French page pointing to an English alternate while the English page points nowhere back. Small string, large consequence.

Accessibility belongs in scope wherever translated web pages do. WCAG 2.2 became a W3C Recommendation on 5 October 2023. Localized alt text, form labels, language attributes, and error messages deserve the same care as visible copy, and each localized build is a new build rather than a copy of the source.

 

What does good program management actually look like?

A defined governance cadence, a tiered escalation path with named owners and response times, and reporting that shows quality and delivery as data rather than reassurance.

This is the area buyers most often score on impression, and the one that decides whether the programme still works in year two. Three things make it assessable.

 

Cadence

Ask what happens weekly, monthly and quarterly, and who attends. A workable default is a weekly check on live work and blockers, a monthly review of quality metrics, volumes and terminology decisions, and a quarterly business review covering roadmap, forecast, cost trend and anything that went wrong. If they cannot describe it without inventing it on the call, they do not have one.

 

Escalation

Escalation should be tiered, time-bound, and owned by name. Who handles a linguistic query during production, and how fast? Who owns the missed delivery? Who is accountable when a defect reaches a live market? The most useful follow-up question is what triggers escalation automatically rather than on request.

 

Service levels that mean something

Ask for turnaround targets by content type and volume band, query response times during business hours in each region, monthly on-time delivery rates, defect rates by severity, and what happens commercially when a level is missed. A provider reporting on-time delivery but not defect severity is reporting the easy half.

 

One question that reveals the whole area

Ask them to walk you through a real incident from the last twelve months. What broke, who was notified, how long the loop took to close, and what changed afterwards. A mature provider has a specific answer and volunteers the uncomfortable parts. A provider who has never had an incident has either never operated at scale, or will not tell you when it happens to you.

 

How should you assess the commercial model?

Look at total programme cost over twelve months. The per-word rate is the least informative number in a translation proposal.

Ask for pricing that separates human translation with revision, MTPE, review-only work, desktop publishing, localization engineering, and rush handling. These are different services, and a single blended rate hides which one you are actually buying.

Then there are three questions that are easy to overlook at the beginning but can become very important later.

How is translation memory reuse handled, and what discount levels apply? If your content has a lot of repetition, the cost should ideally come down over time. By the second year, a mature programme with strong reuse should be noticeably more efficient. If it is not, it is worth asking whether the translation memory is actually being maintained and applied correctly.

What gets billed beyond the translation itself? Make sure you understand all additional charges, including project management, connector maintenance, minimum fees, file preparation, and any other recurring or project-specific costs.

Finally, who owns the translation memory, termbase, and bilingual files? Get this confirmed in writing. If ownership is unclear, switching providers later can become unnecessarily difficult and expensive, and that uncertainty can influence every renewal discussion that follows.

 

How should you score the finalists?

Use a weighted hundred-point scorecard, and put the weight on risk controls and production fit rather than the lowest rate.

The weighting below favors the things that affect launch risk. Adjust it to your content mix, but keep the scoring evidence-based. A cheaper provider that cannot support XLIFF 2.1, ISO 18587 post-editing, or ISO/IEC 27001 security will generate downstream costs in engineering, legal review, and market remediation that dwarf the rate difference.

Area Points Evidence to Require
Quality governance 25 ISO 17100:2015 workflow mapping, revision process, error taxonomy, and a real sample QA report.
Security and privacy 20 ISO/IEC 27001:2022 status, data processing terms, named subprocessors, retention policy, and MT restrictions.
Technology integration 20 Connector plan, XLIFF 2.1 round-trip test, ICU MessageFormat handling, API readiness, and file integrity results.
Domain and locale expertise 15 Linguist qualification model, terminology examples, reviewer management, and locale-specific samples.
Program management 10 Governance cadence, escalation tiers with response times, SLA metrics with remedies, and a real incident walkthrough.
Commercial model 10 Workflow-separated pricing, reuse discount bands, billable extras disclosed, and written asset ownership.

 

Score each area out of its weight, using the evidence column as the test. Where a provider offers reassurance instead of an artifact, score it zero rather than half. That rule is what keeps the exercise honest.

 

What should the pilot include?

Real content across four file types, and measurement of file integrity and query handling alongside language quality.

Use one high-value web page, one software resource file, one regulated or legal excerpt, and one terminology-heavy product asset, wherever those exist in your programme. Score linguistic quality, file integrity, query handling, reviewer experience, turnaround, and reporting clarity separately, because a provider can be excellent at language and weak at everything that makes language deliverable.

Send the same files to every finalist, with the same brief and the same deadline. Ask each to submit its queries along with its output. The query log is often more revealing than the translation, because it shows whether they noticed the ambiguities you left in deliberately.

 

How does GPI approach this?

GPI by the numbers

Operating since 2001. Over 200 languages. More than 500 enterprise clients, including Fortune 1000 companies. 164,000 completed projects informing the ARTEE 1000 engine. Four ISO certifications with certificates published for download: ISO 17100:2015, ISO 18587:2017, ISO/IEC 27001:2022 and ISO/IEC 27017:2015, the last with all 37 cloud controls implemented. Fourteen native CMS and DXP connectors plus a Translation Services API, free to configure.

It would be a strange guide that told you to demand evidence and then offered none. Here is GPI against the same six areas. Every certificate mentioned is published, so you can check it before you talk to anyone.

Area What we can show you
Quality governance ISO 17100:2015 certified through ATC Certification Service, with the certificate published on our ISO certifications page. Scope covers vendor management, project manager training, pre-production processes, and project preparation. Also ISO 18587:2017 certified for full post-editing, covering post-editor qualification, MT suitability assessment, and the difference between light and full post-editing.
Security and privacy ISO/IEC 27001:2022 certified for information security and ISO/IEC 27017:2015 for cloud security, with all 37 cloud controls implemented, covering data in transit and storage, separation of each customer environment, and secure disposal on termination. NDAs are signed with clients and with every subcontractor and partner before they touch a file.
Technology integration Fourteen native CMS and DXP connectors including Adobe Experience Manager, Contentful, Contentstack, Amplience, Drupal, Optimizely, Sitecore, Umbraco, and WordPress, plus a Translation Services API for custom stacks. See the connectors library. Translation memory runs on Trados through our translation memory tools. The connectors are free to configure.
Domain and locale expertise Over 200 languages delivered by native, in-country linguists assigned by subject domain rather than pooled by language. Locale capability is treated as distinct from language coverage: es-MX, es-ES, and es-US are staffed, glossed, and reviewed as separate locales, with terminology, tone, and formatting held per locale. Termbases and style guides are built in Trados before production and maintained across releases, and in-country reviewer changes run through a closed loop so approved terminology holds. The ARTEE 1000 engine draws on 164,000 completed projects to match linguists and prior terminology to your content type.
Program management A documented project management methodology covering everything from account orientation to post-project review, run by a named Globalization Services Team, with checklist-based quality audits at each step through the Globalization Project Management Suite.
Commercial model A Quick Quote Calculator gives you a real-time cost and turnaround estimate before you engage. The GPI Translation Portal gives you files, schedules, billing, and analytics on demand, including spend per language and per time frame.

 

GPI works across enterprise translation, website localization, software localization, and multilingual SEO. Whether you end up choosing us or someone else, insist on the same evidence from everyone on your list.

 

Frequently asked questions

What is an enterprise translation LSP?

A language service provider that manages high-volume multilingual content across departments, systems, and markets. The provider should support documented workflows, qualified linguists, terminology assets, translation memory, secure file handling, and integration with your content or software systems.

 

How is an enterprise LSP different from a freelance translator?

A freelance translator provides linguistic production. An enterprise LSP runs the operating model around it: Project management, multi-locale staffing, revision under ISO 17100:2015, localization engineering, desktop publishing, security controls, and reporting. The difference is not the quality of the language. It is whether there is a system holding it up.

 

Should an enterprise LSP use machine translation?

For the right content, yes, but the provider should define where it is allowed and how output is reviewed. ISO 18587:2017 sets requirements for post-editing, which is what separates controlled MTPE from raw output. Get the distinction in writing, because the two are often priced as though they were the same thing.

 

What should be included in an enterprise translation RFP?

Workflow evidence, ISO alignment, security documentation, technology integration detail, language coverage by locale, sample QA reports, pricing rules, ownership of translation assets, and the proposed governance model. Plus a pilot using real files rather than a sample paragraph.

 

How many LSPs should a global enterprise use?

Many organizations do best with one primary provider and carefully managed specialist backups. A single provider improves terminology consistency and reporting. Multiple providers help with niche domains, regulated markets, or surge capacity. Let content risk decide, not a procurement preference for redundancy.

 

What certifications should a translation provider have?

ISO 17100:2015 for translation service requirements, ISO 18587:2017 where post-editing is offered, and ISO/IEC 27001:2022 for information security. Certification proves an audited management system. It does not prove fit for your content, which is what the pilot is for.

 

How long should an LSP evaluation take?

Most run six to twelve weeks: Two to three weeks to define requirements and issue the RFP, two to three for responses and scoring, two to four for a pilot with real files, and a final week for references and contracting. Compressing the pilot is the most common shortcut and the most expensive one.

 

Who should be involved in selecting a translation vendor?

Localization or content operations owns the process. But the decision needs IT security for data handling, legal for regulated content and contract terms, engineering for file formats, procurement for commercial terms, and in-country stakeholders for locale judgment. Bringing security and engineering in after selection is how a signed contract becomes a stalled implementation.

 

Where to start

The best provider is the one that can prove quality, security, technical execution, and governance before production begins. Not the one with the lowest rate, and not the one with the most confident narrative.

If you do one thing this week, take the scorecard above and apply it to whoever you are currently working with. Not to shortlist a replacement, but to find out which of the six areas you have never actually asked about. That answer usually tells you where your next problem is coming from.

 

Three ways we can help


Talk through your scorecard weighting

Bring your content mix and we will help you set the weights for your programme, whether or not GPI is on the list.

Scope a pilot

We will help define the four files and the acceptance criteria before anyone quotes.

Check our certificates

They are published, so you can verify everything above before we speak.