EPT Explained: How Errors Per Thousand Actually Gets Calculated

In this article

In a localization industry that has historically obsessed over word counts, the conversation is finally shifting toward the only metric that truly impacts the bottom line: quality. Managing a global brand requires more than just knowing how many words were translated. It requires understanding how EPT explains errors per thousand actually gets calculated to provide an objective measurement of accuracy. Errors Per Thousand (EPT) provides this precision. It transforms the subjective act of “checking a translation” into a standardized data point that can be tracked, compared, and optimized across millions of words.

Key takeaways

  • EPT is a standardized accuracy metric that calculates the number of weighted errors found per 1,000 words reviewed during linguistic quality assurance.
  • The formula excludes unmodified ICE matches, ensuring that the final score reflects only the quality of the new translation rather than the health of the existing memory.
  • Weighted severities (Minor, Major, Critical) prevent a high volume of small stylistic preferences from masking a single, business-critical error.
  • Benchmarking EPT alongside TTE allows localization managers to see the full picture of their Human-AI symbiosis, measuring both precision and efficiency.

What EPT measures and why word count alone doesn’t cut it

Word count has long been the primary lever for localization budgeting, but it is a poor indicator of linguistic success. A project can be 100% complete and on budget while still failing to meet the quality standards required for a successful market entry. This is because word count only measures volume, ignoring the “friction” caused by mistranslations, grammatical errors, or terminological inconsistencies. EPT solves this by providing a snapshot of accuracy that is independent of project size. This allows a 500-word product description to be compared directly against a 50,000-word technical manual.

By tracking the number of errors found for every 1,000 words reviewed, EPT allows enterprises to move beyond “good enough” translations. It provides a common language for both localization managers and linguists to discuss performance. Instead of relying on vague feedback like “the quality feels off,” teams can identify that a specific language pair has a high EPT due to recurring terminological errors. This level of granularity is essential for maintaining a consistent global voice, especially when using advanced models like Lara that rely on high-quality data to deliver context-aware results.

The formula behind errors per thousand, step by step

Calculating EPT is a straightforward process, but its accuracy depends on isolating the new translation work from existing assets. In an AI-first localization environment managed through a platform like TranslationOS, the system automatically aggregates data from various audits to generate this score. To calculate it manually, you must follow a three-part calculation that balances the number of errors against the actual effort required to translate the content.

The mathematical core of EPT is expressed as: EPT = (Sum of Weighted Errors / Reviewed Word Count) * 1,000

To reach the final number, localization teams follow these specific steps:

  1. Categorize and Weight Errors: Every error found during the review is assigned a severity (Minor, Major, or Critical) and a corresponding point value.
  2. Define the Denominator: The “Reviewed Word Count” is calculated by taking the total word count and subtracting any unmodified In-Context Exact (ICE) matches.
  3. Final Division and Scaling: The total weighted error points are divided by the reviewed word count, and the result is multiplied by 1,000 to produce the final EPT score.

Defining the denominator: Why ICE matches matter

The most common mistake in calculating translation quality is using the total project word count as the denominator. This approach is misleading because it rewards “perfect” translations that were actually retrieved from a translation memory rather than produced by the linguist or Lara. If a project contains 10,000 words but 8,000 of those are ICE matches that required no human intervention, including them would artificially deflate the error rate.

By excluding unmodified ICE matches, EPT focuses the quality assessment strictly on the performance of Lara and the human linguist on the new content. This ensures that the EPT score is a true reflection of the translation engine’s current accuracy, allowing for more effective fine-tuning and resource allocation to improve data quality where it is most impactful.

How different error types get weighted

Not all errors have the same impact on your global brand. A missing comma in a footnote is a minor distraction; a mistranslated safety instruction in a medical manual is a critical failure. EPT accounts for this variance by applying weights to each error based on its severity. This weighting system ensures that a single, devastating mistake carries more weight than dozens of minor stylistic preferences, preventing the “averaging out” of critical risks.

In a standard linguistic quality evaluation, Translated uses the following weights:

  • Minor (1 point): These are issues that do not affect the meaning of the content. They include minor punctuation errors, spelling mistakes that are easily understood by the reader, or small stylistic deviations from the brand guide.
  • Major (5 points): These errors impact the clarity or accuracy of the message. They include mistranslations of key concepts, significant grammatical failures, or the use of incorrect terminology that violates the established glossary.
  • Critical (10 points): These are “show-stopper” errors. A critical error is one that carries legal risk, poses a safety hazard, uses offensive language, or leaves entire sections of content untranslated in high-visibility areas.

Error categories: From mistranslation to style

While severity measures the depth of the problem, error categories help teams identify the source of the problem. By tagging errors as “Terminology,” “Grammar,” or “Accuracy,” localization managers can spot patterns that might indicate a larger systemic issue. For instance, a high concentration of terminology errors in a project translated by Lara might suggest that the model’s underlying glossary needs to be updated. This category-based feedback loop is what allows the human-AI symbiosis to improve over time, steadily driving down the EPT score with each iteration.

What a “good” EPT score looks like by content type

A “good” EPT score is not a fixed number; it is a moving target that depends entirely on the intent and risk profile of the content. High-stakes technical documentation requires a much higher level of precision than a social media post or an internal email. By setting category-specific thresholds, enterprises can optimize their review resources, focusing their human specialists where accuracy is non-negotiable.

Here are the general benchmarks used to define quality at scale:

  • Regulated and Technical Content (EPT < 1.0): In industries like life sciences, legal, or engineering, there is zero tolerance for error. These projects often require an EPT score below 1.0, meaning less than one weighted point of error per thousand words.
  • Marketing and Branded Assets (EPT < 2.0): While stylistic nuance is important, these assets can tolerate a slightly higher EPT as long as the brand voice remains consistent. A score under 2.0 is generally considered a pass for customer-facing marketing content.
  • Internal and High-Volume Informational Content (EPT < 5.0): For content intended for internal understanding or high-volume e-commerce listings, an EPT of up to 5.0 may be acceptable. The goal here is informational parity, and the cost of achieving a lower EPT may outweigh the business value.

Where EPT falls short as a standalone metric

While EPT is an essential tool for measuring accuracy, it is a “defensive” metric. It tells you what went wrong, but it doesn’t necessarily tell you if the translation was good in a creative or strategic sense. A translation can have an EPT of 0.0, meaning it is grammatically and terminologically perfect, while still feeling flat, uninspired, or culturally tone-deaf. It measures the absence of failure rather than the presence of excellence.

Furthermore, EPT does not account for efficiency. To see the full health of a localization program, EPT must be benchmarked alongside Time to Edit (TTE). TTE measures the cognitive effort required to reach that EPT score. If your EPT is low but your TTE is rising, it means your human linguists are working harder to fix a failing Lara model. True optimization in a human-AI symbiosis occurs when EPT and TTE decrease simultaneously, proving that the technology is delivering contextually accurate results that require minimal human intervention.

Ensure your organization’s resources are deployed for the fine-tuning that makes the most impact on ROI. Start the conversation with a proven strategic partner for localization, Translated, today.

Frequently asked questions

Understanding the nuances of EPT is essential for any localization manager seeking to build a data-driven quality framework. Below are the most common questions regarding the calculation and application of this metric in modern translation workflows.

What is the primary difference between EPT and TTE?

Errors Per Thousand (EPT) is an accuracy metric that counts the number of weighted errors in a translation to determine its precision. Time to Edit (TTE) is an efficiency metric that tracks how many seconds a professional translator spends refining a machine-translated segment. Together, they provide a complete picture of quality and productivity.

Why are unmodified ICE matches excluded from the calculation?

Including In-Context Exact (ICE) matches in the denominator would artificially lower the error rate. These segments are retrieved from a memory and were not actually translated by the linguist or Lara in the current project. Excluding them ensures that the EPT score is a direct measurement of the quality of the new work.

How much content needs to be reviewed to get a reliable EPT score?

For a statistically significant EPT score, localization teams typically audit at least 10% of the project volume, or a minimum of 1,000 words. This sample size is usually sufficient to identify recurring linguistic patterns or systemic errors in the translation engine.

Can EPT be used to compare different Lara models?

Yes. You can use EPT to compare different models by running the same source content through them and performing a linguistic audit. This determines which model, such as Lara or a generic neural machine translation (NMT) engine, delivers the highest initial accuracy for your specific domain and language pair.

How does EPT impact the ranking of translators in T-Rank™?

T-Rank™ uses historical EPT data as a key signal for matching projects with the right professional linguist. Translators who consistently deliver low EPT scores for a specific industry or content type are ranked higher. This ensures that high-stakes projects are always handled by the most accurate specialists.

You might be interested in