Language families
How to Read an ISO 639 Language Code
639-1, 639-2, 639-3 and macrolanguages: what each code length means and why the same language has several codes.

When parsing language codes, it is critical for a system to work properly. ISO 639 does not assign one unique code for each language. Instead, an ISO 639 language code can be two letters, or three, or may map to a macrolanguage that groups related individual languages into one code. Understanding which code an app or system is using, and correctly interpreting the code's length and form, can dramatically affect how well the technology localizes content and accurately identifies languages.
The ISO 639 standard is published in separate parts, each with a distinct purpose and level of specificity. ISO 639-1 is a shorthand for the most commonly used and easily recognized languages. Its codes are always two letters long, like en for English or es for Spanish. ISO 639-1 primarily covers individual languages, and focuses on languages that are widely used.
Consider ISO 639-1 as the introductory set for ISO 639. It is useful as a quick reference, but not comprehensive. For broader coverage, ISO 639-2 offers three-letter identifiers for larger universe of known languages. It includes many of the same languages as ISO 639-1, but also includes additional languages and some language groups, or families. ISO 639-2 codes are three letters, like zho for Chinese.
And ISO 639-2 is not the end either. ISO 639-3 is the most comprehensive collection: it includes both every language that is part of ISO 639-2, and any language that has been documented and attested at least once, ever. That includes many living languages, and some extinct or ancient ones as well. Like ISO 639-2, ISO 639-3 codes are always three letters.
The layers to the ISO 639 standard can make it hard to understand how to use ISO 639 codes correctly. There is no one-to-one relationship between languages and codes in ISO 639-3 and the codes of ISO 639-1 and ISO 639-2. For many of the languages in ISO 639-3, some of the codes in ISO 639-1 and the corresponding codes in ISO 639-2 will represent macrolanguages that map to more than one language in the ISO 639-3 list.
What is considered a macrolanguage depends on ISO standards, not organizational policy, but there are well known examples. Chinese is the most common one. In ISO 639-3, the code for Chinese is zho, but its code under ISO 639-2 is chi, and ISO 639-1 it's zh, and in the ISO 639-3 code tables, Chinese is explicitly defined as a macrolanguage with several individual member languages, rather than one unified language.
Today, ISO 639 describes each of the three major parts it publishes as ISO 639-1, 639-2, and 639-3. However, those codes are just the ones in force as of 2025, which are an ample, or improved, edition of the standards originally published in 2002, 1998, and 2007, respectively. And the status of those earlier versions also seems to have changed: the old versions ISO 639-1:2002 and ISO 639-2:1998 are now described as "withdrawn" by ISO, and the more recent additions (which presumably include updated codes) are simply latest, not fifth, editions of the standard.
A correct process for parsing an unknown ISO 639 code involves three steps. The first is to confirm with the software or system using the code which language code scheme it relies on - that may depend on the organizational policy and the software’s specific requirements. The second step is determining whether the code is five letters or three, to confirm the level of specificity. And the third is looking up each code to understand whether it represents an individual language, or a macrolanguage with many individual languages below it. Together, those three steps allow accurate understanding of how many languages, or language families, the code indicates.
Chinese language is a critical case for understanding ISO 639 correctly. Despite seeming to be multiple codes with the same meaning, under the current ISO 639 definitions: zh, zho, and chi are not the same, and do not mean the same. zh means "Chinese language" on the ISO 639-1 list, while chi is in ISO 639-2. chi is in ISO 639-2, and corresponds to zh in ISO 639-1, but also includes the languages of Cantonese (zho-cmn), Mandarin (zho-han), Min (zho-min), Wu (zho-wuu), Xiang (zho-xlg), and Hakka (zho-hak). zho is the ISO 639-3 macrolanguage code for Chinese, covering subgroups like Mandarin (cmn), Hakka (hak), and others. The ISO 639-3 code tables call out zho/chi/zh as a macrolanguage, with list of the individually coded member languages.
Where this sits in the Register
Written by the Linguasphere desk. Classification data referred to here comes from the Linguasphere Register — see sources and attribution.