Languages
rigour.langs
Language code handling
This library helps to normalise the ISO 639 codes used to describe languages from two-letter codes to three letters, and vice versa.
import rigour.langs as languagecodes
assert 'eng' == languagecodes.iso_639_alpha3('en')
assert 'eng' == languagecodes.iso_639_alpha3('ENG ')
assert 'en' == languagecodes.iso_639_alpha2('ENG ')
Uses data from: https://iso639-3.sil.org/ See also: https://www.loc.gov/standards/iso639-2/php/code_list.php
LangStr
Bases: str
A string carrying an optional language tag.
Use this to keep track of which language a piece of multilingual
content is written in while still passing it around as a str.
The language tag is part of the value's identity:
- With no tag (
lang is None), aLangStris indistinguishable from its content string: it compares equal to the plainstr, hashes the same, and deduplicates against it in sets and dict keys. - With a tag, a
LangStris a distinct value identified by the pair(content, lang). It is not equal to the bare content string, nor to aLangStrwith a different or missing tag, and it never deduplicates against them.
Ordinary str methods (.upper(), slicing, concatenation, …) return
plain str and drop the tag.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
content
|
str
|
The text. |
required |
lang
|
str | None
|
An ISO 639-3 language code, or |
None
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in rigour/langs/text.py
is_lang_better(candidate, baseline)
Decide if the candidate language code is 'better' than the baseline language code, according to the preferred languages list.
is_lang_better('eng', 'deu') True is_lang_better('fra', 'eng') False
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
candidate
|
str
|
The candidate language code. |
required |
baseline
|
str
|
The baseline language code. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
bool |
bool
|
True if the candidate is better than the baseline. |
Source code in rigour/langs/__init__.py
iso_639_alpha2(code)
Convert a language identifier to an ISO 639 Part 1 code, such as "en"
or "de". For languages which do not have a two-letter identifier, or
invalid language codes, None will be returned.
Source code in rigour/langs/__init__.py
iso_639_alpha3(code)
Convert a given language identifier into an ISO 639 Part 2 code, such
as "eng" or "deu". This will accept language codes in the two- or three-
letter format, some language names, and IETF/BCP 47-style tags with a
script or region subtag (e.g. "zh-Hans", "pt-BR"), which resolve to the
code of their primary subtag. If the given string cannot be converted,
None will be returned.
iso_639_alpha3('en') 'eng'
Source code in rigour/langs/__init__.py
list_to_alpha3(languages, synonyms=True)
Parse all the language codes in a given list into ISO 639 Part 2 codes and optionally expand them with synonyms (i.e. other names for the same language).
Synonym groups mix in ISO 639-2/B and Tesseract-style codes (e.g. ger,
chi) which aid input matching but are not valid ISO 639-3 identifiers;
they are filtered out so the returned set only contains canonical codes.