Skywordtech: Independent Language Technology Reference

A Short Glossary of Language Technology

Tokenization is the splitting of text into units — words, subwords, or characters — that a model can process. A corpus is a body of text or audio used for training or evaluation. An embedding is a numerical vector representing a token or sentence such that similar meanings map to nearby points. A language model assigns probabilities to sequences of tokens; a large language model is one trained at sufficient scale to perform many language tasks without task-specific training.

Machine translation converts text from one language to another. Automatic speech recognition transcribes audio into text; text-to-speech does the reverse. A hallucination, in this context, is a fluent but false statement produced by a generative model. A low-resource language is one with little digitized text available for training, which is why model quality varies so widely across the world's languages.

A Short Glossary of Language Technology