A noun. A verb. An adjective. A pronoun.
Identifying these word types is something we do almost without thinking. But when language is being analysed at scale, getting this classification right becomes much more important.
For clinical language research, automated parts of speech tagging can help researchers measure how language changes across conditions. But language is not one size fits all. Grammatical structures, vocabulary and patterns differ considerably between languages and language varieties. So before these methods can be applied confidently across multilingual populations, we need to understand how well they perform across languages.
Led by Redenlab’s clinical linguistics lead, Dr Loretta Gasparini, and an international consortium of 105 linguists across 42 countries, the research team benchmarked widely used natural language processing tools across 52 language varieties. Expert linguistic annotation provided the reference point, allowing the team to compare automated tagging with human analysis.
The results showed that performance varied considerably between languages. 25 language varieties showed strong performance, while 11 produced unclear results and 16 showed relatively poor performance. The findings demonstrate why methods developed or validated in one language cannot simply be assumed to work in another. For clinical language analysis, this is an important distinction.
If speech and language are going to contribute to digital biomarkers and clinical outcomes, researchers need methods that can produce meaningful measurements across different populations.
That starts with understanding the strengths and limitations of the tools being used. Rather than treating multilingual analysis as an afterthought, this research puts language diversity at the centre of the evaluation. It provides a benchmark for identifying where existing tools are suitable, where they need improvement, and where more linguistic resources are needed.
The work also creates a foundation for further research into languages and varieties that remain underrepresented in NLP. For Redenlab, this is part of a broader focus on making speech and language analysis more robust across the diverse populations represented in clinical research.

