Can AI really understand the way we speak? It sounds like a simple question. It isn’t.
When a child’s speech patterns hold clues to their diagnosis, the tools that analyse those patterns must work reliably across the extraordinary diversity of human language. Yet until now, there’s been no transparent way to know how well the automated systems used in clinical research actually perform.
A new study, published in the Natural Language Processing Journal, changes that. Led by Redenlab researchers in partnership with the University of Melbourne, Murdoch Children’s Research Institute, and 105 linguists across 42 countries, the work establishes the first large-scale benchmark of automated parts-of-speech tagging across 52 language varieties. Before automated language analysis can be used confidently in clinical research, we need to know how well it actually works. So we tested it.
Over several years, we benchmarked widely used natural language processing tools using expert human annotation as the ground truth. That geographic and linguistic breadth was deliberate: language structure varies profoundly, as do the challenges of automated analysis.
The results were encouraging in some languages, but more than half produced poor or unclear results. That’s not a failure of the project. It’s exactly why the work needed to be done. Speech and language data are increasingly central to clinical diagnosis, but these NLP tools have rarely been rigorously tested across diverse languages. Without that testing, clinicians can’t confidently know whether their findings hold equally across populations.
Rather than simply producing a report, we did something harder. We developed an openly shared benchmarking methodology, released our data and code, and built a framework designed for ongoing improvement.
AI-enabled healthcare will only be as good as the evidence underneath it. We’re currently contributing to a landmark trial: the first randomised, placebo-controlled precision medicine study in 16p11.2 deletion syndrome, with speech articulation as a primary outcome. If we want digital biomarkers to work for patients around the world, we need to know where the technology works, where it doesn’t, and what needs to improve.
That’s the less visible part of scientific progress: asking difficult questions before making big claims. By openly sharing this work, we’re inviting the broader research community to build on it. We’re proud to be doing that work.

