Open-source translation models are reaching languages commercial tools skipped
Community-built datasets are covering languages with tens of millions of speakers and almost no software support.
The gap was never technical capability. It was that nobody had assembled the training data.
Who built it
University groups and volunteer collectives compiled parallel text, often from public broadcasting archives and government documents already published in multiple languages.
Quality is now good enough for everyday correspondence, and education departments are among the first to deploy it at scale.