Back to Research papers
Research paper index

Language corpora for the Dutch medical domain

B. van Es

arXiv:2604.25374Published April 28, 20260 citations
  • cs.CL
  • cs.AI

Abstract

\textbf{Background:} Dutch medical corpora are scarce, limiting NLP development. \\ \textbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. \\ \textbf{Results:} The resulting corpus comprises $\pm$ 35 billion tokens across the medical domain in about 100 million documents, freely available on Hugging Face. \\ \textbf{Conclusion:} This work establishes the first large-scale Dutch medical language corpus for pre-training and downstream NLP tasks.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.