arXiv Releases Wieszcz-XIX: A 3.1 Billion-Word Historical Polish Corpus and Temporally Bounded Language Models
Summary:arXiv has released Wieszcz-XIX, a historical Polish corpus consisting of 6.75 billion tokens (approximately 3.1 billion words) across 294,369 documents published between 1800 and 1918. The corpus, assembled from sources like Wolne Lektury and the Internet Archive, is over three orders of magnitude larger than existing annotated corpora for the same period. The research team trained a series of decoder-only models (ranging from 47M to 349M parameters) on this corpus and demonstrated their tempora
Key Breakthroughs
-
Large-Scale Historical Corpus Construction: Wieszcz-XIX contains approximately 6.75 billion tokens (3.1 billion words) from 1800 to 1918, making it thousands of times larger than existing annotated corpora for the same period. This provides a rich resource for studying historical Polish.
-
Data Processing and Quality Control: The corpus was built through rigorous filtering, deduplication, and post-1918 leakage detection, ensuring data quality. Despite some recognition errors and uncorrectable text, the overall character error rate is controlled at 0.68%.
-
Temporally Bounded Language Model Training: The research team trained a series of decoder-only models (ranging from 47M to 349M parameters) on this corpus and demonstrated their temporal boundedness. These models outperform modern Polish models in handling historical text, showing lower vocabulary costs and higher spelling consistency.
Technical Highlights
- Corpus Scale and Quality: Wieszcz-XIX's scale far exceeds existing corpora, and its multi-layered quality control ensures data reliability.
- Temporal Boundedness Verification: The models perform well in handling historical text, retaining historical spelling habits, while modern models fail to do so.
- Model Performance Improvement: Increasing parameter size yields about twice the performance gain as a second pass over the data, showcasing the advantages of large-scale data training.
Industry Impact and Developer Recommendations
- Historical Language Research: The corpus provides a valuable resource for studying historical Polish, filling gaps in existing data.
- Temporally Bounded Model Applications: For developers needing to process historical or time-specific data, the model offers a new technical path.
- Ethical and Bias Issues: The model replicates historical biases, including antisemitic statements, and developers should handle related ethical issues with caution.
Future Outlook
The release of Wieszcz-XIX not only provides a new tool for Polish historical language research but also opens new directions for the development of temporally bounded language models. As more similar corpora are released, AI's ability to handle historical text and time-specific data will be further enhanced.
— END —Source: ArXiv NLP/LLM (cs.CL) (2026-10-09)
Tags: #Historical Corpus #Temporally Bounded Model #Language Model #Polish #Data Quality
Editorial & Fact-Checking Note: This article is compiled from primary research, official release documentation, and source papers by the ZICQ Newsroom pipeline with automated entity verification and human editorial review. If you notice any technical inaccuracy, please submit a correction via our corrections policy or email our editorial desk directly.
Community Comments