KazBERT: A Custom BERT Model for the Kazakh Language
Published in Preprint, 2025
Recommended citation: Gainulla, Y. (2025). "KazBERT: A Custom BERT Model for the Kazakh Language."
KazBERT adapts google-bert/bert-base-uncased for Kazakh through continued Masked Language Model (MLM) training with a custom Kazakh WordPiece tokenizer. The training corpus combines Kazakh Wikipedia and Common Crawl text.
Key Features
- Custom WordPiece tokenizer trained for Kazakh
- Continued MLM adaptation from
google-bert/bert-base-uncased - Trained on Kazakh Wikipedia and Common Crawl text
- Supports downstream NLP tasks for Kazakh
Cited By
KazBERT’s architecture and findings have been utilized in several peer-reviewed studies, including:
- LLM-Assisted Weak Supervision for Low-Resource Kazakh Sequence Labeling: Synthetic Annotation and CRF-Refined NER/POS Models (MDPI Applied Sciences, 2026)
- Hybrid artificial intelligence architectures for automatic text correction in the Kazakh language (Frontiers in Artificial Intelligence, 2025)
- Application of Vector Models in Intelligent Information Retrieval Systems (Academic Scientific Journal of Computer Science, 2025)
Blog