KazBERT: A Custom BERT Model for the Kazakh Language
Published in Preprint, 2025
Recommended citation: Gainulla, Y. (2025). "KazBERT: A Custom BERT Model for the Kazakh Language."
KazBERT is a BERT model pretrained from scratch with a custom tokenizer designed for Kazakh. The model is trained using Masked Language Modeling (MLM) on a multilingual text corpus containing Kazakh, Russian, and English texts.
Key Features
- Custom tokenizer optimized for Kazakh language
- Trained on diverse Kazakh text corpus
- Supports downstream NLP tasks for Kazakh
Cited By
KazBERT’s architecture and findings have been utilized in several peer-reviewed studies, including:
- LLM-Assisted Weak Supervision for Low-Resource Kazakh Sequence Labeling: Synthetic Annotation and CRF-Refined NER/POS Models (MDPI Applied Sciences, 2026)
- Hybrid artificial intelligence architectures for automatic text correction in the Kazakh language (Frontiers in Artificial Intelligence, 2025)
- Application of Vector Models in Intelligent Information Retrieval Systems (Academic Scientific Journal of Computer Science, 2025)
X
Blog