KazBERT: A Custom BERT Model for the Kazakh Language

Published in Preprint, 2025

Recommended citation: Gainulla, Y. (2025). "KazBERT: A Custom BERT Model for the Kazakh Language."

KazBERT is a BERT model pretrained from scratch with a custom tokenizer designed for Kazakh. The model is trained using Masked Language Modeling (MLM) on a multilingual text corpus containing Kazakh, Russian, and English texts.

Key Features

  • Custom tokenizer optimized for Kazakh language
  • Trained on diverse Kazakh text corpus
  • Supports downstream NLP tasks for Kazakh

Cited By

KazBERT’s architecture and findings have been utilized in several peer-reviewed studies, including:

  • LLM-Assisted Weak Supervision for Low-Resource Kazakh Sequence Labeling: Synthetic Annotation and CRF-Refined NER/POS Models (MDPI Applied Sciences, 2026)
  • Hybrid artificial intelligence architectures for automatic text correction in the Kazakh language (Frontiers in Artificial Intelligence, 2025)
  • Application of Vector Models in Intelligent Information Retrieval Systems (Academic Scientific Journal of Computer Science, 2025)

Download paper here

View model on Hugging Face