KazBERT: A Custom BERT Model for the Kazakh Language

Published in Preprint, 2025

Recommended citation: Gainulla, Y. (2025). "KazBERT: A Custom BERT Model for the Kazakh Language."

KazBERT adapts google-bert/bert-base-uncased for Kazakh through continued Masked Language Model (MLM) training with a custom Kazakh WordPiece tokenizer. The training corpus combines Kazakh Wikipedia and Common Crawl text.

Key Features

  • Custom WordPiece tokenizer trained for Kazakh
  • Continued MLM adaptation from google-bert/bert-base-uncased
  • Trained on Kazakh Wikipedia and Common Crawl text
  • Supports downstream NLP tasks for Kazakh

Cited By

KazBERT’s architecture and findings have been utilized in several peer-reviewed studies, including:

  • LLM-Assisted Weak Supervision for Low-Resource Kazakh Sequence Labeling: Synthetic Annotation and CRF-Refined NER/POS Models (MDPI Applied Sciences, 2026)
  • Hybrid artificial intelligence architectures for automatic text correction in the Kazakh language (Frontiers in Artificial Intelligence, 2025)
  • Application of Vector Models in Intelligent Information Retrieval Systems (Academic Scientific Journal of Computer Science, 2025)

Download paper here

View model on Hugging Face