Skip to content

Navigation Menu

Sign in
Sign up
#

huggingface-tokenizers

Here are 13 public repositories matching this topic...

HTGM.2 is a Hindi-first BPE tokenizer trained on ~41GB corpus using streaming architecture for scalable Hindi LLMs, Devanagari NLP, and low-memory tokenizer engineering.

  • Updated May 17, 2026
  • Python

Automated glossary generation and QA assistant that extracts technical terms from text corpora using regex and Levenshtein clustering, tokenizes them with custom BPE, and generates definitions and examples using a local Ollama LLM, all accessible through a CLI interface.

  • Updated Feb 20, 2026
  • Python

Natural Language Processing Practice — a hands‐on repository spanning the full spectrum of NLP, from classical algorithms to cutting‐edge large language models (LLMs). Built around the Hugging Face LLM Course, it’s enriched with practical notebooks on foundational libraries and advanced fine‐tuning workflows.

  • Updated Jun 24, 2026
  • Jupyter Notebook

Build, analyze, and visualize a custom Byte Pair Encoding (BPE) tokenizer trained on the WikiText-2 dataset with vocabulary analysis, compression evaluation, and an interactive GitHub Pages demo.

  • Updated Jun 21, 2026
  • HTML

Add this topic to your repo

To associate your repository with the huggingface-tokenizers topic, visit your repo's landing page and select "manage topics."

Learn more

AltStyle によって変換されたページ (->オリジナル) /