Building applications with LLMs through composability, in Kotlin
-
Updated
Oct 14, 2024 - Kotlin
Building applications with LLMs through composability, in Kotlin
Implementation of the LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens Paper
This repository is part of a course on Elasticsearch in Python. It includes notebooks that demonstrate its usage, along with a YouTube series to guide you through the material.
NeuroRVQ: Multi-Scale Biosignal Tokenization for Generative Foundation Models
Develop DL models using Pytorch and Hugging Face
A ridiculously fast Python BPE (Byte Pair Encoder) implementation written in Rust
the small distributed language model toolkit; fine-tune state-of-the-art LLMs anywhere, rapidly
This project shows how to derive the total number of training tokens from a large text dataset from 🤗 datasets with Apache Beam and Dataflow.
Python script for manipulating the existing tokenizer.
Fast, pure-PHP tokenizer library compatible with Hugging Face tokenizers for encoding and decoding text
[Unofficial] Simple .NET wrapper of HuggingFace Tokenizers library
Use custom tokenizers in spacy-transformers
Package to align tokens from different tokenizations.
DeepSeek-OCR-2-Unlimited-OCR is an advanced, experimental visual document processing and open-ended text localization dashboard. This application establishes a unified interface that allows users to swap between two premier vision-language document models: deepseek-ai/DeepSeek-OCR-2 and baidu/Unlimited-OCR.
Use Huggingface Transformer and Tokenizers as Tensorflow Reusable SavedModels
High-performance tokenizers
Python library + CLI for token counting and token distribution analysis in text datasets for LLM data workflows.
Small library that provides functions to tokenize a string into an array of words with or without punctuation
A graphical user interface for the Elasticsearch Analyze API
To associate your repository with the tokenizers topic, visit your repo's landing page and select "manage topics."