Skip to content
#

olmocr

Here are 5 public repositories matching this topic...

Higher performance OpenAI LLM service than vLLM serve: A pure C++ high-performance OpenAI LLM service implemented with GPRS+TensorRT-LLM+Tokenizers.cpp, supporting chat and function call, AI agents, distributed multi-GPU inference, multimodal capabilities, and a Gradio chat interface.

  • Updated Dec 8, 2025
  • Python

Add this topic to your repo

To associate your repository with the olmocr topic, visit your repo's landing page and select "manage topics."

Learn more