The first Task-Aware MCP server and automated VRAM calculator for LLM fine-tuning. Instantly snipe the cheapest, fastest GPUs across 10+ cloud providers.
-
Updated
Aug 4, 2026 - Python
The first Task-Aware MCP server and automated VRAM calculator for LLM fine-tuning. Instantly snipe the cheapest, fastest GPUs across 10+ cloud providers.
Will this LLM fit your GPU or Mac? npx fitllm — accurate memory math for MLA/sliding-window/hybrid/MoE architectures that naive VRAM calculators get wrong by up to 18x. Single file, zero deps, conformance-vector tested. MIT.
A free VRAM calculator for AI models. Not just for LLMs: text generation, embeddings, vision, multimodal, image diffusion, video, audio, and tabular workloads, across inference, LoRA/QLoRA fine-tuning, and full training. Every calculation runs in the browser. Demo app to simultaneously build frontend harness.
🚀 A lightweight CLI to estimate hardware requirements and quantization compatibility for Hugging Face models.
A lightweight CLI tool to inspect ML checkpoints (.safetensors, .gguf, .pt) and calculate inference VRAM, multi-GPU memory splits, and vLLM serving capacity.
Anthropic-standard Skill — decide API-vs-self-host LLM costs and fine-tune ROI from any agent context (Claude Code, Cursor, Codex). Live GPU+API prices, deterministic local math.
macOS menu bar tool to explore Hugging Face models, detect GGUF/Safetensors configs, and calculate precise VRAM footprint and KV cache overhead.
A CLI tool for estimating GPU VRAM requirements for Hugging Face models, supporting various data types, parallelization strategies, and fine-tuning scenarios like LoRA.
CLI toolkit for LLM inference preflight, vLLM serving configuration, benchmarking, telemetry, and capacity analysis.
AI runtime loader and model-management control center for local LLM deployments.
⚡ Accurate LLM VRAM Calculator (Weights + Dynamic KV-Cache + Activations) and Inference Speedtest (TTFT & TPS) for vLLM, llama.cpp, and Ollama.
Will a Hugging Face LLM fit your GPU or DGX? Plan weights, KV cache or recurrent state, context, TP/DP, and measured calibration
A simple CLI tool to fetch Hugging Face model metadata and estimate required VRAM/RAM for inference.
Free MCP server for engineers running AI in production: price a model, size a GPU, audit your MCP config, explain network config drift. No account, no telemetry.
🖥️ Check if any LLM fits your GPU. VRAM calculator with context-aware fitting, quantization support, and real-time performance estimates.
Local LLM VRAM Requirements, Concurrency Sizing & DeepSeek R1 Benchmarks
⚡ Fast, interactive LLM VRAM calculator and real-time cloud GPU price comparison tool built with Astro, React, and Tailwind CSS.
Optimal GPU, VRAM, and RAM configurations for running DeepSeek R1 locally (7B to 671B models).
Can I run this AI model locally? Sourced memory requirements for 150+ local AI models across Mac, PC, GPU and phones.
To associate your repository with the vram-calculator topic, visit your repo's landing page and select "manage topics."