大肥鱼:桌面小鲸鱼,显示 LLM 余额,支持 Codex/本地/云端对话、本地委派分流、文件处理与本地生图。不绑定特定模型,可接任意 OpenAI 兼容接口。
-
Updated
Sep 3, 2026 - JavaScript
大肥鱼:桌面小鲸鱼,显示 LLM 余额,支持 Codex/本地/云端对话、本地委派分流、文件处理与本地生图。不绑定特定模型,可接任意 OpenAI 兼容接口。
A free VRAM calculator for AI models. Not just for LLMs: text generation, embeddings, vision, multimodal, image diffusion, video, audio, and tabular workloads, across inference, LoRA/QLoRA fine-tuning, and full training. Every calculation runs in the browser. Demo app to simultaneously build frontend harness.
Resident local AI inference daemon & multimodal studio in Rust. Subagent model orchestration (DMT), 3D mesh, image, speech & LLMs on consumer GPUs.
MCP server for Wan2GP video generation with GPU detection and VRAM management
Connects remote Ollama servers to local clients over LAN with auto-discovery and VRAM management.
Local AI platform: WebGPU/WGSL browser inference engine + HuggingFace Transformers + Ollama. TurboQuant KV cache compression, GPTQ INT4 fused dequant, mixed-precision BF16/INT4 for hybrid SSM+attention models. 9B parameters in a browser, 8GB VRAM.
Discrete-event simulation of multi-LoRA adapter serving strategies on a single GPU — comparing naive swap, hot-set preloading, and batch-by-adapter under variable VRAM pressure and arrival rates.
[LEGACY PoC] A sovereign, local-first AI reasoning runtime and VRAM orchestration engine built for constrained consumer hardware.
⚡ Fast, interactive LLM VRAM calculator and real-time cloud GPU price comparison tool built with Astro, React, and Tailwind CSS.
High-throughput Paged KV-Cache & speculative inference engine in Rust 2024 + CUDA 12/13 with lock-free continuous batching, O(1) page rollbacks, and a real-time Next.js 15 VRAM telemetry visualizer.
Warm KV-transfer llama.cpp swaps under a sandboxed TypeScript policy, with a per-GPU VRAM budget and load-driven autoscaling.
To associate your repository with the vram-management topic, visit your repo's landing page and select "manage topics."