-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathindex.html
More file actions
96 lines (96 loc) · 13.6 KB
/
Copy pathindex.html
File metadata and controls
96 lines (96 loc) · 13.6 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
<!doctype html>
<html lang="en">
<head>
<meta charset="UTF-8"><meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="theme-color" content="#0c0f0c">
<meta name="description" content="InferCap is a CLI toolkit for LLM inference preflight, vLLM serving configuration, deployment verification, benchmarking, telemetry, and capacity analysis on NVIDIA GPUs.">
<title>InferCap — Know your inference limits.</title>
<link rel="preconnect" href="https://fonts.googleapis.com"><link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=DM+Mono:wght@400;500&family=Manrope:wght@400;500;600;700;800&display=swap" rel="stylesheet">
<link rel="stylesheet" href="styles.css">
<link rel="icon" href="data:image/svg+xml;base64,PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciIHZpZXdCb3g9IjAgMCA2NCA2NCI+PHJlY3Qgd2lkdGg9IjY0IiBoZWlnaHQ9IjY0IiByeD0iMTQiIGZpbGw9IiMwYzBmMGMiLz48cmVjdCB4PSIyMCIgeT0iMTMiIHdpZHRoPSIyNCIgaGVpZ2h0PSIzOCIgcng9IjIiIGZpbGw9IiNjMGY4NzkiIHRyYW5zZm9ybT0icm90YXRlKC0xMiAzMiAzMikiLz48L3N2Zz4=" type="image/svg+xml">
</head>
<body>
<a class="skip-link" href="#main-content">Skip to content</a>
<header class="site-header"><a class="brand" href="#" aria-label="InferCap home"><span class="brand-symbol" aria-hidden="true">▰</span>infercap<span class="brand-period">.</span></a><nav aria-label="Main navigation"><a href="#capabilities">Capabilities</a><a href="#workflow">How it works</a><a href="#quickstart">Quick start <span>→</span></a></nav><a class="github-link" href="https://github.com/nguyenvmthien/infercap"><svg viewBox="0 0 24 24" aria-hidden="true"><path fill="currentColor" d="M12 .8a11.3 11.3 0 0 0-3.6 22c.6.1.8-.2.8-.5v-2.1c-3.3.7-4-1.4-4-1.4-.5-1.3-1.3-1.6-1.3-1.6-1.1-.7.1-.7.1-.7 1.2.1 1.8 1.2 1.8 1.2 1 1.7 2.7 1.2 3.4.9.1-.8.4-1.2.7-1.5-2.6-.3-5.4-1.3-5.4-5.6 0-1.2.4-2.2 1.2-3-.1-.3-.5-1.5.1-3.1 0 0 1-.3 3.1 1.2a10.7 10.7 0 0 1 5.7 0c2.2-1.5 3.1-1.2 3.1-1.2.7 1.6.3 2.8.2 3.1.7.8 1.1 1.8 1.1 3 0 4.3-2.8 5.3-5.4 5.6.4.4.8 1.1.8 2.2v3c0 .3.2.6.8.5A11.3 11.3 0 0 0 12 .8Z"/></svg>GitHub <span>↗</span></a></header>
<main id="main-content">
<section class="hero wrap">
<div class="hero-copy"><a class="release" href="https://github.com/nguyenvmthien/infercap/blob/master/pyproject.toml"><span class="status-dot"></span> OPEN SOURCE / v0.2.1 <span>↗</span></a><h1>Big models.<br>Real hardware.<br><span>Know your limits.</span></h1><p class="hero-description">Find the right model for your GPU. Configure vLLM. Measure real capacity.</p><div class="hero-actions"><a class="button primary" href="#quickstart">Get started <span>↗</span></a><a class="button secondary" href="#capabilities">Explore capabilities <span>→</span></a></div><div class="install-command"><span>$</span><code>uv sync</code><button class="copy-button" data-copy="uv sync" aria-label="Copy installation command">Copy</button></div><p class="command-context">Inside the cloned repo. <a href="#quickstart">See setup instructions →</a></p><p class="hero-note">Python 3.12+ <span>·</span> NVIDIA GPUs <span>·</span> Built for vLLM</p></div>
<div class="hardware-scene" role="img" aria-label="Illustration of a GPU connected to model feasibility, runtime configuration, and capacity measurement">
<div class="scene-grid"></div><div class="orbital orbit-one"></div><div class="orbital orbit-two"></div><div class="scene-coordinate top-coordinate">SYS.01 / INFERENCE ENGINE</div><div class="scene-coordinate bottom-coordinate">HARDWARE → INTELLIGENCE</div>
<div class="signal-line signal-one"></div><div class="signal-line signal-two"></div>
<div class="chip-stack"><div class="chip-under"></div><div class="chip-board"><div class="chip-pins"></div><div class="chip-core"><span class="chip-icon">▰</span><strong>infercap</strong><span class="chip-caption">KNOW YOUR CAPACITY</span></div><span class="board-detail">IC—002 / GPU</span></div></div>
<div class="floating-card hardware-card"><span class="mini-icon">▦</span><div><small>HARDWARE AWARE</small><strong>NVIDIA GPU <span class="status-dot"></span></strong></div></div>
<div class="floating-card fit-card"><span class="check-icon">✓</span><div><small>MODEL FEASIBILITY</small><strong>Find your fit.</strong></div><span class="tiny-bars"><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="floating-card capacity-card"><div><small>CAPACITY, MEASURED</small><strong>Beyond the guesswork.</strong></div><svg viewBox="0 0 140 42" aria-hidden="true"><path d="M1 39 22 34 43 23 64 18 86 7 108 5 139 4"/></svg></div>
</div>
</section>
<section class="ecosystem wrap" aria-label="Supported ecosystem"><p>BUILT ON THE STACK<br><span>YOU ALREADY USE.</span></p><div class="ecosystem-name nvidia">▧ <strong>NVIDIA</strong></div><div class="ecosystem-name vllm">v<span>LLM</span></div><div class="ecosystem-name hugging">🤗 <strong>Hugging Face</strong></div><div class="ecosystem-name python"><span>⌘</span> Python</div><div class="ecosystem-name uv"><span>▥</span> uv</div></section>
<section class="capabilities wrap section" id="capabilities"><div class="section-heading"><div><p class="eyebrow">01 / FROM POSSIBILITY TO PERFORMANCE</p><h2>Less trial and error.<br>Better-informed inference decisions.</h2></div><p>One toolkit to connect your hardware,<br>your model, and your next deployment.</p></div><div class="feature-grid">
<article class="feature-card">
<div class="card-top"><span class="feature-icon" aria-hidden="true">⌕</span><span>01 / RECOMMEND</span></div>
<h3>Find models that fit your GPU.</h3>
<p>Search a model family. Get recommendations matched to your environment.</p>
<div class="feature-example">
<p class="example-label">EXAMPLE · QWEN / 24 GiB FREE / FP16</p>
<div class="candidate-preview"><span>Qwen2.5-7B-Instruct<small>~15.0 GiB estimated weights</small></span><b>FIT</b></div>
<div class="candidate-preview"><span>Qwen2.5-14B-Instruct<small>~30.0 GiB estimated weights</small></span><b class="over-budget">NO FIT</b></div>
<p class="example-caption">Weight estimates only · illustrative.</p>
</div>
<a class="feature-action" href="#workflow" data-workflow="recommend">See model discovery in action <span>→</span></a>
</article>
<article class="feature-card">
<div class="card-top"><span class="feature-icon" aria-hidden="true">⌘</span><span>02 / CONFIGURE</span></div>
<h3>Generate your serving command.</h3>
<p>Pick a profile. Get a vLLM command ready to review and run.</p>
<div class="feature-example">
<p class="example-label">EXAMPLE · BALANCED PROFILE</p>
<pre class="serving-preview"><span>vllm serve</span> Qwen/Qwen2.5-7B-Instruct \
--tensor-parallel-size 1 \
--dtype float16 \
--gpu-memory-utilization 0.9 \
--max-model-len 4096 \
--generation-config vllm \
--kv-cache-dtype auto \
--max-num-seqs 64 \
--max-num-batched-tokens 8192 \
--enable-prefix-caching \
--enable-chunked-prefill</pre>
<p class="example-caption">Example: 1 GPU · balanced profile. Review before running.</p>
</div>
<a class="feature-action" href="#workflow" data-workflow="serve">See the configuration workflow <span>→</span></a>
</article>
<article class="feature-card">
<div class="card-top"><span class="feature-icon" aria-hidden="true">↗</span><span>03 / BENCHMARK</span></div>
<h3>Find your sustainable load.</h3>
<p>Sweep the load. See where throughput stops scaling.</p>
<div class="feature-example">
<p class="example-label">ILLUSTRATIVE CONCURRENCY SWEEP</p>
<div class="load-preview"><div><span>Workers</span><strong>8</strong><strong>16</strong><strong>32</strong></div><div><span>Output TPS</span><b>420</b><b>760</b><b>790</b></div></div>
<p class="saturation-preview">16 workers <span>Last healthy level</span></p>
<p class="example-caption">+3.9% throughput at 32 workers · sample data.</p>
</div>
<a class="feature-action" href="#workflow" data-workflow="benchmark">See the benchmark command <span>→</span></a>
</article>
</div><p class="estimate-note">Estimates exclude KV cache and runtime memory. Verify by running your model.</p></section>
<section class="workflow wrap section" id="workflow"><div class="workflow-copy"><p class="eyebrow">02 / YOUR TERMINAL. YOUR WORKFLOW.</p><h2>From “will it run?”<br>to “how far can it go?”</h2><p>Choose a step. Copy the command. Run it locally.</p><div class="steps" role="tablist" aria-orientation="vertical" aria-label="Workflow commands"><button id="tab-recommend" class="step active" role="tab" aria-selected="true" aria-controls="command-panel" data-step="recommend"><span>01</span><div><strong>Find your model</strong><small>A family name in. A ranked shortlist out.</small></div><b>↗</b></button><button id="tab-check" class="step" role="tab" tabindex="-1" aria-selected="false" aria-controls="command-panel" data-step="check"><span>02</span><div><strong>Check your model</strong><small>Inspect hardware and estimate feasibility.</small></div><b>↗</b></button><button id="tab-serve" class="step" role="tab" aria-selected="false" aria-controls="command-panel" tabindex="-1" data-step="serve"><span>03</span><div><strong>Configure & serve</strong><small>Use the recommended vLLM command.</small></div><b>↗</b></button><button id="tab-benchmark" class="step" role="tab" aria-selected="false" aria-controls="command-panel" tabindex="-1" data-step="benchmark"><span>04</span><div><strong>Find your operating point</strong><small>Run a sweep and explore the results.</small></div><b>↗</b></button></div></div><div class="terminal" id="command-panel" role="tabpanel" aria-labelledby="tab-recommend" tabindex="0"><div class="terminal-bar"><span class="window-dots"><i></i><i></i><i></i></span><span>infercap — terminal</span><button class="copy-button" id="copy-workflow" aria-label="Copy workflow command">Copy command</button></div><div class="terminal-content"><p class="terminal-comment" id="terminal-comment"># Start with the hardware you have.</p><pre id="terminal-command"></pre><div class="benchmark-modes" id="benchmark-modes" role="group" aria-label="Benchmark mode"><span>MODE</span><button type="button" data-mode="concurrency" class="active">concurrency</button><button type="button" data-mode="burst">burst</button><button type="button" data-mode="request-rate">request-rate</button></div><div class="benchmark-answer" id="benchmark-answer" aria-live="polite"></div><div class="recommend-demo" id="recommend-demo">
<div class="shortlist">
<div class="shortlist-header"><span>MODEL SHORTLIST</span><span class="shortlist-example">SAMPLE OUTPUT</span></div>
<div class="shortlist-context"><span>Qwen family</span><span>24 GiB free</span><span>FP16</span></div>
<article class="shortlist-model shortlist-pick">
<div class="shortlist-rank">01</div><div class="shortlist-info"><span class="shortlist-kicker">LIKELY FIT</span><strong>Qwen2.5-7B<span>-Instruct</span></strong><div class="shortlist-memory"><i style="width:62.5%"></i></div><small>~15.0 GiB weights <span>/ 24 GiB free</span></small></div><span class="shortlist-status" aria-label="Estimated fit">✓</span>
</article>
<article class="shortlist-model">
<div class="shortlist-rank">02</div><div class="shortlist-info"><strong>Qwen2.5-14B<span>-Instruct</span></strong><small>~30.0 GiB weights</small></div><span class="shortlist-no-fit">NO FIT</span>
</article>
<div class="shortlist-bottom"><span>Hub discovery</span><span>vLLM compatibility</span><span>Memory estimates</span></div>
</div>
<p class="demo-disclaimer">Illustrative shortlist, not live results. Run locally for your hardware. Estimates exclude KV cache and runtime memory.</p>
</div><div class="terminal-output" id="terminal-output"></div><div class="terminal-prompt"><span>❯</span><i></i></div></div><div class="terminal-footer"><span class="status-dot"></span><span id="terminal-caption">Illustrative output · results depend on your environment</span><span>bash</span></div></div></section>
<section class="quickstart wrap" id="quickstart"><div><p class="eyebrow">03 / MAKE EVERY GPU COUNT</p><h2>Your next inference run<br>starts with a little clarity.</h2><p>Open source. Hardware aware. Built for developers.</p><a class="button primary" href="https://github.com/nguyenvmthien/infercap">Explore the repository <span>↗</span></a><a class="text-link" href="https://github.com/nguyenvmthien/infercap/blob/master/README.md">Read the docs →</a></div><div class="setup"><div><span>01 / CLONE & INSTALL</span><button class="copy-button" data-copy="git clone https://github.com/nguyenvmthien/infercap.git cd infercap uv sync" aria-label="Copy setup commands">Copy setup</button></div><pre><span>$</span> git clone https://github.com/nguyenvmthien/infercap.git
<span>$</span> cd infercap
<span>$</span> uv sync</pre><p>Requires Python 3.12+ and a compatible vLLM / PyTorch platform.</p><a class="setup-next" href="#workflow">02 / Find a model for your GPU <span>→</span></a></div></section>
</main>
<footer class="wrap"><a class="brand" href="#"><span class="brand-symbol" aria-hidden="true">▰</span>infercap<span class="brand-period">.</span></a><p>Know your hardware. Own your inference.</p><div><a href="https://github.com/nguyenvmthien/infercap/blob/master/LICENSE">Apache 2.0</a><a href="https://github.com/nguyenvmthien/infercap/blob/master/CONTRIBUTING.md">Contribute ↗</a><a href="https://github.com/nguyenvmthien/infercap">GitHub ↗</a></div></footer>
<div id="toast" role="status" aria-live="polite"></div><script src="app.js"></script>
</body></html>