Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Search
/
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
vllm-project
/
vllm-gguf-plugin
Public
Notifications
You must be signed in to change notification settings
Fork
54
Star
50
Code
Issues
7
Pull requests
26
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Actions
Projects
Security and quality
Insights
Actions: vllm-project/vllm-gguf-plugin
Actions
All workflows
Workflows
Copilot code review
Copilot code review
pre-commit
pre-commit
Release
Release
Show more workflows...
Management
Caches
Deployments
All workflows
All workflows
Actions
Loading...
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading.
Please reload this page
.
will be ignored since log searching is not yet available
Showing runs from all workflows
will be ignored since log searching is not yet available
297 workflow runs
297 workflow runs
Workflow
Filter by Workflow
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching workflows.
Event
Filter by Event
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching events.
Status
Filter by Status
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching statuses.
Branch
Filter by Branch
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching branches.
Actor
Filter by Actor
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
No matching users.
[Model] Add Qwen-VL GGUF support
pre-commit
#359:
Pull request
#140
opened by
tangzzycc
Action required
tangzzycc:feat-qwen-vl-gguf
tangzzycc:feat-qwen-vl-gguf
Action required
View #140
View workflow file
[Model] Support Wan2.2 TI2V-5B GGUF diffusion weights
pre-commit
#358:
Pull request
#138
synchronize by
0z5a
Action required
0z5a:codex/wan22-gguf
0z5a:codex/wan22-gguf
Action required
View #138
View workflow file
[Bugfix] Preserve non-GGUF diffusion loader keyword arguments
pre-commit
#357:
Pull request
#139
synchronize by
0z5a
Action required
0z5a:codex/diffusion-quant-loader-kwargs
0z5a:codex/diffusion-quant-loader-kwargs
Action required
View #139
View workflow file
[Model] Support Wan2.2 TI2V-5B GGUF diffusion weights
pre-commit
#356:
Pull request
#138
synchronize by
0z5a
Action required
0z5a:codex/wan22-gguf
0z5a:codex/wan22-gguf
Action required
View #138
View workflow file
[Bugfix] Preserve non-GGUF diffusion loader keyword arguments
pre-commit
#355:
Pull request
#139
synchronize by
0z5a
Action required
0z5a:codex/diffusion-quant-loader-kwargs
0z5a:codex/diffusion-quant-loader-kwargs
Action required
View #139
View workflow file
fix(moe_vec): chunk launches when tokens*top_k exceeds gridDim.z limit
pre-commit
#354:
Pull request
#125
synchronize by
BruceLoveDecimal
Action required
BruceLoveDecimal:fix/moe_vec_grid_z
BruceLoveDecimal:fix/moe_vec_grid_z
Action required
View #125
View workflow file
fix(moe_vec): chunk launches when tokens*top_k exceeds gridDim.z limit
pre-commit
#353:
Pull request
#125
synchronize by
BruceLoveDecimal
Action required
BruceLoveDecimal:fix/moe_vec_grid_z
BruceLoveDecimal:fix/moe_vec_grid_z
Action required
View #125
View workflow file
[BugFix] Correct the iq3xs_grid table for IQ3_S dequantization (#136)
pre-commit
#352:
Commit
e2b8ad5
pushed by
Isotr0py
44s
main
main
44s
View workflow file
[Model] Support Wan2.2 TI2V-5B GGUF diffusion weights
pre-commit
#351:
Pull request
#138
synchronize by
0z5a
Action required
0z5a:codex/wan22-gguf
0z5a:codex/wan22-gguf
Action required
View #138
View workflow file
[Bugfix] Preserve non-GGUF diffusion loader keyword arguments
pre-commit
#350:
Pull request
#139
opened by
0z5a
Action required
0z5a:codex/diffusion-quant-loader-kwargs
0z5a:codex/diffusion-quant-loader-kwargs
Action required
View #139
View workflow file
[Model] Support Wan2.2 TI2V-5B GGUF diffusion weights
pre-commit
#349:
Pull request
#138
opened by
0z5a
Action required
0z5a:codex/wan22-gguf
0z5a:codex/wan22-gguf
Action required
View #138
View workflow file
[BugFix] Correct the iq3xs_grid table for IQ3_S dequantization
pre-commit
#348:
Pull request
#136
opened by
axiom-of-choice
36s
axiom-of-choice:fix-iq3-xs-grid
axiom-of-choice:fix-iq3-xs-grid
36s
View #136
View workflow file
[BugFix] Dispatch each GGUF MoE tensor to the best kernel its own quant type supports
pre-commit
#347:
Pull request
#135
opened by
KaigeGao1110
Action required
KaigeGao1110:pr4-per-tensor-moe-dispatch
KaigeGao1110:pr4-per-tensor-moe-dispatch
Action required
View #135
View workflow file
[BugFix] Resolve Triton MoE BLOCK_M from the per-type table
pre-commit
#346:
Pull request
#134
opened by
KaigeGao1110
Action required
KaigeGao1110:pr3-triton-moe-block-m
KaigeGao1110:pr3-triton-moe-block-m
Action required
View #134
View workflow file
[BugFix] Treat negative expert ids as empty routing slots in GGUF MoE
pre-commit
#345:
Pull request
#133
opened by
KaigeGao1110
2m 45s
KaigeGao1110:pr2-moe-negative-expert-ids
KaigeGao1110:pr2-moe-negative-expert-ids
2m 45s
View #133
View workflow file
[BugFix] Split per-row MoE launches at CUDA's grid z limit
pre-commit
#344:
Pull request
#132
opened by
KaigeGao1110
3m 23s
KaigeGao1110:pr1-moe-vec-grid-z
KaigeGao1110:pr1-moe-vec-grid-z
3m 23s
View #132
View workflow file
fix(weight_utils): name weight_type companions after the last weight only
pre-commit
#343:
Pull request
#131
opened by
KaigeGao1110
Action required
KaigeGao1110:fix/gguf-weight-type-companion-name
KaigeGao1110:fix/gguf-weight-type-companion-name
Action required
View #131
View workflow file
fix(fused_moe): treat negative expert ids as empty routing slots
pre-commit
#342:
Pull request
#130
opened by
KaigeGao1110
Action required
KaigeGao1110:fix/gguf-moe-negative-expert-ids
KaigeGao1110:fix/gguf-moe-negative-expert-ids
Action required
View #130
View workflow file
fix(gguf): validate every dim/stride crossing into int kernel params (fixes #127)
pre-commit
#341:
Pull request
#128
synchronize by
x14ngch3n
Action required
x14ngch3n:fix/int32-range-kernel-params
x14ngch3n:fix/int32-range-kernel-params
Action required
View #128
View workflow file
fix(gguf): validate every dim/stride crossing into int kernel params (fixes #127)
pre-commit
#340:
Pull request
#128
synchronize by
x14ngch3n
Action required
x14ngch3n:fix/int32-range-kernel-params
x14ngch3n:fix/int32-range-kernel-params
Action required
View #128
View workflow file
fix(gguf): validate every dim/stride crossing into int kernel params (fixes #127)
pre-commit
#339:
Pull request
#128
synchronize by
x14ngch3n
-1s
x14ngch3n:fix/int32-range-kernel-params
x14ngch3n:fix/int32-range-kernel-params
-1s
View #128
View workflow file
fix(gguf): validate every dim/stride crossing into int kernel params (fixes #127)
pre-commit
#338:
Pull request
#128
opened by
x14ngch3n
Action required
x14ngch3n:fix/int32-range-kernel-params
x14ngch3n:fix/int32-range-kernel-params
Action required
View #128
View workflow file
[Models] Support DeepSeek-V4 GGUF weights mapping and architecture fallback
pre-commit
#337:
Pull request
#126
opened by
DjoserKhemSimeu
Action required
DjoserKhemSimeu:feat/deepseek-v4-support
DjoserKhemSimeu:feat/deepseek-v4-support
Action required
View #126
View workflow file
fix(moe_vec): chunk launches when tokens*top_k exceeds gridDim.z limit
pre-commit
#336:
Pull request
#125
reopened by
BruceLoveDecimal
1m 24s
BruceLoveDecimal:fix/moe_vec_grid_z
BruceLoveDecimal:fix/moe_vec_grid_z
1m 24s
View #125
View workflow file
fix(moe_vec): chunk launches when tokens*top_k exceeds gridDim.z limit
pre-commit
#335:
Pull request
#125
opened by
BruceLoveDecimal
2m 18s
BruceLoveDecimal:fix/moe_vec_grid_z
BruceLoveDecimal:fix/moe_vec_grid_z
2m 18s
View #125
View workflow file
Previous
1
2
3
4
5
…
11
12
Next
You can’t perform that action at this time.