-
Notifications
You must be signed in to change notification settings - Fork 146
Pull requests: vllm-project/vllm-gaudi
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Enable HPU residual fix for qwen3_next (Qwen3-Coder-Next)
#1676
opened Aug 3, 2026 by
libinta
Collaborator
Loading…
NIXL examples: port disagg accuracy test script fixes to v0.26.0
#1675
opened Jul 31, 2026 by
skaulintel
Contributor
Loading…
Port of: Add pure-Python MiniMax-M3 tool-call parser for HPU- #1655
#1674
opened Jul 31, 2026 by
iboiko-habana
Collaborator
Loading…
test: add security validation unit tests
#1669
opened Jul 30, 2026 by
adobrzyn
Collaborator
Loading…
[DO NOT MERGE] Enable Nemotron-H FP8 on HPU: quant-aware Mamba in_proj + non-gated FP8 MoE
#1667
opened Jul 30, 2026 by
rsmyrek
Contributor
Loading…
Warm up native-resolution vision towers at explicit resolutions and item counts
#1661
opened Jul 28, 2026 by
libinta
Collaborator
Loading…
Fix #1612: don't wrap INT8 W8A8 loader with FP8-only gaudi_weight_wrapper
#1659
opened Jul 28, 2026 by
rsmyrek
Contributor
Loading…
Parallelize some full_tests e2e groups per-card
#1658
opened Jul 28, 2026 by
iboiko-habana
Collaborator
Loading…
Improve MiniMax-M3 MoE decode performance after upstream merge
#1640
opened Jul 27, 2026 by
mkrze
Contributor
Loading…
Add pure-Python MiniMax-M3 tool-call parser for HPU
#1639
opened Jul 27, 2026 by
mkrze
Contributor
Loading…
Run Qwen3-30B MoE FP8 load/generate CI tests in parallel
#1633
opened Jul 24, 2026 by
iboiko-habana
Collaborator
Loading…
Fix model banner and runtime logging issues - multi models runs
#1608
opened Jul 8, 2026 by
12010486
Contributor
Loading…
Fix runtime errors on Qwen single process model swapping
#1604
opened Jul 7, 2026 by
12010486
Contributor
Loading…
[ATEA] [google/gemma-4-31B] Gaudi3 /hpu support
#1596
opened Jul 5, 2026 by
sureshnam
Collaborator
Loading…
[kimi2.6] enable model text only initial
#1574
opened Jun 28, 2026 by
sureshnam
Collaborator
Loading…
Pin vLLM core version to VLLM_GAUDI_COMMIT in RHEL UBI Dockerfile
#1571
opened Jun 26, 2026 by
ghandoura
Contributor
Loading…
MXFP4 GPT-OSS SwiGLU-OAI + bias serving
#1567
opened Jun 25, 2026 by
osavchenkox
Contributor
Loading…
2 of 3 tasks
granite-4.0-h-small made compatible with single process models swap
#1565
opened Jun 25, 2026 by
12010486
Contributor
Loading…
[DRAFT][GAUDISW-248216] Flag compact-GDN per-layer states as independent for the Synapse bridge
#1563
opened Jun 25, 2026 by
moshehoori
•
Draft
Port of #1558 Adapt multi_model_api_server to vLLM ServingTokenizatio…
#1559
opened Jun 24, 2026 by
iboiko-habana
Collaborator
Loading…
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.