Popular repositories Loading
-
dspark-mlx
dspark-mlx PublicLossless speculative decoding for Ternary-Bonsai-27B on Apple Silicon: prism-ml's DSpark drafter ported to MLX
Python 1
-
-
mlx-mtp
mlx-mtp PublicMTP speculative decoding for Apple Silicon (MLX) — 1.27x throughput on Qwen3.5-35B-A3B via fused MoE kernels + zero-replay rejection
Python
-
deepseek-v4-flash-2bit-gb10
deepseek-v4-flash-2bit-gb10 PublicServe DeepSeek-V4-Flash 2-bit (299B MoE) on one DGX Spark (GB10/sm_121) — stock vLLM + vq2 plugin, ~22 tok/s
Python 1
-
-
vllm
vllm PublicForked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Python
If the problem persists, check the GitHub status page or contact support.



