MTP speculative decoding for Apple Silicon (MLX) — 1.27x throughput on Qwen3.5-35B-A3B via fused MoE kernels + zero-replay rejection