Damn,27B模型在MacBook上跑到70 tok/s,比Fable和Sol还快, 还完全不降质,感觉本地AI的可用
Damn,27B模型在MacBook上跑到70 tok/s,比Fable和Sol还快, 还完全不降质,感觉本地AI的可用性拐点,可能就是今天了🤔
DFlash 2让Qwen3.8-27B在一台M5 Max MacBook Pro上跑到了70 tok/s,最高达到自回归解码的4.6倍,而且输出完全一致,lossless,一个token都不差,
70
Zhijian Liu @zhijianliu_ DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro.
⚡ Up to 4.6× the speed of autoregressive decoding, with the same output.
This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free!
![]()