Damn,27B模型在MacBook上跑到70 tok/s,比Fable和Sol还快, 还完全不降质,感觉本地AI的可用

Damn,27B模型在MacBook上跑到70 tok/s,比Fable和Sol还快, 还完全不降质,感觉本地AI的可用性拐点,可能就是今天了🤔

DFlash 2让Qwen3.8-27B在一台M5 Max MacBook Pro上跑到了70 tok/s,最高达到自回归解码的4.6倍,而且输出完全一致,lossless,一个token都不差,

70

Zhijian Liu @zhijianliu_ DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro.

⚡ Up to 4.6× the speed of autoregressive decoding, with the same output.

This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free!

原链接