原文整理页

DFlash 2 正式发布,在 M5 Max 上运行 Qwen3.8-27B 速度达 70 tok/s,推理效率提升高达 4.6 倍

来源作者:Zhijian Liu (@zhijianliu_)原始来源:https://x.com/steipete/status/2089957375499464910

中文导读

DFlash 2 正式发布,在 M5 Max 上运行 Qwen3.8-27B 速度达 70 tok/s,推理效率提升高达 4.6 倍。

正文 Markdown

DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro. ⚡ Up to 4.6× the speed of autoregressive decoding, with the same output. This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free! https://t.co/We0lwYPSBl