aifollow.news Search
Back 量子位 AI 报道
量子位 AI 报道· · Original publication time

METASTONE reports inference gains from its Meta-Infer engineMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI-assisted summary

METASTONE says Meta-Infer used kernel fixes, communication changes and parallelism tuning to raise DeepSeek-V4.1-Flash input throughput on eight PCIe-only GPUs from a community Day0 baseline of 1,932 tok/s to 13,274 tok/s, a 6.87-fold increase. The work did not change the model weights or structure; the figures come from tests reported in the article.

The complete source text is not yet available.

Read at the original source
Found an error? Send a correction