METASTONE reports inference gains from its Meta-Infer engineMachine translation
Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.
METASTONE says Meta-Infer used kernel fixes, communication changes and parallelism tuning to raise DeepSeek-V4.1-Flash input throughput on eight PCIe-only GPUs from a community Day0 baseline of 1,932 tok/s to 13,274 tok/s, a 6.87-fold increase. The work did not change the model weights or structure; the figures come from tests reported in the article.
The complete source text is not yet available.
Read at the original source