Skip to content

Question: Prefill on NPU and decode on CPU #4776

Description

@DominikHil

Dear developers,

on your latest release you report significant speed up on prefill times using NPU (~9x) while at the same time decode seems to be much faster on CPU (~1.5x).

Is there a reason why inference is not split such that prefill is done by NPU and decode by CPU, essentially combining the best of both worlds?

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions