<table><tr><th>Metric</th><th>Result</th></tr><tr><td>Dataset</td><td>184 long-form Chinese audio files, 11,539 s total, 192.3 min.</td></tr><tr><td>GPU</td><td>NVIDIA H100 80GB HBM3.</td></tr><tr><td>Best GPU speed</td><td>SenseVoice-Small: 169.6x realtime in the full benchmark, 211.8x in the initial run.</td></tr><tr><td>Best CPU speed</td><td>SenseVoice-Small: 17.2x realtime; Paraformer-Large: 15.6x realtime.</td></tr><tr><td>Baseline</td><td>OpenAI Whisper-large-v3: 13.4x realtime on GPU.</td></tr></table>
0 commit comments