Scoring Modes
audio_samples_qoe.visqol() supports two modes, selected with the mode
keyword argument (default "audio").
Audio Mode
score = visqol(ref, deg) # default
score = visqol(ref, deg, mode="audio") # explicit
Property |
Value |
|---|---|
Operating rate |
48 kHz (both signals resampled) |
Gammatone bands |
32, up to Nyquist |
MOS mapping |
Support vector regression (libSVM) |
Patch selection |
All analysis windows |
Identical score |
~4.73 |
Use audio mode for music, broadcast content, or any full-bandwidth signal. The SVR model is calibrated on codec and bandwidth-limiting degradations at 48 kHz.
Speech Mode
score = visqol(ref, deg, mode="speech")
Property |
Value |
|---|---|
Operating rate |
Reference’s native rate (degraded resampled) |
Gammatone bands |
21, capped at 8 kHz |
MOS mapping |
Exponential NSIM → MOS fit |
Patch selection |
Voice-activity gated (silent frames excluded) |
Identical score |
5.0 (exact) |
Use speech mode for narrowband or wideband telephony, VoIP, or voice codec evaluation. Upstream recommends 16 kHz input. The exponential fit maps a perfect NSIM of 1.0 to exactly 5.0, so identical signals always score 5.0.
Choosing a Mode
Scenario |
Recommended mode |
|---|---|
Music quality (streaming, codec, mastering) |
|
Speech codec / VoIP / ASR pre-processing |
|
Mixed content |
|
Narrowband telephony (≤ 8 kHz) |
|
When in doubt, use "audio". The SVR model generalises better across content
types; speech mode’s voice-activity gating can produce surprising results on
non-speech signals with long silences.
Conformance
Both modes are conformance-tested against Google’s C++ reference. Maximum deviation across the test corpus:
Mode |
Max Δ MOS |
Notes |
|---|---|---|
Audio |
0.025 |
FMA instruction scheduling |
Speech (short clips) |
0.08 |
Short-duration fit sensitivity |