Ready wake-word models: alexa.onnx, hey_mycroft.onnx, hey_mycroft_synthetic.onnx,
wake_up.onnx
Copyright TigreGotico. Distributed with this package under the Apache License,
Version 2.0.

Each model is a recurrent classifier head on features from the WakeHuBERT tiny
featurizer bundled in ../featurizers/wakehubert-tiny (TigreGotico/wakehubert-tiny,
Apache-2.0). Its output is Platt-calibrated: sigmoid(output) is the probability
that the 1.5 s window holds the wake word.

hey_mycroft.onnx was trained on human recordings. alexa.onnx, wake_up.onnx and
hey_mycroft_synthetic.onnx were trained on synthetic speech only. Training data:

  TigreGotico/synthetic-wakeword-alexa        CC BY 4.0
  TigreGotico/synthetic-wakeword-hey_mycroft  CC BY 4.0
  TigreGotico/synthetic-wakeword-wake_up      CC BY 4.0
      https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-alexa
      https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_mycroft
      https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-wake_up
      Wake-word positives made with text-to-speech voices and voice conversion.

  edge-tts voices
      Additional hey_mycroft_synthetic.onnx positives: the wake word spoken by
      edge-tts English voices at several rates and pitches.

  TigreGotico/not-wake-words-speech-en        CC BY 4.0
      https://huggingface.co/datasets/TigreGotico/not-wake-words-speech-en
      Speech that is not the wake word (negatives).

  AudioSet-derived clips                      CC BY 4.0
      https://huggingface.co/datasets/agkphysics/AudioSet
      Non-speech negatives and background noise. The AudioSet labels are
      CC BY 4.0; the audio comes from YouTube videos under their uploaders'
      terms.
