Metadata-Version: 2.4
Name: parakeet-cpp-cuda
Version: 0.2.2
Summary: Python bindings for parakeet.cpp — NVIDIA Parakeet ASR on ggml, with hotwords
Author: parakeet.cpp contributors
Keywords: asr,speech-recognition,parakeet,nemo,ggml,hotwords
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: POSIX :: Linux
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Requires-Python: >=3.8
Description-Content-Type: text/markdown
Provides-Extra: numpy
Requires-Dist: numpy>=1.20; extra == "numpy"

# parakeet-cpp

Python bindings for [parakeet.cpp](https://github.com/mudler/parakeet.cpp) —
NVIDIA Parakeet ASR on ggml, with hotwords.

```python
from parakeet import Model

with Model("model.gguf") as m:
    print(m.transcribe("audio.wav"))
    print(m.transcribe("audio.wav", hotwords=["台積電", "緯創"], score=8))
```

## Install

```sh
pip install parakeet-cpp          # CPU
pip install parakeet-cpp-cuda     # CUDA (needs a CUDA runtime)
```

Both provide the `parakeet` module, so install one or the other, not both.
Linux x86_64 wheels; the bundled `libparakeet.so` is what does the work.

To use a library you built yourself — another accelerator, a debug build — set
`PARAKEET_LIBRARY` to its path and the bundled one is ignored.

## Hotwords

Biasing changes which candidate wins at each step; it does not make the decoder
search more. Compiling a phrase list costs microseconds and decoding costs
nothing extra, so compile once and reuse:

```python
with Model("model.gguf") as m:
    hw = m.compile_hotwords(["台積電", "聯發科", "緯創"], score=8)
    print(hw.info)          # {'states': 21, 'arcs': 20, 'phrases': 3, 'build_ms': 0.01}
    for path in clips:
        print(m.transcribe(path, hotwords=hw))
    hw.close()
```

`score` is **not comparable across decoders**: around 1.5 suits a TDT
checkpoint, around 8 a CTC one, whose log-softmaxed probabilities have much
wider gaps. Sweep on your own audio. Too high and a phrase is inserted into
audio that never contained it, or displaces the words after it.

Per-phrase weights, and the full document with timestamps and graph statistics:

```python
hw = m.compile_hotwords(["台積電", "緯創"], weights=[2.0, 1.0], score=8)
doc = m.transcribe("audio.wav", hotwords=hw, detail=True)
doc["text"], doc["words"], doc["context"]
```

A phrase this checkpoint's tokenizer cannot represent raises `ParakeetError`
naming it, rather than being dropped — a silently missing hotword looks like a
decoder problem.

## Audio in memory

```python
import numpy as np, soundfile as sf

pcm, sr = sf.read("audio.wav", dtype="float32")
m.transcribe_pcm(pcm, sr, hotwords=["台積電"], score=8)
```

A float32 numpy array is passed through without copying.

## Streaming

For cache-aware EOU checkpoints:

```python
with m.stream() as s:
    for block in blocks_of_pcm:
        chunk = s.feed(block)
        if chunk.text:
            print(chunk.text, end="", flush=True)
        if chunk.eou:
            print("  <- your turn")
    print(s.finalize().text)
```

`chunk.events` carries each event's time and whether it was a backchannel
(`<EOB>`, not a turn end) rather than an end of utterance.

## CTC log-probabilities

For an external decoder stack — an n-gram LM, a lattice rescorer — that needs
the distribution rather than this library's decode:

```python
logits = m.ctc_logits(pcm, sr)   # numpy [T, vocab+1], already log-softmaxed
```

## Threading

`Model` is safe to call from several threads. ctypes releases the GIL for the
duration of every call, and a transcribe call holds the library for a second or
more, so concurrent decodes really do run in parallel.

## Why ctypes and not pybind11

The package contains **no CPython extension**: nothing is compiled against the
Python ABI, so one wheel covers every Python 3 rather than one per interpreter
version. The measured cost of that choice is 0.08 µs of call overhead against a
transcribe call of roughly 1.6 s — about five millionths of one percent. See
`docs/decisions/0008-python-bindings-over-the-c-api.md` in the repository.
