Surogate Speech

Romanian, spoken and understood.

A voice model that reads Romanian aloud, and a recognizer that writes it down, from a file or live as you speak. Both run natively in the Surogate engine, on a CPU or a GPU.

Jackrabbit ASR
Jackrabbit Streaming ASR
Amami TTS

Surogate Speech is our line of small, open speech models for agents: a model family per job, Jackrabbit to listen and Amami to speak, shipped language by language. Romanian is the first language out, because our first agents take calls in it. More languages and more families are on the way.

Weights under CC-BY-NC-4.0 for research and non-commercial use; commercial licences through Invergent.

Text to speech · Amami 110M · Romanian

Amami TTS: a Romanian voice on two CPU cores.

A 110M-parameter Romanian voice that runs in real time on two CPU cores, no GPU. It puts the stress where Romanians put it, reads numbers, dates and amounts in words, and dictates codes, IBANs and phone numbers one character at a time from recorded clips, so a voice agent never reads a code wrong.

What to hear

Female voice

Male voice

2.83% WER

word error rate on a set of ordinary Romanian sentences, transcribed back by the Surogate Romanian speech recognizer

Model card

Speech recognition · Romanian

Jackrabbit ASR writes it down.

A 116M-parameter recognizer that writes cased, punctuated Romanian, from a file or live over HTTP or WebSocket.

Live, as you speakIllustration

Bună ziua, aș dori să programez o întâlnire mâine.

Speech recognition · Jackrabbit 110M · Romanian

Jackrabbit ASR

Writes cased, punctuated Romanian from recordings. A 116M-parameter model that runs comfortably on a CPU, with no Python in the serving path.

5.69% WER
word error rate on the FLEURS Romanian test, CTC decoding with a 4-gram language model, Open ASR Leaderboard runner
Model card
Live speech recognition · Jackrabbit 110M Streaming · Romanian

Jackrabbit Streaming ASR

Shows words as you speak, then writes a final, punctuated sentence when you pause. Live audio over HTTP or WebSocket.

0.72 s to text
from the moment you stop talking to the final sentence, on one RTX 5090, including 640 ms to detect the pause
Model card

Run it

One container, one HTTP API.

From the model cards. Amami runs on CPU only and needs Surogate 1.5.5 or later; drop --gpus all to run Jackrabbit on a CPU too.

Speak with Amamitext to speech
hf download surogate/amami-110m-ro --local-dir amami-110m-ro
SUROGATE_TTS_PYTHON=/path/to/venv/bin/python surogate serve --tts amami-110m-ro --port 8080

curl localhost:8080/v1/audio/speech -H "Content-Type: application/json" \
  -d '{"input": "Bună ziua! Codul de confirmare este 4B7X9.", "voice": "male"}' -o out.wav
Listen with Jackrabbitspeech to text
docker run --gpus all -p 8000:8000 ghcr.io/invergent-ai/surogate:1.5.4 \
  serve --stt surogate/jackrabbit-110m-ro-streaming --host 0.0.0.0 --port 8000

curl http://localhost:8000/v1/audio/transcriptions -F [email protected]