Overview
Fusion Runtime is a self-hosted voice agent runtime that runs speech-to-text, a language model, and text-to-speech together in one process. Its streaming pipeline can begin playing a reply while the model is still generating it, and it supports interruptions during speech.
Key Features
- Runs the voice pipeline locally on one machine.
- Supports configurable silence waiting and interruption timing.
- Accepts catalog IDs, local files, Hugging Face models, and OpenAI-compatible endpoints.
- Checks model settings at startup and reports likely corrections for misspelled options.
- Provides a browser client for speaking with an agent.
How to Use
Fusion Runtime requires Python 3.11 through 3.13. An agent is defined in a Python file with its prompt, speech-to-text model, language model, text-to-speech model, and turn settings.
The command-line tools can pull the models named by an agent, start the runtime, open a terminal conversation, and diagnose libraries, GPU support, models, and audio. A default agent is also available when starting the runtime without an agent file.