GitHub - SamarthUrs18/fusion-runtime: Self-hosted voice agent runtime: speech-to-text, an LLM and text-to-speech streaming into each other in one process.

Fusion Runtime is a self-hosted voice agent runtime that streams speech recognition, language model output, and speech synthesis on one machine.

Visit Website
GitHub - SamarthUrs18/fusion-runtime: Self-hosted voice agent runtime: speech-to-text, an LLM and text-to-speech streaming into each other in one process.

Introduction

Overview

Fusion Runtime is a self-hosted voice agent runtime that runs speech-to-text, a language model, and text-to-speech together in one process. Its streaming pipeline can begin playing a reply while the model is still generating it, and it supports interruptions during speech.

Key Features

  • Runs the voice pipeline locally on one machine.
  • Supports configurable silence waiting and interruption timing.
  • Accepts catalog IDs, local files, Hugging Face models, and OpenAI-compatible endpoints.
  • Checks model settings at startup and reports likely corrections for misspelled options.
  • Provides a browser client for speaking with an agent.

How to Use

Fusion Runtime requires Python 3.11 through 3.13. An agent is defined in a Python file with its prompt, speech-to-text model, language model, text-to-speech model, and turn settings.

The command-line tools can pull the models named by an agent, start the runtime, open a terminal conversation, and diagnose libraries, GPU support, models, and audio. A default agent is also available when starting the runtime without an agent file.