Public preview · Try Sentiment, Image Captioning and Housing. Explore Forge & Voice and WiseShield through their project stories.
IAIbrahim Aldulaimi

Audio production

Forge & Voice

A personal audiobook project inspired by my love of books and listening while on the move.

Live app available in a local walkthrough. This public preview contains the project story.

Project
Local audio studio · Python, FastAPI, Fish Speech
My contribution
Script review, voice assignment, audio checks, export recovery, and an MCP interface.
Result & scope
Single-voice and multi-speaker production. Reference-based cloning; no newly trained speech model.

Built from a love of audiobooks

I love books, but I’m often walking or busy doing something else. Audiobooks let me keep following a story when I can’t sit down and read. That is the personal reason behind Forge & Voice: I wanted to explore turning written books into audio I could enjoy while going about my day.

That idea led me to explore a local audiobook pipeline and then build a studio for reviewing text, choosing voices, and producing recordings. Giving characters distinct voices became part of making the listening experience more expressive. The current studio works with shorter passages and scripts; extending it to long-form chapters is still a next step.

The first recording was only the beginning

Generating speech exposed the work around the model. A saved voice sample could be mistaken for a reading of new text. An unavailable speech service could leave someone waiting without knowing why. In a cast recording, one loud reference could also make a character dominate the conversation.

Those problems shaped the interface: reference previews and generated recordings now have separate controls, service readiness is visible, and recording levels are checked and balanced. Users can audition a voice before committing to a longer production.

Keeping the writer in control

I built the script workflow around a review step. Imported text becomes structured JSON, the user checks the spoken words, and each speaker gets a chosen voice. The runner preserves the source through segmentation and stops an incomplete set of audio segments from becoming a finished export. A saved production record keeps the cast, script, and output together after a restart.

What the studio became

The current app supports both a quick single-voice passage and a reviewed multi-speaker production. A local MCP interface exposes the same workflow to an assistant. In the latest end-to-end check, a two-speaker script produced a downloadable recording.

The project taught me to treat previewing, review, failure handling, and recovery as part of audio production. It uses reference-based voice cloning, not newly trained speech weights. Long-form chapter editing and verification of the actual spoken words are the next steps.

From a document to a cast recording

The studio imports TXT, Markdown, DOCX paragraph text, and its own JSON format. It recognizes explicit speaker labels, preserves spoken words, and asks the user to review the structured script. Each speaker gets a chosen saved voice before generation. The production manifest records the cast and each audio segment, making the result easier to inspect and reproduce.

A local Model Context Protocol server exposes the same script conversion, validation, cast generation, voice listing, and job-status actions to an assistant. It has been tested over stdio and registered in Codex. A key lesson was separating three different operations: preparing a voice reference, assigning that reference to a character, and training new model weights. This version implements the first two and provides a training-readiness report; it does not claim a new fine-tuned model.

Make a voice, then audition it

Add a clear 10–30 second recording and its exact transcript. Preview the upload, confirm permission, and prepare a local voice profile. The quality panel reports duration, recording level, and possible clipping. Generate a short test with new words and compare it with the source recording.

Existing profiles can be revised as separate copies: correct the transcript or adjust recording level, then audition both versions. This uses Fish reference-based cloning; it does not fine-tune model weights. Signal checks catch some recording problems but do not establish speaker similarity or pronunciation quality.

How it works

What it does

A local text-to-speech studio with two paths: a quick single-voice recording and a reviewed multi-speaker script. Both expose the structured text before generation. The cast workflow assigns saved or uploaded references to each speaker, offers sample playback, and exports the finished recording with production details.

Tools and architecture

The browser uses HTML, CSS, and JavaScript. A FastAPI service accepts generation requests and queues work through a single-worker Python executor. Fish Speech S2 Pro runs locally for synthesis. Optional Ollama models suggest delivery metadata; FFmpeg normalizes and encodes the audio.

How a recording is made

The runner splits the source into segments and verifies that joining them reproduces the original text. Each segment is sent to Fish with the selected reference audio and its transcript. Manual bracket cues take priority over optional model-generated delivery suggestions. Generated segments are converted to mono 24 kHz PCM, checked, joined, and exported as AAC in an M4A container.

Voice uploads and storage

An upload supplies a name, a 3–30-second recording, its exact spoken transcript, and a speaker-permission confirmation. The service limits file size, decodes the audio, checks duration and signal level, and saves a local profile under a generated identifier. These references are outside the public website. Completed jobs have a manifest containing the source hash, segment text, voice, and signal-check results.

Engineering decisions and verification

Signal checks reject empty, very short, quiet, or excessively silent output. An incomplete job cannot silently become a final mix. Completed manifests allow the recording library to recover after a server restart. These checks measure audio properties; they do not prove correct pronunciation or word-for-word speech accuracy.

Scope and next steps

The original project explored EPUB parsing and character workflows. The current studio supports single-voice passages up to 6,000 characters and reviewed multi-speaker scripts with up to 12 speakers, 60 segments, and 12,000 spoken characters. Long-form chapter editing and transcription-based validation remain future work. No speech foundation model was trained for this app.