[Tool] A Batch Inference Workflow for GPT-SoVITS

Lately I’ve been working on a mod for Magical Girl Witch Trial. Besides writing the script and doing the actual modding, most of the work is voicing it. The voicing is done with GPT-SoVITS — train on the original voice lines, then run inference — but the WebUI that ships with GPT-SoVITS can only do one line at a time, and the batch inference script in the repo is half-finished. I first threw together a script that just ran a for-loop, and quickly found out that the hard part of batch inference isn’t the loop at all — it’s how you pick a reference clip for each line. Get that wrong and the results are pretty surreal.

So I forked upstream GPT-SoVITS and built a whole workflow around batch inference + automatic pre-selection + human review. My first version worked by pre-selecting on textual emotion, which — as it turns out — was optimizing for the wrong thing. That’s what the first section below is about. The second version instead matches on delivery.

First: what is a reference clip actually a reference for?

GPT-SoVITS needs a reference clip for every line it infers. In the first version, my approach was as direct as it gets: happy lines get a happy reference, sad lines get a sad one. So I ran an emotion classifier over both the reference pool and the script, tagged everything with an “emotion label”, and pre-selected by similarity. It works often enough, but the results were always slightly off.

Later I had Claude dig into the question of what a reference clip is actually a reference for. What GPT-SoVITS really “copies” from the reference is its delivery — intonation, rhythm, stress, pauses, emotional energy, even the recording environment. The AR model takes the reference’s semantic tokens as a prefix and “continues writing” from them (much like VALL-E’s in-context learning).

In other words, an “emotion label” and “how the line is actually spoken” are two completely different things. The same 「私、大丈夫です」 can be cold, tearful, or a forced smile. A text emotion model cannot tell those apart — and yet that distinction is exactly what matters most in the output.

The new approach: match on “delivery”

Matching now has two sides:

  • The reference pool: every reference clip gets an acoustic analysis — a dimensional emotion model (SER) gives valence / arousal / dominance, then librosa extracts a set of prosodic features (pitch contour, energy, speaking rate, pause count). These become 7 axes — Warmth, Excitement, Assertiveness, Speed, Loudness, Pitch range, Pauses — roughly in the range −3…+3, where 0 is that character’s normal state. The axis scores are normalized within each character’s own reference pool (so “fast, loud, exaggerated” is always relative to that character’s usual voice). As in the first version, results are of course cached — this step is far too slow to redo.
  • The lines: for each line, an LLM reads its surrounding dialogue (plus an optional “delivery note”) and outputs values on the same 7 axes, describing how the line should be spoken.

Then it’s a weighted distance over those same axes, and the closest clips win. The weights are just the absolute values of the numbers themselves — the more extreme an axis, the more it counts in the match; a 0 means “this axis doesn’t matter” and drops out of the selection entirely.

Setting up

You need two things to get going: a speaker config and a script.

Speaker config inputs/speaker_config.yaml tells the tool which characters exist, where each one’s GPT/SoVITS weights live, and where its reference pool is (a standard .list — your training set works as-is):

speakers:
  ema:
    gpt_path: GPT_weights_v2ProPlus/ema-e15.ckpt
    sovits_path: SoVITS_weights_v2ProPlus/ema_e8_s2232.pth
    ref_list: inputs/ema.list
    ref_audio_dir: inputs/
  meruru:
    gpt_path: GPT_weights_v2ProPlus/meruru-e15.ckpt
    sovits_path: SoVITS_weights_v2ProPlus/meruru_e8_s480.pth
    ref_list: inputs/meruru.list
    ref_audio_dir: inputs/

The script, inputs/scenes/*.scene, is a new input format designed to organize varied input better. Plain text, one line per line of dialogue; the file name becomes the scene name (and the output sub-folder):

# scene01                  (scene name + output sub-folder)
lang: ja                   (default language for the scene)

ema: 心配要らないよ。今回はボクも一緒だから。
  note: 表面は優しく安心させる口調、でも奥に冷たい支配感

meruru: 私も大好きですよ、エマさん。
  note: 素直で温かい、信頼しきった声

narrator: ふたりは指を絡め合った。   (narration — not in the config, so context only)
  • name: text, where name has to match a key in the YAML for the line to be voiced.
  • An indented note: gives a “delivery instruction” (Chinese, Japanese, English — the LLM understands all of them), ref: pins a specific reference clip, and a per-line lang: overrides the scene language.
  • Any speaker not in the config (e.g. narrator:) is only fed to the LLM as context, never voiced — narration and throwaway characters you don’t need voiced can all go here, and they help the LLM understand the scene, which improves its scoring.

There’s also tools/script_to_scene.py for converting existing scripts into .scene files (see its header).

Opening the WebUI

Double-click go-batch-webui.bat in the repo root, or run it manually:

runtime\python.exe -I webui_batch_inference.py zh_CN   :: Chinese UI
runtime\python.exe -I webui_batch_inference.py en_US   :: English UI

Your browser opens http://localhost:9870 automatically. The UI is still three tabs — Config, Review & adjust and Help — covering the two big steps: analyze → review → generate.

Step 1: Analyze and pick references

In the Config tab, fill in a few fields: Scene files to read (the path to your scripts — there’s a preview next to it showing which files matched), Speaker config (auto-scanned from inputs/*.yaml), Output folder, how many candidates to keep per line, and an AI delivery analysis section — which LLM to use (anthropic / openai / a local openai_compatible / or none to skip AI entirely), how many lines of context, and a “Variation (temperature)”.

Once that’s done, click Step 1 · Analyze & pick references. This step doesn’t synthesize any audio. Instead it:

  1. Runs the acoustic analysis over every character’s reference pool;
  2. Has the LLM read the context and your notes line by line to compute the 7 delivery values;
  3. Picks the Top-K candidates by weighted distance within each character’s own pool;
  4. Writes those results and candidates into a <scene>.profiles.json sidecar next to each scene.

The log streams into the Live log box in real time.

Step 2: Review and adjust

Switch to Review & adjust and pick a scene you just analyzed. Three lines per page, each showing:

  • The line itself;
  • The delivery note you wrote;
  • 7 sliders (Warmth / Excitement / Assertiveness / Speed / Loudness / Pitch range / Pauses; 0 = doesn’t matter);
  • A few inline audio players — the candidate clips that got picked — plus a radio button to choose which one to use.
  • There’s also a Lock toggle to freeze your edits to that line.

Click Save changes when you’re done to write it back. If you changed a note and want the AI to re-score, hit Re-analyze this scene (locked lines are preserved). This step takes the longest, but it’s where the quality comes from — in the vast majority of cases one of the candidates is usable, and when none of them is, it usually means the reference pool simply doesn’t have it.

Step 3: Generate audio

Once you’re done reviewing, click Generate audio for this scene. The real run prefers the clip you picked. Then you can hand the machine over to the GPU for the night, and the voiced files will be sitting in the output directory in the morning.

About the LLM and that note

The “analyze” step calls an LLM. Drop a .env in the repo root with your key:

OPENAI_API_KEY=sk-...        (AI service = openai)
ANTHROPIC_API_KEY=sk-...     (AI service = anthropic, the default)

You can also point Server URL at a local Ollama / LM Studio / vLLM, or just pick none to skip the AI entirely and set the sliders by hand.

The note is optional but genuinely useful: with no note, the AI judges purely from context; with a note, the note wins — which puts you in the voice director’s chair.

Finally, the repo: https://github.com/AsaChiri/GPT-SoVITS. Feedback welcome — if you run into anything, just open an issue on GitHub.