Local Voice

Local, validated CosyVoice3 speech generation for reusable MPM briefing and character profiles.

Version: 0.1.0
Author: Bryan Gilliom / Message Point Media
Repository: https://git.mpm.to/mpm/local-voice
CoWork Project: CW-026 — Local Voice

Overview

Local Voice is a shared audio dependency for Claude CoWork, Codex, and other local automation. It converts a canonical script or structured two-role dialogue into WAV or MP3, retaining raw generations and Whisper alignment data so failed seams can be repaired without regenerating successful speech. Voice reference audio is installed separately from the public plugin and remains in the private CW-026 recovery package.

The verified production configuration is Apple Silicon Metal with the CosyVoice3 Candle model. A provisional NVIDIA/CUDA path is documented for the planned Windows host but is not yet certified.

Skill

Skill What it does
local-voice Plans, renders, resumes, aligns, validates, and assembles local speech when a user asks for a voice brief, local TTS message, spoken notification, character dialogue, CosyVoice render, or MP3 briefing.

Commands

Command Purpose
list List bundled voice profiles and their roles.
doctor Verify Python, audio tools, model files, and authorized assets.
plan Convert simple Markdown or Ryan-labelled dialogue into a render plan.
validate Validate a hand-authored render plan without generating audio.
render Generate, align, process, assemble, and quality-check audio.

render --resume reuses completed raw generations. render --assemble-only repairs trimming or seams without rerunning the model.

Supported profiles

  • Donna
  • Chris Engineer
  • Grandpa Bomber
  • Ryan Pitch Meeting: Producer and Writer roles
  • Val Holiday

See skills/local-voice/references/voice-catalog.md for the production rules that distinguish these profiles.

Setup

  1. Follow docs/INSTALL_MACOS.md on the verified Apple Silicon path.
  2. Install the private local-voice-authorized-assets archive from CW-026.
  3. Download the exact Candle model into the runtime models directory.
  4. Run scripts/verify_install.py.
  5. Install the Claude or Codex plugin package.
  6. Start a new task so the host discovers the skill.

The private recovery package and public plugin are intentionally separate: the public repository contains redistributable software, while the private Drive folder contains the authorized voice configuration.

Calling Local Voice

python3 scripts/local_voice.py plan \
  --voice donna \
  --script /absolute/path/brief.md \
  --output-plan /absolute/path/brief-plan.json \
  --output-audio /absolute/path/brief.mp3

python3 scripts/local_voice.py render \
  /absolute/path/brief-plan.json \
  --resume

Set LOCAL_VOICE_RUNTIME when the runtime is not installed at the platform-default location.

Requirements

  • Python 3.11 recommended
  • CosyVoice3 0.1.0 Candle build
  • CosyVoice3-0.5B-Candle model
  • ffmpeg
  • Whisper CLI with word timestamps
  • NumPy and SoundFile
  • Authorized reference audio and transcripts

Operational rules

  • Keep canonical text separate from pronunciation-safe tts_text.
  • Generate the largest safe complete passages; paragraph breaks are candidates, not mandatory cuts.
  • Do not normalize tempo on short greetings, closings, or reactions.
  • Preserve raw WAV, Whisper JSON, processed WAV, and QA reports.
  • Use fixed assets for exact Ryan signatures and very short reactions.
  • Do not silently fall back to a cloud provider.

Troubleshooting

Symptom Resolution
A word is clipped Retain more natural tail or add neighboring sacrificial context, then rerun with --assemble-only when possible.
A seam clicks Move the join to a quiet boundary or add a natural pause; do not regenerate speech that already passed.
A short reaction sounds generic Use an authorized fixed asset from the Ryan signature library.
Opening or closing sounds drunk or rushed Remove tempo adjustment; short frames stay native.
Literal alignment fails Inspect retained Whisper JSON; normalized fuzzy alignment is expected to tolerate ordinary ASR variation.
lead uses the wrong pronunciation Keep canonical text and use a local tts_text override such as led for ammunition.

Security and rights

Local Voice does not require a service credential. Reference audio must only be installed or used when the operator has authorization. The public repository must never contain private voice assets, source recordings, generated briefings, or model binaries.

S
Description
Local CosyVoice production renderer and reusable Claude CoWork and Codex voice-generation plugins.
Readme 61 KiB
Languages
Python 96.3%
Shell 3.7%