feat: add CW-026 Local Voice plugin
This commit is contained in:
@@ -0,0 +1,120 @@
|
||||
# Local Voice
|
||||
|
||||
Local, validated CosyVoice3 speech generation for reusable MPM briefing and
|
||||
character profiles.
|
||||
|
||||
**Version:** 0.1.0
|
||||
**Author:** Bryan Gilliom / Message Point Media
|
||||
**Repository:** https://git.mpm.to/mpm/local-voice
|
||||
**CoWork Project:** CW-026 — Local Voice
|
||||
|
||||
## Overview
|
||||
|
||||
Local Voice is a shared audio dependency for Claude CoWork, Codex, and other
|
||||
local automation. It converts a canonical script or structured two-role
|
||||
dialogue into WAV or MP3, retaining raw generations and Whisper alignment data
|
||||
so failed seams can be repaired without regenerating successful speech. Voice
|
||||
reference audio is installed separately from the public plugin and remains in
|
||||
the private CW-026 recovery package.
|
||||
|
||||
The verified production configuration is Apple Silicon Metal with the
|
||||
CosyVoice3 Candle model. A provisional NVIDIA/CUDA path is documented for the
|
||||
planned Windows host but is not yet certified.
|
||||
|
||||
## Skill
|
||||
|
||||
| Skill | What it does |
|
||||
|---|---|
|
||||
| `local-voice` | Plans, renders, resumes, aligns, validates, and assembles local speech when a user asks for a voice brief, local TTS message, spoken notification, character dialogue, CosyVoice render, or MP3 briefing. |
|
||||
|
||||
## Commands
|
||||
|
||||
| Command | Purpose |
|
||||
|---|---|
|
||||
| `list` | List bundled voice profiles and their roles. |
|
||||
| `doctor` | Verify Python, audio tools, model files, and authorized assets. |
|
||||
| `plan` | Convert simple Markdown or Ryan-labelled dialogue into a render plan. |
|
||||
| `validate` | Validate a hand-authored render plan without generating audio. |
|
||||
| `render` | Generate, align, process, assemble, and quality-check audio. |
|
||||
|
||||
`render --resume` reuses completed raw generations. `render --assemble-only`
|
||||
repairs trimming or seams without rerunning the model.
|
||||
|
||||
## Supported profiles
|
||||
|
||||
- Donna
|
||||
- Chris Engineer
|
||||
- Grandpa Bomber
|
||||
- Ryan Pitch Meeting: Producer and Writer roles
|
||||
- Val Holiday
|
||||
|
||||
See `skills/local-voice/references/voice-catalog.md` for the production rules
|
||||
that distinguish these profiles.
|
||||
|
||||
## Setup
|
||||
|
||||
1. Follow `docs/INSTALL_MACOS.md` on the verified Apple Silicon path.
|
||||
2. Install the private `local-voice-authorized-assets` archive from CW-026.
|
||||
3. Download the exact Candle model into the runtime `models` directory.
|
||||
4. Run `scripts/verify_install.py`.
|
||||
5. Install the Claude or Codex plugin package.
|
||||
6. Start a new task so the host discovers the skill.
|
||||
|
||||
The private recovery package and public plugin are intentionally separate:
|
||||
the public repository contains redistributable software, while the private
|
||||
Drive folder contains the authorized voice configuration.
|
||||
|
||||
## Calling Local Voice
|
||||
|
||||
```bash
|
||||
python3 scripts/local_voice.py plan \
|
||||
--voice donna \
|
||||
--script /absolute/path/brief.md \
|
||||
--output-plan /absolute/path/brief-plan.json \
|
||||
--output-audio /absolute/path/brief.mp3
|
||||
|
||||
python3 scripts/local_voice.py render \
|
||||
/absolute/path/brief-plan.json \
|
||||
--resume
|
||||
```
|
||||
|
||||
Set `LOCAL_VOICE_RUNTIME` when the runtime is not installed at the
|
||||
platform-default location.
|
||||
|
||||
## Requirements
|
||||
|
||||
- Python 3.11 recommended
|
||||
- CosyVoice3 0.1.0 Candle build
|
||||
- CosyVoice3-0.5B-Candle model
|
||||
- ffmpeg
|
||||
- Whisper CLI with word timestamps
|
||||
- NumPy and SoundFile
|
||||
- Authorized reference audio and transcripts
|
||||
|
||||
## Operational rules
|
||||
|
||||
- Keep canonical `text` separate from pronunciation-safe `tts_text`.
|
||||
- Generate the largest safe complete passages; paragraph breaks are candidates,
|
||||
not mandatory cuts.
|
||||
- Do not normalize tempo on short greetings, closings, or reactions.
|
||||
- Preserve raw WAV, Whisper JSON, processed WAV, and QA reports.
|
||||
- Use fixed assets for exact Ryan signatures and very short reactions.
|
||||
- Do not silently fall back to a cloud provider.
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
| Symptom | Resolution |
|
||||
|---|---|
|
||||
| A word is clipped | Retain more natural tail or add neighboring sacrificial context, then rerun with `--assemble-only` when possible. |
|
||||
| A seam clicks | Move the join to a quiet boundary or add a natural pause; do not regenerate speech that already passed. |
|
||||
| A short reaction sounds generic | Use an authorized fixed asset from the Ryan signature library. |
|
||||
| Opening or closing sounds drunk or rushed | Remove tempo adjustment; short frames stay native. |
|
||||
| Literal alignment fails | Inspect retained Whisper JSON; normalized fuzzy alignment is expected to tolerate ordinary ASR variation. |
|
||||
| `lead` uses the wrong pronunciation | Keep canonical text and use a local `tts_text` override such as `led` for ammunition. |
|
||||
|
||||
## Security and rights
|
||||
|
||||
Local Voice does not require a service credential. Reference audio must only be
|
||||
installed or used when the operator has authorization. The public repository
|
||||
must never contain private voice assets, source recordings, generated briefings,
|
||||
or model binaries.
|
||||
Reference in New Issue
Block a user