ZooWork Market · Skill
local-voice
Local speech-to-text (transcribe/translate) and text-to-speech (voice generation) using Whisper.cpp and Sesame CSM-1B. Use when (1) transcribing or translating audio/voice messages, (2) generating voice messages or audio from text, (3) user asks to "say something", "read this aloud", or "send a voice message", (4) processing inbound voice notes, (5) voice cloning from reference audio. Fully local on DGX Spark — no API keys, no cloud. For setup issues or background, see references/SETUP_GUIDE.md.