Private by architecture
Cloud voice-clone services keep your biometric voiceprint on their servers. Here the model, the reference clip and the synthesis all live on your machine — there's nothing to leak.
Register your voice once from a 10–30 second clip. From then on, Claude can narrate any script — a voiceover, a full Smart Video — in your voice, synthesized 100% on-device. The reference clip never leaves your machine.
Clone my voice from ~/Desktop/me-talking.m4a, then narrate this script with it.
…and Claude drives the whole pipeline below on your machine.
Cloud voice-clone services keep your biometric voiceprint on their servers. Here the model, the reference clip and the synthesis all live on your machine — there's nothing to leak.
Cloned-voice synthesis is priced at 2.5¢ per 1,000 characters and hard-capped at 10¢ per generation — a full video's narration never costs more than a dime.
Faceless-channel output with a real, consistent host voice — yours. Scripts Claude writes from trend research get read the way your audience already knows you.
create_cloned_voice takes a 10–30s clip of you speaking clearly (a video file works too — the audio is extracted). Whisper auto-fills the transcript; the voice gets an id in your library.
ViralMint runs VoxCPM (an open-source 0.5B voice-cloning model) locally — your reference audio and every synthesized line stay on-device.
A standalone voiceover mp3, a narrated Smart Video with captions and b-roll, or a re-voiced competitor remix — the cloned voice is a first-class TTS option across the pipeline.
List, reuse and delete cloned voices from chat. Voices survive engine reinstalls — the reference clips are kept separately.
create_cloned_voicelist_cloned_voicestool_voiceovergenerate_smart_video Download the free, open-source desktop app (macOS / Windows / Linux) and open it. It bundles the whole pipeline — yt-dlp, Whisper, FFmpeg — and the MCP server.
It shows your loopback endpoint and a bearer token, with a ready-to-paste Claude Code command and a JSON config snippet for Cursor / Claude Desktop.
Paste the command. Your client loads ViralMint's tools, the built-in workflow recipes (this skill is one) and read-only resources like your balance and recent videos.
Type the prompt above. Claude runs the recipe end-to-end against the local app and hands you the finished mp4 — no clicking through the UI.
| What matters | Typical paid agent tools | ViralMint |
|---|---|---|
| Agent tools / skills | Behind a paid tier | Free & open-source (AGPL) |
| Where it runs | Their cloud | Your machine (loopback MCP) |
| Your footage & data | Uploaded to their servers | Stays local — only AI calls leave |
| Pricing | Monthly subscription | Prepaid, pay-per-render at ~cost |
| Output watermark | On free tiers | Never |
| Script quality | Prompt → generic script | Grounded in real breakouts + your data |
highlighted column = clearer fit.
macOS on Apple Silicon or Linux, plus the Voice Cloning engine — a one-time ~550 MB on-device install from Settings. Windows and Intel Macs aren't supported yet, and Claude will say so rather than fail silently.
10–30 seconds of one person speaking clearly, no background music or crosstalk. Phone recordings are fine. Claude extracts and converts the audio automatically, and Whisper writes the transcript for you.
No. The reference clip is stored locally, the VoxCPM model runs locally, and synthesis happens on your CPU. This is the main difference from ElevenLabs-style cloud cloning.
Only with their permission. The skill is built for cloning your own voice or voices you're authorized to use — not public figures or third parties.
Install the free ViralMint desktop app, connect it to Claude Code or Cursor, and run this skill — no subscription, no watermark, your footage never leaves your machine.