How-to · Voice cloning

How to Clone a Voice with Echo

A practical field guide: clean sample, smart trim, short lines — with the real mobile limits that decide whether a clone sounds like you.

5sMin sample
10 MBUpload cap
240Chars / line
iOS+Android
Echo's waveform trim screen for voice cloning, captured from the app

To clone a voice with AI, you need a clean sample of speech, a model that can learn that voice, and a text-to-speech step that speaks new lines in the same identity. In Echo on iOS or Android, the practical path is short and fixed. Download at get.echofx.ai.

Capture Trim Name Text Generate Play

This is the how-to companion to our deeper overview of how voice cloning works. Here we stay operational: what to record, how long it must be, file size caps, text limits, and when not to bother.

Key takeaways

  • Custom clone flowRecord, upload, video, or Share Extension → waveform trim → name → up to 240 characters of text → generate speech.
  • Hard limits in EchoAt least 5 seconds of selected audio; max 10 MB per cloning upload; 240 characters per spoken line.
  • Quality is the sampleQuiet room, single speaker, no music bed, no heavy reverb — more than “perfect wording.”
  • Two product pathsClone a voice you provide, or speak with a prebuilt famous / voice-over library voice (same text cap).
  • Cost modelGenerations spend credits from your wallet. Free users may see a first-generation paywall shortly after the first success.

What you need before you start

You do not need a studio. You do need controllable audio and clear permission.

  • Echo installed on your phone (iOS or Android). Get it at get.echofx.ai — one link for both stores.
  • Permission to use the voice you are cloning. No consent = hard stop ethically and legally.
  • A quiet space, or an existing clean recording (or a video you can pull speech from).
  • Enough credits for a generation, or a subscription / pack that tops them up. Exact amounts are live-configured in-app — trust the wallet, not blog folklore.
Echo's consent gate before cloning: confirm this is your own voice or you have explicit permission
Echo asks for explicit consent before every custom clone — this is not skippable.

How to clone a voice in Echo

Echo’s custom voice cloning follows a fixed pipeline. Do not skip trim: noisy heads and tails wreck clones more often than “bad models” do.

01

Capture the source audio

Open Add Voice from the tab bar and pick an input method — the four options as they appear in the app:

  • Record Sound — record live with the microphone.
  • Capture Sound from Video — pull the speech out of a clip you have.
  • Share Your Voice With Us — send audio in from another app (Files, Voice Memos, camera roll). Deep link form: echoapp://open?page=voiceCloning&audio=… (or video).
  • YouTube URL — capture speech from a video link.
Echo's Add Voice screen with four input options: Record Sound, Capture Sound from Video, Share Your Voice With Us, and YouTube URL
The Add Voice screen in Echo — pick how the sample comes in.

For a first clone, record yourself reading two short paragraphs of normal speech — the recorder recommends 2+ minutes of continuous talking, and more clean speech helps. Natural pacing beats a flat monotone. Avoid whispering and avoid shouting.

Echo's Record Your Voice screen with the microphone button and the 2-minute recording recommendation
Record Sound: tap the mic, speak continuously, tap again to stop.
02

Trim with the waveform

Echo’s trim UI uses start/end markers on a waveform. Keep only clear speech.

  • Minimum length: 5 seconds for the selected range.
  • Maximum upload size: 10 MB for cloning (voice isolator allows 30 MB for longer mixed tracks).
  • Cut silence, music intros, other speakers, and phone notifications.

If the file is too large, export a shorter clip before re-importing. Prefer a shorter clean take over a crushed long one.

Echo's Trim the Sound screen showing a speech waveform with draggable start and end markers, Confirm Selection and Record Again buttons
Trim the Sound: drag the handles to keep only clear speech, then Confirm Selection.
03

Name the voice

Use a name you’ll recognize later (“Me — quiet desk,” “Podcast host,” not “Voice 3”). Metadata stays with the creation so you can re-generate lines without re-recording.

Echo's Voice Name dialog with a text field filled in as My Studio Voice
The Voice Name dialog appears right after you confirm the trim.
04

Enter the text to speak

Custom voice speech synthesis accepts up to 240 characters per generation — enough for a voice note, ad line, or caption, not a chapter.

Write for the ear: short sentences, natural punctuation, no ALL CAPS spam. Split longer scripts into multiple takes and stitch in an editor if needed.

Echo's Enter Text screen with a name, the text to speak, and the Create the Sound button showing its credit cost
Enter Text: the Create the Sound button shows the credit cost before you commit.
05

Generate, play, and save

Generate speech, then play it in Echo’s shared player (seek, skip, lock-screen controls). From Library you can play, rename, delete, or re-generate with the same voice.

Audio is stored with your account; manage custom voices from the Voices section of Library.

Recording checklist for better clones

Desktop clone guides often push multi-minute datasets. On a phone, smaller clean samples win if they are honest speech.

Do Do not Why it matters
One speaker, steady mic distance Group chatter or TV in the background The model treats the whole spectrum as “the voice.”
Natural pitch range (question + statement) A forced accent you never use You generate speech in the style of the sample.
Dry room, phone 20–30 cm away Bathroom reverb or outdoor wind Room noise becomes part of the clone identity.
Trim to ≥5 s of real speech Pad with silence to “make it longer” Empty audio wastes the minimum without adding identity.
Keep file under 10 MB Full 4K videos with no speech trim Cloning upload limit is 10 MB; extract audio first.

Custom clone vs famous voice library

Not every “clone a voice” job needs a personal sample.

Path Best for Main limits
Custom voice cloning Your voice, a consenting collaborator, a brand narrator you control Need a sample; 5 s min; 10 MB max; 240-char lines
Famous / voice-over library Quick demos and catalog styles already in the app No personal sample; 240-char lines; ads may show for free users; catalog can vary by country

Famous voices still use the same style of generation pipeline and the same text cap. They do not replace consent rules: do not use any AI voice to impersonate someone for fraud or deception.

Credits and free-user friction

Wallet truth

Echo uses subscriptions, credit packs, and (when enabled) ads. Voice generations spend credits. Balances and prices are live-configured — check the in-app wallet rather than assuming a fixed price from an older article.

Non-subscribers may hit a first-generation paywall after creating voice content (voice delay is short while audio can keep playing under a volume duck). Treat every generation as a paid-or-gated action unless your plan says otherwise.

Not for you

Skip this workflow if…

  • You want to clone a voice without clear permission.
  • You need long-form narration in one shot (240-character cap is intentional).
  • Your only sample is a noisy concert, multi-speaker call, or heavy music bed — re-record or isolate first.
  • You expect offline, on-device cloning with no network. Generations use the app’s cloud pipeline.

How Echo fits the job

Echo is one AI music and voice app on iOS and Android with the same product surface. Custom voice cloning sits beside famous voices, AI music, sound effects, and vocal isolation. Download at get.echofx.ai. For short spoken lines, the path is: capture → trim within limits → speak up to 240 characters → manage results in Library.

For ownership, commercial reuse, and platform rules, read the live Terms of Service. Behavior here tracks Echo’s feature reference — not marketing wish lists.

Clone a voice with Echo

Download Echo, record a clean 5+ second sample, and generate your first line of speech.

Download Echo

FAQ

How long does a voice sample need to be?

The waveform trimmer enforces a 5-second minimum for the selected clip. Clean seconds beat noisy minutes.

What file size limit applies when cloning a voice?

Cloning uploads are capped at 10 MB. Extract audio from video first when needed. Voice isolator allows up to 30 MB for longer mixed files — different feature.

How much text can the cloned voice speak at once?

Custom and famous voice text fields accept up to 240 characters per generation. Split longer scripts into multiple takes.

Can I import audio from another app?

Yes. Use Echo’s Share Extension for audio or video, or open cloning with the share deep link. Share launches can skip onboarding.

Is this the same as an AI voice changer?

Related but not identical. Cloning builds a reusable voice identity from a sample, then speaks new text. Effects transform audio from a prompt — see the AI voice changer guide.