Many of them do, and the popular ones say so in their privacy policies. The rule of thumb is simple: anything the app’s own chip can compute stays on the phone, and anything a large AI model produces almost certainly leaves it. Trimming, transitions, filters and colour adjustments are local work. Generative effects, AI upscaling, face swaps, background replacement and most auto-captioning involve sending your footage, or frames and audio from it, to a server.
You can check which camp a specific app falls into in about five minutes, and it’s worth doing before you edit anything with other people’s faces in it.
Which features need the cloud #
| Feature | Usually runs | Why |
|---|---|---|
| Cut, trim, speed, transitions, colour, text overlays | On the phone | Plain video processing the GPU has done for years |
| Stickers, LUT-style filters, basic stabilisation | On the phone | Small, fixed operations |
| Auto captions and transcription | Varies | Phone speech engines can do it locally; many apps use a cloud service for accuracy |
| Background removal, object erase | Varies | Increasingly on-device on recent phones, still cloud in many apps |
| AI upscaling and “enhance” | Cloud | The models are far too large for a phone |
| Text-to-video, style transfer, face swap, relighting | Cloud | Generative models run in data centres |
| Cloud drafts, project sync, collaboration | Cloud by definition | The project is stored remotely |
“Varies” is the important row. Two apps offering the same button can implement it completely differently, and the marketing copy rarely says which.
What a major editor’s policy actually says #
CapCut is the example worth reading because it’s the most used mobile editor with AI features, and because its policy is unusually explicit. In the version published at the time of writing, the CapCut privacy policy states that the service “may pre-upload User Content to improve your experience”, which means footage can reach servers before you ever tap export.
It also describes collecting information about what’s in your content, including “identifying the objects and scenery that appear, the existence and location within an image of face and body features and attributes, the nature of the audio, and the text of the words spoken”. On the training question, it says information is used for “improving our technology, such as our machine learning models and algorithms”, and that for auto-captions specifically “we will collect the transcribed text and the edited text in order to develop and improve this function”. Retention is described as for as long as necessary to provide the service.
None of that is hidden or unusual for a free cloud-backed editor. It’s just more specific than most people expect, and policies change, so check the current version rather than trusting this summary or any other.
Test any editor yourself in five minutes #
- Airplane mode. Turn it on, open the app, and try each feature you care about. Whatever still works runs on the phone. Anything that spins, queues or errors needs a server.
- Watch for a progress bar that doesn’t match your phone. A four-second clip that takes thirty seconds and shows a percentage is usually a round trip, not local compute.
- Read the store privacy label. Both stores require a declaration of what’s collected and whether it’s linked to you. What AI app privacy labels mean covers how to read them, including their limits.
- Search the policy for four words: upload, train, improve, retain. The sentences around them are the ones that matter, and they’re usually clearer than the summary at the top.
- Check whether drafts sync. If your projects appear on another device, your media is on a server, whatever the editing features do.
What’s actually at stake #
Video is denser than most personal data. A single clip can contain other people’s faces who never agreed to anything, a child’s school uniform, your street, a conversation in the background, and metadata about where and when it was shot. Unreleased work, client footage and anything under an NDA carry their own problems, since uploading is a disclosure regardless of whether anyone watches it.
The training clauses are the part worth reading carefully, because they outlive the edit. A clause covering “improving our models” means material derived from your content can shape a system long after you deleted the project. Deleting the video doesn’t unwind that.
The same reasoning applies to stills, which is covered in whether it’s safe to upload photos to a chatbot, and to the photo tools on your phone in which AI photo editing tools work offline.
Practical ways to reduce what you send #
- Do the ordinary edit in an offline editor, including the stock editor built into your phone, and use a cloud tool only for the one clip that genuinely needs a generative effect.
- Trim before you upload. Send the six seconds the effect applies to, not the four-minute source.
- Keep faces out of generative features unless everyone in frame has agreed. Face swap and relighting tools are the ones that draw complaints.
- Turn off cloud drafts and auto-sync if you don’t need editing across devices.
- Check location metadata before sharing exports, since many apps preserve it.
- Treat client and workplace footage as off-limits for consumer tools, which is the same rule as using AI with confidential client data.
Where an on-device model helps, and where it doesn’t #
Worth being direct: a local language model is not a video editor. It cannot cut, render or apply effects, and no phone-sized model generates video.
What it can do is the text and comprehension work around the edit, without sending anything anywhere. Personal LLM runs open models such as Qwen 3.5 and Gemma 4 on the phone’s own chip, and every model in its catalog can take an image, so you can screenshot a frame and ask what’s in it, draft captions, titles and descriptions offline, or paste an app’s privacy policy in and ask what a clause means in plain language. It’s free on iOS and Android, there’s no account, and the chats stay on the device. For the policy-reading job specifically, explaining a document in plain English covers the prompts.
Frequently asked questions #
Does CapCut upload my videos? #
Its privacy policy says user content may be pre-uploaded to improve your experience, and describes analysing content and using information to improve its machine learning models. Read the current version yourself before deciding what to edit there, since the terms change and differ by region.
Does editing on my phone mean my video stays private? #
Only if the features you use run on the phone. A local app with a cloud effect is still a cloud upload for that effect. Airplane mode tells you which is which faster than any policy.
Is background removal done on the device? #
Sometimes. Recent phones can segment a subject locally, and some apps use that. Others send frames to a server for better results, especially for hair and motion. The airplane mode test answers it for your specific app.
Do video apps use my footage to train AI? #
Several say they use content and derived information to improve their models and features, which is the same thing in practice. Look for “improve our services”, “machine learning models” and “algorithms” in the policy, and for a setting that opts out, which some apps offer and many don’t.
What’s the most private way to edit a video with AI features? #
Use on-device features for everything you can, keep generative effects for clips without identifiable people, and avoid apps whose policies claim broad training rights over your content. If an effect only exists in the cloud, decide whether that specific clip is one you’d be comfortable handing over permanently.