Home/Transcribe an iPhone Video
Verified 2026-08-15

How to Transcribe an iPhone Video (2026 Guide)

Four working paths, matched to what you actually have. Apple Notes (iOS 18, iPhone 12+, English only), 8 iOS apps with speaker-label support (Otter, Notta, Speakwise for multi-speaker; Detail and Documents for solo), hosted upload from Photos via Safari for anything over 30 min, and AirDrop-to-Mac with self-hosted Whisper for sensitive content. Includes iPhone-specific file-format gotchas — HEVC vs H.264, Slow-Mo audio issues, Live Photo exports.

By VexaScribe Editorial · Published · Verified

4
workflow paths
Notes · Apps · Hosted · Mac
iOS 18+
Notes transcript
iPhone 12 or later
8
apps compared
incl. diarization
100+
languages
via hosted

Quick answer

Three main paths. Short clips, English: play the video and use Apple Notes' audio-recording + transcript feature (iOS 18, iPhone 12+, on-device). Longer videos, single speaker: use an iOS app that reads directly from Photos — Detail, Documents by Readdle, or Vomo. Multi-speaker or 30+ min or non-English: use Otter.ai (virtual meetings), Speakwise (in-person), Notta (multilingual), or upload from Photos to a hosted service like VexaScribe. Sensitive content: AirDrop to Mac and transcribe locally with Whisper — no data leaves your devices.

Which path fits your video?

Match your video's length, language, speaker count, and privacy needs to the right path. Picking wrong wastes time — Apple Notes doesn't handle a 45-minute interview; a hosted service is overkill for a 30-second clip.

Your videoBest pathWhy
Under 5 min, English, single speakerApple Notes (built-in, free)Fastest, on-device, no upload
5-30 min, single speaker, English or common languageDetail, Documents by Readdle, iScribeDirect Photos access, no export loop
Multi-speaker Zoom/Teams/Meet recordingOtter.ai (OtterPilot bot integration)Best diarization on virtual meetings
In-person multi-speaker (roundtable, focus group)Speakwise or hosted servicePurpose-built for single-device in-person capture
Non-English or multilingual audioNotta (104 languages) or VexaScribeWhisper large-v3 backbone, code-switching handled
30+ min, need SRT/VTT exportHosted service (upload from Photos)Whisper large-v3, subtitle formats, speaker labels
Sensitive content (legal, medical, HR)AirDrop to Mac + local WhisperNo data leaves your devices
Legal / court-of-record accuracyRev (AI + human hybrid)Human review for lowest error rate

Path 1 — Apple Notes (built-in, iOS 18)

iOS 18's Notes app records audio and generates a real-time transcript on-device. It only records new audio though — it doesn't transcribe an existing video file directly. The workaround: play the video on one device (or through a speaker) and let Notes record and transcribe as it plays. Works well for short clips; awkward for longer content.

Requirements

  • iOS 18 or later
  • iPhone 12 or later (transcript feature requires the Neural Engine on A14 Bionic or newer)
  • English audio only as of iOS 18 initial release — Apple has signalled more languages via update
  • Apple Intelligence summaries require iPhone 15 Pro / 15 Pro Max / any iPhone 16

Steps

  1. Open the Notes app on your iPhone. Create a new note.
  2. Tap the paperclip / attachment icon and choose Record Audio.
  3. Start playback of your video on a second device (Mac, iPad, second iPhone, TV, laptop). Or unmute and play the video on the same iPhone in Safari and let the mic pick it up — quality drops noticeably.
  4. While recording, Notes shows a live transcript. When the video ends, tap stop.
  5. Tap the recording, then the three-dot menu → Add Transcript to Note. The full transcript is inserted into the note as searchable text.
  6. Copy and paste to wherever you need it (Mail, Messages, Drafts, Google Docs).

Limitations: no timestamps, no speaker labels, no export to SRT/VTT, English only, and you're capturing through a microphone which introduces artifacts. For anything beyond quick reference, use Path 2 or Path 3.

Path 2 — iOS transcription apps

Dedicated iOS apps read the video directly from your Photos library, skipping the play-and-re-record gymnastics. Eight that actually rank on 2026-08-15 SERPs, grouped by strength:

Otter.ai

Meeting bot + iOS transcription (virtual meetings speciality)

Strengths: Best iOS diarization for Zoom/Teams/Meet via OtterPilot auto-join. iPhone app records live too. 300 min/mo free tier.

Limits: iPhone app is essentially remote-control for the desktop/cloud experience — best paired with virtual meetings

Best for: Zoom/Teams/Meet meeting capture with speaker labels

Notta

Multilingual video + audio transcription (104 languages)

Strengths: Best cross-platform multilingual capture. Handles code-switching. 120 min/mo free tier. Speaker labels supported.

Limits: Free tier caps at 120 min; longer videos need paid plan

Best for: Multi-language interviews, international team meetings

Speakwise

In-person multi-speaker capture (roundtables, focus groups, panels)

Strengths: Best for 3+ person in-person conversations from a single iPhone placed centrally. 95%+ accuracy in optimal conditions. AirPods hands-free recording. 100+ languages. Native Notion integration.

Limits: In-person focus — virtual meetings less optimized than Otter

Best for: Focus groups, roundtables, in-person panels captured on one device

Rev

AI + human hybrid transcription for maximum accuracy

Strengths: AI-only or human-reviewed tiers. Cleanest speaker identification when accuracy matters most. Legal/medical accuracy.

Limits: Paid per-minute (AI $0.25/min, human $1.50/min) — no free tier for files

Best for: Legal, medical, court-of-record accuracy where errors are costly

Detail

Video-first workflow (record + transcribe + edit)

Strengths: Reads directly from Photos, tap-to-export transcript, integrated video editor for creators

Limits: Content-creator focus — no diarization, no multi-language depth

Best for: TikTok/Reels/YouTube Shorts creators editing short-form vertical video

Documents by Readdle

File manager + transcription add-on

Strengths: Free tier includes audio/video transcription. iCloud + Google Drive integration. Handles file management alongside transcription.

Limits: Transcription can be slower than dedicated apps. No speaker labels.

Best for: iPad users who already file-manage in Documents

iScribe

Comprehensive audio-to-text with real-time features

Strengths: Real-time transcription with advanced AI features. Broad format support.

Limits: Subscription model for full features

Best for: Frequent-use transcription across audio and video sources

VexaScribe (via mobile Safari)

Hosted Whisper large-v3 for long / multi-speaker / non-English videos

Strengths: Handles 5 GB / multi-hour videos, speaker labels (up to 50), 100+ languages, SRT/VTT/TXT export. Free tier ~30 min.

Limits: Not a native iOS app — upload via Safari from Photos share-sheet

Best for: Long videos, multi-speaker interviews, non-English audio, SRT/VTT export

Common workflow across most: open the app, tap import, pick your video from Photos, wait 1-5 minutes for transcription (depends on video length and processing tier), export as TXT/DOCX/SRT. For diarization specifically: Otter, Notta, Speakwise, Rev, and VexaScribe support speaker labels; Detail, Documents, and iScribe produce single-track transcripts. Diarization is where iOS-native transcription has caught up significantly in 2026.

Path 3 — upload from Photos to a hosted service

For anything over 30 minutes, multi-speaker with SRT export need, or non-English audio, a hosted service that runs Whisper large-v3 (or equivalent) delivers markedly better accuracy than iOS-native paths. The upload flow from iPhone Photos is quick — no need to move the file to a computer first.

Steps (any hosted service)

  1. Open Photos on your iPhone, tap the video.
  2. Tap the share icon (square with an up arrow).
  3. Scroll to Save to Files (or use the service's share-sheet extension if installed).
  4. Open Safari, navigate to the hosted service, tap upload, pick the video from Files.
  5. Configure language and speaker labels if the service supports it, then start.
  6. Download the transcript as TXT, SRT, or DOCX when processing completes.

VexaScribe accepts iPhone videos (MP4, MOV, HEVC) up to 5 GB from Safari on iPhone. Whisper large-v3 processes a 1-hour video in roughly 5-10 minutes. Speaker labels handle up to 50 speakers (best accuracy on 2-6). Multi-language support covers 100+ languages. Free tier covers ~30 minutes of transcription.

Alternatives with similar iPhone Photos workflows: Rev (AI or human), Otter, Descript, TurboScribe. For accuracy benchmarks across services see how accurate is Whisper.

Path 4 — AirDrop to Mac + local transcription

For sensitive content that shouldn't leave your devices (legal, medical, HR interviews), AirDrop the video to a Mac and transcribe locally with self-hosted Whisper. Higher setup cost, zero data exposure.

  1. On iPhone: Photos → select the video → share icon → AirDrop → pick your Mac.
  2. On Mac: video lands in Downloads.
  3. Follow the Whisper install guide (5 minutes on macOS via Homebrew) if you haven't installed yet.
  4. Run whisper ~/Downloads/video.mov --model base.
  5. Whisper writes video.txt (plus .srt and .vtt if requested) to the same folder.

On Apple Silicon Macs the small/medium models run in real-time. The large model is slower but more accurate — pick--model medium for a good balance. See the model size picker for detailed trade-offs.

Multi-speaker interview workflow

Multi-speaker capture is where iOS-native transcription has changed most in 2026. Apple Notes still produces single-track output. But Otter, Notta, Speakwise, and hosted services (VexaScribe, Rev) all handle diarization on iOS now. Which one is right depends on how the multi-speaker audio was captured.

Virtual meeting (Zoom, Teams, Google Meet)

Each speaker has their own microphone channel — diarization is much easier from the audio. Otter.ai is the clear leader: OtterPilot auto-joins meetings via calendar integration and produces speaker-labeled real-time transcripts. Notta and VexaScribe both handle recorded video files well but Otter's calendar-integrated bot is a category on its own.

In-person capture (roundtable, panel, focus group)

All speakers share one microphone — diarization is harder. Speakwise is purpose-built for this: single iPhone placed centrally on the table, 95%+ accuracy in optimal conditions per their reported testing. For higher-stakes content, VexaScribe with Whisper large-v3 + pyannote diarization is the accuracy leader for hosted iOS uploads.

Legal / court-of-record accuracy

AI-only will not meet legal filing standards. Rev offers human-reviewed transcription at $1.50/min with 99%+ accuracy targets. Otherwise, an AI-first pass through VexaScribe or Otter followed by human proofreading is the practical hybrid.

iPhone video file formats and gotchas

iPhone videos default to two containers depending on your Camera settings. Both work with almost every transcription service, but there are edge cases worth knowing.

  • HEVC (H.265) inside MOV — default on iPhone 7+ with Camera set to "High Efficiency." Smaller files, better quality. Some older tools and desktop editors can't decode HEVC; when in doubt switch Camera to "Most Compatible" in Settings → Camera → Formats before recording.
  • H.264 inside MP4 — the "Most Compatible" setting. Universal support, larger files.
  • Slow-Mo videos (120 or 240 fps)common gotcha: some Slow-Mo captures record no audio at all, especially on older iPhones or specific Camera app modes. Even when audio is present, timing metadata differs and some transcription services stretch or compress the audio incorrectly. Verify audio exists before uploading; convert to normal-speed if possible.
  • Live Photos — the ~3-second video clip attached to a still. Most transcription services can't extract the audio directly. Export as regular video first: Photos → tap Live Photo → share icon → Save as Video.
  • Screen Recordings — .mov files with system + optional mic audio. Work like any other video, but verify mic recording was enabled at capture time (Control Center → Screen Record → long-press → toggle microphone).
  • Cinematic mode videos — dual-layer with focus depth data. Audio is standard, transcription works normally.
  • Portrait mode video — audio is standard, standard transcription workflow.
  • 4K ProRes video (iPhone 13 Pro+) — huge files (multi-GB per minute). May exceed some hosted service upload caps. Check the service's max upload size before starting a long ProRes video upload.

Privacy tiers — pick by content sensitivity

Journalism sources, legal privilege, medical HIPAA-adjacent content, and HR investigations demand different privacy postures than marketing videos and personal memos. Four tiers, ranked by data-locality strength:

On-device only

Tools: Apple Notes (iOS 18, iPhone 12+, English), iOS Voice Memos transcription

Audio never leaves the phone. Neural Engine handles the model locally. Best for sensitive personal content.

App-cloud (check per app)

Tools: Detail, Documents by Readdle, iScribe (read each app's privacy policy)

Some process on-device, some upload to vendor cloud. Vendor claims vary — verify before uploading legal or medical content.

Hosted cloud (encrypted in transit)

Tools: Otter.ai, Notta, Speakwise, VexaScribe

TLS 1.2+ upload, at-rest encryption is standard. Enterprise plans typically add data-residency + non-training contracts. Business meetings, journalism interviews.

Truly private (offline)

Tools: AirDrop to Mac + self-hosted Whisper (see whisper-install-guide)

Zero cloud exposure. Requires macOS with Whisper installed. Best for privileged content (attorney-client, medical HIPAA-adjacent, HR investigations).

Common problems

"No audio found" error on upload

Check whether the video was recorded with the mic muted (screen recording with Do Not Disturb), whether iOS stripped the audio during a share-sheet compression, or whether it's a Slow-Mo capture with no audio track. Re-export from Photos with Options → Most Compatible to preserve audio.

Notes transcript is empty or partial

Apple Notes transcript only works on English audio and only on iPhone 12 or later running iOS 18+. If either condition isn't met, no transcript is generated — the recording still saves as audio. For non-English videos, skip Path 1 and use Path 2 or 3.

Upload from iPhone Safari is slow or stalls

Cellular uploads over 4G/5G can time out on long videos. Switch to Wi-Fi. If the file is over 1 GB, AirDrop to a Mac first and upload from there — Wi-Fi tethering also works but adds a hop.

Transcript quality is poor

Usually the audio, not the model. iPhone videos captured across a noisy room, at distance, or with wind pick up background noise that Whisper handles worse than clean lapel-mic audio. If the audio quality is fixed, expect 5-15% word error rate for real-world speech (see Whisper accuracy benchmarks).

Multi-speaker transcript merges voices

Path 1 (Apple Notes) and some Path 2 apps (Detail, Documents) produce single-track output regardless of speaker count. For speaker labels, use Otter (virtual), Speakwise (in-person), Notta, VexaScribe, or Rev.

Verified sources

Every hardware requirement, app claim, and workflow step on this page was cross-checked against these sources on 2026-08-15:

  • Apple Support — Record and transcribe audio in Notes (iPhone 12+, iOS 18, on-device English)
  • Otter.ai product page + iOS App Store listing (OtterPilot, 300 min/mo free tier)
  • Notta product page (104 languages, 120 min/mo free tier)
  • Speakwise product pages (single-device in-person capture, 100+ languages, Notion integration)
  • Rev pricing page (AI $0.25/min, human $1.50/min)
  • Detail iOS product page (Photos import, video editor integration)
  • Readdle Documents App Store listing (transcription add-on)
  • Apple Camera format documentation (HEVC vs H.264, ProRes, Slow-Mo)
  • Independent 2026 reviews: ticnote.com, livetranscribe.pro, vidnotes.app

Frequently asked questions

How do I pull a transcript from an iPhone video?

Three paths depending on video length. Short clips: play the video and use Apple Notes' audio-recording + transcript feature (iOS 18, iPhone 12+, English only, on-device). Longer videos: use an iOS transcription app that reads directly from Photos (Detail, Documents by Readdle, Vomo). 30+ minutes or multi-speaker: share the video from Photos to a hosted service (VexaScribe, Rev, Otter) via Safari.

Can you get a transcript from an iPhone video recording?

Yes. iPhone videos (MOV or MP4 from the Camera app) contain a standard audio track that any transcription service can read. Apple's built-in Notes app in iOS 18 transcribes new audio recordings on-device (English, iPhone 12+); for existing videos, use an iOS transcription app that imports from Photos or upload the video to a hosted transcription service via Safari.

Is there a way to pull a transcript from a video?

Yes — every modern transcription service extracts the audio track from your video (MP4, MOV, MKV, WebM, AVI) and runs speech-to-text on it. On iPhone: use a native app like Detail or Documents by Readdle that reads from Photos, or upload the video from Photos to a hosted service through Safari. On desktop: AirDrop or import the video, then use a service like VexaScribe or self-host Whisper locally.

Can you get transcripts from an iPhone recording?

Yes. iOS 18 introduced native transcription in the Notes app for new voice recordings on iPhone 12 and later (English initially). For existing recordings — whether from Voice Memos, Camera, or a screen recording — you can either play them back for Notes to transcribe live, use a third-party iOS app that reads the file directly, or upload to a hosted transcription service. See our /voice-memo-to-text page for audio-first workflows.

How do I get a transcript from my own video?

Locate the video in the Photos app or Files app. If under 5 minutes and English-only, playback + Apple Notes works on iPhone 12+ with iOS 18. If longer or non-English, either use a dedicated iOS transcription app (Detail, Vomo, Documents by Readdle) or share the video to Files and upload it via Safari to a hosted service like VexaScribe. The hosted path handles multi-hour videos and multiple speakers.

Which iPhone models support the built-in Notes transcript?

iPhone 12 and later, running iOS 18 or newer. The feature depends on the Neural Engine in the A14 Bionic chip. Older iPhones (11, XR, XS, X, SE 2nd gen) don't get the transcription feature even after updating to iOS 18. Apple Intelligence-powered summaries additionally require iPhone 15 Pro, 15 Pro Max, or any iPhone 16.

Does the Apple Notes transcript work in Spanish or French?

Not at launch. As of iOS 18's initial release, Notes' audio transcription is English-only. Apple has signaled additional language support in future updates but hasn't confirmed timing as of 2026-08-15. For non-English iPhone videos, use a third-party app or hosted service — most support 90-100+ languages via Whisper large-v3 or equivalent.

What file formats do iPhone videos use?

Default is HEVC (H.265) inside a .mov container for High Efficiency mode (iPhone 7 and later), and H.264 inside .mp4 for Most Compatible mode. Both work with almost every transcription service. Slow-Mo videos, Live Photos, and Screen Recordings have quirks — Live Photos need to be exported as a regular video first (share icon → Save as Video), and Slow-Mo audio timing can confuse some services.

How long does it take to transcribe an iPhone video?

Depends on the path. Apple Notes generates transcript in real time as it records (so a 5-minute video takes ~5 minutes). iOS apps like Detail process a 10-minute video in ~1-2 minutes on-device or in the cloud. Hosted services like VexaScribe process a 1-hour video in roughly 5-10 minutes using Whisper large-v3 in the cloud. Upload time depends on your connection — 4G/5G can stall on files over 1 GB.

Is transcribing an iPhone video private?

Apple Notes runs entirely on-device — audio never leaves your iPhone. Third-party iOS apps vary: check the app's privacy policy for whether transcription runs on-device or in the cloud. Hosted services process the audio on their servers (that's the trade-off for higher accuracy). For truly private handling of sensitive video, AirDrop to a Mac and run Whisper locally — no data leaves your Apple devices.

Can I get speaker labels from an iPhone video transcript?

Not from Apple Notes — it produces a single-speaker transcript regardless of how many voices are in the audio. Most iOS transcription apps don't offer diarization either. Hosted services do: VexaScribe supports up to 50 speakers with best accuracy on 2-6 speakers, and most other hosted services (Otter, Rev, Descript) include diarization on paid tiers. Interview and meeting recordings benefit substantially from proper speaker labels.

Related guides