Bulk transcribe Instagram Reels
You can work through a stack of Reels on this page, one file after another, but there is no batch button here and there never will be one — the tool has no server, so every file has to pass through your own browser, and the page can only hold one at a time. What makes a stack practical anyway is that the expensive part is paid once: the speech model downloads on your first file (79.7 MB for the default model, 43.6 MB for the smaller one) and your browser serves it from cache for every file after that. Measured on this tool, a repeat visit transferred zero bytes of model. The thing that actually runs out on batch work here is memory, not download — 23.04 MB for every minute of 48 kHz stereo audio, before the model even starts. All of it is spelled out below, with the measurements it came from.
Last checked: 6 October 2026. Every number on this page was measured on this tool, on a two-core, four-gigabyte machine, or worked out from the page's own source code, which is linked in the sections below. Nothing here is estimated.
Why there is no batch box on this page
Bulk transcription is normally a server feature. You paste twenty links, a machine somewhere fetches twenty videos and runs twenty jobs, and you come back later. That is a real way to do it, and it needs three things this page does not have: a backend to receive the links, an outbound fetcher to pull the videos, and a queue that survives you closing the tab.
This page was built without any of them, on purpose. There is no server, so there is nothing to receive a list of links and nothing left running when you leave. The trade is explicit: no batch, in exchange for no upload. If your Reels are already files on your device, that trade usually costs you less than you would expect — the next three sections are about why.
One file, one buffer, no queue. When a file finishes, only the transcript remains; the decoded audio is released. Even inside a single file the model works through the audio in 30-second chunks and the page asks for 16,000 Hz mono, but that is chunking one file — it is not a queue of several, and it will not start the next file for you.
The download is paid once, not once per file
On a server-based tool, every job is somebody else's machine doing the work, and you never see the cost. Here the cost is visible and front-loaded, which is exactly what makes a stack cheaper than it looks:
- First file. The recognition library comes from
cdn.jsdelivr.netat about 4 MB. The speech model comes fromhuggingface.co. Measured, first visit with the default model transferred 79.7 MB in total; the smaller model transferred 43.6 MB. - Every file after the first. Nothing. The library and the model are ordinary web files, so your browser caches them the way it caches any other web file. Measured on a repeat visit: zero bytes of model transferred.
- Switching models costs a second download. If you start on Better accuracy and then change to Smaller download, that is a separate 43.6 MB. Pick your model before you start the stack and stay on it.
For a batch, that arithmetic matters. Twenty files on a server tool is twenty jobs. Twenty
files here is one download of 79.7 MB and then nineteen files that transfer nothing.
The model files themselves, taken straight from the model repository, are
encoder_model_quantized.onnx at 23.20 MB and
decoder_model_merged_quantized.onnx at 53.69 MB for the default model —
those two are the bulk of that 79.7 MB.
Two measured details are worth knowing if you are going to sit and watch the download. The progress total climbs as it goes rather than being fixed at the start — measured, 26.0 MB at the first sample and 79.7 MB by the third — because the library discovers the model's files one at a time, so the number shown is true of the files it knows about at that moment, not a promise about the final size. And the download is deliberately not started when the page loads. It used to be, on the theory that it would overlap with you reading; measured, it did not pay for itself: choosing the other model afterwards left both downloads running, and a run that should have cost 44 MB cost 118 MB. The download now starts at your first file, and only for the model you actually picked.
What the file becomes before it is transcribed
This is the part that decides how many Reels you get through, and it is invisible unless you know it is happening. Before any speech recognition runs, the browser decodes the file's audio track into raw samples and holds them uncompressed. That decoded buffer — not the size of the file on your disk — is what has to fit in memory.
Those samples are 32-bit floating point, at the file's own sample rate and channel count, so the arithmetic is fixed: bytes per second is sample rate × channels × 4. Written out for the formats a Reel is likely to arrive in:
- 48 kHz stereo — 384,000 bytes a second, 23.04 MB a minute. This is the worst case and the common one for screen recordings.
- 44.1 kHz stereo — 352,800 bytes a second, 21.17 MB a minute.
- 16 kHz stereo — 128,000 bytes a second, 7.68 MB a minute.
- 16 kHz mono — 64,000 bytes a second, 3.84 MB a minute.
- 8 kHz mono — 32,000 bytes a second, 1.92 MB a minute.
Then there is a second buffer. The recognition model wants 16 kHz mono, so the page resamples the decoded audio down to exactly that — 64,000 bytes a second, every time, whatever the source was. So the buffer the model receives is small and the same size for everybody; the buffer that can kill the tab is the first one, and its size is set by the source file's format rather than by its length.
Concretely: twenty minutes of audio at 48 kHz stereo decodes to 460.8 MB. The same twenty minutes, arriving as 16 kHz mono, decodes to 76.8 MB. Same duration, same transcript, six times the memory — and a 20-minute MP3 of that length might be under 19 MB on disk while decoding to 460.8 MB in the tab. The file size you can see is not the number that matters.
What actually runs out first: memory
This is the honest reason a big stack is harder here than on a server, and it is worth understanding before you queue up twenty files.
Measured on this four-gigabyte machine, with the tool running in a normal browser tab:
- A 10-minute file decoded in 10.5 seconds and went through.
- A 20-minute WAV was refused.
- A 30-minute file decoded fine — because its audio was 16 kHz mono, which decodes to 115.2 MB rather than the 460.8 MB the same duration would take at 48 kHz stereo.
So the ceiling is above 115.2 MB of decoded audio, because that much passed. Where exactly it sits I am not going to claim, because the 20-minute WAV's sample rate and channel count were not recorded when it was tried. If it was 48 kHz stereo it decoded to 460.8 MB and the ceiling is somewhere between the two figures; if it was already 16 kHz mono it decoded to 76.8 MB and the refusal had another cause entirely. Both readings fit the evidence I have, so the useful statement is the one you can act on: keep the decoded audio small, and the duration takes care of itself.
One more thing makes the real peak worse than the arithmetic above suggests. The resample to 16 kHz mono happens while the original buffer is still alive, so for a moment the tab is holding both. For a 20-minute 48 kHz stereo file that is 460.8 MB plus 76.8 MB — a peak of 537.6 MB, not 460.8 MB. The peak, not the average, is where a tab dies.
And when it dies it does so silently. A tab killed for memory cannot print an error,
because the thing that would print it has been killed too. What you see is the page simply
stop. Worth knowing if you are troubleshooting: the browser also reports "there is no audio
in this file", "this codec is missing" and "this is too big to hold" identically — the same
EncodingError with nothing in it to tell them apart. Measured in this browser, a 77-byte text
file renamed to .mp4 is rejected in exactly the same way as a valid 20-minute
WAV. That is why the page names all three causes instead of guessing at one.
Nothing is uploaded, and you can check that yourself
"Nothing leaves your device" is the kind of claim every tool makes, so here is the version
you can falsify. The page fetches from exactly two hosts, and both are named in its own
source: cdn.jsdelivr.net for the recognition library, and
huggingface.co for the speech model — 79.7 MB of model files on the default
setting, and nothing else. There is no third host, because there is no server of ours to send
anything to: the whole site is a set of static files with no server-side code at all.
To check it, open your browser's developer tools, switch to the Network tab, tick "Preserve log", and then pick a file. What you should see is a handful of requests for the library and the model, all of them coming towards you. What you will not see is a request whose body is your video. Your file is never in a request body, because nothing here knows how to send one — and that is a property of the design, not a promise about our behaviour.
Two more details from the source, for anyone who wants to verify the shape of it rather
than take it on faith. The page sets allowLocalModels = false before it loads
anything, so it will not go looking for a model on your own disk either. And the recognition
runs on WebAssembly with 8-bit quantised weights: the two model files behind the default
choice are encoder_model_quantized.onnx at 23.20 MB and
decoder_model_merged_quantized.onnx at 53.69 MB. Because all of that runs
on the CPU of the machine you are sitting at, a stack takes as long as it takes — a
server-based tool would finish faster, and you would never see where the time went.
A batch recipe that holds up
- Get the files onto your device first. Your own posts can be saved from Instagram. For someone else's Reel there is no download button, so a screen recording while it plays is the usual route. Do this in one pass, before you start transcribing — switching between apps mid-stack is what makes batch work feel slow.
- Pick your model once, then leave it alone. Default for accuracy, smaller if the first download is a problem. Changing mid-stack costs another 43.6 MB for no benefit.
- Set the language explicitly on every file. This page does not auto-detect — the browser build of the recognition library has no detection step. Get it wrong and you get confident nonsense rather than an error: measured, the same English clip run with the picker set to Spanish came back as 404 words of repeated Spanish instead of the 54 English words it actually contains.
- Watch the decoded size, not the duration. Ten minutes at 48 kHz stereo is 230.4 MB and is the practical ceiling these measurements point at; the same ten minutes as 16 kHz mono is 38.4 MB. If you must do a long one, do it alone rather than at the end of a stack.
- Export as you go. Copy the text or download the
.txt,.srtor.vttbefore you open the next file. This page keeps no history — nothing is stored anywhere, so an unexported transcript is gone when you move on. - Close other tabs before a long stack. The tool competes with everything else for the same memory, and the peak is higher than the decoded size alone: the resample buffer is held alongside it, which for a 20-minute 48 kHz stereo file means 537.6 MB rather than 460.8 MB. If the tab runs short it can be killed outright, and a killed tab cannot show you an error — you would just see it stop.
If you need a link-based batch instead
Sometimes the files are not on your device and screen-recording twenty Reels is not reasonable. Then you need a server-based tool, and there is nothing wrong with that — it is a different job with a different trade. Four things are worth checking on any tool that advertises bulk, because these are the places the claims tend to be thin:
- What "batch" actually means there. Some interfaces show a Batch tab and never explain what it does — how many links at once, whether it is a paid-only feature, and whether the batch runs in parallel or just queues them. If the page does not say, assume nothing.
- Whether the batch is counted or limited. A free tier measured in credits tells you very little until you know how many files one credit buys.
- How long a clip it will accept. A tool that says "no matter the length" without a number has not answered the question. One that names a figure — two minutes, for instance — has.
- What happens to the video. A link-based tool has to fetch the video to a machine it controls. That is inherent to the approach, not a red flag, but it is the reason this page does not do it: 79.7 MB of model is the largest thing that ever moves here, and it moves towards you.
Where this leaves you
For a handful of files, transcribing them one after another here is fast and costs no download after the first 79.7 MB. For twenty files you did not have to download anyway, it is still workable as long as each one is short. For twenty Reels that only exist as links, you want a server-based tool and you should read its limits before you trust its Batch tab.
Everything above is the same tool described on the home page — the file picker, the language setting and the export buttons. See Limits for the file-size and speed measurements, and the privacy page for what does and does not leave your device.