IGTranscript

Bulk transcribe Instagram Reels

You can work through a stack of Reels on this page, one file after another, but there is no batch button here and there never will be one — the tool has no server, so every file has to pass through your own browser, and the page can only hold one at a time. What makes a stack practical anyway is that the expensive part is paid once: the speech model downloads on your first file (79.7 MB for the default model, 43.6 MB for the smaller one) and your browser serves it from cache for every file after that. Measured on this tool, a repeat visit transferred zero bytes of model. The thing that actually runs out on batch work here is memory, not download — 23.04 MB for every minute of 48 kHz stereo audio, before the model even starts. All of it is spelled out below, with the measurements it came from.

Last checked: 6 October 2026. Every number on this page was measured on this tool, on a two-core, four-gigabyte machine, or worked out from the page's own source code, which is linked in the sections below. Nothing here is estimated.

Why there is no batch box on this page

Bulk transcription is normally a server feature. You paste twenty links, a machine somewhere fetches twenty videos and runs twenty jobs, and you come back later. That is a real way to do it, and it needs three things this page does not have: a backend to receive the links, an outbound fetcher to pull the videos, and a queue that survives you closing the tab.

This page was built without any of them, on purpose. There is no server, so there is nothing to receive a list of links and nothing left running when you leave. The trade is explicit: no batch, in exchange for no upload. If your Reels are already files on your device, that trade usually costs you less than you would expect — the next three sections are about why.

One file, one buffer, no queue. When a file finishes, only the transcript remains; the decoded audio is released. Even inside a single file the model works through the audio in 30-second chunks and the page asks for 16,000 Hz mono, but that is chunking one file — it is not a queue of several, and it will not start the next file for you.

The download is paid once, not once per file

What each file costs in downloaded bytes: the first file downloads the 79.7 megabyte speech model, every file after it downloads nothing because the browser has it cached. Decoding and transcription happen on every file regardless. Bytes downloaded, per file The bar only appears once. The work below it appears every time. First file 79.7 MB the speech model, downloaded once then decode + transcribe runs on your own CPU, every file Every file after that 0 bytes already in your browser's cache then decode + transcribe the same work, unchanged
The model download is a one-off per model per browser; the decoding and transcription are not. That is why the tenth Reel in a row costs you time but no bandwidth — and why the thing that eventually stops you is memory, not data.

On a server-based tool, every job is somebody else's machine doing the work, and you never see the cost. Here the cost is visible and front-loaded, which is exactly what makes a stack cheaper than it looks:

For a batch, that arithmetic matters. Twenty files on a server tool is twenty jobs. Twenty files here is one download of 79.7 MB and then nineteen files that transfer nothing. The model files themselves, taken straight from the model repository, are encoder_model_quantized.onnx at 23.20 MB and decoder_model_merged_quantized.onnx at 53.69 MB for the default model — those two are the bulk of that 79.7 MB.

Two measured details are worth knowing if you are going to sit and watch the download. The progress total climbs as it goes rather than being fixed at the start — measured, 26.0 MB at the first sample and 79.7 MB by the third — because the library discovers the model's files one at a time, so the number shown is true of the files it knows about at that moment, not a promise about the final size. And the download is deliberately not started when the page loads. It used to be, on the theory that it would overlap with you reading; measured, it did not pay for itself: choosing the other model afterwards left both downloads running, and a run that should have cost 44 MB cost 118 MB. The download now starts at your first file, and only for the model you actually picked.

What the file becomes before it is transcribed

This is the part that decides how many Reels you get through, and it is invisible unless you know it is happening. Before any speech recognition runs, the browser decodes the file's audio track into raw samples and holds them uncompressed. That decoded buffer — not the size of the file on your disk — is what has to fit in memory.

Those samples are 32-bit floating point, at the file's own sample rate and channel count, so the arithmetic is fixed: bytes per second is sample rate × channels × 4. Written out for the formats a Reel is likely to arrive in:

Then there is a second buffer. The recognition model wants 16 kHz mono, so the page resamples the decoded audio down to exactly that — 64,000 bytes a second, every time, whatever the source was. So the buffer the model receives is small and the same size for everybody; the buffer that can kill the tab is the first one, and its size is set by the source file's format rather than by its length.

Two 20-minute recordings. Decoded from 48 kilohertz stereo, the audio occupies 460.8 megabytes. Decoded from 16 kilohertz mono, the same 20 minutes occupies 76.8 megabytes. Both are resampled to 16 kilohertz mono before the model sees them. The same 20 minutes, two source formats Bar length is the decoded audio sitting in memory. Decoded from 48 kHz stereo 460.8 MB Decoded from 16 kHz mono 76.8 MB Both are then resampled to 16 kHz mono before the model sees them. The top bar is the one that decides whether the tab survives.
Six times the memory for the same duration and the same transcript. Which bar you get is decided by the source file's sample rate and channel count — and nothing about the two files' sizes on disk tells you which one you have.

Concretely: twenty minutes of audio at 48 kHz stereo decodes to 460.8 MB. The same twenty minutes, arriving as 16 kHz mono, decodes to 76.8 MB. Same duration, same transcript, six times the memory — and a 20-minute MP3 of that length might be under 19 MB on disk while decoding to 460.8 MB in the tab. The file size you can see is not the number that matters.

What actually runs out first: memory

This is the honest reason a big stack is harder here than on a server, and it is worth understanding before you queue up twenty files.

Measured on this four-gigabyte machine, with the tool running in a normal browser tab:

So the ceiling is above 115.2 MB of decoded audio, because that much passed. Where exactly it sits I am not going to claim, because the 20-minute WAV's sample rate and channel count were not recorded when it was tried. If it was 48 kHz stereo it decoded to 460.8 MB and the ceiling is somewhere between the two figures; if it was already 16 kHz mono it decoded to 76.8 MB and the refusal had another cause entirely. Both readings fit the evidence I have, so the useful statement is the one you can act on: keep the decoded audio small, and the duration takes care of itself.

One more thing makes the real peak worse than the arithmetic above suggests. The resample to 16 kHz mono happens while the original buffer is still alive, so for a moment the tab is holding both. For a 20-minute 48 kHz stereo file that is 460.8 MB plus 76.8 MB — a peak of 537.6 MB, not 460.8 MB. The peak, not the average, is where a tab dies.

And when it dies it does so silently. A tab killed for memory cannot print an error, because the thing that would print it has been killed too. What you see is the page simply stop. Worth knowing if you are troubleshooting: the browser also reports "there is no audio in this file", "this codec is missing" and "this is too big to hold" identically — the same EncodingError with nothing in it to tell them apart. Measured in this browser, a 77-byte text file renamed to .mp4 is rejected in exactly the same way as a valid 20-minute WAV. That is why the page names all three causes instead of guessing at one.

Nothing is uploaded, and you can check that yourself

"Nothing leaves your device" is the kind of claim every tool makes, so here is the version you can falsify. The page fetches from exactly two hosts, and both are named in its own source: cdn.jsdelivr.net for the recognition library, and huggingface.co for the speech model — 79.7 MB of model files on the default setting, and nothing else. There is no third host, because there is no server of ours to send anything to: the whole site is a set of static files with no server-side code at all.

To check it, open your browser's developer tools, switch to the Network tab, tick "Preserve log", and then pick a file. What you should see is a handful of requests for the library and the model, all of them coming towards you. What you will not see is a request whose body is your video. Your file is never in a request body, because nothing here knows how to send one — and that is a property of the design, not a promise about our behaviour.

Two more details from the source, for anyone who wants to verify the shape of it rather than take it on faith. The page sets allowLocalModels = false before it loads anything, so it will not go looking for a model on your own disk either. And the recognition runs on WebAssembly with 8-bit quantised weights: the two model files behind the default choice are encoder_model_quantized.onnx at 23.20 MB and decoder_model_merged_quantized.onnx at 53.69 MB. Because all of that runs on the CPU of the machine you are sitting at, a stack takes as long as it takes — a server-based tool would finish faster, and you would never see where the time went.

A batch recipe that holds up

  1. Get the files onto your device first. Your own posts can be saved from Instagram. For someone else's Reel there is no download button, so a screen recording while it plays is the usual route. Do this in one pass, before you start transcribing — switching between apps mid-stack is what makes batch work feel slow.
  2. Pick your model once, then leave it alone. Default for accuracy, smaller if the first download is a problem. Changing mid-stack costs another 43.6 MB for no benefit.
  3. Set the language explicitly on every file. This page does not auto-detect — the browser build of the recognition library has no detection step. Get it wrong and you get confident nonsense rather than an error: measured, the same English clip run with the picker set to Spanish came back as 404 words of repeated Spanish instead of the 54 English words it actually contains.
  4. Watch the decoded size, not the duration. Ten minutes at 48 kHz stereo is 230.4 MB and is the practical ceiling these measurements point at; the same ten minutes as 16 kHz mono is 38.4 MB. If you must do a long one, do it alone rather than at the end of a stack.
  5. Export as you go. Copy the text or download the .txt, .srt or .vtt before you open the next file. This page keeps no history — nothing is stored anywhere, so an unexported transcript is gone when you move on.
  6. Close other tabs before a long stack. The tool competes with everything else for the same memory, and the peak is higher than the decoded size alone: the resample buffer is held alongside it, which for a 20-minute 48 kHz stereo file means 537.6 MB rather than 460.8 MB. If the tab runs short it can be killed outright, and a killed tab cannot show you an error — you would just see it stop.

If you need a link-based batch instead

Sometimes the files are not on your device and screen-recording twenty Reels is not reasonable. Then you need a server-based tool, and there is nothing wrong with that — it is a different job with a different trade. Four things are worth checking on any tool that advertises bulk, because these are the places the claims tend to be thin:

Where this leaves you

For a handful of files, transcribing them one after another here is fast and costs no download after the first 79.7 MB. For twenty files you did not have to download anyway, it is still workable as long as each one is short. For twenty Reels that only exist as links, you want a server-based tool and you should read its limits before you trust its Batch tab.

Everything above is the same tool described on the home page — the file picker, the language setting and the export buttons. See Limits for the file-size and speed measurements, and the privacy page for what does and does not leave your device.