12k
All articles

Adding Captions to a Video With AI in the Browser

Generate video captions in the browser with Transformers.js and Whisper. Extract audio, use WebGPU or WASM, create WebVTT subtitles, and keep media on-device.

OpenReplay Team
OpenReplay Team
Adding Captions to a Video With AI in the Browser

Transformers.js runs OpenAI’s Whisper speech-recognition model directly in the browser, so a video’s audio can be transcribed into captions with no upload, no API key, and no server.

Most caption tutorials start by asking you to send the user’s video to a transcription API and pay for every minute. If the video is private, or you just don’t want another bill and a backend to maintain, that approach doesn’t work for you.

This guide builds the whole pipeline on the client. You decode the audio, run Whisper in a Web Worker with a WebGPU-to-WASM fallback, format the output as valid WebVTT, and attach the result to a <video> element.

Key Takeaways

  • Whisper in transformers.js expects mono audio at 16 kHz as a Float32Array, so a video’s audio track has to be decoded and resampled before you transcribe it.
  • Having navigator.gpu does not mean WebGPU will work. Use device: "webgpu" only when navigator.gpu.requestAdapter() returns an adapter, and fall back to WASM otherwise.
  • Run the transcription pipeline in a Web Worker. Model loading and inference would otherwise block the main thread for the whole job.
  • A WebVTT file starts with WEBVTT, separates cues with blank lines, uses hh:mm:ss.ttt timestamps, and needs & and < escaped in cue text.
  • In-browser Whisper is practical for clips a few minutes long. For hour-long recordings or bulk media, use a hosted transcription service.

What You Are Building and Why It Stays On-Device

The finished in-browser captioning app has four parts: a file input, a <video> element, a worker that runs Whisper, and a generated .vtt file attached as a caption track. The media is read from the user’s disk into memory and never leaves the tab. The only network traffic is the one-time download of the model from the Hugging Face Hub and, unless you host them yourself, the ONNX Runtime WASM files, which transformers.js loads from a CDN by default. No audio or video is sent anywhere.

This article doesn’t cover WebCodecs, burned-in captions or MP4 export. Captions stay as a separate text track that the browser renders over the video. Frame-level work belongs in a real-time video processing pipeline built on WebCodecs, which is a different job.

How Do You Extract the Audio Track From a Video?

Whisper in transformers.js takes mono audio sampled at 16 kHz as a Float32Array. The pipeline accepts raw audio as a typed array and assumes it already has the right sampling rate. It doesn’t check, so audio at any other rate produces wrong transcripts and no error.

decodeAudioData() decodes the file’s audio. For video files, whether this works depends on which containers and codecs the browser supports, so an MP4 that plays fine in one browser may fail to decode in another. You then render the decoded buffer through an OfflineAudioContext created with one channel at 16,000 Hz. Web Audio downmixes the source to fit the single-channel destination and resamples it to the context’s rate, so one render gives you both conversions.

async function extractAudio(file) {
  const arrayBuffer = await file.arrayBuffer();
  const decodeCtx = new AudioContext();
  const decoded = await decodeCtx.decodeAudioData(arrayBuffer);
  await decodeCtx.close();

  const targetRate = 16000;
  const offline = new OfflineAudioContext(
    1,
    Math.ceil(decoded.duration * targetRate),
    targetRate
  );
  const source = offline.createBufferSource();
  source.buffer = decoded;
  source.connect(offline.destination);
  source.start();
  const rendered = await offline.startRendering();
  return { audio: rendered.getChannelData(0), duration: decoded.duration };
}

Keep duration. The formatter uses it later.

Run Transformers.js Whisper in a Web Worker

Pass return_timestamps: true when you call the transformers.js ASR pipeline. The output then includes a chunks array, and each entry has a text string and a timestamp: [start, end] pair in seconds. Run the pipeline inside a Web Worker so the page stays responsive: model loading and inference would otherwise block the main thread for the whole job.

Install the package with npm install @huggingface/transformers. The worker creates the pipeline once and reuses it for every file after that:

// worker.js
import { pipeline } from "@huggingface/transformers";

let transcriberPromise;

async function pickDevice() {
  if (!("gpu" in navigator)) return "wasm";
  try {
    const adapter = await navigator.gpu.requestAdapter();
    return adapter ? "webgpu" : "wasm";
  } catch {
    return "wasm";
  }
}

function getTranscriber() {
  transcriberPromise ??= pickDevice().then((device) =>
    pipeline("automatic-speech-recognition", "onnx-community/whisper-tiny.en", { device })
  );
  return transcriberPromise;
}

self.onmessage = async ({ data }) => {
  try {
    const transcriber = await getTranscriber();
    const output = await transcriber(data.audio, {
      return_timestamps: true,
      chunk_length_s: 30,
      stride_length_s: 5,
    });
    self.postMessage({ type: "done", chunks: output.chunks });
  } catch (err) {
    self.postMessage({ type: "error", message: String(err) });
  }
};

On the main thread, send the samples to the worker by transferring the buffer instead of copying it. The new URL(..., import.meta.url) form is the worker pattern Vite recognises, and webpack 5 handles it too:

const worker = new Worker(new URL("./worker.js", import.meta.url), { type: "module" });

function transcribe(audio) {
  return new Promise((resolve, reject) => {
    worker.onmessage = ({ data }) =>
      data.type === "done" ? resolve(data.chunks) : reject(new Error(data.message));
    worker.postMessage({ audio }, [audio.buffer]);
  });
}

Whisper works on 30-second windows. If you leave out chunk_length_s, transformers.js keeps only the first 30 seconds of audio and logs a warning, so the rest of the video gets no captions. With chunk_length_s: 30 and stride_length_s: 5, the pipeline cuts longer audio into overlapping 30-second pieces and joins the results into one set of timestamped chunks.

WebGPU vs WASM: Pick the Device Before Loading

Transformers.js should run Whisper on WebGPU only when navigator.gpu.requestAdapter() returns an adapter, and on WASM otherwise. Switching to the GPU takes one pipeline option, device: "webgpu", set when the model loads. Having navigator.gpu does not guarantee that WebGPU works, though. Call requestAdapter(), which can resolve to null, and choose WebGPU only when it returns an adapter. navigator.gpu is also exposed in workers, so pickDevice() above runs there.

ConditionDeviceWhat the user gets
No navigator.gpu"wasm"Slower transcription, CPU only
navigator.gpu exists, requestAdapter() returns null"wasm"Slower transcription, CPU only
Adapter returned"webgpu"GPU-accelerated inference

A null adapter is common on real machines. Chrome’s WebGPU troubleshooting guide lists the usual causes: the user has turned off graphics acceleration in settings, the GPU is on Chrome’s blocklist, WebGPU isn’t supported on that platform yet, or Chrome can’t find a GPU at all. Checking the adapter yourself before creating the pipeline, as pickDevice() does, covers all of these.

Hugging Face’s guide puts global WebGPU support at around 85% as of March 2026. The WebGPU implementation status page lists it as on by default in Chrome 113 and later on Windows, macOS and ChromeOS, and in Chrome 121 and later on most Android devices. On Linux, Chrome turns it on only for some GPUs. Firefox has it on by default on Windows since version 141 and on Apple Silicon Macs since version 147. Safari 26 supports it on macOS, iOS, iPadOS and visionOS. Firefox on Android still has it off by default: it needs Firefox Beta or Nightly plus the gfx.webgpu.ignore-blocklist setting in about:config.

How Do You Convert Whisper Chunks to WebVTT?

Each Whisper chunk becomes one WebVTT cue: its start and end times go on the timing line, and its text goes on the line below. MDN’s WebVTT reference gives the rules. The file starts with WEBVTT, a blank line separates each cue from the next, and two timestamps joined by --> set when each cue is shown.

The formatter below always writes the long hh:mm:ss.ttt form. Hours get two digits or more, minutes and seconds never go past 59, and the milliseconds part is always three digits. Every cue also has to finish later than it starts. Two characters can’t appear as-is in cue text: & and <. The code swaps & for &amp; first, then < for &lt;. MDN also suggests writing > as &gt;, so the code does that too.

const pad = (n, width = 2) => String(n).padStart(width, "0");

function toVttTime(seconds) {
  const total = Math.round(seconds * 1000);
  const h = Math.floor(total / 3_600_000);
  const m = Math.floor((total % 3_600_000) / 60_000);
  const s = Math.floor((total % 60_000) / 1000);
  const ms = total % 1000;
  return `${pad(h)}:${pad(m)}:${pad(s)}.${pad(ms, 3)}`;
}

function escapeCueText(text) {
  return text.replace(/&/g, "&amp;").replace(/</g, "&lt;").replace(/>/g, "&gt;");
}

function chunksToVtt(chunks, duration) {
  const cues = [];
  for (const { text, timestamp } of chunks) {
    const [start, end] = timestamp;
    const stop = end ?? duration;
    const body = escapeCueText(text.replace(/\s+/g, " ").trim());
    if (!body || stop <= start) continue;
    cues.push(`${toVttTime(start)} --> ${toVttTime(stop)}\n${body}`);
  }
  return `WEBVTT\n\n${cues.join("\n\n")}\n`;
}

A few details matter here:

  • toVttTime rounds to whole milliseconds before splitting the value into hours, minutes and seconds, so it can never output .1000.
  • & is escaped first because the other replacements add ampersands, and escaping it last would double-escape them.
  • Collapsing whitespace removes any newline in the model’s text. A blank line inside a payload ends the cue early.
  • The ?? duration fallback is a precaution in case a final chunk comes back without an end time.

Output from the formatter looks like this (illustrative):

WEBVTT

00:00:00.000 --> 00:00:04.320
Welcome back. Today we are wiring up captions.

00:00:04.320 --> 00:00:09.100
The model runs in a worker &amp; nothing is uploaded.

Once you have a VTT file you can also translate the subtitles into other languages.

Attach the Captions With a Blob URL

To display the captions, wrap the VTT string in a Blob with type text/vtt, create an object URL for it, and use that URL as the src of a <track> element. Give the track kind="captions", srclang and default. The Blob constructor encodes JavaScript strings as UTF-8, which is the encoding WebVTT requires.

function attachCaptions(video, vtt) {
  const url = URL.createObjectURL(new Blob([vtt], { type: "text/vtt" }));
  const track = document.createElement("track");
  track.kind = "captions";
  track.srclang = "en";
  track.label = "English (auto-generated)";
  track.src = url;
  track.default = true;
  video.append(track);
  track.track.mode = "showing";
  return () => URL.revokeObjectURL(url);
}

srclang is "en" because whisper-tiny.en is an English-only model. The track is added after the video has already loaded, so setting mode to "showing" explicitly makes it display straight away. Call the returned function when you replace or remove the video, so the object URL is released. In React, the same track goes inside the <video> you render, as covered in embedding video in React and building a video player with React.

What Are the Limits of In-Browser Whisper?

In-browser Whisper works well for clips a few minutes long and poorly for hour-long recordings.

  • Download size. The first run downloads model weights ranging from tens to hundreds of megabytes, depending on the variant. The whisper-tiny.en model page links to its files. Larger Whisper variants on the Hugging Face Hub are more accurate but bigger, so check each model card before switching.
  • Speed. Long files are slow, especially on the WASM path. A few minutes of audio is fine. An hour of audio ties up the tab for a long time.
  • Memory. The full decoded audio is held in memory as a typed array.
  • Bulk or long media. For hour-long recordings or batch jobs, a hosted transcription service is the better choice.

Conclusion

You can caption a short video entirely in the browser: decode and resample its audio, run Whisper in a worker on WebGPU (or WASM when no adapter is available), format the chunks into WebVTT, and attach the file as a Blob-backed track. Start with whisper-tiny.en on a two-minute clip and check the cue timings against the audio. Move to a larger Hub model only if accuracy isn’t good enough and your users can handle the bigger download.

FAQs

Does transformers.js download the Whisper model on every page load?

No. On the first run, transformers.js downloads the model files and stores them in the browser cache, so later loads read from the cache instead of the network. The env.useBrowserCache setting controls this. In transformers.js v4, ModelRegistry.is_pipeline_cached reports whether a pipeline's files are already cached, and ModelRegistry.clear_pipeline_cache removes them.

How do I show model download progress while Whisper loads in a worker?

Pass a progress_callback function in the pipeline options, next to device. Transformers.js calls it with status updates as each model file downloads. Inside a worker, forward each update to the main thread with postMessage and render the progress bar there. Transformers.js v4 adds a progress_total event, which reports overall loading progress so you do not have to add up per-file updates yourself.

Can I generate captions for videos that are not in English?

Yes, but you need a multilingual Whisper checkpoint instead of an English-only model ending in .en. Pass language and task in the transcriber call, for example language: 'french' with task: 'transcribe'. Setting task to 'translate' makes Whisper produce English text from speech in another language. Set the track's srclang to the language of the caption text, not the language of the audio.

Can the captioning app run without contacting the Hugging Face Hub or a CDN?

Yes, if you self-host every file. Set env.allowRemoteModels to false and point env.localModelPath at a folder on your server that holds the model files. By default the ONNX Runtime WASM binaries also load from a CDN, so set env.backends.onnx.wasm.wasmPaths to your own copies too. For fully offline use, add a service worker so the page itself loads without a connection.

DevTools for the frontend

Gain Debugging Superpowers

Unleash the power of session replay to reproduce bugs, track slowdowns and uncover frustrations in your app. Get complete visibility into your frontend with OpenReplay — the most advanced open-source session replay tool for developers.

Star on GitHub12k

We use cookies to improve your experience. By using our site, you accept cookies.