Skip to content
facecam

Live face swap latency, measured

How long a frame takes to go from your camera to our GPU, get a new face, and come back to your screen. Measured on October 4, 2026 from Texas, USA, through the same live face swap you can use on Facecam today. In Chrome 154.0 on a computer, a swapped frame took a median of 217 ms from camera to screen, at 17–18 frames a second, and 24 ms of that was GPU time.

217ms

Camera to screen, median, in Chrome 154.0 on a computer

17–18fps

Swapped frames shown per second in Chrome (the camera runs at 30)

24ms

GPU time per frame, median, from arrival to reply

On the day, every browser run used a GPU in EU West. A session uses whichever GPU is free, and with a closer one the network part is shorter: see GPU time vs network time below.

Camera to screen, by browser

Each bar is one frame's trip, split into its four steps (the median of each step). The camera frame waits for the page to grab it, gets compressed to a JPEG, goes to the GPU and back, then is decoded and drawn on the page.

  • Waiting to be grabbed
  • Grab and compress
  • Network and GPU
  • Decode and draw
Chrome, computer896×504 frames217 ms
Waiting to be grabbed: 37 msGrab and compress: 12 msNetwork and GPU: 154 msDecode and draw: 2 ms

37 + 12 + 154 + 2 ms

Chrome, phone-size frames (touch emulation)640×360 frames202 ms
Waiting to be grabbed: 36 msGrab and compress: 10 msNetwork and GPU: 156 msDecode and draw: 1 ms

36 + 10 + 156 + 1 ms

WebKit (Safari's engine), computer896×504 frames191 ms
Waiting to be grabbed: 13 msGrab and compress: 9 msNetwork and GPU: 164 msDecode and draw: 3 ms

13 + 9 + 164 + 3 ms

Firefox, computer896×504 frames192 ms
Waiting to be grabbed: 23 msGrab and compress: 4 msNetwork and GPU: 162 msDecode and draw: 2 ms

23 + 4 + 162 + 2 ms

The total on the right is the median of whole trips, so it can differ a little from the sum of the step medians.
Chrome 154.0, WebKit 26.6 (Playwright) and Firefox 155.0 (Playwright). Each run is a new live session of 20 measured seconds; every browser run used a GPU in EU West. The first step isn't exactly comparable between browsers: Chrome read the clip as a camera, WebKit and Firefox got it from a canvas stream, and Firefox only reports when a frame was shown, not when it was captured.
BrowserFrame sentRunsFrames shown / sMedian95th pct
Chrome, computer896×504217–18217 ms318 ms
Chrome, phone-size frames (touch emulation)640×360219–22202 ms233 ms
WebKit (Safari's engine), computer896×504213–21191 ms339 ms
Firefox, computer896×504222192 ms229 ms

GPU time vs network time

The GPU server reports how long it held each frame. The rest of the round trip is the network, both ways. A session starts on whichever GPU is free, so we ran 3 sessions and each one landed somewhere else: Canada, US Central and EU West.

The GPU part stays about the same: 24 ms for a computer-size frame. The network part depends on how far you are from the GPU.

  • GPU server
  • Network, both ways
GPU in Canadasession 180 ms
GPU server: 24 msNetwork, both ways: 57 ms

24 + 57 ms

GPU in US Centralsession 2231 ms
GPU server: 149 msNetwork, both ways: 56 ms

149 + 56 ms

GPU in EU Westsession 3157 ms
GPU server: 24 msNetwork, both ways: 133 ms

24 + 133 ms

Computer-size frames (896×504), medians per session. The total is the median round trip.

One slow stretch, kept in the numbers: in session 2, computer-size frames took 149 ms on the GPU server (median) and only 11 came back each second. Later in the same session, the same frame size took 23 ms.

Network echo: a tiny message sent to the GPU server and back before the swap starts, median. Ready after: from asking for a session to a GPU ready to swap, starting from scratch. That wait happens once, not on every frame.
SessionGPU regionNetwork echoReady after
#1Canada56 ms19.3 s
#2US Central32 ms21.9 s
#3EU West130 ms23.3 s

By frame size

The site sends smaller frames from phones and on slow connections. Smaller frames upload faster and need a little less GPU time, but the distance to the GPU still sets most of the round trip.

Medians. GPU: every frame of every session. Round trip: one value per session, 15 seconds of frames each.
Frame sizeSentJPEGFrames back / sGPURound trip (Canada / US Central / EU West)
Computer896×50467 KB11–30 of 3024 ms80 / 231 / 157 ms
Phone640×36032 KB24 of 2421 ms74 / 57 / 153 ms
Data saver480×27016 KB15 of 1519 ms72 / 54 / 152 ms

Inside the GPU server

Each computer-size frame goes through six steps on a server with one NVIDIA L4 GPU, about 23 ms in all. These are the median of each session's median.

Read the JPEG2.3 ms2.3 ms
Find the face4.8 ms4.8 ms
Swap the face8.1 ms8.1 ms
Mask hair and glasses2.3 ms2.3 ms
Blend it back3.7 ms3.7 ms
Write the JPEG1.9 ms1.9 ms
The steps add up to a little less than the GPU server time above, which also counts the short wait for a free worker.

If you're further from the GPU

At most 4 frames are on their way at once. So when a round trip takes longer than about 133 ms, fewer than 30 frames a second can come back. The delay goes up, and the frame rate goes down.

To show this, our test client held every frame and every reply a little longer, adding the extra time below to each round trip. These rows are simulated on top of the real sessions.

Computer-size frames offered at 30 a second: frames back per second and the median round trip.
Added round tripGPU in CanadaGPU in US CentralGPU in EU West
None (as measured)30 fps · 80 ms11 fps · 231 ms24 fps · 157 ms
+50 ms24 fps · 133 ms30 fps · 110 ms17 fps · 209 ms
+100 ms20 fps · 183 ms24 fps · 159 ms15 fps · 259 ms
+200 ms13 fps · 282 ms15 fps · 260 ms11 fps · 358 ms

How we measured

  • The real service. Every number comes from live sessions on the production live face swap, the same GPUs and servers our users get, opened by a test account. Nothing was run on a special server.
  • The camera. A looping 1280×720 webcam clip of an AI-generated person at 30 frames a second, swapped to one of our built-in faces. Chrome read the clip as its camera. WebKit and Firefox can't, so the page got the same frames from a video stream drawn on a canvas.
  • In the browser. We opened the Live Face Swap tool and timed every frame from inside the page: when the camera captured it, when the page grabbed and sent it, when the swapped frame came back, and when it was drawn on screen. The app itself was not changed.
  • GPU vs network. A scripted client sent the same clip at each frame size and read the server's own time for every frame. Network time is the round trip minus that.
  • The machine. Apple M4, macOS 26.5.1, on a home connection in Texas, USA. One session at a time.
  • The numbers. Medians and 95th percentiles over every frame. 8 browser runs and 3 scripted sessions, all on October 4, 2026.

What this doesn't measure

  • The camera's own delay before a frame reaches the browser, and the screen's delay before it lights up. A real webcam and monitor add their own time on both ends.
  • Real phones. The phone row is Chrome on a computer pretending to be a phone, so it sends phone-size frames. A phone's processor and mobile data will differ.
  • Other places and networks. Everything ran from one home connection in Texas, USA. Further from our GPUs, expect the network part to grow.
  • A slower network in the browser. Chrome's network throttling doesn't slow an open WebSocket (we checked), so the longer paths above were simulated in the scripted client instead.
  • Video calls. Passing the swapped camera into Meet or Zoom adds the call's own delay.
  • Other face swap apps. We only publish numbers we measured ourselves.

Use these numbers

You're welcome to quote and chart them. Please link this page and give the date:

Facecam, "Live face swap latency, measured", October 4, 2026. https://www.facecam.ai/live-face-swap/benchmark

The raw data has every frame. Questions about the method, or want a run from your setup? Email support@facecam.ai. Logos and screenshots are in the press kit.

See the delay for yourself

Start your camera, pick a face and become anyone you want, live.

Try live face swap