Guides · How it works
How on-device background removal works, step by step
What happens between dropping a photo in and downloading a transparent PNG: model download and caching, resizing to 1024 × 1024, the IS-Net segmentation network, the alpha mask, and full-resolution compositing.
By the developer of ImgCutout · Published · 5 min read
Most background removers are web pages in front of a server: your photo is uploaded, a GPU in a data center runs the AI model, and the result is sent back. ImgCutout works differently. The AI model is sent to your browser once, and every image is processed on your own device. This article explains, step by step, what happens from the moment you drop a photo in to the moment you download the result, and why it is built this way.
The building blocks
Four pieces of open technology make in-browser AI practical:
- ormbg, an open-source background-removal model published on Hugging Face under the Apache 2.0 license. It is a segmentation network from the IS-Net family, designed to separate a main subject from its background.
- ONNX, an open file format for AI models. The model is distributed as ONNX files that any compatible runtime can execute.
- ONNX Runtime Web, Microsoft's engine for running ONNX models in a browser, using WebAssembly (fast, portable code that runs on your CPU) or WebGPU (access to your graphics card).
- Transformers.js, Hugging Face's library that downloads the model, prepares images for it, and runs it through ONNX Runtime Web.
None of these require installing anything. They are delivered as ordinary web files.
Step 1: the model is downloaded once
The first time you use the tool, your browser downloads the model weights. The default is the 8-bit quantized version, about 28MB compressed. Optional 16-bit and 32-bit versions (about 78MB and 156MB) can be chosen in Settings; see 8-bit vs 16-bit vs 32-bit for the difference.
The browser stores the downloaded files in its cache. On later visits, the model loads from the cache instead of the network, which is why the first image takes noticeably longer than the following ones. The download contains only the generic model; it carries nothing about you or your photos.
Step 2: your image is read locally
When you drop or paste an image, the browser reads the file into memory. Nothing is sent over the network. The image is decoded into raw pixels, and the original version is kept aside for the final step. On phones, tablets, and computers with 4GB of memory or less, photos larger than about 4 megapixels are first scaled down to that size, because holding several full-size copies of a 12 to 48 megapixel photo can crash a mobile browser tab.
Step 3: the image is resized for the model
The model expects an input of exactly 1024 × 1024 pixels, with color values scaled to the range 0 to 1. A copy of your image is resized to that shape. This is standard practice for segmentation models: they are trained at a fixed resolution, and running them on a 12-megapixel photo directly would be extremely slow and memory-hungry.
Resizing for the model does not mean your result is 1024 pixels. The model only produces a mask; the mask is scaled back up and applied to your original pixels (step 6).
Step 4: the network predicts a mask
The resized image goes through the network. IS-Net-style models work roughly like this: an encoder repeatedly shrinks the image while extracting increasingly abstract features (edges, then textures, then shapes, then "this looks like a person"); a decoder then works back up to full model resolution, combining the abstract understanding with fine detail from earlier layers. The output is a single-channel image of 1024 × 1024 values, one per pixel, each saying how likely that pixel is to belong to the main subject.
This step is where almost all of the computation happens. On a typical laptop CPU with WebAssembly it takes from under a second to a few seconds; with WebGPU on a capable graphics card it can be faster. WebGPU vs WebAssembly covers the trade-offs.
Step 5: the mask becomes an alpha channel
The prediction is turned into an alpha mask: 0 means fully transparent, 255 fully opaque, and values in between are partially transparent. Partial values matter. They are what allow fine hair, fur, and soft edges to blend into a new background instead of being cut along a hard, jagged line.
If you use the edge cleanup setting, it is applied here. "Tighter" shrinks the mask by a few pixels to remove a halo of leftover background; "Wider" grows it to recover parts that were clipped, such as fingers or thin straps. This step is pure image processing and applies instantly, without running the model again.
Step 6: compositing at full resolution
The mask is scaled up to your original image's dimensions and combined with your original pixels. The result is a full-resolution image with transparency. From here, everything else is ordinary canvas drawing in the browser: placing the subject on a background color or image, adding a shadow in the bulk tool, cropping to a preset, and encoding the file as PNG, WEBP, or JPG.
Step 7: refinement when needed
If the model picked the wrong subject or missed part of it, Select object lets you paint over or draw around the object you want. The model then runs again on just that region of the image, which focuses its attention and usually fixes the problem. Your strokes are stored as coordinates on the original image, so they also work at full resolution.
Where this runs: the main thread and a worker
Both tools run the model inside a Web Worker, a background thread, so the page stays responsive while an image is processed, and the bulk tool can work through up to 100 images in a row. On low-memory devices, such as phones, the tool keeps only one AI runtime in memory at a time and handles very large photos more conservatively, because the model, the image, and the mask together can use hundreds of megabytes.
Why build it this way
Running the model on your device has clear advantages: privacy (the photo never leaves your device, which you can verify yourself), no per-image server cost (which is why there are no credits or paid tiers), and no upload wait for large photos. The trade-offs are a one-time model download and speed that depends on your hardware: an old phone is slower than a data-center GPU.
Summary
The tool downloads an open-source IS-Net-family model once and caches it; each image is read locally, a copy is resized to 1024 × 1024, the network predicts a per-pixel foreground probability, and that prediction becomes an alpha mask that is cleaned up, scaled back up, and applied to your original pixels. Everything happens in your browser, using WebAssembly or optionally WebGPU.
Try it on your own photo
Free, no sign-up, and processed on your device. Nothing is uploaded.
Open the background remover