Blog
We Looked at 14,452 Uploads. Seven in Ten Were MP4.
This is the reference post for the rest of the data series on this blog. Every completed upload to VidClean is counted once and tagged with its container format, whether it contains video, coarse duration and size buckets, and the tool that processed it. Nothing else: no filenames, no content hashes, no user identifiers, no per-file rows. What is left is a clean picture of what people actually bring to a web video tool, and it is the largest dataset on the site: 14,452 jobs between July 20 and August 25, 2026.
The other studies in this series are slices of the same stream, which is why some numbers below will look familiar. The stabilization study drew its sub-minute finding from the stabilize slice of this data, the transcription study its language mix, and the audio cleanup study its video-versus-audio split. This post aggregates the whole stream and adds the two dimensions those studies did not use: container format and file size.
The short version is that people upload exactly what their phones and microphones produce, and almost nothing else. Seven in ten files are MP4. Five in six contain video. Just over half are under a minute long, and three quarters are under 100MB. There is no exotic format tail and no hidden population of giant masters. The demand is mundane, and the mundane is the point: it is what a tool like this has to be built for.
THE SHORT VERSION
Every figure below is measured on VidClean's own uploads. Quote any of them with attribution.
- Seven in ten files are MP4 (10,266 of 14,452, 71.0%). Add MOV and the phone-native pair is 83.6% of everything.
- Five in six files contain video (12,141 jobs, 84.0%), against 2,311 audio-only files (16.0%).
- Just over half are under a minute long (7,500 jobs, 51.9%), and only 3.3% run past an hour.
- Three quarters are under 100MB (10,893 of the 14,286 jobs with a recorded size, 76.2%). Eighteen files were over 2GB.
- The format tail is almost empty: everything outside MP4, MOV, MP3, M4A and WAV is 0.5% of uploads combined.
- The destination split is led by stabilization (21.7% of the 9,735 jobs with a recorded tool), followed by transcription at 15.0% and silence removal at 14.6%.
Window is July 20 to August 25, 2026 (n=14,452). File size is recorded for 14,286 jobs; tool is recorded for 9,735 jobs, from August 6 onward. The three samples are different sizes and are never mixed. See methodology.
SEVEN IN TEN ARE MP4
Of the 14,452 files, 10,266 are MP4. That is 71.0%, and it is not a preference; it is the default of every phone on the market. MP4 is what the camera app writes, so MP4 is what arrives.
Bar length is the share of all 14,452 files. Format is read from the upload filename's extension, which is the container, not the codec. Percentages are rounded.
The second row is the one most people get wrong. MOV is 1,815 files, 12.6%, and it is the iPhone's default container, which makes it the second-most-common thing we see. A phone-heavy audience shows up here as an MP4 and MOV pair that together are 83.6% of the entire stream. Everything designed to process internet video already handles those two; the rest of the work is the 16% of files that are not video at all.
The audio files are MP3, M4A and WAV: 2,300 files between them, 15.9% of uploads. M4A is worth a separate mention because it is the iPhone's default audio container, the same phone that produces all that MOV, and people rarely know the difference between M4A and MP3 when they upload one. The format a file arrives in is a fingerprint of the device that made it, and there are only three devices here: an iPhone, an Android phone, and a podcast recorder.
Outside those five formats, the entire tail is 69 files. MKV is 43, mostly screen recordings from a specific desktop workflow, and WebM is 26, mostly short clips grabbed from the web. Everything exotic that a general video tool is supposed to "handle just in case" is a rounding error in the actual stream.
FIVE IN SIX CONTAIN VIDEO
12,141 of the 14,452 files have a video stream. That 84.0% is the load-bearing number for the whole site: video jobs are the ones that carry encode cost, resolution limits and processing caps, while audio-only jobs are cheap by comparison. A tool that sized itself from a sample of audio files would be off by a factor of five.
84.0%
of uploaded files contain video
12,141 of 14,452 jobs, July 20 to August 25, 2026. The other 2,311 (16.0%) are audio-only.
The audio-only files are not evenly spread across the tools. The cleanup tools in the audio cleanup study saw a third of their uploads arrive as audio, and the transcription tool sees more still, because people record voice memos and podcast drafts and send them straight to processing. The aggregate 16% mixes a small group of tools that are majority-audio with a large group that is nearly all video.
The reverse framing is the more useful one for anyone building this kind of product: five out of every six files bring an ffmpeg video decode with them, even to tools whose name says audio. That is why the audio cleanup study's headline was that two thirds of its uploads were videos. The pattern runs through everything: people reach for a web tool, grab whatever file is on their phone, and send it as-is.
JUST OVER HALF ARE UNDER A MINUTE
7,500 files, 51.9%, are under one minute long. This is the number that keeps reappearing across the series: the stabilization study found 87.7% of its videos under a minute, the silence study found the same shape, and here it is in the full stream, diluted by the podcast-length tail the audio tools attract.
Bar length is the share of all 14,452 files. Duration comes from the media probe at processing time. Percentages are rounded and sum to 100.
The distribution falls off fast and then levels out. A quarter of files sit in the 1-to-5-minute bucket, and everything from five minutes up is single digits. The long end is real but small: 470 files, 3.3%, ran past an hour, and they are overwhelmingly podcast episodes and long voice memos going to the transcription, silence and cleanup tools.
Two forces shape this that the data cannot separate. The first is user behaviour, which the stabilize study already argued is genuinely short: phone clips shot and noticed in the same moment. The second is the product's own caps, because several tools limit file length and people stop uploading past the limit. The honest reading is that the sub-minute majority is behaviour, and the precise share at the long end is partly policy. Neither changes the practical conclusion: a web video tool is a short-file business, and every architecture decision, from processing budgets to queue depth, should start from that.
THREE QUARTERS ARE UNDER 100MB
File size is the dimension nobody asks about and the one that decides how much a job costs to move. Of the 14,286 jobs with a recorded size, 10,893, or 76.2%, are under 100MB. The median upload lands in the 10-to-50MB bucket, which is a two-minute 1080p phone clip or a five-minute podcast MP3.
Bar length is the share of the 14,286 files with a recorded size. Size is the uploaded object's byte count. Percentages are rounded and sum to 100.
Under 100MB is the practical definition of "fits through any pipe". It means the typical job can be fetched from object storage in a second or two, processed, and the output delivered without anyone noticing a transfer. The 16.3% in the 100-to-500MB bucket are the same short videos at higher bitrates, and the 411 files between 1GB and 2GB are mostly long recordings and 4K clips. Only 18 files in five weeks exceeded 2GB.
The size distribution is the quiet confirmation of everything else in this post. Short duration, small size and MP4 are three independent counters that all point at the same object: the modern phone video, a few seconds to a couple of minutes, recorded at default settings, small enough to send anywhere. The tools that win on the web are the ones built for that file.
WHERE THE FILES GO
The last dimension is which tool processed each upload. It was only recorded from August 6, so this slice covers the 9,735 jobs since then, and every share below is of that total. This is the reference table the rest of the series draws on: the stabilize study's n, for example, is the stabilize slice of this same stream.
| Tool | Files | Share |
|---|---|---|
| Stabilize video | 2,110 | 21.7% |
| Transcribe | 1,459 | 15.0% |
| Remove silence | 1,424 | 14.6% |
| Remove background noise | 1,297 | 13.3% |
| Merge videos | 744 | 7.7% |
| Repair audio | 464 | 4.8% |
| Enhance speech | 418 | 4.3% |
| Compress video | 329 | 3.4% |
| Resize video | 328 | 3.4% |
| Extract audio | 328 | 3.4% |
| Clip generator | 214 | 2.2% |
| Editor renders | 181 | 1.9% |
| Every other tool (8) | 430 | 4.4% |
Shares are of the 9,735 files whose tool was recorded. Tool recording started August 6; files from the first 17 days of the window carry no tool. The bottom row pools add subtitles (124), mute (101), crop (58), speed (47), rotate (42), trim (31), convert (17) and GIF (9).
The top four tools take 64.6% of the stream between them, and they are the four the stabilize, cleanup, transcription and silence studies have already written about. The long tail of small tools matters for a different reason: it is where the site's own tests and quality work live, which is why the usage studies net out test traffic before quoting shares.
Stabilization being first is worth a second look, because it is the only tool in the top five that changes the picture rather than cleaning it up. People arrive with shaky footage they cannot un-shoot, and they are willing to wait for a two-pass process on it. That is the single most useful fact in this table for anyone deciding what a free video site should be built around.
METHODOLOGY
These are aggregate counters, not per-file records. When a job completes, the backend increments a handful of running totals against one shared key: the format read from the filename extension, whether the probe found a video stream, one coarse duration bucket, one size bucket, and, since August 6, the tool. No filenames, no hashes, no timestamps, no user identifiers and no per-file rows exist anywhere. No individual upload can be reconstructed from the data, and no two dimensions can be crossed: the data cannot say whether the under-a-minute files are also the under-100MB ones, only that both are true of the stream as a whole.
The window is July 20 to August 25, 2026, when the counters shipped and when this pull ran. The job count, format, media kind and duration families cover all 14,452 jobs. The size family covers 14,286 of them, 98.9%, because a small set of jobs complete without a size recorded; every size share above is of that smaller total. The tool family covers 9,735 jobs, from August 6 onward, because earlier jobs never recorded it; every tool share above is of that total.
VidClean's automated production tests run through the same pipeline, roughly one short MP4 per tool per test run, and they inflate the sub-minute, small-size and MP4 buckets. The test volume is small next to 14,452 jobs, so every headline figure above survives the exclusion; the exact shares are the gross ones, with that grain of salt. The same disclosure applies to the per-tool studies in this series, which is why they net out test traffic before quoting shares.
The full counter dump behind this post is available as a CSV file, with the applicable sample size on every row.
This data is licensed CC BY 4.0. Feel free to reuse the numbers with credit to VidClean.
LIMITATIONS
This is VidClean's user base, not a random sample of anyone's video. People who upload to a free web tool already believe their file needs fixing, which skews the stream toward the phone-clip case and away from finished work. The proportions here describe demand on one site, not the world's footage.
The counters cannot be crossed. The two strongest findings in the post, short duration and small size, are measured independently, and the claim that they are the same files is an inference from the phone-clip story, not a join the data supports.
Format is the extension, which is the container, not the codec. An MP4 can hold H.264, H.265 or AV1, and this dataset cannot tell them apart. For the argument here, that the container is what a pipeline must accept, the distinction does not matter; for a codec-level analysis it would.
The duration distribution is partly policy-shaped. Several tools cap file length, so the long end is truncated by the product. The sub-minute majority is too large to be an artifact of any cap, but the exact share at the top of the range would look different on uncapped tools.
And it is five weeks in summer. Long enough for 14,452 jobs, short enough that seasonal drift is real, and the counters keep running. When the stream says something new, we will write it down.
IF YOU ARE BUILDING FOR THIS TRAFFIC
The practical read is short. Build for the MP4 phone clip, a minute long and well under 100MB, because that is five of every six files. Optimize the encode path before the upload path, because video decode is where the cost lives. And treat every other format as a fallback to handle gracefully rather than a population to serve: the exotic tail is 0.5% of uploads.
VidClean is a solo project. Questions about the data or the methodology are welcome at hello@vidclean.net.
Related reading
We Analyzed 2,039 Shaky Videos. Almost All of Them Were Under a Minute.The first study in this series, and the tool slice of this same stream: 87.7% of stabilized videos ran under a minute, one upload in seven was 4K, and stabilizing 4K costs 3.3 times what 1080p costs.