This is not an accuracy score
A returned result means the system produced and persisted a non-Unknown movie candidate. The production table does not contain human ground-truth labels for these attempts. This study therefore measures operational result return and speed, not whether every returned title was correct.
Movie identification systems face very different evidence depending on what a person submits. A written memory may contain rich plot clues but one mistaken detail. An image provides direct visual evidence but may be generic or compressed. A public video URL can add metadata, transcript, thumbnail, and comment context, but only when the source exposes them.
We analyzed one complete calendar month of direct-finder production telemetry to establish a transparent baseline. Text descriptions and YouTube URLs returned candidates most often in this cohort. Social URL-only attempts were far less likely to return a candidate when their platforms exposed no usable evidence.

Results by input type
Across 9,713 eligible model runs, 4,751 returned a candidate and 4,962 ended as Unknown. The overall result-return rate was 48.91%.
| Input | Attempts | Returned | Unknown | Return rate | P50 | P90 |
|---|---|---|---|---|---|---|
| Text description | 3,244 | 2,269 | 975 | 69.94% | 2.00 s | 6.24 s |
| Image upload | 2,271 | 1,072 | 1,199 | 47.20% | 5.36 s | 7.97 s |
| YouTube URL | 1,504 | 1,018 | 486 | 67.69% | 4.19 s | 8.81 s |
| Instagram URL | 928 | 327 | 601 | 35.24% | 4.47 s | 8.31 s |
| Facebook URL | 755 | 0 | 755 | 0.00% | 4.70 s | 7.48 s |
| TikTok URL | 453 | 3 | 450 | 0.66% | 3.41 s | 4.95 s |
| Other URL | 383 | 43 | 340 | 11.23% | 1.88 s | 5.49 s |
| X / Twitter URL | 175 | 19 | 156 | 10.86% | 3.08 s | 6.10 s |
| All eligible runs | 9,713 | 4,751 | 4,962 | 48.91% | 3.80 s | 7.48 s |
What the differences do and do not show
Text descriptions returned a candidate in 69.94% of attempts and YouTube URLs in 67.69%. Instagram URLs returned candidates in 35.24% of attempts, while TikTok and Facebook URL-only attempts were close to zero. That does not prove that platform choice caused the difference. Public sources expose very different amounts of metadata and visual evidence, and users submit different material through each path.
Image uploads returned a result in 47.20% of attempts. Their median latency was 5.36 seconds versus 2.00 seconds for text descriptions, consistent with the additional work required by visual analysis. Again, the traffic was not randomized, so the difference is an observed property of this cohort rather than an isolated treatment effect.
Recorded model agreement
Of the 4,751 returned results, 2,424 recorded title agreement from at least two providers: 51.02% of returned candidates. VidScio normalizes title punctuation, articles, and close title variants when grouping provider answers.
Agreement is useful evidence, but it remains distinct from correctness. Providers can share a misleading caption, a famous visual association, or the same incorrect clue. Some valid production paths also finalize a candidate without recording a two-provider agreement, so the remaining 48.98% should not be read as a measured disagreement rate.
Methodology
Study window
June 1, 2026 at 00:00 UTC through July 1, 2026 at 00:00 UTC. Aggregates were queried on July 30, 2026.
Unit of analysis
One persisted terminal attempt from the direct streaming finder. The table records one final status per handled request when telemetry insertion succeeds; it is not a count of unique people or unique movies.
Cohort construction
The month contained 12,229 direct-finder rows. We excluded 2,516 cache hits so previously stored answers would not be credited as fresh model runs, leaving 9,713 eligible attempts. Deleted rows and correction retries were excluded; this direct-finder cohort contained no correction retries.
Input groups
Image upload takes precedence, followed by YouTube URL. Remaining URLs are classified by domain pattern into Instagram, Facebook, TikTok, X/Twitter, or other URL. A row with no URL pattern is classified as a text description. Only aggregate category counts were selected; no prompt or URL was reviewed.
Result-return rate
returned non-Unknown result / all eligible model runs
A returned result requires a non-empty title other than Unknown and a persisted movie record. All status and movie-record fields agreed in this cohort.
Latency
Server elapsed time from request handling to the terminal result, measured in milliseconds. P50 is the median; P90 is the value at or below which 90% of attempts completed. Background analytics and image persistence are not part of the measured response path. No eligible row lacked latency data.
Agreement
A returned result is counted as having recorded agreement when its stored consensus list contains at least two provider names whose normalized titles matched under the production title-comparison rules.
Limitations
- This is observational production traffic, not a randomized or controlled test set.
- There are no human ground-truth labels for the 9,713 attempts.
- A missing result is not necessarily a model failure; the source may be private, inaccessible, or too ambiguous.
- A returned title may still be wrong, and model agreement can also be wrong.
- Input groups differ in content, user intent, and difficulty, so cross-group differences are not causal estimates.
- The telemetry does not flag staff tests, repeated users, or duplicate inputs, and the analysis does not attempt to infer them.
- Results describe the June 2026 production system and should not be generalized to later model or pipeline versions.
Privacy statement
This publication uses aggregate counts and latency percentiles only. No user identifiers, prompts, URLs, uploaded images, movie titles, IP addresses, or country-level records were selected for publication. No raw identification record was reviewed to write this article.
The practical baseline
The useful benchmark is not a single inflated accuracy number. It is a reproducible operational baseline: 48.91% of fresh direct-finder attempts returned a candidate, the median attempt completed in 3.80 seconds, and 51.02% of returned candidates recorded at least two-provider title agreement.
Future accuracy research requires a separately designed, human-labeled test set with known source titles and balanced input difficulty. Until then, keeping result return, speed, agreement, and user-confirmed correctness separate is the more honest way to describe system performance.
Try the direct finder
Submit a scene description, supported public URL, or screenshot and inspect the returned candidate against the source evidence.
Identify a movie