Skip to content
Zbor Palilula Niš Niš Fortress
AI analysis

Could AI have made this fire?

One fake video — yes, in half an hour. But not a dozen videos from different angles showing the same fire in the same place, at the same time, with the same smoke, light, crowd and firefighters — posted minutes apart from independent accounts.

Conclusion
Between the first two angles (S-01, position A, and S-02, position B) 2 min passed, and 51 min 37 s until the third position (S-07). According to available data on AI tools, a consistent set of such videos takes at least 2–4 hours in the best case, and realistically 1–3 days. Fabrication within the observed time is therefore practically impossible.
Observed posting windows vs. estimated fabrication time
Logarithmic time axis: each grid step is ten times longer. Dots are measured from post IDs; bars are estimates.
First two angles: A, B (S-01 → S-02)
2 min
Three angles: A, B, C (S-01 → S-07)
51 min 37 s
Until the firefighters video (S-01 → S-10)
70 min
One fake clip
20–30 min
Consistent set — best case
2–4 h
Consistent set — realistic
1–3 days
Consistent set — VFX grade
1–4 weeks
Why it is hard

Obstacles to a consistent AI fabrication

  1. Each clip is processed separately: every tool “invents” its own flame shape, height, smoke and fire progression. Matching 10+ videos from different angles is manual work. [source]
  2. Physics: models do not understand combustion and fluid dynamics (Physics-IQ; VideoPhy-2 — best model 47.7% on the hard subset; PhyGenBench includes combustion). [source]
  3. Smoke and wind: no tool takes a shared wind direction; smoke direction and density must be matched by hand in every clip.
  4. Light: a real fire casts flickering orange light on facades, faces and surroundings — at the same intensity in every video at the same moment.
  5. People: the crowd reacts, looks and points at the fire, and firefighters arrive (S-10). Inserting fire does not change the behaviour of people in the base footage.
  6. Length: S-08 runs continuously for 56 s — longer than the maximum editing tools process in one pass (5–30 s); joins would show as jumps in the flame shape.
  7. Sound: crackling, sirens and voices would have to be made separately for every clip.
  8. There is no published case or tool that produced 10+ mutually consistent videos of one event from different angles within an hour or two.
State of the art, October 2026

Leading AI video models

The key question is not “can AI make fire” but whether it can edit existing footage and do so consistently across several recordings from different angles. No tool does that across separate recordings.

ModelMax clip per passGeneration time Edits real footage?Multi-angle consistencyWatermark / status
explained below
Google Veo 3.1 4–8 s (extension up to 148 s, 720p)11 s – 6 min per generationNo — creates new video from text/imageNoSynthID (invisible) on every output
Gemini Omni Flash (maj 2026.) 10 snot disclosedYes — conversational video editingContext within one session, not multiple real camerasSynthID
OpenAI Sora 2 10–25 s—No — video uploads blockedNoVisible moving watermark + C2PA. Discontinued: app 26 Apr 2026, API 24 Sep 2026.
Runway Aleph / Aleph 2.0 5 s (Aleph) / 30 s (Aleph 2.0, 1080p)≈ 220 s per job; often several attemptsYes — adds objects and relights real footageWithin one multi-cut video; weak across separate recordingsUsage policy bans deceptive content
Kling 3.0 / O1 Edit 3–15 s≈ up to 2 minYes (O1 Edit)Only between its own generated shotsInvisible watermark + public detector
Luma Ray3 Modify / Ray 3.2 up to 10–18 snot specifiedYes (video-to-video)Character reference, not sceneNo C2PA found
ByteDance Seedance 2.5 up to 30 s—Partly (modify/extend)Not across real cameras—
Wan 2.2 / VACE (otvoreni kod) ≈ 5 s, 480/720p5 s at 720p: up to 9 min on an RTX 4090Yes (VACE)NoLocal — no enforced watermark
Watermark / status

How AI tools mark their videos — and how to check

The “Watermark / status” column in the table above says two things: whether a tool embeds a mark in its videos showing they were made by artificial intelligence (a watermark), and whether the tool is still available at all (e.g. Sora 2 was shut down). There are four situations:

Invisible watermark
Google Veo, Gemini (SynthID) · Kling
What it is
A signal woven into the pixels (and the sound) of every frame — invisible to the eye. According to Google, SynthID “remains detectable even when the content is shared or undergoes a range of transformations”; Kling says its watermark stays “traceable through compression, editing, and resharing”.
How to check
Google: upload the video to the Gemini app and ask “Was this generated using Google AI?” — Gemini scans the image and sound and states which segments contain SynthID (files up to 100 MB and 90 seconds). Kling: upload the video at kling.ai/detection.
Limit
Each detector only recognises content from its own company’s tools.
Visible watermark
OpenAI Sora 2
What it is
A logo that moves around the frame of every generated video, so anyone can see it.
How to check
With the naked eye — watch the whole video.
Limit
Tools for removing it appeared within days of launch, so the absence of a logo proves nothing. Status: Sora 2 was shut down in 2026.
Content Credentials (C2PA)
industry standard; used by some cameras and tools
What it is
A cryptographically signed record inside the file itself: who made the content, with which tool, and what was changed — “like a nutrition label for digital content”.
How to check
Upload the file at contentcredentials.org/verify.
Limit
Social networks strip or ignore it on upload: the Washington Post uploaded an AI video to 8 social apps and found that major platforms do not use this standard to flag it. A video downloaded from X or Instagram normally has no C2PA, whether it is real or not.
No mandatory watermark
Wan 2.2 / VACE run locally; for Runway and Luma we found none
What it is
A video made with an open model on one’s own computer does not have to carry any mark at all.
How to check
There is no watermark to check.
Limit
That is why no detector can confirm that a video is real.
What a check proves — and what it does not
  • Watermark found → strong evidence that the video (or part of it) was made with that AI tool.
  • No watermark found → only means it was not made with that company’s tools — not that the video is real.

A watermark can therefore prove that something is fake, but not that something is real. That is why this site relies on evidence that does not depend on watermarks: server posting times, the agreement of several camera positions and the intersection of their sightlines, the firefighters’ and prosecutor’s confirmations, and the soot found by the police.

We have not run the videos through detectors: copies downloaded from social networks are re-encoded files, and a negative result would prove nothing. Whoever claims the videos are AI can show it — for example by finding a SynthID watermark in an original.

Check it yourself
  1. Get the original file from the author — not a screen recording or a copy downloaded from social media.
  2. Google tools: in the Gemini app upload the video and ask “Was this generated using Google AI?” (up to 100 MB and 90 s).
  3. Kling: kling.ai/detection.
  4. Certificate of origin (C2PA): contentcredentials.org/verify.
  5. Visible watermark: watch the whole video for a moving logo.
Estimate

How long fabricating 10–12 consistent videos would take

All values in the table are our estimate, derived from data on the tools (time per job, clip length, number of retries) and typical VFX work durations.

StepBest caseRealisticVFX grade
Filming base footage from 10–12 positions10–20 min (people pre-positioned)45–90 minhalf a day
Inserting fire, with retries (2–4 min per job, 3–5 attempts)40–60 min, in parallel5–15 hsome angles never converge
Matching flames, smoke, light and timing across clips30–60 min (fails side-by-side)3–8 hdays (simulation + camera tracking)
Compositing and sound10–20 min per clip1–4 h per clipseveral days per shot
Posting via 10+ accounts and outletsminutes — only if every account is controlledindependent outlets with their own reporters would have to be complicitsame
Total for a consistent set of 10–12 clips≈ 2–4 h≈ 1–3 days≈ 1–4 weeks
What if it was made in advance?

That does not explain the evidence either

  • The videos would have to match the exact time, place, weather, crowd, flags and fireworks of that rally — and nobody disputes the fireworks and flares (M-01, I-01).
  • Soot on the rampart was found by the police investigation (P-03, P-06), and N1’s reporter saw “the remains of where it burned” (P-02).
  • Firefighters confirmed the fire to N1 the same evening (M-05); the prosecutor qualified it as a criminal offence (P-07).
  • The sightlines of all 18 videos intersect at one point on the southern rampart of the Fortress, within ±15 m — from different places they all film the same flames. Pre-made videos would also have to hit exactly that spot from every position (map).
  • Južne vesti had its own reporter and photos from the scene (M-02).
About detectors. Google SynthID and the Kling detector only recognise content from their own tools. Social networks strip C2PA labels, so their absence proves nothing. That is why we rely on what cannot be faked: server posting times, the agreement of several independent angles, and physical traces.