Principles
Both studios follow three rules. Independent methods, combined. No single technique is proof; the tools report each method separately as no signal, weak or notable and leave the conclusion to the examiner, who should look for two or more methods agreeing on the same region. Documented parameters. Every threshold below is printed in the report so another examiner can reproduce the result. Local processing. Hashes are computed over the original bytes before decoding and nothing leaves the browser.
Verdict levels mean: none = the statistic stayed within the image's own norm; weak = scattered outliers or a small cluster, usually explained by content; notable = a contiguous region behaves differently from the rest of the image by the stated margin; info = descriptive, not scored.
File hashes
SHA-256, SHA-1 and MD5 are computed over the exact bytes of the dropped file (Web Crypto for SHA, an in-page implementation for MD5) before the image is decoded or the video demuxed, so the report describes a specific file that anyone can re-hash. For videos over 1.5 GB hashing is skipped to protect memory and the report says so.
Error Level Analysis (ELA)
Idea. Re-saving a JPEG at a known quality changes regions with a different compression history by a different amount (Krawetz, A Picture's Worth, Black Hat 2007).
Computation. The image (downscaled to at most 1400 px on the long side) is re-encoded as JPEG at the chosen quality (default 75, range 50–95) using the browser's encoder; the absolute per-pixel RGB difference is averaged into 16 px blocks. Because edges and text legitimately produce high error, each block's error is divided by 1 + (mean gradient magnitude / 6). Blocks above max(median + 6·MAD, 2.2·median) of the normalised error are "hot"; textured blocks below 0.35·median are "cold" (content that was already heavily compressed before being pasted in). Clusters are measured with 4-connectivity (hot) or gap-tolerant connectivity (cold).
Verdicts. ≤2 hot blocks and no cold cluster: none. Hot cluster < 6 blocks: weak. Hot cluster ≥ 6 or cold cluster ≥ 6: notable.
Limits. Weak on heavily recompressed or resized images, screenshots and social-media downloads; sky and flat areas are always low-error and are excluded from the cold test by the texture requirement.
JPEG ghost sweep
Idea. A region that was once saved at quality q shows a local minimum of re-compression error at q (Farid, Exposing digital forgeries from JPEG ghosts, IEEE TIFS 2009).
Computation. At native resolution (centre crop of at most 2200 px if larger, because resampling destroys the 8×8 grid), the image is re-encoded at qualities 40, 45, … 100 and the mean error per 16 px block recorded. A block has a dip at q when its error at q is lower than at both neighbouring qualities by at least 12% (25% below quality 65, where curves are noisier). The file's own quality, estimated from its luminance quantization table against the IJG scale (Kee, Johnson & Farid, Digital image authentication from JPEG headers, 2011), is excluded with a ±7 margin because every block dips there. For each remaining dip quality, blocks within ±5 are clustered with one-block gap tolerance; the quality with the largest cluster is the reported ghost.
Verdicts. Largest cluster < 8: none. 8–19: weak. ≥ 20: notable. More than 60% of blocks with a foreign dip: info — uniform double compression (the whole image was re-saved), not localised editing.
Limits. Pasted regions must be roughly 200 px or larger and aligned to the 8×8 grid; PNG and WebP sources have no own-quality reference, so the mode is used instead when it covers more than 30% of blocks.
Copy-move (clone) detection
Idea. Clone-stamp and healing tools copy blocks within the same image; duplicated blocks share a common displacement (Fridrich, Soukal & Lukáš, Detection of copy-move forgery in digital images, DFRWS 2003).
Computation. On a grayscale copy at most 900 px wide, 16 px blocks are taken every 4 px; blocks with variance below 60 (flat) are skipped. Each block is described by its 4×4 grid of sub-block means, brightness-normalised (mean removed, divided by 2, rounded) so recoloured clones still match. Features are sorted lexicographically and each block is compared with its next five neighbours; pairs within L1 distance 4 whose displacement is at least 32 px vote for that displacement. A displacement counts when it has at least max(20, 0.4% of candidate blocks) votes and at least 60% of its source blocks lie within two steps of another source block (coherence). Up to six displacements are drawn in distinct colours on both source and destination.
Verdicts. No qualifying displacement: none. Fewer than 40 matched pairs: weak. 40 or more: notable.
Limits. Rotated, rescaled or heavily blended clones are not matched; strongly periodic textures (brick, tiles, text) can produce legitimate matches, which the coherence rule only partly removes.
Noise variance map
Idea. Sensor noise is roughly uniform across an untouched capture; spliced content, generative fill and local denoising change it (Mahdian & Saic, Using noise inconsistencies for blind image forensics, 2009).
Computation. The residual between each pixel and the mean of its 8 neighbours is averaged per 8 px block. Because texture inflates the residual, the noise floor is estimated only on blocks whose gradient energy is at or below the image median; outliers are flat blocks above median + 3.5·MAD (noisy) or below min(median − 3.5·MAD, 0.3·median) (very smooth). A noise floor below 0.4 is reported as info: rendered graphic, upscaled or generative image.
Verdicts. Under 2% outliers or largest cluster < 4: none. Largest cluster < 12 or noisy cluster < 6: weak. Otherwise notable.
Luminance gradient
Central-difference gradient of luminance; direction is mapped to hue and magnitude to brightness. A visual aid for lighting consistency (an inserted object lit from a different direction shows a different hue pattern at its edges). Not scored.
Bit planes and chi-square
The least-significant bit of each channel is shown as a binary image. The pairs-of-values chi-square statistic (Westfeld & Pfitzmann, Attacks on steganographic systems, 1999) is computed on the blue channel: sequential LSB embedding equalises the counts of values 2k and 2k+1. Chi-square per degree of freedom below 0.7 is weak, below 0.3 notable, with at least 40 degrees of freedom. Only meaningful for lossless formats; JPEG decoding destroys LSB payloads.
EXIF thumbnail comparison
The camera writes a small thumbnail into EXIF at capture; editors often leave it untouched. The embedded thumbnail is decoded and correlated with the image downscaled to the same size. Correlation above 0.95: none; 0.85–0.95: weak; below 0.85: notable. Some cameras and transfer utilities write differently processed thumbnails, so this is corroborating evidence.
Metadata and file-structure audit
Parsed. JPEG segments (JFIF, EXIF/TIFF IFD0, ExifIFD, GPS, Interop, IFD1 thumbnail, XMP and extended XMP, ICC, Photoshop IRB with IPTC, Adobe APP14, APP11 JUMBF/C2PA, COM, DQT, SOF); PNG chunks (IHDR, tEXt/iTXt/zTXt including XMP and generator parameters, eXIf, caBX/jumb); WebP chunks (EXIF, XMP, C2PA); TIFF-based RAW (IFD0 with DNG tags, SubIFDs, embedded previews); HEIF/AVIF EXIF and XMP by scanning.
Derived. Decimal GPS; JPEG quality from the quantization tables and whether they match the IJG scale; progressive flag and chroma subsampling; bytes after the end-of-image marker; count of embedded JPEG headers; XMP xmpMM:History steps; IPTC DigitalSourceType and generator mentions; mismatches between EXIF dimensions and actual pixels and between capture and modification dates.
Findings are labelled info, weak or notable: an editor in the Software tag, an XMP edit history or a generative-AI marker is notable; stripped EXIF, date mismatches and appended data are weak; everything else is descriptive. Metadata is trivially editable, so its absence proves nothing and its presence is a claim to be checked.
C2PA / Content Credentials
The studios detect the JUMBF/C2PA container (APP11 in JPEG, caBX in PNG, C2PA chunk in WebP, uuid box in MP4; see the C2PA Technical Specification 2.4 (April 2026)) and report its presence and size. Validating the signature chain requires the C2PA trust list and is left to the AI provenance & C2PA hub and the official verifier. A valid credential is a positive provenance signal; its absence proves nothing, since most cameras and platforms still strip or never write one.
RAW files
DNG, CR2, NEF, ARW, PEF, ORF, RW2 and other TIFF-based RAW files are parsed for their full TIFF/EXIF metadata. The largest complete JPEG stream inside the file (the camera's embedded preview) is extracted and analysed with every pixel-level method; the report states this and that sensor data was not demosaiced. RAF and CR3 are handled by preview extraction only.
Video Forensics Studio
Container audit. MP4/MOV: ftyp brand and compatible brands, mvhd creation/modification time and duration, per-track tkhd (dimensions, rotation matrix), mdhd (timescale, duration, language), hdlr, stsd (codec four-character codes, compressor name, audio channels and rate), stts (frame count and average rate), stss (keyframes), elst (edit lists), udta/ilst (©too encoder, ©xyz GPS, device make/model, Apple and Android keys), XMP uuid box, C2PA box, top-level box order (fast-start), mdat count, fragmentation and trailing bytes. Matroska/WebM: DocType, Info (muxing/writing app, date, duration), Tracks (codec, dimensions, default duration, audio) and Tags.
Findings. Encoder strings are matched against lists of editors, platforms and generative-video services; zeroed (1904) timestamps, fast-start layout, fragmented files, multiple mdat boxes, trailing data, non-standard frame rates, edit lists, irregular frame timing, missing audio and video/audio length mismatches are reported with their usual explanations.
Temporal analysis. 24, 48 or 96 frames sampled at even intervals and downscaled to 160 px width. Metrics per step: mean absolute luminance difference, block-matching motion magnitude (8 px blocks, ±2 px search), mean luminance and a duplicate flag when the mean difference is below 1.2. Spikes beyond mean + 2.5 SD are flagged as possible cuts or inserts; motion spikes as possible warps; duplicate runs as frozen frames.
Per-frame analysis. The chosen frame is re-grabbed at up to 480 px and analysed with ELA at the selected quality, a Laplacian noise residual and 8×8 block-grid energy. Every inspected frame is stored for the report at reduced size.
Frame export. Rendered at native resolution to PNG; the SHA-256 of the PNG is shown and recorded.
Limits. Sampled, not frame-accurate; decoded frames only (no codec-level structure); HEVC and some MOV variants depend on the browser's codecs; the optional AI detector (onnx-community/deepfake-detection-model via Transformers.js) is a model's opinion and is reported as such.
References
- N. Krawetz, A Picture's Worth: Digital Image Analysis and Forensics, Black Hat Briefings, 2007.
- H. Farid, Exposing digital forgeries from JPEG ghosts, IEEE Transactions on Information Forensics and Security 4(1), 2009.
- J. Fridrich, D. Soukal, J. Lukáš, Detection of copy-move forgery in digital images, Digital Forensic Research Workshop, 2003.
- B. Mahdian, S. Saic, Using noise inconsistencies for blind image forensics, Image and Vision Computing 27(10), 2009.
- E. Kee, M. K. Johnson, H. Farid, Digital image authentication from JPEG headers, IEEE Transactions on Information Forensics and Security 6(3), 2011.
- A. Westfeld, A. Pfitzmann, Attacks on steganographic systems, Information Hiding Workshop, 1999.
- W. Wang, H. Farid, Exposing digital forgeries in video by detecting double MPEG compression, ACM Multimedia and Security Workshop, 2006.
- Scientific Working Group on Digital Evidence, Best Practices for Image Authentication, SWGDE 18-I-001, version 2.0, 3 March 2025.
- Coalition for Content Provenance and Authenticity, C2PA Technical Specification, version 2.4, April 2026.
Engine source: forensics-engine.js and video-forensics-engine.js are served unminified and can be read in your browser's developer tools. Version history: changelog.
Frequently Asked Questions
Why does the tool say "weak" so often?
Because most real images contain content that triggers one method legitimately: edges and text in ELA, skies in noise maps, repeating textures in copy-move. "Weak" means a small or scattered signal that the examiner should look at and will usually explain by content. Only coherent regions that several methods agree on should change a conclusion.
Can I change the thresholds?
The ELA quality is adjustable and each method can be switched off. The statistical thresholds are fixed so that results are comparable between examinations and reports; they are listed on this page and in every report, and the engine source is readable in the browser.
Why is the ghost sweep run at native resolution when everything else is downscaled?
JPEG compression history lives on the 8×8 pixel grid. Resampling the image destroys that alignment and erases ghosts, so the sweep uses native pixels (a centre crop for very large images) while the other passes use a downscaled copy for speed.
Does a clean result mean the image is authentic?
No. It means none of the methods found a trace. Recompression, resizing, screenshots and generative models can leave no detectable trace, which is why the report wording is "no signal" rather than "authentic".
How were the thresholds chosen?
They were calibrated so that untouched camera photographs (including a 3264×2448 phone capture and re-saved originals) return no signal on every method, while cloned regions, pasted elements with a different compression history and LSB-embedded files are flagged. The calibration set and its results are described in the changelog entry for engine 2.0.0.