How to Capture Web Source Notes & Provenance When Archiving Online Video
In digital preservation and open-source intelligence (OSINT), a video file without documented provenance is considered an evidentiary orphan. Years from now—or even weeks from now—when you review a saved video clip, you may ask yourself: Who originally published this? At what exact UTC hour was it posted? Was it a re-upload of an older event? What commentary or hashtags accompanied the media?
When you download a video file from Twitter (X), the downloaded MP4 file only contains the raw audio/video frames and generic encoder metadata. The post text, author handle, replies, geolocation hints, and platform timestamps are left behind on the web page. This guide provides a battle-tested protocol for capturing and storing comprehensive source notes alongside your media archive.
The Critical Need for Verifiable Digital Provenance
Digital provenance refers to the documented chronology of the origin, custody, and contextual ownership of a digital asset. When saving social video, capturing provenance is essential for three primary reasons:
- Legal & Evidentiary Integrity: In legal proceedings or journalistic reporting, undocumented clips can easily be challenged as decontextualized, misattributed, or doctored deepfakes. Detailed capture logs establish chain of custody.
- Guarding Against Deletion: Authors frequently delete posts after receiving unexpected viral attention or legal pushback. Once the original tweet is removed, your local capture note becomes the sole surviving record of what was originally claimed.
- Disinformation Protection: Re-uploaders routinely take old footage from past conflicts or sporting events and present it as breaking news. A proper source note records the earliest discovered instance of the clip.
The Essential 8 Metadata Fields to Record
Whenever you archive an important social media video, document the following eight data points immediately:
| Field | Description | Example |
|---|---|---|
| Source URL | Canonical permalink to the original post | https://x.com/space_desk/status/192837465 |
| Author Handle & ID | Screen name and immutable platform user ID | @space_desk (User ID: 849201948) |
| Post Timestamp (UTC) | Exact publication time extracted from page DOM | 2026-09-23T14:15:30Z |
| Capture Timestamp (UTC) | When the video file was saved locally | 2026-09-23T16:48:10Z |
| Verbatim Caption Text | Exact post copy, hashtags, and mentions | "Falcon Heavy booster separation viewed from..." |
| Direct CDN URL | Raw video link on platform content delivery network | https://video.twimg.com/amplify_video/...mp4 |
| Cryptographic Hash | SHA-256 digest of the downloaded MP4 file | a1b2c3d4e5f6... (64 hex characters) |
| Third-Party Archive URL | Wayback Machine or Archive.today snapshot link | https://web.archive.org/web/20260923... |
Recommended Capture Tools: WARC, JSON & Snapshots
Depending on the scope of your archival workflow, several free tools can streamline source note documentation:
- Archive.today / Wayback Machine: Before or immediately after downloading the video, submit the post URL to
https://archive.todayandhttps://web.archive.org. This creates an independent, globally verified third-party snapshot of the post interface. - SingleFile (Browser Extension): An open-source extension that packages the entire web page (including HTML, styles, fonts, and inline images) into a single, tamper-evident
.htmlfile. - ExifTool: An open-source command-line tool that can embed source URLs directly into the MP4 container's comment metadata atom without re-encoding:
exiftool -comment="Archived from https://x.com/username/status/123456789 on 2026-09-23" saved_video.mp4
Creating Standardized .nfo and .json Sidecars
For maximum long-term durability, store your notes in a plain-text sidecar file sharing the exact same basename as your video:
20260923_SpaceDesk_FalconSeparation.mp420260923_SpaceDesk_FalconSeparation.json
By pairing files using uniform basenames, any future asset management software or command-line script can automatically parse and index your entire video library.
{
"schema_version": "1.0",
"provenance": {
"platform": "Twitter/X",
"source_url": "https://x.com/space_desk/status/192837465",
"author_handle": "@space_desk",
"published_utc": "2026-09-23T14:15:30Z",
"captured_utc": "2026-09-23T16:48:10Z"
},
"media_integrity": {
"filename": "20260923_SpaceDesk_FalconSeparation.mp4",
"sha256": "4f53cda18c2baa0c0354bb5f9a3ecbe5ed12ab4d8e11ba873c2f11161202b945",
"resolution": "1920x1080",
"duration_seconds": 45.2
},
"content": {
"caption": "Falcon Heavy booster separation viewed from high altitude cameras.",
"hashtags": ["#SpaceX", "#FalconHeavy"]
}
}
Frequently Asked Questions
Why not just save a screenshot of the tweet?
A screenshot is a useful supplementary visual record, but it cannot be easily parsed by automated databases, search engines, or command-line tools. Text-based sidecar files (JSON or NFO) allow instantaneous full-text search across thousands of files.
What is the best way to extract the exact UTC timestamp of a tweet?
On desktop, hover your mouse over the relative timestamp (e.g., "2h ago") and right-click to inspect the DOM element. The HTML <time datetime="..."> attribute contains the official server UTC timestamp string.
Will sidecar JSON files cause compatibility issues on external hard drives?
Not at all. JSON files are plain UTF-8 text files and work seamlessly across Windows, macOS, Linux, and all external drive filesystems (exFAT, NTFS, APFS).