Skip to main content

Synthetic Media Markers: How Digital Signatures and Watermarks Secure Content Provenance

Generative AI capabilities has led to influx of synthetic content on the web. It is becoming hard to distinguish between authentic human content and AI generated content. The need for content provenance has let to synthetic media markers being actively used by companies.

Synthetic Media Markers: How Digital Signatures and Watermarks Secure Content Provenance

Synthetic media markers are technical indicators embedded in AI generated content like text, images, audio and video. They identify the origin of content as well as track its edit history. They can be looked at as “digital nutrition labels” that help platforms, search engines and users identify good quality content, combat deepfakes and misinformation.

Types of Synthetic media markers

There are essentially two types of synthetic media markers:

1. Cryptographic Metadata:

Cryptographic metadata is a digital signature attached to the file information of AI generated and AI modified content. This technical information within the digital signature is called the manifest or content credentials.

The industry standard for tracking synthetic media is Coalition for Content Provenance and Authenticity (C2PA). It is actively used by companies like OpenAI, Google, Adobe and Microsoft. C2PA ensures that the content generated and uploaded for usage by the people, complies with the required industry standards.

Content credentials stores timestamps, software and model version of the generative AI tool used, prompt logs and edit histories. The manifest is signed using asymmetric private key owned by the AI provider (eg. Claude or OpenAI) and anyone viewing the content can use the public key to see the content credentials (eg. LinkedIn marks CR on AI generated content making it easy for the public to view).

C2PA enforces content transparency. Eg. The EU AI Act requires companies to legally label synthetic content. However, this is not a foolproof solution. This metadata can easily be stripped from the content. For example, if someone screenshots an image or uploads it to platforms like Canva to get a new artwork altogether.

2. Invisible Pixel / Audio Watermarks (Steganography):

Steganography conceals a message, a file, a pattern or a watermark within the very core of the generated media file. Example of strategies used for these embeddings are modifying subtle noise patterns within AV files, modifying audio wave frequency or pixel frequencies. These watermarks are undetectable by humans.

SynthID is a digital watermarking technology developed by Google’s DeepMind. Undetectable by humans, they can be detected by the SynthID technology. These synthIDs are designed to withhold when the AI generated content is modified. 

Example, AI generated visual content like image or video has watermarks which hold up during modifications like cropping, screenshots, adding filters, changing frame rates etc. Audio generated content has a watermark which is inaudible to the human ear. It cannot be modified by adding noise, mp3 compressions or change in speed. AI generation of text happens token by token, each word is assigned probability score based on its probability of coming after a particular word. SynthID generates a watermark based on these probability scores.

SynthID detector can be used to identify the watermarks in AI generated content having SynthIDs.

Difference between cryptography and steganography: Cryptography relies on asymmetric compression, using a private-public key. It protects identity and makes it verifiable. Steganography binds origin markers directly into the content, concealing the evidence of provenance.

Modern provenance architecture: It uses both C2PA and SynthID together in a robust provenance system. Cryptography validates authenticity and detailed edit histories when the metadata is intact while Steganographic watermarks become temper proof anchors pointing back to cryptographic registry even if the metadata is removed.

Many companies still use visible AI disclosure markers and overlays e.g. on-screen watermarks, semi-transparent logos, burn-in icons. These are instantly visible to the human eye without specialised software. There was a point when Gemini generated media files had clear visible gemini logo in them. These disclosure markers can easily be cropped out or painted over.

Importance of synthetic media markers:

Synthetic media markers are important because they provide authenticity and verification when the amount of deepfakes and disinformation on the internet is a lot. When looking for information on the web, newsrooms, media platforms and users are able to confirm whether the information can be trusted or not, whether a video clip is real or generated by a software.

Synthetic media markers also play an important role in search and indexing. Web crawlers and search algorithms utilise content credentials as trust signals to evaluate quality and origin of the content. EEAT is an important framework for AI to attribute trust points to a brand, synthetic media markers form an important signal for the same.

Global legal regulatory frameworks and legal compliance are strengthening. Synthetic media markers act as tool to implement AI transparency and deepfake disclosures.