DEDA toolkit reads and anonymises printer tracking dots

DEDA toolkit reads and anonymises printer tracking dots

DEDA is a Python tool published on GitHub. Its README starts from a simple fact: document colour tracking dots, also called yellow dots, are small systematic dots that encode information about the printer and/or the printout itself. According to the README, this process is built into almost every commercial colour laser printer, so almost every printout carries coded information about the source device, such as the serial number.

The tool works in two directions. It can read out and decode these forensic features from a scanned page, and it can anonymise a document so the printer cannot be used for arbitrary tracking. Users who rely on the software are asked to cite the 2018 paper "Forensic Analysis and Anonymisation of Printed Documents" by Timo Richter, Stephan Escher, Dagmar Schönfeld and Thorsten Strufe, published in the Proceedings of the 6th ACM Workshop on Information Hiding and Multimedia Security (IH&MMSec '18), pages 127-138.

Installation needs Python 3 and a pip install: pip3 install --user deda, or pip3 install --user . from a local checkout. A graphical interface opens with deda_gui. An optional dependency, Wand (Unix and GNU/Linux only), is needed by deda_anonmask_apply; without it, pages with white areas on images cannot be anonymised.

For reading dots, the README says tracking data can be read and only sometimes decoded from a scanned image. Good results need a lossless format such as png at 300 dpi, with a neutral contrast setting. deda_parse_print INPUTFILE decodes a single scan, and deda_compare_prints takes two or more scans for comparison. New patterns might not be recognised by parse_print; in that case deda_extract_yd INPUTFILE pulls out the dots for further analysis. deda_create_dots lets users build their own tracking-dot matrix and add it to a PDF, and the calibration page from deda_anonmask_create -w can serve as input.

There are two ways to clean things up. deda_clean_document INPUTFILE OUTPUTFILE (mostly) removes tracking data from a scan. For printing, the README gives a per-printer workflow. Save the document as DOCUMENT.PDF. Print the testpage.pdf created by deda_anonmask_create -w without any page margin. Scan that page at 300 dpi in a lossless format and pass it to deda_anonmask_create -r INPUTFILE, which creates mask.json, the individual printer's anonymisation mask. Then run deda_anonmask_apply mask.json DOCUMENT.PDF to produce masked.pdf, which may be printed with a zero page margin. The mask's dot radius and x and y offsets can be customised through parameters.

The README ends with troubleshooting. If the commands are not found, add the user's Python bin directory to PATH. If the GUI fails, the eel dependency may be at fault, and the README suggests installing build packages and eel. A Wand PolicyError is caused by ImageMagick; the fix is either to uninstall Wand or to add a PDF coder policy line to ImageMagick's policy.xml. If a scan shows no tracking dots, the scan program may be eliminating them, so settings should be changed; the README also reminds users that monochrome pages and inkjet prints might not contain tracking dots. In that case users can create their own dots, or print the calibration page on another printer and use that mask, either in anonymised form or as a straight copy (deda_anonmask_create --copy).

Key facts

  • DEDA is a Python toolkit that reads, decodes and anonymises the colour tracking dots (yellow dots) that, per its README, almost every commercial colour laser printer adds to printouts.
  • Per the README, almost every printout carries coded information about the source device, such as the serial number.
  • Decoding works from scans in a lossless format such as png at 300 dpi, but tracking data can only sometimes be decoded.
  • Anonymisation uses a per-printer mask: print a test page, scan it at 300 dpi, create mask.json, apply it to a PDF to get masked.pdf.
  • The README asks users to cite the 2018 ACM IH&MMSec '18 paper by Richter, Escher, Schönfeld and Strufe, and says anonymisation (mostly) removes tracking data and should be checked with a microscope.

Why it matters

Most people never see the yellow dots, yet the README says they are built into almost every commercial colour laser printer and can encode details such as the printer's serial number. DEDA puts both sides of that into one open-source tool: reading and decoding the dots as a forensic feature, and masking them to prevent arbitrary tracking. The approach rests on a 2018 academic paper presented at ACM IH&MMSec '18.

Who it affects

Anyone who prints documents on a colour laser printer, since the README says almost every such printout contains the coded information. The tool is aimed at people who want to read the dots from scans or anonymise documents before printing, and the paper citation request points to researchers working on forensic analysis of printed documents.

How to use it

Install Python 3, then run pip3 install --user deda, or install from a local checkout with pip3 install --user .. Open the GUI with deda_gui, or use the command line. To read a page, scan it as a lossless file such as png at 300 dpi with neutral contrast and run deda_parse_print INPUTFILE; deda_compare_prints compares several scans, and deda_extract_yd extracts the dots when a pattern is not recognised. To anonymise a document, print the test page from deda_anonmask_create -w with no margin, scan it at 300 dpi, run deda_anonmask_create -r INPUTFILE to get mask.json, then run deda_anonmask_apply mask.json DOCUMENT.PDF and print the resulting masked.pdf with a zero page margin. Install Wand if the document has white areas on images.

How solid is it

The source is the project's README, which explains the method and links the 2018 paper by Richter, Escher, Schönfeld and Strufe (DOI 10.1145/3206004.3206019) as the academic basis. The README gives no test results, success rates or accuracy figures for decoding or anonymisation, so how well it works on a given printer is not quantified there.

Risks and caveats

The README is open about the limits. Tracking data can only sometimes be decoded, and new patterns might not be recognised by parse_print. Cleaning (mostly) removes tracking data rather than guaranteeing it, so users are told to check with a microscope whether a masked page covers their printer's dots. Without Wand, pages with white areas on images cannot be anonymised, and Wand itself works only on Unix and GNU/Linux. Monochrome pages and inkjet prints might not contain tracking dots at all. Setup can also trip on an ImageMagick PolicyError or on the eel dependency for the GUI.

“almost every printout contains coded information about the source device, such as the serial number”

— DEDA README