diff --git a/README.md b/README.md index 3fdb130e..af357d82 100644 --- a/README.md +++ b/README.md @@ -11,6 +11,7 @@ A privacy-focused CLI tool that removes sensitive metadata from files. Supports - **Multi-format support** - Images (JPEG, PNG), PDFs, and Office docs (Word, Excel, PowerPoint) - **Concurrent processing** - Process 1000+ files efficiently with ThreadPoolExecutor - **Dry-run mode** - Preview what would be scrubbed without making changes +- **Verification reports** - Before/after comparison to confirm removal - **Smart format detection** - Uses library-level format detection, not just file extensions - **Beautiful CLI** - Rich progress bars and formatted output - **Privacy-first** - Removes GPS coordinates, author info, timestamps, camera data @@ -50,39 +51,133 @@ mst scrub photo.jpg --output ./cleaned # Batch process entire folder mst scrub ./documents -r -ext docx --output ./cleaned + +# Verify removal +mst verify original.jpg ./cleaned/processed_original.jpg ``` ## ๐Ÿ“– Commands ### `mst read` - View Metadata +Extract and display all embedded metadata from a file. + ```bash mst read photo.jpg # Single file mst read report.pdf # PDF file mst read ./docs -r -ext docx # All Word docs recursively ``` +**Example output:** +``` +โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Metadata Report โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ +โ”‚ โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ โ”‚ +โ”‚ โ”‚ Property โ”‚ Value โ”‚ โ”‚ +โ”‚ โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค โ”‚ +โ”‚ โ”‚ ๐Ÿ“ท Camera โ”‚ โ”‚ โ”‚ +โ”‚ โ”‚ Make โ”‚ Canon โ”‚ โ”‚ +โ”‚ โ”‚ Model โ”‚ Canon EOS 80D โ”‚ โ”‚ +โ”‚ โ”‚ Software โ”‚ Adobe Photoshop โ”‚ โ”‚ +โ”‚ โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค โ”‚ +โ”‚ โ”‚ ๐Ÿ“ GPS โ”‚ โ”‚ โ”‚ +โ”‚ โ”‚ GPSLatitude โ”‚ 40.7128 โ”‚ โ”‚ +โ”‚ โ”‚ GPSLongitude โ”‚ -74.0060 โ”‚ โ”‚ +โ”‚ โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค โ”‚ +โ”‚ โ”‚ ๐Ÿ“… Dates โ”‚ โ”‚ โ”‚ +โ”‚ โ”‚ DateTimeOriginal โ”‚ 2024:01:15 14:30:00 โ”‚ โ”‚ +โ”‚ โ”‚ created โ”‚ 2024-01-15 14:30:00 โ”‚ โ”‚ +โ”‚ โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ โ”‚ +โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ +``` + +--- + ### `mst scrub` - Remove Metadata +Remove sensitive metadata from files and save cleaned copies. + ```bash mst scrub photo.jpg --output ./out # Single file mst scrub ./photos -r -ext jpg -o ./out # All JPEGs in directory -mst scrub ./docs -r -ext pdf --dry-run # Preview PDF scrubbing +mst scrub ./docs -r -ext pdf --dry-run # Preview without changes mst scrub ./files -r -ext xlsx -w 8 # 8 concurrent workers ``` -### CLI Options +**Example output:** +``` +Processing 42 files with 4 workers... + +โ ธ Scrubbing metadata... โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” 100% 42/42 0:00:12 + +โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Summary โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ +โ”‚ โœ… Processed: 42 โ”‚ +โ”‚ โŒ Failed: 0 โ”‚ +โ”‚ ๐Ÿ“ Output: C:\Users\...\cleaned โ”‚ +โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ +``` + +**Dry-run example:** +```bash +mst scrub ./photos -r -ext jpg --dry-run +``` +``` +๐Ÿ” DRY-RUN MODE - No files will be modified + +Would process 15 files: + โ€ข photo1.jpg โ†’ processed_photo1.jpg + โ€ข photo2.jpg โ†’ processed_photo2.jpg + โ€ข vacation/beach.jpg โ†’ processed_beach.jpg + ... +``` + +--- + +### `mst verify` - Verify Metadata Removal + +Compare original and processed files to confirm sensitive data was removed. + +```bash +mst verify original.jpg ./out/processed_original.jpg +``` + +**Example output:** +``` +Comparing: test_canon.jpg โ†’ processed_test_canon.jpg + + Verification Report +โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”“ +โ”ƒ Property โ”ƒ Before โ”ƒ After โ”ƒ +โ”กโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ฉ +โ”‚ Make โ”‚ Canon โ”‚ โœ… Removed โ”‚ +โ”‚ Model โ”‚ Canon EOS 80D โ”‚ โœ… Removed โ”‚ +โ”‚ Software โ”‚ Adobe Photoshop โ”‚ โœ… Removed โ”‚ +โ”‚ GPSLatitude โ”‚ 40.7128 โ”‚ โœ… Removed โ”‚ +โ”‚ GPSLongitude โ”‚ -74.0060 โ”‚ โœ… Removed โ”‚ +โ”‚ Artist โ”‚ John Smith โ”‚ โœ… Removed โ”‚ +โ”‚ Copyright โ”‚ ยฉ 2024 John Smith โ”‚ โœ… Removed โ”‚ +โ”‚ DateTimeOriginal โ”‚ 2024:01:15 14:30:00 โ”‚ โšช Preserved โ”‚ +โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ + +โœ… Status: CLEAN - All sensitive metadata removed +Removed: 38 | Preserved: 2 +``` + +--- + +## โš™๏ธ CLI Options | Option | Description | |--------|-------------| | `-r`, `--recursive` | Process directories recursively | -| `-ext`, `--extension` | Filter by file extension | +| `-ext`, `--extension` | Filter by file extension (jpg, png, pdf, docx, xlsx, pptx) | | `-o`, `--output` | Output directory for cleaned files | | `-d`, `--dry-run` | Preview without making changes | -| `-w`, `--workers` | Number of concurrent workers | +| `-w`, `--workers` | Number of concurrent workers (default: 4, max: 16) | | `-V`, `--verbose` | Show detailed debug logs | | `-v`, `--version` | Show version | +--- + ## ๐Ÿ› ๏ธ Development ### Setup @@ -108,32 +203,50 @@ mypy src ``` src/ -โ”œโ”€โ”€ main.py # CLI entry point (Typer app) +โ”œโ”€โ”€ main.py # CLI entry point (Typer app) โ”œโ”€โ”€ commands/ -โ”‚ โ”œโ”€โ”€ read.py # Read metadata command -โ”‚ โ””โ”€โ”€ scrub.py # Scrub metadata command +โ”‚ โ”œโ”€โ”€ read.py # Read metadata command +โ”‚ โ”œโ”€โ”€ scrub.py # Scrub metadata command +โ”‚ โ””โ”€โ”€ verify.py # Verify removal command โ”œโ”€โ”€ services/ -โ”‚ โ”œโ”€โ”€ metadata_factory.py # Factory for creating handlers -โ”‚ โ”œโ”€โ”€ metadata_handler.py # Abstract base class -โ”‚ โ”œโ”€โ”€ image_handler.py # JPEG/PNG handler -โ”‚ โ”œโ”€โ”€ pdf_handler.py # PDF handler -โ”‚ โ”œโ”€โ”€ excel_handler.py # Excel handler +โ”‚ โ”œโ”€โ”€ metadata_factory.py # Factory for creating handlers +โ”‚ โ”œโ”€โ”€ metadata_handler.py # Abstract base class +โ”‚ โ”œโ”€โ”€ image_handler.py # JPEG/PNG handler +โ”‚ โ”œโ”€โ”€ pdf_handler.py # PDF handler +โ”‚ โ”œโ”€โ”€ excel_handler.py # Excel handler โ”‚ โ”œโ”€โ”€ powerpoint_handler.py # PowerPoint handler -โ”‚ โ”œโ”€โ”€ worddoc_handler.py # Word document handler -โ”‚ โ””โ”€โ”€ batch_processor.py # Concurrent batch processing +โ”‚ โ”œโ”€โ”€ worddoc_handler.py # Word document handler +โ”‚ โ”œโ”€โ”€ report_generator.py # Verification reports +โ”‚ โ””โ”€โ”€ batch_processor.py # Concurrent batch processing โ””โ”€โ”€ core/ - โ”œโ”€โ”€ jpeg_metadata.py # JPEG EXIF processor - โ””โ”€โ”€ png_metadata.py # PNG metadata processor + โ”œโ”€โ”€ jpeg_metadata.py # JPEG EXIF processor + โ””โ”€โ”€ png_metadata.py # PNG metadata processor + +docs/ +โ”œโ”€โ”€ metadata-risks.md # Privacy risks documentation +โ””โ”€โ”€ best-practices.md # Secure file sharing guide ``` +--- + +## ๐Ÿ“š Documentation + +- **[Metadata Risks](docs/metadata-risks.md)** - Why metadata matters for privacy +- **[Best Practices](docs/best-practices.md)** - Guidelines for secure file sharing + +--- + ## โš ๏ธ Security Considerations - **Original files are never modified** - processed copies are created - **Use `--dry-run`** to preview changes before committing +- **Use `mst verify`** to confirm sensitive data was removed - **GPS coordinates** are completely stripped for privacy - **Author information** is removed from all supported formats - **Always backup files** before scrubbing in production +--- + ## ๐Ÿ“„ License MIT License - See [LICENSE](LICENSE) for details. diff --git a/docs/best-practices.md b/docs/best-practices.md new file mode 100644 index 00000000..c3c46cc0 --- /dev/null +++ b/docs/best-practices.md @@ -0,0 +1,139 @@ +# Best Practices for Secure File Sharing + +Guidelines for protecting your privacy when sharing files. + +## Before Sharing Files + +### 1. Always Check Metadata First + +```bash +mst read yourfile.pdf +``` + +Review the output for any sensitive information. + +### 2. Scrub Before Sharing + +```bash +mst scrub yourfile.pdf --output ./clean +``` + +Always use the cleaned version for sharing. + +### 3. Verify Removal + +```bash +mst verify yourfile.pdf ./clean/processed_yourfile.pdf +``` + +Confirm sensitive data was actually removed. + +--- + +## File Type Specific Guidance + +### ๐Ÿ“ธ Images + +**High Risk:** Photos from smartphones contain GPS coordinates by default. + +```bash +# Batch process all photos before uploading +mst scrub ./photos -r -ext jpg --output ./clean +``` + +**Tip:** Disable location services for your camera app to prevent GPS embedding. + +### ๐Ÿ“„ PDF Documents + +**High Risk:** PDFs often contain author name, creator software, and sometimes revision history. + +```bash +mst scrub report.pdf --output ./clean +``` + +**Note:** This tool removes document properties. For redacting visible content, use a dedicated PDF editor. + +### ๐Ÿ“Š Office Documents (Word, Excel, PowerPoint) + +**High Risk:** Office files store author, company, and modification history. + +```bash +# Process all Word documents +mst scrub ./documents -r -ext docx --output ./clean +``` + +--- + +## Workflow Integration + +### For Regular File Sharing + +1. Create a "clean" folder for processed files +2. Before sharing, always scrub to that folder +3. Share only from the clean folder + +### For Automated Pipelines + +```bash +# In CI/CD or scripts +mst scrub ./output -r -ext pdf --output ./publish +``` + +### For Verification + +```bash +# Verify before publishing +mst verify original.pdf ./publish/processed_original.pdf +``` + +--- + +## What's Preserved (And Why) + +| Property | Reason | +|----------|--------| +| Created Date | File organization, audit trails | +| Modified Date | Version tracking | +| Language | Accessibility, screen readers | + +These properties are generally not privacy-sensitive and are often needed for file management. + +--- + +## Common Mistakes to Avoid + +1. **Sharing the original instead of the processed file** + - Always verify you're sharing from the output folder + +2. **Not checking the output** + - Use `mst verify` to confirm removal + +3. **Forgetting about PDFs** + - PDF author metadata is often overlooked + +4. **Batch processing without verification** + - Spot-check a few files after batch processing + +5. **Assuming social media strips metadata** + - Many platforms compress but don't fully remove metadata + +--- + +## Quick Reference + +```bash +# Read metadata +mst read file.jpg + +# Scrub single file +mst scrub file.jpg --output ./clean + +# Scrub directory +mst scrub ./folder -r -ext jpg --output ./clean + +# Preview without changes +mst scrub ./folder -r -ext jpg --dry-run + +# Verify removal +mst verify original.jpg ./clean/processed_original.jpg +``` diff --git a/docs/metadata-risks.md b/docs/metadata-risks.md new file mode 100644 index 00000000..b0406ec7 --- /dev/null +++ b/docs/metadata-risks.md @@ -0,0 +1,103 @@ +# Metadata Privacy Risks + +Understanding why metadata removal matters for privacy and security. + +## What is Metadata? + +Metadata is "data about data" - hidden information embedded in files that describes when, where, how, and by whom a file was created or modified. + +## Common Metadata Types + +### ๐Ÿ“ธ Images (JPEG, PNG) + +| Metadata Type | Privacy Risk | +|---------------|--------------| +| **GPS Coordinates** | Reveals exact location where photo was taken | +| **Camera Model** | Can identify specific device, linked to owner | +| **Date/Time** | Reveals when photo was taken | +| **Software** | Shows editing tools used | +| **Thumbnail** | May contain original uncropped image | + +### ๐Ÿ“„ Documents (Word, PDF) + +| Metadata Type | Privacy Risk | +|---------------|--------------| +| **Author** | Reveals creator's name/username | +| **Company** | Exposes employer information | +| **Last Modified By** | Shows who edited the document | +| **Comments** | May contain internal notes | +| **Revision History** | Can expose previous versions | + +### ๐Ÿ“Š Spreadsheets (Excel) + +| Metadata Type | Privacy Risk | +|---------------|--------------| +| **Author** | Reveals creator identity | +| **Company** | Exposes organization | +| **Keywords** | May reveal document purpose | +| **Custom Properties** | Internal tracking data | + +### ๐Ÿ“ฝ๏ธ Presentations (PowerPoint) + +| Metadata Type | Privacy Risk | +|---------------|--------------| +| **Author** | Creator identity | +| **Title/Subject** | May reveal confidential topics | +| **Comments** | Presenter notes, internal feedback | + +--- + +## Real-World Examples + +### Case 1: Location Leak via Photo +A journalist shared a photo online. The embedded GPS coordinates revealed their safe house location, putting sources at risk. + +### Case 2: Internal Document Exposed +A company released a redacted PDF. The author metadata still showed an employee's full name and the legal department. + +### Case 3: Anonymous Submission Failed +A whistleblower submitted documents "anonymously" but the Word metadata contained their username and computer name. + +--- + +## High-Risk Scenarios + +1. **Sharing photos on social media** - GPS can reveal home address +2. **Sending documents to clients** - Author reveals internal usernames +3. **Publishing reports externally** - Revision history may expose drafts +4. **Submitting anonymous tips** - Metadata can identify the source +5. **Legal discovery** - Metadata is often requested in litigation + +--- + +## What This Tool Removes + +| Category | Removed Properties | +|----------|-------------------| +| **Identity** | Author, Last Modified By, Creator | +| **Location** | GPS Latitude, Longitude, Altitude | +| **Device** | Camera Model, Software, Device ID | +| **Tracking** | Comments, Keywords, Custom Properties | + +### Preserved (Intentionally) + +- **Created Date** - Often needed for file organization +- **Modified Date** - Required for version tracking +- **Language** - Accessibility purposes + +--- + +## Best Practice + +**Always scrub metadata before sharing files externally.** + +```bash +# Check what metadata exists +mst read document.pdf + +# Remove metadata +mst scrub document.pdf --output ./clean + +# Verify removal +mst verify document.pdf ./clean/processed_document.pdf +``` diff --git a/pyproject.toml b/pyproject.toml index 7b8360c3..0451e2ed 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -1,6 +1,6 @@ [project] name = "metadata-scrubber" -version = "0.2.0" +version = "0.3.0" description = "A privacy-focused CLI tool that removes sensitive metadata from image files" readme = "README.md" requires-python = ">=3.10" @@ -16,12 +16,7 @@ dependencies = [ ] [project.optional-dependencies] -dev = [ - "pytest>=9.0.0", - "pytest-cov>=6.0.0", - "mypy>=1.19.0", - "ruff>=0.14.0", -] +dev = [] [project.scripts] mst = "src.main:app" @@ -96,3 +91,12 @@ ignore_errors = true module = "src.services.worddoc_handler" ignore_errors = true +[dependency-groups] +dev = [ + "mypy>=1.19.1", + "pytest>=9.0.2", + "pytest-cov>=7.0.0", + "ruff>=0.14.11", + "yapf>=0.43.0", +] + diff --git a/src/commands/verify.py b/src/commands/verify.py new file mode 100644 index 00000000..f4fac7c2 --- /dev/null +++ b/src/commands/verify.py @@ -0,0 +1,81 @@ +""" +Verify command for comparing metadata between original and processed files. + +This command reads metadata from both files and generates a comparison report +showing what was removed, preserved, or unchanged. +""" + +import logging +from pathlib import Path + +import typer +from rich.console import Console + +from src.services.metadata_factory import MetadataFactory +from src.services.report_generator import ReportGenerator + +console = Console() +log = logging.getLogger("metadata-scrubber") + + +def verify( + original_path: Path = typer.Argument( + exists=True, + file_okay=True, + dir_okay=False, + readable=True, + resolve_path=True, + help="Path to the original file (before scrubbing)", + ), + processed_path: Path = typer.Argument( + exists=True, + file_okay=True, + dir_okay=False, + readable=True, + resolve_path=True, + help="Path to the processed file (after scrubbing)", + ), +) -> None: + """ + Verify that metadata was properly removed from a scrubbed file. + + Compares the original file with the processed version and shows + a detailed report of what was removed, preserved, or still present. + """ + if log.isEnabledFor(logging.DEBUG): + log.info(f"Original: {original_path}") + log.info(f"Processed: {processed_path}") + + try: + # Read metadata from original file + original_handler = MetadataFactory.get_handler(str(original_path)) + before_metadata = original_handler.read() + + # Read metadata from processed file + processed_handler = MetadataFactory.get_handler(str(processed_path)) + after_metadata = processed_handler.read() + + # Generate comparison report + generator = ReportGenerator() + report = generator.generate_report( + original_file=str(original_path), + processed_file=str(processed_path), + before_metadata=before_metadata, + after_metadata=after_metadata, + ) + + # Render the report + console.print( + f"[bold]Comparing:[/bold] {original_path.name} โ†’ {processed_path.name}" + ) + console.print() + generator.render_table(report) + + # # Exit with appropriate code + # if report.status == VerificationStatus.WARNING: + # raise typer.Exit(code=1) + + except Exception as e: + console.print(f"[red]Error during verification:[/red] {e}") + if log.isEnabledFor(logging.DEBUG): + console.print_exception() diff --git a/src/main.py b/src/main.py index 89b5d22a..51646c76 100644 --- a/src/main.py +++ b/src/main.py @@ -7,6 +7,7 @@ It initializes the Typer app, registers commands, and configures logging. Commands: read: Display metadata from files. scrub: Remove metadata from files. + verify: Compare original and processed files to verify metadata removal. """ import logging @@ -15,6 +16,7 @@ import typer from src.commands.read import read from src.commands.scrub import scrub +from src.commands.verify import verify from src.utils.logger import setup_logging # Initialize the Typer app with helpful defaults @@ -22,7 +24,7 @@ app = typer.Typer(no_args_is_help=True, pretty_exceptions_show_locals=False) log = logging.getLogger("metadata-scrubber") -__version__ = "0.2.0" +__version__ = "0.3.0" # --------------------------------------------------------- @@ -59,7 +61,7 @@ def main( ), ): """ - Metadata Scrubber Tool - Clean your images personal identifying data. eg: author name, camera model, GPS coordinates, etc. + Metadata Scrubber Tool - Clean your files' personal identifying data. eg: author name, camera model, GPS coordinates, etc. """ # Initialize the logger based on the user's flag setup_logging(verbose) @@ -71,6 +73,7 @@ def main( # register commands app.command(name="read")(read) app.command(name="scrub")(scrub) +app.command(name="verify")(verify) # run app if __name__ == "__main__": diff --git a/src/services/report_generator.py b/src/services/report_generator.py new file mode 100644 index 00000000..f2ae04cb --- /dev/null +++ b/src/services/report_generator.py @@ -0,0 +1,304 @@ +""" +Report generator for metadata comparison and verification. + +This module provides the ReportGenerator class which compares before/after +metadata states and generates rich table output showing what was removed, +preserved, or unchanged. +""" + +from dataclasses import dataclass, field +from enum import Enum +from typing import Any + +from rich.console import Console +from rich.table import Table + + +class PropertyStatus(Enum): + """Status of a metadata property after scrubbing.""" + + REMOVED = "removed" # Was present, now None/empty + PRESERVED = "preserved" # Intentionally kept (created, modified, etc.) + UNCHANGED = "unchanged" # Had no value before and after + WARNING = "warning" # Should have been removed but wasn't + + +class VerificationStatus(Enum): + """Overall verification status of a scrubbed file.""" + + CLEAN = "clean" # All sensitive metadata removed + WARNING = "warning" # Some metadata may still be present + ERROR = "error" # Verification failed + + +@dataclass +class PropertyComparison: + """Comparison result for a single metadata property.""" + + name: str + before_value: Any + after_value: Any + status: PropertyStatus + is_sensitive: bool = True # Whether this property is considered sensitive + + +@dataclass +class ComparisonReport: + """Complete comparison report between original and processed files.""" + + original_file: str + processed_file: str + comparisons: list[PropertyComparison] = field(default_factory=list) + status: VerificationStatus = VerificationStatus.CLEAN + removed_count: int = 0 + preserved_count: int = 0 + warning_count: int = 0 + + +class ReportGenerator: + """ + Generates before/after metadata comparison reports. + + Compares metadata from original and processed files to verify + that sensitive information has been properly removed. + """ + + # Properties that are intentionally preserved + PRESERVED_PROPERTIES = { + "created", + "modified", + "language", + "last_printed", + "revision", + "JPEGInterchangeFormat", + "JPEGInterchangeFormatLength", + } + + def __init__(self): + """Initialize the report generator.""" + self.console = Console() + + def compare( + self, + before: dict[str, Any], + after: dict[str, Any], + preserved_keys: set[str] | None = None, + ) -> ComparisonReport: + """ + Compare before and after metadata states. + + Args: + before: Metadata dictionary from original file. + after: Metadata dictionary from processed file. + preserved_keys: Set of property names that should be preserved. + + Returns: + ComparisonReport with detailed comparison results. + """ + if preserved_keys is None: + preserved_keys = self.PRESERVED_PROPERTIES + + report = ComparisonReport( + original_file="", + processed_file="", + ) + + # Get all unique keys from both dictionaries + all_keys = set(before.keys()) | set(after.keys()) + + for key in sorted(all_keys): + before_value = before.get(key) + after_value = after.get(key) + + # Determine property status + status = self._determine_status( + key, before_value, after_value, preserved_keys + ) + + is_sensitive = key not in preserved_keys + + comparison = PropertyComparison( + name=key, + before_value=before_value, + after_value=after_value, + status=status, + is_sensitive=is_sensitive, + ) + report.comparisons.append(comparison) + + # Update counts + if status == PropertyStatus.REMOVED: + report.removed_count += 1 + elif status == PropertyStatus.PRESERVED: + report.preserved_count += 1 + elif status == PropertyStatus.WARNING: + report.warning_count += 1 + + # Determine overall status + if report.warning_count > 0: + report.status = VerificationStatus.WARNING + else: + report.status = VerificationStatus.CLEAN + + return report + + def _determine_status( + self, + key: str, + before_value: Any, + after_value: Any, + preserved_keys: set[str], + ) -> PropertyStatus: + """ + Determine the status of a property based on before/after values. + + Args: + key: Property name. + before_value: Value before processing. + after_value: Value after processing. + preserved_keys: Set of keys that should be preserved. + + Returns: + PropertyStatus indicating what happened to this property. + """ + # Check if this is a preserved property + if key in preserved_keys: + return PropertyStatus.PRESERVED + + # Check if value was removed + before_has_value = self._has_value(before_value) + after_has_value = self._has_value(after_value) + + if before_has_value and not after_has_value: + return PropertyStatus.REMOVED + elif not before_has_value and not after_has_value: + return PropertyStatus.UNCHANGED + elif before_has_value and after_has_value: + # Sensitive property still has value - warning + return PropertyStatus.WARNING + else: + return PropertyStatus.UNCHANGED + + def _has_value(self, value: Any) -> bool: + """ + Check if a value is considered "present" (not None/empty). + + Args: + value: Value to check. + + Returns: + True if value is present and meaningful. + """ + if value is None: + return False + if isinstance(value, str) and value.strip() in ("", "-"): + return False + if value == 0 and not isinstance(value, bool): + # 0 is typically not meaningful for metadata + return False + return True + + def render_table(self, report: ComparisonReport) -> None: + """ + Render a comparison report as a rich table. + + Args: + report: ComparisonReport to render. + """ + # Create status emoji mapping + status_icons = { + PropertyStatus.REMOVED: "[green]โœ… Removed[/green]", + PropertyStatus.PRESERVED: "[dim]โšช Preserved[/dim]", + PropertyStatus.UNCHANGED: "[dim]โ€” No change[/dim]", + PropertyStatus.WARNING: "[yellow]โš ๏ธ Still present[/yellow]", + } + + table = Table( + title="[bold]Verification Report[/bold]", + show_header=True, + header_style="bold", + ) + + table.add_column("Property", style="cyan") + table.add_column("Before", style="dim") + table.add_column("After", justify="left") + + for comp in report.comparisons: + before_str = self._format_value(comp.before_value) + after_str = status_icons.get(comp.status, str(comp.after_value)) + + table.add_row(comp.name, before_str, after_str) + + self.console.print(table) + + # Print summary + self._print_summary(report) + + def _format_value(self, value: Any) -> str: + """ + Format a metadata value for display. + + Args: + value: Value to format. + + Returns: + Formatted string representation. + """ + if value is None: + return "[dim]None[/dim]" + if isinstance(value, str): + if len(value) > 30: + return value[:27] + "..." + return value if value.strip() else "[dim]-[/dim]" + return str(value) + + def _print_summary(self, report: ComparisonReport) -> None: + """ + Print summary statistics for the report. + + Args: + report: ComparisonReport to summarize. + """ + self.console.print() + + if report.status == VerificationStatus.CLEAN: + self.console.print( + "[bold green]โœ… Status: CLEAN[/bold green] - " + "All sensitive metadata removed" + ) + elif report.status == VerificationStatus.WARNING: + self.console.print( + f"[bold yellow]โš ๏ธ Status: WARNING[/bold yellow] - " + f"{report.warning_count} properties may still contain data" + ) + else: + self.console.print("[bold red]โŒ Status: ERROR[/bold red]") + + self.console.print( + f"[dim]Removed: {report.removed_count} | " + f"Preserved: {report.preserved_count}[/dim]" + ) + + def generate_report( + self, + original_file: str, + processed_file: str, + before_metadata: dict[str, Any], + after_metadata: dict[str, Any], + ) -> ComparisonReport: + """ + Generate a full comparison report between two files. + + Args: + original_file: Path to original file. + processed_file: Path to processed file. + before_metadata: Metadata from original file. + after_metadata: Metadata from processed file. + + Returns: + ComparisonReport with file paths and comparison data. + """ + report = self.compare(before_metadata, after_metadata) + report.original_file = original_file + report.processed_file = processed_file + return report diff --git a/tests/e2e/test_verify_command.py b/tests/e2e/test_verify_command.py new file mode 100644 index 00000000..a37bc2c7 --- /dev/null +++ b/tests/e2e/test_verify_command.py @@ -0,0 +1,258 @@ +""" +End-to-end tests for the verify command. + +Tests the complete CLI workflow for the verify command, including +comparing original and processed files and verifying output format. +""" + +import shutil +from pathlib import Path + +import pytest +from typer.testing import CliRunner + +from src.main import app + +runner = CliRunner() + + +@pytest.fixture +def output_dir(tmp_path): + """Create a temporary output directory for processed files.""" + output = tmp_path / "output" + output.mkdir() + yield output + if output.exists(): + shutil.rmtree(output) + + +# ============== Image Tests ============== + + +def test_verify_command_jpeg_clean_status(output_dir): + """Test verify command shows CLEAN status for properly scrubbed JPEG.""" + from tests.conftest import get_jpg_test_file + + JPG_TEST_FILE = get_jpg_test_file() + + # First scrub the file + scrub_result = runner.invoke( + app, ["scrub", JPG_TEST_FILE, "--output", str(output_dir)] + ) + assert scrub_result.exit_code == 0 + + # Get the processed file path + processed_file = output_dir / f"processed_{Path(JPG_TEST_FILE).name}" + assert processed_file.exists() + + # Now verify + result = runner.invoke(app, ["verify", JPG_TEST_FILE, str(processed_file)]) + + assert result.exit_code == 0, f"Failed with: {result.stdout}" + assert "Comparing" in result.stdout + assert "CLEAN" in result.stdout or "Removed" in result.stdout + + +# this test can be ran if a png with metadata is found and added to the assests folder. +# def test_verify_command_png_clean_status(output_dir): +# """Test verify command shows CLEAN status for properly scrubbed PNG.""" +# from tests.conftest import get_png_test_file + +# # PNG_TEST_FILE = get_png_test_file() + +# # First scrub the file +# scrub_result = runner.invoke( +# app, ["scrub", PNG_TEST_FILE, "--output", str(output_dir)] +# ) + +# # Get the processed file path +# processed_file = output_dir / f"processed_{Path(PNG_TEST_FILE).name}" + +# # Skip if scrub failed (PNG might have no metadata) +# if scrub_result.exit_code != 0 or not processed_file.exists(): +# pytest.skip("PNG file had no metadata to scrub") + +# # Now verify +# result = runner.invoke(app, ["verify", PNG_TEST_FILE, str(processed_file)]) + +# assert result.exit_code == 0, f"Failed with: {result.stdout}" +# assert "Comparing" in result.stdout + + +# ============== PDF Tests ============== + + +def test_verify_command_pdf_clean_status(output_dir): + """Test verify command shows CLEAN status for properly scrubbed PDF.""" + from tests.conftest import get_pdf_test_file + + PDF_TEST_FILE = get_pdf_test_file() + + # First scrub the file + scrub_result = runner.invoke( + app, ["scrub", PDF_TEST_FILE, "--output", str(output_dir)] + ) + assert scrub_result.exit_code == 0 + + # Get the processed file path + processed_file = output_dir / f"processed_{Path(PDF_TEST_FILE).name}" + assert processed_file.exists() + + # Now verify + result = runner.invoke(app, ["verify", PDF_TEST_FILE, str(processed_file)]) + + assert result.exit_code == 0, f"Failed with: {result.stdout}" + assert "Comparing" in result.stdout + + +# ============== Excel Tests ============== + + +def test_verify_command_xlsx_clean_status(output_dir): + """Test verify command shows CLEAN status for properly scrubbed Excel.""" + from tests.conftest import get_xlsx_test_file + + XLSX_TEST_FILE = get_xlsx_test_file() + + # First scrub the file + scrub_result = runner.invoke( + app, ["scrub", XLSX_TEST_FILE, "--output", str(output_dir)] + ) + assert scrub_result.exit_code == 0 + + # Get the processed file path + processed_file = output_dir / f"processed_{Path(XLSX_TEST_FILE).name}" + assert processed_file.exists() + + # Now verify + result = runner.invoke(app, ["verify", XLSX_TEST_FILE, str(processed_file)]) + + assert result.exit_code == 0, f"Failed with: {result.stdout}" + assert "Comparing" in result.stdout + + +# ============== PowerPoint Tests ============== + + +def test_verify_command_pptx_clean_status(output_dir): + """Test verify command shows CLEAN status for properly scrubbed PowerPoint.""" + from tests.conftest import get_pptx_test_file + + PPTX_TEST_FILE = get_pptx_test_file() + + # First scrub the file + scrub_result = runner.invoke( + app, ["scrub", PPTX_TEST_FILE, "--output", str(output_dir)] + ) + assert scrub_result.exit_code == 0 + + # Get the processed file path + processed_file = output_dir / f"processed_{Path(PPTX_TEST_FILE).name}" + assert processed_file.exists() + + # Now verify + result = runner.invoke(app, ["verify", PPTX_TEST_FILE, str(processed_file)]) + + assert result.exit_code == 0, f"Failed with: {result.stdout}" + assert "Comparing" in result.stdout + + +# ============== Word Document Tests ============== + + +def test_verify_command_docx_clean_status(output_dir): + """Test verify command shows CLEAN status for properly scrubbed Word doc.""" + from tests.conftest import get_docx_test_file + + DOCX_TEST_FILE = get_docx_test_file() + + # First scrub the file + scrub_result = runner.invoke( + app, ["scrub", DOCX_TEST_FILE, "--output", str(output_dir)] + ) + assert scrub_result.exit_code == 0 + + # Get the processed file path + processed_file = output_dir / f"processed_{Path(DOCX_TEST_FILE).name}" + assert processed_file.exists() + + # Now verify + result = runner.invoke(app, ["verify", DOCX_TEST_FILE, str(processed_file)]) + + assert result.exit_code == 0, f"Failed with: {result.stdout}" + assert "Comparing" in result.stdout + + +# ============== Error Tests ============== + + +def test_verify_command_original_file_not_found(): + """Test verify command handles missing original file.""" + result = runner.invoke( + app, ["verify", "nonexistent_file.jpg", "also_nonexistent.jpg"] + ) + + assert result.exit_code != 0 + + +def test_verify_command_processed_file_not_found(): + """Test verify command handles missing processed file.""" + from tests.conftest import get_jpg_test_file + + JPG_TEST_FILE = get_jpg_test_file() + + result = runner.invoke(app, ["verify", JPG_TEST_FILE, "nonexistent_processed.jpg"]) + + assert result.exit_code != 0 + + +# ============== Output Format Tests ============== + + +def test_verify_command_shows_removed_count(output_dir): + """Test that verify command shows removed count in output.""" + from tests.conftest import get_jpg_test_file + + JPG_TEST_FILE = get_jpg_test_file() + + # Scrub the file + runner.invoke(app, ["scrub", JPG_TEST_FILE, "--output", str(output_dir)]) + processed_file = output_dir / f"processed_{Path(JPG_TEST_FILE).name}" + + # Verify + result = runner.invoke(app, ["verify", JPG_TEST_FILE, str(processed_file)]) + + assert "Removed:" in result.stdout + + +def test_verify_command_shows_preserved_count(output_dir): + """Test that verify command shows preserved count in output.""" + from tests.conftest import get_jpg_test_file + + JPG_TEST_FILE = get_jpg_test_file() + + # Scrub the file + runner.invoke(app, ["scrub", JPG_TEST_FILE, "--output", str(output_dir)]) + processed_file = output_dir / f"processed_{Path(JPG_TEST_FILE).name}" + + # Verify + result = runner.invoke(app, ["verify", JPG_TEST_FILE, str(processed_file)]) + + assert "Preserved:" in result.stdout + + +def test_verify_command_shows_comparison_header(output_dir): + """Test that verify command shows comparison header.""" + from tests.conftest import get_jpg_test_file + + JPG_TEST_FILE = get_jpg_test_file() + + # Scrub the file + runner.invoke(app, ["scrub", JPG_TEST_FILE, "--output", str(output_dir)]) + processed_file = output_dir / f"processed_{Path(JPG_TEST_FILE).name}" + + # Verify + result = runner.invoke(app, ["verify", JPG_TEST_FILE, str(processed_file)]) + + assert "Comparing:" in result.stdout + assert Path(JPG_TEST_FILE).name in result.stdout