Use this protocol before telling the user that a PDF is final, correct, ready, fixed, redacted, merged, split, or visually polished.
Always confirm these points:
- The output file exists.
- The PDF opens without parser errors.
- The page count matches the requested output.
- Page size, orientation, and rotation are intentional.
- The output file size is reasonable for the content.
- Any user-requested operation is reflected in the final file.
Recommended command:
python scripts/validate_pdf.py output.pdfWith an expected page count:
python scripts/validate_pdf.py output.pdf --expected-pages 5Use visual validation when the user asks for design, layout, branding, forms, tables, charts, posters, certificates, one-pagers, handouts, scanned pages, redactions, stamps, signatures, or page-level changes.
Recommended command:
python scripts/validate_pdf.py output.pdf --render --out-dir validation/output --dpi 160Check rendered PNG files for:
- clipped text
- overlapping objects
- missing glyphs
- broken images
- incorrect page order
- weak hierarchy
- poor contrast
- margins that are too tight
- headers, footers, tables, and charts crossing page boundaries
- watermarks covering important content
- redaction marks that do not fully cover the intended visual region
Use text-layer validation when the PDF must be searchable, copyable, accessible, or properly redacted.
Recommended command:
python scripts/validate_pdf.py output.pdf --extract-text --text-out validation/output.txtCheck:
- whether text extraction returns meaningful text
- whether expected phrases appear
- whether removed or redacted sensitive text is absent
- whether page ordering in extracted text is correct
- whether scanned pages are correctly identified as having limited or no text layer
For true redaction, never rely only on drawing a black box, white box, or shape over text.
A valid redaction workflow must verify both:
- Visual layer: the sensitive content is not visible in rendered pages.
- Text layer: the sensitive content cannot be extracted from the PDF.
Recommended checks:
python scripts/validate_pdf.py redacted.pdf --render --extract-text --text-out validation/redacted.txtThen search the extracted text for the sensitive terms. If any sensitive term remains, the PDF is not safely redacted.
For merge operations:
- confirm final page count equals the sum of included pages
- render the first page of each source boundary when practical
- confirm bookmarks or metadata limitations if relevant
For split operations:
- confirm selected pages match the requested page range
- confirm no unrequested pages remain
For rotation operations:
- inspect rotation metadata
- render affected pages to confirm orientation visually
For PDF forms:
- confirm fields are filled with the intended values
- confirm whether the form should remain editable or be flattened
- render final pages to check field overflow and alignment
- avoid claiming a form is flattened unless verified
For professional visual PDFs:
- use readable type size
- preserve adequate foreground/background contrast
- avoid colour-only meaning
- label charts clearly
- use consistent spacing, alignment, and page rhythm
- verify page renders at the target size
When delivering a finished PDF, include:
- The file link.
- A brief summary of what was created or changed.
- What validation was completed.
- Any limitations, such as scanned pages, OCR uncertainty, missing fonts, unavailable source files, or third-party service dependency.
For repository or release validation, run:
python scripts/validate_package.py .
python -m compileall scripts srcIf installed as a package, also run:
pdf-pro --help
pdf-pro validate --helpFor CLI-generated PDFs, include the command used and the validation result in the delivery summary. Prefer writing validation outputs to validation/<document-name>/.
Recommended pattern:
pdf-pro inspect input.pdf
pdf-pro <operation> input.pdf --out output.pdf
pdf-pro validate output.pdf --render --extract-text --out-dir validation/output --text-out validation/output.txt