Document is not tagged
The PDF has no logical structure tree, so assistive technology cannot navigate the content semantically.
From raw PDFs to well-done.
The current PDF accessibility check library ChefPDF is being built to evaluate.
214 planned checks
The checks below are scaffolded as a deterministic PDF/UA-oriented library. They do not run real PDF parsing yet, but each item has a stable ID for future implementation, reporting, and remediation work.
No checks match the current filters.
The PDF has no logical structure tree, so assistive technology cannot navigate the content semantically.
The catalog does not identify the document as tagged even when structural tags may be present.
The StructTreeRoot entry is absent or invalid, preventing a reliable logical reading model.
Marked content references cannot be resolved back to the structure tree through a valid parent tree.
MCID values are reused in a way that makes content-to-tag associations ambiguous.
Visible or meaningful content exists in marked content sequences that are not connected to structural elements.
Tags point to content items that cannot be found on the referenced page.
Visible content appears on a page without marked content identifiers or structural ownership.
Meaningful text, images, or controls are hidden from assistive technology as artifacts.
Headers, footers, page numbers, backgrounds, borders, or purely decorative marks are exposed as meaningful content.
Custom structure types are not mapped to valid standard structure types.
Custom tags are used without enough namespace information to interpret them reliably.
Page dictionaries do not expose the structural parent links needed to resolve tagged content.
The PDF contains old structure references from prior edits that no longer match current page content.
Content is tagged with a structure type that misrepresents its actual role.
Meaningful body text is exposed through Div, Span, or other generic tags instead of paragraph-level semantics.
The visible document title is missing or not exposed as an appropriate heading or title structure.
Block quotations or inline quotations are tagged as plain paragraphs or spans.
Program code, commands, or literal text are not exposed with a suitable code-related structure.
Footnote content is not linked to the corresponding reference in the reading order.
Endnote content and its invocation are not structurally connected.
Caption text is detached from the figure, chart, or table it describes.
Bibliographic references, citations, or cross-references are exposed as plain text without reference semantics.
Layout-only table structures are exposed as semantic tables, adding noisy navigation for assistive technology.
Text is fragmented into excessive inline tags that interrupt meaningful reading and selection.
Spacing characters, empty paragraphs, or layout-only text runs are exposed as content.
Abbreviations or acronyms are not expanded when their meaning is not clear from context.
Symbols, ligatures, or custom glyphs do not provide ActualText or equivalent semantic text.
The logical structure order differs from the sequence needed to understand the document.
Columnar layouts are ordered horizontally across columns instead of within each column sequence.
Pull quotes, sidebars, marginalia, or callouts are inserted in a way that breaks the primary reading flow.
Running headers, footers, or repeated navigation content are exposed repeatedly in the logical flow.
Line-break hyphenation causes words to be read, copied, or searched incorrectly.
The content stream order and the structure tree order conflict in ways that can affect assistive technology.
When content is reflowed, labels, values, captions, or grouped items are separated from each other.
Invisible, clipped, or off-page content appears in the accessible reading stream.
Heading hierarchy jumps over levels without a meaningful structural reason.
Text styled as a heading is exposed as a paragraph or generic tag.
Large, bold, or decorative text is incorrectly exposed as a heading.
The document exposes multiple top-level headings where the hierarchy suggests a single title should lead the structure.
A heading tag has no meaningful text or accessible replacement.
Headings are structurally ordered in a way that does not match the document sections users see.
Visually grouped list items are exposed as plain paragraphs or lines.
Bullets, numbers, or item markers are not represented in the list item structure.
List items do not contain a valid body element for the item content.
Nested lists are exposed as a single flat list, losing hierarchy.
Ordered list labels do not match the visible numbering or sequence.
Repeated label-value or step patterns are not grouped in a navigable list-like structure.
Rows and columns are visually present but missing table structure semantics.
Header cells are not identified for a data table that requires row or column headers.
Header cells do not define whether they apply to rows, columns, or both.
Complex table data cells cannot be programmatically associated with all relevant headers.
Table cells are not contained in valid table row structures.
RowSpan or ColSpan values do not match the visual table grid.
Blank cells are not identified as intentionally empty or are used only for layout spacing.
A visible table title or caption is not associated with the table structure.
A complex table lacks explanatory text needed to understand its organization.
A table split across pages or sections does not preserve header and continuation relationships.
Cells that visually span a region are represented as multiple unrelated cells.
Lines, shading, or positioning carry meaning that is not represented in the table structure.
An image, figure, or graphic that conveys content has no accessible description.
Purely decorative imagery is announced instead of being artifacted.
Image descriptions such as image, photo, or graphic do not convey the image purpose.
The figure description duplicates nearby caption or body text without adding useful meaning.
A chart, graph, or visualization lacks a text equivalent for trends, values, and conclusions.
A process, flow, map, or technical diagram is not described in a meaningful sequence.
Text rendered inside an image is not available as real text or equivalent replacement text.
Icons that communicate status, action, or category lack an accessible name or explanation.
Watermarks or background marks are announced as part of the document reading flow.
Recognized text over an image is inaccurate, duplicated, or ordered differently from the visible text.
A link annotation does not expose meaningful link text or alternate description.
Repeated text such as click here or read more does not identify the destination or action.
A clickable annotation exists without a corresponding logical Link structure element.
A logical Link structure exists but is not connected to an actionable PDF annotation.
The annotation rectangle is too small, too large, or offset from the visible link text.
The destination, named target, page reference, or external URL cannot be resolved.
Comments, popups, stamps, attachments, or other annotations do not expose a meaningful label.
Keyboard traversal through annotations does not follow a meaningful order.
An interactive field has no programmatic label, tooltip, or alternate name.
The label text near a form field is not connected to the field object.
A required input is visually indicated but not programmatically exposed.
Help text, format requirements, or constraints are not connected to the relevant field.
Error messages are not programmatically connected to the field they describe.
Related radio buttons do not expose the shared question or group label.
A checkbox has an ambiguous label or depends only on visual layout for meaning.
Combo box or list box options do not expose clear text labels.
Keyboard navigation through fields does not match the intended completion order.
Button controls do not expose the action they perform.
The PDF does not define a default natural language for assistive technology.
The document language value is malformed or not a recognized language tag.
Passages in another language are not tagged with their own language.
Fonts or glyphs do not map reliably to Unicode, causing incorrect reading, search, or copy behavior.
Ligatures or presentation glyphs do not expose the intended underlying characters.
Equations or formulas are not exposed with usable text, MathML, alternate text, or a clear explanation.
Symbols, units, or specialized notation lack expansion where pronunciation would be unclear.
Image-only pages contain text that is not available to assistive technology or text search.
Recognized text contains errors that change meaning or prevent reliable reading.
Characters or words are stored in an order that reads incorrectly despite appearing visually correct.
Status, grouping, required state, or meaning depends only on color without text or structural support.
Text and background colors may not provide enough contrast for readable presentation.
Meaningful icons, chart lines, form boundaries, or controls may not have enough contrast.
Text or symbols may be visually present but difficult to read due to small size or scaling.
Layering, clipping, transparency, or positioning makes content difficult to read visually or programmatically.
Focusable annotations or fields do not provide a visible focus state in common PDF viewers.
Chart categories or series cannot be distinguished without color perception.
Rasterized text, scans, or graphics are too degraded for reliable visual or OCR interpretation.
The PDF metadata does not provide a meaningful document title.
The viewer is not instructed to display the document title instead of the file name.
A long or complex document lacks bookmarks or outlines for efficient navigation.
Document outline entries are missing, mislabeled, or inconsistent with structural headings.
Logical page labels do not match visible page numbering or expected navigation.
The document opens at an inappropriate page, zoom, or pane for accessible use.
Metadata and catalog language values disagree or are inconsistently applied.
Right-to-left, vertical, or mixed-direction content lacks explicit direction information.
Embedded audio content has no text transcript.
Embedded video with speech has no synchronized captions.
Important visual-only information in video is not provided through audio description or text alternative.
Attachments or file annotations do not expose meaningful names, descriptions, or purpose.
Embedded 3D objects do not provide an equivalent description or fallback.
JavaScript actions interfere with keyboard access, reading, focus, or assistive technology behavior.
Animated or multimedia content may flash at unsafe frequencies or intensities.
Moving, auto-updating, or timed embedded content cannot be paused, stopped, or controlled.
Permissions or encryption prevent text extraction, screen reader access, copying for accessibility, or remediation.
Password, certificate, or launch behavior prevents accessible opening or reading.
The PDF depends on remote content, external files, or network access to convey accessible information.
Hidden or redacted content is still present in text, tags, metadata, or alternate descriptions.
Document metadata contains private data that should not be included in reports or remediation output.
The document does not declare PDF/UA conformance metadata.
The conformance claim is malformed, inconsistent, or not supported by the file structure.
Low-level syntax errors may prevent reliable parsing, validation, or assistive technology access.
Object references cannot be resolved reliably because xref data is missing or corrupt.
Fonts needed for consistent text rendering and Unicode mapping are missing.
A future remediation pass must not discard valid tags, metadata, alt text, or form labels already present.
The PDF version, declared standards, or metadata are insufficient to choose the correct validation strategy.
Page rotation metadata and visible content orientation disagree, causing awkward reading, OCR, or navigation.
The visible page crop clips text, controls, figures, or annotations that remain structurally present.
Printer marks, bleed content, or trim artifacts are exposed in ways that can confuse reading or remediation.
Unexpected portrait and landscape changes make navigation, magnification, and reflow harder.
Mixed page sizes create unpredictable zoom, scrolling, and output behavior.
A scanned page is tilted enough to reduce readability or OCR accuracy.
A scan is too dark, too light, washed out, or noisy for comfortable reading or reliable OCR.
Text, annotations, or tagged objects are positioned off the visible page or partly outside readable bounds.
Optional content groups or layers hide information that remains necessary for understanding.
Stacking order causes labels, highlights, or annotations to obscure or alter the meaning of content.
Opacity, blend modes, or overlays make text and graphics hard to perceive.
Crop marks, registration marks, color bars, or production notes are included in the accessible reading model.
The PDF file name does not help users identify the document, version, language, or purpose.
The opening content does not quickly explain what the document is for or who should use it.
A long or complex document has no usable contents list for orientation and navigation.
Contents entries do not link to the correct sections or cannot be activated reliably.
Dense or complex sections lack summaries that help users understand key points quickly.
Sentences, terminology, or structure may be unnecessarily hard to understand for the intended audience.
Specialized terms, acronyms, or internal labels are used without nearby explanation.
Procedures or tasks are written as dense prose instead of clear ordered steps.
Dates, deadlines, or time-sensitive actions are buried in surrounding text.
Users cannot easily determine what they need to do, by when, and through which channel.
The document does not provide a clear way to ask questions, request help, or get an accessible alternative.
Email addresses, phone numbers, or URLs are plain text when they should be usable links.
Users cannot tell whether the PDF is current, draft, archived, or superseded.
Alternative language versions exist but are not discoverable from the document.
A complex or hard-to-remediate PDF does not tell users how to get an accessible alternative.
Crowded content, weak grouping, or poor spacing makes the document harder to scan and understand.
Cards, panels, callouts, or grouped fields rely on visual proximity without structural relationships.
Important notices, warnings, or errors do not have clear text and structural treatment.
Statuses such as approved, rejected, overdue, optional, or complete are not explicit in text.
Instructions such as see left, below, red box, or above are used without structural or textual alternatives.
At high magnification, relationships between labels, values, captions, and sections become hard to follow.
Users must pan back and forth repeatedly to read normal text at larger zoom levels.
Paragraph lines are long enough to make tracking, magnification, and reading harder.
Meaning relies on seeing facing pages together, which may fail on small screens or assistive setups.
Notes are difficult to use when zoomed because references and note text are separated without navigation support.
Magnification or small-screen viewing makes callouts, labels, or legends hard to connect with their targets.
Interactive annotations, buttons, or form controls are difficult to activate accurately on touch devices.
Adjacent interactive areas increase the risk of accidental activation.
Viewer reflow presents content in a different sequence from the tagged structure.
Labels, fields, help text, and errors cannot be kept in context while filling out the form at high zoom.
Sections with the same visual role are exposed at different heading levels.
Repeated callouts, cards, lists, or tables are tagged differently across the document.
Links to the same destination use different accessible names without a meaningful reason.
Identical link names lead to different locations or actions, creating ambiguity.
Fields that ask for similar information use inconsistent names, labels, or instructions.
The same status or category is described with different words, icons, or colors across the document.
Repeated page regions appear in inconsistent positions or with inconsistent artifact treatment.
Bookmarks, table of contents entries, headings, and page labels do not use consistent naming.
The image purpose is contextual enough that deterministic checks cannot confirm whether the description is sufficient.
The layout is complex enough that tag order should be compared with page screenshots.
Complex, irregular, or nested table structures need review beyond simple header detection.
A chart summary may need domain knowledge to confirm that the right insight is described.
The document may be understandable only after a human reviews audience, terminology, and legal or technical constraints.
Some repeated or decorative-looking content may still be meaningful and should be reviewed before hiding it.
A form may be technically labeled but still confusing to complete without human usability review.
Automated remediation would affect meaningful structure or content and should require review before export.
The visual layout, OCR text, metadata, and structure tree disagree about the content meaning.
Suggested changes could alter regulated, contractual, or legal text and should be reviewed by an owner.
Fixing tags or metadata should not alter the visible document unless explicitly requested.
Automated fixes must not rewrite, delete, duplicate, or reorder visible text unintentionally.
Fixes must preserve link targets, destinations, and annotation actions.
Fixes must preserve form field values, validation, calculation, submission, and keyboard behavior.
Output generation must warn before invalidating or removing signatures, seals, or certification.
Output metadata must match the actual accessibility and PDF standards state after remediation.
Automated tagging should not duplicate content in the structure tree or reading stream.
Title, author, subject, language, identifiers, and other important metadata must be preserved or intentionally updated.
Output should avoid unnecessary rasterization, embedded duplicates, or bloated resources that harm download and use.
The remediated PDF must remain valid and readable in common PDF viewers and validators.
Findings and fixes should include enough evidence for a user to understand what changed and why.
Automated decisions should expose confidence, uncertainty, and whether human review is recommended.
The exported file name does not distinguish remediated, reviewed, draft, or original versions clearly.
The workflow does not make it clear which file is original, remediated, verified, or ready to publish.
The generated report itself must be structured, readable, and accessible.
Users cannot distinguish blocking issues, warnings, review items, and informational results.
Findings do not identify the relevant page, object, annotation, or structure element.
Users cannot inspect what changed between the original and remediated PDF.
Human-review findings must remain visible until accepted, dismissed, or resolved.
The exported PDF must be revalidated after fixes rather than assuming remediation succeeded.