What's Inside a PDF: Objects, Streams, and the Cross-Reference Table
A PDF is a graph of numbered objects, compressed streams, and an xref of byte offsets. Get Info inspects it; Repair rebuilds structure when that map breaks.

A PDF is not a flat page dump. It is a small database of objects, many stored as compressed streams, located by a cross-reference table (xref) so a reader can jump to object 17 without scanning the whole file. When that map breaks, viewers call the file damaged even if page bytes still sit on disk.
Objects
Everything interesting (pages, fonts, images, metadata, the document catalog) is an object with a number. Pages point at content streams and resources; resources point at fonts and XObjects; the catalog points at the page tree. Tools that merge, rotate, or stamp mostly rewrite that graph rather than redrawing pixels.
Streams
A stream is an object whose payload is a byte sequence, often Flate-compressed: page content operators, embedded font files, image data. Looking at a PDF in a text editor shows a mix of ASCII tokens and binary blobs for that reason. Expanding streams (without changing how pages look) is useful for debugging; Compress on this site recompresses those streams with qpdf—it does not downsample images or subset fonts.
The cross-reference table
Near the end of a typical file, an xref section (or an xref stream in newer files) lists byte offsets for each object number. Readers seek to those offsets. Truncated downloads, bad merges, and half-written saves often leave the trailer pointing at offsets that no longer match. The pixels of page one may still exist; the map that finds "page one" does not.
That is why structural repair is a different job from OCR or unlock: repair rebuilds internal structure from valid remnants; it does not invent missing page art or guess a password.
Inspect before you rewrite
Get Info on PDF returns a JSON report (page count, metadata, fonts, structural details), not a modified PDF. Use it when you need to know what you are holding before compress, convert, or repair.
If a viewer refuses to open the file at all, try Repair PDF, which rebuilds structure including cross-reference recovery. If the file opens but scanned text is not selectable, that is an OCR problem, not an xref problem.
A short decision order
- File will not open → Repair first.
- File opens; you need facts (pages, fonts, encryption hints) → Get Info.
- Structure is fine; you want a smaller file or readable text from a scan → Compress for size, or OCR for Markdown text, not another repair pass.
Both Get Info and Repair run without an account. For the same ops over HTTP, see Developers. For how Repair differs from branded desktop "recovery" products, see PDF Repair Online vs Recovery Toolbox searches.