Skip to content
@go-pdfkit

go-pdfkit

go-pdfkit

go-pdfkit

The whole of PDF in pure Go, with no C anywhere.

Read a file, write one, rearrange it, draw it, read it back as words and pictures,
and edit it with other people — in a browser tab if you like.

🌐 Website · 📚 Documentation

Docs License: BSD-3-Clause Go 1.26.4+ Coverage 100%


go-pdfkit is the whole of PDF in Go with zero C dependencies anywhere (CGO_ENABLED=0), and every library here builds for GOOS=js/wasm. Fonts are parsed and shaped by go-opentype; pages are rasterised by go-gfx.

// Take three pages out of a report, turn one the right way up, stamp the lot.
doc, _ := ops.Open(bytes)
doc.Select("4-6")
doc.SetRotation("2", 90)
doc.Watermark("all", "DRAFT")
out, _ := doc.Bytes()

// Draw a page, and read one back.
img, _ := render.Page(src, 1, render.Options{DPI: 150})
text, _ := extract.Text(src, 1)
$ pdfops merge out.pdf a.pdf b.pdf
$ pdfops nup -n 4 slides.pdf handout.pdf
$ pdfops encrypt -user letmein -allow print,copy plain.pdf locked.pdf
$ pdfops text -layout -pages 1 paper.pdf

Repositories

Repo What it is
reader reads and writes the format itself: objects, xref tables and streams, object streams, repair, every filter, encryption to AES-256, content streams, the page tree — and a writer that produces the same bytes twice
ops the verbs, and the pdfops command: merge, split, rotate, crop, n-up, booklet, watermark, stamp, sanitize, compress, encrypt, and reading a page back
render turns a page into pixels: paths, images, text in every font flavour, PDF functions, all seven shadings and tiling patterns
pdffont what a document says about a font — encodings, widths, and what text a code stands for
extract reads a page back: the text with where it sits, and the pictures with the box each covers
coedit a PDF several people edit at once — the plan is shared, not the file
app a PDF workbench that runs in a browser tab and nowhere else
pdfkit the document builder: pages, vector graphics, and text in embedded subsetted fonts
docs the documentation site (MkDocs Material, versioned with mike)
go-pdfkit.github.io the org landing page (Hugo)

Measured against 118 863 real files

Not against files written to pass a test. The corpus is arXiv's figures — from Matplotlib, Mathematica, pdfTeX, Ghostscript and Adobe — and every wave of work is checked by hashing the pixels of every page before and after and putting the biggest changes beside what the operating system's own renderer draws.

files that open 118 833 of 118 835 (the rest are PNGs named .pdf)
content operations read 1 536 769 753
rewritten to an identical fingerprint 118 833 of 118 833
compressed, page for page identical 118 833 — 35.2 GB → 33.6 GB
encrypted and decrypted both ways 118 833
pages drawn, no panics 4 108
pages read back as text 121 946 — 34 million characters
pictures located 300 685

That is what found a stroke that came out at half its colour, a font that took the dots off every i, a colour transform that turned every plot's paper yellow, a pattern that rubbed the page out, and every gradient inside a figure landing in the corner of the page rather than on the panel it was meant to fill. Each is written up in the release notes of the library it was found in.

Honest about scope

  • An ExtGState soft mask — a luminosity mask, the way Illustrator writes a faded surface — is not applied, so a page that uses one comes out with that part drawn at full strength. Constant alpha, /ca and /CA, is honoured.
  • A page whose text the document gives no way to read — a Type 3 dvips font with no /ToUnicode, whose glyphs are called a1 — comes back marked unreadable rather than guessed at. That is 1.89% of runs.
  • Eight of 14 614 embedded Type 1 programs have a private half that decrypts cleanly for eighty bytes and then does not: one byte of each was altered before it was embedded, and no reader can recover that.
  • Tagged PDF and PDF/A, forms and interactive annotations are not implemented.

Principles

  • Pure Go, zero cgo, zero C dependencies. No bundled libpoppler, libharfbuzz or libfreetype — cross-compiles anywhere Go runs, and all of it runs in a browser tab.
  • Built on go-opentype, not a second font stack: sfnt parsing and complex-script shaping (GSUB/GPOS) come from go-opentype/opentype, the same engine behind go-opentype and go-ruby-prawn.
  • Correctness checked against an independent parser — generated documents are re-opened with rsc.io/pdf and their structure verified; the embedded TrueType subset is re-parsed with go-opentype to confirm it still contains the glyphs that were drawn.
  • 100% test coverage, deterministic and network-free — tests use a synthesised TrueType font and a bundled OFL OpenType/CFF font, never a network fetch.

Status

Every library here is released, runs at 100% statement coverage including the branches that report a file saying something it should not, and builds for the six 64-bit architectures the fleet targets plus js/wasm, macOS and Windows.

BSD-3-Clause.

Popular repositories Loading

  1. pdfkit pdfkit Public

    Pure-Go, zero-C PDF 1.7 writer: font subsetting/embedding (TrueType+CFF), graphics, images, shaped text. 100% test coverage.

    Go

  2. go-pdfkit.github.io go-pdfkit.github.io Public

    Landing page for go-pdfkit — pure-Go, CGO-free PDF 1.7 writer

    HTML

  3. docs docs Public

    Documentation for go-pdfkit (MkDocs Material + mike)

  4. .github .github Public

    Org profile for go-pdfkit

  5. reader reader Public

    Pure-Go (CGO=0) PDF reader: lexer, object model, xref (tables + streams), filters, encryption, page tree — the parsing half of go-pdfkit

    Go

  6. ops ops Public

    Pure-Go (CGO=0) PDF operations: merge, split, extract, reorder, rotate, crop, watermark, metadata — the verbs, over go-pdfkit/reader

    Go

Repositories

Showing 10 of 13 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…