Bundle a static pdftotext into the binary

The PDF parsers shell out to pdftotext, so every machine running money needed
poppler-utils installed. scripts/build-bundled.sh now builds a deployable
binary that carries its own: pdftotext is compiled in a container from a
checksum-pinned poppler release as a fully static musl executable, then
embedded with `go build -tags bundled`. The result is one file that runs on
any Linux of that architecture with nothing installed alongside it.

It is still the real pdftotext, run as a subprocess. Linking poppler through
cgo would have cost the pure-Go build, and its C++ text API is not guaranteed
to space columns the way pdftotext -layout does, which is what the parsers
were tuned on. Only what text extraction needs is compiled in -- no
fontconfig, cairo or image codecs -- and its output is byte-identical to a
full distro build on the same PDF.

At runtime the embedded copy is written to the user cache directory, not
/tmp, which servers often mount noexec. It is named by content hash, so a
newer build never runs an older copy, and verified before reuse, so a write cut
short by a killed process is replaced rather than trusted. `money config` says
which pdftotext is in use.

The tag is opt-in: plain go build and go test never need the 5 MB executable,
which is gitignored rather than committed. Building with the tag for anything
but linux/amd64 or linux/arm64 fails with a message saying so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-10-02 18:02:46 +02:00
co-authored by Claude Opus 5.5
parent 6faf99719a
commit 5847245638
12 changed files with 398 additions and 3 deletions
+11
View File
@@ -261,6 +261,17 @@ Keep PDF text extraction separate from parsing: the bank parsers expose a pure
a runtime dependency of `nlb` and `traderepublic`; no Go library reconstructs
column layout as well.
It can be carried inside the binary instead: `scripts/build-bundled.sh` builds
a static musl `pdftotext` in a container and embeds it under `-tags bundled`,
and `pdftotextCommand` writes it to the user cache dir (named by content hash,
verified before reuse) and runs that. It is still the real executable run as a
subprocess, deliberately — not poppler linked in through cgo, which would cost
the pure-Go build, and not a different extraction API whose spacing might not
match what the parsers were tuned on. The tag is opt-in so `go build`/`go test`
never need the 5 MB file, which is gitignored rather than committed. Bumping
`POPPLER_VERSION` means updating its pinned SHA-256 too, and checking real
statements still parse — that output is what the parsers depend on.
Do not hardcode absolute column positions from a sample PDF. `pdftotext`
compresses runs of spaces, so columns shift with font and page size — derive
positions from the header line (`traderepublic`) or from the line being parsed