Trade Republic statements all download as document-N.pdf, which says
nothing about what they hold. traderepublic now implements parser.Namer:
it reads the period from the address block at the top of the first page
("DATE 01 May 2025 - 31 Jul 2026") and names the statement
traderepublic_2025_05_01_2026_07_31, lowercase like the NLB names, so the
folder sorts by start date and two downloads of one statement meet under
one name. The transaction table's own DATE heading never matches, since
no date range follows it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Four changes to statement handling in the web app, made together and
touching the same upload and statements-list code.
Name NLB uploads by statement date. parser.Namer is an optional
interface, like Warner, through which a parser names its statements;
nlb reads the "Datum izpiska" from the izpisek header and names it
izpisek_YYYY_MM_DD, lowercase, extension included -- ported from the
rename_izpiski.py it replaces. Uploads are staged as dotfiles, invisible
to import, so the parser can read them; two downloads of one statement
then meet under one name and the second is recognised as already there,
while a different statement of the same date is numbered _2 as the
script did. Only uploads are named: source_files records statements by
path, so renaming a file already in a folder would orphan its rows.
Delete a statement from the statements list. The file is removed from
disk for good -- the page says so before it asks -- and
store.ForgetSourceFile drops its transactions and their transfer rows.
A row two overlapping statements share is stored once, under the file
imported first, so it goes too; the account's other statements forget
their checksums and show as changed until the next Import re-reads them
and restores it. A file already gone from disk can be forgotten.
Upload and delete no longer import. Importing stays the user's call,
made with the Import button, so a batch can be put together and looked
over first. Delete still re-pairs transfers, which reads no statement.
Show rows and new rows per statement. The list read "0" for a file
whose rows an earlier, overlapping statement already held, which looked
like a file that failed to parse. source_files now records how many
transactions each statement holds, and the list reads "3 rows · 0 new".
This adds a column the code reads, so an index built by an earlier
version fails with "no such column: s.rows": delete index.db and import
again.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The PDF parsers shell out to pdftotext, so every machine running money needed
poppler-utils installed. scripts/build-bundled.sh now builds a deployable
binary that carries its own: pdftotext is compiled in a container from a
checksum-pinned poppler release as a fully static musl executable, then
embedded with `go build -tags bundled`. The result is one file that runs on
any Linux of that architecture with nothing installed alongside it.
It is still the real pdftotext, run as a subprocess. Linking poppler through
cgo would have cost the pure-Go build, and its C++ text API is not guaranteed
to space columns the way pdftotext -layout does, which is what the parsers
were tuned on. Only what text extraction needs is compiled in -- no
fontconfig, cairo or image codecs -- and its output is byte-identical to a
full distro build on the same PDF.
At runtime the embedded copy is written to the user cache directory, not
/tmp, which servers often mount noexec. It is named by content hash, so a
newer build never runs an older copy, and verified before reuse, so a write cut
short by a killed process is replaced rather than trusted. `money config` says
which pdftotext is in use.
The tag is opt-in: plain go build and go test never need the 5 MB executable,
which is gitignored rather than committed. Building with the tag for anything
but linux/amd64 or linux/arm64 fails with a message saying so.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
counterparty was a structured field only nlb could fill honestly. revolut
and traderepublic invented one by running an IBAN-shaped regex over the
description they had just built, and the two spellings disagreed --
SI56 1234 5678 9012 345 against SI56123456789012345 -- so a literal rule
pattern that worked on one account silently matched nothing on another. It
is gone from the model, the index, the rule keys, ls --wide and the rules
screen. nlb now appends its IBAN column to the end of the description,
where the other two already keep theirs, so match = "*SI56*" works
everywhere. That changes those descriptions and with them their
fingerprints, so a statement overlapping an already-imported period will
re-add rather than dedupe those rows until the index is rebuilt. An index
built by an older binary drops the column when it is opened.
The index itself moves from .money/index.db up to index.db beside
rules.toml. Nothing looks in the old location, so an existing one has to be
moved by hand -- otherwise the tool quietly starts a fresh index and the
manual tags in the old file, the only thing statements cannot reproduce,
stay behind in it.
The csv and cmd parsers are gone along with the [csv] and [cmd] config they
carried. cmd shelled out to the Python extractors, which were ported to Go
and deleted, so it bridged to nothing; csv was a generic column-mapped
fallback that no account used, and between them they were the largest
configuration surface in the tool. A bank is now described in Go, where it
can be tested. The importer tests register their own three-column parser
rather than borrow a bank's, so they stay about the directory walk, dedupe
and per-file error reporting.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A data directory holds one folder per account. Statements dropped into
those folders are parsed into a rebuildable SQLite index, categorised by
ordered glob rules in rules.toml, and browsed or hand-tagged in a Bubble
Tea TUI. Movements between the user's own accounts are marked as
transfers by the same rules and excluded from spending totals.
Manual tags and transfer marks are stored separately from the rule-derived
ones and always win, so editing rules.toml and re-running retag never
destroys hand edits.
Parsers are pluggable. Three are ported from the Python extractors they
replace -- nlb and traderepublic read PDFs via pdftotext -layout, revolut
reads the CSV export -- alongside a configurable-column CSV parser and a
cmd parser that shells out to an external script.
Both ports fix two latent bugs in the originals: the sign character class
rejected the typographic minus U+2212 that some PDF fonts emit, and NLB's
hardcoded continuation indent broke when pdftotext compressed runs of
spaces, so the threshold is now measured from the description column.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>