Trade Republic statements all download as document-N.pdf, which says
nothing about what they hold. traderepublic now implements parser.Namer:
it reads the period from the address block at the top of the first page
("DATE 01 May 2025 - 31 Jul 2026") and names the statement
traderepublic_2025_05_01_2026_07_31, lowercase like the NLB names, so the
folder sorts by start date and two downloads of one statement meet under
one name. The transaction table's own DATE heading never matches, since
no date range follows it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
21 KiB
money
Statement-driven personal finance tracker. A data directory holds one folder per
account; statements dropped into those folders are parsed into a rebuildable
SQLite index, categorised by ordered glob rules, and browsed in a web app
(money serve, which is also the default command).
Read README.md first — it is the user-facing reference for every config key.
This file covers what the code assumes and why.
Layout
cmd/money/main.go subcommands; serve is the default
internal/config rules.toml (rules + transfers), account.toml, XDG config, data-root resolution
internal/model Account, Transaction, amount formatting, description normalisation
internal/glob the `*` / `?` matcher used by rules (two-pointer, no exponential blowup)
internal/parser Parser interface + registry; nlb, revolut, traderepublic
internal/store SQLite index (modernc.org/sqlite, no cgo)
internal/importer directory walk, dedupe, balance checks
internal/rules applies ordered rules, writing rule_tag
internal/transfers pairs the two legs of a movement, writing the transfers table
internal/report per-tag aggregation
internal/web `money serve`: JSON API + embedded single page (static/)
Invariants
Every tag comes from rules.toml. There is no way to tag a transaction by
hand; rule_tag is derived state that Engine.Retag rewrites wholesale, which
is what makes money retag safe to run at any time. Nothing may write a tag
from anywhere else — the moment something does, the index stops being
reproducible and retag stops being safe. Covered by TestRetagRewritesEveryTag.
A transfer is a pair or it is nothing. A [[transfer]] block names both
legs (from_account + from_desc, to_account + to_desc), and only a
matched pair is dropped from the report — both legs together, never one. This
is the whole reason the earlier boolean transfer flag was removed and this
replaced it: a one-sided verdict let half a movement vanish and left the report
unbalanced. An unmatched leg therefore keeps counting, and is surfaced as a
warning instead: ⚠ on the transfers screen, on stderr from import/retag.
Nothing may start excluding a single leg on the strength of one side matching.
Legs pair within transfers.WindowDays, nearest date first. Within one
currency the amount is the evidence and must be the exact opposite, unless
the definition sets tolerance_pct — a per-definition opt-in for a route where
the bank takes a fee on the way, so the two statements genuinely disagree. It
defaults to zero and belongs on the one definition that charges: a global or
default tolerance would loosen every route that does not. Across currencies
the amount is not checked at all: there are no exchange rates here, so the two
numbers are unrelated and the dates carry the pairing alone. A tolerance means
nothing there and is ignored. That asymmetry is the design, not an oversight —
do not "fix" the cross-currency case by inventing a rate, and do not turn the
tolerance into the default. Like rules, the first definition claims a leg and a
transaction belongs to at most one transfer.
A tolerated mismatch is a fee, and a fee is money, so it is reported rather
than forgiven. That is the condition on the tolerance existing at all: the
pair leaves the report entirely, so a difference swallowed inside one would be
spending that never appears anywhere. Pair.Fee is what left less what
arrived, and report.Excluded carries it out per currency alongside the legs —
named on the money report line and on its own row under the web report.
Nothing may pair on a mismatch without that difference reaching Excluded.
It is a fee and not a tag: no rule produces it, report.ByTag never sees it,
and it must not be turned into a synthetic transaction to make the total
reconcile. Excluded counts it only for pairs whose legs are both in view,
for the same reason it counts legs and not pairs — half a pair cannot say what
the other half received.
(transfer) in a tag column is display only. Transaction.DisplayTag
falls back to model.TransferTag for a paired leg so the column does not read
as a blank waiting for a rule, but nothing writes it: rule_tag stays what
rules.toml made it, store.Tags never returns it, and no rule can be built
from it. It is the same kind of label as report.Untagged. Do not "persist"
it — that is precisely the second verdict this design exists to avoid.
A paired leg is not untagged. store.Filter{Untagged: true} means "no tag
and no transfer", so a leg never turns up in money ls --untagged, the u
view or the rule builder's preview asking to be tagged — it is spoken for, and
the report drops it regardless. Covered by TestPairedLegsAreNotUntagged. An
unpaired leg is untagged like anything else, which is how it gets noticed.
The pairing is derived state, exactly like the tags. Engine.Link rewrites
the whole transfers table from rules.toml, so money retag re-derives both
halves of what that file decides. It runs over the whole index deliberately —
pairing within a filtered view would let a movement count as a transfer in one
report and not in another. Note the asymmetry with a schema column: a new
table is created by CREATE TABLE IF NOT EXISTS in store.Open, so adding one
does not force an index rebuild.
Rules and transfers share rules.toml, so textual deletion is block-aware.
config.deleteBlocks finds every [[...]] header, not only the kind being
deleted, because a block ends where the next block of any kind begins —
otherwise deleting a rule would swallow a transfer that follows it. Covered by
TestDeleteLeavesTheOtherKindAlone.
Statements are the source of truth; the index is disposable. Deleting
index.db at the root of the data directory and re-importing must reproduce
everything, with nothing lost.
The index is a cache, so there are no migrations. store.Open applies
schema and nothing else — no ALTER TABLE, no version column, no repair of
an index an older build wrote. Changing the schema costs one line in the
release note: delete index.db and import again. The two directions are not
symmetric, which is worth knowing before you assume something is broken:
- Removing a column leaves an older index still working, carrying the dead column and its data unread.
- Adding one the code reads breaks every command against an older index with
query transactions: SQL logic error: no such column: …until it is deleted.
That asymmetry holds only because every statement names its columns —
no SELECT *, and nothing may depend on column order. Keep it that way.
Do not reintroduce migration machinery either; if re-parsing ever becomes too
expensive to ask for, that is a decision to revisit deliberately rather than a
helper to slip back in.
Dedupe is by fingerprint: sha256(date | amount | normalised description | ordinal), where the ordinal distinguishes identical lines within one
statement. Two identical purchases on one day both survive; the same line in an
overlapping statement does not duplicate. Unchanged files are skipped by
checksum before parsing at all — so an import with nothing new finishes in
milliseconds. That is the checksum skip working, not a failure.
Money is int64 minor units, never a float. Per-account currency, no
conversion, and totals are never summed across currencies.
The most specific rule wins, not the topmost. rules.Engine sorts the
rules once in New and matches in that order: most literal characters first,
then fewest *, then account-scoped over unscoped, with a stable sort so
equally specific rules keep file order and the earlier one still wins. That is
what lets *NIKOLA* carve an exception out of *NIK* from anywhere in the
file, and a catch-all * sit wherever it reads best. A rule setting both
match and type requires both of them, and both count towards its literals.
The two orders must not be confused. Engine.rules stays in file order and
Rules(), Usage and MatchIndex all speak in file positions, because that
is what the rules screen numbers, what config.DeleteRules deletes by, and
what the user can point at in rules.toml; only Engine.order is sorted.
Anything new that reports a rule must report its file position too.
config.AppendRule still appends rather than prepends, but that now only
settles ties: a saved rule cannot displace an equally specific one written by
hand, while a narrower one is meant to take precedence and does.
The report's period narrows the report and nothing else. The screen opens on
last month, and ←/→ step along report.Periods — the named windows, then
the months the index holds. The report and the transaction list share a scope
(account, search, untagged), which web.filterFrom reads and state.filter
holds in the page; the period is read by the report handler alone and kept in
state.report, so opening on last month cannot hide the rest of the index from
the list. Nothing may put the period into that shared scope, which is exactly
what would make it leak. Covered by TestReportPeriodLeavesTransactionsAlone.
The axis is built from store.Months over the whole index, not from the rows
in view, or it would grow and shrink as the account or search filter changed and
move under the cursor. It is rebuilt on every request (an import can reach
further back), and the page names the period by its label rather than its
position, so a grown axis keeps the window the user was on. The windows are
relative to today, never to the newest statement: "last month" with nothing in
it reports nothing and says so, because silently answering for a month nobody
asked for is worse than an empty screen. That is also why an empty period and
an empty index give different messages — one asks for another period, the
other for an import.
The report's sort rearranges rows; it never changes which rows there are.
report.Order is passed to ByTag and is deliberately not part of
store.Filter or the shared scope — the period decides what is counted, the
order only how it is listed, and merging the two would make a sort able to hide
a tag. Currency stays the outer sort key under every order, because there are no
exchange rates to compare two currencies by, and every order falls back to the
tag so ties keep a fixed position instead of shuffling between reloads.
Order.Column is what lets the page mark the sorted heading without keeping its
own copy of that mapping; a new order needs a column of its own, which
TestOrderColumnsAreDistinct checks.
The single-key shortcuts must not fire while typing. The page's keydown
handler ignores them whenever the target is an input, or typing i in the tag
field would start an import. A view's own key hook runs first and is told
whether the user is typing, so a command inside a builder has to be a chord —
ctrl+s re-sorts the preview — because every printable key belongs to the
field being typed in.
A rule's usage count is how many transactions it wins, not how many its glob
could match — Engine.Usage counts by MatchIndex, so a rule shadowed by a
more specific one correctly reports zero. That is what makes the rules screen able to
find dead rules at all.
A rule's note is documentation that round-trips. It is a TOML key rather
than a # comment so LoadRules can return it, the builder can write it and
the rules screen can show it. It never takes part in matching — rules.Engine
does not look at it — and it must stay that way.
config.DeleteRules and config.ReplaceRule edit rules.toml textually, never
by re-serialising the parsed rules, because comments and formatting are not
recoverable from []Rule. A rule owns the comment lines directly above it, so
deleting takes them with it; a comment separated by a blank line is a heading
for what follows and stays. Replacing keeps them — they say why the rule is
there, which editing its glob rarely changes — and rewrites only the rule's own
lines, leaving its position alone: position still breaks ties, so a rule that
moved could start beating an equally specific one it never used to. Every
writer goes through writeFileAtomic, and both editors re-parse the result
before replacing the file.
An edit must not lose what the builder does not show. The form has four
fields and a Rule has five, so editRule takes Type from the rule on disk —
never from the request — and the form says it is keeping it. Nothing may
round-trip a rule through those four fields alone — a pattern the user was
never shown is not a pattern they chose to remove. A new field on Rule needs
the same treatment or a field of its own. Covered by TestEditKeepsTypePattern.
The rule builder's preview is what the rule is judged against, which is not
always "what is untagged". For a new rule those are the same thing. For an
edit they are not: the rule's own transactions are tagged, so an untagged-only
preview would be empty for a rule that works. previewGroups adds them back,
and rulePreview keeps the ones the new glob stops catching on screen marked
− instead of dropping them silently, because giving one up is the decision
being made. Covered by TestEditPreviewShowsWhatIsLetGo.
The web app is a frontend, not a second implementation. Everything that
decides something stays in Go: amounts are formatted server-side (money is
text plus a sign, never a number the browser could add up), the rule builder's
preview is matched by glob.Match on the server rather than re-implemented in
JS, and the transfer preview runs transfers.Analyze. Keep it that way — a
JS glob or a JS sum is a second copy of a rule that would drift from the
first.
static/app.js only lays out what the API returns.
The server outlives edits to rules.toml in a way a one-shot command does not,
so /api/retag and /api/import re-read it first (import also re-reads the
account folders), and /api/overview reports stale when the file on disk no
longer equals what the engines hold. Every edit and delete sends the rule or
transfer it showed alongside the position, and is refused with 409 unless
rules.toml still holds exactly that there — a position from a stale page can
otherwise name a different rule. Covered by TestStalePositionIsRefused.
There is no auth by design (--addr defaults to loopback). Non-GET requests
must be application/json, which is what keeps a cross-site form from posting
to it; do not relax that without putting something else in its place.
That is why /api/upload takes files base64 inside JSON rather than as
multipart: multipart is precisely what a cross-site form can send.
An upload only puts a file where money import looks. It writes into an
existing account folder and stops there: importing stays the user's call, made
with the Import button, so the statements stay the source of truth and nothing
reaches the index any other way. It never overwrites a statement (identical
contents are a no-op, different ones a 409), refuses any name import would not
read back — not a plain file name, a dotfile, account.toml, outside the
account's include — and checks a whole batch before writing any of it. Covered
by TestUploadSavesWithoutImporting, TestUploadNeverReplacesAStatement and
TestUploadRefusesBadNames.
The statements list is read from the folders, not the index. /api/files
walks importer.StatementFiles — the same list import reads — and only then
looks each file up in store.SourceFiles, comparing importer.Checksum, so a
file is listed exactly when import would read it and a file the index remembers
but the disk lost shows as missing instead of vanishing.
/api/files/{account}/{name} serves a file only by finding it in that list,
never by joining the name onto a path. A statement is served from the app's
origin, where a script could drive the API, so nothing is ever rendered as a
page: text is text/plain under CSP: sandbox, anything not text or PDF is a
sandboxed download, and PDFs — whose viewers refuse a sandbox — open in the
browser's own isolated viewer. Covered by TestServeFileServesOnlyStatements.
Deleting a statement marks its neighbours to be re-read. /api/files/delete
finds the file the same way serveFile does and deletes it — permanently, by
the user's choice; the page says so before it asks. store.ForgetSourceFile
then drops its transactions and their transfer rows, and Link re-pairs, which
reads no statement. It does not import, any more than an upload does. A
transaction two overlapping statements share is stored once, under the file
imported first, so it goes too; ForgetSourceFile therefore blanks the
checksums of the account's other statements, which show as changed until the
next Import re-reads them and restores what they hold. Without that the checksum
skip would leave those rows gone for good. Covered by
TestRemoveFileKeepsWhatAnotherStatementHolds.
Adding a bank parser
Implement parser.Parser and call parser.Register from an init. Nothing
else changes; the name becomes usable in an account.toml. Use
parser.ParseAmount rather than hand-rolling decimal handling — it copes with
1.234,56, trailing minus, parenthesised negatives, currency codes and both
the ASCII hyphen and U+2212.
Keep PDF text extraction separate from parsing: the bank parsers expose a pure
parseXText(text string, digits int) so they can be tested against captured
pdftotext output without a PDF fixture. pdftotext -layout (poppler-utils) is
a runtime dependency of nlb and traderepublic; no Go library reconstructs
column layout as well.
It can be carried inside the binary instead: scripts/build-bundled.sh builds
a static musl pdftotext in a container and embeds it under -tags bundled,
and pdftotextCommand writes it to the user cache dir (named by content hash,
verified before reuse) and runs that. It is still the real executable run as a
subprocess, deliberately — not poppler linked in through cgo, which would cost
the pure-Go build, and not a different extraction API whose spacing might not
match what the parsers were tuned on. The tag is opt-in so go build/go test
never need the 5 MB file, which is gitignored rather than committed. Bumping
POPPLER_VERSION means updating its pinned SHA-256 too, and checking real
statements still parse — that output is what the parsers depend on.
Do not hardcode absolute column positions from a sample PDF. pdftotext
compresses runs of spaces, so columns shift with font and page size — derive
positions from the header line (traderepublic) or from the line being parsed
(nlb).
The other side's account number belongs in the description, not in a field
of its own. There used to be a Counterparty field; only nlb could fill it
honestly, revolut and traderepublic invented it with an IBAN-shaped regex
over the description, and the two disagreed on spacing, so one rule pattern
could not serve both. Now nlb appends its IBAN column to the end of the
description — at the end, and not in the position it held on the page, so an
IBAN wrapped across continuation lines stays contiguous for a glob to match.
A new parser must do the same rather than reintroduce a structured field.
A parser may name its statements, and only uploads use it. parser.Namer is
optional, like parser.Warner: nlb names an izpisek izpisek_YYYY_MM_DD
after its Datum izpiska, ported from the rename_izpiski.py it replaced, and
traderepublic names a statement traderepublic_YYYY_MM_DD_YYYY_MM_DD after
the period in its header. A name is derived from what the statement says, never
from the upload's own name, and comes back without an extension; the upload's is
kept, lowercased. /api/upload stages each file as a dotfile — invisible to
import — so the parser can read it, then numbers a chosen name that is taken by
different contents (_2, as the script did) where a name the user gave is
refused with 409. Nothing renames a file already in a folder: source_files
records statements by path, so a rename would orphan its rows as missing until
the index is rebuilt. Covered by TestUploadNamesStatementsTheParserCanName.
Verifying
go build ./... && go test ./... && go vet ./... && gofmt -l .
For end-to-end checks, build a throwaway data root under the scratchpad rather than touching real data:
money --root /tmp/.../demo import
money --root /tmp/.../demo import # must report 0 new
money --root /tmp/.../demo ls --wide
If you changed the schema, delete that root's index.db first. Nothing
migrates it, so a demo root left over from an earlier build either carries dead
columns or fails with no such column, depending on which way the schema
moved.
Testing the web app
Drive the API through Server.Handler() with httptest — newTestServer in
internal/web/server_test.go builds a data root and an index tagged and paired
as an import would leave them, with "today" fixed so the report's default
window is predictable. That is how every web test works, and since the page
only lays out what the API returns, it is where behaviour gets tested.
For a look at the real page, run money --root <demo> serve --addr 127.0.0.1:<port> against a throwaway root. Delete the index first if you want
to watch an import do work, or it will be skipped by checksum and finish
instantly.
Conventions
Comments explain why, not what. Errors name the file and the offending row or key so a bad statement is actionable. Per-file import failures are reported and the run continues; only a broken data root aborts.