Files
money/CLAUDE.md
T
nikolaandClaude Opus 5.5 ed52be3f1d Name NLB uploads izpisek_YYYY-MM-DD
Back to the dashed date rename_izpiski.py used, so a statement the script
already named is recognised as already there when it is uploaded again,
rather than landing as a second copy under the underscored name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 22:49:41 +02:00

21 KiB
Raw Blame History

money

Statement-driven personal finance tracker. A data directory holds one folder per account; statements dropped into those folders are parsed into a rebuildable SQLite index, categorised by ordered glob rules, and browsed in a web app (money serve, which is also the default command).

Read README.md first — it is the user-facing reference for every config key. This file covers what the code assumes and why.

Layout

cmd/money/main.go     subcommands; serve is the default
internal/config       rules.toml (rules + transfers), account.toml, XDG config, data-root resolution
internal/model        Account, Transaction, amount formatting, description normalisation
internal/glob         the `*` / `?` matcher used by rules (two-pointer, no exponential blowup)
internal/parser       Parser interface + registry; nlb, revolut, traderepublic
internal/store        SQLite index (modernc.org/sqlite, no cgo)
internal/importer     directory walk, dedupe, balance checks
internal/rules        applies ordered rules, writing rule_tag
internal/transfers    pairs the two legs of a movement, writing the transfers table
internal/report       per-tag aggregation
internal/web          `money serve`: JSON API + embedded single page (static/)

Invariants

Every tag comes from rules.toml. There is no way to tag a transaction by hand; rule_tag is derived state that Engine.Retag rewrites wholesale, which is what makes money retag safe to run at any time. Nothing may write a tag from anywhere else — the moment something does, the index stops being reproducible and retag stops being safe. Covered by TestRetagRewritesEveryTag.

A transfer is a pair or it is nothing. A [[transfer]] block names both legs (from_account + from_desc, to_account + to_desc), and only a matched pair is dropped from the report — both legs together, never one. This is the whole reason the earlier boolean transfer flag was removed and this replaced it: a one-sided verdict let half a movement vanish and left the report unbalanced. An unmatched leg therefore keeps counting, and is surfaced as a warning instead: ⚠ on the transfers screen, on stderr from import/retag. Nothing may start excluding a single leg on the strength of one side matching.

Legs pair within transfers.WindowDays, nearest date first. Within one currency the amount is the evidence and must be the exact opposite, unless the definition sets tolerance_pct — a per-definition opt-in for a route where the bank takes a fee on the way, so the two statements genuinely disagree. It defaults to zero and belongs on the one definition that charges: a global or default tolerance would loosen every route that does not. Across currencies the amount is not checked at all: there are no exchange rates here, so the two numbers are unrelated and the dates carry the pairing alone. A tolerance means nothing there and is ignored. That asymmetry is the design, not an oversight — do not "fix" the cross-currency case by inventing a rate, and do not turn the tolerance into the default. Like rules, the first definition claims a leg and a transaction belongs to at most one transfer.

A tolerated mismatch is a fee, and a fee is money, so it is reported rather than forgiven. That is the condition on the tolerance existing at all: the pair leaves the report entirely, so a difference swallowed inside one would be spending that never appears anywhere. Pair.Fee is what left less what arrived, and report.Excluded carries it out per currency alongside the legs — named on the money report line and on its own row under the web report. Nothing may pair on a mismatch without that difference reaching Excluded. It is a fee and not a tag: no rule produces it, report.ByTag never sees it, and it must not be turned into a synthetic transaction to make the total reconcile. Excluded counts it only for pairs whose legs are both in view, for the same reason it counts legs and not pairs — half a pair cannot say what the other half received.

(transfer) in a tag column is display only. Transaction.DisplayTag falls back to model.TransferTag for a paired leg so the column does not read as a blank waiting for a rule, but nothing writes it: rule_tag stays what rules.toml made it, store.Tags never returns it, and no rule can be built from it. It is the same kind of label as report.Untagged. Do not "persist" it — that is precisely the second verdict this design exists to avoid.

A paired leg is not untagged. store.Filter{Untagged: true} means "no tag and no transfer", so a leg never turns up in money ls --untagged, the u view or the rule builder's preview asking to be tagged — it is spoken for, and the report drops it regardless. Covered by TestPairedLegsAreNotUntagged. An unpaired leg is untagged like anything else, which is how it gets noticed.

The pairing is derived state, exactly like the tags. Engine.Link rewrites the whole transfers table from rules.toml, so money retag re-derives both halves of what that file decides. It runs over the whole index deliberately — pairing within a filtered view would let a movement count as a transfer in one report and not in another. Note the asymmetry with a schema column: a new table is created by CREATE TABLE IF NOT EXISTS in store.Open, so adding one does not force an index rebuild.

Rules and transfers share rules.toml, so textual deletion is block-aware. config.deleteBlocks finds every [[...]] header, not only the kind being deleted, because a block ends where the next block of any kind begins — otherwise deleting a rule would swallow a transfer that follows it. Covered by TestDeleteLeavesTheOtherKindAlone.

Statements are the source of truth; the index is disposable. Deleting index.db at the root of the data directory and re-importing must reproduce everything, with nothing lost.

The index is a cache, so there are no migrations. store.Open applies schema and nothing else — no ALTER TABLE, no version column, no repair of an index an older build wrote. Changing the schema costs one line in the release note: delete index.db and import again. The two directions are not symmetric, which is worth knowing before you assume something is broken:

  • Removing a column leaves an older index still working, carrying the dead column and its data unread.
  • Adding one the code reads breaks every command against an older index with query transactions: SQL logic error: no such column: … until it is deleted.

That asymmetry holds only because every statement names its columns — no SELECT *, and nothing may depend on column order. Keep it that way. Do not reintroduce migration machinery either; if re-parsing ever becomes too expensive to ask for, that is a decision to revisit deliberately rather than a helper to slip back in.

Dedupe is by fingerprint: sha256(date | amount | normalised description | ordinal), where the ordinal distinguishes identical lines within one statement. Two identical purchases on one day both survive; the same line in an overlapping statement does not duplicate. Unchanged files are skipped by checksum before parsing at all — so an import with nothing new finishes in milliseconds. That is the checksum skip working, not a failure.

Money is int64 minor units, never a float. Per-account currency, no conversion, and totals are never summed across currencies.

The most specific rule wins, not the topmost. rules.Engine sorts the rules once in New and matches in that order: most literal characters first, then fewest *, then account-scoped over unscoped, with a stable sort so equally specific rules keep file order and the earlier one still wins. That is what lets *NIKOLA* carve an exception out of *NIK* from anywhere in the file, and a catch-all * sit wherever it reads best. A rule setting both match and type requires both of them, and both count towards its literals.

The two orders must not be confused. Engine.rules stays in file order and Rules(), Usage and MatchIndex all speak in file positions, because that is what the rules screen numbers, what config.DeleteRules deletes by, and what the user can point at in rules.toml; only Engine.order is sorted. Anything new that reports a rule must report its file position too.

config.AppendRule still appends rather than prepends, but that now only settles ties: a saved rule cannot displace an equally specific one written by hand, while a narrower one is meant to take precedence and does.

The report's period narrows the report and nothing else. The screen opens on last month, and ←/→ step along report.Periods — the named windows, then the months the index holds. The report and the transaction list share a scope (account, search, untagged), which web.filterFrom reads and state.filter holds in the page; the period is read by the report handler alone and kept in state.report, so opening on last month cannot hide the rest of the index from the list. Nothing may put the period into that shared scope, which is exactly what would make it leak. Covered by TestReportPeriodLeavesTransactionsAlone.

The axis is built from store.Months over the whole index, not from the rows in view, or it would grow and shrink as the account or search filter changed and move under the cursor. It is rebuilt on every request (an import can reach further back), and the page names the period by its label rather than its position, so a grown axis keeps the window the user was on. The windows are relative to today, never to the newest statement: "last month" with nothing in it reports nothing and says so, because silently answering for a month nobody asked for is worse than an empty screen. That is also why an empty period and an empty index give different messages — one asks for another period, the other for an import.

The report's sort rearranges rows; it never changes which rows there are. report.Order is passed to ByTag and is deliberately not part of store.Filter or the shared scope — the period decides what is counted, the order only how it is listed, and merging the two would make a sort able to hide a tag. Currency stays the outer sort key under every order, because there are no exchange rates to compare two currencies by, and every order falls back to the tag so ties keep a fixed position instead of shuffling between reloads. Order.Column is what lets the page mark the sorted heading without keeping its own copy of that mapping; a new order needs a column of its own, which TestOrderColumnsAreDistinct checks.

The single-key shortcuts must not fire while typing. The page's keydown handler ignores them whenever the target is an input, or typing i in the tag field would start an import. A view's own key hook runs first and is told whether the user is typing, so a command inside a builder has to be a chord — ctrl+s re-sorts the preview — because every printable key belongs to the field being typed in.

A rule's usage count is how many transactions it wins, not how many its glob could match — Engine.Usage counts by MatchIndex, so a rule shadowed by a more specific one correctly reports zero. That is what makes the rules screen able to find dead rules at all.

A rule's note is documentation that round-trips. It is a TOML key rather than a # comment so LoadRules can return it, the builder can write it and the rules screen can show it. It never takes part in matching — rules.Engine does not look at it — and it must stay that way.

config.DeleteRules and config.ReplaceRule edit rules.toml textually, never by re-serialising the parsed rules, because comments and formatting are not recoverable from []Rule. A rule owns the comment lines directly above it, so deleting takes them with it; a comment separated by a blank line is a heading for what follows and stays. Replacing keeps them — they say why the rule is there, which editing its glob rarely changes — and rewrites only the rule's own lines, leaving its position alone: position still breaks ties, so a rule that moved could start beating an equally specific one it never used to. Every writer goes through writeFileAtomic, and both editors re-parse the result before replacing the file.

An edit must not lose what the builder does not show. The form has four fields and a Rule has five, so editRule takes Type from the rule on disk — never from the request — and the form says it is keeping it. Nothing may round-trip a rule through those four fields alone — a pattern the user was never shown is not a pattern they chose to remove. A new field on Rule needs the same treatment or a field of its own. Covered by TestEditKeepsTypePattern.

The rule builder's preview is what the rule is judged against, which is not always "what is untagged". For a new rule those are the same thing. For an edit they are not: the rule's own transactions are tagged, so an untagged-only preview would be empty for a rule that works. previewGroups adds them back, and rulePreview keeps the ones the new glob stops catching on screen marked − instead of dropping them silently, because giving one up is the decision being made. Covered by TestEditPreviewShowsWhatIsLetGo.

The web app is a frontend, not a second implementation. Everything that decides something stays in Go: amounts are formatted server-side (money is text plus a sign, never a number the browser could add up), the rule builder's preview is matched by glob.Match on the server rather than re-implemented in JS, and the transfer preview runs transfers.Analyze. Keep it that way — a JS glob or a JS sum is a second copy of a rule that would drift from the first. static/app.js only lays out what the API returns.

The server outlives edits to rules.toml in a way a one-shot command does not, so /api/retag and /api/import re-read it first (import also re-reads the account folders), and /api/overview reports stale when the file on disk no longer equals what the engines hold. Every edit and delete sends the rule or transfer it showed alongside the position, and is refused with 409 unless rules.toml still holds exactly that there — a position from a stale page can otherwise name a different rule. Covered by TestStalePositionIsRefused.

There is no auth by design (--addr defaults to loopback). Non-GET requests must be application/json, which is what keeps a cross-site form from posting to it; do not relax that without putting something else in its place. That is why /api/upload takes files base64 inside JSON rather than as multipart: multipart is precisely what a cross-site form can send.

An upload only puts a file where money import looks. It writes into an existing account folder and stops there: importing stays the user's call, made with the Import button, so the statements stay the source of truth and nothing reaches the index any other way. It never overwrites a statement (identical contents are a no-op, different ones a 409), refuses any name import would not read back — not a plain file name, a dotfile, account.toml, outside the account's include — and checks a whole batch before writing any of it. Covered by TestUploadSavesWithoutImporting, TestUploadNeverReplacesAStatement and TestUploadRefusesBadNames.

The statements list is read from the folders, not the index. /api/files walks importer.StatementFiles — the same list import reads — and only then looks each file up in store.SourceFiles, comparing importer.Checksum, so a file is listed exactly when import would read it and a file the index remembers but the disk lost shows as missing instead of vanishing. /api/files/{account}/{name} serves a file only by finding it in that list, never by joining the name onto a path. A statement is served from the app's origin, where a script could drive the API, so nothing is ever rendered as a page: text is text/plain under CSP: sandbox, anything not text or PDF is a sandboxed download, and PDFs — whose viewers refuse a sandbox — open in the browser's own isolated viewer. Covered by TestServeFileServesOnlyStatements.

Deleting a statement marks its neighbours to be re-read. /api/files/delete finds the file the same way serveFile does and deletes it — permanently, by the user's choice; the page says so before it asks. store.ForgetSourceFile then drops its transactions and their transfer rows, and Link re-pairs, which reads no statement. It does not import, any more than an upload does. A transaction two overlapping statements share is stored once, under the file imported first, so it goes too; ForgetSourceFile therefore blanks the checksums of the account's other statements, which show as changed until the next Import re-reads them and restores what they hold. Without that the checksum skip would leave those rows gone for good. Covered by TestRemoveFileKeepsWhatAnotherStatementHolds.

Adding a bank parser

Implement parser.Parser and call parser.Register from an init. Nothing else changes; the name becomes usable in an account.toml. Use parser.ParseAmount rather than hand-rolling decimal handling — it copes with 1.234,56, trailing minus, parenthesised negatives, currency codes and both the ASCII hyphen and U+2212.

Keep PDF text extraction separate from parsing: the bank parsers expose a pure parseXText(text string, digits int) so they can be tested against captured pdftotext output without a PDF fixture. pdftotext -layout (poppler-utils) is a runtime dependency of nlb and traderepublic; no Go library reconstructs column layout as well.

It can be carried inside the binary instead: scripts/build-bundled.sh builds a static musl pdftotext in a container and embeds it under -tags bundled, and pdftotextCommand writes it to the user cache dir (named by content hash, verified before reuse) and runs that. It is still the real executable run as a subprocess, deliberately — not poppler linked in through cgo, which would cost the pure-Go build, and not a different extraction API whose spacing might not match what the parsers were tuned on. The tag is opt-in so go build/go test never need the 5 MB file, which is gitignored rather than committed. Bumping POPPLER_VERSION means updating its pinned SHA-256 too, and checking real statements still parse — that output is what the parsers depend on.

Do not hardcode absolute column positions from a sample PDF. pdftotext compresses runs of spaces, so columns shift with font and page size — derive positions from the header line (traderepublic) or from the line being parsed (nlb).

The other side's account number belongs in the description, not in a field of its own. There used to be a Counterparty field; only nlb could fill it honestly, revolut and traderepublic invented it with an IBAN-shaped regex over the description, and the two disagreed on spacing, so one rule pattern could not serve both. Now nlb appends its IBAN column to the end of the description — at the end, and not in the position it held on the page, so an IBAN wrapped across continuation lines stays contiguous for a glob to match. A new parser must do the same rather than reintroduce a structured field.

A parser may name its statements, and only uploads use it. parser.Namer is optional, like parser.Warner: nlb names an izpisek izpisek_YYYY-MM-DD after its Datum izpiska, ported from the rename_izpiski.py it replaced, and traderepublic names a statement traderepublic_YYYY_MM_DD_YYYY_MM_DD after the period in its header. A name is derived from what the statement says, never from the upload's own name, and comes back without an extension; the upload's is kept, lowercased. /api/upload stages each file as a dotfile — invisible to import — so the parser can read it, then numbers a chosen name that is taken by different contents (_2, as the script did) where a name the user gave is refused with 409. Nothing renames a file already in a folder: source_files records statements by path, so a rename would orphan its rows as missing until the index is rebuilt. Covered by TestUploadNamesStatementsTheParserCanName.

Verifying

go build ./... && go test ./... && go vet ./... && gofmt -l .

For end-to-end checks, build a throwaway data root under the scratchpad rather than touching real data:

money --root /tmp/.../demo import
money --root /tmp/.../demo import          # must report 0 new
money --root /tmp/.../demo ls --wide

If you changed the schema, delete that root's index.db first. Nothing migrates it, so a demo root left over from an earlier build either carries dead columns or fails with no such column, depending on which way the schema moved.

Testing the web app

Drive the API through Server.Handler() with httptest — newTestServer in internal/web/server_test.go builds a data root and an index tagged and paired as an import would leave them, with "today" fixed so the report's default window is predictable. That is how every web test works, and since the page only lays out what the API returns, it is where behaviour gets tested.

For a look at the real page, run money --root <demo> serve --addr 127.0.0.1:<port> against a throwaway root. Delete the index first if you want to watch an import do work, or it will be skipped by checksum and finish instantly.

Conventions

Comments explain why, not what. Errors name the file and the offending row or key so a bad statement is actionable. Per-file import failures are reported and the run continues; only a broken data root aborts.