counterparty was a structured field only nlb could fill honestly. revolut and traderepublic invented one by running an IBAN-shaped regex over the description they had just built, and the two spellings disagreed -- SI56 1234 5678 9012 345 against SI56123456789012345 -- so a literal rule pattern that worked on one account silently matched nothing on another. It is gone from the model, the index, the rule keys, ls --wide and the rules screen. nlb now appends its IBAN column to the end of the description, where the other two already keep theirs, so match = "*SI56*" works everywhere. That changes those descriptions and with them their fingerprints, so a statement overlapping an already-imported period will re-add rather than dedupe those rows until the index is rebuilt. An index built by an older binary drops the column when it is opened. The index itself moves from .money/index.db up to index.db beside rules.toml. Nothing looks in the old location, so an existing one has to be moved by hand -- otherwise the tool quietly starts a fresh index and the manual tags in the old file, the only thing statements cannot reproduce, stay behind in it. The csv and cmd parsers are gone along with the [csv] and [cmd] config they carried. cmd shelled out to the Python extractors, which were ported to Go and deleted, so it bridged to nothing; csv was a generic column-mapped fallback that no account used, and between them they were the largest configuration surface in the tool. A bank is now described in Go, where it can be tested. The importer tests register their own three-column parser rather than borrow a bank's, so they stay about the directory walk, dedupe and per-file error reporting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
7.1 KiB
money
Statement-driven personal finance tracker. A data directory holds one folder per account; statements dropped into those folders are parsed into a rebuildable SQLite index, categorised by ordered glob rules, and browsed in a Bubble Tea TUI.
Read README.md first — it is the user-facing reference for every config key.
This file covers what the code assumes and why.
Layout
cmd/money/main.go subcommands; the TUI is the default
internal/config rules.toml, account.toml, XDG config, data-root resolution
internal/model Account, Transaction, amount formatting, description normalisation
internal/glob the `*` / `?` matcher used by rules (linear time, no backtracking)
internal/parser Parser interface + registry; nlb, revolut, traderepublic
internal/store SQLite index (modernc.org/sqlite, no cgo)
internal/importer directory walk, dedupe, balance checks
internal/rules applies ordered rules to rule_* columns only
internal/report per-tag aggregation, transfers excluded
internal/tui Bubble Tea models
Invariants
Rule verdicts and manual edits never share a column. rule_tag /
rule_transfer are rewritten wholesale on every retag; manual_tag /
manual_transfer are only ever written by the user. Effective values are
COALESCE(manual_*, rule_*). This is what makes money retag safe to run at
any time, and it is the first thing to preserve when touching the schema or the
rules engine. Covered by TestManualTagSurvivesRetag.
Statements are the source of truth; the index is disposable. Deleting
index.db at the root of the data directory and re-importing must reproduce
everything except manual tags and manual transfer marks.
Dedupe is by fingerprint: sha256(date | amount | normalised description | ordinal), where the ordinal distinguishes identical lines within one
statement. Two identical purchases on one day both survive; the same line in an
overlapping statement does not duplicate. Unchanged files are skipped by
checksum before parsing at all — so an import with nothing new finishes in
milliseconds. That is the checksum skip working, not a failure.
Money is int64 minor units, never a float. Per-account currency, no
conversion, and totals are never summed across currencies.
First matching rule wins, so transfer rules belong above general tag rules.
A rule setting both match and type requires both of them.
config.AppendRule therefore appends — never prepends — so saving from the
rule builder cannot shadow a rule the user wrote by hand.
The rule builder is a form, so the global keymap must not apply there.
Update routes to updateRules before updateNormal whenever the view is
viewRules, or typing q would quit and i would start an import. Any new
full-screen input needs the same treatment.
In the rule builder, tab completes first and moves focus second. The
account and tag fields use textinput.ShowSuggestions, whose own AcceptSuggestion
key is tab and whose NextSuggestion/PrevSuggestion are up/down — all
three already meant something here. So updateRules intercepts tab and calls
acceptCompletion before falling back to setRuleFocus, keeps up/down on
field navigation, and lets ctrl+n / ctrl+p through to the input for cycling.
SetValue does not re-match the suggestion list, so acceptCompletion re-sets
it afterwards or ctrl+n would offer candidates that no longer fit the value.
A rule's usage count is how many transactions it wins, not how many its glob
could match — Engine.Usage counts by MatchIndex, so a rule shadowed by an
earlier one correctly reports zero. That is what makes the rules screen able to
find dead rules at all.
A rule's note is documentation that round-trips. It is a TOML key rather
than a # comment so LoadRules can return it, the builder can write it and
the rules screen can show it. It never takes part in matching — rules.Engine
does not look at it — and it must stay that way.
config.DeleteRules edits rules.toml textually, never by re-serialising the
parsed rules, because comments and formatting are not recoverable from
[]Rule. A rule owns the comment lines directly above it; a comment separated
by a blank line is a heading for what follows and stays. Both writers go
through writeFileAtomic, and deletion re-parses the result before replacing
the file.
Adding a bank parser
Implement parser.Parser and call parser.Register from an init. Nothing
else changes; the name becomes usable in an account.toml. Use
parser.ParseAmount rather than hand-rolling decimal handling — it copes with
1.234,56, trailing minus, parenthesised negatives, currency codes and both
the ASCII hyphen and U+2212.
Keep PDF text extraction separate from parsing: the bank parsers expose a pure
parseXText(text string, digits int) so they can be tested against captured
pdftotext output without a PDF fixture. pdftotext -layout (poppler-utils) is
a runtime dependency of nlb and traderepublic; no Go library reconstructs
column layout as well.
Do not hardcode absolute column positions from a sample PDF. pdftotext
compresses runs of spaces, so columns shift with font and page size — derive
positions from the header line (traderepublic) or from the line being parsed
(nlb).
The other side's account number belongs in the description, not in a field
of its own. There used to be a Counterparty field; only nlb could fill it
honestly, revolut and traderepublic invented it with an IBAN-shaped regex
over the description, and the two disagreed on spacing, so one rule pattern
could not serve both. Now nlb appends its IBAN column to the end of the
description — at the end, and not in the position it held on the page, so an
IBAN wrapped across continuation lines stays contiguous for a glob to match.
A new parser must do the same rather than reintroduce a structured field.
Verifying
go build ./... && go test ./... && go vet ./... && gofmt -l .
For end-to-end checks, build a throwaway data root under the scratchpad rather than touching real data:
money --root /tmp/.../demo import
money --root /tmp/.../demo import # must report 0 new
money --root /tmp/.../demo ls --wide
Driving the TUI in tests
Prefer feeding tea.Msg values to Model.Update directly — that is how every
existing TUI test works, and it covers the keymap without a terminal.
If a real terminal is genuinely needed, note that piping into script does not
deliver keystrokes. Use a pty and answer the two capability queries Bubble Tea
sends on startup, or the program blocks before its first render:
ESC]11;?(background colour) → replyESC]11;rgb:1e1e/1e1e/1e1e ESC\ESC[6n(cursor position) → replyESC[1;1R
Also delete the index first, or the import you are trying to observe will be skipped by checksum and finish instantly.
Conventions
Comments explain why, not what. Errors name the file and the offending row or key so a bad statement is actionable. Per-file import failures are reported and the run continues; only a broken data root aborts.