# money Statement-driven personal finance tracker. A data directory holds one folder per account; statements dropped into those folders are parsed into a rebuildable SQLite index, categorised by ordered glob rules, and browsed in a Bubble Tea TUI. Read `README.md` first — it is the user-facing reference for every config key. This file covers what the code assumes and why. ## Layout ``` cmd/money/main.go subcommands; the TUI is the default internal/config rules.toml (rules + transfers), account.toml, XDG config, data-root resolution internal/model Account, Transaction, amount formatting, description normalisation internal/glob the `*` / `?` matcher used by rules (two-pointer, no exponential blowup) internal/parser Parser interface + registry; nlb, revolut, traderepublic internal/store SQLite index (modernc.org/sqlite, no cgo) internal/importer directory walk, dedupe, balance checks internal/rules applies ordered rules, writing rule_tag internal/transfers pairs the two legs of a movement, writing the transfers table internal/report per-tag aggregation internal/tui Bubble Tea models ``` ## Invariants **Every tag comes from rules.toml.** There is no way to tag a transaction by hand; `rule_tag` is derived state that `Engine.Retag` rewrites wholesale, which is what makes `money retag` safe to run at any time. Nothing may write a tag from anywhere else — the moment something does, the index stops being reproducible and retag stops being safe. Covered by `TestRetagRewritesEveryTag`. **A transfer is a pair or it is nothing.** A `[[transfer]]` block names both legs (`from_account` + `from_desc`, `to_account` + `to_desc`), and only a matched pair is dropped from the report — both legs together, never one. This is the whole reason the earlier boolean `transfer` flag was removed and this replaced it: a one-sided verdict let half a movement vanish and left the report unbalanced. An unmatched leg therefore keeps counting, and is surfaced as a warning instead: `⚠` on the transfers screen, on stderr from `import`/`retag`. Nothing may start excluding a single leg on the strength of one side matching. Legs pair within `transfers.WindowDays`, nearest date first. **Within one currency the amount is the evidence and must be the exact opposite**, unless the definition sets `tolerance_pct` — a per-definition opt-in for a route where the bank takes a fee on the way, so the two statements genuinely disagree. It defaults to zero and belongs on the one definition that charges: a global or default tolerance would loosen every route that does not. **Across currencies the amount is not checked at all**: there are no exchange rates here, so the two numbers are unrelated and the dates carry the pairing alone. A tolerance means nothing there and is ignored. That asymmetry is the design, not an oversight — do not "fix" the cross-currency case by inventing a rate, and do not turn the tolerance into the default. Like rules, the first definition claims a leg and a transaction belongs to at most one transfer. **A tolerated mismatch is a fee, and a fee is money, so it is reported rather than forgiven.** That is the condition on the tolerance existing at all: the pair leaves the report entirely, so a difference swallowed inside one would be spending that never appears anywhere. `Pair.Fee` is what left less what arrived, and `report.Excluded` carries it out per currency alongside the legs — named on the `money report` line and on its own row under the TUI's report. Nothing may pair on a mismatch without that difference reaching `Excluded`. It is a fee and not a tag: no rule produces it, `report.ByTag` never sees it, and it must not be turned into a synthetic transaction to make the total reconcile. `Excluded` counts it only for pairs whose legs are *both* in view, for the same reason it counts legs and not pairs — half a pair cannot say what the other half received. **`(transfer)` in a tag column is display only.** `Transaction.DisplayTag` falls back to `model.TransferTag` for a paired leg so the column does not read as a blank waiting for a rule, but nothing writes it: `rule_tag` stays what rules.toml made it, `store.Tags` never returns it, and no rule can be built from it. It is the same kind of label as `report.Untagged`. Do not "persist" it — that is precisely the second verdict this design exists to avoid. **A paired leg is not untagged.** `store.Filter{Untagged: true}` means "no tag *and* no transfer", so a leg never turns up in `money ls --untagged`, the `u` view or the rule builder's preview asking to be tagged — it is spoken for, and the report drops it regardless. Covered by `TestPairedLegsAreNotUntagged`. An *unpaired* leg is untagged like anything else, which is how it gets noticed. **The pairing is derived state, exactly like the tags.** `Engine.Link` rewrites the whole `transfers` table from rules.toml, so `money retag` re-derives both halves of what that file decides. It runs over the whole index deliberately — pairing within a filtered view would let a movement count as a transfer in one report and not in another. Note the asymmetry with a schema *column*: a new table is created by `CREATE TABLE IF NOT EXISTS` in `store.Open`, so adding one does not force an index rebuild. **Rules and transfers share rules.toml, so textual deletion is block-aware.** `config.deleteBlocks` finds every `[[...]]` header, not only the kind being deleted, because a block ends where the *next* block of any kind begins — otherwise deleting a rule would swallow a transfer that follows it. Covered by `TestDeleteLeavesTheOtherKindAlone`. **Statements are the source of truth; the index is disposable.** Deleting `index.db` at the root of the data directory and re-importing must reproduce everything, with nothing lost. **The index is a cache, so there are no migrations.** `store.Open` applies `schema` and nothing else — no `ALTER TABLE`, no version column, no repair of an index an older build wrote. Changing the schema costs one line in the release note: delete `index.db` and import again. The two directions are not symmetric, which is worth knowing before you assume something is broken: - *Removing* a column leaves an older index still working, carrying the dead column and its data unread. - *Adding* one the code reads breaks every command against an older index with `query transactions: SQL logic error: no such column: …` until it is deleted. That asymmetry holds only because every statement names its columns — **no `SELECT *`, and nothing may depend on column order.** Keep it that way. Do not reintroduce migration machinery either; if re-parsing ever becomes too expensive to ask for, that is a decision to revisit deliberately rather than a helper to slip back in. **Dedupe is by fingerprint**: `sha256(date | amount | normalised description | ordinal)`, where the ordinal distinguishes identical lines *within one statement*. Two identical purchases on one day both survive; the same line in an overlapping statement does not duplicate. Unchanged files are skipped by checksum before parsing at all — so an import with nothing new finishes in milliseconds. That is the checksum skip working, not a failure. **Money is `int64` minor units**, never a float. Per-account currency, no conversion, and totals are never summed across currencies. **The most specific rule wins, not the topmost.** `rules.Engine` sorts the rules once in `New` and matches in that order: most literal characters first, then fewest `*`, then account-scoped over unscoped, with a *stable* sort so equally specific rules keep file order and the earlier one still wins. That is what lets `*NIKOLA*` carve an exception out of `*NIK*` from anywhere in the file, and a catch-all `*` sit wherever it reads best. A rule setting both `match` and `type` requires both of them, and both count towards its literals. The two orders must not be confused. `Engine.rules` stays in file order and `Rules()`, `Usage` and `MatchIndex` all speak in file positions, because that is what the rules screen numbers, what `config.DeleteRules` deletes by, and what the user can point at in rules.toml; only `Engine.order` is sorted. Anything new that reports a rule must report its file position too. `config.AppendRule` still appends rather than prepends, but that now only settles ties: a saved rule cannot displace an equally specific one written by hand, while a narrower one is meant to take precedence and does. **The builders are forms, so the global keymap must not apply there.** `Update` routes to `updateRules` / `updateTransfers` before `updateNormal` whenever the view is `viewRules` or `viewTransfers`, or typing `q` would quit and `i` would start an import. Any new full-screen input needs the same treatment. **In the rule builder, `tab` completes first and moves focus second.** The account and tag fields use `textinput.ShowSuggestions`, whose own `AcceptSuggestion` key is `tab` and whose `NextSuggestion`/`PrevSuggestion` are `up`/`down` — all three already meant something here. So `updateRules` intercepts `tab` and calls `acceptCompletion` before falling back to `setRuleFocus`, keeps `up`/`down` on field navigation, and lets `ctrl+n` / `ctrl+p` through to the input for cycling. `SetValue` does not re-match the suggestion list, so `acceptCompletion` re-sets it afterwards or `ctrl+n` would offer candidates that no longer fit the value. **A rule's usage count is how many transactions it wins, not how many its glob could match** — `Engine.Usage` counts by `MatchIndex`, so a rule shadowed by a more specific one correctly reports zero. That is what makes the rules screen able to find dead rules at all. **A rule's `note` is documentation that round-trips.** It is a TOML key rather than a `#` comment so `LoadRules` can return it, the builder can write it and the rules screen can show it. It never takes part in matching — `rules.Engine` does not look at it — and it must stay that way. **`config.DeleteRules` and `config.ReplaceRule` edit rules.toml textually, never by re-serialising the parsed rules**, because comments and formatting are not recoverable from `[]Rule`. A rule owns the comment lines directly above it, so deleting takes them with it; a comment separated by a blank line is a heading for what follows and stays. Replacing keeps them — they say why the rule is there, which editing its glob rarely changes — and rewrites only the rule's own lines, leaving its *position* alone: position still breaks ties, so a rule that moved could start beating an equally specific one it never used to. Every writer goes through `writeFileAtomic`, and both editors re-parse the result before replacing the file. **An edit must not lose what the builder does not show.** The form has four fields and a `Rule` has five, so `saveRule` carries `Type` through from the rule being edited and the form says it is doing so. Nothing may round-trip a rule through those four fields alone — a pattern the user was never shown is not a pattern they chose to remove. A new field on `Rule` needs the same treatment or a field of its own. Covered by `TestRuleEditKeepsTypePattern`. **The rule builder's preview is what the rule is judged against, which is not always "what is untagged".** For a new rule those are the same thing. For an edit they are not: the rule's own transactions are tagged, so an untagged-only preview would be empty for a rule that works. `reloadPreviewGroups` adds them back, and `refreshRulePreview` keeps the ones the new glob stops catching on screen marked `−` instead of dropping them silently, because giving one up is the decision being made. ## Adding a bank parser Implement `parser.Parser` and call `parser.Register` from an `init`. Nothing else changes; the name becomes usable in an `account.toml`. Use `parser.ParseAmount` rather than hand-rolling decimal handling — it copes with `1.234,56`, trailing minus, parenthesised negatives, currency codes and both the ASCII hyphen and U+2212. Keep PDF text extraction separate from parsing: the bank parsers expose a pure `parseXText(text string, digits int)` so they can be tested against captured `pdftotext` output without a PDF fixture. `pdftotext -layout` (poppler-utils) is a runtime dependency of `nlb` and `traderepublic`; no Go library reconstructs column layout as well. Do not hardcode absolute column positions from a sample PDF. `pdftotext` compresses runs of spaces, so columns shift with font and page size — derive positions from the header line (`traderepublic`) or from the line being parsed (`nlb`). **The other side's account number belongs in the description, not in a field of its own.** There used to be a `Counterparty` field; only `nlb` could fill it honestly, `revolut` and `traderepublic` invented it with an IBAN-shaped regex over the description, and the two disagreed on spacing, so one rule pattern could not serve both. Now `nlb` appends its IBAN column to the end of the description — at the end, and not in the position it held on the page, so an IBAN wrapped across continuation lines stays contiguous for a glob to match. A new parser must do the same rather than reintroduce a structured field. ## Verifying ``` go build ./... && go test ./... && go vet ./... && gofmt -l . ``` For end-to-end checks, build a throwaway data root under the scratchpad rather than touching real data: ``` money --root /tmp/.../demo import money --root /tmp/.../demo import # must report 0 new money --root /tmp/.../demo ls --wide ``` If you changed the schema, delete that root's `index.db` first. Nothing migrates it, so a demo root left over from an earlier build either carries dead columns or fails with `no such column`, depending on which way the schema moved. ### Driving the TUI in tests Prefer feeding `tea.Msg` values to `Model.Update` directly — that is how every existing TUI test works, and it covers the keymap without a terminal. If a real terminal is genuinely needed, note that piping into `script` does not deliver keystrokes. Use a pty and answer the two capability queries Bubble Tea sends on startup, or the program blocks before its first render: - `ESC]11;?` (background colour) → reply `ESC]11;rgb:1e1e/1e1e/1e1e ESC\` - `ESC[6n` (cursor position) → reply `ESC[1;1R` Also delete the index first, or the import you are trying to observe will be skipped by checksum and finish instantly. ## Conventions Comments explain why, not what. Errors name the file and the offending row or key so a bad statement is actionable. Per-file import failures are reported and the run continues; only a broken data root aborts.