Drop counterparty, the generic parsers and .money/

counterparty was a structured field only nlb could fill honestly. revolut
and traderepublic invented one by running an IBAN-shaped regex over the
description they had just built, and the two spellings disagreed --
SI56 1234 5678 9012 345 against SI56123456789012345 -- so a literal rule
pattern that worked on one account silently matched nothing on another. It
is gone from the model, the index, the rule keys, ls --wide and the rules
screen. nlb now appends its IBAN column to the end of the description,
where the other two already keep theirs, so match = "*SI56*" works
everywhere. That changes those descriptions and with them their
fingerprints, so a statement overlapping an already-imported period will
re-add rather than dedupe those rows until the index is rebuilt. An index
built by an older binary drops the column when it is opened.

The index itself moves from .money/index.db up to index.db beside
rules.toml. Nothing looks in the old location, so an existing one has to be
moved by hand -- otherwise the tool quietly starts a fresh index and the
manual tags in the old file, the only thing statements cannot reproduce,
stay behind in it.

The csv and cmd parsers are gone along with the [csv] and [cmd] config they
carried. cmd shelled out to the Python extractors, which were ported to Go
and deleted, so it bridged to nothing; csv was a generic column-mapped
fallback that no account used, and between them they were the largest
configuration surface in the tool. A bank is now described in Go, where it
can be tested. The importer tests register their own three-column parser
rather than borrow a bank's, so they stay about the directory walk, dedupe
and per-file error reporting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-11 22:51:20 +02:00
co-authored by Claude Opus 5
parent 2a2c0c1db9
commit 16c2585637
23 changed files with 182 additions and 627 deletions
+22 -59
View File
@@ -12,7 +12,7 @@ transfers so they never count as spending.
```
~/money/ # the data root (see "Where the data root lives")
rules.toml # tag + transfer rules, in order
.money/index.db # SQLite index (rebuildable; safe to delete*)
index.db # SQLite index (rebuildable; safe to delete*)
checking/
account.toml # currency + how to parse this bank's exports
2026-01.csv
@@ -63,7 +63,7 @@ money import # extract new transactions from every statement
money import --force # re-parse statements even if unchanged
money retag # re-apply rules.toml to everything already imported
money ls --untagged # what still needs a tag
money ls --wide # also show type, counterparty and reported balance
money ls --wide # also show type and reported balance
money ls --account checking --month 2026-01
money report --month 2026-01 # spending by tag
money accounts # balances
@@ -175,9 +175,8 @@ Rules are evaluated in file order and the **first match wins**, so put transfer
rules above general tag rules. Patterns are globs (`*` and `?`) matched
case-insensitively, with whitespace collapsed.
A rule matches on `match` (the description), `counterparty` (the other side's
account number) and `type` (the bank's own classification). Setting several is
an "and": all must match.
A rule matches on `match` (the description) and `type` (the bank's own
classification). Setting both is an "and": both must match.
`note` is free text for you, never for the matcher: why the rule is there, or
what the unrecognisable payee behind the glob actually is. It shows in the last
@@ -206,12 +205,13 @@ match = "*FROM CHECKING*"
transfer = true
tag = "transfer"
# Transfers are often only identifiable by the counterparty IBAN, whatever
# the description happens to say.
# Transfers are often only identifiable by the other side's account number,
# whatever the rest of the description happens to say. Every bank parser keeps
# that number in the description, so an ordinary glob finds it.
[[rule]]
counterparty = "SI56*"
transfer = true
tag = "transfer"
match = "*SI56*"
transfer = true
tag = "transfer"
# A rule can be limited to one account, and can require several patterns.
[[rule]]
@@ -236,43 +236,20 @@ Every account folder needs one. `currency` and `parser` are required.
```toml
name = "Main Checking"
currency = "EUR"
parser = "csv"
parser = "nlb"
# minor_digits = 2 # decimal places for this currency
# include = ["*.csv"] # only treat matching files as statements
```
### `parser = "csv"`
### Parsers
Column positions are 0-based.
```toml
[csv]
delimiter = "," # default ","
skip_rows = 1 # header rows to drop
date = { col = 0, layout = "02.01.2006" } # Go reference layout
description = { col = 3 }
amount = { col = 4, decimal = ",", thousands = "." }
# invert = true # if outflows are written as positive
```
For statements with separate debit and credit columns, replace `amount`:
```toml
debit = { col = 4 } # both written as positive numbers
credit = { col = 5 }
```
Amount parsing is deliberately tolerant: `1.234,56`, `-45.20`, `45,20-`,
`(45.20)` and `45.20 EUR` all work.
### Bank-specific parsers
Three are built in, ported from the original Python extractors. They need no
`[csv]` block — the layout is baked in.
Three are built in, ported from the original Python extractors. A statement
layout is described in Go rather than in a table of column indexes, so there is
nothing else to configure — `parser` names one of these and that is all.
| `parser` | Statement | Notes |
| --- | --- | --- |
| `nlb` | NLB izpisek PDF | Wrapped descriptions and the counterparty IBAN are folded in from continuation lines. |
| `nlb` | NLB izpisek PDF | Wrapped descriptions are folded in from continuation lines, and the IBAN column is appended to the description. |
| `traderepublic` | Trade Republic PDF | Handles both the single-line and the stacked layout by measuring column positions. |
| `revolut` | `account-statement*.csv` | Skips non-COMPLETED rows, folds the fee into the amount. |
@@ -286,33 +263,19 @@ The two PDF parsers shell out to `pdftotext -layout` (poppler-utils), exactly
as the Python versions did; its layout reconstruction is what makes the
column-based parsing work.
Amount parsing is shared and deliberately tolerant: `1.234,56`, `-45.20`,
`45,20-`, `(45.20)` and `45.20 EUR` all work, whichever parser reads them.
**Revolut and multiple currencies.** One export can hold several currencies,
but an account here has exactly one. Rows in other currencies are skipped and
reported at import. To keep them, give that currency its own account folder
with its own `currency` and `minor_digits` (JPY wants `minor_digits = 0`) and
put a copy or symlink of the export in it.
### `parser = "cmd"`
## Adding a parser
Runs an external extractor and reads normalised CSV (`date,description,amount`)
from its stdout. This is how PDF statements and any existing Python extractor
are handled without porting them first.
```toml
parser = "cmd"
[cmd]
argv = ["python3", "../extract_bankx.py", "{{file}}"]
skip_rows = 1 # if the script prints a header
# layout = "2006-01-02" # date format the script emits (default)
```
`{{file}}` is replaced with the statement's path, and the command runs with the
account folder as its working directory, so relative script paths work.
## Adding a native parser
When a Python extractor is ported to Go, drop it in `internal/parser` and
register it — no other package changes:
A new bank means a new parser. Drop it in `internal/parser` and register it —
no other package changes:
```go
func init() {