Four changes to statement handling in the web app, made together and
touching the same upload and statements-list code.
Name NLB uploads by statement date. parser.Namer is an optional
interface, like Warner, through which a parser names its statements;
nlb reads the "Datum izpiska" from the izpisek header and names it
izpisek_YYYY_MM_DD, lowercase, extension included -- ported from the
rename_izpiski.py it replaces. Uploads are staged as dotfiles, invisible
to import, so the parser can read them; two downloads of one statement
then meet under one name and the second is recognised as already there,
while a different statement of the same date is numbered _2 as the
script did. Only uploads are named: source_files records statements by
path, so renaming a file already in a folder would orphan its rows.
Delete a statement from the statements list. The file is removed from
disk for good -- the page says so before it asks -- and
store.ForgetSourceFile drops its transactions and their transfer rows.
A row two overlapping statements share is stored once, under the file
imported first, so it goes too; the account's other statements forget
their checksums and show as changed until the next Import re-reads them
and restores it. A file already gone from disk can be forgotten.
Upload and delete no longer import. Importing stays the user's call,
made with the Import button, so a batch can be put together and looked
over first. Delete still re-pairs transfers, which reads no statement.
Show rows and new rows per statement. The list read "0" for a file
whose rows an earlier, overlapping statement already held, which looked
like a file that failed to parse. source_files now records how many
transactions each statement holds, and the list reads "3 rows · 0 new".
This adds a column the code reads, so an index built by an earlier
version fails with "no such column: s.rows": delete index.db and import
again.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Accounts screen gets a Statements list: every file in every account
folder beside what the index made of it -- imported, changed since,
not imported yet (or failed), or gone from disk while its rows remain --
with its size, modified and import times, and how many transactions it
brought in. Clicking a name opens the file.
The list is read from the folders, not the index, through
importer.StatementFiles, so a file shows exactly when import would read
it; store.SourceFiles and importer.Checksum then say how far each has
got. A file the index remembers but the disk lost is listed as missing
rather than vanishing, since its rows would not survive a rebuild.
A file is served only by finding it in that list, never by joining the
requested name onto a path. Statements come from outside and are served
from the app's origin, so none is rendered as a page: text is text/plain
under CSP sandbox, anything not text or PDF is a sandboxed download, and
PDFs -- whose viewers refuse a sandbox -- open in the browser's own
isolated viewer.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The boolean transfer flag went two commits ago because a one-sided verdict let
half a movement vanish and left the report unbalanced. This is what replaces
it: a [[transfer]] block names both legs, and only a matched pair is dropped
from the report -- both legs together, never one.
Legs pair within five days, nearest date first, and a transaction belongs to at
most one transfer, so the first definition to claim a leg keeps it, exactly as
the first matching rule keeps a tag. The pairing is derived state like the tags:
Engine.Link rewrites the whole transfers table from rules.toml, which is why
retag re-derives both halves of what that file decides, and why it runs over
the whole index rather than a filtered view -- pairing inside one would let a
movement count as a transfer in one report and not in another. An unmatched leg
is not a transfer and keeps counting, surfaced as a warning instead.
Within one currency the amount is the evidence and must be the exact opposite.
Across currencies it is not checked at all: there are no rates here, so the two
numbers are unrelated and the dates carry the pairing alone.
tolerance_pct is the one exception, per definition, for a route where the bank
takes a fee and the two statements genuinely disagree. It defaults to zero and
belongs on the one definition that charges; a global or default tolerance would
loosen every route that does not. The difference it admits is not forgiven --
the pair leaves the report entirely, so a fee hidden inside one would be
spending that appears nowhere. Pair.Fee is what left less what arrived, and
report.Excluded carries it out per currency alongside the legs. It counts only
pairs whose legs are both in view, for the same reason it counts legs and not
transfers: half a pair cannot say what the other half received.
The screens:
- 6 builds a definition against the index as you type, showing the pairs it
would form and the legs it would catch but leave unpaired. Six fields need
more room than the rule builder's four, so the form sheds its spacing, then
its hints, then the borders on unfocused fields.
- 7 lists every definition with what it pairs. Two counts, because they mean
different things: an unpaired leg is a definition doing something and not
finishing it, no pairs at all is dead weight. Tol names the tolerance, blank
where amounts must agree.
- 3 grows a (transfers) row under TOTAL, and a fees row beneath it, or the
report silently disagrees with the account balances.
Two things that are not part of transfers but are the same day's work:
- ls --uniq lists each account and description once, normalised the way a glob
sees them, which is the shape of "what still needs a rule?" -- fifty visits
to one shop are one pattern to write, not fifty rows to read.
- The rule builder's preview now filters to what the glob matches instead of
marking matches in a full list. The count carries the context the rows no
longer can: 2 of 7, measured against everything still in view.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A transfer was a second verdict carried alongside the tag: a boolean set by
transfer = true in a rule or by x in the TUI, kept in its own pair of rule_
and manual_ columns, whose one real effect was to hold the row out of the
report. The rest of it was display -- a T column in the transaction list, in
the rules screen and in money ls.
Money moved between your own accounts is now tagged like anything else and
counts like anything else. The leg leaving checking is an outflow and the leg
arriving in savings is an inflow, so a report over the whole data root roughly
nets out while one scoped to a single account or month does not. That is the
price of one verdict per transaction instead of two.
A rule now needs a tag, and one that set only transfer = true is refused by
number on load. A leftover transfer key beside a tag is ignored, as unknown
TOML keys always were, and an index built by an older binary drops both
columns when it is opened.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
counterparty was a structured field only nlb could fill honestly. revolut
and traderepublic invented one by running an IBAN-shaped regex over the
description they had just built, and the two spellings disagreed --
SI56 1234 5678 9012 345 against SI56123456789012345 -- so a literal rule
pattern that worked on one account silently matched nothing on another. It
is gone from the model, the index, the rule keys, ls --wide and the rules
screen. nlb now appends its IBAN column to the end of the description,
where the other two already keep theirs, so match = "*SI56*" works
everywhere. That changes those descriptions and with them their
fingerprints, so a statement overlapping an already-imported period will
re-add rather than dedupe those rows until the index is rebuilt. An index
built by an older binary drops the column when it is opened.
The index itself moves from .money/index.db up to index.db beside
rules.toml. Nothing looks in the old location, so an existing one has to be
moved by hand -- otherwise the tool quietly starts a fresh index and the
manual tags in the old file, the only thing statements cannot reproduce,
stay behind in it.
The csv and cmd parsers are gone along with the [csv] and [cmd] config they
carried. cmd shelled out to the Python extractors, which were ported to Go
and deleted, so it bridged to nothing; csv was a generic column-mapped
fallback that no account used, and between them they were the largest
configuration surface in the tool. A bank is now described in Go, where it
can be tested. The importer tests register their own three-column parser
rather than borrow a bank's, so they stay about the directory walk, dedupe
and per-file error reporting.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A data directory holds one folder per account. Statements dropped into
those folders are parsed into a rebuildable SQLite index, categorised by
ordered glob rules in rules.toml, and browsed or hand-tagged in a Bubble
Tea TUI. Movements between the user's own accounts are marked as
transfers by the same rules and excluded from spending totals.
Manual tags and transfer marks are stored separately from the rule-derived
ones and always win, so editing rules.toml and re-running retag never
destroys hand edits.
Parsers are pluggable. Three are ported from the Python extractors they
replace -- nlb and traderepublic read PDFs via pdftotext -layout, revolut
reads the CSV export -- alongside a configurable-column CSV parser and a
cmd parser that shells out to an external script.
Both ports fix two latent bugs in the originals: the sign character class
rejected the typographic minus U+2212 that some PDF fonts emit, and NLB's
hardcoded continuation indent broke when pdftotext compressed runs of
spaces, so the threshold is now measured from the description column.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>