Initial import: 10 Radix skills across 3 categories
Kategoriserer danske regnskabs-/ERP-skills, ERP-data-afstemnings-
skill og to devops-workflows (gitea-issue-agent, kanban-workflows)
som Radix-medlemmers AI-agenter kan dele.
Struktur:
skills/<kategori>/<skill-navn>/SKILL.md
skills/<kategori>/<skill-navn>/{references,templates,scripts}/
Indekseret af index.json (genereret af scripts/rebuild_index.py).
Kør `python3 scripts/rebuild_index.py --check` i CI for at fange
synkroniseringsfejl.
This commit is contained in:
@@ -0,0 +1,314 @@
|
||||
---
|
||||
name: erp-data-reconciliation
|
||||
description: |
|
||||
Analyze and reconcile financial/accounting data between two ERP systems
|
||||
(e.g. SAP/Coop and Uniconta) using Excel exports. Covers file inspection,
|
||||
column mapping, Danish-format normalization, match-engine design, and
|
||||
iterative build workflow.
|
||||
trigger: |
|
||||
User asks to reconcile, match, compare, or afstemme data between two
|
||||
accounting/ERP systems, typically via Excel files. Also relevant when
|
||||
importing/exporting financial ledgers, kontoudtog, vendor/customer
|
||||
statements, or bonus calculations across systems.
|
||||
---
|
||||
|
||||
# ERP Data Reconciliation Skill
|
||||
|
||||
## Overview
|
||||
|
||||
Reconciling data between two ERP/accounting systems is a multi-phase task.
|
||||
The typical deliverable is a webapp or script that can:
|
||||
1. Import Excel files from both systems
|
||||
2. Normalize dates, amounts, and text
|
||||
3. Match postings between systems
|
||||
4. Show unmatched, partial matches, and differences
|
||||
5. Support manual approval/rejection
|
||||
6. Generate reports
|
||||
|
||||
## Phase 1: File Analysis (CRITICAL — do not skip)
|
||||
|
||||
Before writing any import code, run a structured analysis of the Excel files.
|
||||
|
||||
### Steps
|
||||
1. List all files in both source directories
|
||||
2. For each file, identify:
|
||||
- Sheet names
|
||||
- Actual header row location (may NOT be row 1!)
|
||||
- Column names and data types
|
||||
- Date columns and their formats
|
||||
- Amount columns and their formats
|
||||
- Text columns
|
||||
- Reference/invoice/document columns
|
||||
- Null percentages
|
||||
- Duplicate rows
|
||||
- Data range (min/max dates)
|
||||
|
||||
### Key Pitfalls
|
||||
- **Header row is NOT always row 1.** Uniconta exports often have account
|
||||
info in rows 1-2, blank row 3, headers in row 4.
|
||||
- **Amounts mixed formats.** Danish Excel files often have amounts as both
|
||||
Excel floats AND as Danish text strings like `-2.543.803,93`.
|
||||
Always normalize using a robust parser:
|
||||
```python
|
||||
def parse_danish_number(val):
|
||||
if pd.isna(val): return float('nan')
|
||||
if isinstance(val, (int, float)): return float(val)
|
||||
s = str(val).strip()
|
||||
if not s: return float('nan')
|
||||
if '.' in s and ',' in s:
|
||||
s = s.replace('.', '').replace(',', '.')
|
||||
elif ',' in s: s = s.replace(',', '.')
|
||||
return float(s)
|
||||
```
|
||||
- **Invoice/reference numbers may NOT match between systems.** This is the
|
||||
#1 cause of failed reconciliation. Always check overlap BEFORE assuming
|
||||
IDs can be used as join keys. Use amount overlap as a quick viability check.
|
||||
|
||||
### Analysis Script Template
|
||||
Use or adapt `scripts/analyze_excel.py` from the Coop-Uniconta project:
|
||||
- Reads all sheets in all files
|
||||
- Prints column stats, null counts, sample values
|
||||
- Detects Danish number formats
|
||||
- Identifies empty/duplicate rows
|
||||
|
||||
## Phase 2: Determine Match Strategy
|
||||
|
||||
After analysis, document the match strategy BEFORE coding the engine.
|
||||
|
||||
### Decision Tree
|
||||
1. **Do invoice/reference numbers overlap?**
|
||||
- If yes → use `Faktura + abs(Amount) + Date` as PRIMARY signal (two-phase).
|
||||
First phase: Reference==Faktura + abs(Amount) within tolerance (±5 DKK) + Date±21d.
|
||||
Second phase: abs(Amount) + Date±21d for remaining rows.
|
||||
- If no → use `Amount + Date + Text similarity`
|
||||
2. **Are dates exact or within tolerance?**
|
||||
- Same day → exact
|
||||
- ±1-3 days → high confidence partial
|
||||
- ±7 days → medium confidence
|
||||
- ±21 days → general match tolerance
|
||||
- >21 days → fuzzy/manual only
|
||||
3. **Is text available on both sides?**
|
||||
- Both have text → Levenshtein/rapidfuzz similarity score
|
||||
- Only one side has text → amount+date only
|
||||
4. **Do signs differ?**
|
||||
- Same sign → same direction transaction
|
||||
- Opposite sign → same transaction viewed from creditor vs debtor side.
|
||||
Match on `abs(amount)` but label "MODSAT" in output.
|
||||
*Example: Coop RE (payment) is negative; Uniconta Faktura is positive.*
|
||||
|
||||
### Tolerance Matching (V6 innovation)
|
||||
|
||||
Instead of exact amount match (±0.01), use a tolerance (±5 DKK):
|
||||
- Catches rounding differences between systems
|
||||
- Catches partial payments (one system may record 440.000, the other 430.732)
|
||||
- Uses bucket-based indexing for fast lookup:
|
||||
```python
|
||||
AMT_TOL = 5.0
|
||||
bucket = round(abs_amount / AMT_TOL) * AMT_TOL
|
||||
# Search adjacent buckets: [bucket-AMT_TOL, bucket, bucket+AMT_TOL]
|
||||
```
|
||||
|
||||
### Coop Fakturaliste + Uniconta Betalingsfordeling as the Simple Forward Workflow
|
||||
|
||||
When the user wants the process to be as simple as possible, prefer a forward-looking workflow based on:
|
||||
1. `Payment_advice`
|
||||
2. latest Uniconta Coop debtor/payment allocation export
|
||||
3. the complete Coop invoice list
|
||||
|
||||
Model the invoice list and payment allocation as separate import sources. Use invoice number as the primary join between Coop invoices and Uniconta allocation lines, and calculate paid/open/partial status from payment/settlement lines rather than only from original invoice postings. Treat the payment allocation as an explanation layer between a bank deposit and many invoice/credit lines.
|
||||
|
||||
See `references/coop-invoice-payment-allocation.md` for the concrete Django model/import/status pattern, UI shape, and tests.
|
||||
|
||||
### Coop Payment Advice / Settlement Files as a Fourth Source
|
||||
|
||||
When the user provides Coop betalingsadviseringer/opgørelser for individual bank
|
||||
payments, treat them as a separate reconciliation source, not just attachments.
|
||||
They are often the missing bridge:
|
||||
|
||||
`Bank deposit → Coop payment advice total → advice invoice/credit lines → Coop ledger / Uniconta invoices`
|
||||
|
||||
For Coop-Uniconta style systems, the import UI should expose the full operational set, not just "Coop" and "Uniconta":
|
||||
1. Uniconta Coop Debitor / payment allocation
|
||||
2. Uniconta/ERP invoice list: all Coop invoices/credit notes for the year
|
||||
3. Coop account statement / kontoudtog
|
||||
4. Bank statement containing all Coop postings
|
||||
5. Coop Payment_advice documents, uploadable continuously as they are collected
|
||||
|
||||
Recommended source/model names:
|
||||
- `coop_settlement` or `coop_payment_advice`
|
||||
- `PaymentAdvice` / `PaymentAdviceDocument` / `SettlementFile`: file name, unique id/bilag number, file type
|
||||
(PDF/XLSX), document date, payment date, total amount, period, parse status, matched bank transaction, raw
|
||||
text/raw JSON.
|
||||
- `PaymentAdviceLine` / `SettlementLine`: invoice number, invoice date, company
|
||||
code, amount, currency, line type (invoice/credit/bonus/modregning/unknown),
|
||||
links to bank/Coop/Uniconta transactions, match status, explanation.
|
||||
|
||||
If the user does not yet have all payment advice documents, still allow upload now. Store raw PDF text or Excel row JSON with best-effort date/amount/reference/classification so later parsers can be tightened without re-requesting files.
|
||||
|
||||
Parser patterns observed:
|
||||
- XLSX files may have header row 3 with columns `Fakturanummer`, `Fakturadato`,
|
||||
`Firmakode`, `Beløb`, `Valuta` and supplier metadata in rows 1-2.
|
||||
- Text PDFs often extract as repeated triples:
|
||||
invoice number line, `dd.mm.yyyy firmakode` line, amount/`DKK` line.
|
||||
- Some PDFs split sign into a separate `-` column/line; treat `123,45`, `-`,
|
||||
`DKK` as `-123,45`.
|
||||
- Some PDFs are bank statement screenshots, not payment advice. Classify and do
|
||||
not parse them as settlement lines.
|
||||
- Some files include manual middle calculations or special layouts. Keep raw text
|
||||
and mark parse confidence/status instead of silently forcing a schema.
|
||||
|
||||
Matching/reporting logic:
|
||||
1. Match payment-advice total to bank deposit using ±5 DKK tolerance.
|
||||
2. Then match advice lines to Uniconta by invoice number first, then amount/date.
|
||||
3. Match advice lines to Coop ledger by reference/doc number; in the Coop-Uniconta
|
||||
project this was a very strong signal.
|
||||
4. Treat payment-advice matches as an evidence/explanation layer, not as automatic
|
||||
ledger settlement. An invoice can remain open in Uniconta while Payment_advice
|
||||
proves Coop included it in a settlement, credit, bonus, modregning, or manual
|
||||
middle calculation.
|
||||
5. Report per bank deposit/open item: bank date/amount, matched advice file,
|
||||
difference, line count, Uniconta hit count, Coop hit count, credit/bonus/
|
||||
modregning total, advice documents/vouchers/classifications, and status OK /
|
||||
partial / needs review.
|
||||
6. Keep 2026 or out-of-period advice files in the database, but exclude or label
|
||||
them separately when current bank/Uniconta data only covers 2024-2025.
|
||||
|
||||
### Invoice Status: Net-based + Deduplication
|
||||
|
||||
**Critical pitfall:** The same payment may exist in BOTH systems (Coop RE and Uniconta Betaling). If you sum both, you get double the actual paid amount.
|
||||
|
||||
**Solution:** Calculate net payments per invoice:
|
||||
```python
|
||||
# Net Coop payments (NOT abs-sum!)
|
||||
coop_net = sum(c['amount'] for c in coop
|
||||
if c['reference'] == fnr and c['doc_type'] in ('RE','ZP'))
|
||||
coop_paid = abs(coop_net) if coop_net < 0 else 0
|
||||
|
||||
# Net Uniconta payments
|
||||
uni_net = sum(u['amount'] for u in uni
|
||||
if u['faktura'] == fnr and u['konteringstype'] == 'Betaling')
|
||||
uni_paid = abs(uni_net) if uni_net < 0 else 0
|
||||
|
||||
# Deduplicate: if both exist, use only one
|
||||
if coop_paid > 0 and uni_paid > 0:
|
||||
if abs(coop_paid - uni_paid) <= AMT_TOL:
|
||||
betalt = coop_paid # Same payment recorded twice
|
||||
else:
|
||||
betalt = max(coop_paid, uni_paid) # Partial, take largest
|
||||
```
|
||||
|
||||
**Also handle:** Positive Coop RE amounts are credit memos (modregninger) that should be SUBTRACTED from the net, not added with abs().
|
||||
|
||||
### Match Scoring Formula (default)
|
||||
```
|
||||
Score = 0
|
||||
# Amount match (40%)
|
||||
if abs(a-b) < AMT_TOL: Score += 40
|
||||
|
||||
# Date match (35%)
|
||||
days = abs(d1 - d2).days
|
||||
if days == 0: Score += 35
|
||||
elif days <= 3: Score += 25
|
||||
elif days <= 7: Score += 15
|
||||
elif days <= 14: Score += 5
|
||||
|
||||
# Text match (25%)
|
||||
if both_have_text:
|
||||
similarity = fuzzy_ratio(text1, text2) # 0-100
|
||||
if similarity > 80: Score += 25
|
||||
elif similarity > 50: Score += 12
|
||||
|
||||
# Reference bonus (extra)
|
||||
if reference_match: Score += 40
|
||||
|
||||
# Thresholds
|
||||
if Score >= 85: exact_match, auto_approvable
|
||||
if Score >= 60: partial_match, requires_manual
|
||||
if Score >= 40: fuzzy_match, requires_manual
|
||||
```
|
||||
|
||||
| Table | Key Fields |
|
||||
|-------|-----------|
|
||||
| `ImportBatch` | source (coop/uniconta), filename, imported_at, raw_row_count |
|
||||
| `Posting` | batch_id, source, fiscal_year, posting_date, amount, original_text, normalized_text, reference, document_number, original_data (JSON) |
|
||||
| `Match` | posting_a_id, posting_b_id, score, match_type (exact/partial/fuzzy/manual), status (proposed/approved/rejected), matched_by (auto/user), matched_at |
|
||||
| `AuditLog` | user, action, match_id, old_status, new_status, comment, timestamp |
|
||||
|
||||
## Phase 4: Build Iteratively
|
||||
|
||||
Follow this order strictly:
|
||||
1. Analysis scripts and documentation → `ANALYSE.md`
|
||||
2. Import parser with preview → validate before commit
|
||||
3. Normalization layer (dates, amounts, text)
|
||||
4. Match engine (exact → partial → fuzzy)
|
||||
5. Simple web UI (dashboard → import → matching table)
|
||||
6. Manual approve/reject + audit log
|
||||
7. Reports and exports
|
||||
8. Bonus calculation (if needed)
|
||||
9. Tests + README + Docker
|
||||
|
||||
### Revisor/Auditor Package Pattern
|
||||
|
||||
When the user asks for output that can be sent to an accountant/revisor, produce a deterministic evidence package rather than relying on dashboard screenshots. First verify and remove duplicate import runs after a DB backup; then export a Markdown summary plus an Excel workbook with source files, month totals, open invoice/credit residuals, unexplained bank items, ignored bank items, payment advice rows, and raw normalized postings. Use explicit status language such as `AFSTEMT_MED_ÅBNE_POSTER` when the bank layer reconciles but invoice/credit-note residuals remain.
|
||||
|
||||
See `references/auditor-pack-and-dedup.md` for the duplicate-import cleanup sequence, auditor workbook sheet shape, status semantics, and verification checklist.
|
||||
|
||||
### Idempotent SourceFile Imports
|
||||
|
||||
For accounting import UIs, especially Coop `Payment_advice` bulk uploads, make uploads idempotent. Add a SHA256 `content_hash` on the source-file model, compute it before creating a new row, skip identical reuploads, and update/reparse an existing same-named source file when the content changed. Wrap child-row replacement in a database transaction and test both duplicate-skip and update-replace paths so repeated user uploads do not inflate reconciliation totals.
|
||||
|
||||
See `references/idempotent-source-file-imports.md` for the Django model/view/parser pattern, UI counters, bulk upload setting, and regression checklist.
|
||||
|
||||
## Danish-specific Conventions
|
||||
|
||||
- Dates: use ISO format internally (`YYYY-MM-DD`), but display in Danish
|
||||
(`dd-MM-yyyy`) per user preference. Parse ambiguous Danish text dates with day-first semantics (`dayfirst=True`).
|
||||
- Amounts: always normalize to float/Decimal; display with 2 decimals, Danish comma.
|
||||
- Fiscal year: calendar year (1/1 – 31/12) unless specified otherwise.
|
||||
- For Coop workflows, model invoice year/month separately from payment year/month: Coop can have ~3 months of credit, so December invoices may be paid in March of the following fiscal year. Reports must not treat that as an error by default.
|
||||
- Text: normalize by lowercasing, stripping extra spaces, removing punctuation before fuzzy comparison.
|
||||
- Regulation/bonus/fee/fine classification should be explicit and reviewable. Start with keyword classification (`bonus`, `rabat`, `gebyr`, `bøde`, `regulering`, `korrektion`, `modregning`) but keep raw data and manual review because wording/layout varies.
|
||||
|
||||
## Django Accounting UI Display Formatting
|
||||
|
||||
When the reconciliation deliverable is a Django webapp, treat amount display as part of the reconciliation contract rather than a cosmetic detail. Use one shared template filter/localization utility for money columns in dashboards, match tables, invoice/payment status pages, and reports.
|
||||
|
||||
For Danish accounting UIs, monetary amounts should normally render with `.` thousands separators, `,` decimal separators, exactly two decimals, and a leading minus sign for negative values (for example `1.234.567,89`, `-1.234.567,89`, `0,00`). Prefer `Decimal`-based formatting over Python floats. Replace monetary `|floatformat:2` usage, but do not reformat percentages, match scores, counts, or raw imported/audit JSON unless explicitly requested.
|
||||
|
||||
Minimum verification for UI formatting changes:
|
||||
- Add a unit/template test for the formatting helper.
|
||||
- Add at least one rendered-page regression test proving formatted money appears and raw `1234567.89`-style output is absent.
|
||||
- Confirm templates using the filter load its tag module, and restart the web container if needed.
|
||||
|
||||
See `references/django-accounting-ui-formatting.md` for the concrete Django template-filter implementation, search targets, Docker test commands, and pitfalls.
|
||||
|
||||
## References
|
||||
- See `references/coop-uniconta-analysis.md` for the concrete discovery
|
||||
patterns from the Coop ↔ Uniconta project (header row locations,
|
||||
column mappings, SAP document type codes, etc.).
|
||||
- See `references/coop-uniconta-match-engine.md` for the 5-phase match engine
|
||||
implementation, tolerance matching, net-based invoice status, and score formulas.
|
||||
- See `references/bank-reconciliation.md` for three-source reconciliation extending
|
||||
Coop ↔ Uniconta with bank CSV — month-level validation, bonus reconciliation,
|
||||
and many-to-one (bank payment = sum of Coop invoices) matching strategy.
|
||||
- See `references/coop-payment-advice.md` for using Coop betalingsadviseringer
|
||||
(PDF/XLSX settlement files) as the bridge between bank deposits and invoice
|
||||
lines, including parser patterns, coverage metrics, and reporting shape.
|
||||
- See `references/three-way-django-analysis-engine.md` for Django implementation
|
||||
structure: separate analysis module, batch import command, three-way report
|
||||
outputs, interpretation pitfalls, and minimum tests.
|
||||
- See `references/coop-invoice-payment-allocation.md` for the simplified forward
|
||||
workflow using Coop invoice lists + Uniconta payment allocation + Payment_advice,
|
||||
including separate models, invoice paid/open status logic, employee-expense
|
||||
exceptions, UI badges, and minimum tests.
|
||||
- See `references/coop-uniconta-import-ui-and-payment-advice.md` for the five-source
|
||||
import UI pattern, flexible PDF/Excel Payment_advice raw import, monthly dashboard
|
||||
fields, 3-month Coop credit handling, and regression tests.
|
||||
- See `references/auditor-pack-and-dedup.md` for revisor-ready exports, duplicate
|
||||
import cleanup after DB backup, status semantics such as `AFSTEMT_MED_ÅBNE_POSTER`,
|
||||
and Excel/Markdown verification.
|
||||
- See `references/idempotent-source-file-imports.md` for SHA256-based idempotent
|
||||
accounting source-file uploads, update-vs-skip semantics, Django transaction
|
||||
boundaries, bulk upload limits, and regression tests.
|
||||
- See `references/django-accounting-ui-formatting.md` for Django money-formatting
|
||||
implementation details, template search targets, regression tests, and pitfalls.
|
||||
@@ -0,0 +1,66 @@
|
||||
# Auditor-ready reconciliation packs and duplicate-import cleanup
|
||||
|
||||
Use this when a reconciliation webapp must produce material that can be sent to an accountant/revisor, not just a green dashboard.
|
||||
|
||||
## Revisor-grade status language
|
||||
|
||||
Do not collapse every successful bank reconciliation into “fully reconciled”. Use explicit statuses:
|
||||
|
||||
- `AFSTEMT`: all required sources loaded, relevant bank movements explained, and no open invoice/credit-note residuals.
|
||||
- `AFSTEMT_MED_ÅBNE_POSTER`: bank layer is explained, but invoice/credit-note residuals remain and must be reviewed by accounting/revisor.
|
||||
- `BANK_DIFFERENCER`: required sources loaded, but relevant bank postings are unexplained.
|
||||
- `MANGLER_DATA`: one or more required sources are missing.
|
||||
|
||||
This prevents the dangerous interpretation “0 unexplained bank postings = no accounting work remains”.
|
||||
|
||||
## Duplicate import cleanup pattern
|
||||
|
||||
Historical accounting imports are often rerun. Before trusting totals, check for duplicate `SourceFile` rows by source/file name/parsed count and repeated downstream row counts.
|
||||
|
||||
Safe cleanup sequence:
|
||||
|
||||
1. Create a database backup first.
|
||||
2. Identify duplicate import runs, keeping the canonical/original `SourceFile` rows.
|
||||
3. Delete duplicate `SourceFile` rows through the ORM so related imported rows cascade consistently.
|
||||
4. Record an audit log entry with reason and deleted source file IDs.
|
||||
5. Recompute counts and reports after cleanup.
|
||||
|
||||
Do not modify original source files. The goal is to remove duplicate imported rows, not change evidence.
|
||||
|
||||
## Auditor pack contents
|
||||
|
||||
Generate a deterministic export package with at least:
|
||||
|
||||
- Markdown executive summary in Danish.
|
||||
- Excel workbook with sheets:
|
||||
- `Resume`: data completeness, bank coverage, open balances, status.
|
||||
- `Kildefiler`: imported files, parsed/error counts, timestamps.
|
||||
- `Måneder`: month-level totals across Coop, bank, invoice list, and payment allocation.
|
||||
- `Åbne poster`: invoice/credit-note residuals with invoice number, date, customer/account, total, paid, residual, vouchers, explanation.
|
||||
- `Uforklaret bank`: relevant bank movements that could not be explained.
|
||||
- `Ignoreret bank`: bank movements outside the debtor reconciliation scope, with reason.
|
||||
- `Payment advice`: uploaded settlement/advice records.
|
||||
- `Rå posteringer`: normalized postings for traceability.
|
||||
|
||||
Use Danish number formatting in human-facing Markdown and UI (`1.234.567,89`). Excel cells may stay numeric where useful, but headings and sheet names should be accountant-readable.
|
||||
|
||||
## Implementation notes
|
||||
|
||||
- Keep the analysis/export layer read-only; it should not mutate match status.
|
||||
- Convert timezone-aware datetimes to ISO strings before writing with pandas/openpyxl; Excel rejects timezone-aware datetimes.
|
||||
- Include both gross open amount and net open balance. Accountants often need both:
|
||||
- gross open review amount = sum of absolute residuals;
|
||||
- positive residuals = potential receivables;
|
||||
- negative residuals = open credits/modregninger;
|
||||
- net residual = positive + negative.
|
||||
- Add tests that read the generated workbook back and assert key sheets/rows exist.
|
||||
- If visual browser automation is unavailable, verify dashboard output with Django’s test client/HTTP response text and keep a separate browser/UI check when the environment supports it.
|
||||
|
||||
## Minimum verification
|
||||
|
||||
- Framework check passes.
|
||||
- Full reconciliation/core tests pass.
|
||||
- Auditor export command runs idempotently.
|
||||
- Workbook exists and contains all expected sheets.
|
||||
- Summary status and open totals match the recomputed analysis function.
|
||||
- Dashboard shows bank coverage and open-poster status consistently with the exported package.
|
||||
@@ -0,0 +1,73 @@
|
||||
# Bank Reconciliation — Django Implementation Guide
|
||||
|
||||
Practical implementation notes for extending a two-source Coop↔Uniconta reconciliation webapp with bank CSV support.
|
||||
|
||||
## Model extensions
|
||||
|
||||
Add these fields to `Transaction` (one migration):
|
||||
```python
|
||||
account_number = models.CharField(max_length=50, blank=True, db_index=True)
|
||||
bank_text = models.TextField(blank=True)
|
||||
valuta_date = models.DateField(null=True, blank=True)
|
||||
```
|
||||
|
||||
Also extend `SourceFile.source` choices and `Transaction.source`/`match_type` choices to include `("bank", "Bank")` and `("bank_total", "Bank-samlet")`.
|
||||
|
||||
## Bank CSV parser
|
||||
|
||||
Key implementation in `apps/import_/parsers.py`:
|
||||
|
||||
- **Encodings**: Try `latin-1` first, then `utf-8`. Most Danish bank exports use latin-1.
|
||||
- **Delimiter**: `;` (semicolon)
|
||||
- **Header mapping**: Fuzzy-match by lowercasing, stripping null bytes and quotes:
|
||||
- `dato` → date column
|
||||
- `valør`/`valuta` → value date column
|
||||
- `tekst`/`beskrivelse` → text column
|
||||
- `bel` (prefix) → amount column
|
||||
- `saldo` → balance column
|
||||
- `reference`/`bilagsrefer` → reference column
|
||||
- **Date parse**: `DD.MM.YYYY` → split by dots → `datetime(int(parts[2]), int(parts[1]), int(parts[0])).date()`
|
||||
- **Amount parse**:
|
||||
1. Strip `\xa0` (nbsp), `\u202f` (narrow nbsp), `" DKK"`
|
||||
2. If `.` and `,` both present → remove dots, replace comma with dot
|
||||
3. If only comma → replace with dot
|
||||
4. Convert to `Decimal`
|
||||
- **Bulk create**: Use `Transaction.objects.bulk_create(txns_to_create)` for performance
|
||||
- **Account extraction**: Regex `r"(\d{8,})"` on text field to find kundenummer/kontonummer
|
||||
|
||||
## Bank ↔ Coop match (many-to-one)
|
||||
|
||||
Group bank payments by `(account_number, year, month)` and match against the sum of unmatched Coop transactions in the same bucket.
|
||||
|
||||
```python
|
||||
bank_txns = Transaction.objects.filter(source="bank", year=year).values(
|
||||
"account_number", "year", "month"
|
||||
).annotate(total=Sum("amount"), count=Count("id"))
|
||||
|
||||
for bt in bank_txns:
|
||||
coop_sum = Transaction.objects.filter(
|
||||
source="coop", year=bt["year"], month=bt["month"],
|
||||
match_status="unmatched", account_number=bt["account_number"],
|
||||
).aggregate(s=Sum("amount"))["s"]
|
||||
|
||||
if coop_sum is not None and abs(bt["total"] - coop_sum) <= TOLERANCE:
|
||||
# Create MatchGroup linking bank bucket to Coop bucket
|
||||
...
|
||||
```
|
||||
|
||||
**Note**: Bank amounts are positive (indbetalinger), Coop RE/Bankoverførsel amounts are negative. Compare absolute values. Set `match_type="bank_total"`.
|
||||
|
||||
## Views/templates changes
|
||||
|
||||
- **Dashboard**: Add `bank_count`, `bank_sum` cards; add `bank_count` / `bank_sum` columns to monthly overview table
|
||||
- **Upload form**: Add `<option value="bank">Bank (Kontobevaegelser .csv)</option>`
|
||||
- **Reconciliation list**: Add `bank` to source filter; render bank badge with `.badge-warning`
|
||||
- **Run match page**: Document match phases: Reference → ZV↔Faktura → RE/Betaling → General → Many-to-one → Bank-total
|
||||
|
||||
## Discovered pitfalls
|
||||
|
||||
1. **Migrations not detected**: If `makemigrations` reports "No changes detected", check whether a prior manual migration already captured the changes. Run `migrate --check` first.
|
||||
2. **Docker file path mismatch**: When testing via `docker compose exec`, container filesystem may differ from host. Use the web UI upload form for file import testing.
|
||||
3. **Bank text homogeneity**: All transactions from the same counterparty often have identical text (e.g. `"Coop 0010014808 -SE MEDD."`). Text similarity is useless for matching. Rely on `account_number` + date+amount grouping.
|
||||
4. **Bank negatives are rare but critical**: Negative bank amounts (~1-2%) represent bonus/modregning/retur. Do NOT match these against positive Coop invoices. Flag separately for bonus reconciliation.
|
||||
5. **Bank total may differ slightly from Coop sum**: Normal difference is ±0.1% (fees, rounding). Use tolerance matching (e.g. ±5 DKK) for bank-total matches.
|
||||
@@ -0,0 +1,222 @@
|
||||
# Bank Reconciliation — Three-Source Analysis (2026-05-26)
|
||||
|
||||
Reference for matching bank transactions against Coop SAP exports and Uniconta
|
||||
ledgers. This extends the two-source reconciliation with a third (bank) dimension
|
||||
that serves as the "ground truth" for actual cash flow.
|
||||
|
||||
---
|
||||
|
||||
## Bank file format (Danske Bank CSV export)
|
||||
|
||||
| Kolonne | Type | Brug |
|
||||
|---------|------|------|
|
||||
| Dato | DD.MM.YYYY | Primær dato |
|
||||
| Valør | DD.MM.YYYY | Sekundær dato |
|
||||
| Tekst | string | **Identifikation** |
|
||||
| (tom) | — | Padding |
|
||||
| Beløb | +/- flot/DK | Primær beløb |
|
||||
| Saldo | flot/DK | Sekundær — løbende saldo |
|
||||
| Egen bilagsreference | string | Audit/trace |
|
||||
|
||||
**CSV separator**: `;` (semikolon)
|
||||
**Encoding**: `latin-1` (CP1252)
|
||||
**Decimal**: Danish comma, Danish point-as-thousand-separator
|
||||
**Negative**: explicit minus in front
|
||||
|
||||
Typisk tekstmønster for Coop-indbetalinger:
|
||||
- `"Coop 0010014808 -SE MEDD."`
|
||||
- Samme kundenummer for ALLE transaktioner
|
||||
- Stort set alle transaktioner har identisk tekst
|
||||
|
||||
---
|
||||
|
||||
## Bank transaktionskategorier
|
||||
|
||||
### Standard indbetalinger (~98%)
|
||||
- Beløb fra ~30.000 kr til ~800.000 kr
|
||||
- Sum 2024: +9.225.626 kr (129 stk)
|
||||
- Sum 2025: +10.031.368 kr (118 stk)
|
||||
- Udgør den "faktiske" Coop-omsætning der rammer banken
|
||||
|
||||
### Negative poster — bonus/modregning/retur (~1-2%)
|
||||
| Dato | Beløb | Tekst | Type |
|
||||
|------|-------|-------|------|
|
||||
| 13.06.2024 | -18.750,00 | COOP DANMARK AS | Bonus-tilbagebetaling |
|
||||
| 25.11.2024 | -31,25 | Coop 365 | Lokal retur/modregning |
|
||||
| 22.04.2025 | -128,00 | DK 57014 Coop Kvickly Slagelse | Butiksretur |
|
||||
| 13.05.2025 | -40,00 | DK 50960 Coop Kvickly Slagelse | Butiksretur |
|
||||
|
||||
**Total negative: -18.949,25 kr**
|
||||
|
||||
---
|
||||
|
||||
## THREE-SOURCE MATCHING CONCEPT
|
||||
|
||||
```
|
||||
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
|
||||
│ Coop Excel │ │ Uniconta Excel │ │ Bank CSV │
|
||||
│ (fakturaer) │ │ (bogføring) │ │ (indbetalinger)│
|
||||
└────────┬────────┘ └────────┬────────┘ └────────┬────────┘
|
||||
│ │ │
|
||||
▼ ▼ ▼
|
||||
┌──────────────────────────────────────────────────────────────────┐
|
||||
│ MATCH LAYER 1: Post-for-post (Coop ↔ Uniconta) │
|
||||
│ Reference+abs(amount)+date → 68% match rate │
|
||||
├──────────────────────────────────────────────────────────────────┤
|
||||
│ MATCH LAYER 2: Month-level (All 3 sources) │
|
||||
│ Sum(Coop) ≈ Sum(Uniconta) ≈ Sum(Bank) per month │
|
||||
├──────────────────────────────────────────────────────────────────┤
|
||||
│ MATCH LAYER 3: Bonus reconciliation (separate) │
|
||||
│ Sum(KG + Kreditnota) per period ≈ Bank negative per release │
|
||||
│ (tolerate time lag — bank credit arrives later than postings) │
|
||||
└──────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Bank-to-Coop / Bank-to-Uniconta matching
|
||||
|
||||
### Month-level validation (PRIMARY)
|
||||
|
||||
Compare month totals across all three sources. If Coop and Uniconta diverge,
|
||||
the bank settles the dispute.
|
||||
|
||||
| Måned | Bank (kr) | Coop (kr) | Uniconta (kr) | Status |
|
||||
|-------|-----------|-----------|---------------|--------|
|
||||
| 2024-04 | 2.139.610 | TBD | TBD | To verify |
|
||||
| 2024-05 | 1.632.814 | TBD | TBD | To verify |
|
||||
| 2024-09 | 1.053.598 | TBD | TBD | To verify |
|
||||
|
||||
**Interpretation rule:**
|
||||
- Bank ≈ Coop: period is captured correctly in Coop
|
||||
- Bank ≈ Uniconta: period is correct in Uniconta
|
||||
- Bank >> Coop: missing Coop postings (likely late/omitted)
|
||||
- Bank >> Uniconta: missing Uniconta postings
|
||||
- Bank << both: bonus/modregning not yet reflected in accounting
|
||||
|
||||
### Many-to-one: Bank single payment = sum of Coop invoices
|
||||
|
||||
Each bank posting (e.g. 641.168,06 kr) is a **consolidated transfer** that
|
||||
aggregates many individual Coop invoices. Algorithm:
|
||||
1. Take bank amount B on date D
|
||||
2. Find Coop transactions in [D-30 days, D+3 days] that sum to B ± tolerance
|
||||
3. This is a "subset-sum" / knapsack search:
|
||||
```python
|
||||
def find_combo(target, candidates, tol=5.0):
|
||||
# Greedy for small candidate sets
|
||||
# DP for larger sets
|
||||
candidates = sorted([c for c in candidates
|
||||
if abs(c.amount) <= target + tol])
|
||||
# ... subset sum with tolerance
|
||||
```
|
||||
4. If exact match: mark all matched Coop invoices as "grouped under bank payment"
|
||||
5. If near-match (±500 kr): flag for manual review — likely partial payment or fees
|
||||
|
||||
---
|
||||
|
||||
## Bonus / Rebate reconciliation
|
||||
|
||||
### The mismatch:why Bank negatives don't match individual KG/Kreditnota
|
||||
|
||||
Coop KG-posteringer: 486 stk, total -1.101.849,97 kr
|
||||
Bank negative total: -18.949,25 kr (4 stk)
|
||||
Uniconta Kreditnota total: -4.593.829,17 kr
|
||||
Uniconta Rabat: 1 stk, -178.194,55 kr
|
||||
|
||||
**Observation:** Bank -18.750 kr does NOT match ANY individual Coop KG posting
|
||||
(nor any Uniconta Kreditnota). The nearest Uniconta Kreditnota is -18.961 (2025-03-21).
|
||||
|
||||
**Interpretation:** Bank negatives are **consolidated credit notes or bonus
|
||||
payments** that aggregate MANY smaller individual credit postings across a period
|
||||
(quarter or year). They arrive in the bank later than the individual postings.
|
||||
|
||||
### Suggested bonus reconciliation strategy
|
||||
|
||||
1. **Period-based grouping** (not transaction-based):
|
||||
- Group Coop KG by quarter/year
|
||||
- Group Uniconta Kreditnota by linked Faktura-number
|
||||
- Group Uniconta Rabat by year
|
||||
2. **Compare aggregated totals per period** with bank negatives per period:
|
||||
```
|
||||
Coop KG per period = sum(KG amounts where Posting Date in period)
|
||||
Uniconta credit per period = sum(Kreditnota + Rabat amounts in period)
|
||||
Bank negative per period = sum(bank negative amounts with Value Date in period)
|
||||
```
|
||||
3. **Differences indicate:
|
||||
- Timing: bank credit arrives 1-3 months after postings
|
||||
- Pooling: quarterly/yearly bonus vs monthly individual credits
|
||||
- Source mismatch: some credits originate from Uniconta only, some from Coop only
|
||||
4. **Manual review** for amounts > 50.000 kr discrepancies per period
|
||||
|
||||
### Special: Uniconta "Rabat" -178.194,55 "Coop bonus 2023"
|
||||
|
||||
- Single large posting on 2024-01-16
|
||||
- No matching bank negative from 2023 or early 2024 found
|
||||
- Likely an **accrual/adjustment** recognized in 2024 for 2023 bonus, but the
|
||||
actual cash flow may have been handled differently (e.g. applied as discount
|
||||
on future invoices rather than direct bank transfer)
|
||||
- **Flag:** accrual-type postings should NOT be matched against bank — they
|
||||
are non-cash adjustments. Separate them into "bonus accrued" vs "bonus paid".
|
||||
|
||||
---
|
||||
|
||||
## Key insight: Three sources triangulate
|
||||
|
||||
| Question | How to answer with 3 sources |
|
||||
|----------|-------------------------------|
|
||||
| "Is this invoice paid?" | Uniconta Faktura + Coop RE + Bank amount all agree |
|
||||
| "Did bonus arrive in cash?" | Bank negative confirms, Coop KG shows breakdown |
|
||||
| "Is Uniconta missing a posting?" | Bank has it but Uniconta doesn't → missing entry |
|
||||
| "Is Coop missing a posting?" | Bank has it but Coop doesn't → missing entry |
|
||||
| "Is this payment split?" | Bank single amount = sum of multiple Coop invoices |
|
||||
| "Is there a bonus adjustment?" | Coop KG present, Uniconta Kreditnota present, but bank negative missing → accrual |
|
||||
|
||||
---
|
||||
|
||||
## Practical notes for the webapp
|
||||
|
||||
### Bank import
|
||||
- Import as a third source alongside Coop and Uniconta
|
||||
- Store as `Transaction(source="bank", ...)` with minimal fields:
|
||||
- date, amount, text, original_value_date, original_balance
|
||||
- Apply the same year filter (2024-2025) to exclude stray transactions
|
||||
|
||||
### Bank-specific views
|
||||
1. **Month-level dashboard**: Three-bar chart per month (Coop, Uniconta, Bank)
|
||||
2. **Divergence report**: Months where |Coop - Bank| > tolerance
|
||||
3. **Bonus reconciliation**: Separate tab for aggregated KG/Kreditnota/Rabat vs bank negatives
|
||||
4. **Many-to-one explorer**: Click a bank transaction to see suggested Coop invoice combinations that sum to the bank amount
|
||||
|
||||
### Data model extension for bank
|
||||
```python
|
||||
class BankTransaction(models.Model):
|
||||
date = models.DateField()
|
||||
value_date = models.DateField(null=True)
|
||||
amount = models.DecimalField(max_digits=18, decimal_places=2)
|
||||
text = models.TextField()
|
||||
balance = models.DecimalField(max_digits=18, decimal_places=2)
|
||||
source_file = models.ForeignKey(SourceFile, ...)
|
||||
```
|
||||
Or reuse existing `Transaction` model with `source="bank"`.
|
||||
|
||||
---
|
||||
|
||||
## Performance / verification script
|
||||
|
||||
Use `docker exec` with the container's Python + pandas to parse bank CSVs
|
||||
and cross-check against already-loaded Coop/Uniconta data.
|
||||
|
||||
```python
|
||||
# Quick verification of month-level totals
|
||||
import pandas as pd
|
||||
|
||||
def parse_bank_csv(path):
|
||||
df = pd.read_csv(path, sep=';', encoding='latin-1',
|
||||
header=0, names=['Dato','Valør','Tekst','_','Beløb','Saldo','Ref'])
|
||||
df['Dato'] = pd.to_datetime(df['Dato'], format='%d.%m.%Y', errors='coerce')
|
||||
df['Beløb'] = df['Beløb'].str.replace('.', '').str.replace(',', '.').astype(float)
|
||||
return df
|
||||
|
||||
# Summarize by month and compare with Coop/month, Uniconta/month
|
||||
# Run this as a Django management command or script inside the container
|
||||
```
|
||||
+109
@@ -0,0 +1,109 @@
|
||||
# Coop fakturaliste + Uniconta betalingsfordeling workflow
|
||||
|
||||
Session-derived pattern from the Coop-Uniconta project: simplify reconciliation by using three forward-looking sources instead of many historical exports.
|
||||
|
||||
## Target workflow
|
||||
|
||||
Use these files as the primary operational import set:
|
||||
1. Coop invoice list (`Coop Fakturaer YYYY.xlsx`) — all invoices/credits issued to Coop.
|
||||
2. Uniconta payment allocation / debtor ledger (`Coop Debitor YYYY.xlsx`) — invoice and settlement/payment lines.
|
||||
3. Coop `Payment_advice` — later bridge source for explicit settlement detail.
|
||||
|
||||
This supports the user question: “Mangler Coop at betale nogle fakturaer?”
|
||||
|
||||
## Models / concepts
|
||||
|
||||
Recommended separate models, not generic postings only:
|
||||
|
||||
### `CoopInvoice`
|
||||
- source file, fiscal year/accounting year
|
||||
- date
|
||||
- account number/name
|
||||
- customer reference (`Deres ref`)
|
||||
- order number
|
||||
- invoice number
|
||||
- requisition
|
||||
- total, net, VAT amounts
|
||||
- payment terms / sent dates if present
|
||||
- raw row JSON + sheet/row reference
|
||||
|
||||
### `PaymentAllocationLine`
|
||||
- source file, fiscal/accounting year
|
||||
- date
|
||||
- account
|
||||
- invoice number
|
||||
- voucher/document number
|
||||
- posting number
|
||||
- text
|
||||
- debit, credit, VAT
|
||||
- physical voucher/origin/account type/debtor-creditor/name
|
||||
- `is_payment_line` boolean
|
||||
- raw row JSON + sheet/row reference
|
||||
|
||||
## Payment-line detection
|
||||
|
||||
In Uniconta debtor/payment allocation files, not every line is a payment. They may contain both original invoice postings and later settlement lines.
|
||||
|
||||
Observed signals for payment/settlement lines:
|
||||
- `Bilag != Faktura`
|
||||
- text contains `Coop 0010014808`
|
||||
- text contains `Omp. Coop`
|
||||
|
||||
Do not infer paid status from invoice existence alone; calculate it from settlement/payment lines linked by invoice number.
|
||||
|
||||
## Invoice payment status algorithm
|
||||
|
||||
For each `CoopInvoice.invoice_number`:
|
||||
1. Find `PaymentAllocationLine` rows with same invoice number.
|
||||
2. Sum payment/settlement effect from rows marked `is_payment_line`.
|
||||
3. Compare against `CoopInvoice.total_amount`.
|
||||
4. Emit status:
|
||||
- `paid`: residual within tolerance
|
||||
- `partial`: some payment but residual remains
|
||||
- `open`: no settlement/payment lines found
|
||||
5. Keep explanatory fields: payment vouchers, dates, amount paid, residual, source rows.
|
||||
|
||||
Use Danish accounting display, but ISO dates and Decimal internally.
|
||||
|
||||
## Bank reconciliation with allocation groups
|
||||
|
||||
Use the payment allocation as an explanation layer between bank deposits and invoices:
|
||||
|
||||
`Bank payment → allocation group/voucher → multiple invoice lines → Coop invoices`
|
||||
|
||||
This is especially useful when one bank deposit equals many invoices/credits. Build groups by voucher/payment text/date and compare the group total to the bank payment.
|
||||
|
||||
## Employee expense exception
|
||||
|
||||
Negative bank postings whose text indicates purchases at Coop/Kvickly may be employee expenses, not Coop customer payments. Treat them as ignored/out-of-scope only when the user confirms the business meaning. In the Coop-Uniconta project, two negative bank lines with Coop/Kvickly text were confirmed as employee outlays and excluded from relevant bank coverage.
|
||||
|
||||
Implementation pattern:
|
||||
- keep ignored bank payments visible in reports
|
||||
- do not count them as unmatched customer-payment differences
|
||||
- report both total bank count and relevant bank coverage percentage
|
||||
|
||||
## UI/reporting shape
|
||||
|
||||
Add a dedicated invoice-payment view or dashboard section:
|
||||
- paid invoices
|
||||
- open/different invoices
|
||||
- total invoices
|
||||
- open residual amount
|
||||
- status badges: `BETALT`, `ÅBEN`, `DIFFERENCE`, `FAKTURAER`
|
||||
- filters for year, customer/account, amount, residual, status
|
||||
- export of open/difference invoices to Excel/CSV
|
||||
|
||||
## Minimum tests
|
||||
|
||||
Add tests for:
|
||||
1. A Coop invoice is marked paid via matching Uniconta payment allocation lines.
|
||||
2. Partial/open invoices calculate residual correctly.
|
||||
3. Confirmed Coop/Kvickly employee expense bank lines are ignored but still reported.
|
||||
4. Relevant bank coverage is 100% when all customer-payment bank lines are explained and only confirmed out-of-scope lines remain.
|
||||
|
||||
## Pitfalls
|
||||
|
||||
- Do not hardcode confidential file paths or sample data into committed code.
|
||||
- Keep raw Excel/source folders gitignored.
|
||||
- Do not collapse invoice list, payment allocation, and payment advice into one generic parser; their semantics differ.
|
||||
- Payment advice parser can be added later, but wire the import type early so UI/workflow is ready.
|
||||
@@ -0,0 +1,125 @@
|
||||
# Coop betalingsadviseringer / payment advice as reconciliation bridge
|
||||
|
||||
Use this when Coop provides per-payment settlement/advice documents (PDF/XLSX)
|
||||
for bank deposits. These documents can turn previously unexplained bank deposits
|
||||
into explainable groups of invoice, credit, bonus, and modregning lines.
|
||||
|
||||
## Durable workflow
|
||||
|
||||
1. Inventory files by extension and classify each file:
|
||||
- `payment_advice` — Coop betalingsadvisering/payment advice with invoice lines.
|
||||
- `bank_statement` — bank print/screenshot containing a Coop deposit message.
|
||||
- `unknown` — manual calculation, unsupported layout, or needs OCR/manual review.
|
||||
2. Extract payment advice rows:
|
||||
- invoice/fakturanummer
|
||||
- invoice date/fakturadato
|
||||
- company code/firmakode
|
||||
- amount/beløb
|
||||
- currency/valuta
|
||||
- raw row/text and parse confidence
|
||||
3. Compute advice totals from parsed rows; do not rely solely on a trailing total
|
||||
because some PDFs have manually added middle calculations or split sign columns.
|
||||
4. Match advice total to bank transaction with ±5 DKK tolerance.
|
||||
5. Match advice lines to Coop ledger by reference/doc number and to Uniconta by
|
||||
invoice number. Fall back to amount/date only for remaining lines.
|
||||
6. Produce a bank-deposit-level report for accounting review.
|
||||
|
||||
## PDF/XLSX parser patterns
|
||||
|
||||
XLSX pattern:
|
||||
- Supplier metadata may be rows 1-2.
|
||||
- Header row often contains: `Fakturanummer`, `Fakturadato`, `Firmakode`, `Beløb`,
|
||||
`Valuta`.
|
||||
- Dates can already be Excel datetimes.
|
||||
|
||||
Text PDF pattern A:
|
||||
```text
|
||||
Payment Advice
|
||||
Fakturanummer
|
||||
Fakturadato
|
||||
Firmakode Beløb
|
||||
Valuta
|
||||
5119751312
|
||||
17.11.2023 2245
|
||||
-1556,98 DKK
|
||||
...
|
||||
```
|
||||
|
||||
Text PDF pattern B with split sign:
|
||||
```text
|
||||
5126294822
|
||||
24.10.2024 1000
|
||||
2.119,86
|
||||
-
|
||||
DKK
|
||||
```
|
||||
Treat the amount as negative when a standalone `-` follows the amount.
|
||||
|
||||
Bank-statement PDFs:
|
||||
- May include `Kontobevægelser`, `Coop Danmark A/S -SE MEDD.`, and a bank amount.
|
||||
- Do not parse these as settlement invoice lines; classify separately and optionally
|
||||
link to the bank transaction/message.
|
||||
|
||||
## Reporting shape
|
||||
|
||||
Per bank deposit show:
|
||||
- bank date
|
||||
- bank amount
|
||||
- matched payment advice file(s)
|
||||
- total according to advice
|
||||
- difference
|
||||
- number of advice lines
|
||||
- number of lines found in Coop ledger
|
||||
- number of invoice numbers found in Uniconta
|
||||
- credit/negative line total
|
||||
- status: `OK`, `partial`, `needs_review`, `out_of_period`
|
||||
|
||||
## Interpretation pitfalls
|
||||
|
||||
- Advice files may include invoice dates from prior years but payment dates in the
|
||||
current year. Derive both invoice year and payment/bank year.
|
||||
- A low overlap with existing “likely Coop owes” flags does not mean the advice
|
||||
files are useless; it means the analysis must model settlement files as their
|
||||
own source rather than using them as loose attachments.
|
||||
- 2026 advice files should be imported but labelled out-of-period when bank or
|
||||
Uniconta data for 2026 is incomplete.
|
||||
- Keep original PDFs/XLSX confidential and out of git; only commit parser code,
|
||||
schemas, docs, and non-sensitive aggregate findings.
|
||||
|
||||
## Coop-Uniconta session observations
|
||||
|
||||
A mixed folder of 125 files contained 116 PDFs and 9 XLSX files. A first pass found
|
||||
122 payment advice files, 2 bank-statement PDFs, and 1 unknown/special PDF. Parsed
|
||||
advice lines were strong as a bridge: most lines matched Coop references/doc
|
||||
numbers, many matched Uniconta invoice numbers, and many previously unexplained
|
||||
bank deposits could be explained by matching advice totals to bank amounts within
|
||||
±5 DKK.
|
||||
|
||||
Later production import of the same class of files used PyMuPDF (`fitz`) for PDFs
|
||||
and pandas/openpyxl-style Excel parsing, creating one source-file record per
|
||||
Payment_advice file and line-level records with raw JSON/text. A folder with 125
|
||||
files (116 PDF, 9 Excel) produced 2,727 Payment_advice lines without mutating
|
||||
original files. For a 2025 auditor pack, 1,253 lines were relevant and 78 of 296
|
||||
open invoice/credit-note residuals had a Payment_advice trail.
|
||||
|
||||
Important interpretation: a Payment_advice line must not automatically close or
|
||||
"pay" an invoice in the ledger analysis. Treat it as an evidence/explanation layer
|
||||
for open items: the invoice can remain open in Uniconta while the advice proves
|
||||
Coop included it in a settlement, modregning, credit, bonus, or manual middle
|
||||
calculation. Reports should therefore add fields such as advice line count, advice
|
||||
amount, document list, voucher/bilag list, classifications, and a human-readable
|
||||
explanation rather than changing ledger status silently.
|
||||
|
||||
Implementation pattern that worked:
|
||||
- Add `PyMuPDF>=1.24` to backend dependencies before relying on PDF import in Docker.
|
||||
- Extend the batch importer to include `Entydigt*.pdf`, `Entydigt*.xlsx`, `*.xlsm`,
|
||||
and `*.xls` from the payment-advice folder.
|
||||
- Parse file names like `Entydigt id_4765_Bilag_7283.xlsx` into unique/document id
|
||||
and voucher/bilag metadata when present.
|
||||
- Classify lines as payment, credit, adjustment/manual total, or unknown, while
|
||||
preserving raw row/text for manual review.
|
||||
- Add regression tests that cover the case "open invoice has Payment_advice trail
|
||||
but no Uniconta payment line"; expected output is explanatory evidence, not an
|
||||
auto-paid status.
|
||||
- Regenerate Markdown + Excel auditor packs after import and include Payment_advice
|
||||
coverage counts in the summary.
|
||||
@@ -0,0 +1,235 @@
|
||||
# Coop ↔ Uniconta Reconciliation Reference (Aktualiseret 2025-05-26)
|
||||
|
||||
Konkrete felter, mappings og pitfalls fra det aktuelle Coop-Uniconta projekt.
|
||||
|
||||
---
|
||||
|
||||
## Coop (SAP export) — Kontoudtog.xlsx
|
||||
|
||||
| Kolonne | Navn | Type | Match-brug |
|
||||
|---------|------|------|------------|
|
||||
| Entry Date | Entry Date | datetime | Sekundær |
|
||||
| Document Type | Document Type | **string** | **Filter + match-type** |
|
||||
| Document Date | Document Date | datetime | Sekundær |
|
||||
| **Posting Date** | Posting Date | **datetime** | **Primær dato** |
|
||||
| **Reference** | Reference | **string (75% non-null)** | **Primær match-nøgle** |
|
||||
| Net Due Date | Net Due Date | datetime | Forfaldsdato |
|
||||
| **Amount** | Amount in Local Currency | **float** | **Primær beløb** |
|
||||
| Document Currency | Local Currency | string ('DKK') | Filter |
|
||||
| **Document Number** | Document Number | int | **Audit/trace** |
|
||||
| Text | Text | string (20% non-null) | Sekundær match |
|
||||
| Payment Reference | Payment Reference | string (12% non-null) | OCR/trace |
|
||||
| User Name | User Name | string | Audit |
|
||||
|
||||
### Coop Document Types (2024-2025)
|
||||
| Type | Andel | Beskrivelse | Uniconta-modpart |
|
||||
|------|-------|-------------|----------------|
|
||||
| RE | ~50% | Regning/betaling | Betaling/Faktura |
|
||||
| ZV | ~15% | Udbetalingsforslag/faktura | Betaling |
|
||||
| RG | ~6% | Kontoafstemning | Kreditnota |
|
||||
| KG | ~5% | Kontoafstemning gruppe (bonus/rabat) | Kreditnota |
|
||||
| ZC | ~3% | Kreditering | Kreditnota |
|
||||
| ZP | ~0.4% | Betaling | Betaling |
|
||||
| XX | ~0% (filtreret fra) | Manuelt bogført/migrering | — |
|
||||
|
||||
**VIGTIGT:** ZV-poster er positive fakturaer udstedt til kunder. De har typisk **ingen Reference/Faktura-nummer** og matches via `abs(beløb) + dato` mod Uniconta Betalinger.
|
||||
|
||||
---
|
||||
|
||||
## Uniconta — Coop danmark.xlsx
|
||||
|
||||
### Fil-struktur
|
||||
- Række 1: Konto-info (Checked, Konto, Kontonavn, Adresse, Bynavn)
|
||||
- Række 2-3: Tomme
|
||||
- Række 4: Kolonne-headers: Dato, Faktura, Forfaldsdato, Bilag, Tekst, Beløb, Resterende, Forfalden, Konteringstype, Sum
|
||||
- Række 5+: Data
|
||||
|
||||
| Kolonne | Navn | Type | Match-brug |
|
||||
|---------|------|------|------------|
|
||||
| **Dato** | Dato | **datetime** | **Primær dato** |
|
||||
| **Faktura** | Faktura | **mixed int/str, 0=ingen** | **Primær match-nøgle** |
|
||||
| Forfaldsdato | Forfaldsdato | datetime | Sekundær |
|
||||
| Bilag | Bilag | mixed int/str | Audit |
|
||||
| **Tekst** | Tekst | **string** | **Tekst-match** |
|
||||
| **Beløb** | Beløb | **float/tekst** | **Primær beløb** |
|
||||
| Resterende | Resterende | float | Filter |
|
||||
| Forfalden | Forfalden | float | Filter |
|
||||
| **Konteringstype** | Konteringstype | **string** | **Filter + match-bonus** |
|
||||
| Sum | Sum | float | IGNORE — løbende saldo |
|
||||
|
||||
### Uniconta Konteringstyper
|
||||
| Type | Beskrivelse | Match-modpart (Coop) |
|
||||
|------|-------------|---------------------|
|
||||
| Faktura | Faktura udstedt | RE (betaling) eller ZV |
|
||||
| Betaling | Betaling modtaget | RE, ZP, ZV |
|
||||
| Kreditnota | Kreditnota | RG, ZC, KG |
|
||||
| Primo | Primo-saldo | XX (filtreres fra) |
|
||||
| Afslutning | Årsafslutning | — |
|
||||
| Overførsel | Overførselsbilag | — |
|
||||
| Manuel | Manuel posting | — |
|
||||
| Rabat | Rabat | KG |
|
||||
|
||||
---
|
||||
|
||||
## Match-nøgle-observationer
|
||||
|
||||
### Fakturanumre overlapper IDENTISK
|
||||
- Coop `Reference` og Uniconta `Faktura` bruger **identiske 4-5 cifrede tal**.
|
||||
- Eksempler: 46344, 44907, 48456, 49495, 51519
|
||||
- **2.970 unikke RE-references findes som Faktura-numre i Uniconta.**
|
||||
- Dette er den stærkeste match-signal — langt stærkere end beløb+dato alene.
|
||||
|
||||
### Modsat fortegn er NORMALT
|
||||
| Coop type | Coop fortegn | Uniconta type | Uniconta fortegn | Match |
|
||||
|-----------|-------------|---------------|-----------------|-------|
|
||||
| RE (betaling) | Negativ (-) | Faktura | Positiv (+) | abs(beløb) |
|
||||
| ZV (faktura) | Positiv (+) | Betaling | Negativ (-) | abs(beløb) |
|
||||
| RG (kreditnota) | Positiv (+) | Kreditnota | Negativ (-) | abs(beløb) |
|
||||
| KG (rabat) | Positiv (+) | Kreditnota | Negativ (-) | abs(beløb) |
|
||||
|
||||
**Implikation:** Match-motoren SKAL sammenligne `abs(beløb)`, ikke rå beløb.
|
||||
|
||||
### Dato-tolerance
|
||||
- Faktura-reference matches: ±21 dage fungerer fint
|
||||
- Generelle beløb+dato matches: ±21 dage fanger de fleste
|
||||
- ZV↔Betaling matches: ±30 dage anbefales (betalingsforslag vs. faktura)
|
||||
|
||||
---
|
||||
|
||||
## Normalisering
|
||||
|
||||
### Beløb (dansk format)
|
||||
Nogle Uniconta-celler indeholder tekststrenge som `-2.543.803,93`:
|
||||
```python
|
||||
def parse_danish_number(val):
|
||||
if pd.isna(val) or val is None: return None
|
||||
if isinstance(val, (int, float)): return float(val)
|
||||
s = str(val).strip().replace('\xa0','').replace('\u202f','').replace(' DKK','')
|
||||
if s.lower() in ('nan',''): return None
|
||||
if s.endswith('.0'): s = s[:-2]
|
||||
if '.' in s and ',' in s:
|
||||
s = s.replace('.','').replace(',','.')
|
||||
elif ',' in s:
|
||||
s = s.replace(',','.')
|
||||
try: return float(s)
|
||||
except: return None
|
||||
```
|
||||
|
||||
### Tekst
|
||||
```python
|
||||
def normalize_text(t):
|
||||
if not t or pd.isna(t): return ''
|
||||
return re.sub(r'\s+',' ',str(t).lower().strip()).strip()
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Konto-specifikt mønster
|
||||
|
||||
### "Coop 0010014808"
|
||||
- Uniconta indeholder mange poster med teksten "Coop 0010014808"
|
||||
- Beløb varierer fra -1.155.458,24 til -65.000
|
||||
- Nogle har fakturanummer, andre har ikke
|
||||
- Dette ligner en **samlet konto-kontering** for Coop-omsætning, ikke enkeltfakturaer
|
||||
- Disse poster er svære at matche automatisk og bør behandles særskilt
|
||||
|
||||
### Bonus/rabat-poster
|
||||
- Coop KG: "CCM Purchasing Rebate Credit"
|
||||
- Uniconta Rabat: sjælden, men findes
|
||||
- Disse bør matches separat via beløb + dato + bonus-tekst
|
||||
|
||||
---
|
||||
|
||||
## Specifikke problemer opdaget (V6.1)
|
||||
|
||||
1. **Faktura 46344** — Uniconta Faktura +441.720, Coop RE -441.720, men også Uniconta Betaling -430.732,50. Delvis kreditnota.
|
||||
|
||||
2. **Faktura 51519** — Uniconta Faktura +244.926, Coop RE -244.926, Uniconta Betaling -244.926. Dobbeltregistrering (Coop+Uni betaling).
|
||||
|
||||
3. **Faktura 48323** — Ingen Coop-betaling, men Uniconta Betaling -240.000 findes. Faktura er BETALT via Uniconta-betaling.
|
||||
|
||||
4. **Faktura 47140/47347 (DELVIST BETALT)** — Coop har både negativ og positiv RE med samme reference. Netto-sum giver korrekt status.
|
||||
|
||||
5. **Dobbelt-registrering** — Samme betaling findes som både Coop RE og Uniconta Betaling. Ved summing af begge får man 2x betalt. Løsning: brug netto-sum og dedupliker.
|
||||
|
||||
6. **MR8M-modregninger** — Positive Coop RE med "MR8M Credit memo" tekst er modregninger. De skal trækkes fra i netto-beregning, ikke lægges til med `abs()`.
|
||||
|
||||
---
|
||||
|
||||
## Afstemningsstrategi der virkede (V6.1 anbefalet)
|
||||
|
||||
**Fase 1: Reference-match (RE↔Faktura, RG↔Kreditnota)**
|
||||
- Krav: Coop.Reference == Uniconta.Faktura
|
||||
- Krav: abs(Coop.Amount) ≈ abs(Uniconta.Amount) ±5 kr (tolerance)
|
||||
- Krav: Dato ±21 dage
|
||||
- Bonus: Korrekt konteringstype
|
||||
- Score ~250
|
||||
- Resultat 2024-2025: ~2.421 matches
|
||||
|
||||
**Fase 2: ZV↔Faktura**
|
||||
- Coop ZV (positive fakturaer) matcher Uniconta Faktura
|
||||
- abs(beløb) ±5 kr + dato ±21 dage
|
||||
- Bonus: Document Number == Faktura
|
||||
- Resultat 2024-2025: ~173 matches
|
||||
|
||||
**Fase 3: RE↔Betaling uden reference**
|
||||
- Coop RE uden fakturanummer matcher Uniconta Betaling
|
||||
- abs(beløb) ±5 kr + dato ±21 dage
|
||||
- Bonus: tekst-lighed
|
||||
- Resultat 2024-2025: ~192 matches
|
||||
|
||||
**Fase 4: Generel abs(beløb)+dato**
|
||||
- abs(beløb) ±5 kr + dato ±21 dage
|
||||
- Konteringstype-bonus, tekst-lighed, reference-bonus
|
||||
- Score ≥ 40
|
||||
- Resultat 2024-2025: ~1.939 matches
|
||||
|
||||
**Fase 5: Mange-til-én**
|
||||
- Sum af 2-3 Uniconta-poster == Coop-postering ±5 kr
|
||||
- Dato ±14 dage, max 25 kandidater
|
||||
- Resultat 2024-2025: ~276 matches
|
||||
|
||||
**Total matches 2024-2025: 5.001 ud af 7.328 Coop / 6.717 Uniconta (~68%)**
|
||||
|
||||
**Uafstemt: ~2.327 Coop (31,8%) + ~2.211 Uniconta (32,9%)**
|
||||
|
||||
---
|
||||
|
||||
## Faktura-status: Netto-baseret + deduplikering
|
||||
|
||||
### Dobbelt-registrering
|
||||
Samme betaling findes som både Coop RE og Uniconta Betaling.
|
||||
|
||||
```python
|
||||
# Beregn NETTO Coop betalinger (ikke abs-sum!)
|
||||
coop_net = sum(c['amount'] for c in coop
|
||||
if c['reference'] == fnr and c['doc_type'] in ('RE','ZP'))
|
||||
coop_paid = abs(coop_net) if coop_net < 0 else 0
|
||||
|
||||
# Beregn NETTO Uniconta betalinger
|
||||
uni_net = sum(u['amount'] for u in uni
|
||||
if u['faktura'] == fnr and u['konteringstype'] == 'Betaling')
|
||||
uni_paid = abs(uni_net) if uni_net < 0 else 0
|
||||
|
||||
# Dedupliker
|
||||
if coop_paid > 0 and uni_paid > 0:
|
||||
if abs(coop_paid - uni_paid) <= AMT_TOL:
|
||||
betalt = coop_paid # Samme betaling
|
||||
else:
|
||||
betalt = max(coop_paid, uni_paid)
|
||||
elif coop_paid > 0: betalt = coop_paid
|
||||
elif uni_paid > 0: betalt = uni_paid
|
||||
else: betalt = 0
|
||||
```
|
||||
|
||||
**Resultat:** 2.968 BETALT, 2 DELVIST, 186 UBETALT
|
||||
|
||||
---
|
||||
|
||||
## Bevar original rådata
|
||||
|
||||
Exportér ALTID:
|
||||
- Original række-nummer (Excel-rækkeindeks)
|
||||
- Filnavn og ark-navn
|
||||
- Rå JSON med alle kolonner
|
||||
Så brugeren kan spore enhver match tilbage til kildefilen.
|
||||
+89
@@ -0,0 +1,89 @@
|
||||
# Coop-Uniconta import UI + Payment_advice lessons
|
||||
|
||||
Use this reference when building or fixing a Coop/Uniconta reconciliation webapp import flow.
|
||||
|
||||
## Required import sources
|
||||
|
||||
The user expects five uploadable source types, not a generic two-source UI:
|
||||
|
||||
1. **Uniconta Coop Debitor / betalingsfordeling**
|
||||
- Shows how Coop payments are allocated across invoices/credits.
|
||||
- Primary bridge from payment to invoice lines.
|
||||
2. **Uniconta/ERP fakturaliste — all Coop invoices for the year**
|
||||
- Complete invoice/credit-note list for the relevant year.
|
||||
- Primary source for "has Coop paid all invoices?".
|
||||
3. **Coop kontoudtog**
|
||||
- Account statement from Coop's system.
|
||||
4. **Bankkontoudtog — all Coop postings**
|
||||
- Bank statement filtered to Coop-related transactions.
|
||||
5. **Coop Payment_advice documents**
|
||||
- PDF and Excel files, uploaded continuously as they are collected.
|
||||
- Needed to identify and explain regulations, fees, bonuses, fines, credits, and offsets.
|
||||
|
||||
If a UI only shows "Coop" and "Uniconta", treat that as a workflow bug even if the backend has more enum choices.
|
||||
|
||||
## Fiscal year and monthly reporting
|
||||
|
||||
- Fiscal year is calendar year: 1/1–31/12.
|
||||
- Store/report month for every data source.
|
||||
- Keep invoice month/year separate from payment month/year.
|
||||
- Coop can have around 3 months of credit: December 2025 invoices may be paid in March 2026. Cross-year payment timing is expected and should be visible, not automatically flagged as an error.
|
||||
|
||||
## Payment_advice ingestion pattern
|
||||
|
||||
Implement Payment_advice as a real source early, even before final matching logic is known.
|
||||
|
||||
Recommended flexible model fields:
|
||||
|
||||
- source file
|
||||
- accounting year (nullable when no date can be inferred)
|
||||
- document date
|
||||
- payment date
|
||||
- document/bilag number
|
||||
- voucher number
|
||||
- invoice number
|
||||
- amount
|
||||
- classification: `unknown`, `payment`, `bonus`, `fee`, `fine`, `adjustment`, `credit`
|
||||
- raw text
|
||||
- raw JSON / raw row data
|
||||
- raw sheet name / row index
|
||||
|
||||
## Parser approach
|
||||
|
||||
For Excel:
|
||||
- Read all sheets.
|
||||
- Store raw row JSON.
|
||||
- Best-effort infer date from date-like columns (`Dato`, `Date`, `Betaling`, `Forfald`).
|
||||
- Best-effort infer amount from amount-like columns (`Beløb`, `Amount`, `Total`, `Netto`, `Moms`); if multiple amounts appear, pick the largest absolute amount as a temporary headline amount.
|
||||
- Extract invoice/voucher/document IDs from matching column names.
|
||||
|
||||
For PDF:
|
||||
- Store file even if text extraction is poor.
|
||||
- Try `pypdf` / `PyPDF2` text extraction if available.
|
||||
- Infer date and amount from filename + extracted text.
|
||||
- Mark whether text was extracted.
|
||||
|
||||
Do not reject Payment_advice uploads just because final parser rules are not complete. The key is preserving raw documents/data so later parser improvements can be applied.
|
||||
|
||||
## UI copy / workflow guidance
|
||||
|
||||
On the import page, list the five source types in the same order as the user's workflow. Explain that data is split into accounting years by date, and that Payment_advice can be uploaded gradually.
|
||||
|
||||
For monthly dashboard rows, include at least:
|
||||
|
||||
- Coop postings count/sum
|
||||
- Uniconta postings count/sum
|
||||
- Bank postings count/sum
|
||||
- Invoice list count/sum
|
||||
- Payment allocation count, debit, credit
|
||||
- Payment_advice count/sum
|
||||
- Payment_advice classifications per month
|
||||
|
||||
## Minimum regression tests
|
||||
|
||||
Add tests that verify:
|
||||
|
||||
1. The upload page renders all five source labels and accepts PDF for Payment_advice.
|
||||
2. Payment_advice Excel import preserves raw row data and derives accounting year/month.
|
||||
3. Payment_advice classification detects at least a bonus/regulation example.
|
||||
4. Existing invoice/payment allocation tests still pass after adding the new source.
|
||||
@@ -0,0 +1,236 @@
|
||||
# Coop ↔ Uniconta Match Engine Reference (Aktualiseret 2025-05-26)
|
||||
|
||||
Konkrete match-engine-mønstre fra iteration V1 → V6.1.
|
||||
|
||||
---
|
||||
|
||||
## Match-filosofi: Faser i prioriteret rækkefølge
|
||||
|
||||
Fakturanummer-match (Reference ↔ Faktura) er det STÆRKESTE signal.
|
||||
Beløb+dato er sekundært. Sekvens:
|
||||
|
||||
### V6.1 (anbefalet)
|
||||
|
||||
**Fase 1: Reference + abs(beløb) + dato**
|
||||
- Coop `Reference` == Uniconta `Faktura` (identiske 4-5 cifrede tal)
|
||||
- `abs(Coop.Amount) == abs(Uniconta.Amount)` ±5 kr (tolerance for afrunding)
|
||||
- Dato ±21 dage
|
||||
- Konteringstype-bonus (RE↔Faktura, RG↔Kreditnota, KG↔Kreditnota)
|
||||
- Score ~250
|
||||
- Resultat 2024-2025: ~2.421 matches
|
||||
|
||||
**Fase 2: ZV↔Faktura (Coop positive fakturaer)**
|
||||
- Coop `Document Type == ZV` (positive fakturaer udstedt til kunder)
|
||||
- Match mod Uniconta `Konteringstype == Faktura`
|
||||
- `abs(beløb)` ±5 kr + dato ±21 dage
|
||||
- Bonus: `Document Number == Faktura` nummer
|
||||
- Resultat 2024-2025: ~173 matches
|
||||
|
||||
**Fase 3: RE↔Betaling uden reference**
|
||||
- Coop `Document Type == RE` (betalinger) uden fakturanummer
|
||||
- Match mod Uniconta `Konteringstype == Betaling`
|
||||
- `abs(beløb)` ±5 kr + dato ±21 dage
|
||||
- Bonus: tekst-lighed ("coop", "0010014808")
|
||||
- Resultat 2024-2025: ~192 matches
|
||||
|
||||
**Fase 4: Generel abs(beløb) + dato (fallback)**
|
||||
- `abs(beløb)` matcher ±5 kr
|
||||
- Dato ±21 dage
|
||||
- Konteringstype-bonus, tekst-lighed, reference-bonus
|
||||
- Score ≥ 40
|
||||
- Resultat 2024-2025: ~1.939 matches
|
||||
|
||||
**Fase 5: Mange-til-én**
|
||||
- Sum af 2-3 Uniconta-poster == Coop-postering ±5 kr
|
||||
- Dato ±14 dage, max 25 kandidater
|
||||
- Resultat 2024-2025: ~276 matches
|
||||
|
||||
**Total matches 2024-2025: 5.001 ud af 7.328 Coop / 6.717 Uniconta (~68%)**
|
||||
|
||||
---
|
||||
|
||||
## Tolerance matching (V6 innovation)
|
||||
|
||||
I stedet for eksakt beløbs-match med ±0,01 kr, brug ±5 kr tolerance:
|
||||
|
||||
```python
|
||||
AMT_TOL = 5.0 # DKK tolerance
|
||||
|
||||
# Index-building: bucket per tolerance interval
|
||||
def build_index(transactions, max_days=21):
|
||||
idx = defaultdict(lambda: defaultdict(list))
|
||||
for t in transactions:
|
||||
for delta in range(-max_days, max_days + 1):
|
||||
d = t['date'] + timedelta(days=delta)
|
||||
bucket = round(t['abs_amount'] / AMT_TOL) * AMT_TOL
|
||||
idx[d][bucket].append(t)
|
||||
return idx
|
||||
|
||||
# Lookup: søg naboliggende buckets
|
||||
bucket = round(c['abs_amount'] / AMT_TOL) * AMT_TOL
|
||||
for b in [bucket - AMT_TOL, bucket, bucket + AMT_TOL]:
|
||||
candidates.extend(idx.get(c['date'], {}).get(b, []))
|
||||
```
|
||||
|
||||
**Fordel:** Fanger afrundingsforskelle og delvise betalinger som
|
||||
ellers ville være uafstemte.
|
||||
|
||||
**Bivirkning:** Kan skabe flere kandidater → brug score-threshold
|
||||
og konteringstype-filter til at vælge den bedste.
|
||||
|
||||
---
|
||||
|
||||
## Fortegn: MODSAT er normalt og korrekt
|
||||
|
||||
| Coop type | Fortegn | Uniconta type | Fortegn | Match-mode |
|
||||
|-----------|---------|---------------|---------|------------|
|
||||
| RE (betaling) | Negativ | Faktura | Positiv | abs(beløb) |
|
||||
| ZV (faktura) | Positiv | Betaling | Negativ | abs(beløb) |
|
||||
| RG (kreditnota) | Positiv | Kreditnota | Negativ | abs(beløb) |
|
||||
| KG (rabat) | Positiv | Kreditnota | Negativ | abs(beløb) |
|
||||
|
||||
```python
|
||||
# Match via abs(), men marker med fortegn-kolonne
|
||||
amt_match = abs(coop_amount - uniconta_amount) <= AMT_TOL
|
||||
fortegn = "SAMME" if (coop_amount > 0) == (uniconta_amount > 0) else "MODSAT"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Filtrering før match
|
||||
|
||||
**Afgræns til regnskabsår** i stedet for at filtrere på Doc_Type:
|
||||
- 2024+2025 fjerner 8.861 Coop og 5.684 Uniconta poster
|
||||
- Primært primo/XX/saldo-poster fra 2023 og 2026
|
||||
- Langt mere effektiv end Doc_Type-filtrering
|
||||
|
||||
| Filter | Kilde | Hvad | Konsekvens |
|
||||
|--------|-------|------|------------|
|
||||
| År 2024-2025 | Begge | Regnskabsår-afgrænsning | 56% færre poster |
|
||||
| `Primo` | Uniconta | Startsaldo | Fjernes ved års-filter |
|
||||
| `Afslutning` | Uniconta | Årsafslutning | Fjernes ved års-filter |
|
||||
|
||||
---
|
||||
|
||||
## Mange-til-én: Praktisk approach
|
||||
|
||||
Én Coop-postering kan svare til SUM af flere Uniconta-poster.
|
||||
|
||||
- **AVOID:** `itertools.combinations` over 7.328 × 6.717 rækker = O(n²) timeout
|
||||
- **USE:** Index Uniconta per `(dato, abs_beløb)` og søg kombinationer indenfor ±14 dage
|
||||
- **LIMIT:** Maks 3 Uniconta-poster per Coop-postering, max 25 kandidater
|
||||
- **TOLERANCE:** ±5 kr (ikke ±0,01 kr)
|
||||
- Resultat: ~276 matches i 2024-2025 — større effekt end V5.1
|
||||
|
||||
---
|
||||
|
||||
## Faktura-status: Netto-baseret + deduplikering
|
||||
|
||||
### Problemet: Dobbelt-registrering
|
||||
|
||||
Samme betaling kan findes som både:
|
||||
- **Coop RE** (betaling) med fakturanummer i Reference
|
||||
- **Uniconta Betaling** med fakturanummer i Faktura-kolonnen
|
||||
|
||||
Hvis man summerer begge får man **dobbelt så meget betalt** som faktisk.
|
||||
|
||||
### Løsning: Netto-sum per faktura
|
||||
|
||||
```python
|
||||
# 1. Beregn NETTO Coop betalinger per faktura
|
||||
# (positive RE trækker fra — det er modregninger)
|
||||
coop_net = sum(c['amount'] for c in coop
|
||||
if c['reference'] == fnr
|
||||
and c['doc_type'] in ('RE', 'ZP'))
|
||||
coop_paid = abs(coop_net) if coop_net < 0 else 0
|
||||
|
||||
# 2. Beregn NETTO Uniconta betalinger per faktora
|
||||
uni_net = sum(u['amount'] for u in uni
|
||||
if u['faktura'] == fnr
|
||||
and u['konteringstype'] == 'Betaling')
|
||||
uni_paid = abs(uni_net) if uni_net < 0 else 0
|
||||
|
||||
# 3. Dedupliker: hvis begge findes, brug kun én
|
||||
if coop_paid > 0 and uni_paid > 0:
|
||||
if abs(coop_paid - uni_paid) <= AMT_TOL:
|
||||
betalt = coop_paid # Samme betaling, tag én
|
||||
else:
|
||||
betalt = max(coop_paid, uni_paid) # Delvise, tag største
|
||||
```
|
||||
|
||||
### MR8M-modregninger
|
||||
|
||||
Nogle fakturaer har både negativ og positiv RE med samme reference:
|
||||
- Faktura 47140: RE -8.984 + RE +8.984 (MR8M modregning) = netto 0
|
||||
- Med `abs()` ville dette tælle som 17.968 betalt (FORKERT)
|
||||
- Med netto-sum: `abs(-8.984 + 8.984) = 0` (KORREKT)
|
||||
|
||||
---
|
||||
|
||||
## Score-formel (V6.1 konkret)
|
||||
|
||||
```
|
||||
# Fase 1 (Reference-match)
|
||||
Score = 250 - dato_diff * 4 - abs_amount_diff * 2
|
||||
|
||||
# Fase 2 (ZV↔Faktura)
|
||||
Score = 150 - dato_diff * 3 - abs_amount_diff * 2
|
||||
if doc_number == faktura: +50
|
||||
if tekst_identisk: +25
|
||||
if tekst_delvis: +12
|
||||
Threshold: ≥ 60
|
||||
|
||||
# Fase 3 (RE↔Betaling)
|
||||
Score = 120 - dato_diff * 3 - abs_amount_diff * 2
|
||||
if reference == faktura: +40
|
||||
if "coop" in tekst: +15
|
||||
Threshold: ≥ 60
|
||||
|
||||
# Fase 4 (Generel)
|
||||
Score = 100 - dato_diff * 3 - abs_amount_diff * 0.5
|
||||
if konteringstype_match: +25
|
||||
if tekst_identisk: +25
|
||||
if tekst_delvis: +12
|
||||
if reference_match: +40
|
||||
Threshold: ≥ 40
|
||||
|
||||
# Fase 5 (Mange-til-én)
|
||||
Score = 150 (fast, kræver manuel godkendelse)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Rapport-struktur (Excel-ark)
|
||||
|
||||
Rækkefølgen er vigtig — brugeren vil se opsummering først:
|
||||
|
||||
0. `0_Opsummering` — nøgletal, antal matches, totalbeløb, difference, faktura-status
|
||||
1. `1_Faktura_Status` — alle Uniconta fakturaer med BETALT/DELVIST/UBETALT
|
||||
2. `2_Matches` — alle matchede par med type, score, fortegn, beløbs-diff, dato-diff
|
||||
3. `3_Alle_Coop` — alle rå Coop-rækker med match-status kolonne
|
||||
4. `4_Alle_Uniconta` — alle rå Uniconta-rækker med match-status kolonne
|
||||
5. `5_Kun_Coop` — uafstemte Coop-poster
|
||||
6. `6_Kun_Uniconta` — uafstemte Uniconta-poster
|
||||
7. `7_Maaneds_Oversigt` — månedlig opsummering pr. år
|
||||
|
||||
**Nøgle:** Ark 3 og 4 bevarer overblik — brugeren vil se ALLE poster.
|
||||
Ark 5+6 viser kun dem der stadig mangler.
|
||||
|
||||
---
|
||||
|
||||
## Bevar original rådata
|
||||
|
||||
Exportér ALTID:
|
||||
- Original række-nummer (Excel-rækkeindeks)
|
||||
- Filnavn og ark-navn
|
||||
- Rå JSON med alle kolonner
|
||||
Så brugeren kan spore enhver match tilbage til kildefilen.
|
||||
|
||||
---
|
||||
|
||||
## Performance-tips
|
||||
|
||||
- V6.1 kører på ~10 sekunder for 7.328 Coop + 6.717 Uniconta poster
|
||||
- Brug `defaultdict(lambda: defaultdict(list))` til dato→beløb→post indeks
|
||||
- Uniconta parsing: `pd.read_excel(file, header=3)` for at læse fra række 4
|
||||
- Python 3.9 med pandas 2.3.3 — set `PYTHONPATH` hvis user-site-packages
|
||||
+122
@@ -0,0 +1,122 @@
|
||||
# Django Accounting UI Formatting
|
||||
|
||||
Former skill: `django-accounting-ui-formatting`.
|
||||
|
||||
Use this when changing how money, balances, invoices, payments, totals, or differences are displayed in a Django accounting/reconciliation webapp.
|
||||
|
||||
## Core rule
|
||||
|
||||
Accounting amounts must be formatted deliberately and consistently. Do not rely on ad-hoc `floatformat:2` everywhere if the app has locale-specific expectations.
|
||||
|
||||
For Danish Radix accounting/reconciliation UIs, render amounts with:
|
||||
|
||||
- thousands separator: `.`
|
||||
- decimal separator: `,`
|
||||
- exactly 2 decimals
|
||||
- minus sign before the formatted number
|
||||
|
||||
Example:
|
||||
|
||||
```text
|
||||
1234567.89 -> 1.234.567,89
|
||||
-1234567.89 -> -1.234.567,89
|
||||
0 -> 0,00
|
||||
```
|
||||
|
||||
## Recommended Django implementation
|
||||
|
||||
Create a shared template filter in a common app, e.g.:
|
||||
|
||||
```text
|
||||
backend/apps/common/templatetags/amount_format.py
|
||||
```
|
||||
|
||||
Example filter:
|
||||
|
||||
```python
|
||||
from decimal import Decimal, InvalidOperation, ROUND_HALF_UP
|
||||
|
||||
from django import template
|
||||
|
||||
register = template.Library()
|
||||
|
||||
|
||||
@register.filter
|
||||
def amount_dk(value):
|
||||
if value is None or value == "":
|
||||
return "0,00"
|
||||
|
||||
try:
|
||||
amount = Decimal(str(value)).quantize(Decimal("0.01"), rounding=ROUND_HALF_UP)
|
||||
except (InvalidOperation, ValueError, TypeError):
|
||||
return value
|
||||
|
||||
sign = "-" if amount < 0 else ""
|
||||
amount = abs(amount)
|
||||
formatted = f"{amount:,.2f}"
|
||||
return sign + formatted.replace(",", "X").replace(".", ",").replace("X", ".")
|
||||
```
|
||||
|
||||
Ensure `templatetags/__init__.py` exists.
|
||||
|
||||
In templates:
|
||||
|
||||
```django
|
||||
{% load amount_format %}
|
||||
{{ amount|amount_dk }} kr
|
||||
```
|
||||
|
||||
## What to update
|
||||
|
||||
Search all relevant templates for amount displays, especially:
|
||||
|
||||
- `floatformat:2`
|
||||
- table columns named Beløb, Sum, Difference, Rest, Betalt
|
||||
- dashboard cards
|
||||
- report tables
|
||||
- invoice/payment status pages
|
||||
- transaction detail pages
|
||||
|
||||
Replace monetary `floatformat:2` with the shared amount filter. Leave percentages, scores, counts, and raw JSON/raw imported data alone unless the user explicitly asks otherwise.
|
||||
|
||||
## Tests
|
||||
|
||||
Follow TDD for behavior changes:
|
||||
|
||||
1. Add a failing unit/template test for the filter:
|
||||
|
||||
```python
|
||||
template = Template("{% load amount_format %}{{ value|amount_dk }}")
|
||||
assert template.render(Context({"value": Decimal("1234567.89")})) == "1.234.567,89"
|
||||
```
|
||||
|
||||
2. Add at least one rendered-page regression test for the main UI path, asserting:
|
||||
|
||||
- formatted amount is present, e.g. `1.234.567,89`
|
||||
- raw unformatted amount is absent, e.g. `1234567.89`
|
||||
|
||||
3. Run the relevant app tests and system check in the project runtime, usually Docker Compose for Radix internal apps:
|
||||
|
||||
```bash
|
||||
docker compose exec -T web python manage.py test apps.core --settings=config.settings.test -v 2
|
||||
docker compose exec -T web python manage.py check
|
||||
```
|
||||
|
||||
## Verification
|
||||
|
||||
Before finalizing:
|
||||
|
||||
- Confirm there are no remaining monetary `|floatformat:2` usages in the changed templates.
|
||||
- Confirm templates that use `amount_dk` include `{% load amount_format %}`.
|
||||
- Restart the web container if the running app needs to load a new template tag module:
|
||||
|
||||
```bash
|
||||
docker compose restart web
|
||||
```
|
||||
|
||||
## Pitfalls
|
||||
|
||||
- Do not alter raw JSON/raw import displays; those are audit data and should preserve original values.
|
||||
- Do not apply money formatting to percentages or match scores.
|
||||
- Avoid Python float math for money formatting; use `Decimal` to avoid binary rounding surprises.
|
||||
- If the app already has a localization/formatting utility, extend that instead of adding a second competing filter.
|
||||
+122
@@ -0,0 +1,122 @@
|
||||
# Idempotent SourceFile imports for accounting/reconciliation apps
|
||||
|
||||
Use this when users can upload the same ERP/accounting source files repeatedly, especially bulk settlement files such as Coop `Payment_advice` PDFs/XLSX.
|
||||
|
||||
## Problem
|
||||
|
||||
Financial import apps often store one `SourceFile` plus many parsed rows/documents. If a user uploads the same file again, the app must not create:
|
||||
- duplicate `SourceFile` records
|
||||
- duplicate parsed settlement/payment/invoice lines
|
||||
- duplicated totals in reconciliation reports
|
||||
|
||||
If the same named file has changed, treat it as an update of that source, not as a new independent source file.
|
||||
|
||||
## Durable pattern
|
||||
|
||||
Add a content hash to the source-file model:
|
||||
|
||||
```python
|
||||
class SourceFile(models.Model):
|
||||
source = models.CharField(max_length=64)
|
||||
file_name = models.CharField(max_length=255)
|
||||
file = models.FileField(upload_to="imports/")
|
||||
content_hash = models.CharField(max_length=64, blank=True, db_index=True)
|
||||
```
|
||||
|
||||
Compute SHA256 from the uploaded bytes before creating a new `SourceFile`. Always rewind the file afterwards:
|
||||
|
||||
```python
|
||||
import hashlib
|
||||
|
||||
def uploaded_file_sha256(uploaded_file):
|
||||
h = hashlib.sha256()
|
||||
for chunk in uploaded_file.chunks():
|
||||
h.update(chunk)
|
||||
uploaded_file.seek(0)
|
||||
return h.hexdigest()
|
||||
```
|
||||
|
||||
Decision table:
|
||||
|
||||
| Existing row? | Same file name? | Same hash? | Action |
|
||||
|---|---:|---:|---|
|
||||
| no | n/a | n/a | create new `SourceFile`, parse rows |
|
||||
| yes | yes | yes | skip import; report as unchanged/skipped |
|
||||
| yes | yes | no | update existing `SourceFile`, delete old parsed child rows, reparse |
|
||||
| yes | no | yes | usually skip as duplicate content; optionally report duplicate content under different name |
|
||||
| yes | no | no | create new source file |
|
||||
|
||||
For source types that represent operational settlement documents, prefer matching on `(source, file_name)` for update semantics, plus `content_hash` for duplicate detection.
|
||||
|
||||
## Transaction boundary
|
||||
|
||||
Wrap the parser/import in a transaction so child rows are not partially replaced:
|
||||
|
||||
```python
|
||||
from django.db import transaction
|
||||
|
||||
@transaction.atomic
|
||||
def import_payment_advice(source_file):
|
||||
PaymentAdviceDocument.objects.filter(source_file=source_file).delete()
|
||||
# parse and recreate child rows/documents
|
||||
```
|
||||
|
||||
If the upload flow creates a new file and parsing fails, delete only that newly-created `SourceFile`. Do not delete/revert an existing source record unless you have an explicit rollback strategy.
|
||||
|
||||
## UI behaviour
|
||||
|
||||
Return counters that distinguish:
|
||||
- imported files
|
||||
- updated files
|
||||
- skipped unchanged files
|
||||
- failed files
|
||||
|
||||
Use user-facing messages like:
|
||||
- `1 uændret fil sprunget over`
|
||||
- `1 fil opdateret`
|
||||
- `23 filer importeret`
|
||||
|
||||
This matters for accounting users: they need confidence that re-uploading a folder does not inflate totals.
|
||||
|
||||
## Bulk upload settings
|
||||
|
||||
Django defaults may reject large batches with `TooManyFilesSent`. For workflows where users upload many small settlement documents, make the limit explicit in settings/env, for example:
|
||||
|
||||
```python
|
||||
DATA_UPLOAD_MAX_NUMBER_FILES = int(os.getenv("DATA_UPLOAD_MAX_NUMBER_FILES", "2000"))
|
||||
```
|
||||
|
||||
Add the setting to `.env.example` and document that it controls maximum files per HTTP upload.
|
||||
|
||||
## Tests to add
|
||||
|
||||
Minimum regression tests:
|
||||
|
||||
1. Upload same file twice with identical bytes:
|
||||
- `SourceFile.objects.count()` unchanged after second upload
|
||||
- child document/line count unchanged
|
||||
- response contains skipped/unchanged message
|
||||
|
||||
2. Upload same file name with changed bytes:
|
||||
- same `SourceFile` primary key reused
|
||||
- `content_hash` changes
|
||||
- old child rows are deleted/replaced, not appended
|
||||
- response contains updated message
|
||||
|
||||
3. Bulk upload many files if the UI supports multi-file import:
|
||||
- no `TooManyFilesSent`
|
||||
- all valid new files imported
|
||||
- repeated files skipped
|
||||
|
||||
4. Parser failure path:
|
||||
- failed new file does not leave orphan `SourceFile`
|
||||
- failed update does not silently erase previous valid data
|
||||
|
||||
## Documentation checklist
|
||||
|
||||
Update README/import docs with:
|
||||
- duplicate uploads are skipped by content hash
|
||||
- same file name with changed content updates the existing source
|
||||
- updates replace parsed rows for that source file
|
||||
- original source files are not edited in place; imported copies live under media/storage
|
||||
- bulk upload file-count limit and env var
|
||||
+167
@@ -0,0 +1,167 @@
|
||||
# Reconciliation Webapp Architecture (Django)
|
||||
|
||||
Reference architecture for building a webapp on top of a reconciliation engine.
|
||||
Based on Coop-Uniconta project (Radix), May 2026.
|
||||
|
||||
## Stack
|
||||
|
||||
- **Backend**: Django 4.2 + Django REST Framework
|
||||
- **Database**: MariaDB 11.4
|
||||
- **Cache/Queue**: Redis + Celery (for background matching)
|
||||
- **Frontend**: Django Templates + vanilla JS (no React needed for internal tools)
|
||||
- **Deployment**: Docker Compose (web, db, redis)
|
||||
- **Auth**: Django built-in + admin, later MS Entra/OIDC
|
||||
|
||||
## Project Structure
|
||||
|
||||
```
|
||||
backend/
|
||||
config/
|
||||
settings/base.py # Shared config
|
||||
settings/local.py # Dev overrides (DEBUG=True, admin enabled)
|
||||
urls.py # URL routing
|
||||
wsgi.py / asgi.py # Entry points
|
||||
celery.py # Celery app
|
||||
apps/
|
||||
core/ # Models: Transaction, MatchResult, InvoiceStatus, AuditLog, BonusRule
|
||||
models.py
|
||||
admin.py
|
||||
matching.py # 5-phase match engine
|
||||
import_/ # Excel parsers
|
||||
parsers.py # Coop + Uniconta import with preview
|
||||
reconciliation/ # Views: dashboard, import, matching, reports
|
||||
views.py
|
||||
templates/
|
||||
templates/
|
||||
base.html # Dark-themed layout (sidebar nav + main content)
|
||||
static/
|
||||
css/app.css # Dark theme: --bg-primary: #0f172a, --accent: #3b82f6
|
||||
manage.py
|
||||
entrypoint.sh # wait-for-db + migrate + runserver
|
||||
Dockerfile # python:3.12-slim
|
||||
requirements.txt
|
||||
```
|
||||
|
||||
## Data Model
|
||||
|
||||
```
|
||||
AccountingYear (year, start_date, end_date)
|
||||
SourceFile (source, file_name, file_path, row_count, parsed_count, error_count, accounting_year)
|
||||
Transaction (source_file, source, year, month, date, amount, abs_amount,
|
||||
doc_type, reference, faktura, text, text_normalized,
|
||||
konteringstype, bilag, doc_number,
|
||||
match_status, match_type, match_score, match_group,
|
||||
original_data JSON)
|
||||
MatchResult (group_id, match_type, score, status,
|
||||
transaction_a, transaction_b, additional_b_ids JSON,
|
||||
amount_diff, date_diff_days, explanation, comment,
|
||||
approved_by, approved_at)
|
||||
InvoiceStatus (faktura_number, accounting_year, invoice_total, paid_total, remaining,
|
||||
status, invoice_date, first_invoice_text,
|
||||
coop_payment_count, uniconta_payment_count, invoice_count)
|
||||
AuditLog (action, user, match_result, details JSON)
|
||||
BonusRule (name, year, period_start, period_end, conditions JSON, percentage, min_amount)
|
||||
BonusCalculation (rule, transaction, calculated_amount, status)
|
||||
```
|
||||
|
||||
## Key Design Decisions
|
||||
|
||||
1. **Store original_data as JSON** — Always preserve raw Excel row data for traceability.
|
||||
2. **Separate `amount` and `abs_amount`** — `amount` keeps original sign for net calculations; `abs_amount` for matching.
|
||||
3. **MatchResult captures both sides** — Primary match (transaction_a → transaction_b) + optional additional_b_ids for many-to-one.
|
||||
4. **InvoiceStatus is computed, not stored per-transaction** — Recalculated after each matching run.
|
||||
5. **AuditLog for every manual action** — Match approval, rejection, bonus adjustment.
|
||||
|
||||
## Match Engine Integration
|
||||
|
||||
The match engine runs as a Celery task triggered from the web UI:
|
||||
|
||||
```python
|
||||
# apps/core/matching.py
|
||||
AMT_TOL = Decimal("5.00")
|
||||
DATE_TOL_REF = 21
|
||||
DATE_TOL_GEN = 21
|
||||
|
||||
def run_matching(accounting_year=None):
|
||||
# Phase 1: Reference match (RE↔Faktura, RG↔Kreditnota)
|
||||
# Phase 2: ZV↔Faktura
|
||||
# Phase 3: RE↔Betaling without reference
|
||||
# Phase 4: General abs(amount)+date
|
||||
# Phase 5: Many-to-one (sum of 2-3 Uni = 1 Coop)
|
||||
# Bulk create MatchResult, bulk update Transaction statuses
|
||||
```
|
||||
|
||||
## UI Pages
|
||||
|
||||
| Page | Purpose |
|
||||
|------|---------|
|
||||
| Dashboard | Year selector, stat cards, invoice status summary, monthly overview |
|
||||
| Import | Upload Excel, show preview, map columns if auto-detection fails |
|
||||
| Reconciliation List | Filterable table of all transactions, paginated (50/page) |
|
||||
| Transaction Detail | Raw data, normalized data, potential matches, match history |
|
||||
| Faktura Status | All invoices with BETALT/DELVIST/UBETALT, filterable |
|
||||
| Reports | By type, by month, match summary, export to Excel |
|
||||
| Run Match | Trigger background matching for selected year |
|
||||
|
||||
## Dark Theme CSS Variables
|
||||
|
||||
```css
|
||||
:root {
|
||||
--bg-primary: #0f172a;
|
||||
--bg-secondary: #1e293b;
|
||||
--bg-tertiary: #334155;
|
||||
--text-primary: #f8fafc;
|
||||
--text-secondary: #94a3b8;
|
||||
--accent: #3b82f6;
|
||||
--success: #22c55e;
|
||||
--warning: #f59e0b;
|
||||
--danger: #ef4444;
|
||||
--border: #334155;
|
||||
--radius: 8px;
|
||||
}
|
||||
```
|
||||
|
||||
Inspired by Radix-ERP visual style: sidebar navigation, card-based stats,
|
||||
data tables with badges, filter bars above tables.
|
||||
|
||||
## Docker Compose
|
||||
|
||||
```yaml
|
||||
services:
|
||||
web:
|
||||
build: ./backend
|
||||
command: ["web-dev"]
|
||||
ports: ["8000:8000"]
|
||||
depends_on:
|
||||
db: {condition: service_healthy}
|
||||
env_file: [.env]
|
||||
db:
|
||||
image: mariadb:11.4
|
||||
environment:
|
||||
MARIADB_DATABASE: ${DB_NAME}
|
||||
MARIADB_USER: ${DB_USER}
|
||||
MARIADB_PASSWORD: ${DB_PASSWORD}
|
||||
MARIADB_ROOT_PASSWORD: ${DB_ROOT_PASSWORD}
|
||||
healthcheck:
|
||||
test: ["CMD", "healthcheck.sh", "--connect", "--innodb_initialized"]
|
||||
redis:
|
||||
image: redis:7-alpine
|
||||
```
|
||||
|
||||
## Deployment Notes
|
||||
|
||||
- `python manage.py migrate` runs on container startup via entrypoint.sh
|
||||
- `python manage.py createsuperuser` for first admin login
|
||||
- Uploads stored in MEDIA_ROOT (mounted volume in production)
|
||||
- Static files collected via `collectstatic` for production (nginx)
|
||||
|
||||
## Next Steps for a New Reconciliation Project
|
||||
|
||||
1. Copy project structure from template
|
||||
2. Adapt parsers for the specific Excel formats
|
||||
3. Run analysis script to determine match strategy
|
||||
4. Configure AMT_TOL and DATE_TOL based on data quality
|
||||
5. Build import UI with column mapping fallback
|
||||
6. Implement match engine phases iteratively
|
||||
7. Add manual match/approve/reject with audit log
|
||||
8. Add reports and bonus calculation framework
|
||||
+146
@@ -0,0 +1,146 @@
|
||||
# Three-Way Django Analysis Engine — Coop ↔ Uniconta ↔ Bank
|
||||
|
||||
Session-specific implementation notes for building a maintainable three-source reconciliation layer in a Django-based internal finance app.
|
||||
|
||||
## When to use this pattern
|
||||
|
||||
Use when the task has three sources with different semantic roles:
|
||||
|
||||
- ERP/accounting ledger, e.g. Uniconta = invoiced/booked postings
|
||||
- Counterparty/customer statement, e.g. Coop kontoudtog = settlement/detail source
|
||||
- Bank CSV = actual cash movement / reality check
|
||||
|
||||
Do **not** treat the result as a legal claim automatically. The engine should produce review flags and explainable evidence.
|
||||
|
||||
## Recommended architecture
|
||||
|
||||
Keep the analysis engine separate from import parsing and HTTP views:
|
||||
|
||||
```text
|
||||
apps/core/parsers.py # source-specific import/normalization
|
||||
apps/core/three_way.py # pure-ish analysis functions
|
||||
apps/core/management/commands/*.py # import/analyze CLI commands
|
||||
apps/reconciliation/views.py # dashboard/report presentation only
|
||||
apps/reconciliation/templates/... # HTML tables/cards
|
||||
```
|
||||
|
||||
Good public functions in the analysis module:
|
||||
|
||||
```python
|
||||
analyze_three_way_year(year)
|
||||
build_monthly_summary(year)
|
||||
build_invoice_statuses(year)
|
||||
find_unexplained_bank_payments(year)
|
||||
build_bonus_adjustment_summary(year)
|
||||
```
|
||||
|
||||
Views should call these functions and render results; avoid embedding reconciliation logic in templates or views.
|
||||
|
||||
## Batch import command
|
||||
|
||||
For folder-based source projects, add a deterministic command that imports all configured source directories:
|
||||
|
||||
```bash
|
||||
python manage.py import_all_sources --base-dir /app/source_data
|
||||
```
|
||||
|
||||
Expected source layout:
|
||||
|
||||
```text
|
||||
Fra Coop/
|
||||
Fra Uniconta/
|
||||
Fra bank/
|
||||
```
|
||||
|
||||
The command should:
|
||||
|
||||
1. Walk each known directory.
|
||||
2. Infer source from directory, not filename alone.
|
||||
3. Preserve original file, sheet and row identifiers.
|
||||
4. Store row-level import errors.
|
||||
5. Print counts per file and total imported/error counts.
|
||||
6. Be safe to rerun or clearly document duplicate behavior.
|
||||
|
||||
## Three-way reporting outputs
|
||||
|
||||
A useful first report page should show:
|
||||
|
||||
- possible amounts the counterparty may owe (`Mulige beløb Coop skylder`)
|
||||
- unexplained bank postings (`Uforklarede bankposter`)
|
||||
- invoice statuses analyzed
|
||||
- months with material differences
|
||||
- month-level totals across all three sources
|
||||
- bonus/credit/offset summary by period
|
||||
|
||||
For each proposed issue, include evidence:
|
||||
|
||||
- amount difference
|
||||
- relevant date range
|
||||
- invoice/reference, if present
|
||||
- source counts
|
||||
- which fields matched and which did not
|
||||
- score/status and whether manual review is required
|
||||
|
||||
## Interpretation pitfall
|
||||
|
||||
Large totals such as “possible amount owed” are **analysis flags**, not final claims.
|
||||
|
||||
Reasons:
|
||||
|
||||
- One bank payment can cover many Coop/Uniconta postings.
|
||||
- One Coop settlement can cover many invoices.
|
||||
- Dates may differ between invoice date, statement date, settlement date and bank date.
|
||||
- Bonus/credit notes may be non-cash accruals or period-based offsets.
|
||||
- Open-post lists are status snapshots and should be treated separately from ledger/history.
|
||||
|
||||
The UI and README should label these as review candidates until group matching and manual approval confirm them.
|
||||
|
||||
## Matching strategy after basic import
|
||||
|
||||
Start with conservative layers:
|
||||
|
||||
1. Invoice/reference + absolute amount + date tolerance.
|
||||
2. Remaining amount + date tolerance.
|
||||
3. Month-level totals across all three sources.
|
||||
4. Many-to-one and one-to-many group matching.
|
||||
5. Bonus/credit period reconciliation.
|
||||
6. Manual approval/rejection with audit log.
|
||||
|
||||
For bank matching, prioritize group matching:
|
||||
|
||||
```text
|
||||
bank payment = sum of many Coop settlement lines = sum of many Uniconta invoices
|
||||
```
|
||||
|
||||
Use the bank as cash-flow validation, not necessarily as a row-level join target.
|
||||
|
||||
## Tests to include
|
||||
|
||||
Minimum tests for this class of feature:
|
||||
|
||||
- invoice is paid when Uniconta, Coop and bank agree
|
||||
- invoice is flagged when Coop/bank evidence indicates a residual amount
|
||||
- bank payment without Coop/Uniconta support is flagged as unexplained
|
||||
- monthly summary calculates Bank-Coop and Bank-Uniconta differences correctly
|
||||
|
||||
For Django, add a lightweight test settings module if the normal dev database requires services/permissions:
|
||||
|
||||
```text
|
||||
config/settings/test.py # SQLite, fast local tests
|
||||
```
|
||||
|
||||
Then run:
|
||||
|
||||
```bash
|
||||
DJANGO_SETTINGS_MODULE=config.settings.test python manage.py test
|
||||
```
|
||||
|
||||
## Documentation to update
|
||||
|
||||
When this pattern is implemented, update:
|
||||
|
||||
- README.md: installation, Docker, import flow, analysis commands, limitations
|
||||
- ANALYSE.md or ANALYSE_TREVEJS.md: source findings and interpretation
|
||||
- HANDOFF.md: current status and next recommended steps
|
||||
- TODO.md/ROADMAP.md: group matching, manual review, bonus expansion
|
||||
|
||||
@@ -0,0 +1,211 @@
|
||||
"""
|
||||
Generic Excel file analyzer for ERP reconciliation projects.
|
||||
|
||||
Usage:
|
||||
python analyze_excel.py <directory> [--output analysis.json]
|
||||
|
||||
This script:
|
||||
1. Scans a directory for .xlsx files
|
||||
2. Identifies sheets, headers, data ranges
|
||||
3. Detects Danish number formats
|
||||
4. Prints null counts and sample values
|
||||
5. Saves structured JSON for downstream processing
|
||||
|
||||
To customize for a project, adjust:
|
||||
- EXPECTED_COLS: columns that must exist for a row to be considered a header
|
||||
- AMOUNT_COL_PATTERNS: regex patterns for amount columns
|
||||
- DATE_COL_PATTERNS: regex patterns for date columns
|
||||
"""
|
||||
import pandas as pd
|
||||
import openpyxl
|
||||
import os
|
||||
import json
|
||||
import re
|
||||
from collections import OrderedDict
|
||||
|
||||
# --- CUSTOMIZE PER PROJECT ---
|
||||
EXPECTED_COLS = {'Dato', 'Faktura', 'Bilag', 'Tekst', 'Beløb'}
|
||||
AMOUNT_COL_PATTERNS = [
|
||||
r'(?i)bel[oø]b',
|
||||
r'(?i)amount',
|
||||
r'(?i)sum',
|
||||
r'(?i)rest',
|
||||
r'(?i)forfalden',
|
||||
]
|
||||
DATE_COL_PATTERNS = [
|
||||
r'(?i)dato',
|
||||
r'(?i)date',
|
||||
r'(?i)posting',
|
||||
r'(?i)entry',
|
||||
]
|
||||
REFERENCE_COL_PATTERNS = [
|
||||
r'(?i)reference',
|
||||
r'(?i)faktura',
|
||||
r'(?i)bilag',
|
||||
r'(?i)document number',
|
||||
r'(?i)invoice',
|
||||
]
|
||||
# --- END CUSTOMIZE ---
|
||||
|
||||
|
||||
def classify_column(col_name):
|
||||
"""Guess the semantic type of a column from its name."""
|
||||
cn = str(col_name)
|
||||
for p in AMOUNT_COL_PATTERNS:
|
||||
if re.search(p, cn):
|
||||
return 'amount'
|
||||
for p in DATE_COL_PATTERNS:
|
||||
if re.search(p, cn):
|
||||
return 'date'
|
||||
for p in REFERENCE_COL_PATTERNS:
|
||||
if re.search(p, cn):
|
||||
return 'reference'
|
||||
return 'other'
|
||||
|
||||
|
||||
def find_header_row(filepath, sheet=0, required_cols=None):
|
||||
"""Find the first row that looks like a header."""
|
||||
if required_cols is None:
|
||||
required_cols = EXPECTED_COLS
|
||||
wb = openpyxl.load_workbook(filepath, data_only=True)
|
||||
ws = wb[wb.sheetnames[sheet]]
|
||||
for i, row in enumerate(ws.iter_rows(values_only=True), start=0):
|
||||
vals = set(str(v) for v in row if v is not None)
|
||||
if required_cols.issubset(vals):
|
||||
return i
|
||||
return None
|
||||
|
||||
|
||||
def parse_danish_number(val):
|
||||
"""Normalize Danish text numbers like '-2.543.803,93' to float."""
|
||||
if pd.isna(val):
|
||||
return float('nan')
|
||||
if isinstance(val, (int, float)):
|
||||
return float(val)
|
||||
s = str(val).strip()
|
||||
if not s:
|
||||
return float('nan')
|
||||
# Danish format: thousands sep = '.', decimal = ','
|
||||
if '.' in s and ',' in s:
|
||||
s = s.replace('.', '').replace(',', '.')
|
||||
elif ',' in s:
|
||||
s = s.replace(',', '.')
|
||||
try:
|
||||
return float(s)
|
||||
except ValueError:
|
||||
return float('nan')
|
||||
|
||||
|
||||
def analyze_file(filepath):
|
||||
info = {
|
||||
"filepath": filepath,
|
||||
"filename": os.path.basename(filepath),
|
||||
"size_kb": round(os.path.getsize(filepath) / 1024, 1),
|
||||
"sheets": OrderedDict(),
|
||||
}
|
||||
xl = pd.ExcelFile(filepath)
|
||||
for sheet in xl.sheet_names:
|
||||
info["sheets"][sheet] = analyze_sheet(filepath, sheet)
|
||||
return info
|
||||
|
||||
|
||||
def analyze_sheet(filepath, sheet):
|
||||
# Try default header=0 first; if columns look wrong, try header detection
|
||||
df = pd.read_excel(filepath, sheet_name=sheet, engine='openpyxl')
|
||||
header_row = None
|
||||
# Simple heuristic: if ALL columns are unnamed, try detecting header
|
||||
unnamed = [c for c in df.columns if 'Unnamed' in str(c)]
|
||||
if len(unnamed) > len(df.columns) * 0.5:
|
||||
hr = find_header_row(filepath, sheet)
|
||||
if hr is not None:
|
||||
df = pd.read_excel(filepath, sheet_name=sheet, engine='openpyxl', header=hr)
|
||||
header_row = hr
|
||||
|
||||
cols = OrderedDict()
|
||||
for col in df.columns:
|
||||
col_type = classify_column(str(col))
|
||||
is_date = pd.api.types.is_datetime64_any_dtype(df[col])
|
||||
is_num = pd.api.types.is_numeric_dtype(df[col])
|
||||
null_count = int(df[col].isna().sum())
|
||||
unique_count = int(df[col].nunique())
|
||||
samples = []
|
||||
for v in df[col].dropna().head(5):
|
||||
if len(samples) >= 3:
|
||||
break
|
||||
try:
|
||||
if is_date:
|
||||
samples.append(str(pd.Timestamp(v).date()))
|
||||
else:
|
||||
samples.append(str(v)[:60])
|
||||
except Exception:
|
||||
samples.append(str(v)[:60])
|
||||
|
||||
# Detect danish number format
|
||||
danish_flag = False
|
||||
if is_num:
|
||||
for v in df[col].dropna().head(10):
|
||||
if isinstance(v, str) and (',' in v or ('.' in v and ',' in v.replace('.', ''))):
|
||||
danish_flag = True
|
||||
break
|
||||
|
||||
cols[str(col)] = {
|
||||
"semantic_type": col_type,
|
||||
"pandas_dtype": str(df[col].dtype),
|
||||
"null_count": null_count,
|
||||
"null_pct": round(null_count / len(df) * 100, 1),
|
||||
"unique_count": unique_count,
|
||||
"is_date": bool(is_date),
|
||||
"is_numeric": bool(is_num),
|
||||
"danish_number_format": danish_flag,
|
||||
"sample_values": samples,
|
||||
}
|
||||
|
||||
empty_rows = int((df.isna().all(axis=1)).sum())
|
||||
dupes = int(df.duplicated().sum())
|
||||
|
||||
# Try to find date range
|
||||
date_range = None
|
||||
for col, meta in cols.items():
|
||||
if meta["is_date"]:
|
||||
non_null = df[col].dropna()
|
||||
if len(non_null) > 0:
|
||||
date_range = {
|
||||
"min": str(non_null.min().date()),
|
||||
"max": str(non_null.max().date()),
|
||||
"column": col,
|
||||
}
|
||||
break
|
||||
|
||||
return {
|
||||
"rows": len(df),
|
||||
"columns": len(df.columns),
|
||||
"detected_header_row": header_row,
|
||||
"empty_rows": empty_rows,
|
||||
"duplicate_rows": dupes,
|
||||
"date_range": date_range,
|
||||
"columns": cols,
|
||||
}
|
||||
|
||||
|
||||
def main():
|
||||
import argparse
|
||||
parser = argparse.ArgumentParser(description="Analyze Excel files for data reconciliation")
|
||||
parser.add_argument("directory", help="Directory containing .xlsx files")
|
||||
parser.add_argument("--output", default="excel_analysis.json", help="Output JSON file")
|
||||
args = parser.parse_args()
|
||||
|
||||
results = OrderedDict()
|
||||
for f in sorted(os.listdir(args.directory)):
|
||||
if f.endswith('.xlsx') and not f.startswith('~'):
|
||||
fp = os.path.join(args.directory, f)
|
||||
key = os.path.splitext(f)[0]
|
||||
results[key] = analyze_file(fp)
|
||||
print(f"Analyzed: {f}")
|
||||
|
||||
with open(args.output, 'w', encoding='utf-8') as fh:
|
||||
json.dump(results, fh, indent=2, ensure_ascii=False)
|
||||
print(f"\nSaved: {args.output}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user