Files
dennis 236053cd7f Initial import: 10 Radix skills across 3 categories
Kategoriserer danske regnskabs-/ERP-skills, ERP-data-afstemnings-
skill og to devops-workflows (gitea-issue-agent, kanban-workflows)
som Radix-medlemmers AI-agenter kan dele.

Struktur:
  skills/<kategori>/<skill-navn>/SKILL.md
  skills/<kategori>/<skill-navn>/{references,templates,scripts}/

Indekseret af index.json (genereret af scripts/rebuild_index.py).
Kør `python3 scripts/rebuild_index.py --check` i CI for at fange
synkroniseringsfejl.
2026-06-15 14:41:04 +02:00

123 lines
4.2 KiB
Markdown

# Idempotent SourceFile imports for accounting/reconciliation apps
Use this when users can upload the same ERP/accounting source files repeatedly, especially bulk settlement files such as Coop `Payment_advice` PDFs/XLSX.
## Problem
Financial import apps often store one `SourceFile` plus many parsed rows/documents. If a user uploads the same file again, the app must not create:
- duplicate `SourceFile` records
- duplicate parsed settlement/payment/invoice lines
- duplicated totals in reconciliation reports
If the same named file has changed, treat it as an update of that source, not as a new independent source file.
## Durable pattern
Add a content hash to the source-file model:
```python
class SourceFile(models.Model):
source = models.CharField(max_length=64)
file_name = models.CharField(max_length=255)
file = models.FileField(upload_to="imports/")
content_hash = models.CharField(max_length=64, blank=True, db_index=True)
```
Compute SHA256 from the uploaded bytes before creating a new `SourceFile`. Always rewind the file afterwards:
```python
import hashlib
def uploaded_file_sha256(uploaded_file):
h = hashlib.sha256()
for chunk in uploaded_file.chunks():
h.update(chunk)
uploaded_file.seek(0)
return h.hexdigest()
```
Decision table:
| Existing row? | Same file name? | Same hash? | Action |
|---|---:|---:|---|
| no | n/a | n/a | create new `SourceFile`, parse rows |
| yes | yes | yes | skip import; report as unchanged/skipped |
| yes | yes | no | update existing `SourceFile`, delete old parsed child rows, reparse |
| yes | no | yes | usually skip as duplicate content; optionally report duplicate content under different name |
| yes | no | no | create new source file |
For source types that represent operational settlement documents, prefer matching on `(source, file_name)` for update semantics, plus `content_hash` for duplicate detection.
## Transaction boundary
Wrap the parser/import in a transaction so child rows are not partially replaced:
```python
from django.db import transaction
@transaction.atomic
def import_payment_advice(source_file):
PaymentAdviceDocument.objects.filter(source_file=source_file).delete()
# parse and recreate child rows/documents
```
If the upload flow creates a new file and parsing fails, delete only that newly-created `SourceFile`. Do not delete/revert an existing source record unless you have an explicit rollback strategy.
## UI behaviour
Return counters that distinguish:
- imported files
- updated files
- skipped unchanged files
- failed files
Use user-facing messages like:
- `1 uændret fil sprunget over`
- `1 fil opdateret`
- `23 filer importeret`
This matters for accounting users: they need confidence that re-uploading a folder does not inflate totals.
## Bulk upload settings
Django defaults may reject large batches with `TooManyFilesSent`. For workflows where users upload many small settlement documents, make the limit explicit in settings/env, for example:
```python
DATA_UPLOAD_MAX_NUMBER_FILES = int(os.getenv("DATA_UPLOAD_MAX_NUMBER_FILES", "2000"))
```
Add the setting to `.env.example` and document that it controls maximum files per HTTP upload.
## Tests to add
Minimum regression tests:
1. Upload same file twice with identical bytes:
- `SourceFile.objects.count()` unchanged after second upload
- child document/line count unchanged
- response contains skipped/unchanged message
2. Upload same file name with changed bytes:
- same `SourceFile` primary key reused
- `content_hash` changes
- old child rows are deleted/replaced, not appended
- response contains updated message
3. Bulk upload many files if the UI supports multi-file import:
- no `TooManyFilesSent`
- all valid new files imported
- repeated files skipped
4. Parser failure path:
- failed new file does not leave orphan `SourceFile`
- failed update does not silently erase previous valid data
## Documentation checklist
Update README/import docs with:
- duplicate uploads are skipped by content hash
- same file name with changed content updates the existing source
- updates replace parsed rows for that source file
- original source files are not edited in place; imported copies live under media/storage
- bulk upload file-count limit and env var