Files
Radix-Skills/skills/data-science/erp-data-reconciliation/references/idempotent-source-file-imports.md
T
dennis 236053cd7f Initial import: 10 Radix skills across 3 categories
Kategoriserer danske regnskabs-/ERP-skills, ERP-data-afstemnings-
skill og to devops-workflows (gitea-issue-agent, kanban-workflows)
som Radix-medlemmers AI-agenter kan dele.

Struktur:
  skills/<kategori>/<skill-navn>/SKILL.md
  skills/<kategori>/<skill-navn>/{references,templates,scripts}/

Indekseret af index.json (genereret af scripts/rebuild_index.py).
Kør `python3 scripts/rebuild_index.py --check` i CI for at fange
synkroniseringsfejl.
2026-06-15 14:41:04 +02:00

4.2 KiB

Idempotent SourceFile imports for accounting/reconciliation apps

Use this when users can upload the same ERP/accounting source files repeatedly, especially bulk settlement files such as Coop Payment_advice PDFs/XLSX.

Problem

Financial import apps often store one SourceFile plus many parsed rows/documents. If a user uploads the same file again, the app must not create:

  • duplicate SourceFile records
  • duplicate parsed settlement/payment/invoice lines
  • duplicated totals in reconciliation reports

If the same named file has changed, treat it as an update of that source, not as a new independent source file.

Durable pattern

Add a content hash to the source-file model:

class SourceFile(models.Model):
    source = models.CharField(max_length=64)
    file_name = models.CharField(max_length=255)
    file = models.FileField(upload_to="imports/")
    content_hash = models.CharField(max_length=64, blank=True, db_index=True)

Compute SHA256 from the uploaded bytes before creating a new SourceFile. Always rewind the file afterwards:

import hashlib

def uploaded_file_sha256(uploaded_file):
    h = hashlib.sha256()
    for chunk in uploaded_file.chunks():
        h.update(chunk)
    uploaded_file.seek(0)
    return h.hexdigest()

Decision table:

Existing row? Same file name? Same hash? Action
no n/a n/a create new SourceFile, parse rows
yes yes yes skip import; report as unchanged/skipped
yes yes no update existing SourceFile, delete old parsed child rows, reparse
yes no yes usually skip as duplicate content; optionally report duplicate content under different name
yes no no create new source file

For source types that represent operational settlement documents, prefer matching on (source, file_name) for update semantics, plus content_hash for duplicate detection.

Transaction boundary

Wrap the parser/import in a transaction so child rows are not partially replaced:

from django.db import transaction

@transaction.atomic
def import_payment_advice(source_file):
    PaymentAdviceDocument.objects.filter(source_file=source_file).delete()
    # parse and recreate child rows/documents

If the upload flow creates a new file and parsing fails, delete only that newly-created SourceFile. Do not delete/revert an existing source record unless you have an explicit rollback strategy.

UI behaviour

Return counters that distinguish:

  • imported files
  • updated files
  • skipped unchanged files
  • failed files

Use user-facing messages like:

  • 1 uændret fil sprunget over
  • 1 fil opdateret
  • 23 filer importeret

This matters for accounting users: they need confidence that re-uploading a folder does not inflate totals.

Bulk upload settings

Django defaults may reject large batches with TooManyFilesSent. For workflows where users upload many small settlement documents, make the limit explicit in settings/env, for example:

DATA_UPLOAD_MAX_NUMBER_FILES = int(os.getenv("DATA_UPLOAD_MAX_NUMBER_FILES", "2000"))

Add the setting to .env.example and document that it controls maximum files per HTTP upload.

Tests to add

Minimum regression tests:

  1. Upload same file twice with identical bytes:

    • SourceFile.objects.count() unchanged after second upload
    • child document/line count unchanged
    • response contains skipped/unchanged message
  2. Upload same file name with changed bytes:

    • same SourceFile primary key reused
    • content_hash changes
    • old child rows are deleted/replaced, not appended
    • response contains updated message
  3. Bulk upload many files if the UI supports multi-file import:

    • no TooManyFilesSent
    • all valid new files imported
    • repeated files skipped
  4. Parser failure path:

    • failed new file does not leave orphan SourceFile
    • failed update does not silently erase previous valid data

Documentation checklist

Update README/import docs with:

  • duplicate uploads are skipped by content hash
  • same file name with changed content updates the existing source
  • updates replace parsed rows for that source file
  • original source files are not edited in place; imported copies live under media/storage
  • bulk upload file-count limit and env var