Metadata-Version: 2.5
Name: feincms3-data
Version: 0.11.1
Project-URL: Homepage, https://github.com/matthiask/feincms3-data/
Author-email: Matthias Kestenholz <mk@feinheit.ch>
License: BSD-3-Clause
License-File: LICENSE
Classifier: Environment :: Web Environment
Classifier: Framework :: Django
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: BSD License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Internet :: WWW/HTTP :: Dynamic Content
Classifier: Topic :: Software Development
Classifier: Topic :: Software Development :: Libraries :: Application Frameworks
Requires-Python: >=3.10
Requires-Dist: django>=4.2
Provides-Extra: tests
Requires-Dist: coverage; extra == 'tests'
Description-Content-Type: text/x-rst

=============
feincms3-data
=============

.. image:: https://github.com/matthiask/feincms3-data/actions/workflows/tests.yml/badge.svg
    :target: https://github.com/matthiask/feincms3-data/
    :alt: CI Status


Why
===

Utilities for loading and dumping database data as JSON.

These utilities (partially) replace Django's built-in ``dumpdata`` and
``loaddata`` management commands.

Suppose you want to move data between systems incrementally. In this case it
isn't sufficient to only know the data which has been created or updated; you
also want to know which data has been deleted in the meantime. Django's
``dumpdata`` and ``loaddata`` management commands only support the former case,
not the latter. They also do not including dependent objects in the dump.

This package offers utilities and management commands to address these
shortcomings.


How
===

``pip install feincms3-data``.

Add ``feincms3_data`` to ``INSTALLED_APPS`` so that the included management
commands are discovered.

Add datasets somewhere describing the models and relationships you want to
dump, e.g. in a module named ``app.f3datasets``:

.. code-block:: python

    from feincms3_data.data import (
        specs_for_app_models,
        specs_for_derived_models,
        specs_for_models,
    )
    from app.dashboard import models as dashboard_models
    from app.world import models as world_models


    def districts(args):
        pks = [int(arg) for arg in args.split(",") if arg]
        return [
            *specs_for_models(
                [world_models.District],
                {
                    "filter": {"pk__in": pks},
                    "delete_missing": True,
                },
            ),
            *specs_for_models(
                [world_models.Exercise],
                {
                    "filter": {"district__in": pks},
                    "delete_missing": True,
                },
            ),
            # All derived non-abstract models which aren't proxies:
            *specs_for_derived_models(
                world_models.ExercisePlugin,
                {
                    "filter": {"parent__district__in": pks},
                    "delete_missing": True,
                },
            ),
        ]


    def datasets():
        return {
            "articles": {
                "specs": lambda args: specs_for_app_models(
                    "articles",
                    {"delete_missing": True},
                ),
            },
            "pages": {
                "specs": lambda args: specs_for_app_models(
                    "pages",
                    {"delete_missing": True},
                ),
            },
            "teachingmaterials": {
                "specs": lambda args: specs_for_models(
                    [
                        dashboard_models.TeachingMaterialGroup,
                        dashboard_models.TeachingMaterial,
                    ],
                    {"delete_missing": True},
                ),
            },
            "districts": {
                "specs": districts,
            },
        }

Add a setting with the Python module path to the specs function:

.. code-block:: python

    FEINCMS3_DATA_DATASETS = "app.f3datasets.datasets"


Now, to dump e.g. pages you would run::

    ./manage.py f3dumpdata pages > tmp/pages.json

To dump the districts with the primary key of 42 and 43 you would run::

    ./manage.py f3dumpdata districts:42,43 > tmp/districts.json

The resulting JSON file has three top-level keys:

- ``"version": 1``: The version of the dump, because not versioning dumps is a
  recipe for pain down the road.
- ``"specs": [...]``: A list of model specs.
- ``"objects": [...]``: A list of model instances; uses the same serializer as
  Django's ``dumpdata``, everything looks the same.

Model specs consist of the following fields:

- ``"model"``: The lowercased label (``app_label.model_name``) of a model.
- ``"filter"``: A dictionary which can be passed to the ``.filter()`` queryset
  method as keyword arguments; used for determining the objects to dump and the
  objects to remove after loading.
- ``"delete_missing"``: This flag makes the loader delete all objects matching
  ``"filter"`` which do not exist in the dump. Those objects whose deletion is
  a precondition for loading the dump at all -- because they hold unique values
  which an object from the dump is claiming -- are deleted *before* loading
  instead of at the end, right before that spec's own objects are saved.
  Exactly the same objects are deleted, only earlier -- but earlier means a
  ``CASCADE`` from that deletion may still reach objects of specs listed
  *after* this one, even ones which the dump only intends to update (e.g. an
  object which keeps its primary key but has a foreign key repointed to the
  recreated object). Listing dependent specs *before* the one doing the
  deleting avoids this, since their objects get saved -- and repointed -- first.
- ``"ignore_missing_m2m"``: A list of field names where deletions of related
  models should be ignored when restoring. This may be especially useful when
  only transferring content partially between databases.
- ``"save_as_new"``: If present and truish, objects are inserted using new
  primary keys into the database instead of (potentially) overwriting
  pre-existing objects.
- ``"defer_values"``: A list of fields which should receive random garbage when
  loading initially and only receive their real value later. This is especially
  useful to avoid unique constraint errors when loading partial graphs.

.. note::
   Multi table inheritance children share the primary key of their parent. If
   the database says an object is a different concrete model than the dump does
   -- which happens once the databases drift apart, e.g. because objects are
   created on the target as well -- loading is refused with an
   ``InconsistentModelError``.

   Loading anyway would leave the stale row of the other type behind: the
   parent row is shared, so nothing ever removes it, and the result would be an
   object which is two things at once. Removing it automatically isn't an
   option either -- deleting the stale child takes the shared parent row with
   it, and since Django is perfectly happy with a parent having several
   children the row may not even be stale. Delete the offending objects
   yourself and load again.

.. note::
   Objects which have been deleted and recreated on the source database arrive
   with a new primary key, while the target database still holds the row with
   the same unique values. Databases don't allow both rows to exist at the same
   time, so the old row has to go before the dump can be loaded.

   ``"delete_missing"`` handles this by itself, as long as the old row matches
   the spec's ``"filter"``. If you cannot use ``"delete_missing"`` for a model
   -- typically because deletions shouldn't be propagated to the target as soon
   as anything else is transferred -- restrict the filter to the unique values
   contained in the dump instead:

   .. code-block:: python

       specs_for_models(
           [Identifier],
           {
               "filter": {"identifier__in": identifiers},
               "delete_missing": True,
           },
       )

   This only ever deletes rows claiming one of the dumped identifiers (and
   everything hanging off them) and leaves all other identifiers alone. Keep
   the filter in sync with the objects you're actually dumping.

   "Everything hanging off them" includes objects of specs listed *after* the
   one doing the deleting, even if those objects are also part of the dump and
   would otherwise simply have their foreign key repointed to the recreated
   row. List such dependent specs *before* the spec that deletes conflicting
   rows to avoid losing anything -- e.g. local-only data -- attached to them.

.. note::
   When using ``save_as_new`` and ``delete_missing`` together, you may need to
   specify how primary keys should be mapped to avoid inadvertent deletion of
   objects. If you have a parent model with ``save_as_new`` and child models
   with both ``save_as_new`` and ``delete_missing``, you should use the
   dictionary form of ``delete_missing`` with a ``map`` parameter to map old
   primary keys to new ones. For example:

   .. code-block:: python

       {
           "filter": {"parent__in": [42]},
           "save_as_new": True,
           "delete_missing": {
               "map": [
                   ("parent__in", "app.Parent"),
               ],
           },
       }

   This ensures that child objects are matched against the new parent primary
   keys rather than the old ones, preventing old data from being kept and new
   data from being inadvertently deleted.

   Note that the mapping fails loudly on purpose if the primary key mapping
   isn't available. The reason could be that the parent model doesn't actually
   use ``save_as_new``. Since this is a potentially destructive operation it's
   better to fail loudly than to silently eat data.

The dumps can be loaded back into the database by running::

    ./manage.py f3loaddata -v2 tmp/pages.json tmp/districts.json

Each dump is processed in an individual transaction. The data is first loaded
into the database; at the end, data *matching* the filters but whose primary
key wasn't contained in the dump is deleted from the database (if
``"delete_missing": True``). The only exception are objects holding unique
values which the dump's data claims -- those are removed upfront, since
databases do not allow the old and the new row to exist at the same time.

Both deletions are restricted to the spec's ``"filter"``. An object outside of
the filter which holds a unique value claimed by the dump therefore still makes
the load fail; widen the filter (or dump fewer objects) in that case.
