napari_track_edit.import_export.sql_io ====================================== .. py:module:: napari_track_edit.import_export.sql_io .. autoapi-nested-parse:: Reading and writing tracks as an on-disk tracksdata SQL database. A funtracks ``Tracks`` holds a tracksdata graph, which may be an in-memory ``IndexedRXGraph`` or a database-backed ``SQLGraph``. The SQL backend is interesting for three reasons: the file on disk is always in sync, so a crash loses nothing; several annotators can work against one database; and the candidate graph (the ``solution=False`` nodes) does not have to fit in RAM. Import never converts: CSV and geff always build in-memory graphs. The only way into the SQL backend is to open a database that already exists, or to export one and optionally carry on editing in it. A database written elsewhere - say Ultrack - can be opened too, and the ``_sniff_*`` helpers below exist for that case: they recover from the graph and its metadata what a database nTE wrote would have stated outright. Note that opening such a database writes to it; see :func:`tracks_from_sql`. Note that SQL does not make the *solution* out-of-core. ``Tracks`` builds ``graph_solution`` as a ``GraphView``, which subclasses ``RustWorkXGraph``, so for a SQL root it is materialised in memory. What lives on disk is the full graph. Restricting the solution view to a time window is the follow-up that makes the memory benefit real. This module deliberately holds no Qt, so the round trip can be tested without a running application. Attributes ---------- .. autoapisummary:: napari_track_edit.import_export.sql_io.SQL_SUFFIX napari_track_edit.import_export.sql_io.DRIVERNAME napari_track_edit.import_export.sql_io._POS_SAMPLE_SIZE napari_track_edit.import_export.sql_io.META_KEY Functions --------- .. autoapisummary:: napari_track_edit.import_export.sql_io.is_sql_backed napari_track_edit.import_export.sql_io.sql_database_path napari_track_edit.import_export.sql_io.is_same_database napari_track_edit.import_export.sql_io.write_tracks_to_sql napari_track_edit.import_export.sql_io.close_database napari_track_edit.import_export.sql_io._write_new_database napari_track_edit.import_export.sql_io.tracks_from_sql napari_track_edit.import_export.sql_io.rebind_tracks_to_graph napari_track_edit.import_export.sql_io._describe napari_track_edit.import_export.sql_io._sniff_time_attr napari_track_edit.import_export.sql_io._sniff_scale napari_track_edit.import_export.sql_io._sniff_pos_attr napari_track_edit.import_export.sql_io._pos_is_populated Module Contents --------------- .. py:data:: SQL_SUFFIX :value: '.db' .. py:data:: DRIVERNAME :value: 'sqlite' .. py:data:: _POS_SAMPLE_SIZE :value: 512 .. py:data:: META_KEY :value: 'napari_track_edit' .. py:function:: is_sql_backed(tracks: funtracks.data_model.Tracks) -> bool Whether the tracks are stored in a database rather than in memory. :param tracks: The tracks to inspect. :type tracks: Tracks .. py:function:: sql_database_path(tracks: funtracks.data_model.Tracks) -> pathlib.Path | None The database file backing the given tracks, or None if in memory. Reaches into ``SQLGraph._url``, which tracksdata does not expose publicly. Kept in one place so there is a single site to update if it ever does. :param tracks: The tracks to inspect. :type tracks: Tracks .. py:function:: is_same_database(path: pathlib.Path, tracks: funtracks.data_model.Tracks) -> bool Whether `path` names the database the given tracks already live in. Exporting a database over itself would clear the way before reading it, so this has to catch the same file spelled differently: a symlinked parent (on macOS /tmp is a link to /private/tmp) or, on a case-insensitive filesystem, different capitalisation. ``samefile`` compares the underlying file and is the reliable test when both paths exist; resolving covers the case where the destination does not exist yet. :param path: The proposed destination. :type path: Path :param tracks: The tracks being exported. :type tracks: Tracks .. py:function:: write_tracks_to_sql(tracks: funtracks.data_model.Tracks, path: pathlib.Path, overwrite: bool = False) -> tracksdata.graph.SQLGraph Write tracks to a SQLite database at the given path. Writes ``graph_full``, not ``graph_solution``, so soft-deleted candidates survive. A database exists to be reopened and edited, and a geff round trip already drops candidates and marks everything ``solution=True`` again; the database format should not repeat that. tracksdata copies a SQLite source to a SQLite destination with ``ATTACH DATABASE``, never materialising the graph, but only when the destination is absent or empty. So replacing an existing database means clearing the way first, and doing that in place would destroy the old file before knowing the new one can be written. Instead the graph is written to a temporary file beside the destination and moved over it once complete: a failed export leaves whatever was there untouched. :param tracks: The tracks to write. :type tracks: Tracks :param path: The database file to create. :type path: Path :param overwrite: Whether to replace a database already at `path`. :type overwrite: bool :returns: The graph that was written, open at `path` and ready to be handed to :func:`rebind_tracks_to_graph`. :rtype: td.graph.SQLGraph :raises FileExistsError: If something is already at `path` and `overwrite` is False. .. py:function:: close_database(graph_or_tracks: tracksdata.graph.BaseGraph | funtracks.data_model.Tracks) -> None Release a database's connections. SQLAlchemy engines are not closed when the object goes out of scope, and while an engine is alive Windows will not let the file be renamed over or deleted. Anything that stops using a database should say so. No-op for an in-memory graph, so callers need not check first. :param graph_or_tracks: A tracksdata graph, or tracks holding one. .. py:function:: _write_new_database(tracks: funtracks.data_model.Tracks, path: pathlib.Path) -> tracksdata.graph.SQLGraph Copy the full graph into a database at a path known to be free. .. py:function:: tracks_from_sql(path: pathlib.Path, scale: list[float] | None = None) -> funtracks.data_model.Tracks Open an existing SQLite database as tracks. The database is opened in place, not copied: every later edit is written straight to this file. ``SQLGraph`` reflects the existing schema back when ``overwrite`` is false, which is the default. Opening is **not** read-only. ``Tracks`` adds the ``solution`` node and edge attribute keys if the graph has none, which is an ``ALTER TABLE``, and it computes and writes back any track ids the graph only pretends to have. A database written by nTE already has everything, so nothing happens; a foreign database is written to by being opened. Copy the file first if that matters. A database that records no scale opens without one. Tracks with no scale are an ordinary state throughout the application - loading a geff produces them too - so there is nothing to ask the user about here. :param path: An existing database file. :type path: Path :param scale: Scale to use, overriding whatever the database records. :type scale: list[float] | None :returns: Tracks backed by the database. :rtype: Tracks .. py:function:: rebind_tracks_to_graph(tracks: funtracks.data_model.Tracks, graph: tracksdata.graph.BaseGraph) -> funtracks.data_model.Tracks Return tracks of the same kind, backed by the given graph. Used after exporting to a database, when the user asked to carry on editing in it. Building a new object rather than swapping ``graph_full`` in place keeps the graph/view/annotator wiring entirely in funtracks' hands. The new object starts with an empty action history, so **undo and redo are cleared**: the actions on the old stack hold references into the old graph and cannot be replayed against the new one. Callers must tell the user. :param tracks: The tracks to rebind. Its own graph is left alone. :type tracks: Tracks :param graph: The graph to bind to. :type graph: td.graph.BaseGraph :returns: A new object carrying over everything about `tracks` that is not stored in the graph. A MotileRun rebinds to a MotileRun so its solver params survive; anything else rebinds to a plain Tracks, which is what every other loader in this package returns. :rtype: Tracks .. py:function:: _describe(tracks: funtracks.data_model.Tracks) -> dict[str, Any] What to record in the database beyond the graph itself. The feature keys matter as much as the scale: ``Tracks`` defaults its time attribute to "time", but every funtracks graph calls it "t", so a database reopened without them would be given the wrong FeatureDict. .. py:function:: _sniff_time_attr(graph: tracksdata.graph.BaseGraph) -> str Guess the time attribute of a database written by something else. funtracks graphs use "t"; "time" is the funtracks default and worth trying second so a graph built elsewhere still opens. .. py:function:: _sniff_scale(graph: tracksdata.graph.BaseGraph) -> list[float] | None The scale of a database that spells the scale metadata differently. Nothing is needed here for a database that follows the tracksdata convention: funtracks already backs ``Tracks.scale`` with ``graph_full.metadata["scale"]``, treats it as **spatial-only** and adds the dummy time entry itself. Ultrack writes the same key with the time scale included - one entry per axis of the segmentation, e.g. ``[1, 1.97, 0.485, 0.485]`` for t/z/y/x. Handing that to funtracks unchanged makes it prepend a further entry and then refuse to open the database at all ("Dimensions from segmentation 4, scale 5, and ndim 4 must match"), so the clash has to be resolved here. Detected by length: as long as it equals the number of axes in the segmentation shape, the leading entry is a time scale. Passing the value through as the time-first scale ``Tracks`` accepts also rewrites the metadata into the spatial-only form, so a database only needs this once. Returns None when the metadata already follows the funtracks convention, or when there is no scale to read; in both cases funtracks does the right thing unaided. .. py:function:: _sniff_pos_attr(graph: tracksdata.graph.BaseGraph) -> str | list[str] Guess the position attribute(s) of a database written by something else. A single "pos" array is the funtracks convention; one column per axis is the other shape funtracks accepts. With only one of the two present the answer is obvious. With both, "pos" wins - unless it was never filled in, which is what an Ultrack database looks like: a zeroed "pos" column next to real z/y/x columns. Reading positions from that one puts every node at the origin, and nothing says so until the points are on screen in the wrong place. .. py:function:: _pos_is_populated(graph: tracksdata.graph.BaseGraph) -> bool True if the "pos" column holds anything other than zeros. Only a sample is read: "pos" holds a pickled array per node, so reading the whole column would pull hundreds of megabytes off disk to answer a question a few hundred rows already answer. The sample is spread over the id range rather than taken from the front, so a database that merely starts with unpopulated rows is not misjudged.