# Checkpoints

Squash migration history into a cumulative-schema checkpoint that fresh databases bootstrap from.

Source: https://docs.ptah.run/v0.8.1/versioned/checkpoints/

As a migration directory grows, every fresh database — a CI job, a new
developer machine, a throwaway shadow database — replays the whole history,
including long-dead intermediate DDL for columns and tables that were later
dropped or renamed. A **checkpoint** captures the cumulative schema at a version
so a fresh database can bootstrap from that snapshot and skip the squashed
history, while an already-migrated database ignores the checkpoint and keeps
applying only genuinely pending migrations.

Atlas keeps `migrate checkpoint` in its proprietary Pro build (an Atlas account
and the closed-source binary). Ptah provides checkpoints as an MIT, local,
no-account, embeddable capability.

## Create a checkpoint

`ptah migrations checkpoint` replays the whole directory on an ephemeral
**shadow database**, introspects the resulting schema, and writes a checkpoint
migration whose up body is the full cumulative schema:

The shadow connection's live capabilities and server-resolved identifier
equivalence snapshot are retained through checkpoint planning, including SQL
Server locale, accent, case, kana, and width semantics.

```bash
ptah migrations checkpoint \
  --migrations-dir ./migrations \
  --shadow-db postgres://user:pass@localhost/shadow
```

The shadow database is dropped and replayed from scratch, so it must be an
ephemeral, disposable database — never a real environment. The dialect is
inferred from the shadow database URL; pass `--dialect` only to assert it
explicitly. Use `--dry-run` to print the checkpoint SQL without writing files:

```bash
ptah migrations checkpoint --migrations-dir ./migrations \
  --shadow-db "sqlite://$(mktemp -u).db" --dry-run
```

| Flag | Purpose |
| --- | --- |
| `--shadow-db` | Ephemeral database the directory is replayed into (required). |
| `--migrations-dir` | Directory to checkpoint (default `./migrations`). |
| `--version` | Checkpoint version; defaults to one above the newest migration (`ptah` format) or a UTC timestamp raised to that bound when the bound is higher (`atlas` format). Must be positive and above every existing version; at most ten digits under `ptah`, and never exactly ten rendered digits under `atlas`. |
| `--description` | Description used in the file name (default `checkpoint`). |
| `--dialect` | Asserted dialect; inferred from the shadow database when omitted. |
| `--schemas` | Comma-separated schemas to introspect. |
| `--dir-format` | Checkpoint convention: `ptah` (default) or `atlas`. `auto` is refused. |
| `--data-table` | Reference table whose rows the checkpoint carries, as `table` or `schema.table`. Repeatable. |
| `--dry-run` | Print the checkpoint SQL instead of writing files. |

## Reference rows a fresh database needs

A schema-only checkpoint is enough until a migration after it reads a row an
earlier migration inserted. A fresh database bootstrapped from the checkpoint
never ran that earlier migration, so the row is not there, and the later
migration updates nothing while reporting success.

`--data-table` puts those rows in the checkpoint:

```bash
ptah migrations checkpoint \
  --migrations-dir ./migrations \
  --shadow-db "sqlite://$(mktemp -u).db" \
  --data-table regions \
  --data-table reference.currencies
```

The rows are read from the shadow database **after the replay**, so they are the
rows this history produces at the checkpoint's own version. This matters more
than it sounds: a migration after the checkpoint may read a column later
renamed, or a reference value later retired, and a checkpoint that carried
today's declarations would hand a fresh database values that no migration after
it was written against.

The statements are `INSERT`s after the schema, in the same file, introduced by a
`-- ptah:checkpoint-data <table> rows=<n>` marker, so a checkpoint that carries
data says so in its own bytes. Row order is sorted rather than whatever the
database returned, because a checkpoint whose bytes change between runs is one
whose checksum nobody can reproduce.

A binary column is written as a hexadecimal literal in the engine's own
spelling: `'\x5cff41'::bytea` on PostgreSQL, CockroachDB and Spanner,
`X'5cff41'` on MySQL, MariaDB and SQLite, and `0x5cff41` on SQL Server. A fresh
database receives the bytes the history wrote, including bytes that are not
valid text.

A table that does not exist at the checkpoint's version is refused, and so is a
table with no primary key: rows with no identity are not rows this mechanism can
promise to reproduce.

A checkpoint written without `--data-table` is schema-only, exactly as before.

## File format

A checkpoint is an ordinary Ptah migration pair carrying a `.checkpoint` marker
between the description and the direction:

```text
0000000042_squash.checkpoint.up.sql
0000000042_squash.checkpoint.down.sql
```

The description (`squash` here) comes from `--description` and defaults to
`checkpoint`.

The up body is the full cumulative `CREATE` schema in dependency order; the down
body drops it in reverse. The marker is recognized by discovery and parsing
without changing how ordinary `up`/`down` files are read, so a checkpoint and
the historical migrations it squashes coexist in one directory.

`--dir-format atlas` writes the Atlas convention instead: a single up-only file
named `<version>_<description>.sql` whose **first line** is the
`-- atlas:checkpoint` directive, with `atlas.sum` refreshed rather than
`ptah.sum`. The version is a UTC timestamp (`20060102150405`), as Atlas writes
it, raised to one above the newest migration in the directory whenever that is
higher. Subdirectories count: the replay and the reader both descend into them,
so a nested migration dated after the timestamp would otherwise sort *after* a
checkpoint whose body already contains its SQL, and a fresh database would run
both. There is no down body — the Atlas format is up-only, so an Atlas-format
checkpoint is not reversible. See
[Atlas-compatible surface](#atlas-compatible-surface) below.

## How a checkpoint applies

The behavior depends entirely on whether the target database has any applied
migrations, so the same directory does the right thing everywhere.

- **Fresh database** (empty revision table): `ptah migrations up` runs the
  newest checkpoint's up body, then applies only migrations at or after the
  checkpoint version. The squashed pre-checkpoint migrations are never run
  individually — the checkpoint's own revision row records that history as
  satisfied, so no per-migration rows are written for them.
- **Already-migrated database** (non-empty revision table): the checkpoint is
  ignored entirely and history is applied unchanged. A checkpoint never runs on
  a database that already has migrations applied.

`ptah migrations status` reflects the same decision, so a fresh database lists
the checkpoint plus post-checkpoint migrations as pending, and an already-migrated
database does not list the checkpoint as pending.

Because selection is driven by the checkpoint's own applied state rather than a
separately written baseline, the model is crash-safe: an interrupted bootstrap
resumes to the same end state on the next `up`.

## Integrity

Checkpoint files are covered by the directory's integrity file exactly like
ordinary migrations, so `ptah migrations checkpoint` rewrites it after writing
the checkpoint and a tampered checkpoint fails `ptah migrations validate`. That
is `ptah.sum` for the Ptah convention and `atlas.sum` for `--dir-format atlas`.
Commit the checkpoint files and the updated sum together.

The command captures the directory once before the shadow replay, verifies that
snapshot, and executes the same bytes. The rooted checkpoint writer compares
the directory it holds with the authorized snapshot before it creates the
checkpoint, then compares the expected snapshot-plus-checkpoint before it
publishes the sum. The sum is computed from that expected state. If prior
history changed during the run, the write is refused so a new checksum cannot
launder unverified bytes.

With `--edit`, the command binds the migration directory before opening the
editor. After the editor exits, it accepts changed bytes only in the checkpoint
files created by this run. Every preexisting migration and metadata file must still
match the snapshot that produced the checkpoint; otherwise the command refuses
instead of refreshing the checksum over unrelated concurrent changes. The
edited checkpoint bytes and the replacement checksum are read and published
through that same bound directory.

A directory must carry only one integrity file. Checkpointing a `ptah.sum`
directory under `--dir-format atlas` (or an `atlas.sum` directory under
`--dir-format ptah`) would leave both behind, which `--dir-format auto` refuses
to read, so the command rejects it before writing anything.

An **unhashed** Ptah directory has no sum file to detect, so the same refusal is
also made on file shape: `--dir-format atlas` is rejected whenever the directory
holds Ptah-convention migrations at all. Otherwise the checkpoint would be
written and then stay permanently invisible — discovery reads the Ptah pair and
never sees it, while `validate` finds the `atlas.sum` and reports the directory
as sound.

## Rollback boundary

A checkpoint's down body is meaningful only for a database that bootstrapped
from it. History below the checkpoint boundary no longer has individually
applied migrations to reverse, so `ptah migrations down` to a version between
the checkpoint and the squashed history fails with a clear error rather than
silently doing nothing. You can roll back to the checkpoint boundary, or all the
way to `0` to drop everything (which runs the checkpoint's down body); to land on
an intermediate pre-checkpoint version, restore from a backup or rebuild.

## Checkpoint versus baseline

Both let a database skip running historical migrations, but they solve opposite
problems:

- `ptah migrations baseline` records existing migrations as already applied
  **without executing their SQL** — for adopting Ptah on a database whose schema
  already exists.
- `ptah migrations checkpoint` **executes** a cumulative snapshot on a fresh
  database and then continues with post-checkpoint migrations — for shrinking
  replay time as history grows.

They compose: baseline onto an existing database, checkpoint to keep fresh-setup
fast.

## Atlas-compatible surface

The `ptah-compat` binary's `migrate checkpoint [name]` forwards to the native
command for drop-in Atlas familiarity: `--dir` maps to the migrations directory, `--dev-url`
to the shadow database, and the optional positional name to the checkpoint
description.

On the compat surface `--dir-format` defaults to `atlas`, matching the default
Atlas registers, so an unflagged Atlas pipeline gets an Atlas-format checkpoint
back. The native command keeps `ptah` as its default. Pass `--dir-format ptah`
on compat, or `--dir-format atlas` natively, to select the other convention.

Ptah reads Atlas-format checkpoints: a migration whose first line is
the `-- atlas:checkpoint` file directive (as written by Atlas's own
`migrate checkpoint`) gets the same semantics from both
`ptah-compat migrate apply` and `ptah migrations up` — a fresh database
bootstraps from the latest checkpoint, and a database that already applied
pre-checkpoint history skips the checkpoint silently, matching measured Atlas
behavior.
