Skip to content
PtahPtah

Plan provider capacity

The backfill is one provider call per batch, repeated until the corpus is done. The provider is the bottleneck and the bill.

What leaves the database:
- text-embedding-3-small at https://api.openai.com/v1, declared hosted
- the text of title, body
- for 482310 rows

That row count is the number of texts that will be sent. If it says for a number of rows nobody counted, the plan could not measure the source — usually because the table is not there yet — and you are sizing blind.

Terminal window
ptah inference backfill --spec spec.yaml --db-url "$DB" --run-id "$RUN" \
--batch-rows 500 \
--batch-inputs 64 \
--provider-timeout 60s
  • --batch-inputs is how many texts go in one request. Your provider’s documented limit is the ceiling. Larger batches mean fewer round trips and a longer wait before anything commits.
  • --batch-rows is how many source rows are read in one query. It bounds how long a cancellation waits, because a batch in flight finishes before the run stops.
  • --provider-timeout bounds one request. Too short and a large batch fails repeatedly; too long and a hung endpoint stalls the run.

A reasonable starting point for a hosted provider: --batch-inputs 64, --batch-rows 500, --provider-timeout 60s. Then measure.

Ptah does not implement backoff against a provider’s rate limiter. A request that is refused fails the batch, and the run stops with the provider’s own message. Everything committed stays committed, and running backfill again resumes.

If you are being rate-limited, lower --batch-inputs and run the backfill in sessions rather than expecting one invocation to absorb the limit.

The token counts the provider reports are recorded on the run and shown by status:

- 655360 prompt tokens, 655360 total, as the provider reported them

That is what the provider charged for, and it is the number to compare against your invoice. The batches line above it in the report counts provider round trips, not tokens. Ptah prices nothing and counts no tokens of its own — it has no idea what your contract is.

Not every endpoint reports usage. Where none of a run’s answers carried one, that line reads the provider reported no token usage instead of two zeros, because a provider that charged nothing and one that said nothing are not the same fact and the counts alone cannot tell them apart.

Two things drive the bill more than anything else:

  • The row count. source.filter is the lever. Embedding rows nobody searches is the most common avoidable cost.
  • The input length. Every column in input_fields is tokens. A body column with an entire document in it costs many times what a title does, and preprocessing.max_input_bytes with truncate: bytes is how you cap it.
preprocessing:
max_input_bytes: 8000
truncate: refuse # refuse | bytes

refuse stops the run on a row that is too long. bytes cuts the input and embeds what is left.

There is no default. truncate is required, and a specification omitting it is refused before any verb does work – which is the same reason stated as a rule rather than as a default: a silently truncated document produces a vector for its first half, which searches plausibly and is wrong. refuse is the answer to write unless you mean otherwise, and bytes is a decision to take deliberately, knowing that is what it means.

truncate is required even where max_input_bytes names no cap for it to act at. The block above is two fields of a preprocessing: section rather than a whole one, so pasting it over a specification’s own block is itself a refusal – see the specification reference for the fields a run cannot start without.

The backfill is resumable, so the safe shape for a large corpus is a loop that runs it, and runs it again if it stopped:

Terminal window
until ptah inference backfill --spec spec.yaml --db-url "$DB" --run-id "$RUN"; do
echo "stopped; retrying in 60s"
sleep 60
done

Each iteration continues from the checkpoint. Nothing is re-embedded and nothing is paid for twice.