Plan provider capacity
The backfill is one provider call per batch, repeated until the corpus is done. The provider is the bottleneck and the bill.
What the plan tells you
Section titled “What the plan tells you”What leaves the database: - text-embedding-3-small at https://api.openai.com/v1, declared hosted - the text of title, body - for 482310 rowsThat row count is the number of texts that will be sent. If it says
for a number of rows nobody counted, the plan could not measure the source —
usually because the table is not there yet — and you are sizing blind.
The two flags
Section titled “The two flags”ptah inference backfill --spec spec.yaml --db-url "$DB" --run-id "$RUN" \ --batch-rows 500 \ --batch-inputs 64 \ --provider-timeout 60s--batch-inputsis how many texts go in one request. Your provider’s documented limit is the ceiling. Larger batches mean fewer round trips and a longer wait before anything commits.--batch-rowsis how many source rows are read in one query. It bounds how long a cancellation waits, because a batch in flight finishes before the run stops.--provider-timeoutbounds one request. Too short and a large batch fails repeatedly; too long and a hung endpoint stalls the run.
A reasonable starting point for a hosted provider: --batch-inputs 64,
--batch-rows 500, --provider-timeout 60s. Then measure.
Rate limits
Section titled “Rate limits”Ptah does not implement backoff against a provider’s rate limiter. A request
that is refused fails the batch, and the run stops with the provider’s own
message. Everything committed stays committed, and running backfill again
resumes.
If you are being rate-limited, lower --batch-inputs and run the backfill in
sessions rather than expecting one invocation to absorb the limit.
The token counts the provider reports are recorded on the run and shown by
status:
- 655360 prompt tokens, 655360 total, as the provider reported themThat is what the provider charged for, and it is the number to compare against your invoice. The batches line above it in the report counts provider round trips, not tokens. Ptah prices nothing and counts no tokens of its own — it has no idea what your contract is.
Not every endpoint reports usage. Where none of a run’s answers carried one,
that line reads the provider reported no token usage instead of two zeros,
because a provider that charged nothing and one that said nothing are not the
same fact and the counts alone cannot tell them apart.
Two things drive the bill more than anything else:
- The row count.
source.filteris the lever. Embedding rows nobody searches is the most common avoidable cost. - The input length. Every column in
input_fieldsis tokens. Abodycolumn with an entire document in it costs many times what a title does, andpreprocessing.max_input_byteswithtruncate: bytesis how you cap it.
Truncation is a decision
Section titled “Truncation is a decision”preprocessing: max_input_bytes: 8000 truncate: refuse # refuse | bytesrefuse stops the run on a row that is too long. bytes cuts the input and
embeds what is left.
There is no default. truncate is required, and a specification omitting it is
refused before any verb does work – which is the same reason stated as a rule
rather than as a default: a silently truncated document produces a vector for
its first half, which searches plausibly and is wrong. refuse is the answer to
write unless you mean otherwise, and bytes is a decision to take deliberately,
knowing that is what it means.
truncate is required even where max_input_bytes names no cap for it to act
at. The block above is two fields of a preprocessing: section rather than a
whole one, so pasting it over a specification’s own block is itself a refusal –
see the specification reference for the fields a
run cannot start without.
Running it overnight
Section titled “Running it overnight”The backfill is resumable, so the safe shape for a large corpus is a loop that runs it, and runs it again if it stopped:
until ptah inference backfill --spec spec.yaml --db-url "$DB" --run-id "$RUN"; do echo "stopped; retrying in 60s" sleep 60doneEach iteration continues from the checkpoint. Nothing is re-embedded and nothing is paid for twice.