Runnable Recipes
Every recipe under cookbook/recipes/
ships as a runnable example.py next to a markdown README and is wired
into CI via tests/cookbook_smoke.py — a broken recipe blocks the
merge. These recipes are the OSS source of truth; this page mirrors each
README below.
For the long-form, measured companion — The Cookbook, a Theory↔Computation book that shows one recipe = one equation in the graph-signal-processing monograph = one line of the GNN canon, executed against committed goldens — see The Cookbook. The two are complementary: these How-To Guides are the short, dual-language (Rust + Python), compile-tested “how do I call this verb” reference; The Cookbook is the long-form, Python, executed-and-measured narrative.
The recipes shipped at MVP:
| Recipe | Demonstrates |
|---|---|
mutable_tables | Create/insert/select/drop on a mutable companion table |
trigger_streams | Publish + subscribe on a topic via the in-process broker |
eval_embeddings | recall@k, MRR, nDCG against a golden set |
image_search | Image-to-image search + Recall@K / MRR eval; vision-tower LoRA fine-tune with a served-change assertion; refusal on an unmatched selector |
eval_inference | Accuracy + macro F1 against gold labels |
eval_inference_ner | Entity-level precision / recall / F1 against gold spans |
fine_tune | LoRA fine-tune end-to-end |
flight_sql | Query a remote jammi-server over Arrow Flight SQL |
audio_search | Audio-to-audio search + Recall@K / MRR eval; audio-tower LoRA fine-tune with a served-change assertion; refusal on an unmatched selector |
search_audit | Per-query provenance audit of a search |
session_lifecycle | Ephemeral session storage with scoped cleanup |
Mutable tables
End-to-end create / insert / select / drop on a Jammi mutable table — the OSS primitive for state that needs to live alongside read-only result tables.
When to use this pattern. You need a writable table that sits in the same SQL catalog as your registered sources and embedding tables — for caching enriched rows, holding cursor state, recording user feedback, or any “small table I want to UPDATE / DELETE / INSERT from SQL” workload — without standing up an external Postgres.
What example.py does
- Connects to a temporary artifact dir
- Creates a
notesmutable table with anint64primary key +utf8body column - Inserts three rows through DataFusion DML (
INSERT INTO ...) - Verifies count and ordering via
SELECT - Drops the table, then asserts a
SELECTafter the drop raises - Demonstrates the idempotent
drop_mutable_table(..., if_exists=True)
API surface exercised
Database.create_mutable_table(name, *, schema, primary_key, ...)Database.sql("INSERT INTO mutable.public.<name> ...")Database.sql("SELECT ... FROM mutable.public.<name>")Database.drop_mutable_table(name, *, if_exists=False)
The DataFusion namespace for mutable tables is always
mutable.public.<name> — distinct from registered sources, which live
under <source>.public.<source>.
Run it
python cookbook/recipes/mutable_tables/example.py
Exits 0 on success, prints mutable_tables: OK on the last line.
Trigger streams
End-to-end publish + subscribe on a Jammi topic, plus the registration and listing surface. Uses the embedded in-process broker — no NATS or external broker needed.
When to use this pattern. You need a low-friction event bus inside your application — for fan-out to downstream consumers, fan-in from batch jobs, or replay-from-offset semantics — without bringing up Kafka or NATS in dev/test. The same surface scales out to NATS JetStream by flipping a config flag at deploy time.
What example.py does
- Connects to a temporary artifact dir
- Registers a topic
events.demowith a typed schema and broker metadata - Confirms
list_topics()returns the new topic - Publishes a 3-row batch through
publish_topic— captures the broker-assigned offset - Subscribes from
from_offset=0and round-trips the same rows back - Drops the topic, confirms it’s gone from
list_topics() - Demonstrates idempotent
drop_topic(..., if_exists=True)and strict-mode failure when dropping a missing topic
API surface exercised
Database.register_topic(name, *, schema, broker_metadata=None)Database.list_topics()Database.publish_topic(name, *, batch)— returns the assigned offsetDatabase.subscribe_collect(name, *, from_offset, max_batches)Database.drop_topic(name, *, if_exists=False)
The subscribe_collect path drives the replay-from-backing-table flow
when from_offset=0; the live-tail flow is exercised in the broker
integration suite.
Run it
python cookbook/recipes/trigger_streams/example.py
Exits 0 on success, prints trigger_streams: OK on the last line.
Evaluate retrieval quality
Measure recall@k, precision@k, MRR, and nDCG of an embedding index against a golden relevance set.
When to use this pattern. You have a corpus and a small set of (query, expected document) judgments, and you need a number that tells you “is my new encoder better than the one I shipped last month?” The same loop powers nightly regression dashboards and A/B model comparison.
What example.py does
- Connects to a temporary artifact dir
- Registers the tiny corpus as a Parquet source
- Builds 32-dim embeddings over the
contentcolumn with the localtiny_bertfixture - Reads
cookbook/fixtures/tiny_golden.json, expands it into the(query_id, query_text, relevant_id)CSV shapeeval_embeddingsconsumes, and registers it as agoldensource - Calls
db.eval_embeddings(source="corpus", golden_source="golden.public.golden", k=5) - Asserts each aggregate metric is in
[0.0, 1.0]and the per-query records carry their golden-setquery_id
API surface exercised
Database.generate_embeddings(*, source, model, columns, key, modality="text")Database.eval_embeddings(*, source, golden_source, model=None, k=10)
The returned dict carries aggregate (mean across queries — recall_at_k,
precision_at_k, mrr, ndcg) and per_query (one entry per query with
query_id and a metrics sub-dict of the same four names, un-averaged).
Golden source shape
eval_embeddings requires a registered source with these columns:
| column | type | example |
|---|---|---|
query_id | utf8 | q1 |
query_text | utf8 | quantum computing applications |
relevant_id | utf8 | 1 (matches corpus.id as a string) |
Image queries are supported via a query_image BLOB column instead of
query_text; cross-modal eval is out of scope for this recipe.
Run it
python cookbook/recipes/eval_embeddings/example.py
Exits 0 on success, prints the metrics dict + eval_embeddings: OK.
Image search
Run image-to-image semantic search over a corpus with an OpenCLIP-format vision model, measure retrieval quality, and adapt the vision tower itself on caller-supplied image triplets.
When to use this pattern. You have a corpus of images (figures, drawings,
photos) and want to find the ones most similar to a query image — and a number
that tells you how good the retrieval is. This is the image counterpart of the
text eval_embeddings recipe.
Flow
- Load a small image corpus (inline image bytes in a Parquet source)
- Generate L2-normalized vision embeddings over the image column
- Search the index with an encoded image query (cosine ANN)
- Eval retrieval quality (Recall@K / MRR) against a held-out golden set
- Fine-tune LoRA adapters inside the vision tower on image triplets (adapted ≠ base), then watch a wrong selector get refused
Model
A domain-specialized CLIP checkpoint is a drop-in for the reference model
when your corpus is technical drawings or diagrams rather than photographs — a
generic CLIP has seen few of them, and a checkpoint tuned on that kind of
imagery separates them far better. patentclip/PatentCLIP_Vit_B on the Hugging
Face Hub is one such checkpoint:
JAMMI_IMAGE_MODEL=patentclip/PatentCLIP_Vit_B \
python cookbook/recipes/image_search/example.py
patentclip/PatentCLIP_Vit_B is pulled from the Hugging Face Hub on first use
and produces 512-dim L2-normalized embeddings. Any OpenCLIP-format model
works the same way — OpenAI CLIP, LAION CLIP-ViT-B-32-*, EVA-CLIP, etc. — the
encoder is auto-detected from the model’s open_clip_config.json.
By default (no env var) the recipe runs against the hermetic
cookbook/fixtures/tiny_open_clip fixture so it runs offline in CI — about two
seconds on a laptop, nearly all of it the tower-LoRA and refusal legs (the
search-and-eval flow alone is a fraction of a second). That fixture has random
weights, so its retrieval numbers are meaningless — it exercises the full
pipeline, not model quality. Point JAMMI_IMAGE_MODEL at any real checkpoint
for real numbers.
What example.py does
- Connects to a temporary artifact dir
- Reads the 20 committed 224×224 PNGs under
cookbook/fixtures/tiny_image_corpus/into a Parquetcorpussource (image_id,imagebytes) db.generate_embeddings(source="corpus", model=MODEL, columns=["image"], key="image_id", modality="image")db.encode_query(model=MODEL, query=png_bytes, modality="image")→db.search("corpus", query=vec, k=5)(returns apyarrow.Table)- Builds the image-query golden source from
tiny_image_golden.jsonand callsdb.eval_embeddings(source="corpus", golden_source="golden.public.golden", k=5) - Prints the aggregate Recall@K / precision@K / MRR / nDCG and the per-query records. It reports the metrics; it does not assert a quality bar.
- Builds synthetic
(anchor, positive, negative)image triplets from the corpus (positive = same shape family, negative = a different family) and callsdb.fine_tune(source="triplets", base_model=MODEL, columns=["anchor","positive","negative"], method="lora", task="image_embedding", target_modules=["in_proj","c_fc"], ...). A non-emptytarget_modulesputs LoRA inside the vision tower —in_projis the transformer block’s fused-QKV projection,c_fcthe MLP’s first linear — so the tower’s own representation moves. (The audio recipe runs both modes side by side: an empty list instead trains a projection head on a frozen tower — cheaper, less capacity.) It then re-encodes the same query image through the adapted model and asserts the embedding vector changed (max elementwise|Δ| > 1e-4versus the base encoding). - Submits one more job with
target_modules=["q_proj"]— a real selector on plenty of decoder checkpoints and on nothing in an OpenCLIP tower — and assertsjob.wait()raisesjammi.errors.TrainingErrorwhose message echoesq_projand names this tower’s real sites (in_proj, …). It prints the message.
What each leg proves, and the honesty rule
- The tower leg proves the adapter is trained and applied when the model
is served: an adapter that trained but was silently dropped at serve time
leaves the two query vectors bit-identical, and that is what the
|Δ|check catches. It asserts change, not improvement — the default fixture has random weights, so the direction of the change carries no information. The vector check is also the deterministic one: asserting a top-k metric moved is flaky, because on this tiny eval set the rankings rarely flip even when the vectors do. What the tower leg canNOT check from here is the saved adapter’s kind — that it is an encoder-adapters bundle carrying the vision tower’s id. That is engine-internal and is pinned by the engine’s own integration tests; the client surface (describe_model) reports only the model’s id, backend, task and status, so the recipe asserts the task and leans on the|Δ|check for the rest. - The refusal leg proves a selector that matches no site fails the job rather than publishing an adapter that changes nothing — and that the message is actionable, carrying this architecture’s own site vocabulary.
- The independently-known improvement number — tuned retrieval quality beating the base by a measured margin — is not this recipe’s to claim. It belongs to a real-checkpoint chapter (built from a committed GPU-produced cache; planned under issue #421), which does not exist yet. A recipe running a random-weight fixture on a laptop can honestly prove mechanism; it cannot prove quality.
The pairing semantics (what a “positive” means) are the caller’s training data, not the trainer’s: the trainer only minimizes the contrastive triplet loss over whatever images you pair.
Stepwise scripts
example.py runs every phase in one process (this is the version wired into
tests/cookbook_smoke.py). The numbered scripts decompose the search-and-eval
flow and share a persistent workdir, so run them in order:
python cookbook/recipes/image_search/01-load-corpus.py
python cookbook/recipes/image_search/02-generate-embeddings.py
python cookbook/recipes/image_search/03-search.py
python cookbook/recipes/image_search/04-eval.py
API surface exercised
Database.generate_embeddings(*, source, model, columns, key, modality="image")Database.encode_query(*, model, query, modality="image")→list[float]Database.search(source, *, query, k, filter=None, select=None)→pyarrow.TableDatabase.eval_embeddings(*, source, golden_source, model=None, k=10)Database.fine_tune(*, source, base_model, columns, method, task="image_embedding", target_modules=[...], ...)→TrainingJobDatabase.describe_model(model_id)→dict | None
Image triplet schema (fine-tune input)
| column | type | notes |
|---|---|---|
anchor | binary | encoded image |
positive | binary | an image the caller deems related |
negative | binary | an image the caller deems unrelated |
Same column shape as text and audio triplets — task="image_embedding" is what
tells the loader to read the three columns as encoded images rather than text.
Vision-tower LoRA sites
target_modules names sites on this architecture. The OpenCLIP towers
offer in_proj (fused QKV), out_proj, c_fc and c_proj; all-linear
selects every one. A selector matches a site name exactly or as a suffix of it.
A list matching nothing fails the job with a message that echoes what you
submitted and lists the tower’s real names.
Input schema
| column | type | notes |
|---|---|---|
image_id | utf8 | per-row key |
image | binary | raw PNG/JPEG/TIFF bytes (decoded by the encoder) |
Preprocessing (pad-to-square, no center crop, normalization, L2-normalized
output) is handled inside the encoder per the model’s preprocess_cfg.
Golden source shape (image mode)
eval_embeddings switches to image-query mode when the golden source carries a
query_image (binary) column instead of query_text:
| column | type | example |
|---|---|---|
query_id | utf8 | q_circle |
query_image | binary | raw PNG bytes of the query image |
relevant_id | utf8 | img_circle_0 (matches image_id) |
Fixtures
cookbook/fixtures/tiny_image_corpus/— 20 synthetic 224×224 PNGs in 5 shape families (circle / triangle / square / hexagon / grating), 4 per family, plus a held-out query image per family underqueries/. Rendered programmatically bycookbook/fixtures/generate.py— no real-world imagery (licensing).cookbook/fixtures/tiny_image_golden.json— per-query → expected corpus IDs (same shape family).cookbook/fixtures/tiny_open_clip/— tiny offline OpenCLIP fixture used as the default CI model.
Run it
python cookbook/recipes/image_search/example.py
Exits 0 on success, prints the top-K and the metrics dict + image_search: OK.
Evaluate inference (classification)
Run a classifier over a registered source and score its predictions against gold labels.
When to use this pattern. You have a labelled holdout set and you want a single number — accuracy, macro F1, per-class F1 — to compare two classifiers, or to track drift over time on the same classifier.
What example.py does
- Connects to a temporary artifact dir
- Registers the tiny corpus as
corpus(parquet) - Registers
tiny_labels.csvasgolden(csv) —(id, label)rows - Runs
db.eval_inferencewith the localtiny_modernbert_classifierfixture against thecontentcolumn - Prints the returned aggregate
accuracy, macrof1, per-class metrics, and the count of per-record predictions - Asserts every reported rate is in
[0.0, 1.0]
API surface exercised
Database.eval_inference(*, model, source, columns, task, golden_source, label_column)
The returned dict carries aggregate (tagged by "task" — currently
"classification") with accuracy, f1, and per_class, plus
per_record (one entry per aligned {record_id, predicted, gold}).
The task argument is the string form of the inference task —
"classification" here. For NER, see ../eval_inference_ner/.
Golden source shape
eval_inference requires a registered source with these columns:
| column | type | example |
|---|---|---|
id | utf8 | "1" |
<label_column> | utf8 | physics |
label_column is the kwarg you pass at call time — label in this
recipe. Every id in the golden source must resolve to a row in the
input source; rows without a gold label are silently dropped from the
metric.
Run it
python cookbook/recipes/eval_inference/example.py
Exits 0 on success, prints the metrics dict + eval_inference: OK.
Evaluate inference (NER)
Run a token-classification model over a registered source and score its predicted entity spans against gold spans.
When to use this pattern. You have a labelled NER holdout set (one gold span per row) and you want strict entity-level precision, recall, and F1 — both overall and per entity type — to compare two NER models or to track regressions on the same one.
What example.py does
- Connects to a temporary artifact dir
- Registers
tiny_ner_corpus.parquetascorpus(parquet) - Registers
tiny_ner_gold.csvasgolden(csv) — one row per gold entity span:(id, label, start, end) - Runs
db.eval_inferencewith the localtiny_modernbert_nerfixture against thetextcolumn,task="ner" - Prints the returned aggregate
precision,recall,f1, the per-type breakdown, and the count of per-record predictions - Asserts every reported rate is in
[0.0, 1.0]
API surface exercised
Database.eval_inference(*, model, source, columns, task, golden_source, label_column)
The returned dict carries aggregate (tagged by "task" — "ner" for
this recipe) with precision, recall, f1, and per_type (one
breakdown per entity type the model emitted or the gold set carried),
plus per_record (one entry per aligned {record_id, predicted, gold}
where predicted and gold are entity-span lists, each tagged
"task": "ner").
The task argument is the string form of the inference task — "ner"
here. For classification, see ../eval_inference/.
Golden source shape
eval_inference with task="ner" requires a registered source with
these columns — one row per entity span (multiple spans on the same
id accumulate into one per-row gold set):
| column | type | example |
|---|---|---|
id | utf8 | "1" |
<label_column> | utf8 | PER |
start | i64 | 0 |
end | i64 | 13 |
label_column is the kwarg you pass at call time — label in this
recipe. start is inclusive, end is exclusive, both byte offsets
into the source row’s text column. The label set must match the
shipped model’s id2label minus the B-/I- prefixes —
tiny_modernbert_ner knows PER and ORG only.
Rows in the source without a matching gold id are silently dropped
from the metric (same alignment rule the classification recipe uses).
Run it
python cookbook/recipes/eval_inference_ner/example.py
Exits 0 on success, prints the metrics dict + eval_inference (ner): OK.
Fine-tune an encoder
Run a LoRA fine-tune on top of an existing text encoder, poll the job to completion, and use the resulting checkpoint to encode a query.
When to use this pattern. Your domain (legal contracts, medical abstracts, patent claims, internal product docs) doesn’t match the distribution the base encoder was trained on, and you have a few hundred to a few thousand labelled or contrastive pairs. LoRA gets you ~80% of the lift of a full fine-tune at a fraction of the cost; the resulting adapter is small enough to ship as an attachment to the base model rather than a re-distributed full checkpoint.
What example.py does
- Connects to a temporary artifact dir
- Registers
tiny_pairs.csv(30 contrastive pairs) astraining - Calls
db.fine_tune(...)with the localtiny_bertbase, a small LoRA rank, and one epoch — kept fast for CI - Waits for terminal status via
job.wait() - Asserts the resulting
model_idstarts withjammi:fine-tuned: - Encodes a query through the fine-tuned model to confirm it loads
API surface exercised
Database.fine_tune(*, source, base_model, columns, method, task=..., ...)Job.wait()Job.job_id,Job.output_model_idDatabase.encode_query(*, model, query, modality="text")
The full keyword list on fine_tune covers LoRA rank/alpha/dropout,
learning rate, epochs, batch size, max sequence length, validation
fraction, early-stopping patience/metric, warmup, gradient accumulation,
backbone dtype, weight decay, and gradient clipping — the recipe uses
the defaults for everything except rank and epochs.
Performance note
This recipe is excluded from the per-PR smoke matrix because even at one
epoch it runs ~30 seconds on CPU. The nightly cron with
JAMMI_COOKBOOK_SLOW=1 includes it. Override the gate locally:
JAMMI_COOKBOOK_SLOW=1 python tests/cookbook_smoke.py
Run it
python cookbook/recipes/fine_tune/example.py
Exits 0 on success, prints job_id, model_id, and fine_tune: OK.
Connect via Flight SQL
Run a query against a remote jammi-server over Arrow Flight SQL.
When to use this pattern. You’re connecting from a non-Python client
(Tableau, dbt, JDBC tools, Rust binaries), or you want to expose Jammi
to multiple readers without each one holding an embedded session. The
same protocol is what dbt-flightsql, the official Flight SQL JDBC
driver, and BI tools speak natively.
What example.py does
- Spawns
target/release/jammi-serveras a child process pointed at a tempartifact_dir - Polls the health endpoint (
http://127.0.0.1:8080/healthz) until the server is ready (5 s budget) - Opens a
pyarrow.flight.FlightClientagainstgrpc://127.0.0.1:8081 - Submits
SELECT 1 AS oneover Flight SQL and confirms the response - Tears down the server process cleanly
This recipe is gated out of the per-PR CI matrix — it depends on the
jammi-server binary being built (cargo build --release -p jammi-server),
and the build cost dominates the test wall-clock. The nightly cookbook job
builds the binary and runs the recipe behind JAMMI_COOKBOOK_SLOW=1.
Prerequisites
cargo build --release -p jammi-server— producestarget/release/jammi-serverpip install pyarrow(already ajammi-aidependency)
The script auto-detects JAMMI_BIN (env var) or falls back to the
workspace’s target/release/jammi-server.
API surface exercised
pyarrow.flight.FlightClient.execute(query)over the Flight SQL command dialectjammi-server— the OSS deployment-shape binary entrypoint
Run it
cargo build --release -p jammi-server # one-time build
python cookbook/recipes/flight_sql/example.py
Exits 0 on success, prints the query result + flight_sql: OK.
Audio search
Run audio-to-audio similarity search over a corpus with a CLAP-format audio model, measure retrieval quality, and domain-tune the audio embeddings on caller-supplied triplets.
When to use this pattern. You have a corpus of sounds (clips, stems,
loops, recordings) and want to find the ones most similar to a query clip — and
a number that tells you how good the retrieval is. This is the audio
counterpart of the image eval_embeddings recipe; audio is simply the third
embedding modality the engine supports alongside text and images.
Flow
- Load a small audio corpus (inline audio bytes in a Parquet source)
- Generate L2-normalized audio embeddings over the audio column
- Search the index with an encoded audio query (cosine ANN)
- Eval retrieval quality (Recall@K / MRR) against a held-out golden set
- Fine-tune a projection head on audio triplets and re-eval (tuned ≠ base)
- Fine-tune LoRA adapters inside the audio tower on the same triplets (adapted ≠ base), then watch a wrong selector get refused
Model
Any HuggingFace CLAP audio model works — its config.json declares
model_type = "clap_audio_model" (or lists ClapModel /
ClapAudioModelWithProjection in architectures), its checkpoint exposes the
audio_model.audio_encoder.* + audio_projection.* HTSAT-Swin tower keys, and
a preprocessor_config.json carries the feature-extractor geometry. The encoder
is auto-detected from that config, exactly as the image recipe auto-detects
OpenCLIP:
JAMMI_AUDIO_MODEL=<hf-repo-id-or-local-path> \
python cookbook/recipes/audio_search/example.py
By default (no env var) the recipe runs against the hermetic
cookbook/fixtures/htsat_clap_tiny fixture so it runs offline in CI — on the
order of ten seconds on a laptop, dominated by the three training legs (the
projection head, the tower LoRA, and the refused job). That fixture has random
weights, so its retrieval numbers are meaningless — it exercises the full
pipeline, not model quality. Point JAMMI_AUDIO_MODEL at a real CLAP
checkpoint for real numbers.
What example.py does
- Connects to a temporary artifact dir
- Reads the 20 committed mono WAV clips under
cookbook/fixtures/tiny_audio_corpus/into a Parquetcorpussource (clip_id,audiobytes) db.generate_embeddings(source="corpus", model=MODEL, columns=["audio"], key="clip_id", modality="audio")db.encode_query(model=MODEL, query=wav_bytes, modality="audio")→db.search("corpus", query=vec, k=5)(returns apyarrow.Table)- Builds the audio-query golden source from
tiny_audio_golden.jsonand callsdb.eval_embeddings(source="corpus", golden_source="golden.public.golden", k=5) - Prints the base aggregate Recall@K / precision@K / MRR / nDCG and the per-query records. It reports the metrics; it does not assert a quality bar.
- Builds synthetic
(anchor, positive, negative)audio triplets from the corpus (positive = same timbre family, negative = a different family) and callsdb.fine_tune(source="triplets", base_model=MODEL, columns=["anchor","positive","negative"], method="lora", task="audio_embedding", ...). Emptytarget_modules⇒ a trainable projection head on the frozen CLAP audio tower (the cheap, low-risk lightweight mode). It then re-embeds the corpus with the tuned model, re-evals, and prints base-vs-tuned metrics for narrative. For correctness it re-encodes the same query clip through the tuned model and asserts the embedding vector changed (max elementwise|Δ| > 1e-4versus the base encoding) — the real invariant fine-tuning guarantees, and a deterministic check. (Asserting on the coarse top-k metrics instead is flaky: on this tiny eval set the rankings rarely flip even when the vectors move.) It proves the adapter alters audio retrieval — not that it improves it; the random-weight fixture’s direction is not meaningful, real lift comes from a real checkpoint. - Runs the other fine-tune mode on the same triplets:
target_modules=["query", "value", "linear1"]puts LoRA inside the HTSAT-Swin tower itself —query/valueare the Swin blocks’ attention projections (indexed by stage),linear1the audio projection head’s first linear (an unindexed site). It re-encodes the same query clip through the adapted model and asserts the same|Δ| > 1e-4change. - Submits one more job with
target_modules=["q_proj"]— a real selector on plenty of decoder checkpoints and on nothing in an HTSAT-Swin tower — and assertsjob.wait()raisesjammi.errors.TrainingErrorwhose message echoesq_projand names this tower’s real sites (query,linear1, …). It prints the message.
What each leg proves, and the honesty rule
Both fine-tune legs ship and both are real; they are different capabilities:
- The projection-head leg (empty
target_modules) trains a new map on top of a tower whose weights never move — cheap, low-risk, no site names needed. - The tower leg (non-empty
target_modules) moves the tower’s own representation — more capacity for a domain the base checkpoint never saw, at more compute, and it needs the site vocabulary of this architecture.
Both prove the adapter is trained and applied when the model is served: an
adapter that trained but was silently dropped at serve time leaves the two query
vectors bit-identical, and that is what the |Δ| check catches. Both assert
change, not improvement — the default fixture has random weights, so the
direction of the change carries no information.
What neither leg can check from here is the saved adapter’s kind — that it is
an encoder-adapters bundle carrying the audio tower’s id. That is
engine-internal and is pinned by the engine’s own integration tests; the client
surface (describe_model) reports only the model’s id, backend, task and
status, so the recipe asserts the task and leans on the |Δ| check for the rest.
The refusal leg proves a selector matching no site fails the job rather than publishing an adapter that changes nothing, and that the message is actionable — it carries this architecture’s own site names.
The independently-known improvement number — tuned retrieval quality beating the base by a measured margin — is not this recipe’s to claim. It belongs to a real-checkpoint chapter (built from a committed GPU-produced cache; planned under issue #421), which does not exist yet. A recipe running a random-weight fixture on a laptop can honestly prove mechanism; it cannot prove quality.
The pairing semantics (what a “positive” means) are the caller’s training data, not the trainer’s: the trainer only minimizes the contrastive triplet loss over whatever clips you pair.
Stepwise scripts
example.py runs every phase in one process (this is the version wired into
tests/cookbook_smoke.py). The numbered scripts decompose the search-and-eval
flow and share a persistent workdir, so run them in order:
python cookbook/recipes/audio_search/01-load-corpus.py
python cookbook/recipes/audio_search/02-generate-embeddings.py
python cookbook/recipes/audio_search/03-search.py
python cookbook/recipes/audio_search/04-eval.py
API surface exercised
Database.generate_embeddings(*, source, model, columns, key, modality="audio")Database.encode_query(*, model, query, modality="audio")→list[float]Database.search(source, *, query, k, filter=None, select=None)→pyarrow.TableDatabase.eval_embeddings(*, source, golden_source, model=None, k=10)Database.fine_tune(*, source, base_model, columns, method, task="audio_embedding", target_modules=[...], ...)→TrainingJobDatabase.describe_model(model_id)→dict | None
Audio triplet schema (fine-tune input)
| column | type | notes |
|---|---|---|
anchor | binary | encoded audio clip |
positive | binary | a clip the caller deems related |
negative | binary | a clip the caller deems unrelated |
Same column shape as text triplets — task="audio_embedding" is what tells the
loader to read the three columns as encoded audio rather than text.
Audio-tower LoRA sites
target_modules names sites on this architecture. An empty list means
“no tower sites” and selects the projection-head mode instead. The HTSAT-Swin
audio tower offers query, key, value, attention_output,
intermediate_dense, output_dense, reduction, linear1 and linear2;
all-linear selects every one. A selector matches a site name exactly or as a
suffix of it. A non-empty list matching nothing fails the job with a message
that echoes what you submitted and lists the tower’s real names.
Input schema
| column | type | notes |
|---|---|---|
clip_id | utf8 | per-row key |
audio | binary | raw WAV/FLAC/MP3/Ogg bytes (decoded by the encoder) |
Preprocessing (decode → resample to the model’s sample rate → CLAP fusion
log-mel spectrogram → HTSAT-Swin tower → L2-normalized output) is handled inside
the encoder per the model’s preprocessor_config.json feature-extractor
geometry. The audio column may also hold file-path strings instead of inline
bytes.
Golden source shape (audio mode)
eval_embeddings switches to audio-query mode when the golden source carries a
query_audio (binary) column instead of query_text / query_image:
| column | type | example |
|---|---|---|
query_id | utf8 | q_sine |
query_audio | binary | raw WAV bytes of the query clip |
relevant_id | utf8 | clip_sine_0 (matches clip_id) |
Fixtures
cookbook/fixtures/tiny_audio_corpus/— 20 synthetic mono WAV clips in 5 timbre families (sine / harmonic / square / saw / noise), 4 per family, plus a held-out query clip per family underqueries/. Synthesised programmatically bycookbook/fixtures/generate.py— no recorded audio (licensing), no tenant data.cookbook/fixtures/tiny_audio_golden.json— per-query → expected corpus IDs (same timbre family).cookbook/fixtures/htsat_clap_tiny/— tiny offline HTSAT-Swin CLAP fixture used as the default CI model, generated bytests/fixtures/generate_htsat_clap.py.
Run it
python cookbook/recipes/audio_search/example.py
Exits 0 on success, prints the top-K and the metrics dict + audio_search: OK.
Per-query search audit
Record a tamper-evident audit row for every search: what was queried, with what model, what came back, and when. The substrate signs each record, stores it tenant-scoped, and publishes it to a trigger topic — so you do not hand-roll an audit schema, a signature scheme, and a stream integration in every project.
This is the primitive every audited-ML deployment in a regulated setting (finance, healthcare, legal, and the like) needs to answer “show me exactly what this model returned for this query, and prove the record hasn’t been altered.”
What this recipe shows
- Build a
PerQueryAuditrecord (query id, model id/version, query lineage, top-K result ids, retrieval scores). db.audit.log([...])— the substrate injectstenant_id, signs the record with a per-tenant HMAC-SHA256 key, stores it, and publishes it.db.audit.fetch_by_query_id(...)/db.audit.fetch_recent(...)— typed reads, tenant-scoped.record.verify()— re-derive the key and check the signature.- Plain SQL over
mutable.public."_jammi_search_audit"— same tenant scope. db.subscribe_collect("jammi.audit.search.v1", ...)— every logged record is also delivered on a trigger topic for alerting / analytics / warehouse sinks.
Run it
The audit master key is required — the substrate refuses to sign without it:
export JAMMI_AUDIT_MASTER_KEY=$(python -c "import secrets; print(secrets.token_hex(32))")
python cookbook/recipes/search_audit/example.py
The key derives a distinct signing secret per tenant via HKDF-SHA256 and is deterministic across restarts, so signatures written today verify after a redeploy. Source it from your secret manager — never hard-code it.
Key points
- Lineage is capped.
query_lineageJSON may not exceed 8 KiB (override withJAMMI_AUDIT_MAX_LINEAGE_BYTES). Store image hashes and row IDs, not raw payloads — compliance posture is structural, not advisory. top_k_result_idsandretrieval_scoresmust be the same length. This is checked when you construct the record.- The table is reserved.
_jammi_search_auditis created implicitly on the first log; you cannot create or directlyINSERTinto it (that would bypass signing). Read it freely via SQL. - Tenant isolation is automatic. A record logged under tenant A is invisible to tenant B, through both the typed API and raw SQL.
Ephemeral session storage
A session-scoped storage context whose tables are auto-deleted when the session
ends — on explicit close(), on context-manager exit, or when the 60-second
timeout scanner force-closes a session past its deadline. Every transition
publishes to the jammi.audit.session_lifecycle.v1 trigger topic, giving an
audit-log aggregator durable proof that the data was deleted.
Run it:
python cookbook/recipes/session_lifecycle/example.py
When to use it
Use an ephemeral session for sensitive transient data that must not outlive the request that produced it: uploaded images, derived embeddings, draft model inputs. The session is always tenant-scoped — tenant A can never see tenant B’s ephemeral tables.
When NOT to use it
Do not store long-lived data in an ephemeral session. The audit record, the persistent corpus, and anything compliance needs to read later belong in ordinary mutable tables. The pattern is: keep the throwaway working set (raw bytes, embeddings) in the ephemeral session, and write only durable lineage (hashes, ids, scores) to a persistent table — before you close the session, while the working data still exists.
API
with db.ephemeral_session(timeout_seconds=3600) as ephem:
ephem.create_ephemeral_table("imgs", schema=schema, primary_key=["image_id"])
ephem.insert("imgs", batch=table)
rows = ephem.sql("imgs", "SELECT image_hash FROM {table}")
# close() runs on exit: tables dropped, `closed` event published
{table} in a sql query is replaced by the tenant-scoped reference to the
named ephemeral table. The context manager is the recommended path; Drop is
best-effort. Lifecycle events (opened, closed, timed_out,
partial_deletion_failure) carry the session id, tenant, table count, and
deleted-row count.