Assemble a Context Set for Conditioned Prediction
Measured companion: for the long-form, executed-and-measured Python treatment, see The Cookbook → Retrieval.
A prediction is often best made conditioned on a neighbourhood: not “what is
the label of this row” in the abstract, but “given the k most similar
labelled rows, what is the label of this row.” That neighbourhood is a context
set — C = {(xᵢ, yᵢ)} — and in a database it is not an abstraction to invent.
It is a search joined to its labels.
assemble_context makes that first-class. It retrieves a target’s k nearest
neighbours, pairs them with their outcome columns, and pools the neighbour
vectors — permutation-invariantly — into one fixed-width context vector a
predictor can condition on. It is the encode-and-aggregate half of a Neural
Process; the decode half (a learned predictor over the representation) composes
on top.
When this is the right tool
Reach for assemble_context when you want a reusable set representation of a
target’s retrieved neighbourhood — a conditioning vector, a prototype/centroid,
a bag-of-evidence summary. If you only need to aggregate one specific, already
known set of rows once, that is a SQL GROUP BY you already have; this is the
operator that turns any target’s retrieval into a representation,
reproducibly, with the leakage guards a prediction context needs.
Assemble and encode
Rust
#![allow(unused)]
fn main() {
extern crate jammi_db;
extern crate jammi_ai;
extern crate tokio;
use jammi_ai::session::InferenceSession;
use jammi_ai::pipeline::context_set::{ContextRequest, SetAggregator};
async fn ex(session: &std::sync::Arc<InferenceSession>, query: Vec<f32>) -> jammi_db::error::Result<()> {
let mut request = ContextRequest::new("patents", query, 10);
request.value_columns = vec!["category".into()]; // the labels carried per context row
request.aggregator = SetAggregator::Mean; // mean | sum | max pooling
let context = session.assemble_context(&request).await?;
if let Some(vector) = &context.context_vector {
// condition a predictor on `vector` + `context.context_size`
let _ = (vector, context.context_size);
}
Ok(()) }
}
The result carries:
context_vector— the pooledρ(Σ φ(xᵢ))representation, orNonefor an empty context (no neighbour survived the guards below — treat as low-confidence / fall back to the prior, never as a one-element average).context_size— the number of neighbours that entered the pool, carried separately so a decoder can use the count signal without it corrupting the pooled vector.context_keys— the context members’ keys, in retrieval (descending similarity) order.value_rows— the requestedvalue_columnsof each member, in the same order.
The leakage guards (on by default)
A target that retrieves itself as its own context trivially leaks the answer when a value column is the prediction target. So:
exclude_selfdefaults on. Setexclude_keyto the target’s own row key (when the query vector belongs to a stored row) and that same-key neighbour is dropped before pooling; the retrieval over-fetches by one so a self-hit never shrinks the context belowk.splitscopes the context to a train split. When the context feeds a training or evaluation target, pass e.g.split = Some("split = 'train'".into())so the target’s own outcome stays held out — the same train/target line the evaluation harness enforces.
#![allow(unused)]
fn main() {
extern crate jammi_db;
extern crate jammi_ai;
use jammi_ai::pipeline::context_set::ContextRequest;
fn ex(query: Vec<f32>) {
let mut request = ContextRequest::new("patents", query, 10);
request.exclude_key = Some("US-1234567".into()); // drop the target's own row
request.split = Some("split = 'train'".into()); // context from the train split only
let _ = request;
}
}
Pooling: fixed, permutation-invariant, deterministic
The encoder pools through the engine’s vector-aggregation functions
(vector_mean / vector_sum / vector_max) — the same element-wise
aggregation the engine ships for grouped vector reduction. The pool is:
- permutation-invariant — shuffling the context rows yields a byte-identical vector (the aggregate folds with a commutative, associative operator);
- deterministic under exact retrieval — the pooled vector is reproducible across runs.
mean discards set size; sum encodes it; max is robust but lossy. None is
universally right, which is why the aggregator is a knob and context_size is
always carried alongside.
This is fixed pooling — the DeepSets / Conditional-Neural-Process expressiveness ceiling. It cannot model which context element matters; learned attention pooling (the AttnCNP point on the spectrum) is a separate, downstream capability, not a silent extension of this one.
Materialise for batch workflows
For batch pipelines, pool every target once and land the results as a normal embedding-shaped result table — searchable and joinable like any other embedding table, with its own sidecar ANN index:
#![allow(unused)]
fn main() {
extern crate jammi_db;
extern crate jammi_ai;
extern crate tokio;
use jammi_ai::session::InferenceSession;
use jammi_ai::pipeline::context_set::{ContextRequest, MaterializedContext};
use jammi_db::store::CachePolicy;
async fn ex(session: &std::sync::Arc<InferenceSession>, rows: Vec<(String, Vec<f32>)>) -> jammi_db::error::Result<()> {
// The recipe carries *how* every row was pooled — the source, candidate set,
// value columns, aggregator, self-exclusion, and split — and names the source.
let recipe = ContextRequest::new("patents", vec![0.0_f32; 32], 5);
let (table, _outcome) = session
.materialize_context(
MaterializedContext {
rows: &rows,
dimensions: 32,
recipe: &recipe,
// These targets are `patents` rows, so their keys are `patents.id`.
key_column: Some("id"),
},
CachePolicy::Bypass,
)
.await?;
// `table` is a normal embedding result table: search it, join it, index it.
println!("materialised context table: {}", table.table_name);
Ok(()) }
}
Each target’s key becomes the table’s _row_id; the pooled context vector
becomes its vector. A materialised context set is a first-class member of the
same table family every embedding table belongs to. The recipe materialize_context
takes is the batch’s shared assembly definition (the source comes from it); the
per-target query vectors are the inputs the recipe ran over, which is why they are
the rows, not part of the recipe.
key_column names which column of the source those target keys came from, and is
recorded as the table’s provenance: it is what lets a reader join
patents.id = <context table>._row_id to get back to the rows the targets are.
You declare it because the targets are yours — materialize_context receives only
(key, vector) pairs and cannot see where the keys came from — but it does check
the half it can: a column the source does not have is rejected rather than
recorded. Pass None when the targets are free vectors that correspond to no
stored row; the table is still searchable, but nothing joins it back to a source.