Caching General Key/Value Outputs¶
KeyValueCache is a general-purpose cache for
Transformer outputs. It stores values based on one or more key columns,
and reuses cached values whenever those keys are encountered again.
Example use case: I have an expensive transformer that produces features from query and
docno, and I want to reuse those feature values across multiple experiments.
You use a KeyValueCache in place of the transformer in a pipeline. It holds a reference to
the transformer so that it can compute values that are missing from the cache.
Warning
Important Caveats:
The cache key is based only on the columns listed in
key. If other columns affect outputs, they must be included inkeytoo.For a new cache,
keyandvaluemust be provided when constructing the cache.If a key is missing and no backing transformer is provided, a
LookupErroris raised.A
KeyValueCacherepresents the combination of a transformer and data distribution. Reusing the same cache with a different transformer can produce invalid results.
Example 1: Cache feature values by query and docno¶
KeyValueCache¶import numpy as np
import pyterrier as pt
from pyterrier_caching import KeyValueCache
feature_transformer = pt.apply.generic(lambda df: df.assign(
features=[np.array([1.0], dtype=np.float32) for _ in range(len(df))],
))
cached_features = KeyValueCache(
'path/to/cache',
feature_transformer,
key=['query', 'docno'],
value=['features'],
)
# First call computes and stores feature vectors
cached_features(topics_and_results_df)
# Second call reuses cached values for seen (query, docno) pairs
cached_features(topics_and_results_df)
Example 2: Cache multiple outputs per key¶
KeyValueCache¶import pyterrier as pt
from pyterrier_caching import KeyValueCache
feature_builder = pt.apply.generic(lambda df: df.assign(
qlen=df['query'].str.len(),
has_question_word=df['query'].str.contains(r'\b(what|when|where|which|who|how|why)\b', case=False, regex=True).astype(int),
))
cached_features = KeyValueCache(
'path/to/feature-cache',
feature_builder,
key=['query', 'docno'],
value=['qlen', 'has_question_word'],
)
# Works as a normal transformer, but only computes misses
features_df = cached_features(topics_and_results_df)
API Documentation¶
- pyterrier_caching.KeyValueCache¶
alias of
Sqlite3KeyValueCache
- class pyterrier_caching.Sqlite3KeyValueCache(path, transformer=None, *, key=None, value=None, verbose=False)[source]¶
A cache for storing and retrieving scores for documents, backed by a SQLite3 database.
This is a sparse scorer cache, meaning that space is only allocated for documents that have been scored. If a large proportion of the corpus is expected to be scored, a dense cache (e.g.,
Hdf5ScorerCache) may be more appropriate.- Parameters:
path (str) – The path to the directory where the cache should be stored.
transformer (Transformer) – The transformer to apply to get misses from the cache when value isn’t found for key.
key (str | List[str] | Tuple[str] | None) – The name of the column(s) in the input DataFrame that contains the key for looking up value. Can be omitted for caches that have already been created.
value (str | List[str] | Tuple[str] | None) – The name of the column(s) in the input DataFrame that contains the value value to cache. Can be omitted for caches that have already been created.
verbose (bool) – Whether to print verbose output when scoring documents.
If a cache does not yet exist at the provided
path, a new one is created.Added in version 0.4.0.