Outputs and inspection
A representation is only useful if you can read what it produced. PreTab returns model-ready output in the format you ask for, names every column, and can trace each output column back to the exact input and component that created it. This page covers output shapes, formats, feature names, lineage, and the output budget.
Output shapes
fit_transform and transform return a single stacked numpy.ndarray by default. Pass
output_structure="blocks" (or return_array=False for a single call) to receive a
dictionary that maps each feature to its transformed block instead, with keys prefixed
num_ or cat_.
X_array = pre.fit_transform(df, y) # one stacked ndarray
X_dict = Preprocessor(output_structure="blocks").fit_transform(df, y) # {"num_age": ..., "cat_city": ...}
Note
The array form is what a plain scikit-learn estimator expects, and is what lets a
Preprocessor drop straight into a Pipeline. The dict form is convenient for inspection and
for feeding blocks to different model heads; pass return_array explicitly to override
output_structure for a single call without changing the estimator’s configuration.
Example: one column expands into several
A single input column rarely maps to a single output column. Most numerical methods (splines,
feature maps, PLE) expand one feature into output_dim basis columns, so the matrix width
grows well beyond the number of input columns. A B-spline makes this concrete: output_dim=6
turns the one age column into 6 local basis functions, each capturing a different region of
its range.
df = pd.DataFrame({"age": [25, 40, 63, 51]})
y = [0.1, 0.9, 0.3, 0.6]
pre = Preprocessor(numerical_method="bspline", output_dim=6,
target_aware=False, placement_strategy="quantile").fit(df, y)
out = pre.transform(df)
out.shape
(4, 6)
array([[1. , 0. , 0. , 0. , 0. , 0. ],
[0. , 0.179, 0.593, 0.228, 0. , 0. ],
[0. , 0. , 0. , 0. , 0. , 1. ],
[0. , 0. , 0.165, 0.607, 0.229, 0. ]])
4 input rows, 1 input column, but 6 output columns: every row’s age value is spread across
the basis functions whose local support it falls into (each row sums to 1, since a B-spline
basis is a partition of unity). get_feature_names_out() shows exactly where each column came
from:
list(pre.get_feature_names_out())
['num_age_bs0', 'num_age_bs1', 'num_age_bs2', 'num_age_bs3', 'num_age_bs4', 'num_age_bs5']
All six still trace back to the single age input, which is exactly what output_structure= "blocks" reflects: one dict entry per input feature, holding its full expanded block, not
one entry per output column.
pre_blocks = Preprocessor(numerical_method="bspline", output_dim=6,
target_aware=False, placement_strategy="quantile",
output_structure="blocks").fit(df, y)
pre_blocks.transform(df)["num_age"].shape
(4, 6)
Tip
This is why a wide expansion (a spline or feature map with a large output_dim, or several
expanded columns) can produce a much wider matrix than the input DataFrame had columns.
Use get_feature_info(verbose=True) (below) or estimate_output_shape(df) to see the total
width before committing, especially with several expanded columns at once.
Output format and dtype
Two parameters control the physical layout of the stacked output.
output_formatOne of
"dense","sparse", or"auto"."auto"picks sparse when it saves memory (for example wide one-hot blocks) and dense otherwise. Default"dense".dtypeThe floating-point precision of the output, for example
numpy.float32to halve memory.
pre = Preprocessor(output_format="auto", dtype="float32")
After fitting, output_report_ summarizes what was produced: the chosen format, dimensions,
density, and memory saved.
pre.fit(df, y)
pre.output_report_
DataFrame output
PreTab honours the scikit-learn output API, so you can request pandas or polars frames.
pre.set_output(transform="pandas") # or "polars"
Note
Polars output requires the optional polars extra (pip install "pretab[polars]") and is
loaded lazily: if polars is not installed, requesting it raises a clear
OptionalDependencyError rather than failing deep in the call stack. See
Installation.
Choosing your output settings
Four things independently affect what transform / fit_transform return: output_structure
(top-level shape), return_array (a per-call override of it), output_format (dense vs.
sparse), and set_output (scikit-learn’s own DataFrame protocol). This section is the
decision guide: what each one controls, how they interact, and which combination fits a given
use case, with the literal output shown for each rather than just a description of it.
The examples below all share the same tiny, fixed input so the printed output is directly comparable:
import numpy as np
import pandas as pd
from pretab import Preprocessor
df = pd.DataFrame({"age": [25, 40, 63], "city": ["A", "B", "A"]})
y = [0.1, 0.9, 0.3]
Parameter impact at a glance
Parameter |
Values |
Controls |
Set at |
|---|---|---|---|
|
|
Whether |
Constructor |
|
|
Overrides |
Per call |
|
|
Whether the array (or each block) is a NumPy array or a SciPy CSR matrix. |
Constructor |
|
e.g. |
Floating-point precision of the output; also what |
Constructor |
|
|
Wraps the array in a DataFrame. Takes priority over |
Method call, before |
Which setting should I use?
- Feeding a plain scikit-learn estimator or building a
Pipeline Use the default (
output_structure="matrix", noreturn_arrayoverride).Preprocessor()composes directly:Pipeline([("pretab", Preprocessor(...)), ("model", Ridge())])just works, sincetransform(X)already returns a single array.
pre = Preprocessor(numerical_method="minmax", categorical_method="one-hot").fit(df, y)
out = pre.transform(df)
type(out), out.shape
(<class 'numpy.ndarray'>, (3, 3))
array([[0. , 1. , 0. ],
[0.39473684, 0. , 1. ],
[1. , 1. , 0. ]])
Column order matches get_feature_names_out(): the scaled age, then the one-hot city_A
/ city_B columns.
- Inspecting per-feature blocks, or feeding different blocks to different model heads
Set
output_structure="blocks"on the constructor (so it stays the estimator’s default everywhere it’s reused), or passreturn_array=Falsefor a one-off call without changing the estimator’s configuration.
pre = Preprocessor(numerical_method="minmax", categorical_method="one-hot",
output_structure="blocks").fit(df, y)
out = pre.transform(df)
out
{'num_age': array([[0. ],
[0.39473684],
[1. ]]),
'cat_city': array([[1., 0.],
[0., 1.],
[1., 0.]])}
The same estimator still returns an array for a single call if you ask for one:
pre.transform(df, return_array=True) # ndarray, shape (3, 3); this call only
- Passing external
embeddings Embeddings require dict output. They are separate named blocks, so they cannot be stacked into a single matrix. Use
output_structure="blocks"(orreturn_array=False) on everytransformcall that passesembeddings; the default"matrix"raisesIncompatibleParamsErrorthe momentembeddingsis supplied.
emb = np.array([[0.1, 0.2], [0.3, 0.4], [0.5, 0.6]])
pre = Preprocessor(numerical_method="minmax", categorical_method="one-hot",
output_structure="blocks").fit(df, y, embeddings=emb)
out = pre.transform(df, embeddings=emb)
sorted(out.keys()), out["embedding_1"].shape
(['cat_city', 'embedding_1', 'num_age'], (3, 2))
- Large, mostly one-hot-encoded categorical data
Set
output_format="sparse"(or"auto"to let PreTab decide per fit). This applies independently ofoutput_structure: a sparse"matrix"is a single stackedscipy.sparse.csr_matrix, and a sparse"blocks"dict holds a CSR matrix per feature.
pre = Preprocessor(numerical_method="minmax", categorical_method="one-hot",
output_format="sparse").fit(df, y)
pre.transform(df)
<Compressed Sparse Row sparse matrix of dtype 'float64'
with 5 stored elements and shape (3, 3)>
- Halving memory on a large dataset
Set
dtype=numpy.float32.estimate_memory()and themax_dense_memorybudget both reflect the configureddtype, not a hardcodedfloat64assumption, so budgets sized for the real, cast output are honored correctly.
pre = Preprocessor(numerical_method="minmax", categorical_method="one-hot",
dtype=np.float32).fit(df, y)
pre.transform(df).dtype
dtype('float32')
- Feeding a library that expects a DataFrame, or wanting column names attached to the output
Call
pre.set_output(transform="pandas")(or"polars"). This overrides everything else:transform()always returns a DataFrame from that point on, regardless ofoutput_structureor anyreturn_arraypassed to an individual call.
pre = Preprocessor(numerical_method="minmax", categorical_method="one-hot").fit(df, y)
pre.set_output(transform="pandas").transform(df)
num_age cat_city_A cat_city_B
0 0.000000 1.0 0.0
1 0.394737 0.0 1.0
2 1.000000 1.0 0.0
Warning
set_output always wins. If a downstream step unexpectedly receives a DataFrame instead of
the array or dict you configured, check whether set_output was called anywhere upstream
(including by a cloned copy inside a Pipeline/GridSearchCV).
Feature names
Every representation names its output columns, and the names are stable and descriptive.
get_feature_names_out() returns them in output order.
pre.get_feature_names_out()
Use get_feature_info(verbose=True) for a human-readable table of the resolved per-feature
pipeline, output width, and category count.
feature kind pipeline dim cats
----------------------------------------------------------------
age numerical imputer -> minmax -> bspline 13 -
income numerical imputer -> minmax -> ple 12 -
city categorical imputer -> onehot -> to_float 4 4
verbose=True is the default, so get_feature_info() both logs that table and returns the
same data as a tuple of dicts, one per feature kind, keyed by input column name:
pre.get_feature_info()
feature kind pipeline dim cats
----------------------------------------------------------------
age numerical imputer -> minmax -> bspline 7 -
income numerical imputer -> minmax -> rbf 7 -
city categorical imputer -> onehot -> to_float 3 3
(
{"age": {"preprocessing": "imputer -> minmax -> bspline", "dimension": 7, "categories": None},
"income": {"preprocessing": "imputer -> minmax -> rbf", "dimension": 7, "categories": None}},
{"city": {"preprocessing": "imputer -> onehot -> to_float", "dimension": 3, "categories": 3}},
{},
)
Tip
get_feature_info() is a handy, no-setup way to check exactly what happened to each feature:
which pipeline it went through, how many output columns it produced, and (for categoricals)
how many categories were seen. Pass verbose=False to get just the returned dicts, silently.
For other fitted details (total output width, per-feature output counts, the last
transform’s memory report), see the Preprocessor Attributes in the
API reference.
Fit-time logging
verbose controls how much fit reports through the shared "pretab" logger, useful when
running PreTab inside a larger training script or notebook.
Level |
What is logged |
|---|---|
|
Nothing, aside from |
|
One summary line: feature counts, resolved method(s) per kind, total output width, duration. |
|
Also logs the same table |
|
Also logs internal fitted decisions (bins, knots, centers). |
Preprocessor(numerical_method="ple", verbose=1).fit(df, y)
fit complete: 2 numerical (ple) + 1 categorical (int) feature(s) -> 15 output columns in 0.02s
Note
When feature_preprocessing overrides make columns of the same kind resolve to different
methods, the summary reports mixed: <method>, <method>, ... instead of a single name, so the
log line never misrepresents which method actually ran on a given feature. For the definitive,
per-column answer, use get_feature_info(verbose=True) (level 2 logs the same table).
Tip
This also covers columns left out of feature_preprocessing entirely: they resolve to the
global numerical_method / categorical_method for their kind, and show up in the same
mixed: ... summary whenever a sibling column of that kind was overridden. See
Columns not listed in feature_preprocessing
for a worked example.
Feature lineage
Lineage is the flagship inspection feature. get_feature_lineage() returns one record per
output column, mapping it back to its origin.
lineage = pre.get_feature_lineage()
lineage[0]
Each FeatureLineage record carries:
the source input column(s) the output came from,
the representation family that produced it,
the component it corresponds to (a basis function, knot, center, frequency, interval, or category),
whether the target was used to fit it,
whether it is an interaction across several inputs.
Tip
Lineage covers every output column and the names line up with get_feature_names_out(). This
makes a fitted Preprocessor fully auditable, which is invaluable when you interpret a linear
model fit on top of the expansion.
Output budget
Expansions can multiply columns quickly, especially wide splines or high-cardinality one-hot. The output budget lets you cap the blast radius and estimate cost before committing.
Parameter |
Effect |
|---|---|
|
Cap on total output columns. |
|
Cap on columns produced from any single input. |
|
Cap on dense output memory. |
|
What to do on overflow, default |
pre = Preprocessor(max_output_features=500, overflow_policy="error")
pre.estimate_output_shape(df) # predicted (n_rows, n_cols) without transforming
pre.estimate_memory(df) # predicted dense memory in bytes
Warning
With overflow_policy="error", exceeding a budget raises OutputBudgetError at fit. Use
estimate_output_shape and estimate_memory first when you work with wide expansions or
large data.
Where to go next
Reproducibility to serialize and fingerprint the fitted output.
Representations for what each family emits.
Comparing representations to measure width and memory.