Configuration
The Preprocessor is configured through a small, predictable set of parameters. This page
covers the four ways to express intent: global defaults, per-feature overrides, presets, and
reading back the resolved configuration. The mechanics of width and placement live in
Resolution and placement, and target usage in
Target awareness.
Global defaults
The simplest configuration sets one strategy for every numerical column and one for every categorical column.
from pretab import Preprocessor
pre = Preprocessor(
numerical_method="ple", # applied to every numerical column
categorical_method="int", # applied to every categorical column
)
The defaults are numerical_method="ple" and categorical_method="int". The full list of
strategy strings is in the representation comparison.
NumPy array input
Preprocessor also accepts a plain numpy.ndarray, not just a DataFrame. Columns are
named feature_0, feature_1, … in position order, then detected as numerical or
categorical exactly as they would be for a DataFrame.
import numpy as np
from pretab import Preprocessor
X = np.random.default_rng(0).normal(size=(100, 3))
y = np.random.default_rng(0).normal(size=100)
pre = Preprocessor(numerical_method="ple").fit(X, y)
pre.numerical_features_
['feature_0', 'feature_1', 'feature_2']
Tip
feature_preprocessing works the same way on array input: target the synthetic name, for
example {"feature_0": "rbf"}. Check numerical_features_ / categorical_features_ after
fit (or get_feature_info()) to confirm the names PreTab assigned before writing the
overrides, rather than guessing the column order.
Per-feature overrides
Columns rarely want identical treatment. The feature_preprocessing dict assigns a strategy
to individual columns and takes precedence over the global defaults for those columns.
pre = Preprocessor(
numerical_method="ple", # default for numerical columns not listed
feature_preprocessing={
"age": "naturalspline",
"income": "rbf",
"city": "one-hot",
},
)
Note
A per-feature entry is resolved in the correct namespace for its detected column kind. You do not need to state whether a column is numerical or categorical; PreTab already knows from feature-type detection.
Columns not listed in feature_preprocessing
feature_preprocessing only overrides the columns it names. Any numerical column left out
falls back to numerical_method, and any categorical column left out falls back to
categorical_method. This is real method resolution, not just a logging detail: the fallback
column is fit with the global default’s transformer, exactly as if you had listed it yourself.
pre = Preprocessor(
numerical_method="bspline", # applies to every numerical column not listed below
feature_preprocessing={
"income": "rbf", # overrides the default for this one column
"city": "one-hot",
},
).fit(df, y)
Column |
Kind |
Listed in |
Resolved method |
|---|---|---|---|
|
numerical |
no |
|
|
numerical |
yes, |
|
|
categorical |
yes, |
|
|
categorical |
no |
|
Tip
Don’t infer the resolved method per column from the constructor arguments alone. Call
get_feature_info(verbose=True) after fit for the definitive per-column table, or set
verbose=2 (or higher) on the Preprocessor to log the same table at fit time. See
Fit-time logging.
Presets
Presets are transparent, named bundles of parameters for common intents. They set the same
knobs you could set by hand, so nothing is hidden, and each one resolves to a fixed,
documented set of values. numerical_method is the one exception: every preset resolves it
from task instead of a fixed value, since a spline basis and piecewise-linear encoding suit
regression and classification differently.
Preset |
|
|
|
|
|
|
|---|---|---|---|---|---|---|
|
task-dependent |
|
|
|
- |
- |
|
task-dependent |
|
|
|
- |
- |
|
task-dependent |
|
- |
|
|
|
numerical_method resolves to "bspline" when task="regression" (the default) and to
"ple" when task="classification", for every preset:
standard_regression = Preprocessor(preset="standard", task="regression")
standard_classification = Preprocessor(preset="standard", task="classification")
standard_regression.get_resolved_config()["numerical_method"] # "bspline"
standard_classification.get_resolved_config()["numerical_method"] # "ple"
Note
The regression/classification split follows Kumar et al. (2026), “From Uniform to Learned Knots: A Study of Spline-Based Numerical Encodings for Tabular Deep Learning”, Transactions on Machine Learning Research. See the representations overview for the full citation list.
standard = Preprocessor(preset="standard")
expanded = Preprocessor(preset="expanded")
standard.get_resolved_config()["categorical_method"] # "int": compact integer codes
expanded.get_resolved_config()["categorical_method"] # "one-hot": one column per category
expanded.get_resolved_config()["output_dim"] # 10: wider than "standard"
So "standard" is the balanced default (output_dim=7, integer-coded categoricals),
"expanded" widens the representation to output_dim=10 and one-hot-encodes categoricals
instead, and "adaptive" lets each feature pick its own width between min_output_dim=7 and
max_output_dim=15 rather than using a fixed output_dim.
Tip
A preset is a starting point, not a lock. Any parameter you pass alongside a preset overrides
the preset’s value for that knob, including numerical_method itself, for example
Preprocessor(preset="expanded", numerical_method="rbf").
Reading the resolved configuration
Because global defaults, per-feature overrides, and presets interact, PreTab lets you read
back exactly what will be used. get_resolved_config() returns the fully resolved settings
as a plain dict.
pre = Preprocessor(preset="expanded", feature_preprocessing={"age": "bspline"})
pre.get_resolved_config()
This is the authoritative answer to “what did my configuration actually become”, and it is useful in tests and reproducible experiments.
How the layers combine
The resolution order is deterministic. Later layers win.
Library defaults.
A
preset, if given.Explicit constructor arguments (
numerical_method,output_dim, and so on).Per-column
feature_preprocessingentries, for the columns they name.
Warning
Configuration is validated at fit time, not silently coerced. An invalid combination, such
as a method that requires the target used with target_aware=False, raises a typed error.
This is intentional: it surfaces mistakes early rather than producing a quietly wrong
representation.
Key parameters at a glance
The parameters below are the ones you reach for most. Each links to the page that explains it in depth.
Parameter |
Default |
Covered in |
|---|---|---|
|
|
this page |
|
|
this page |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
Where to go next
Resolution and placement for width and location.
Target awareness for supervised placement.
Representations for what each method does.