DeepMTP.data package
The data namespace owns built-in dataset helpers, interaction normalization, feature validation, splitting, preparation, model-ready loading, and structured progress events:
from DeepMTP.data import (
DataLoaderFactory,
DataProgressObserver,
TrainingDataLoaders,
data_process,
)
Batches
Model-ready dataloaders return DeepMTP.data.batches.MTPBatch. It is a
mapping compatible with the historical tensor keys and also exposes typed
dense, graph, sparse, ID, image, sequence, tabular, composite, or custom branch
inputs. Graph, tabular, and sequence inputs remain structured; sparse inputs
use PyTorch COO or CSR tensors; composite inputs keep named component values
together. Optional composite components carry their present sub-batch and a
full-batch Boolean mask. Calling batch.to(device) moves every component
together.
Typed model-input batches shared by data loading and training.
- class DeepMTP.data.batches.CompositeBranchBatch(values: CompositeInput)
Bases:
objectA named collection of inputs consumed by a composite encoder.
- property batch_size: int
Number of observations shared by every component.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'composite'
- pin_memory() CompositeBranchBatch
Pin every component in the structured input.
- to(device: torch.device | str, *, non_blocking: bool = False) CompositeBranchBatch
Move every component to one device.
- values: CompositeInput
- class DeepMTP.data.batches.CompositeInput(components: Mapping[str, Any])
Bases:
Mapping[str,Any]Named model inputs for multiple encoders on one entity axis.
- property batch_size: int
Number of observations shared by every component.
- components: Mapping[str, Any]
- pin_memory() CompositeInput
Pin every component for asynchronous host-to-device transfer.
- to(device: torch.device | str, *, non_blocking: bool = False) CompositeInput
Move every component to one device.
- class DeepMTP.data.batches.CustomBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA tensor batch passed unchanged to a user-supplied branch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'custom'
- class DeepMTP.data.batches.DenseBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA batch of dense feature vectors shaped
[batch, features].- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'dense'
- class DeepMTP.data.batches.GraphBranchBatch(values: GraphInput)
Bases:
objectA structured batch consumed by a graph neural network encoder.
- property batch_size: int
Number of graphs represented by this branch batch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'graph'
- pin_memory() GraphBranchBatch
Pin every tensor in the graph input.
- to(device: torch.device | str, *, non_blocking: bool = False) GraphBranchBatch
Move the graph input to one device.
- values: GraphInput
- class DeepMTP.data.batches.GraphInput(node_features: torch.Tensor, edge_index: torch.Tensor, edge_features: torch.Tensor | None, graph_index: torch.Tensor, num_graphs: int)
Bases:
objectOne batch of homogeneous graphs represented by DeepMTP tensors.
- property batch_size: int
Number of graphs represented by this input.
- edge_features: torch.Tensor | None
- edge_index: torch.Tensor
- graph_index: torch.Tensor
- node_features: torch.Tensor
- num_graphs: int
- pin_memory() GraphInput
Pin every graph tensor for asynchronous device transfer.
- to(device: torch.device | str, *, non_blocking: bool = False) GraphInput
Move every graph tensor to one device.
- class DeepMTP.data.batches.IDBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA batch of zero-based integer entity IDs shaped
[batch].- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'id'
- class DeepMTP.data.batches.ImageBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA batch of images shaped
[batch, channels, height, width].- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'image'
- class DeepMTP.data.batches.MTPBatch(instance_id: torch.Tensor, target_id: torch.Tensor, instance_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch, target_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch, score: torch.Tensor)
Bases:
Mapping[str,Any]A complete interaction batch with mapping compatibility.
Attribute access exposes the typed branch inputs. Historical mapping keys continue to return the underlying tensors so existing dataloader consumers remain compatible.
- property batch_size: int
Number of interactions represented by the batch.
- property instance_features: torch.Tensor | CompositeInput | GraphInput | SequenceInput | TabularInput
Compatibility view of the instance branch input.
- instance_id: torch.Tensor
- instance_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch
- score: torch.Tensor
- property target_features: torch.Tensor | CompositeInput | GraphInput | SequenceInput | TabularInput
Compatibility view of the target branch input.
- target_id: torch.Tensor
- class DeepMTP.data.batches.MTPBatchCollator(instance_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite'], target_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite'], instance_sequence_padding_idx: int = 0, target_sequence_padding_idx: int = 0, instance_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None, target_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None, instance_component_padding_indices: Mapping[str, int] | None = None, target_component_padding_indices: Mapping[str, int] | None = None, instance_component_optional: Mapping[str, bool] | None = None, target_component_optional: Mapping[str, bool] | None = None)
Bases:
objectCollate dataset items into a typed interaction batch.
- instance_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None
- instance_component_optional: Mapping[str, bool] | None = None
- instance_component_padding_indices: Mapping[str, int] | None = None
- instance_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']
- instance_sequence_padding_idx: int = 0
- target_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None
- target_component_optional: Mapping[str, bool] | None = None
- target_component_padding_indices: Mapping[str, int] | None = None
- target_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']
- target_sequence_padding_idx: int = 0
- class DeepMTP.data.batches.MaskedComponentInput(values: Any | None, presence_mask: torch.Tensor)
Bases:
objectPresent values plus a full-batch mask for one optional component.
- property batch_size: int
Full batch size including rows where this component is missing.
- pin_memory() MaskedComponentInput
Pin the present values and full-batch mask.
- presence_mask: torch.Tensor
- property present_count: int
Number of rows containing this component.
- to(device: torch.device | str, *, non_blocking: bool = False) MaskedComponentInput
Move the values and presence mask to one device.
- values: Any | None
- class DeepMTP.data.batches.SequenceBranchBatch(values: SequenceInput)
Bases:
objectA structured batch consumed by a token-sequence encoder.
- property batch_size: int
Number of observations represented by this branch batch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'sequence'
- pin_memory() SequenceBranchBatch
Pin every tensor in the structured sequence input.
- to(device: torch.device | str, *, non_blocking: bool = False) SequenceBranchBatch
Move the structured sequence input to one device.
- values: SequenceInput
- class DeepMTP.data.batches.SequenceInput(token_ids: torch.Tensor, attention_mask: torch.Tensor, lengths: torch.Tensor, padding_idx: int = 0)
Bases:
objectPadded token IDs, attention mask, and original sequence lengths.
- attention_mask: torch.Tensor
- property batch_size: int
Number of token sequences represented by this input.
- lengths: torch.Tensor
- padding_idx: int = 0
- pin_memory() SequenceInput
Pin every sequence tensor for asynchronous device transfer.
- to(device: torch.device | str, *, non_blocking: bool = False) SequenceInput
Move every sequence tensor to one device.
- token_ids: torch.Tensor
- class DeepMTP.data.batches.SparseBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA sparse feature batch shaped
[batch, features].- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'sparse'
- pin_memory() SparseBranchBatch
Pin sparse indices and values for asynchronous device transfer.
- class DeepMTP.data.batches.TabularBranchBatch(values: TabularInput)
Bases:
objectA structured batch of numeric values and categorical indices.
- property batch_size: int
Number of observations represented by this branch batch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'tabular'
- pin_memory() TabularBranchBatch
Pin both tensors in the structured input.
- to(device: torch.device | str, *, non_blocking: bool = False) TabularBranchBatch
Move the structured input to one device.
- values: TabularInput
- class DeepMTP.data.batches.TabularInput(numeric: torch.Tensor, categorical: torch.Tensor)
Bases:
objectNumeric and categorical tensors consumed by a tabular encoder.
- property batch_size: int
Number of observations in the structured input.
- categorical: torch.Tensor
- numeric: torch.Tensor
- pin_memory() TabularInput
Pin both tensors for asynchronous host-to-device transfer.
- to(device: torch.device | str, *, non_blocking: bool = False) TabularInput
Move both tensors to one device.
- class DeepMTP.data.batches.TensorBranchBatch(values: torch.Tensor)
Bases:
objectOne branch’s tensor payload with an explicit modality contract.
- property batch_size: int
Number of observations represented by this branch batch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'custom'
- pin_memory() TensorBranchBatch
Return the same typed batch backed by pinned host memory.
- to(device: torch.device | str, *, non_blocking: bool = False) TensorBranchBatch
Return the same typed batch with its tensor moved to
device.
- values: torch.Tensor
Tabular preprocessing
Tabular schemas declare numeric columns, category vocabularies, missing-value policies, and normalization. Fitted training statistics are serializable and are reused for validation, testing, and restored-model prediction.
Schemas and reproducible preprocessing for mixed tabular branch inputs.
- class DeepMTP.data.tabular.CategoricalColumn(name: str, categories: tuple[str | int, ...], embedding_dim: int)
Bases:
objectOne categorical feature and its fixed training vocabulary.
- categories: tuple[str | int, ...]
- embedding_dim: int
- classmethod from_config(name: object, value: object) CategoricalColumn
Validate one categorical-column declaration.
- name: str
- to_dict() dict[str, Any]
Return a JSON-compatible public configuration mapping.
- property vocabulary_size: int
Embedding-table size, including index zero for unknown values.
- class DeepMTP.data.tabular.TabularPreprocessingState(numeric_columns: tuple[str, ...], categorical_columns: tuple[str, ...], numeric_fill_values: tuple[float, ...], numeric_offsets: tuple[float, ...], numeric_scales: tuple[float, ...])
Bases:
objectNumeric statistics fitted only on training entities.
- categorical_columns: tuple[str, ...]
- classmethod from_config(value: object, *, schema: TabularSchema, branch: str) TabularPreprocessingState
Restore and validate preprocessing metadata from a checkpoint config.
- numeric_columns: tuple[str, ...]
- numeric_fill_values: tuple[float, ...]
- numeric_offsets: tuple[float, ...]
- numeric_scales: tuple[float, ...]
- to_dict() dict[str, Any]
Return JSON-compatible checkpoint metadata.
- class DeepMTP.data.tabular.TabularPreprocessor(schema: TabularSchema, state: TabularPreprocessingState, branch: str)
Bases:
objectTransform entity feature frames using one schema and fitted state.
- branch: str
- classmethod fit(frame: DataFrame, schema: TabularSchema, *, branch: str) TabularPreprocessor
Fit numeric imputation and normalization on a training frame.
- classmethod from_state(schema: TabularSchema, state: Mapping[str, Any] | TabularPreprocessingState, *, branch: str) TabularPreprocessor
Build a preprocessor from saved training metadata.
- schema: TabularSchema
- state: TabularPreprocessingState
- transform(frame: DataFrame) DataFrame
Return an indexed feature frame with structured tensor-ready values.
- class DeepMTP.data.tabular.TabularSchema(numeric_columns: tuple[str, ...], categorical_columns: tuple[CategoricalColumn, ...], numeric_normalization: Literal['none', 'standard', 'minmax'] = 'standard', numeric_missing: Literal['error', 'mean', 'zero'] = 'mean', categorical_missing: Literal['error', 'unknown'] = 'unknown', categorical_unknown: Literal['error', 'unknown'] = 'unknown', feature_gating: bool = False)
Bases:
objectExplicit numeric and categorical layout for one entity branch.
- categorical_columns: tuple[CategoricalColumn, ...]
- categorical_missing: Literal['error', 'unknown'] = 'unknown'
- property categorical_names: tuple[str, ...]
Categorical column names in their stable tensor order.
- categorical_unknown: Literal['error', 'unknown'] = 'unknown'
- property encoded_width: int
Width after concatenating numeric values and category embeddings.
- feature_gating: bool = False
- classmethod from_config(value: object, *, branch: str) TabularSchema
Validate a public tabular schema mapping.
- numeric_columns: tuple[str, ...]
- numeric_missing: Literal['error', 'mean', 'zero'] = 'mean'
- numeric_normalization: Literal['none', 'standard', 'minmax'] = 'standard'
- to_dict() dict[str, Any]
Return a JSON-compatible public configuration mapping.
Dataset helpers
Dataset-specific dependencies remain lazy until this module or one of its package-level exports is used.
Built-in, generated, and downloadable dataset helpers.
- class DeepMTP.data.datasets.DatasetBundle
Bases:
TypedDictStable outer structure returned by legacy dataset helpers.
- test: DatasetSplit
- train: DatasetSplit
- val: DatasetSplit
- class DeepMTP.data.datasets.DatasetSplit
Bases:
TypedDictOne raw train, validation, or test split.
- X_instance: ndarray | DataFrame | None
- X_target: ndarray | DataFrame | None
- y: ndarray | DataFrame | None
- class DeepMTP.data.datasets.ProcessedDatasetBundle
Bases:
TypedDictProcessed train, validation, and test data.
- test: ProcessedDatasetSplit
- train: ProcessedDatasetSplit
- class DeepMTP.data.datasets.ProcessedDatasetSplit
Bases:
TypedDictOne processed split with optional interactions and side features.
- X_instance: ProcessedDatasetValue | None
- X_target: ProcessedDatasetValue | None
- y: ProcessedDatasetValue | None
- class DeepMTP.data.datasets.ProcessedDatasetValue
Bases:
TypedDictProcessed data container consumed by legacy data utilities.
- data: DataFrame
- DeepMTP.data.datasets.format_mtr_datasets() str
Format the catalog of supported multivariate regression datasets.
- DeepMTP.data.datasets.generate_MTP_dataset(num_instances: int, num_targets: int, num_instance_features: int | None = None, num_target_features: int | None = None, split_instances: Mapping[str, float] | None = None, split_targets: Mapping[str, float] | None = None, return_static_features_data: bool = False) ProcessedDatasetBundle
Generate deterministic triplet data with optional axis-level splits.
Each split mapping requires a test ratio and may include a validation ratio. The train ratio is derived when omitted and verified when supplied. A test-only axis reuses its training IDs in a validation split created on the other axis.
- Parameters:
num_instances – Number of instances.
num_targets – Number of targets.
num_instance_features – Number of generated instance features, or none.
num_target_features – Number of generated target features, or none.
split_instances – Optional instance-axis split ratios.
split_targets – Optional target-axis split ratios.
return_static_features_data – Repeat non-novel side features in holdouts.
- Returns:
Processed triplet and feature frames for train, validation, and test.
- Raises:
TypeError – If a split or boolean flag has the wrong type.
ValueError – If a dimension or split ratio is invalid.
- DeepMTP.data.datasets.generate_dummy_dataset(num_instances: int, num_targets: int, num_instance_features: int, num_target_features: int, error_mu: float, error_sigma: float, sklearn_version: bool = False, seed: int = 42, mode: str = 'u+logv', split_ratio: Mapping[str, float] | None = None) DatasetBundle
Generate a reproducible dyadic regression dataset.
sklearn_versionstarts from normally distributed features. For modes involving logarithms or fractional powers, the relevant features are transformed to a positive domain so every generated score remains real. The error distribution applies to the four legacy noisy relationship modes.- Parameters:
num_instances – Number of instances.
num_targets – Number of targets.
num_instance_features – Number of features per instance.
num_target_features – Number of features per target.
error_mu – Mean of the Gaussian error added by noisy modes.
error_sigma – Non-negative standard deviation of that Gaussian error.
sklearn_version – Generate normally distributed rather than uniform features.
seed – Non-negative 32-bit seed used for generation and splitting.
mode – Mathematical relationship used to generate scores.
split_ratio – Train, validation, and test ratios.
- Returns:
A dataset bundle containing aligned train, validation, and test rows.
- Raises:
TypeError – If a boolean or error parameter has the wrong type.
ValueError – If a dimension, mode, seed, ratio, or generated value is invalid.
- DeepMTP.data.datasets.generate_interaction_matrix(input_path: str | PathLike[str], output_path: str | PathLike[str]) None
Convert annotator labels into an atomically written correctness matrix.
- DeepMTP.data.datasets.load_process_DP(path: str | PathLike[str] = './data', dataset_name: str = 'ern', variant: Literal['undivided', 'divided'] = 'undivided', random_state: int | None = 42, split_ratio: Mapping[str, float] | None = None, split_instance_features: bool = False, split_target_features: bool = False, validation_setting: str = 'B', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load a dyadic-prediction dataset and optionally create held-out splits.
- Parameters:
path – Directory used for the local dataset cache.
dataset_name – One of
"ern","srn","dpie", or"dpii".variant – Return the complete matrices or create train/validation/test splits.
random_state – Seed used by the split operations, or
None.split_ratio – Train, validation, and test proportions.
split_instance_features – Split instance features with interaction rows.
split_target_features – Split target features with interaction columns.
validation_setting – Novelty setting
"B","C", or"D".print_mode –
"basic"for ordinary messages or"dev"for messages prefixed for the development application.progress – Receives dataset loading and download status events.
- Raises:
ValueError – If an option or dataset matrix is invalid.
TypeError – If a feature-splitting flag is not boolean.
DatasetDownloadError – If a required dataset file cannot be downloaded.
FileNotFoundError – If a download does not create all required files.
- DeepMTP.data.datasets.load_process_MC(path: str | PathLike[str] = './data', dataset_name: str = 'ml-100k', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load a matrix-completion dataset from the MovieLens repository.
- Parameters:
path – Directory used for the local dataset cache.
dataset_name – Matrix-completion dataset to load.
print_mode –
"basic"for ordinary messages or"dev"for messages prefixed for the development application.progress – Receives dataset download status events.
- Raises:
ValueError – If an option or the ratings data is invalid.
DatasetDownloadError – If the archive cannot be downloaded or extracted.
FileNotFoundError – If the archive does not contain the expected ratings file.
- Returns:
The ratings triplets and empty side-feature and held-out splits.
- DeepMTP.data.datasets.load_process_MLC(path: str | PathLike[str] = './data', dataset_name: str = 'bibtex', variant: Literal['undivided', 'divided'] = 'undivided', features_type: Literal['numpy', 'dataframe'] = 'numpy', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load a cached or downloadable scikit-multilearn benchmark dataset.
- Parameters:
path – Directory used for the local dataset cache.
dataset_name – Supported multi-label benchmark name.
variant – Return the complete dataset or its published train/test split.
features_type – Return instance features as an array or feature DataFrame.
print_mode –
"basic"for ordinary messages or"dev"for messages prefixed for the development application.progress – Receives dataset loading and download status events.
- Raises:
ValueError – If an option or cached dataset is invalid.
DatasetDownloadError – If required cache files cannot be downloaded safely.
IsADirectoryError – If an expected cache file path is a directory.
- DeepMTP.data.datasets.load_process_MTL(path: str | PathLike[str] = './data', dataset_name: str = 'dog', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load multi-task learning datasets from my custom repository.
- Parameters:
path (str, optional) – The path where the datasets should be stored. If it doesn’t exist, download and store it in this directory. Defaults to ‘./data’.
dataset_name (str, optional) – The name of the multi-task learning dataset. Defaults to ‘dog’.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives dataset download status events. Defaults to None.
- Raises:
ValueError – If a public option or dataset file is invalid.
DatasetDownloadError – If the dataset archive cannot be downloaded safely.
FileNotFoundError – If required score or image files are missing.
- Returns:
A dictionary with all the available data for the multi-task learning dataset.
- Return type:
dict
- DeepMTP.data.datasets.load_process_MTR(path: str | PathLike[str] = './data', dataset_name: str = 'enb', features_type: Literal['numpy', 'dataframe'] = 'numpy', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load multivariate regression datasets from the mulan repository.
- Parameters:
path (str, optional) – The path where the datasets should be stored. If it doesn’t exist, download and store it in this directory. Defaults to ‘./data’.
dataset_name (str, optional) – The name of the multivariate regression dataset. Defaults to ‘enb’.
features_type (str, optional) – The format of the instance features. There are two possible values, numpy and dataframe. This is intended to test the functionality of the data_process function. Defaults to ‘numpy’.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives dataset loading and download status events. Defaults to None.
- Raises:
ValueError – If a public option or downloaded dataset is invalid.
DatasetDownloadError – If the dataset archive cannot be downloaded safely.
FileNotFoundError – If the archive does not contain the requested dataset.
- Returns:
A dictionary with all the available data for the multivariate regression dataset.
- Return type:
dict
- DeepMTP.data.datasets.print_MTR_datasets(*, progress: DataProgressObserver | None = None) None
Present the supported multivariate regression dataset catalog.
The default observer preserves the historical console table. Inject
NullDataProgressObserverto suppress output or another data progress observer to render the catalog elsewhere.
- DeepMTP.data.datasets.process_dummy_DP(num_instance_features: int = 10, num_target_features: int = 3, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', instance_features_format: Literal['numpy', 'dataframe'] = 'numpy', target_features_format: Literal['numpy', 'dataframe'] = 'numpy', random_state: int | None = 42) DatasetBundle
Generate a reproducible dummy dyadic prediction dataset.
- Parameters:
num_instance_features – Number of features per instance.
num_target_features – Number of features per target.
num_instances – Number of instances.
num_targets – Number of targets.
interaction_matrix_format – Return scores as a dense matrix or triplets.
instance_features_format – Return instance features as an array or DataFrame.
target_features_format – Return target features as an array or DataFrame.
random_state – Seed for an isolated NumPy generator. Use
Nonefor non-deterministic data.
- Returns:
A dataset bundle containing an undivided training split.
- Raises:
ValueError – If a format, dimension, or random state is invalid.
- DeepMTP.data.datasets.process_dummy_MLC(num_features: int = 10, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', features_format: Literal['numpy', 'dataframe'] = 'numpy', random_state: int | None = 42) DatasetBundle
Generate a reproducible dummy multi-label classification dataset.
- Parameters:
num_features – Number of features per instance.
num_instances – Number of instances.
num_targets – Number of binary labels.
interaction_matrix_format – Return labels as a dense matrix or triplets.
features_format – Return features as a dense matrix or feature DataFrame.
random_state – Seed for an isolated NumPy generator. Use
Nonefor non-deterministic data.
- Returns:
A dataset bundle containing an undivided training split.
- Raises:
ValueError – If a format, dimension, or random state is invalid.
- DeepMTP.data.datasets.process_dummy_MTR(num_features: int = 10, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', features_format: Literal['numpy', 'dataframe'] = 'numpy', variant: Literal['undivided', 'divided'] = 'undivided', split_ratio: Mapping[str, float] | None = None, random_state: int | None = 42) DatasetBundle
Generate a reproducible dummy multivariate regression dataset.
- Parameters:
num_features – Number of features per instance.
num_instances – Number of instances.
num_targets – Number of regression targets.
interaction_matrix_format – Return scores as a dense matrix or triplets.
features_format – Return features as a dense matrix or feature DataFrame.
variant – Return an undivided dataset or train/validation/test splits.
split_ratio – Ratios used for the divided variant.
random_state – Seed used for both data generation and splitting. Use
Nonefor non-deterministic data.
- Returns:
A dataset bundle containing the generated scores and features.
- Raises:
TypeError – If
split_ratiois not a mapping.ValueError – If an option, dimension, ratio, or random state is invalid.
Interactions
Interaction normalization, validation, and setting inference.
- class DeepMTP.data.interactions.InteractionInfo
Bases:
TypedDictNormalized interaction data and its detected schema.
- data: DataFrame
- instance_id_type: Literal['int']
- missing_values: bool
- original_format: Literal['triplets', 'numpy']
- target_id_type: Literal['int']
- exception DeepMTP.data.interactions.NoveltyInferenceWarning
Bases:
UserWarningWarning emitted when entity identity was lost in matrix input.
- DeepMTP.data.interactions.check_interaction_files_column_type_format(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
Checks the type of the instance and target ids in the interaction data. If any inconsistencies are detected (like different id types between train and test interaction data), an exception is raised
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose ID-type validation messages. Defaults to None.
- DeepMTP.data.interactions.check_interaction_files_format(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
Checks the format of the interaction data. If any inconsistencies are detected (like different formats between train and test interaction data), an exception is raised
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose format-validation messages. Defaults to None.
- DeepMTP.data.interactions.check_novel_instances(train: Mapping[str, Any] | None, test: Mapping[str, Any] | None, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) bool | None
Return whether test interactions contain previously unseen instances.
progressreceives messages whenverboseis enabled.
- DeepMTP.data.interactions.check_novel_targets(train: Mapping[str, Any] | None, test: Mapping[str, Any] | None, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) bool | None
Return whether test interactions contain previously unseen targets.
progressreceives messages whenverboseis enabled.
- DeepMTP.data.interactions.check_target_variable_type(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, classification_mode: Literal['binary', 'multiclass'] | None = None, num_classes: int | None = None, progress: DataProgressObserver | None = None) Literal['binary', 'multiclass', 'real-valued']
Checks the type of the target variable in the interaction data. If any inconsistencies are detected (like different id types between train and test interaction data), an exception is raised
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose target-type validation messages. Defaults to None.
- DeepMTP.data.interactions.check_variable_type(samples_arr: DataFrame, *, classification_mode: Literal['binary', 'multiclass'] | None = None, num_classes: int | None = None) Literal['binary', 'multiclass', 'real-valued']
Detect binary or real-valued scores, or validate declared class IDs.
- Parameters:
samples_arr (numpy.array) – A numpy array with the target variables
- Returns:
binary,multiclass, orreal-valued.- Return type:
str
Multiclass targets are intentionally opt-in because an integer-valued regression target is otherwise indistinguishable from class IDs.
- DeepMTP.data.interactions.get_estimated_validation_setting(novel_instances_flag: bool | None, novel_targets_flag: bool | None, verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) Literal['A', 'B', 'C', 'D'] | None
Uses the combination of infrmation about novel instances and novel targets to determine the validation setting that is possible.
- Parameters:
novel_instances_flag (bool) – A boolean that indicates whether or not the test set contains novel instances.
novel_targets_flag (bool) – A boolean that indicates whether or not the test set contains novel targets.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose validation-setting inference messages. Defaults to None.
- Returns:
The validation setting. Possible values are A, B, C, D
- Return type:
str
- DeepMTP.data.interactions.process_interaction_data(interaction_data: DataFrame | ndarray, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) InteractionInfo
A function that processes the interaction data. It is called separately for the train, val and test interaction data. There are two main types of interaction data formats that are supported: -> numpy format: This is a 2d numpy array that represents the interaction (also called score) matrix usually found in problems settings with fully observed matrices (multi-label classification, multivariate regression) -> triplet format: The most flexible format as it can be used to represent every possible problem setting
- Parameters:
interaction_data (_type_) – a numpy array or a dataframe with the interaction data
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose interaction format messages. Defaults to None.
- Returns:
a dictionary with the interaction data and additional information that could be detected (format, type of instance and target ids, etc.)
- Return type:
dict
Features
Feature normalization and cross-input consistency validation.
- class DeepMTP.data.features.CompositeFeatureSpec
Bases:
TypedDictData-preparation options for one named composite component.
- kind: Literal['graph', 'sequence'] | None
- optional: bool
- class DeepMTP.data.features.FeatureComponentInfo
Bases:
_FeatureComponentInfoRequiredRepresentation metadata for one component of a composite feature.
- graph_edge_features: int | None
- info: Literal['numpy', 'dataframe', 'graph', 'sequence', 'scipy_sparse', 'torch_sparse']
- num_features: int | None
- class DeepMTP.data.features.FeatureInfo
Bases:
_FeatureInfoRequiredNormalized entity features and their detected representation.
- components: dict[str, FeatureComponentInfo] | None
- data: DataFrame
- graph_edge_features: int | None
- info: Literal['numpy', 'dataframe', 'images', 'graph', 'sequence', 'tabular', 'scipy_sparse', 'torch_sparse', 'composite']
- num_features: int | None
- DeepMTP.data.features.cross_input_consistency_check_instances(data: Mapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
Checks the consistency of instance ids in the interaction data and instance features. The requirements to pass this check change depending on the format of the interaction data and instance features
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
validation_setting (str) – The validation setting of the current problem.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose instance-consistency messages. Defaults to None.
- DeepMTP.data.features.cross_input_consistency_check_targets(data: Mapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
Checks the consistency of target ids in the interaction data and target features. The requirements to pass this check change depending on the format of the interaction data and target features
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
validation_setting (str) – The validation setting of the current problem.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose target-consistency messages. Defaults to None.
- DeepMTP.data.features.process_instance_features(instance_features: Any, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None, feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None) FeatureInfo | None
Normalize instance features without mutating caller-owned inputs.
Composite component specifications may set
optional=Trueto align a subset of entity IDs and preserve missing rows asNone.progressreceives messages whenverboseis enabled.
- DeepMTP.data.features.process_target_features(target_features: Any, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None, feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None) FeatureInfo | None
Normalize target features without mutating caller-owned inputs.
Composite component specifications may set
optional=Trueto align a subset of entity IDs and preserve missing rows asNone.progressreceives messages whenverboseis enabled.
Graph utilities
PyTorch Geometric is imported only when graph inputs are used. Its Data
objects are validated, cloned, and converted to DeepMTP’s graph batch
contract.
Optional PyTorch Geometric normalization and batching helpers.
- class DeepMTP.data.graphs.CollatedGraph(node_features: torch.Tensor, edge_index: torch.Tensor, edge_features: torch.Tensor | None, graph_index: torch.Tensor, num_graphs: int)
Bases:
objectFramework-neutral tensors extracted from one PyG graph batch.
- edge_features: torch.Tensor | None
- edge_index: torch.Tensor
- graph_index: torch.Tensor
- node_features: torch.Tensor
- num_graphs: int
- DeepMTP.data.graphs.collate_graph_rows(rows: Sequence[object]) CollatedGraph
Batch graph rows with PyG and expose only DeepMTP’s tensor contract.
- DeepMTP.data.graphs.copy_graph_value(value: object) Any
Clone one validated graph without sharing its tensors.
- DeepMTP.data.graphs.graph_object_array(values: Sequence[object]) ndarray
Store graph objects without invoking third-party array protocols.
- DeepMTP.data.graphs.is_graph_value(value: object) bool
Return whether
valueis a PyTorch Geometric graph object.
- DeepMTP.data.graphs.normalize_graph_rows(rows: Iterable[object], *, entity_name: str) tuple[list[Any], int, int | None]
Validate homogeneous PyG graphs and return independent normalized copies.
- DeepMTP.data.graphs.require_pyg_data() ModuleType
Return PyG’s data module or explain how to install graph support.
Sparse utilities
The sparse input adapter resolves SciPy only when inspecting a SciPy-backed value.
Sparse feature normalization and batching utilities.
- DeepMTP.data.sparse.collate_sparse_rows(rows: Sequence[object]) torch.Tensor
Stack sparse rows into one floating-point PyTorch COO batch.
- DeepMTP.data.sparse.copy_feature_value(value: object) Any
Copy one dense or sparse feature value without densifying it.
- DeepMTP.data.sparse.is_scipy_sparse(value: object) bool
Return whether
valueis a SciPy sparse matrix.
- DeepMTP.data.sparse.is_sparse_value(value: object) bool
Return whether
valueuses a supported sparse representation.
- DeepMTP.data.sparse.is_torch_sparse(value: object) bool
Return whether
valueis a supported PyTorch sparse tensor.
- DeepMTP.data.sparse.normalize_sparse_matrix(matrix: object, *, entity_name: str) tuple[list[Any], int, Literal['scipy', 'torch']]
Validate a 2-D sparse matrix and return independent sparse rows.
- DeepMTP.data.sparse.normalize_sparse_rows(rows: Iterable[object], *, entity_name: str) tuple[list[Any], int, Literal['scipy', 'torch']]
Validate and copy sparse feature rows with one consistent width/backend.
- DeepMTP.data.sparse.pin_sparse_tensor(values: torch.Tensor) torch.Tensor
Pin sparse tensor components because
Tensor.pin_memorylacks support.
- DeepMTP.data.sparse.sparse_backend(value: object) Literal['scipy', 'torch']
Return the sparse backend used by
value.
- DeepMTP.data.sparse.sparse_object_array(values: Sequence[object]) ndarray
Return an object array without invoking sparse tensor array protocols.
Splitting
Dataset splitting, defensive copying, and result validation.
- DeepMTP.data.splitting.split_data(data: MutableMapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], split_method: str, ratio: object, shuffle: bool, seed: int | None, verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
- Splits the dataset and offers two main functionalities:
split based on the 4 different validation settings (A, B, C, D)
if a test set already exists it separates a validation set, otherwise it first creates a test set and then a validation set.
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
validation_setting (str) – The validation setting of the current problem.
split_method (str) – The splitting method used. The current implementation only supports the ‘random split’ using a specific seed but a future goal is to also offer a stratified option.
ratio (dict) – The train, val and test ratios used to split the data.
seed (int) – The seed used to initiate the randomized split
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose split lifecycle messages. Defaults to None.
Preparation
End-to-end data preparation and feature scaling orchestration.
- class DeepMTP.data.preparation.DataInfo
Bases:
TypedDictMetadata inferred while preparing a multi-target dataset.
- data_preparation_state: dict[str, Any]
- detected_classification_mode: Literal['binary', 'multiclass'] | None
- detected_num_classes: int | None
- detected_problem_mode: Literal['classification', 'regression']
- detected_validation_setting: Literal['A', 'B', 'C', 'D']
- instance_branch_component_graph_edge_dims: dict[str, int | None] | None
- instance_branch_component_input_dims: dict[str, int | None] | None
- instance_branch_graph_edge_dim: int | None
- instance_branch_input_dim: int | None
- target_branch_component_graph_edge_dims: dict[str, int | None] | None
- target_branch_component_input_dims: dict[str, int | None] | None
- target_branch_graph_edge_dim: int | None
- target_branch_input_dim: int | None
- class DeepMTP.data.preparation.Transformer(*args, **kwargs)
Bases:
ProtocolMinimal interface implemented by supported feature scalers.
- transform(values: Any) Any
Transform values using a previously fitted scaler.
- DeepMTP.data.preparation.data_process(data: Mapping[str, Any], validation_setting: str | None = None, split_method: str = 'random', ratio: object = None, shuffle: bool = True, seed: int | None = 42, verbose: bool = False, print_mode: str = 'basic', scale_instance_features: str | None = None, scale_target_features: str | None = None, *, classification_mode: str | None = None, num_classes: int | None = None, instance_feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None, target_feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None, progress: DataProgressObserver | None = None, dense_preprocessing_state: Mapping[str, Any] | DataPreparationState | DensePreprocessingState | None = None) tuple[dict[str, Any], dict[str, Any], dict[str, Any], DataInfo]
The main function that handles all the preprocessing steps and checks needed to prepare the dataset to be used by the model.
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
validation_setting (str, optional) – The validation setting of the current problem. Defaults to None.
split_method (str, optional) – The splitting method used. The current implementation only supports the ‘random split’ using a specific seed but a future goal is to also offer a stratified option. Defaults to ‘random’.
ratio (dict, optional) – The train, val and test ratios used to split the data. Defaults to {‘train’: 0.7, ‘test’: 0.2, ‘val’: 0.1}.
shuffle (bool, optional) – Whether or not the dataset will be shuffled before the split. If is set to False, the seed value is not used.
seed (int, optional) – The seed used to initiate the randomized split. Defaults to 42.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
scale_instance_features (str, optional) – The scaler used for the instance features. Possible values are ‘MinMax’ for the MinMax scaler and ‘Standard’ for the standard scaler. Defaults to None.
scale_target_features (str, optional) – The scaler used for the target features. Possible values are ‘MinMax’ for the MinMax scaler and ‘Standard’ for the standard scaler. Defaults to None.
classification_mode (str, optional) – Set to
multiclassto interpret zero-based integer scores as class IDs. Binary scores continue to be detected automatically. Defaults to None.num_classes (int, optional) – Declared multiclass count. When omitted in multiclass mode, it is inferred from all supplied score splits.
instance_feature_kind (str, optional) – Set to
sequencewhen instance features are already-tokenized integer sequences. For composite features, pass a mapping from component name tosequence,None, or a mapping withkindandoptionalfields. Defaults to None.target_feature_kind (str, optional) – Set to
sequencewhen target features are already-tokenized integer sequences. For composite features, pass a mapping from component name tosequence,None, or a mapping withkindandoptionalfields. Defaults to None.progress (DataProgressObserver | None, optional) – Receives verbose messages from the complete data-processing pipeline. Defaults to None.
dense_preprocessing_state (mapping, optional) – Previously saved dense scaler parameters. Accepts either the
densesection or the fulldata_preparation_statereturned indata_info. When supplied, the saved training statistics are replayed instead of fitting new scalers. Defaults to None.
- Returns:
- Four different dictionaries containing:
Train processed data
Validation processed data
Test processed data
general information about the datasets
- Return type:
dict, dict, dict, dict
- DeepMTP.data.preparation.normalize(row: Any, scaler: Transformer) Any
Just normalizes a row or features
- Parameters:
row (numpy.array) – an array of features
scaler (sklearn.scaler) – the scaler that will be used to scale the features
- Returns:
a scaled array of feature
- Return type:
numpy.array
- DeepMTP.data.preparation.validate_split_ratio(ratio: object) dict[str, float]
Validate and copy a train/validation/test split ratio mapping.
Preprocessing metadata
Legacy dense scaling parameters and split provenance are represented by validated, JSON-compatible state objects that can be replayed without a fitted scikit-learn object.
Serializable preprocessing state shared by data preparation and checkpoints.
- class DeepMTP.data.preprocessing.DataPreparationState(validation_setting: str, split_method: str, split_ratio: tuple[tuple[str, float], ...], shuffle: bool, seed: int | None, dense: DensePreprocessingState)
Bases:
objectSplit provenance and fitted dense preprocessing from
data_process.- dense: DensePreprocessingState
- classmethod from_config(value: object) DataPreparationState
Validate a saved data-preparation state mapping.
- seed: int | None
- shuffle: bool
- split_method: str
- split_ratio: tuple[tuple[str, float], ...]
- to_dict() dict[str, Any]
Return JSON-compatible checkpoint metadata.
- validation_setting: str
- class DeepMTP.data.preprocessing.DensePreprocessingState(instance: DenseScalingState | None = None, target: DenseScalingState | None = None)
Bases:
objectDense scaling state for the instance and target feature axes.
- for_branch(branch: str) DenseScalingState | None
Return the scaler state for one entity axis.
- classmethod from_config(value: object) DensePreprocessingState
Validate a combined dense-preprocessing state mapping.
- instance: DenseScalingState | None = None
- target: DenseScalingState | None = None
- to_dict() dict[str, Any]
Return JSON-compatible checkpoint metadata.
- class DeepMTP.data.preprocessing.DenseScalingState(method: Literal['MinMax', 'Standard'], feature_count: int, offsets: tuple[float, ...], scales: tuple[float, ...], data_minimums: tuple[float, ...] = (), data_maximums: tuple[float, ...] = (), feature_range: tuple[float, float] = (0.0, 1.0))
Bases:
objectParameters needed to replay one fitted dense-feature transformation.
- data_maximums: tuple[float, ...] = ()
- data_minimums: tuple[float, ...] = ()
- feature_count: int
- feature_range: tuple[float, float] = (0.0, 1.0)
- classmethod from_config(value: object) DenseScalingState
Validate and restore a JSON-compatible scaling-state mapping.
- classmethod from_fitted_scaler(scaler: object, *, method: Literal['MinMax', 'Standard']) DenseScalingState
Capture explicit numeric state from a fitted scikit-learn scaler.
- method: Literal['MinMax', 'Standard']
- offsets: tuple[float, ...]
- scales: tuple[float, ...]
- to_dict() dict[str, Any]
Return JSON-compatible checkpoint metadata.
- transform(values: object) ndarray
Transform one vector or matrix using the saved training statistics.
- DeepMTP.data.preprocessing.transform_dense_features(values: object, state: Mapping[str, Any] | DenseScalingState) ndarray
Replay a saved dense-feature transformation.
Loading
Dataset and dataloader construction for DeepMTP experiments.
- class DeepMTP.data.loading.BaseDataset(*args: Any, **kwargs: Any)
Bases:
DatasetResolve interaction rows to model-ready entity features.
- class DeepMTP.data.loading.DataLoaderFactory(config: DeepMTPConfig | Mapping[str, Any], device: torch.device | str)
Bases:
objectBuild configured dataloaders without trainer-specific orchestration.
- for_prediction(data: Mapping[str, Any]) torch.utils.data.DataLoader
Build a deterministic inference dataloader.
- for_training(train_data: Mapping[str, Any], validation_data: Mapping[str, Any] | None, test_data: Mapping[str, Any] | None) TrainingDataLoaders
Build the train, validation, and test dataloaders.
- class DeepMTP.data.loading.DatasetItem
Bases:
TypedDictOne model-ready interaction returned by
BaseDataset.- instance_features: Any
- instance_id: int
- score: Any
- target_features: Any
- target_id: int
- class DeepMTP.data.loading.MTPSampler(*args: Any, **kwargs: Any)
Bases:
SamplerUndersample classes and provide a fresh order for every iteration.
microbalances classes globally,macrobalances them independently for each target, andinstancebalances them independently for each instance. Grouped balancing can omit entities that are absent from the retained observations.- get_balanced_df(data: DataFrame) DataFrame
Return one frame undersampled to its smallest class size.
- class DeepMTP.data.loading.TrainingDataLoaders(train: torch.utils.data.DataLoader, validation: torch.utils.data.DataLoader, test: torch.utils.data.DataLoader)
Bases:
objectDataloaders required by one training run.
- test: torch.utils.data.DataLoader
- train: torch.utils.data.DataLoader
- validation: torch.utils.data.DataLoader
Progress
Progress messages shared by data preparation and validation.
- class DeepMTP.data.progress.ConsoleDataProgressObserver(print_mode: Literal['basic', 'dev'] = 'basic')
Bases:
objectRender data progress using the historical basic or developer format.
- on_event(event: DataProgressEvent) None
- class DeepMTP.data.progress.DataProgressEvent(message: str, kind: Literal['message', 'operation_started', 'operation_completed'] = 'message', subject: str | None = None, level: Literal['info', 'warning', 'error'] = 'info', end: str = '\n', value: object | None = None)
Bases:
objectOne user-facing data preparation or validation message.
- end: str = '\n'
- kind: Literal['message', 'operation_started', 'operation_completed'] = 'message'
- level: Literal['info', 'warning', 'error'] = 'info'
- message: str
- subject: str | None = None
- value: object | None = None
- class DeepMTP.data.progress.DataProgressObserver(*args, **kwargs)
Bases:
ProtocolReceives data preparation and validation messages.
- on_event(event: DataProgressEvent) None
Handle one data progress event.
- class DeepMTP.data.progress.NullDataProgressObserver
Bases:
objectNo-op observer used when data progress is disabled.
- on_event(event: DataProgressEvent) None
- DeepMTP.data.progress.build_data_progress_observer(enabled: bool, *, print_mode: Literal['basic', 'dev'] = 'basic') DataProgressObserver
Select the default data progress observer.
Package contents
Data loading, preparation, validation, and progress services.
- class DeepMTP.data.BaseDataset(*args: Any, **kwargs: Any)
Bases:
DatasetResolve interaction rows to model-ready entity features.
- class DeepMTP.data.CategoricalColumn(name: str, categories: tuple[str | int, ...], embedding_dim: int)
Bases:
objectOne categorical feature and its fixed training vocabulary.
- categories: tuple[str | int, ...]
- embedding_dim: int
- classmethod from_config(name: object, value: object) CategoricalColumn
Validate one categorical-column declaration.
- name: str
- to_dict() dict[str, Any]
Return a JSON-compatible public configuration mapping.
- property vocabulary_size: int
Embedding-table size, including index zero for unknown values.
- class DeepMTP.data.CompositeBranchBatch(values: CompositeInput)
Bases:
objectA named collection of inputs consumed by a composite encoder.
- property batch_size: int
Number of observations shared by every component.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'composite'
- pin_memory() CompositeBranchBatch
Pin every component in the structured input.
- to(device: torch.device | str, *, non_blocking: bool = False) CompositeBranchBatch
Move every component to one device.
- values: CompositeInput
- class DeepMTP.data.CompositeFeatureSpec
Bases:
TypedDictData-preparation options for one named composite component.
- kind: Literal['graph', 'sequence'] | None
- optional: bool
- class DeepMTP.data.CompositeInput(components: Mapping[str, Any])
Bases:
Mapping[str,Any]Named model inputs for multiple encoders on one entity axis.
- property batch_size: int
Number of observations shared by every component.
- components: Mapping[str, Any]
- pin_memory() CompositeInput
Pin every component for asynchronous host-to-device transfer.
- to(device: torch.device | str, *, non_blocking: bool = False) CompositeInput
Move every component to one device.
- class DeepMTP.data.ConsoleDataProgressObserver(print_mode: Literal['basic', 'dev'] = 'basic')
Bases:
objectRender data progress using the historical basic or developer format.
- on_event(event: DataProgressEvent) None
- class DeepMTP.data.CustomBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA tensor batch passed unchanged to a user-supplied branch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'custom'
- values: torch.Tensor
- class DeepMTP.data.DataInfo
Bases:
TypedDictMetadata inferred while preparing a multi-target dataset.
- data_preparation_state: dict[str, Any]
- detected_classification_mode: Literal['binary', 'multiclass'] | None
- detected_num_classes: int | None
- detected_problem_mode: Literal['classification', 'regression']
- detected_validation_setting: Literal['A', 'B', 'C', 'D']
- instance_branch_component_graph_edge_dims: dict[str, int | None] | None
- instance_branch_component_input_dims: dict[str, int | None] | None
- instance_branch_graph_edge_dim: int | None
- instance_branch_input_dim: int | None
- target_branch_component_graph_edge_dims: dict[str, int | None] | None
- target_branch_component_input_dims: dict[str, int | None] | None
- target_branch_graph_edge_dim: int | None
- target_branch_input_dim: int | None
- class DeepMTP.data.DataLoaderFactory(config: DeepMTPConfig | Mapping[str, Any], device: torch.device | str)
Bases:
objectBuild configured dataloaders without trainer-specific orchestration.
- for_prediction(data: Mapping[str, Any]) torch.utils.data.DataLoader
Build a deterministic inference dataloader.
- for_training(train_data: Mapping[str, Any], validation_data: Mapping[str, Any] | None, test_data: Mapping[str, Any] | None) TrainingDataLoaders
Build the train, validation, and test dataloaders.
- class DeepMTP.data.DataPreparationState(validation_setting: str, split_method: str, split_ratio: tuple[tuple[str, float], ...], shuffle: bool, seed: int | None, dense: DensePreprocessingState)
Bases:
objectSplit provenance and fitted dense preprocessing from
data_process.- dense: DensePreprocessingState
- classmethod from_config(value: object) DataPreparationState
Validate a saved data-preparation state mapping.
- seed: int | None
- shuffle: bool
- split_method: str
- split_ratio: tuple[tuple[str, float], ...]
- to_dict() dict[str, Any]
Return JSON-compatible checkpoint metadata.
- validation_setting: str
- class DeepMTP.data.DataProgressEvent(message: str, kind: Literal['message', 'operation_started', 'operation_completed'] = 'message', subject: str | None = None, level: Literal['info', 'warning', 'error'] = 'info', end: str = '\n', value: object | None = None)
Bases:
objectOne user-facing data preparation or validation message.
- end: str = '\n'
- kind: Literal['message', 'operation_started', 'operation_completed'] = 'message'
- level: Literal['info', 'warning', 'error'] = 'info'
- message: str
- subject: str | None = None
- value: object | None = None
- class DeepMTP.data.DataProgressObserver(*args, **kwargs)
Bases:
ProtocolReceives data preparation and validation messages.
- on_event(event: DataProgressEvent) None
Handle one data progress event.
- class DeepMTP.data.DatasetBundle
Bases:
TypedDictStable outer structure returned by legacy dataset helpers.
- test: DatasetSplit
- train: DatasetSplit
- val: DatasetSplit
- class DeepMTP.data.DatasetItem
Bases:
TypedDictOne model-ready interaction returned by
BaseDataset.- instance_features: Any
- instance_id: int
- score: Any
- target_features: Any
- target_id: int
- class DeepMTP.data.DatasetSplit
Bases:
TypedDictOne raw train, validation, or test split.
- X_instance: ndarray | DataFrame | None
- X_target: ndarray | DataFrame | None
- y: ndarray | DataFrame | None
- class DeepMTP.data.DenseBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA batch of dense feature vectors shaped
[batch, features].- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'dense'
- values: torch.Tensor
- class DeepMTP.data.DensePreprocessingState(instance: DenseScalingState | None = None, target: DenseScalingState | None = None)
Bases:
objectDense scaling state for the instance and target feature axes.
- for_branch(branch: str) DenseScalingState | None
Return the scaler state for one entity axis.
- classmethod from_config(value: object) DensePreprocessingState
Validate a combined dense-preprocessing state mapping.
- instance: DenseScalingState | None = None
- target: DenseScalingState | None = None
- to_dict() dict[str, Any]
Return JSON-compatible checkpoint metadata.
- class DeepMTP.data.DenseScalingState(method: Literal['MinMax', 'Standard'], feature_count: int, offsets: tuple[float, ...], scales: tuple[float, ...], data_minimums: tuple[float, ...] = (), data_maximums: tuple[float, ...] = (), feature_range: tuple[float, float] = (0.0, 1.0))
Bases:
objectParameters needed to replay one fitted dense-feature transformation.
- data_maximums: tuple[float, ...] = ()
- data_minimums: tuple[float, ...] = ()
- feature_count: int
- feature_range: tuple[float, float] = (0.0, 1.0)
- classmethod from_config(value: object) DenseScalingState
Validate and restore a JSON-compatible scaling-state mapping.
- classmethod from_fitted_scaler(scaler: object, *, method: Literal['MinMax', 'Standard']) DenseScalingState
Capture explicit numeric state from a fitted scikit-learn scaler.
- method: Literal['MinMax', 'Standard']
- offsets: tuple[float, ...]
- scales: tuple[float, ...]
- to_dict() dict[str, Any]
Return JSON-compatible checkpoint metadata.
- transform(values: object) ndarray
Transform one vector or matrix using the saved training statistics.
- class DeepMTP.data.FeatureComponentInfo
Bases:
_FeatureComponentInfoRequiredRepresentation metadata for one component of a composite feature.
- graph_edge_features: int | None
- info: Literal['numpy', 'dataframe', 'graph', 'sequence', 'scipy_sparse', 'torch_sparse']
- num_features: int | None
- class DeepMTP.data.FeatureInfo
Bases:
_FeatureInfoRequiredNormalized entity features and their detected representation.
- components: dict[str, FeatureComponentInfo] | None
- data: DataFrame
- graph_edge_features: int | None
- info: Literal['numpy', 'dataframe', 'images', 'graph', 'sequence', 'tabular', 'scipy_sparse', 'torch_sparse', 'composite']
- num_features: int | None
- class DeepMTP.data.GraphBranchBatch(values: GraphInput)
Bases:
objectA structured batch consumed by a graph neural network encoder.
- property batch_size: int
Number of graphs represented by this branch batch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'graph'
- pin_memory() GraphBranchBatch
Pin every tensor in the graph input.
- to(device: torch.device | str, *, non_blocking: bool = False) GraphBranchBatch
Move the graph input to one device.
- values: GraphInput
- class DeepMTP.data.GraphInput(node_features: torch.Tensor, edge_index: torch.Tensor, edge_features: torch.Tensor | None, graph_index: torch.Tensor, num_graphs: int)
Bases:
objectOne batch of homogeneous graphs represented by DeepMTP tensors.
- property batch_size: int
Number of graphs represented by this input.
- edge_features: torch.Tensor | None
- edge_index: torch.Tensor
- graph_index: torch.Tensor
- node_features: torch.Tensor
- num_graphs: int
- pin_memory() GraphInput
Pin every graph tensor for asynchronous device transfer.
- to(device: torch.device | str, *, non_blocking: bool = False) GraphInput
Move every graph tensor to one device.
- class DeepMTP.data.IDBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA batch of zero-based integer entity IDs shaped
[batch].- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'id'
- values: torch.Tensor
- class DeepMTP.data.ImageBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA batch of images shaped
[batch, channels, height, width].- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'image'
- values: torch.Tensor
- class DeepMTP.data.InteractionInfo
Bases:
TypedDictNormalized interaction data and its detected schema.
- data: DataFrame
- instance_id_type: Literal['int']
- missing_values: bool
- original_format: Literal['triplets', 'numpy']
- target_id_type: Literal['int']
- class DeepMTP.data.MTPBatch(instance_id: torch.Tensor, target_id: torch.Tensor, instance_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch, target_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch, score: torch.Tensor)
Bases:
Mapping[str,Any]A complete interaction batch with mapping compatibility.
Attribute access exposes the typed branch inputs. Historical mapping keys continue to return the underlying tensors so existing dataloader consumers remain compatible.
- property batch_size: int
Number of interactions represented by the batch.
- property instance_features: torch.Tensor | CompositeInput | GraphInput | SequenceInput | TabularInput
Compatibility view of the instance branch input.
- instance_id: torch.Tensor
- instance_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch
- pin_memory() MTPBatch
Pin every tensor for asynchronous host-to-device transfer.
- score: torch.Tensor
- property target_features: torch.Tensor | CompositeInput | GraphInput | SequenceInput | TabularInput
Compatibility view of the target branch input.
- target_id: torch.Tensor
- target_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch
- to(device: torch.device | str, *, non_blocking: bool = False) MTPBatch
Move every tensor in the batch to one device.
- class DeepMTP.data.MTPBatchCollator(instance_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite'], target_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite'], instance_sequence_padding_idx: int = 0, target_sequence_padding_idx: int = 0, instance_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None, target_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None, instance_component_padding_indices: Mapping[str, int] | None = None, target_component_padding_indices: Mapping[str, int] | None = None, instance_component_optional: Mapping[str, bool] | None = None, target_component_optional: Mapping[str, bool] | None = None)
Bases:
objectCollate dataset items into a typed interaction batch.
- instance_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None
- instance_component_optional: Mapping[str, bool] | None = None
- instance_component_padding_indices: Mapping[str, int] | None = None
- instance_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']
- instance_sequence_padding_idx: int = 0
- target_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None
- target_component_optional: Mapping[str, bool] | None = None
- target_component_padding_indices: Mapping[str, int] | None = None
- target_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']
- target_sequence_padding_idx: int = 0
- class DeepMTP.data.MTPSampler(*args: Any, **kwargs: Any)
Bases:
SamplerUndersample classes and provide a fresh order for every iteration.
microbalances classes globally,macrobalances them independently for each target, andinstancebalances them independently for each instance. Grouped balancing can omit entities that are absent from the retained observations.- get_balanced_df(data: DataFrame) DataFrame
Return one frame undersampled to its smallest class size.
- class DeepMTP.data.MaskedComponentInput(values: Any | None, presence_mask: torch.Tensor)
Bases:
objectPresent values plus a full-batch mask for one optional component.
- property batch_size: int
Full batch size including rows where this component is missing.
- pin_memory() MaskedComponentInput
Pin the present values and full-batch mask.
- presence_mask: torch.Tensor
- property present_count: int
Number of rows containing this component.
- to(device: torch.device | str, *, non_blocking: bool = False) MaskedComponentInput
Move the values and presence mask to one device.
- values: Any | None
- exception DeepMTP.data.NoveltyInferenceWarning
Bases:
UserWarningWarning emitted when entity identity was lost in matrix input.
- class DeepMTP.data.NullDataProgressObserver
Bases:
objectNo-op observer used when data progress is disabled.
- on_event(event: DataProgressEvent) None
- class DeepMTP.data.ProcessedDatasetBundle
Bases:
TypedDictProcessed train, validation, and test data.
- test: ProcessedDatasetSplit
- train: ProcessedDatasetSplit
- class DeepMTP.data.ProcessedDatasetSplit
Bases:
TypedDictOne processed split with optional interactions and side features.
- X_instance: ProcessedDatasetValue | None
- X_target: ProcessedDatasetValue | None
- y: ProcessedDatasetValue | None
- class DeepMTP.data.ProcessedDatasetValue
Bases:
TypedDictProcessed data container consumed by legacy data utilities.
- data: DataFrame
- class DeepMTP.data.SequenceBranchBatch(values: SequenceInput)
Bases:
objectA structured batch consumed by a token-sequence encoder.
- property batch_size: int
Number of observations represented by this branch batch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'sequence'
- pin_memory() SequenceBranchBatch
Pin every tensor in the structured sequence input.
- to(device: torch.device | str, *, non_blocking: bool = False) SequenceBranchBatch
Move the structured sequence input to one device.
- values: SequenceInput
- class DeepMTP.data.SequenceInput(token_ids: torch.Tensor, attention_mask: torch.Tensor, lengths: torch.Tensor, padding_idx: int = 0)
Bases:
objectPadded token IDs, attention mask, and original sequence lengths.
- attention_mask: torch.Tensor
- property batch_size: int
Number of token sequences represented by this input.
- lengths: torch.Tensor
- padding_idx: int = 0
- pin_memory() SequenceInput
Pin every sequence tensor for asynchronous device transfer.
- to(device: torch.device | str, *, non_blocking: bool = False) SequenceInput
Move every sequence tensor to one device.
- token_ids: torch.Tensor
- class DeepMTP.data.SparseBranchBatch(values: torch.Tensor)
Bases:
TensorBranchBatchA sparse feature batch shaped
[batch, features].- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'sparse'
- pin_memory() SparseBranchBatch
Pin sparse indices and values for asynchronous device transfer.
- values: torch.Tensor
- class DeepMTP.data.TabularBranchBatch(values: TabularInput)
Bases:
objectA structured batch of numeric values and categorical indices.
- property batch_size: int
Number of observations represented by this branch batch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'tabular'
- pin_memory() TabularBranchBatch
Pin both tensors in the structured input.
- to(device: torch.device | str, *, non_blocking: bool = False) TabularBranchBatch
Move the structured input to one device.
- values: TabularInput
- class DeepMTP.data.TabularInput(numeric: torch.Tensor, categorical: torch.Tensor)
Bases:
objectNumeric and categorical tensors consumed by a tabular encoder.
- property batch_size: int
Number of observations in the structured input.
- categorical: torch.Tensor
- numeric: torch.Tensor
- pin_memory() TabularInput
Pin both tensors for asynchronous host-to-device transfer.
- to(device: torch.device | str, *, non_blocking: bool = False) TabularInput
Move both tensors to one device.
- class DeepMTP.data.TabularPreprocessingState(numeric_columns: tuple[str, ...], categorical_columns: tuple[str, ...], numeric_fill_values: tuple[float, ...], numeric_offsets: tuple[float, ...], numeric_scales: tuple[float, ...])
Bases:
objectNumeric statistics fitted only on training entities.
- categorical_columns: tuple[str, ...]
- classmethod from_config(value: object, *, schema: TabularSchema, branch: str) TabularPreprocessingState
Restore and validate preprocessing metadata from a checkpoint config.
- numeric_columns: tuple[str, ...]
- numeric_fill_values: tuple[float, ...]
- numeric_offsets: tuple[float, ...]
- numeric_scales: tuple[float, ...]
- to_dict() dict[str, Any]
Return JSON-compatible checkpoint metadata.
- class DeepMTP.data.TabularPreprocessor(schema: TabularSchema, state: TabularPreprocessingState, branch: str)
Bases:
objectTransform entity feature frames using one schema and fitted state.
- branch: str
- classmethod fit(frame: DataFrame, schema: TabularSchema, *, branch: str) TabularPreprocessor
Fit numeric imputation and normalization on a training frame.
- classmethod from_state(schema: TabularSchema, state: Mapping[str, Any] | TabularPreprocessingState, *, branch: str) TabularPreprocessor
Build a preprocessor from saved training metadata.
- schema: TabularSchema
- state: TabularPreprocessingState
- transform(frame: DataFrame) DataFrame
Return an indexed feature frame with structured tensor-ready values.
- class DeepMTP.data.TabularSchema(numeric_columns: tuple[str, ...], categorical_columns: tuple[CategoricalColumn, ...], numeric_normalization: Literal['none', 'standard', 'minmax'] = 'standard', numeric_missing: Literal['error', 'mean', 'zero'] = 'mean', categorical_missing: Literal['error', 'unknown'] = 'unknown', categorical_unknown: Literal['error', 'unknown'] = 'unknown', feature_gating: bool = False)
Bases:
objectExplicit numeric and categorical layout for one entity branch.
- categorical_columns: tuple[CategoricalColumn, ...]
- categorical_missing: Literal['error', 'unknown'] = 'unknown'
- property categorical_names: tuple[str, ...]
Categorical column names in their stable tensor order.
- categorical_unknown: Literal['error', 'unknown'] = 'unknown'
- property encoded_width: int
Width after concatenating numeric values and category embeddings.
- feature_gating: bool = False
- classmethod from_config(value: object, *, branch: str) TabularSchema
Validate a public tabular schema mapping.
- numeric_columns: tuple[str, ...]
- numeric_missing: Literal['error', 'mean', 'zero'] = 'mean'
- numeric_normalization: Literal['none', 'standard', 'minmax'] = 'standard'
- to_dict() dict[str, Any]
Return a JSON-compatible public configuration mapping.
- class DeepMTP.data.TensorBranchBatch(values: torch.Tensor)
Bases:
objectOne branch’s tensor payload with an explicit modality contract.
- property batch_size: int
Number of observations represented by this branch batch.
- kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'custom'
- pin_memory() TensorBranchBatch
Return the same typed batch backed by pinned host memory.
- to(device: torch.device | str, *, non_blocking: bool = False) TensorBranchBatch
Return the same typed batch with its tensor moved to
device.
- values: torch.Tensor
- class DeepMTP.data.TrainingDataLoaders(train: torch.utils.data.DataLoader, validation: torch.utils.data.DataLoader, test: torch.utils.data.DataLoader)
Bases:
objectDataloaders required by one training run.
- test: torch.utils.data.DataLoader
- train: torch.utils.data.DataLoader
- validation: torch.utils.data.DataLoader
- class DeepMTP.data.Transformer(*args, **kwargs)
Bases:
ProtocolMinimal interface implemented by supported feature scalers.
- transform(values: Any) Any
Transform values using a previously fitted scaler.
- DeepMTP.data.build_data_progress_observer(enabled: bool, *, print_mode: Literal['basic', 'dev'] = 'basic') DataProgressObserver
Select the default data progress observer.
- DeepMTP.data.check_interaction_files_column_type_format(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
Checks the type of the instance and target ids in the interaction data. If any inconsistencies are detected (like different id types between train and test interaction data), an exception is raised
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose ID-type validation messages. Defaults to None.
- DeepMTP.data.check_interaction_files_format(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
Checks the format of the interaction data. If any inconsistencies are detected (like different formats between train and test interaction data), an exception is raised
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose format-validation messages. Defaults to None.
- DeepMTP.data.check_novel_instances(train: Mapping[str, Any] | None, test: Mapping[str, Any] | None, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) bool | None
Return whether test interactions contain previously unseen instances.
progressreceives messages whenverboseis enabled.
- DeepMTP.data.check_novel_targets(train: Mapping[str, Any] | None, test: Mapping[str, Any] | None, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) bool | None
Return whether test interactions contain previously unseen targets.
progressreceives messages whenverboseis enabled.
- DeepMTP.data.check_target_variable_type(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, classification_mode: Literal['binary', 'multiclass'] | None = None, num_classes: int | None = None, progress: DataProgressObserver | None = None) Literal['binary', 'multiclass', 'real-valued']
Checks the type of the target variable in the interaction data. If any inconsistencies are detected (like different id types between train and test interaction data), an exception is raised
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose target-type validation messages. Defaults to None.
- DeepMTP.data.check_variable_type(samples_arr: DataFrame, *, classification_mode: Literal['binary', 'multiclass'] | None = None, num_classes: int | None = None) Literal['binary', 'multiclass', 'real-valued']
Detect binary or real-valued scores, or validate declared class IDs.
- Parameters:
samples_arr (numpy.array) – A numpy array with the target variables
- Returns:
binary,multiclass, orreal-valued.- Return type:
str
Multiclass targets are intentionally opt-in because an integer-valued regression target is otherwise indistinguishable from class IDs.
- DeepMTP.data.cross_input_consistency_check_instances(data: Mapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
Checks the consistency of instance ids in the interaction data and instance features. The requirements to pass this check change depending on the format of the interaction data and instance features
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
validation_setting (str) – The validation setting of the current problem.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose instance-consistency messages. Defaults to None.
- DeepMTP.data.cross_input_consistency_check_targets(data: Mapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
Checks the consistency of target ids in the interaction data and target features. The requirements to pass this check change depending on the format of the interaction data and target features
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
validation_setting (str) – The validation setting of the current problem.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose target-consistency messages. Defaults to None.
- DeepMTP.data.data_process(data: Mapping[str, Any], validation_setting: str | None = None, split_method: str = 'random', ratio: object = None, shuffle: bool = True, seed: int | None = 42, verbose: bool = False, print_mode: str = 'basic', scale_instance_features: str | None = None, scale_target_features: str | None = None, *, classification_mode: str | None = None, num_classes: int | None = None, instance_feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None, target_feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None, progress: DataProgressObserver | None = None, dense_preprocessing_state: Mapping[str, Any] | DataPreparationState | DensePreprocessingState | None = None) tuple[dict[str, Any], dict[str, Any], dict[str, Any], DataInfo]
The main function that handles all the preprocessing steps and checks needed to prepare the dataset to be used by the model.
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
validation_setting (str, optional) – The validation setting of the current problem. Defaults to None.
split_method (str, optional) – The splitting method used. The current implementation only supports the ‘random split’ using a specific seed but a future goal is to also offer a stratified option. Defaults to ‘random’.
ratio (dict, optional) – The train, val and test ratios used to split the data. Defaults to {‘train’: 0.7, ‘test’: 0.2, ‘val’: 0.1}.
shuffle (bool, optional) – Whether or not the dataset will be shuffled before the split. If is set to False, the seed value is not used.
seed (int, optional) – The seed used to initiate the randomized split. Defaults to 42.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
scale_instance_features (str, optional) – The scaler used for the instance features. Possible values are ‘MinMax’ for the MinMax scaler and ‘Standard’ for the standard scaler. Defaults to None.
scale_target_features (str, optional) – The scaler used for the target features. Possible values are ‘MinMax’ for the MinMax scaler and ‘Standard’ for the standard scaler. Defaults to None.
classification_mode (str, optional) – Set to
multiclassto interpret zero-based integer scores as class IDs. Binary scores continue to be detected automatically. Defaults to None.num_classes (int, optional) – Declared multiclass count. When omitted in multiclass mode, it is inferred from all supplied score splits.
instance_feature_kind (str, optional) – Set to
sequencewhen instance features are already-tokenized integer sequences. For composite features, pass a mapping from component name tosequence,None, or a mapping withkindandoptionalfields. Defaults to None.target_feature_kind (str, optional) – Set to
sequencewhen target features are already-tokenized integer sequences. For composite features, pass a mapping from component name tosequence,None, or a mapping withkindandoptionalfields. Defaults to None.progress (DataProgressObserver | None, optional) – Receives verbose messages from the complete data-processing pipeline. Defaults to None.
dense_preprocessing_state (mapping, optional) – Previously saved dense scaler parameters. Accepts either the
densesection or the fulldata_preparation_statereturned indata_info. When supplied, the saved training statistics are replayed instead of fitting new scalers. Defaults to None.
- Returns:
- Four different dictionaries containing:
Train processed data
Validation processed data
Test processed data
general information about the datasets
- Return type:
dict, dict, dict, dict
- DeepMTP.data.format_mtr_datasets() str
Format the catalog of supported multivariate regression datasets.
- DeepMTP.data.generate_MTP_dataset(num_instances: int, num_targets: int, num_instance_features: int | None = None, num_target_features: int | None = None, split_instances: Mapping[str, float] | None = None, split_targets: Mapping[str, float] | None = None, return_static_features_data: bool = False) ProcessedDatasetBundle
Generate deterministic triplet data with optional axis-level splits.
Each split mapping requires a test ratio and may include a validation ratio. The train ratio is derived when omitted and verified when supplied. A test-only axis reuses its training IDs in a validation split created on the other axis.
- Parameters:
num_instances – Number of instances.
num_targets – Number of targets.
num_instance_features – Number of generated instance features, or none.
num_target_features – Number of generated target features, or none.
split_instances – Optional instance-axis split ratios.
split_targets – Optional target-axis split ratios.
return_static_features_data – Repeat non-novel side features in holdouts.
- Returns:
Processed triplet and feature frames for train, validation, and test.
- Raises:
TypeError – If a split or boolean flag has the wrong type.
ValueError – If a dimension or split ratio is invalid.
- DeepMTP.data.generate_dummy_dataset(num_instances: int, num_targets: int, num_instance_features: int, num_target_features: int, error_mu: float, error_sigma: float, sklearn_version: bool = False, seed: int = 42, mode: str = 'u+logv', split_ratio: Mapping[str, float] | None = None) DatasetBundle
Generate a reproducible dyadic regression dataset.
sklearn_versionstarts from normally distributed features. For modes involving logarithms or fractional powers, the relevant features are transformed to a positive domain so every generated score remains real. The error distribution applies to the four legacy noisy relationship modes.- Parameters:
num_instances – Number of instances.
num_targets – Number of targets.
num_instance_features – Number of features per instance.
num_target_features – Number of features per target.
error_mu – Mean of the Gaussian error added by noisy modes.
error_sigma – Non-negative standard deviation of that Gaussian error.
sklearn_version – Generate normally distributed rather than uniform features.
seed – Non-negative 32-bit seed used for generation and splitting.
mode – Mathematical relationship used to generate scores.
split_ratio – Train, validation, and test ratios.
- Returns:
A dataset bundle containing aligned train, validation, and test rows.
- Raises:
TypeError – If a boolean or error parameter has the wrong type.
ValueError – If a dimension, mode, seed, ratio, or generated value is invalid.
- DeepMTP.data.generate_interaction_matrix(input_path: str | PathLike[str], output_path: str | PathLike[str]) None
Convert annotator labels into an atomically written correctness matrix.
- DeepMTP.data.get_estimated_validation_setting(novel_instances_flag: bool | None, novel_targets_flag: bool | None, verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) Literal['A', 'B', 'C', 'D'] | None
Uses the combination of infrmation about novel instances and novel targets to determine the validation setting that is possible.
- Parameters:
novel_instances_flag (bool) – A boolean that indicates whether or not the test set contains novel instances.
novel_targets_flag (bool) – A boolean that indicates whether or not the test set contains novel targets.
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose validation-setting inference messages. Defaults to None.
- Returns:
The validation setting. Possible values are A, B, C, D
- Return type:
str
- DeepMTP.data.load_process_DP(path: str | PathLike[str] = './data', dataset_name: str = 'ern', variant: Literal['undivided', 'divided'] = 'undivided', random_state: int | None = 42, split_ratio: Mapping[str, float] | None = None, split_instance_features: bool = False, split_target_features: bool = False, validation_setting: str = 'B', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load a dyadic-prediction dataset and optionally create held-out splits.
- Parameters:
path – Directory used for the local dataset cache.
dataset_name – One of
"ern","srn","dpie", or"dpii".variant – Return the complete matrices or create train/validation/test splits.
random_state – Seed used by the split operations, or
None.split_ratio – Train, validation, and test proportions.
split_instance_features – Split instance features with interaction rows.
split_target_features – Split target features with interaction columns.
validation_setting – Novelty setting
"B","C", or"D".print_mode –
"basic"for ordinary messages or"dev"for messages prefixed for the development application.progress – Receives dataset loading and download status events.
- Raises:
ValueError – If an option or dataset matrix is invalid.
TypeError – If a feature-splitting flag is not boolean.
DatasetDownloadError – If a required dataset file cannot be downloaded.
FileNotFoundError – If a download does not create all required files.
- DeepMTP.data.load_process_MC(path: str | PathLike[str] = './data', dataset_name: str = 'ml-100k', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load a matrix-completion dataset from the MovieLens repository.
- Parameters:
path – Directory used for the local dataset cache.
dataset_name – Matrix-completion dataset to load.
print_mode –
"basic"for ordinary messages or"dev"for messages prefixed for the development application.progress – Receives dataset download status events.
- Raises:
ValueError – If an option or the ratings data is invalid.
DatasetDownloadError – If the archive cannot be downloaded or extracted.
FileNotFoundError – If the archive does not contain the expected ratings file.
- Returns:
The ratings triplets and empty side-feature and held-out splits.
- DeepMTP.data.load_process_MLC(path: str | PathLike[str] = './data', dataset_name: str = 'bibtex', variant: Literal['undivided', 'divided'] = 'undivided', features_type: Literal['numpy', 'dataframe'] = 'numpy', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load a cached or downloadable scikit-multilearn benchmark dataset.
- Parameters:
path – Directory used for the local dataset cache.
dataset_name – Supported multi-label benchmark name.
variant – Return the complete dataset or its published train/test split.
features_type – Return instance features as an array or feature DataFrame.
print_mode –
"basic"for ordinary messages or"dev"for messages prefixed for the development application.progress – Receives dataset loading and download status events.
- Raises:
ValueError – If an option or cached dataset is invalid.
DatasetDownloadError – If required cache files cannot be downloaded safely.
IsADirectoryError – If an expected cache file path is a directory.
- DeepMTP.data.load_process_MTL(path: str | PathLike[str] = './data', dataset_name: str = 'dog', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load multi-task learning datasets from my custom repository.
- Parameters:
path (str, optional) – The path where the datasets should be stored. If it doesn’t exist, download and store it in this directory. Defaults to ‘./data’.
dataset_name (str, optional) – The name of the multi-task learning dataset. Defaults to ‘dog’.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives dataset download status events. Defaults to None.
- Raises:
ValueError – If a public option or dataset file is invalid.
DatasetDownloadError – If the dataset archive cannot be downloaded safely.
FileNotFoundError – If required score or image files are missing.
- Returns:
A dictionary with all the available data for the multi-task learning dataset.
- Return type:
dict
- DeepMTP.data.load_process_MTR(path: str | PathLike[str] = './data', dataset_name: str = 'enb', features_type: Literal['numpy', 'dataframe'] = 'numpy', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) DatasetBundle
Load multivariate regression datasets from the mulan repository.
- Parameters:
path (str, optional) – The path where the datasets should be stored. If it doesn’t exist, download and store it in this directory. Defaults to ‘./data’.
dataset_name (str, optional) – The name of the multivariate regression dataset. Defaults to ‘enb’.
features_type (str, optional) – The format of the instance features. There are two possible values, numpy and dataframe. This is intended to test the functionality of the data_process function. Defaults to ‘numpy’.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives dataset loading and download status events. Defaults to None.
- Raises:
ValueError – If a public option or downloaded dataset is invalid.
DatasetDownloadError – If the dataset archive cannot be downloaded safely.
FileNotFoundError – If the archive does not contain the requested dataset.
- Returns:
A dictionary with all the available data for the multivariate regression dataset.
- Return type:
dict
- DeepMTP.data.normalize(row: Any, scaler: Transformer) Any
Just normalizes a row or features
- Parameters:
row (numpy.array) – an array of features
scaler (sklearn.scaler) – the scaler that will be used to scale the features
- Returns:
a scaled array of feature
- Return type:
numpy.array
- DeepMTP.data.print_MTR_datasets(*, progress: DataProgressObserver | None = None) None
Present the supported multivariate regression dataset catalog.
The default observer preserves the historical console table. Inject
NullDataProgressObserverto suppress output or another data progress observer to render the catalog elsewhere.
- DeepMTP.data.process_dummy_DP(num_instance_features: int = 10, num_target_features: int = 3, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', instance_features_format: Literal['numpy', 'dataframe'] = 'numpy', target_features_format: Literal['numpy', 'dataframe'] = 'numpy', random_state: int | None = 42) DatasetBundle
Generate a reproducible dummy dyadic prediction dataset.
- Parameters:
num_instance_features – Number of features per instance.
num_target_features – Number of features per target.
num_instances – Number of instances.
num_targets – Number of targets.
interaction_matrix_format – Return scores as a dense matrix or triplets.
instance_features_format – Return instance features as an array or DataFrame.
target_features_format – Return target features as an array or DataFrame.
random_state – Seed for an isolated NumPy generator. Use
Nonefor non-deterministic data.
- Returns:
A dataset bundle containing an undivided training split.
- Raises:
ValueError – If a format, dimension, or random state is invalid.
- DeepMTP.data.process_dummy_MLC(num_features: int = 10, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', features_format: Literal['numpy', 'dataframe'] = 'numpy', random_state: int | None = 42) DatasetBundle
Generate a reproducible dummy multi-label classification dataset.
- Parameters:
num_features – Number of features per instance.
num_instances – Number of instances.
num_targets – Number of binary labels.
interaction_matrix_format – Return labels as a dense matrix or triplets.
features_format – Return features as a dense matrix or feature DataFrame.
random_state – Seed for an isolated NumPy generator. Use
Nonefor non-deterministic data.
- Returns:
A dataset bundle containing an undivided training split.
- Raises:
ValueError – If a format, dimension, or random state is invalid.
- DeepMTP.data.process_dummy_MTR(num_features: int = 10, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', features_format: Literal['numpy', 'dataframe'] = 'numpy', variant: Literal['undivided', 'divided'] = 'undivided', split_ratio: Mapping[str, float] | None = None, random_state: int | None = 42) DatasetBundle
Generate a reproducible dummy multivariate regression dataset.
- Parameters:
num_features – Number of features per instance.
num_instances – Number of instances.
num_targets – Number of regression targets.
interaction_matrix_format – Return scores as a dense matrix or triplets.
features_format – Return features as a dense matrix or feature DataFrame.
variant – Return an undivided dataset or train/validation/test splits.
split_ratio – Ratios used for the divided variant.
random_state – Seed used for both data generation and splitting. Use
Nonefor non-deterministic data.
- Returns:
A dataset bundle containing the generated scores and features.
- Raises:
TypeError – If
split_ratiois not a mapping.ValueError – If an option, dimension, ratio, or random state is invalid.
- DeepMTP.data.process_instance_features(instance_features: Any, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None, feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None) FeatureInfo | None
Normalize instance features without mutating caller-owned inputs.
Composite component specifications may set
optional=Trueto align a subset of entity IDs and preserve missing rows asNone.progressreceives messages whenverboseis enabled.
- DeepMTP.data.process_interaction_data(interaction_data: DataFrame | ndarray, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) InteractionInfo
A function that processes the interaction data. It is called separately for the train, val and test interaction data. There are two main types of interaction data formats that are supported: -> numpy format: This is a 2d numpy array that represents the interaction (also called score) matrix usually found in problems settings with fully observed matrices (multi-label classification, multivariate regression) -> triplet format: The most flexible format as it can be used to represent every possible problem setting
- Parameters:
interaction_data (_type_) – a numpy array or a dataframe with the interaction data
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose interaction format messages. Defaults to None.
- Returns:
a dictionary with the interaction data and additional information that could be detected (format, type of instance and target ids, etc.)
- Return type:
dict
- DeepMTP.data.process_target_features(target_features: Any, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None, feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None) FeatureInfo | None
Normalize target features without mutating caller-owned inputs.
Composite component specifications may set
optional=Trueto align a subset of entity IDs and preserve missing rows asNone.progressreceives messages whenverboseis enabled.
- DeepMTP.data.split_data(data: MutableMapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], split_method: str, ratio: object, shuffle: bool, seed: int | None, verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) None
- Splits the dataset and offers two main functionalities:
split based on the 4 different validation settings (A, B, C, D)
if a test set already exists it separates a validation set, otherwise it first creates a test set and then a validation set.
- Parameters:
data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.
validation_setting (str) – The validation setting of the current problem.
split_method (str) – The splitting method used. The current implementation only supports the ‘random split’ using a specific seed but a future goal is to also offer a stratified option.
ratio (dict) – The train, val and test ratios used to split the data.
seed (int) – The seed used to initiate the randomized split
verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.
print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.
progress (DataProgressObserver | None, optional) – Receives verbose split lifecycle messages. Defaults to None.
- DeepMTP.data.transform_dense_features(values: object, state: Mapping[str, Any] | DenseScalingState) ndarray
Replay a saved dense-feature transformation.
- DeepMTP.data.validate_split_ratio(ratio: object) dict[str, float]
Validate and copy a train/validation/test split ratio mapping.