DeepMTP.data package

The data namespace owns built-in dataset helpers, interaction normalization, feature validation, splitting, preparation, model-ready loading, and structured progress events:

from DeepMTP.data import (
    DataLoaderFactory,
    DataProgressObserver,
    TrainingDataLoaders,
    data_process,
)

Batches

Model-ready dataloaders return DeepMTP.data.batches.MTPBatch. It is a mapping compatible with the historical tensor keys and also exposes typed dense, graph, sparse, ID, image, sequence, tabular, composite, or custom branch inputs. Graph, tabular, and sequence inputs remain structured; sparse inputs use PyTorch COO or CSR tensors; composite inputs keep named component values together. Optional composite components carry their present sub-batch and a full-batch Boolean mask. Calling batch.to(device) moves every component together.

Typed model-input batches shared by data loading and training.

class DeepMTP.data.batches.CompositeBranchBatch(values: CompositeInput)

Bases: object

A named collection of inputs consumed by a composite encoder.

property batch_size: int

Number of observations shared by every component.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'composite'
pin_memory() → CompositeBranchBatch

Pin every component in the structured input.

to(device: torch.device | str, *, non_blocking: bool = False) → CompositeBranchBatch

Move every component to one device.

values: CompositeInput
class DeepMTP.data.batches.CompositeInput(components: Mapping[str, Any])

Bases: Mapping[str, Any]

Named model inputs for multiple encoders on one entity axis.

property batch_size: int

Number of observations shared by every component.

components: Mapping[str, Any]
pin_memory() → CompositeInput

Pin every component for asynchronous host-to-device transfer.

to(device: torch.device | str, *, non_blocking: bool = False) → CompositeInput

Move every component to one device.

class DeepMTP.data.batches.CustomBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A tensor batch passed unchanged to a user-supplied branch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'custom'
class DeepMTP.data.batches.DenseBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A batch of dense feature vectors shaped [batch, features].

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'dense'
class DeepMTP.data.batches.GraphBranchBatch(values: GraphInput)

Bases: object

A structured batch consumed by a graph neural network encoder.

property batch_size: int

Number of graphs represented by this branch batch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'graph'
pin_memory() → GraphBranchBatch

Pin every tensor in the graph input.

to(device: torch.device | str, *, non_blocking: bool = False) → GraphBranchBatch

Move the graph input to one device.

values: GraphInput
class DeepMTP.data.batches.GraphInput(node_features: torch.Tensor, edge_index: torch.Tensor, edge_features: torch.Tensor | None, graph_index: torch.Tensor, num_graphs: int)

Bases: object

One batch of homogeneous graphs represented by DeepMTP tensors.

property batch_size: int

Number of graphs represented by this input.

edge_features: torch.Tensor | None
edge_index: torch.Tensor
graph_index: torch.Tensor
node_features: torch.Tensor
num_graphs: int
pin_memory() → GraphInput

Pin every graph tensor for asynchronous device transfer.

to(device: torch.device | str, *, non_blocking: bool = False) → GraphInput

Move every graph tensor to one device.

class DeepMTP.data.batches.IDBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A batch of zero-based integer entity IDs shaped [batch].

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'id'
class DeepMTP.data.batches.ImageBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A batch of images shaped [batch, channels, height, width].

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'image'
class DeepMTP.data.batches.MTPBatch(instance_id: torch.Tensor, target_id: torch.Tensor, instance_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch, target_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch, score: torch.Tensor)

Bases: Mapping[str, Any]

A complete interaction batch with mapping compatibility.

Attribute access exposes the typed branch inputs. Historical mapping keys continue to return the underlying tensors so existing dataloader consumers remain compatible.

property batch_size: int

Number of interactions represented by the batch.

property instance_features: torch.Tensor | CompositeInput | GraphInput | SequenceInput | TabularInput

Compatibility view of the instance branch input.

instance_id: torch.Tensor
instance_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch
pin_memory() → MTPBatch

Pin every tensor for asynchronous host-to-device transfer.

score: torch.Tensor
property target_features: torch.Tensor | CompositeInput | GraphInput | SequenceInput | TabularInput

Compatibility view of the target branch input.

target_id: torch.Tensor
target_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch
to(device: torch.device | str, *, non_blocking: bool = False) → MTPBatch

Move every tensor in the batch to one device.

class DeepMTP.data.batches.MTPBatchCollator(instance_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite'], target_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite'], instance_sequence_padding_idx: int = 0, target_sequence_padding_idx: int = 0, instance_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None, target_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None, instance_component_padding_indices: Mapping[str, int] | None = None, target_component_padding_indices: Mapping[str, int] | None = None, instance_component_optional: Mapping[str, bool] | None = None, target_component_optional: Mapping[str, bool] | None = None)

Bases: object

Collate dataset items into a typed interaction batch.

instance_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None
instance_component_optional: Mapping[str, bool] | None = None
instance_component_padding_indices: Mapping[str, int] | None = None
instance_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']
instance_sequence_padding_idx: int = 0
target_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None
target_component_optional: Mapping[str, bool] | None = None
target_component_padding_indices: Mapping[str, int] | None = None
target_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']
target_sequence_padding_idx: int = 0
class DeepMTP.data.batches.MaskedComponentInput(values: Any | None, presence_mask: torch.Tensor)

Bases: object

Present values plus a full-batch mask for one optional component.

property batch_size: int

Full batch size including rows where this component is missing.

pin_memory() → MaskedComponentInput

Pin the present values and full-batch mask.

presence_mask: torch.Tensor
property present_count: int

Number of rows containing this component.

to(device: torch.device | str, *, non_blocking: bool = False) → MaskedComponentInput

Move the values and presence mask to one device.

values: Any | None
class DeepMTP.data.batches.SequenceBranchBatch(values: SequenceInput)

Bases: object

A structured batch consumed by a token-sequence encoder.

property batch_size: int

Number of observations represented by this branch batch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'sequence'
pin_memory() → SequenceBranchBatch

Pin every tensor in the structured sequence input.

to(device: torch.device | str, *, non_blocking: bool = False) → SequenceBranchBatch

Move the structured sequence input to one device.

values: SequenceInput
class DeepMTP.data.batches.SequenceInput(token_ids: torch.Tensor, attention_mask: torch.Tensor, lengths: torch.Tensor, padding_idx: int = 0)

Bases: object

Padded token IDs, attention mask, and original sequence lengths.

attention_mask: torch.Tensor
property batch_size: int

Number of token sequences represented by this input.

lengths: torch.Tensor
padding_idx: int = 0
pin_memory() → SequenceInput

Pin every sequence tensor for asynchronous device transfer.

to(device: torch.device | str, *, non_blocking: bool = False) → SequenceInput

Move every sequence tensor to one device.

token_ids: torch.Tensor
class DeepMTP.data.batches.SparseBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A sparse feature batch shaped [batch, features].

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'sparse'
pin_memory() → SparseBranchBatch

Pin sparse indices and values for asynchronous device transfer.

class DeepMTP.data.batches.TabularBranchBatch(values: TabularInput)

Bases: object

A structured batch of numeric values and categorical indices.

property batch_size: int

Number of observations represented by this branch batch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'tabular'
pin_memory() → TabularBranchBatch

Pin both tensors in the structured input.

to(device: torch.device | str, *, non_blocking: bool = False) → TabularBranchBatch

Move the structured input to one device.

values: TabularInput
class DeepMTP.data.batches.TabularInput(numeric: torch.Tensor, categorical: torch.Tensor)

Bases: object

Numeric and categorical tensors consumed by a tabular encoder.

property batch_size: int

Number of observations in the structured input.

categorical: torch.Tensor
numeric: torch.Tensor
pin_memory() → TabularInput

Pin both tensors for asynchronous host-to-device transfer.

to(device: torch.device | str, *, non_blocking: bool = False) → TabularInput

Move both tensors to one device.

class DeepMTP.data.batches.TensorBranchBatch(values: torch.Tensor)

Bases: object

One branch’s tensor payload with an explicit modality contract.

property batch_size: int

Number of observations represented by this branch batch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'custom'
pin_memory() → TensorBranchBatch

Return the same typed batch backed by pinned host memory.

to(device: torch.device | str, *, non_blocking: bool = False) → TensorBranchBatch

Return the same typed batch with its tensor moved to device.

values: torch.Tensor

Tabular preprocessing

Tabular schemas declare numeric columns, category vocabularies, missing-value policies, and normalization. Fitted training statistics are serializable and are reused for validation, testing, and restored-model prediction.

Schemas and reproducible preprocessing for mixed tabular branch inputs.

class DeepMTP.data.tabular.CategoricalColumn(name: str, categories: tuple[str | int, ...], embedding_dim: int)

Bases: object

One categorical feature and its fixed training vocabulary.

categories: tuple[str | int, ...]
embedding_dim: int
classmethod from_config(name: object, value: object) → CategoricalColumn

Validate one categorical-column declaration.

name: str
to_dict() → dict[str, Any]

Return a JSON-compatible public configuration mapping.

property vocabulary_size: int

Embedding-table size, including index zero for unknown values.

class DeepMTP.data.tabular.TabularPreprocessingState(numeric_columns: tuple[str, ...], categorical_columns: tuple[str, ...], numeric_fill_values: tuple[float, ...], numeric_offsets: tuple[float, ...], numeric_scales: tuple[float, ...])

Bases: object

Numeric statistics fitted only on training entities.

categorical_columns: tuple[str, ...]
classmethod from_config(value: object, *, schema: TabularSchema, branch: str) → TabularPreprocessingState

Restore and validate preprocessing metadata from a checkpoint config.

numeric_columns: tuple[str, ...]
numeric_fill_values: tuple[float, ...]
numeric_offsets: tuple[float, ...]
numeric_scales: tuple[float, ...]
to_dict() → dict[str, Any]

Return JSON-compatible checkpoint metadata.

class DeepMTP.data.tabular.TabularPreprocessor(schema: TabularSchema, state: TabularPreprocessingState, branch: str)

Bases: object

Transform entity feature frames using one schema and fitted state.

branch: str
classmethod fit(frame: DataFrame, schema: TabularSchema, *, branch: str) → TabularPreprocessor

Fit numeric imputation and normalization on a training frame.

classmethod from_state(schema: TabularSchema, state: Mapping[str, Any] | TabularPreprocessingState, *, branch: str) → TabularPreprocessor

Build a preprocessor from saved training metadata.

schema: TabularSchema
state: TabularPreprocessingState
transform(frame: DataFrame) → DataFrame

Return an indexed feature frame with structured tensor-ready values.

class DeepMTP.data.tabular.TabularSchema(numeric_columns: tuple[str, ...], categorical_columns: tuple[CategoricalColumn, ...], numeric_normalization: Literal['none', 'standard', 'minmax'] = 'standard', numeric_missing: Literal['error', 'mean', 'zero'] = 'mean', categorical_missing: Literal['error', 'unknown'] = 'unknown', categorical_unknown: Literal['error', 'unknown'] = 'unknown', feature_gating: bool = False)

Bases: object

Explicit numeric and categorical layout for one entity branch.

categorical_columns: tuple[CategoricalColumn, ...]
categorical_missing: Literal['error', 'unknown'] = 'unknown'
property categorical_names: tuple[str, ...]

Categorical column names in their stable tensor order.

categorical_unknown: Literal['error', 'unknown'] = 'unknown'
property encoded_width: int

Width after concatenating numeric values and category embeddings.

feature_gating: bool = False
classmethod from_config(value: object, *, branch: str) → TabularSchema

Validate a public tabular schema mapping.

numeric_columns: tuple[str, ...]
numeric_missing: Literal['error', 'mean', 'zero'] = 'mean'
numeric_normalization: Literal['none', 'standard', 'minmax'] = 'standard'
to_dict() → dict[str, Any]

Return a JSON-compatible public configuration mapping.

Dataset helpers

Dataset-specific dependencies remain lazy until this module or one of its package-level exports is used.

Built-in, generated, and downloadable dataset helpers.

class DeepMTP.data.datasets.DatasetBundle

Bases: TypedDict

Stable outer structure returned by legacy dataset helpers.

test: DatasetSplit
train: DatasetSplit
val: DatasetSplit
class DeepMTP.data.datasets.DatasetSplit

Bases: TypedDict

One raw train, validation, or test split.

X_instance: ndarray | DataFrame | None
X_target: ndarray | DataFrame | None
y: ndarray | DataFrame | None
class DeepMTP.data.datasets.ProcessedDatasetBundle

Bases: TypedDict

Processed train, validation, and test data.

test: ProcessedDatasetSplit
train: ProcessedDatasetSplit
val: ProcessedDatasetSplit
class DeepMTP.data.datasets.ProcessedDatasetSplit

Bases: TypedDict

One processed split with optional interactions and side features.

X_instance: ProcessedDatasetValue | None
X_target: ProcessedDatasetValue | None
y: ProcessedDatasetValue | None
class DeepMTP.data.datasets.ProcessedDatasetValue

Bases: TypedDict

Processed data container consumed by legacy data utilities.

data: DataFrame
DeepMTP.data.datasets.format_mtr_datasets() → str

Format the catalog of supported multivariate regression datasets.

DeepMTP.data.datasets.generate_MTP_dataset(num_instances: int, num_targets: int, num_instance_features: int | None = None, num_target_features: int | None = None, split_instances: Mapping[str, float] | None = None, split_targets: Mapping[str, float] | None = None, return_static_features_data: bool = False) → ProcessedDatasetBundle

Generate deterministic triplet data with optional axis-level splits.

Each split mapping requires a test ratio and may include a validation ratio. The train ratio is derived when omitted and verified when supplied. A test-only axis reuses its training IDs in a validation split created on the other axis.

Parameters:
  • num_instances – Number of instances.

  • num_targets – Number of targets.

  • num_instance_features – Number of generated instance features, or none.

  • num_target_features – Number of generated target features, or none.

  • split_instances – Optional instance-axis split ratios.

  • split_targets – Optional target-axis split ratios.

  • return_static_features_data – Repeat non-novel side features in holdouts.

Returns:

Processed triplet and feature frames for train, validation, and test.

Raises:
  • TypeError – If a split or boolean flag has the wrong type.

  • ValueError – If a dimension or split ratio is invalid.

DeepMTP.data.datasets.generate_dummy_dataset(num_instances: int, num_targets: int, num_instance_features: int, num_target_features: int, error_mu: float, error_sigma: float, sklearn_version: bool = False, seed: int = 42, mode: str = 'u+logv', split_ratio: Mapping[str, float] | None = None) → DatasetBundle

Generate a reproducible dyadic regression dataset.

sklearn_version starts from normally distributed features. For modes involving logarithms or fractional powers, the relevant features are transformed to a positive domain so every generated score remains real. The error distribution applies to the four legacy noisy relationship modes.

Parameters:
  • num_instances – Number of instances.

  • num_targets – Number of targets.

  • num_instance_features – Number of features per instance.

  • num_target_features – Number of features per target.

  • error_mu – Mean of the Gaussian error added by noisy modes.

  • error_sigma – Non-negative standard deviation of that Gaussian error.

  • sklearn_version – Generate normally distributed rather than uniform features.

  • seed – Non-negative 32-bit seed used for generation and splitting.

  • mode – Mathematical relationship used to generate scores.

  • split_ratio – Train, validation, and test ratios.

Returns:

A dataset bundle containing aligned train, validation, and test rows.

Raises:
  • TypeError – If a boolean or error parameter has the wrong type.

  • ValueError – If a dimension, mode, seed, ratio, or generated value is invalid.

DeepMTP.data.datasets.generate_interaction_matrix(input_path: str | PathLike[str], output_path: str | PathLike[str]) → None

Convert annotator labels into an atomically written correctness matrix.

DeepMTP.data.datasets.load_process_DP(path: str | PathLike[str] = './data', dataset_name: str = 'ern', variant: Literal['undivided', 'divided'] = 'undivided', random_state: int | None = 42, split_ratio: Mapping[str, float] | None = None, split_instance_features: bool = False, split_target_features: bool = False, validation_setting: str = 'B', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load a dyadic-prediction dataset and optionally create held-out splits.

Parameters:
  • path – Directory used for the local dataset cache.

  • dataset_name – One of "ern", "srn", "dpie", or "dpii".

  • variant – Return the complete matrices or create train/validation/test splits.

  • random_state – Seed used by the split operations, or None.

  • split_ratio – Train, validation, and test proportions.

  • split_instance_features – Split instance features with interaction rows.

  • split_target_features – Split target features with interaction columns.

  • validation_setting – Novelty setting "B", "C", or "D".

  • print_mode – "basic" for ordinary messages or "dev" for messages prefixed for the development application.

  • progress – Receives dataset loading and download status events.

Raises:
  • ValueError – If an option or dataset matrix is invalid.

  • TypeError – If a feature-splitting flag is not boolean.

  • DatasetDownloadError – If a required dataset file cannot be downloaded.

  • FileNotFoundError – If a download does not create all required files.

DeepMTP.data.datasets.load_process_MC(path: str | PathLike[str] = './data', dataset_name: str = 'ml-100k', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load a matrix-completion dataset from the MovieLens repository.

Parameters:
  • path – Directory used for the local dataset cache.

  • dataset_name – Matrix-completion dataset to load.

  • print_mode – "basic" for ordinary messages or "dev" for messages prefixed for the development application.

  • progress – Receives dataset download status events.

Raises:
  • ValueError – If an option or the ratings data is invalid.

  • DatasetDownloadError – If the archive cannot be downloaded or extracted.

  • FileNotFoundError – If the archive does not contain the expected ratings file.

Returns:

The ratings triplets and empty side-feature and held-out splits.

DeepMTP.data.datasets.load_process_MLC(path: str | PathLike[str] = './data', dataset_name: str = 'bibtex', variant: Literal['undivided', 'divided'] = 'undivided', features_type: Literal['numpy', 'dataframe'] = 'numpy', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load a cached or downloadable scikit-multilearn benchmark dataset.

Parameters:
  • path – Directory used for the local dataset cache.

  • dataset_name – Supported multi-label benchmark name.

  • variant – Return the complete dataset or its published train/test split.

  • features_type – Return instance features as an array or feature DataFrame.

  • print_mode – "basic" for ordinary messages or "dev" for messages prefixed for the development application.

  • progress – Receives dataset loading and download status events.

Raises:
  • ValueError – If an option or cached dataset is invalid.

  • DatasetDownloadError – If required cache files cannot be downloaded safely.

  • IsADirectoryError – If an expected cache file path is a directory.

DeepMTP.data.datasets.load_process_MTL(path: str | PathLike[str] = './data', dataset_name: str = 'dog', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load multi-task learning datasets from my custom repository.

Parameters:
  • path (str, optional) – The path where the datasets should be stored. If it doesn’t exist, download and store it in this directory. Defaults to ‘./data’.

  • dataset_name (str, optional) – The name of the multi-task learning dataset. Defaults to ‘dog’.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives dataset download status events. Defaults to None.

Raises:
  • ValueError – If a public option or dataset file is invalid.

  • DatasetDownloadError – If the dataset archive cannot be downloaded safely.

  • FileNotFoundError – If required score or image files are missing.

Returns:

A dictionary with all the available data for the multi-task learning dataset.

Return type:

dict

DeepMTP.data.datasets.load_process_MTR(path: str | PathLike[str] = './data', dataset_name: str = 'enb', features_type: Literal['numpy', 'dataframe'] = 'numpy', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load multivariate regression datasets from the mulan repository.

Parameters:
  • path (str, optional) – The path where the datasets should be stored. If it doesn’t exist, download and store it in this directory. Defaults to ‘./data’.

  • dataset_name (str, optional) – The name of the multivariate regression dataset. Defaults to ‘enb’.

  • features_type (str, optional) – The format of the instance features. There are two possible values, numpy and dataframe. This is intended to test the functionality of the data_process function. Defaults to ‘numpy’.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives dataset loading and download status events. Defaults to None.

Raises:
  • ValueError – If a public option or downloaded dataset is invalid.

  • DatasetDownloadError – If the dataset archive cannot be downloaded safely.

  • FileNotFoundError – If the archive does not contain the requested dataset.

Returns:

A dictionary with all the available data for the multivariate regression dataset.

Return type:

dict

DeepMTP.data.datasets.print_MTR_datasets(*, progress: DataProgressObserver | None = None) → None

Present the supported multivariate regression dataset catalog.

The default observer preserves the historical console table. Inject NullDataProgressObserver to suppress output or another data progress observer to render the catalog elsewhere.

DeepMTP.data.datasets.process_dummy_DP(num_instance_features: int = 10, num_target_features: int = 3, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', instance_features_format: Literal['numpy', 'dataframe'] = 'numpy', target_features_format: Literal['numpy', 'dataframe'] = 'numpy', random_state: int | None = 42) → DatasetBundle

Generate a reproducible dummy dyadic prediction dataset.

Parameters:
  • num_instance_features – Number of features per instance.

  • num_target_features – Number of features per target.

  • num_instances – Number of instances.

  • num_targets – Number of targets.

  • interaction_matrix_format – Return scores as a dense matrix or triplets.

  • instance_features_format – Return instance features as an array or DataFrame.

  • target_features_format – Return target features as an array or DataFrame.

  • random_state – Seed for an isolated NumPy generator. Use None for non-deterministic data.

Returns:

A dataset bundle containing an undivided training split.

Raises:

ValueError – If a format, dimension, or random state is invalid.

DeepMTP.data.datasets.process_dummy_MLC(num_features: int = 10, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', features_format: Literal['numpy', 'dataframe'] = 'numpy', random_state: int | None = 42) → DatasetBundle

Generate a reproducible dummy multi-label classification dataset.

Parameters:
  • num_features – Number of features per instance.

  • num_instances – Number of instances.

  • num_targets – Number of binary labels.

  • interaction_matrix_format – Return labels as a dense matrix or triplets.

  • features_format – Return features as a dense matrix or feature DataFrame.

  • random_state – Seed for an isolated NumPy generator. Use None for non-deterministic data.

Returns:

A dataset bundle containing an undivided training split.

Raises:

ValueError – If a format, dimension, or random state is invalid.

DeepMTP.data.datasets.process_dummy_MTR(num_features: int = 10, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', features_format: Literal['numpy', 'dataframe'] = 'numpy', variant: Literal['undivided', 'divided'] = 'undivided', split_ratio: Mapping[str, float] | None = None, random_state: int | None = 42) → DatasetBundle

Generate a reproducible dummy multivariate regression dataset.

Parameters:
  • num_features – Number of features per instance.

  • num_instances – Number of instances.

  • num_targets – Number of regression targets.

  • interaction_matrix_format – Return scores as a dense matrix or triplets.

  • features_format – Return features as a dense matrix or feature DataFrame.

  • variant – Return an undivided dataset or train/validation/test splits.

  • split_ratio – Ratios used for the divided variant.

  • random_state – Seed used for both data generation and splitting. Use None for non-deterministic data.

Returns:

A dataset bundle containing the generated scores and features.

Raises:
  • TypeError – If split_ratio is not a mapping.

  • ValueError – If an option, dimension, ratio, or random state is invalid.

Interactions

Interaction normalization, validation, and setting inference.

class DeepMTP.data.interactions.InteractionInfo

Bases: TypedDict

Normalized interaction data and its detected schema.

data: DataFrame
instance_id_type: Literal['int']
missing_values: bool
original_format: Literal['triplets', 'numpy']
target_id_type: Literal['int']
exception DeepMTP.data.interactions.NoveltyInferenceWarning

Bases: UserWarning

Warning emitted when entity identity was lost in matrix input.

DeepMTP.data.interactions.check_interaction_files_column_type_format(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None

Checks the type of the instance and target ids in the interaction data. If any inconsistencies are detected (like different id types between train and test interaction data), an exception is raised

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose ID-type validation messages. Defaults to None.

DeepMTP.data.interactions.check_interaction_files_format(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None

Checks the format of the interaction data. If any inconsistencies are detected (like different formats between train and test interaction data), an exception is raised

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose format-validation messages. Defaults to None.

DeepMTP.data.interactions.check_novel_instances(train: Mapping[str, Any] | None, test: Mapping[str, Any] | None, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → bool | None

Return whether test interactions contain previously unseen instances.

progress receives messages when verbose is enabled.

DeepMTP.data.interactions.check_novel_targets(train: Mapping[str, Any] | None, test: Mapping[str, Any] | None, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → bool | None

Return whether test interactions contain previously unseen targets.

progress receives messages when verbose is enabled.

DeepMTP.data.interactions.check_target_variable_type(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, classification_mode: Literal['binary', 'multiclass'] | None = None, num_classes: int | None = None, progress: DataProgressObserver | None = None) → Literal['binary', 'multiclass', 'real-valued']

Checks the type of the target variable in the interaction data. If any inconsistencies are detected (like different id types between train and test interaction data), an exception is raised

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose target-type validation messages. Defaults to None.

DeepMTP.data.interactions.check_variable_type(samples_arr: DataFrame, *, classification_mode: Literal['binary', 'multiclass'] | None = None, num_classes: int | None = None) → Literal['binary', 'multiclass', 'real-valued']

Detect binary or real-valued scores, or validate declared class IDs.

Parameters:

samples_arr (numpy.array) – A numpy array with the target variables

Returns:

binary, multiclass, or real-valued.

Return type:

str

Multiclass targets are intentionally opt-in because an integer-valued regression target is otherwise indistinguishable from class IDs.

DeepMTP.data.interactions.get_estimated_validation_setting(novel_instances_flag: bool | None, novel_targets_flag: bool | None, verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → Literal['A', 'B', 'C', 'D'] | None

Uses the combination of infrmation about novel instances and novel targets to determine the validation setting that is possible.

Parameters:
  • novel_instances_flag (bool) – A boolean that indicates whether or not the test set contains novel instances.

  • novel_targets_flag (bool) – A boolean that indicates whether or not the test set contains novel targets.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose validation-setting inference messages. Defaults to None.

Returns:

The validation setting. Possible values are A, B, C, D

Return type:

str

DeepMTP.data.interactions.process_interaction_data(interaction_data: DataFrame | ndarray, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → InteractionInfo

A function that processes the interaction data. It is called separately for the train, val and test interaction data. There are two main types of interaction data formats that are supported: -> numpy format: This is a 2d numpy array that represents the interaction (also called score) matrix usually found in problems settings with fully observed matrices (multi-label classification, multivariate regression) -> triplet format: The most flexible format as it can be used to represent every possible problem setting

Parameters:
  • interaction_data (_type_) – a numpy array or a dataframe with the interaction data

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose interaction format messages. Defaults to None.

Returns:

a dictionary with the interaction data and additional information that could be detected (format, type of instance and target ids, etc.)

Return type:

dict

Features

Feature normalization and cross-input consistency validation.

class DeepMTP.data.features.CompositeFeatureSpec

Bases: TypedDict

Data-preparation options for one named composite component.

kind: Literal['graph', 'sequence'] | None
optional: bool
class DeepMTP.data.features.FeatureComponentInfo

Bases: _FeatureComponentInfoRequired

Representation metadata for one component of a composite feature.

graph_edge_features: int | None
info: Literal['numpy', 'dataframe', 'graph', 'sequence', 'scipy_sparse', 'torch_sparse']
num_features: int | None
class DeepMTP.data.features.FeatureInfo

Bases: _FeatureInfoRequired

Normalized entity features and their detected representation.

components: dict[str, FeatureComponentInfo] | None
data: DataFrame
graph_edge_features: int | None
info: Literal['numpy', 'dataframe', 'images', 'graph', 'sequence', 'tabular', 'scipy_sparse', 'torch_sparse', 'composite']
num_features: int | None
DeepMTP.data.features.cross_input_consistency_check_instances(data: Mapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None

Checks the consistency of instance ids in the interaction data and instance features. The requirements to pass this check change depending on the format of the interaction data and instance features

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • validation_setting (str) – The validation setting of the current problem.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose instance-consistency messages. Defaults to None.

DeepMTP.data.features.cross_input_consistency_check_targets(data: Mapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None

Checks the consistency of target ids in the interaction data and target features. The requirements to pass this check change depending on the format of the interaction data and target features

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • validation_setting (str) – The validation setting of the current problem.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose target-consistency messages. Defaults to None.

DeepMTP.data.features.process_instance_features(instance_features: Any, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None, feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None) → FeatureInfo | None

Normalize instance features without mutating caller-owned inputs.

Composite component specifications may set optional=True to align a subset of entity IDs and preserve missing rows as None. progress receives messages when verbose is enabled.

DeepMTP.data.features.process_target_features(target_features: Any, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None, feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None) → FeatureInfo | None

Normalize target features without mutating caller-owned inputs.

Composite component specifications may set optional=True to align a subset of entity IDs and preserve missing rows as None. progress receives messages when verbose is enabled.

Graph utilities

PyTorch Geometric is imported only when graph inputs are used. Its Data objects are validated, cloned, and converted to DeepMTP’s graph batch contract.

Optional PyTorch Geometric normalization and batching helpers.

class DeepMTP.data.graphs.CollatedGraph(node_features: torch.Tensor, edge_index: torch.Tensor, edge_features: torch.Tensor | None, graph_index: torch.Tensor, num_graphs: int)

Bases: object

Framework-neutral tensors extracted from one PyG graph batch.

edge_features: torch.Tensor | None
edge_index: torch.Tensor
graph_index: torch.Tensor
node_features: torch.Tensor
num_graphs: int
DeepMTP.data.graphs.collate_graph_rows(rows: Sequence[object]) → CollatedGraph

Batch graph rows with PyG and expose only DeepMTP’s tensor contract.

DeepMTP.data.graphs.copy_graph_value(value: object) → Any

Clone one validated graph without sharing its tensors.

DeepMTP.data.graphs.graph_object_array(values: Sequence[object]) → ndarray

Store graph objects without invoking third-party array protocols.

DeepMTP.data.graphs.is_graph_value(value: object) → bool

Return whether value is a PyTorch Geometric graph object.

DeepMTP.data.graphs.normalize_graph_rows(rows: Iterable[object], *, entity_name: str) → tuple[list[Any], int, int | None]

Validate homogeneous PyG graphs and return independent normalized copies.

DeepMTP.data.graphs.require_pyg_data() → ModuleType

Return PyG’s data module or explain how to install graph support.

Sparse utilities

The sparse input adapter resolves SciPy only when inspecting a SciPy-backed value.

Sparse feature normalization and batching utilities.

DeepMTP.data.sparse.collate_sparse_rows(rows: Sequence[object]) → torch.Tensor

Stack sparse rows into one floating-point PyTorch COO batch.

DeepMTP.data.sparse.copy_feature_value(value: object) → Any

Copy one dense or sparse feature value without densifying it.

DeepMTP.data.sparse.is_scipy_sparse(value: object) → bool

Return whether value is a SciPy sparse matrix.

DeepMTP.data.sparse.is_sparse_value(value: object) → bool

Return whether value uses a supported sparse representation.

DeepMTP.data.sparse.is_torch_sparse(value: object) → bool

Return whether value is a supported PyTorch sparse tensor.

DeepMTP.data.sparse.normalize_sparse_matrix(matrix: object, *, entity_name: str) → tuple[list[Any], int, Literal['scipy', 'torch']]

Validate a 2-D sparse matrix and return independent sparse rows.

DeepMTP.data.sparse.normalize_sparse_rows(rows: Iterable[object], *, entity_name: str) → tuple[list[Any], int, Literal['scipy', 'torch']]

Validate and copy sparse feature rows with one consistent width/backend.

DeepMTP.data.sparse.pin_sparse_tensor(values: torch.Tensor) → torch.Tensor

Pin sparse tensor components because Tensor.pin_memory lacks support.

DeepMTP.data.sparse.sparse_backend(value: object) → Literal['scipy', 'torch']

Return the sparse backend used by value.

DeepMTP.data.sparse.sparse_object_array(values: Sequence[object]) → ndarray

Return an object array without invoking sparse tensor array protocols.

Splitting

Dataset splitting, defensive copying, and result validation.

DeepMTP.data.splitting.split_data(data: MutableMapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], split_method: str, ratio: object, shuffle: bool, seed: int | None, verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None
Splits the dataset and offers two main functionalities:
  • split based on the 4 different validation settings (A, B, C, D)

  • if a test set already exists it separates a validation set, otherwise it first creates a test set and then a validation set.

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • validation_setting (str) – The validation setting of the current problem.

  • split_method (str) – The splitting method used. The current implementation only supports the ‘random split’ using a specific seed but a future goal is to also offer a stratified option.

  • ratio (dict) – The train, val and test ratios used to split the data.

  • seed (int) – The seed used to initiate the randomized split

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose split lifecycle messages. Defaults to None.

Preparation

End-to-end data preparation and feature scaling orchestration.

class DeepMTP.data.preparation.DataInfo

Bases: TypedDict

Metadata inferred while preparing a multi-target dataset.

data_preparation_state: dict[str, Any]
detected_classification_mode: Literal['binary', 'multiclass'] | None
detected_num_classes: int | None
detected_problem_mode: Literal['classification', 'regression']
detected_validation_setting: Literal['A', 'B', 'C', 'D']
instance_branch_component_graph_edge_dims: dict[str, int | None] | None
instance_branch_component_input_dims: dict[str, int | None] | None
instance_branch_graph_edge_dim: int | None
instance_branch_input_dim: int | None
target_branch_component_graph_edge_dims: dict[str, int | None] | None
target_branch_component_input_dims: dict[str, int | None] | None
target_branch_graph_edge_dim: int | None
target_branch_input_dim: int | None
class DeepMTP.data.preparation.Transformer(*args, **kwargs)

Bases: Protocol

Minimal interface implemented by supported feature scalers.

transform(values: Any) → Any

Transform values using a previously fitted scaler.

DeepMTP.data.preparation.data_process(data: Mapping[str, Any], validation_setting: str | None = None, split_method: str = 'random', ratio: object = None, shuffle: bool = True, seed: int | None = 42, verbose: bool = False, print_mode: str = 'basic', scale_instance_features: str | None = None, scale_target_features: str | None = None, *, classification_mode: str | None = None, num_classes: int | None = None, instance_feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None, target_feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None, progress: DataProgressObserver | None = None, dense_preprocessing_state: Mapping[str, Any] | DataPreparationState | DensePreprocessingState | None = None) → tuple[dict[str, Any], dict[str, Any], dict[str, Any], DataInfo]

The main function that handles all the preprocessing steps and checks needed to prepare the dataset to be used by the model.

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • validation_setting (str, optional) – The validation setting of the current problem. Defaults to None.

  • split_method (str, optional) – The splitting method used. The current implementation only supports the ‘random split’ using a specific seed but a future goal is to also offer a stratified option. Defaults to ‘random’.

  • ratio (dict, optional) – The train, val and test ratios used to split the data. Defaults to {‘train’: 0.7, ‘test’: 0.2, ‘val’: 0.1}.

  • shuffle (bool, optional) – Whether or not the dataset will be shuffled before the split. If is set to False, the seed value is not used.

  • seed (int, optional) – The seed used to initiate the randomized split. Defaults to 42.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • scale_instance_features (str, optional) – The scaler used for the instance features. Possible values are ‘MinMax’ for the MinMax scaler and ‘Standard’ for the standard scaler. Defaults to None.

  • scale_target_features (str, optional) – The scaler used for the target features. Possible values are ‘MinMax’ for the MinMax scaler and ‘Standard’ for the standard scaler. Defaults to None.

  • classification_mode (str, optional) – Set to multiclass to interpret zero-based integer scores as class IDs. Binary scores continue to be detected automatically. Defaults to None.

  • num_classes (int, optional) – Declared multiclass count. When omitted in multiclass mode, it is inferred from all supplied score splits.

  • instance_feature_kind (str, optional) – Set to sequence when instance features are already-tokenized integer sequences. For composite features, pass a mapping from component name to sequence, None, or a mapping with kind and optional fields. Defaults to None.

  • target_feature_kind (str, optional) – Set to sequence when target features are already-tokenized integer sequences. For composite features, pass a mapping from component name to sequence, None, or a mapping with kind and optional fields. Defaults to None.

  • progress (DataProgressObserver | None, optional) – Receives verbose messages from the complete data-processing pipeline. Defaults to None.

  • dense_preprocessing_state (mapping, optional) – Previously saved dense scaler parameters. Accepts either the dense section or the full data_preparation_state returned in data_info. When supplied, the saved training statistics are replayed instead of fitting new scalers. Defaults to None.

Returns:

Four different dictionaries containing:
  • Train processed data

  • Validation processed data

  • Test processed data

  • general information about the datasets

Return type:

dict, dict, dict, dict

DeepMTP.data.preparation.normalize(row: Any, scaler: Transformer) → Any

Just normalizes a row or features

Parameters:
  • row (numpy.array) – an array of features

  • scaler (sklearn.scaler) – the scaler that will be used to scale the features

Returns:

a scaled array of feature

Return type:

numpy.array

DeepMTP.data.preparation.validate_split_ratio(ratio: object) → dict[str, float]

Validate and copy a train/validation/test split ratio mapping.

Preprocessing metadata

Legacy dense scaling parameters and split provenance are represented by validated, JSON-compatible state objects that can be replayed without a fitted scikit-learn object.

Serializable preprocessing state shared by data preparation and checkpoints.

class DeepMTP.data.preprocessing.DataPreparationState(validation_setting: str, split_method: str, split_ratio: tuple[tuple[str, float], ...], shuffle: bool, seed: int | None, dense: DensePreprocessingState)

Bases: object

Split provenance and fitted dense preprocessing from data_process.

dense: DensePreprocessingState
classmethod from_config(value: object) → DataPreparationState

Validate a saved data-preparation state mapping.

seed: int | None
shuffle: bool
split_method: str
split_ratio: tuple[tuple[str, float], ...]
to_dict() → dict[str, Any]

Return JSON-compatible checkpoint metadata.

validation_setting: str
class DeepMTP.data.preprocessing.DensePreprocessingState(instance: DenseScalingState | None = None, target: DenseScalingState | None = None)

Bases: object

Dense scaling state for the instance and target feature axes.

for_branch(branch: str) → DenseScalingState | None

Return the scaler state for one entity axis.

classmethod from_config(value: object) → DensePreprocessingState

Validate a combined dense-preprocessing state mapping.

instance: DenseScalingState | None = None
target: DenseScalingState | None = None
to_dict() → dict[str, Any]

Return JSON-compatible checkpoint metadata.

class DeepMTP.data.preprocessing.DenseScalingState(method: Literal['MinMax', 'Standard'], feature_count: int, offsets: tuple[float, ...], scales: tuple[float, ...], data_minimums: tuple[float, ...] = (), data_maximums: tuple[float, ...] = (), feature_range: tuple[float, float] = (0.0, 1.0))

Bases: object

Parameters needed to replay one fitted dense-feature transformation.

data_maximums: tuple[float, ...] = ()
data_minimums: tuple[float, ...] = ()
feature_count: int
feature_range: tuple[float, float] = (0.0, 1.0)
classmethod from_config(value: object) → DenseScalingState

Validate and restore a JSON-compatible scaling-state mapping.

classmethod from_fitted_scaler(scaler: object, *, method: Literal['MinMax', 'Standard']) → DenseScalingState

Capture explicit numeric state from a fitted scikit-learn scaler.

method: Literal['MinMax', 'Standard']
offsets: tuple[float, ...]
scales: tuple[float, ...]
to_dict() → dict[str, Any]

Return JSON-compatible checkpoint metadata.

transform(values: object) → ndarray

Transform one vector or matrix using the saved training statistics.

DeepMTP.data.preprocessing.transform_dense_features(values: object, state: Mapping[str, Any] | DenseScalingState) → ndarray

Replay a saved dense-feature transformation.

Loading

Dataset and dataloader construction for DeepMTP experiments.

class DeepMTP.data.loading.BaseDataset(*args: Any, **kwargs: Any)

Bases: Dataset

Resolve interaction rows to model-ready entity features.

class DeepMTP.data.loading.DataLoaderFactory(config: DeepMTPConfig | Mapping[str, Any], device: torch.device | str)

Bases: object

Build configured dataloaders without trainer-specific orchestration.

for_prediction(data: Mapping[str, Any]) → torch.utils.data.DataLoader

Build a deterministic inference dataloader.

for_training(train_data: Mapping[str, Any], validation_data: Mapping[str, Any] | None, test_data: Mapping[str, Any] | None) → TrainingDataLoaders

Build the train, validation, and test dataloaders.

class DeepMTP.data.loading.DatasetItem

Bases: TypedDict

One model-ready interaction returned by BaseDataset.

instance_features: Any
instance_id: int
score: Any
target_features: Any
target_id: int
class DeepMTP.data.loading.MTPSampler(*args: Any, **kwargs: Any)

Bases: Sampler

Undersample classes and provide a fresh order for every iteration.

micro balances classes globally, macro balances them independently for each target, and instance balances them independently for each instance. Grouped balancing can omit entities that are absent from the retained observations.

get_balanced_df(data: DataFrame) → DataFrame

Return one frame undersampled to its smallest class size.

class DeepMTP.data.loading.TrainingDataLoaders(train: torch.utils.data.DataLoader, validation: torch.utils.data.DataLoader, test: torch.utils.data.DataLoader)

Bases: object

Dataloaders required by one training run.

test: torch.utils.data.DataLoader
train: torch.utils.data.DataLoader
validation: torch.utils.data.DataLoader

Progress

Progress messages shared by data preparation and validation.

class DeepMTP.data.progress.ConsoleDataProgressObserver(print_mode: Literal['basic', 'dev'] = 'basic')

Bases: object

Render data progress using the historical basic or developer format.

on_event(event: DataProgressEvent) → None
class DeepMTP.data.progress.DataProgressEvent(message: str, kind: Literal['message', 'operation_started', 'operation_completed'] = 'message', subject: str | None = None, level: Literal['info', 'warning', 'error'] = 'info', end: str = '\n', value: object | None = None)

Bases: object

One user-facing data preparation or validation message.

end: str = '\n'
kind: Literal['message', 'operation_started', 'operation_completed'] = 'message'
level: Literal['info', 'warning', 'error'] = 'info'
message: str
subject: str | None = None
value: object | None = None
class DeepMTP.data.progress.DataProgressObserver(*args, **kwargs)

Bases: Protocol

Receives data preparation and validation messages.

on_event(event: DataProgressEvent) → None

Handle one data progress event.

class DeepMTP.data.progress.NullDataProgressObserver

Bases: object

No-op observer used when data progress is disabled.

on_event(event: DataProgressEvent) → None
DeepMTP.data.progress.build_data_progress_observer(enabled: bool, *, print_mode: Literal['basic', 'dev'] = 'basic') → DataProgressObserver

Select the default data progress observer.

Package contents

Data loading, preparation, validation, and progress services.

class DeepMTP.data.BaseDataset(*args: Any, **kwargs: Any)

Bases: Dataset

Resolve interaction rows to model-ready entity features.

class DeepMTP.data.CategoricalColumn(name: str, categories: tuple[str | int, ...], embedding_dim: int)

Bases: object

One categorical feature and its fixed training vocabulary.

categories: tuple[str | int, ...]
embedding_dim: int
classmethod from_config(name: object, value: object) → CategoricalColumn

Validate one categorical-column declaration.

name: str
to_dict() → dict[str, Any]

Return a JSON-compatible public configuration mapping.

property vocabulary_size: int

Embedding-table size, including index zero for unknown values.

class DeepMTP.data.CompositeBranchBatch(values: CompositeInput)

Bases: object

A named collection of inputs consumed by a composite encoder.

property batch_size: int

Number of observations shared by every component.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'composite'
pin_memory() → CompositeBranchBatch

Pin every component in the structured input.

to(device: torch.device | str, *, non_blocking: bool = False) → CompositeBranchBatch

Move every component to one device.

values: CompositeInput
class DeepMTP.data.CompositeFeatureSpec

Bases: TypedDict

Data-preparation options for one named composite component.

kind: Literal['graph', 'sequence'] | None
optional: bool
class DeepMTP.data.CompositeInput(components: Mapping[str, Any])

Bases: Mapping[str, Any]

Named model inputs for multiple encoders on one entity axis.

property batch_size: int

Number of observations shared by every component.

components: Mapping[str, Any]
pin_memory() → CompositeInput

Pin every component for asynchronous host-to-device transfer.

to(device: torch.device | str, *, non_blocking: bool = False) → CompositeInput

Move every component to one device.

class DeepMTP.data.ConsoleDataProgressObserver(print_mode: Literal['basic', 'dev'] = 'basic')

Bases: object

Render data progress using the historical basic or developer format.

on_event(event: DataProgressEvent) → None
class DeepMTP.data.CustomBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A tensor batch passed unchanged to a user-supplied branch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'custom'
values: torch.Tensor
class DeepMTP.data.DataInfo

Bases: TypedDict

Metadata inferred while preparing a multi-target dataset.

data_preparation_state: dict[str, Any]
detected_classification_mode: Literal['binary', 'multiclass'] | None
detected_num_classes: int | None
detected_problem_mode: Literal['classification', 'regression']
detected_validation_setting: Literal['A', 'B', 'C', 'D']
instance_branch_component_graph_edge_dims: dict[str, int | None] | None
instance_branch_component_input_dims: dict[str, int | None] | None
instance_branch_graph_edge_dim: int | None
instance_branch_input_dim: int | None
target_branch_component_graph_edge_dims: dict[str, int | None] | None
target_branch_component_input_dims: dict[str, int | None] | None
target_branch_graph_edge_dim: int | None
target_branch_input_dim: int | None
class DeepMTP.data.DataLoaderFactory(config: DeepMTPConfig | Mapping[str, Any], device: torch.device | str)

Bases: object

Build configured dataloaders without trainer-specific orchestration.

for_prediction(data: Mapping[str, Any]) → torch.utils.data.DataLoader

Build a deterministic inference dataloader.

for_training(train_data: Mapping[str, Any], validation_data: Mapping[str, Any] | None, test_data: Mapping[str, Any] | None) → TrainingDataLoaders

Build the train, validation, and test dataloaders.

class DeepMTP.data.DataPreparationState(validation_setting: str, split_method: str, split_ratio: tuple[tuple[str, float], ...], shuffle: bool, seed: int | None, dense: DensePreprocessingState)

Bases: object

Split provenance and fitted dense preprocessing from data_process.

dense: DensePreprocessingState
classmethod from_config(value: object) → DataPreparationState

Validate a saved data-preparation state mapping.

seed: int | None
shuffle: bool
split_method: str
split_ratio: tuple[tuple[str, float], ...]
to_dict() → dict[str, Any]

Return JSON-compatible checkpoint metadata.

validation_setting: str
class DeepMTP.data.DataProgressEvent(message: str, kind: Literal['message', 'operation_started', 'operation_completed'] = 'message', subject: str | None = None, level: Literal['info', 'warning', 'error'] = 'info', end: str = '\n', value: object | None = None)

Bases: object

One user-facing data preparation or validation message.

end: str = '\n'
kind: Literal['message', 'operation_started', 'operation_completed'] = 'message'
level: Literal['info', 'warning', 'error'] = 'info'
message: str
subject: str | None = None
value: object | None = None
class DeepMTP.data.DataProgressObserver(*args, **kwargs)

Bases: Protocol

Receives data preparation and validation messages.

on_event(event: DataProgressEvent) → None

Handle one data progress event.

class DeepMTP.data.DatasetBundle

Bases: TypedDict

Stable outer structure returned by legacy dataset helpers.

test: DatasetSplit
train: DatasetSplit
val: DatasetSplit
class DeepMTP.data.DatasetItem

Bases: TypedDict

One model-ready interaction returned by BaseDataset.

instance_features: Any
instance_id: int
score: Any
target_features: Any
target_id: int
class DeepMTP.data.DatasetSplit

Bases: TypedDict

One raw train, validation, or test split.

X_instance: ndarray | DataFrame | None
X_target: ndarray | DataFrame | None
y: ndarray | DataFrame | None
class DeepMTP.data.DenseBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A batch of dense feature vectors shaped [batch, features].

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'dense'
values: torch.Tensor
class DeepMTP.data.DensePreprocessingState(instance: DenseScalingState | None = None, target: DenseScalingState | None = None)

Bases: object

Dense scaling state for the instance and target feature axes.

for_branch(branch: str) → DenseScalingState | None

Return the scaler state for one entity axis.

classmethod from_config(value: object) → DensePreprocessingState

Validate a combined dense-preprocessing state mapping.

instance: DenseScalingState | None = None
target: DenseScalingState | None = None
to_dict() → dict[str, Any]

Return JSON-compatible checkpoint metadata.

class DeepMTP.data.DenseScalingState(method: Literal['MinMax', 'Standard'], feature_count: int, offsets: tuple[float, ...], scales: tuple[float, ...], data_minimums: tuple[float, ...] = (), data_maximums: tuple[float, ...] = (), feature_range: tuple[float, float] = (0.0, 1.0))

Bases: object

Parameters needed to replay one fitted dense-feature transformation.

data_maximums: tuple[float, ...] = ()
data_minimums: tuple[float, ...] = ()
feature_count: int
feature_range: tuple[float, float] = (0.0, 1.0)
classmethod from_config(value: object) → DenseScalingState

Validate and restore a JSON-compatible scaling-state mapping.

classmethod from_fitted_scaler(scaler: object, *, method: Literal['MinMax', 'Standard']) → DenseScalingState

Capture explicit numeric state from a fitted scikit-learn scaler.

method: Literal['MinMax', 'Standard']
offsets: tuple[float, ...]
scales: tuple[float, ...]
to_dict() → dict[str, Any]

Return JSON-compatible checkpoint metadata.

transform(values: object) → ndarray

Transform one vector or matrix using the saved training statistics.

class DeepMTP.data.FeatureComponentInfo

Bases: _FeatureComponentInfoRequired

Representation metadata for one component of a composite feature.

graph_edge_features: int | None
info: Literal['numpy', 'dataframe', 'graph', 'sequence', 'scipy_sparse', 'torch_sparse']
num_features: int | None
class DeepMTP.data.FeatureInfo

Bases: _FeatureInfoRequired

Normalized entity features and their detected representation.

components: dict[str, FeatureComponentInfo] | None
data: DataFrame
graph_edge_features: int | None
info: Literal['numpy', 'dataframe', 'images', 'graph', 'sequence', 'tabular', 'scipy_sparse', 'torch_sparse', 'composite']
num_features: int | None
class DeepMTP.data.GraphBranchBatch(values: GraphInput)

Bases: object

A structured batch consumed by a graph neural network encoder.

property batch_size: int

Number of graphs represented by this branch batch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'graph'
pin_memory() → GraphBranchBatch

Pin every tensor in the graph input.

to(device: torch.device | str, *, non_blocking: bool = False) → GraphBranchBatch

Move the graph input to one device.

values: GraphInput
class DeepMTP.data.GraphInput(node_features: torch.Tensor, edge_index: torch.Tensor, edge_features: torch.Tensor | None, graph_index: torch.Tensor, num_graphs: int)

Bases: object

One batch of homogeneous graphs represented by DeepMTP tensors.

property batch_size: int

Number of graphs represented by this input.

edge_features: torch.Tensor | None
edge_index: torch.Tensor
graph_index: torch.Tensor
node_features: torch.Tensor
num_graphs: int
pin_memory() → GraphInput

Pin every graph tensor for asynchronous device transfer.

to(device: torch.device | str, *, non_blocking: bool = False) → GraphInput

Move every graph tensor to one device.

class DeepMTP.data.IDBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A batch of zero-based integer entity IDs shaped [batch].

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'id'
values: torch.Tensor
class DeepMTP.data.ImageBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A batch of images shaped [batch, channels, height, width].

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'image'
values: torch.Tensor
class DeepMTP.data.InteractionInfo

Bases: TypedDict

Normalized interaction data and its detected schema.

data: DataFrame
instance_id_type: Literal['int']
missing_values: bool
original_format: Literal['triplets', 'numpy']
target_id_type: Literal['int']
class DeepMTP.data.MTPBatch(instance_id: torch.Tensor, target_id: torch.Tensor, instance_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch, target_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch, score: torch.Tensor)

Bases: Mapping[str, Any]

A complete interaction batch with mapping compatibility.

Attribute access exposes the typed branch inputs. Historical mapping keys continue to return the underlying tensors so existing dataloader consumers remain compatible.

property batch_size: int

Number of interactions represented by the batch.

property instance_features: torch.Tensor | CompositeInput | GraphInput | SequenceInput | TabularInput

Compatibility view of the instance branch input.

instance_id: torch.Tensor
instance_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch
pin_memory() → MTPBatch

Pin every tensor for asynchronous host-to-device transfer.

score: torch.Tensor
property target_features: torch.Tensor | CompositeInput | GraphInput | SequenceInput | TabularInput

Compatibility view of the target branch input.

target_id: torch.Tensor
target_input: CompositeBranchBatch | DenseBranchBatch | GraphBranchBatch | IDBranchBatch | ImageBranchBatch | SequenceBranchBatch | SparseBranchBatch | TabularBranchBatch | CustomBranchBatch
to(device: torch.device | str, *, non_blocking: bool = False) → MTPBatch

Move every tensor in the batch to one device.

class DeepMTP.data.MTPBatchCollator(instance_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite'], target_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite'], instance_sequence_padding_idx: int = 0, target_sequence_padding_idx: int = 0, instance_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None, target_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None, instance_component_padding_indices: Mapping[str, int] | None = None, target_component_padding_indices: Mapping[str, int] | None = None, instance_component_optional: Mapping[str, bool] | None = None, target_component_optional: Mapping[str, bool] | None = None)

Bases: object

Collate dataset items into a typed interaction batch.

instance_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None
instance_component_optional: Mapping[str, bool] | None = None
instance_component_padding_indices: Mapping[str, int] | None = None
instance_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']
instance_sequence_padding_idx: int = 0
target_component_kinds: Mapping[str, Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] | None = None
target_component_optional: Mapping[str, bool] | None = None
target_component_padding_indices: Mapping[str, int] | None = None
target_kind: Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']
target_sequence_padding_idx: int = 0
class DeepMTP.data.MTPSampler(*args: Any, **kwargs: Any)

Bases: Sampler

Undersample classes and provide a fresh order for every iteration.

micro balances classes globally, macro balances them independently for each target, and instance balances them independently for each instance. Grouped balancing can omit entities that are absent from the retained observations.

get_balanced_df(data: DataFrame) → DataFrame

Return one frame undersampled to its smallest class size.

class DeepMTP.data.MaskedComponentInput(values: Any | None, presence_mask: torch.Tensor)

Bases: object

Present values plus a full-batch mask for one optional component.

property batch_size: int

Full batch size including rows where this component is missing.

pin_memory() → MaskedComponentInput

Pin the present values and full-batch mask.

presence_mask: torch.Tensor
property present_count: int

Number of rows containing this component.

to(device: torch.device | str, *, non_blocking: bool = False) → MaskedComponentInput

Move the values and presence mask to one device.

values: Any | None
exception DeepMTP.data.NoveltyInferenceWarning

Bases: UserWarning

Warning emitted when entity identity was lost in matrix input.

class DeepMTP.data.NullDataProgressObserver

Bases: object

No-op observer used when data progress is disabled.

on_event(event: DataProgressEvent) → None
class DeepMTP.data.ProcessedDatasetBundle

Bases: TypedDict

Processed train, validation, and test data.

test: ProcessedDatasetSplit
train: ProcessedDatasetSplit
val: ProcessedDatasetSplit
class DeepMTP.data.ProcessedDatasetSplit

Bases: TypedDict

One processed split with optional interactions and side features.

X_instance: ProcessedDatasetValue | None
X_target: ProcessedDatasetValue | None
y: ProcessedDatasetValue | None
class DeepMTP.data.ProcessedDatasetValue

Bases: TypedDict

Processed data container consumed by legacy data utilities.

data: DataFrame
class DeepMTP.data.SequenceBranchBatch(values: SequenceInput)

Bases: object

A structured batch consumed by a token-sequence encoder.

property batch_size: int

Number of observations represented by this branch batch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'sequence'
pin_memory() → SequenceBranchBatch

Pin every tensor in the structured sequence input.

to(device: torch.device | str, *, non_blocking: bool = False) → SequenceBranchBatch

Move the structured sequence input to one device.

values: SequenceInput
class DeepMTP.data.SequenceInput(token_ids: torch.Tensor, attention_mask: torch.Tensor, lengths: torch.Tensor, padding_idx: int = 0)

Bases: object

Padded token IDs, attention mask, and original sequence lengths.

attention_mask: torch.Tensor
property batch_size: int

Number of token sequences represented by this input.

lengths: torch.Tensor
padding_idx: int = 0
pin_memory() → SequenceInput

Pin every sequence tensor for asynchronous device transfer.

to(device: torch.device | str, *, non_blocking: bool = False) → SequenceInput

Move every sequence tensor to one device.

token_ids: torch.Tensor
class DeepMTP.data.SparseBranchBatch(values: torch.Tensor)

Bases: TensorBranchBatch

A sparse feature batch shaped [batch, features].

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'sparse'
pin_memory() → SparseBranchBatch

Pin sparse indices and values for asynchronous device transfer.

values: torch.Tensor
class DeepMTP.data.TabularBranchBatch(values: TabularInput)

Bases: object

A structured batch of numeric values and categorical indices.

property batch_size: int

Number of observations represented by this branch batch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'tabular'
pin_memory() → TabularBranchBatch

Pin both tensors in the structured input.

to(device: torch.device | str, *, non_blocking: bool = False) → TabularBranchBatch

Move the structured input to one device.

values: TabularInput
class DeepMTP.data.TabularInput(numeric: torch.Tensor, categorical: torch.Tensor)

Bases: object

Numeric and categorical tensors consumed by a tabular encoder.

property batch_size: int

Number of observations in the structured input.

categorical: torch.Tensor
numeric: torch.Tensor
pin_memory() → TabularInput

Pin both tensors for asynchronous host-to-device transfer.

to(device: torch.device | str, *, non_blocking: bool = False) → TabularInput

Move both tensors to one device.

class DeepMTP.data.TabularPreprocessingState(numeric_columns: tuple[str, ...], categorical_columns: tuple[str, ...], numeric_fill_values: tuple[float, ...], numeric_offsets: tuple[float, ...], numeric_scales: tuple[float, ...])

Bases: object

Numeric statistics fitted only on training entities.

categorical_columns: tuple[str, ...]
classmethod from_config(value: object, *, schema: TabularSchema, branch: str) → TabularPreprocessingState

Restore and validate preprocessing metadata from a checkpoint config.

numeric_columns: tuple[str, ...]
numeric_fill_values: tuple[float, ...]
numeric_offsets: tuple[float, ...]
numeric_scales: tuple[float, ...]
to_dict() → dict[str, Any]

Return JSON-compatible checkpoint metadata.

class DeepMTP.data.TabularPreprocessor(schema: TabularSchema, state: TabularPreprocessingState, branch: str)

Bases: object

Transform entity feature frames using one schema and fitted state.

branch: str
classmethod fit(frame: DataFrame, schema: TabularSchema, *, branch: str) → TabularPreprocessor

Fit numeric imputation and normalization on a training frame.

classmethod from_state(schema: TabularSchema, state: Mapping[str, Any] | TabularPreprocessingState, *, branch: str) → TabularPreprocessor

Build a preprocessor from saved training metadata.

schema: TabularSchema
state: TabularPreprocessingState
transform(frame: DataFrame) → DataFrame

Return an indexed feature frame with structured tensor-ready values.

class DeepMTP.data.TabularSchema(numeric_columns: tuple[str, ...], categorical_columns: tuple[CategoricalColumn, ...], numeric_normalization: Literal['none', 'standard', 'minmax'] = 'standard', numeric_missing: Literal['error', 'mean', 'zero'] = 'mean', categorical_missing: Literal['error', 'unknown'] = 'unknown', categorical_unknown: Literal['error', 'unknown'] = 'unknown', feature_gating: bool = False)

Bases: object

Explicit numeric and categorical layout for one entity branch.

categorical_columns: tuple[CategoricalColumn, ...]
categorical_missing: Literal['error', 'unknown'] = 'unknown'
property categorical_names: tuple[str, ...]

Categorical column names in their stable tensor order.

categorical_unknown: Literal['error', 'unknown'] = 'unknown'
property encoded_width: int

Width after concatenating numeric values and category embeddings.

feature_gating: bool = False
classmethod from_config(value: object, *, branch: str) → TabularSchema

Validate a public tabular schema mapping.

numeric_columns: tuple[str, ...]
numeric_missing: Literal['error', 'mean', 'zero'] = 'mean'
numeric_normalization: Literal['none', 'standard', 'minmax'] = 'standard'
to_dict() → dict[str, Any]

Return a JSON-compatible public configuration mapping.

class DeepMTP.data.TensorBranchBatch(values: torch.Tensor)

Bases: object

One branch’s tensor payload with an explicit modality contract.

property batch_size: int

Number of observations represented by this branch batch.

kind: ClassVar[Literal['dense', 'id', 'image', 'graph', 'sequence', 'sparse', 'tabular', 'custom', 'composite']] = 'custom'
pin_memory() → TensorBranchBatch

Return the same typed batch backed by pinned host memory.

to(device: torch.device | str, *, non_blocking: bool = False) → TensorBranchBatch

Return the same typed batch with its tensor moved to device.

values: torch.Tensor
class DeepMTP.data.TrainingDataLoaders(train: torch.utils.data.DataLoader, validation: torch.utils.data.DataLoader, test: torch.utils.data.DataLoader)

Bases: object

Dataloaders required by one training run.

test: torch.utils.data.DataLoader
train: torch.utils.data.DataLoader
validation: torch.utils.data.DataLoader
class DeepMTP.data.Transformer(*args, **kwargs)

Bases: Protocol

Minimal interface implemented by supported feature scalers.

transform(values: Any) → Any

Transform values using a previously fitted scaler.

DeepMTP.data.build_data_progress_observer(enabled: bool, *, print_mode: Literal['basic', 'dev'] = 'basic') → DataProgressObserver

Select the default data progress observer.

DeepMTP.data.check_interaction_files_column_type_format(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None

Checks the type of the instance and target ids in the interaction data. If any inconsistencies are detected (like different id types between train and test interaction data), an exception is raised

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose ID-type validation messages. Defaults to None.

DeepMTP.data.check_interaction_files_format(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None

Checks the format of the interaction data. If any inconsistencies are detected (like different formats between train and test interaction data), an exception is raised

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose format-validation messages. Defaults to None.

DeepMTP.data.check_novel_instances(train: Mapping[str, Any] | None, test: Mapping[str, Any] | None, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → bool | None

Return whether test interactions contain previously unseen instances.

progress receives messages when verbose is enabled.

DeepMTP.data.check_novel_targets(train: Mapping[str, Any] | None, test: Mapping[str, Any] | None, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → bool | None

Return whether test interactions contain previously unseen targets.

progress receives messages when verbose is enabled.

DeepMTP.data.check_target_variable_type(data: Mapping[str, Any], verbose: bool = False, print_mode: str = 'basic', *, classification_mode: Literal['binary', 'multiclass'] | None = None, num_classes: int | None = None, progress: DataProgressObserver | None = None) → Literal['binary', 'multiclass', 'real-valued']

Checks the type of the target variable in the interaction data. If any inconsistencies are detected (like different id types between train and test interaction data), an exception is raised

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose target-type validation messages. Defaults to None.

DeepMTP.data.check_variable_type(samples_arr: DataFrame, *, classification_mode: Literal['binary', 'multiclass'] | None = None, num_classes: int | None = None) → Literal['binary', 'multiclass', 'real-valued']

Detect binary or real-valued scores, or validate declared class IDs.

Parameters:

samples_arr (numpy.array) – A numpy array with the target variables

Returns:

binary, multiclass, or real-valued.

Return type:

str

Multiclass targets are intentionally opt-in because an integer-valued regression target is otherwise indistinguishable from class IDs.

DeepMTP.data.cross_input_consistency_check_instances(data: Mapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None

Checks the consistency of instance ids in the interaction data and instance features. The requirements to pass this check change depending on the format of the interaction data and instance features

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • validation_setting (str) – The validation setting of the current problem.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose instance-consistency messages. Defaults to None.

DeepMTP.data.cross_input_consistency_check_targets(data: Mapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None

Checks the consistency of target ids in the interaction data and target features. The requirements to pass this check change depending on the format of the interaction data and target features

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • validation_setting (str) – The validation setting of the current problem.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose target-consistency messages. Defaults to None.

DeepMTP.data.data_process(data: Mapping[str, Any], validation_setting: str | None = None, split_method: str = 'random', ratio: object = None, shuffle: bool = True, seed: int | None = 42, verbose: bool = False, print_mode: str = 'basic', scale_instance_features: str | None = None, scale_target_features: str | None = None, *, classification_mode: str | None = None, num_classes: int | None = None, instance_feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None, target_feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None, progress: DataProgressObserver | None = None, dense_preprocessing_state: Mapping[str, Any] | DataPreparationState | DensePreprocessingState | None = None) → tuple[dict[str, Any], dict[str, Any], dict[str, Any], DataInfo]

The main function that handles all the preprocessing steps and checks needed to prepare the dataset to be used by the model.

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • validation_setting (str, optional) – The validation setting of the current problem. Defaults to None.

  • split_method (str, optional) – The splitting method used. The current implementation only supports the ‘random split’ using a specific seed but a future goal is to also offer a stratified option. Defaults to ‘random’.

  • ratio (dict, optional) – The train, val and test ratios used to split the data. Defaults to {‘train’: 0.7, ‘test’: 0.2, ‘val’: 0.1}.

  • shuffle (bool, optional) – Whether or not the dataset will be shuffled before the split. If is set to False, the seed value is not used.

  • seed (int, optional) – The seed used to initiate the randomized split. Defaults to 42.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • scale_instance_features (str, optional) – The scaler used for the instance features. Possible values are ‘MinMax’ for the MinMax scaler and ‘Standard’ for the standard scaler. Defaults to None.

  • scale_target_features (str, optional) – The scaler used for the target features. Possible values are ‘MinMax’ for the MinMax scaler and ‘Standard’ for the standard scaler. Defaults to None.

  • classification_mode (str, optional) – Set to multiclass to interpret zero-based integer scores as class IDs. Binary scores continue to be detected automatically. Defaults to None.

  • num_classes (int, optional) – Declared multiclass count. When omitted in multiclass mode, it is inferred from all supplied score splits.

  • instance_feature_kind (str, optional) – Set to sequence when instance features are already-tokenized integer sequences. For composite features, pass a mapping from component name to sequence, None, or a mapping with kind and optional fields. Defaults to None.

  • target_feature_kind (str, optional) – Set to sequence when target features are already-tokenized integer sequences. For composite features, pass a mapping from component name to sequence, None, or a mapping with kind and optional fields. Defaults to None.

  • progress (DataProgressObserver | None, optional) – Receives verbose messages from the complete data-processing pipeline. Defaults to None.

  • dense_preprocessing_state (mapping, optional) – Previously saved dense scaler parameters. Accepts either the dense section or the full data_preparation_state returned in data_info. When supplied, the saved training statistics are replayed instead of fitting new scalers. Defaults to None.

Returns:

Four different dictionaries containing:
  • Train processed data

  • Validation processed data

  • Test processed data

  • general information about the datasets

Return type:

dict, dict, dict, dict

DeepMTP.data.format_mtr_datasets() → str

Format the catalog of supported multivariate regression datasets.

DeepMTP.data.generate_MTP_dataset(num_instances: int, num_targets: int, num_instance_features: int | None = None, num_target_features: int | None = None, split_instances: Mapping[str, float] | None = None, split_targets: Mapping[str, float] | None = None, return_static_features_data: bool = False) → ProcessedDatasetBundle

Generate deterministic triplet data with optional axis-level splits.

Each split mapping requires a test ratio and may include a validation ratio. The train ratio is derived when omitted and verified when supplied. A test-only axis reuses its training IDs in a validation split created on the other axis.

Parameters:
  • num_instances – Number of instances.

  • num_targets – Number of targets.

  • num_instance_features – Number of generated instance features, or none.

  • num_target_features – Number of generated target features, or none.

  • split_instances – Optional instance-axis split ratios.

  • split_targets – Optional target-axis split ratios.

  • return_static_features_data – Repeat non-novel side features in holdouts.

Returns:

Processed triplet and feature frames for train, validation, and test.

Raises:
  • TypeError – If a split or boolean flag has the wrong type.

  • ValueError – If a dimension or split ratio is invalid.

DeepMTP.data.generate_dummy_dataset(num_instances: int, num_targets: int, num_instance_features: int, num_target_features: int, error_mu: float, error_sigma: float, sklearn_version: bool = False, seed: int = 42, mode: str = 'u+logv', split_ratio: Mapping[str, float] | None = None) → DatasetBundle

Generate a reproducible dyadic regression dataset.

sklearn_version starts from normally distributed features. For modes involving logarithms or fractional powers, the relevant features are transformed to a positive domain so every generated score remains real. The error distribution applies to the four legacy noisy relationship modes.

Parameters:
  • num_instances – Number of instances.

  • num_targets – Number of targets.

  • num_instance_features – Number of features per instance.

  • num_target_features – Number of features per target.

  • error_mu – Mean of the Gaussian error added by noisy modes.

  • error_sigma – Non-negative standard deviation of that Gaussian error.

  • sklearn_version – Generate normally distributed rather than uniform features.

  • seed – Non-negative 32-bit seed used for generation and splitting.

  • mode – Mathematical relationship used to generate scores.

  • split_ratio – Train, validation, and test ratios.

Returns:

A dataset bundle containing aligned train, validation, and test rows.

Raises:
  • TypeError – If a boolean or error parameter has the wrong type.

  • ValueError – If a dimension, mode, seed, ratio, or generated value is invalid.

DeepMTP.data.generate_interaction_matrix(input_path: str | PathLike[str], output_path: str | PathLike[str]) → None

Convert annotator labels into an atomically written correctness matrix.

DeepMTP.data.get_estimated_validation_setting(novel_instances_flag: bool | None, novel_targets_flag: bool | None, verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → Literal['A', 'B', 'C', 'D'] | None

Uses the combination of infrmation about novel instances and novel targets to determine the validation setting that is possible.

Parameters:
  • novel_instances_flag (bool) – A boolean that indicates whether or not the test set contains novel instances.

  • novel_targets_flag (bool) – A boolean that indicates whether or not the test set contains novel targets.

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose validation-setting inference messages. Defaults to None.

Returns:

The validation setting. Possible values are A, B, C, D

Return type:

str

DeepMTP.data.load_process_DP(path: str | PathLike[str] = './data', dataset_name: str = 'ern', variant: Literal['undivided', 'divided'] = 'undivided', random_state: int | None = 42, split_ratio: Mapping[str, float] | None = None, split_instance_features: bool = False, split_target_features: bool = False, validation_setting: str = 'B', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load a dyadic-prediction dataset and optionally create held-out splits.

Parameters:
  • path – Directory used for the local dataset cache.

  • dataset_name – One of "ern", "srn", "dpie", or "dpii".

  • variant – Return the complete matrices or create train/validation/test splits.

  • random_state – Seed used by the split operations, or None.

  • split_ratio – Train, validation, and test proportions.

  • split_instance_features – Split instance features with interaction rows.

  • split_target_features – Split target features with interaction columns.

  • validation_setting – Novelty setting "B", "C", or "D".

  • print_mode – "basic" for ordinary messages or "dev" for messages prefixed for the development application.

  • progress – Receives dataset loading and download status events.

Raises:
  • ValueError – If an option or dataset matrix is invalid.

  • TypeError – If a feature-splitting flag is not boolean.

  • DatasetDownloadError – If a required dataset file cannot be downloaded.

  • FileNotFoundError – If a download does not create all required files.

DeepMTP.data.load_process_MC(path: str | PathLike[str] = './data', dataset_name: str = 'ml-100k', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load a matrix-completion dataset from the MovieLens repository.

Parameters:
  • path – Directory used for the local dataset cache.

  • dataset_name – Matrix-completion dataset to load.

  • print_mode – "basic" for ordinary messages or "dev" for messages prefixed for the development application.

  • progress – Receives dataset download status events.

Raises:
  • ValueError – If an option or the ratings data is invalid.

  • DatasetDownloadError – If the archive cannot be downloaded or extracted.

  • FileNotFoundError – If the archive does not contain the expected ratings file.

Returns:

The ratings triplets and empty side-feature and held-out splits.

DeepMTP.data.load_process_MLC(path: str | PathLike[str] = './data', dataset_name: str = 'bibtex', variant: Literal['undivided', 'divided'] = 'undivided', features_type: Literal['numpy', 'dataframe'] = 'numpy', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load a cached or downloadable scikit-multilearn benchmark dataset.

Parameters:
  • path – Directory used for the local dataset cache.

  • dataset_name – Supported multi-label benchmark name.

  • variant – Return the complete dataset or its published train/test split.

  • features_type – Return instance features as an array or feature DataFrame.

  • print_mode – "basic" for ordinary messages or "dev" for messages prefixed for the development application.

  • progress – Receives dataset loading and download status events.

Raises:
  • ValueError – If an option or cached dataset is invalid.

  • DatasetDownloadError – If required cache files cannot be downloaded safely.

  • IsADirectoryError – If an expected cache file path is a directory.

DeepMTP.data.load_process_MTL(path: str | PathLike[str] = './data', dataset_name: str = 'dog', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load multi-task learning datasets from my custom repository.

Parameters:
  • path (str, optional) – The path where the datasets should be stored. If it doesn’t exist, download and store it in this directory. Defaults to ‘./data’.

  • dataset_name (str, optional) – The name of the multi-task learning dataset. Defaults to ‘dog’.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives dataset download status events. Defaults to None.

Raises:
  • ValueError – If a public option or dataset file is invalid.

  • DatasetDownloadError – If the dataset archive cannot be downloaded safely.

  • FileNotFoundError – If required score or image files are missing.

Returns:

A dictionary with all the available data for the multi-task learning dataset.

Return type:

dict

DeepMTP.data.load_process_MTR(path: str | PathLike[str] = './data', dataset_name: str = 'enb', features_type: Literal['numpy', 'dataframe'] = 'numpy', print_mode: Literal['basic', 'dev'] = 'basic', *, progress: DataProgressObserver | None = None) → DatasetBundle

Load multivariate regression datasets from the mulan repository.

Parameters:
  • path (str, optional) – The path where the datasets should be stored. If it doesn’t exist, download and store it in this directory. Defaults to ‘./data’.

  • dataset_name (str, optional) – The name of the multivariate regression dataset. Defaults to ‘enb’.

  • features_type (str, optional) – The format of the instance features. There are two possible values, numpy and dataframe. This is intended to test the functionality of the data_process function. Defaults to ‘numpy’.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives dataset loading and download status events. Defaults to None.

Raises:
  • ValueError – If a public option or downloaded dataset is invalid.

  • DatasetDownloadError – If the dataset archive cannot be downloaded safely.

  • FileNotFoundError – If the archive does not contain the requested dataset.

Returns:

A dictionary with all the available data for the multivariate regression dataset.

Return type:

dict

DeepMTP.data.normalize(row: Any, scaler: Transformer) → Any

Just normalizes a row or features

Parameters:
  • row (numpy.array) – an array of features

  • scaler (sklearn.scaler) – the scaler that will be used to scale the features

Returns:

a scaled array of feature

Return type:

numpy.array

DeepMTP.data.print_MTR_datasets(*, progress: DataProgressObserver | None = None) → None

Present the supported multivariate regression dataset catalog.

The default observer preserves the historical console table. Inject NullDataProgressObserver to suppress output or another data progress observer to render the catalog elsewhere.

DeepMTP.data.process_dummy_DP(num_instance_features: int = 10, num_target_features: int = 3, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', instance_features_format: Literal['numpy', 'dataframe'] = 'numpy', target_features_format: Literal['numpy', 'dataframe'] = 'numpy', random_state: int | None = 42) → DatasetBundle

Generate a reproducible dummy dyadic prediction dataset.

Parameters:
  • num_instance_features – Number of features per instance.

  • num_target_features – Number of features per target.

  • num_instances – Number of instances.

  • num_targets – Number of targets.

  • interaction_matrix_format – Return scores as a dense matrix or triplets.

  • instance_features_format – Return instance features as an array or DataFrame.

  • target_features_format – Return target features as an array or DataFrame.

  • random_state – Seed for an isolated NumPy generator. Use None for non-deterministic data.

Returns:

A dataset bundle containing an undivided training split.

Raises:

ValueError – If a format, dimension, or random state is invalid.

DeepMTP.data.process_dummy_MLC(num_features: int = 10, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', features_format: Literal['numpy', 'dataframe'] = 'numpy', random_state: int | None = 42) → DatasetBundle

Generate a reproducible dummy multi-label classification dataset.

Parameters:
  • num_features – Number of features per instance.

  • num_instances – Number of instances.

  • num_targets – Number of binary labels.

  • interaction_matrix_format – Return labels as a dense matrix or triplets.

  • features_format – Return features as a dense matrix or feature DataFrame.

  • random_state – Seed for an isolated NumPy generator. Use None for non-deterministic data.

Returns:

A dataset bundle containing an undivided training split.

Raises:

ValueError – If a format, dimension, or random state is invalid.

DeepMTP.data.process_dummy_MTR(num_features: int = 10, num_instances: int = 50, num_targets: int = 5, interaction_matrix_format: Literal['numpy', 'dataframe'] = 'numpy', features_format: Literal['numpy', 'dataframe'] = 'numpy', variant: Literal['undivided', 'divided'] = 'undivided', split_ratio: Mapping[str, float] | None = None, random_state: int | None = 42) → DatasetBundle

Generate a reproducible dummy multivariate regression dataset.

Parameters:
  • num_features – Number of features per instance.

  • num_instances – Number of instances.

  • num_targets – Number of regression targets.

  • interaction_matrix_format – Return scores as a dense matrix or triplets.

  • features_format – Return features as a dense matrix or feature DataFrame.

  • variant – Return an undivided dataset or train/validation/test splits.

  • split_ratio – Ratios used for the divided variant.

  • random_state – Seed used for both data generation and splitting. Use None for non-deterministic data.

Returns:

A dataset bundle containing the generated scores and features.

Raises:
  • TypeError – If split_ratio is not a mapping.

  • ValueError – If an option, dimension, ratio, or random state is invalid.

DeepMTP.data.process_instance_features(instance_features: Any, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None, feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None) → FeatureInfo | None

Normalize instance features without mutating caller-owned inputs.

Composite component specifications may set optional=True to align a subset of entity IDs and preserve missing rows as None. progress receives messages when verbose is enabled.

DeepMTP.data.process_interaction_data(interaction_data: DataFrame | ndarray, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → InteractionInfo

A function that processes the interaction data. It is called separately for the train, val and test interaction data. There are two main types of interaction data formats that are supported: -> numpy format: This is a 2d numpy array that represents the interaction (also called score) matrix usually found in problems settings with fully observed matrices (multi-label classification, multivariate regression) -> triplet format: The most flexible format as it can be used to represent every possible problem setting

Parameters:
  • interaction_data (_type_) – a numpy array or a dataframe with the interaction data

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose interaction format messages. Defaults to None.

Returns:

a dictionary with the interaction data and additional information that could be detected (format, type of instance and target ids, etc.)

Return type:

dict

DeepMTP.data.process_target_features(target_features: Any, verbose: bool = False, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None, feature_kind: Literal['graph', 'sequence'] | None | Mapping[str, Literal['graph', 'sequence'] | None | CompositeFeatureSpec] = None) → FeatureInfo | None

Normalize target features without mutating caller-owned inputs.

Composite component specifications may set optional=True to align a subset of entity IDs and preserve missing rows as None. progress receives messages when verbose is enabled.

DeepMTP.data.split_data(data: MutableMapping[str, Any], validation_setting: Literal['A', 'B', 'C', 'D'], split_method: str, ratio: object, shuffle: bool, seed: int | None, verbose: bool, print_mode: str = 'basic', *, progress: DataProgressObserver | None = None) → None
Splits the dataset and offers two main functionalities:
  • split based on the 4 different validation settings (A, B, C, D)

  • if a test set already exists it separates a validation set, otherwise it first creates a test set and then a validation set.

Parameters:
  • data (dict) – The dictionary that store all possible data available in a multi-target prediction dataset.

  • validation_setting (str) – The validation setting of the current problem.

  • split_method (str) – The splitting method used. The current implementation only supports the ‘random split’ using a specific seed but a future goal is to also offer a stratified option.

  • ratio (dict) – The train, val and test ratios used to split the data.

  • seed (int) – The seed used to initiate the randomized split

  • verbose (bool, optional) – Whether or not to print usefull info in the terminal. Defaults to False.

  • print_mode (str, optional) – The mode of printing. Two values are possible. If ‘basic’, the prints are just regural python prints. If ‘dev’ then a prefix is used so that the streamlit application can print more usefull messages. Defaults to ‘basic’.

  • progress (DataProgressObserver | None, optional) – Receives verbose split lifecycle messages. Defaults to None.

DeepMTP.data.transform_dense_features(values: object, state: Mapping[str, Any] | DenseScalingState) → ndarray

Replay a saved dense-feature transformation.

DeepMTP.data.validate_split_ratio(ratio: object) → dict[str, float]

Validate and copy a train/validation/test split ratio mapping.