On the benefits of self-taught learning for brain decoding - Data

Élodie Germani, Fromont Elisa, Camille Maumet · Zenodo (CERN European Organization for Nuclear Research) · 2022

DERIVED DATA FROM PAPER "On the benefits of self-taught learning for brain decoding" Here are stored the data necessary to reproduce the full analysis of the paper "On the benefits of self-taught learning for brain decoding". We study the benefits of using a large public neuroimaging database composed of fMRI statistic maps, in a self-taught learning framework, for improving brain decoding on new tasks. First, we leverage the NeuroVault database to train, on a selection of relevant statistic maps, a convolutional autoencoder to reconstruct these maps. Then, we use this trained encoder to initialize a supervised convolutional neural network to classify tasks or cognitive processes of unseen statistic maps from large collections of the NeuroVault database. We show that such a self-taught learning process always improves the performance of the classifiers but the magnitude of the benefits strongly depends on the number of data available both for pre-training and finetuning the models and on the complexity of the targeted downstream task. Contents overview 1. original The original directory contains 3 subdirectories: - NeuroVault dataset - HCP dataset - BrainPedia dataset Each subdirectory contains: - text files with NeuroVault IDs of statistic maps selected in the global datasets, in the test and validation datasets and in each fold of these datasets ; - csv files corresponding to informations on each statistic map of the datasets (classification labels, subject IDs for split...) ; - an `original` directory in which original statistic maps downloaded from NeuroVault will be stored when executing the `src/download_and_preprocess_data_notebook.ipynb`. 2. preprocessed The preprocessed directory contains 3 subdirectories: - NeuroVault dataset - HCP dataset - BrainPedia dataset Each subdirectory contains: - text files with NeuroVault IDs of statistic maps selected in the global datasets, in the test and validation datasets and in each fold of these datasets ; - csv files corresponding to informations on each statistic map of the datasets (classification labels, subject IDs for split...) ; - several subdirectores (`resampled`, `resampled masked`...) in which preprocessed statistic maps will be stored when executing the `src/download_and_preprocess_data_notebook.ipynb`. 3. derived The derived directory contains 3 subdirectories: - NeuroVault dataset - HCP dataset - BrainPedia dataset Each subdirectory contains subdirectories in which the parameters of models trained on the different datasets are stored. These subdirectories are named in the following way: {name_of_the_dataset}_maps_classification_{classification_task}_model_cnn_{model_architecture}_valid_{type_of_experiment}_retrain_{type_of_initialization}_{preprocessing_type}_epochs_{number_of_epochs}_batch_size_{batch_size}_lr_{learning_rate} For instance, parameters for the following experiment: - Dataset: HCP Dataset subset 50 subjects - Classification task: contrast classification - Model: 4 layers CNN - Type of experiment: Performance evaluation - Initialization: Default - Preprocessing type: Resampled masked normalized - Epochs: 500 - Batch: 32 - Learning rate: 1e-04 will be contained in the directory: hcp_dataset_50_maps_classification_contrast_model_cnn_4layers_valid_perf_retrain_no_resampled_masked_normalized_epochs_500_batch_size_32_lr_1e-04

Read the paper · More papers on PaperTik