test_dataset_RAW

Erik Képeš · Figshare · 2022

## 1 FILE CONTENTS ############################################# 751 x 40003 values: - 1 header row, which corresponds to the 40002 resolved wavelengths and 1 target identifier. - 750 spectra (rows) of the targets whose content is to be predicted: 50 spectra of 15 distinct targets. Note that the dataset occupies significantly less RAM when loaded than the file's disk sizes (cca. 50 %). Note that the column names may be modified by the software/library used for processing the spectra. Hence, we advise increased caution when extracting the wavelengths. Alternatively, you can use the wavelength values provided separately in the wavelengths.csv file. Note that the target identifier column is a character vector of the format "target_i", i = {1,2,...,15}. Thus, we recommended loading in the dataset as a dataframe instead of a matrix (a dataframe can contain both numeric and character values). ## 2 DATA ACQUISITION ############################################# The dataset was collected from metallic targets. Each target's surface was homogenized using a 800 grid sandpaper and cleaned with isopropyl alcohol using paper towels. The targets were sampled at distinct spots in single-shot mode, collecting a single spectrum from each individual spot. There are a total of 50 spectra available for each target. A 1064 nm Nd:YAG laser was used for ablation with a pulse energy of 95 mJ, 10 ns pulse width. The laser pulse was focused into a spot with a diameter of 0.2 mm. The collected emission was resolved with an echelle spectrograph in the 240--1000 nm spectral range (40002 resolved wavelength values), resolving power {\lambda} / {\Delta\lambda} = 6000. The resolved emission was recorded using an EMCCD camera with a 1.5 us delay and 50 us gate width (exposition time). The spectra were not intensity calibrated! This, in combination with the long gate width is part of the challenge.

Read the paper · More papers on PaperTik