Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “input data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Considerations for using Privacy Preserving Machine Learning Techniques for Safeguards

In international nuclear safeguards, the International Atomic Energy Agency (IAEA) is tasked with inspecting and verifying nuclear facilities and their activities. Data analytics and machine learning to support inspections require large amounts of data that nuclear facility operators may consider proprietary or sensitive, so the IAEA may not have full access. Allowing computation over private data without compromising its security therefore has value for safeguards inspections and analysis. Privacy-preserving machine learning (PPML) consists of security-focused techniques that allow data analytics and machine learning algorithms to run on sensitive data without revealing it. This includes ideas like homomorphic encryption (HE), secure multiparty computation (SMPC), and secure enclaves. HE allows algorithms and mathematical operations to be conducted directly on the encrypted data instead of first decrypting it. With SMPC, multiple entities collaboratively compute over distributed data such that no party is able to directly view any others’ original data. Secure enclaves allow computation to take place in a separate and heavily blocked-off section of a CPU. Techniques like these allow for several potential use cases in which the security of data is essential. With SMPC, machine learning models can be trained over the input data from multiple entities, resulting in a model that all users can benefit from without leaking the input data from any particular entity. With SMPC or a zero-knowledge proof (ZKP), an algorithm returning some single answer or truth value can be run on someone else’s data without ever needing to see that data, potentially allowing for verification or proof of some underlying question. HE can allow for outsourcing computation on data to a hostile or untrusted environment. Although most of the research in this field resides within the health and financial domains, tools from PPML may have similar applications in nuclear safeguards. Allowing the IAEA to compute over proprietary information, such as process models and raw sensor data using PPML techniques, provides the baseline for running complex analytics without needing direct unencrypted access to the underlying data, maintaining its privacy. Important limitations to consider for these techniques include the efficiency and level of security required. The security of HE and SMPC come at the cost of speed—the significant amount of overhead means that algorithms implemented in these protocols and encryption schemes are slower than when run on plaintext. Additionally, several important parameters determine what techniques or protocols are used based on the security requirements. SMPC protocols may need to be selected for resistance against a party that attempts to deviate from the protocol to distort the result or gain access to additional information, and a protocol secure against these attacks may further increase the overhead of the algorithm.

97 MATHEMATICS AND COMPUTING↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

A comparison of template vs. direct model fitting for redshift-space distortions in BOSS

The growth of large-scale structure, as revealed in the anisotropic of clustering of galaxies in the low redshift Universe at z < 2, provides a stringent test of our cosmological model. The strongest current constraints come from the BOSS and eBOSS surveys, with uncertainties on σ 8 , the amplitude of clustering on an 8 h –1 Mpc scale, of less than 10 per cent. A number of different approaches have been taken to fitting this signal, leading to discrepancies of up to 1σ in the measurements of the amplitude of fluctuations at late times. We compare in some detail two of the leading approaches, one based on fitting a template cosmology whose amplitude and length scales are allowed to float with one based on varying the underlying parameters of a cosmological model directly, when fitting to the BOSS DR12 data. Holding the input data, scale cuts, window functions and modeling framework fixed we are able to isolate the cause of the differences and discuss the implications for future surveys.

redshift surveys↗

FPGA-based tracking for the CMS Level-1 trigger using the tracklet algorithm

The high instantaneous luminosities expected following the upgrade of the Large Hadron Collider (LHC) to the High-Luminosity LHC (HL-LHC) pose major experimental challenges for the CMS experiment. A central component to allow efficient operation under these conditions is the reconstruction of charged particle trajectories and their inclusion in the hardware based trigger system. There are many challenges involved in achieving this: a large input data rate of about 20–40 Tb/s; processing a new batch of input data every 25 ns, each consisting of about 15,000 precise position measurements and rough transverse momentum measurements of particles (“stubs”); performing the pattern recognition on these stubs to find the trajectories; and producing the list of trajectory parameters within 4 µs. Here, this paper describes a proposed solution to this problem, specifically, it presents a novel approach to pattern recognition and charged particle trajectory reconstruction using an all-FPGA solution. The results of an end-to-end demonstrator system, based on Xilinx Virtex-7 FPGAs, that meets timing and performance requirements are presented along with a further improved, optimized version of the algorithm together with its corresponding expected performance.

47 OTHER INSTRUMENTATION↗

Global maximum carboxylation rate trait scaling in the Sheffield Dynamic Global Vegetation Model 1901-2012

This data package contains model output data and climate input data associated with the New Phytologist publication: Walker, A.P., Quaife, T., van Bodegom, P.M., De Kauwe, M.G., Keenan, T.F., Joiner, J., Lomas, M.R., MacBean, N., Xu, C., Yang, X., Woodward, F.I., 2017. The impact of alternative trait-scaling hypotheses for the maximum photosynthetic carboxylation rate (Vcmax) on global gross primary production. New Phytologist 215, 1370–1386. https://doi.org/10.1111/nph.14623. This dataset contains maximum carboxylation rate at 25 ºC (Vcmax,25), Gross Primary Production (GPP), Net Biome Production (NBP), vegetation carbon, and soil carbon data from Sheffield Dynamic Global Vegetation Model (SDGVM) for multiple trait-scaling methods for scaling Vcmax,25 across Earth, and alternative instantaneous temperature scaling functions. Vcmax,25 values are mean top-leaf values taken during the growing seasons, which were defined as periods during which leaf area index (LAI) was greater than one. Vcmax,25 values were reported prior to scaling of Vcmax by water-stress or leaf-age. The attached metadata file (SDGVM_vcmax_dataset.docx) describes the methods used and explains the organization of data within the included .zip file. This metadata file is also included in PDF format.

54 ENVIRONMENTAL SCIENCES↗

Identifying the nature of the QCD transition in relativistic collision of heavy nuclei with deep learning

Using deep convolutional neural network (CNN), the nature of the QCD transition can be identified from the final-state pion spectra from hybrid model simulations of heavy-ion collisions that combines a viscous hydrodynamic model with a hadronic cascade “after-burner”. Two different types of equations of state (EoS) of the medium are used in the hydrodynamic evolution. The resulting spectra in transverse momentum and azimuthal angle are used as the input data to train the neural network to distinguish different EoS. Different scenarios for the input data are studied and compared in a systematic way. A clear hierarchy is observed in the prediction accuracy when using the event-by-event, cascade-coarse-grained and event-fine-averaged spectra as input for the network, which are about 80%, 90% and 99%, respectively. A comparison with the prediction performance by deep neural network (DNN) with only the normalized pion transverse momentum spectra is also made. High-level features of pion spectra captured by a carefully-trained neural network were found to be able to distinguish the nature of the QCD transition even in a simulation scenario which is close to the experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrated GW Farm ABM

This Data Repository includes data used for the integrated groundwater- farm ABM model, raw model output from scenario ensemble, and processed outputs that isolate the groundwater storage depletion outcomes for the 35,000 farm cells. Model Inputs: Farm ABM Inputs: This folder contains the input data used by the integrated groundwater - farm ABM modelling script (Python file) used for the high performance computing (HPC) experiments. The sub-folder "data inputs" contains all of the farm attribute data, while the three files in the folder have the hydrogeological data lookup table (NLDAS Cost Curve Attributes.csv), a lookup table (Theis well function table.csv) for the groundwater cost curve function, and the farm indexes and corresponding NLDAS ids for all of the cells run in this experiment (nldas farms subset final.csv). NLDAS Cost curve hydrogeological data: Hydrogeological data aggregated to 1/8 degree resolution and aligned with the NLDAS grid. Parameters include: water depth below ground surface [meters], subsurface porosity [unitless], aquifer depth from ground surface to aquifer bottom [meters], annual average recharge (USGS: mm, Doll: meters), and three different hydraulic conductivity (K) values (meters/day). The three K values represent the mean value from Gleeson et al. (2018), one standard deviation above the mean from Gleeson et al. (2018), and the de Graaf et al. 2020 modifications to certain lithologies. Additional information about these datasets and their processing are documented in the supplement to Yoon et al. 2025 (in review). Output: Raw outputs: This folder contains a .zip file that has model outputs for the entire scenario ensemble. There is one csv for each farm id, using the format "farm farmid cases.csv". The relationship between the farm id and NLDAS id is defined by the "nldas farms subset final.csv" located in the Farm ABM Inputs folder. Each csv has 625 rows, corresponding to 625 combinations of different scenario parameter values. Each row (scenario) represents the outcome of a 100 year simulation. Columns define scenario settings and summary statistics for each scenario. The first four columns define the scenario settings: "hydro ratio," "econ ratio," "K scenario," and "gamma scenario." The hydro and econ ratios are values passed to the modeling script that influence multipliers for other model parameters, as documented in the supplement to Yoon et al. 2025 (in review). The gamma multiplier is a coefficient multiplier applied to the baseline gamma values (values below 1 represent lower unobserved costs compared to baseline, values above 1 represent higher costs). The K scenario names represent K values of: "low": 0.5 m/d, "int 1": 2.5 m/d, "int 2": 10 m/d, "high": 50 m/d, and "gleeson": mean Gleeson K value. "Perc vol depleted" is the fraction of groundwater depleted at the end of the 100 simulation. Processed Output: Derived depletion outcomes from raw outputs: All of the individual csv files from the Raw outputs were aggregated into a single file that has the scenario settings and fraction depletion "Perc vol depleted" for every farm cell, for every scenario. The other two files define relationships between the farm id, NLDAS id, and local and major aquifer units, used for aquifer-level depletion analysis.

Agent based modeling↗

Mixture Model Framework for Traumatic Brain Injury Prognosis Using Heterogeneous Clinical and Outcome Data

Prognoses of Traumatic Brain Injury (TBI) outcomes are neither easily nor accurately determined from clinical indicators. This is due in part to the heterogeneity of damage inflicted to the brain, ultimately resulting in diverse and complex outcomes. Using a data-driven approach on many distinct data elements may be necessary to describe this large set of outcomes and thereby robustly depict the nuanced differences among TBI patients’ recovery. In this work, we develop a method for modeling large heterogeneous data types relevant to TBI. Our approach is geared toward the probabilistic representation of mixed continuous and discrete variables with missing values. The model is trained on a dataset encompassing a variety of data types, including demographics, blood-based biomarkers, and imaging findings. In addition, it includes a set of clinical outcome assessments at 3, 6, and 12 months post-injury. The model is used to stratify patients into distinct groups in an unsupervised learning setting. We use the model to infer outcomes using input data, and show that the collection of input data reduces uncertainty of outcomes over a baseline approach. In addition, we quantify the performance of a likelihood scoring technique that can be used to self-evaluate the extrapolation risk of prognosis on unseen patients.

97 MATHEMATICS AND COMPUTING↗

malbacR: A Package for Standardized Implementation of Batch Correction Methods for Omics Data

Mass spectrometry is a powerful tool for identifying and analyzing small molecules, such as metabolites and lipids, in com-plex biological samples. Liquid chromatography and gas chromatography mass spectrometry studies quite commonly in-volve large numbers of samples, which can require significant time for sample preparation and analyses. To accommodate such studies, the samples are commonly split into batches. Inevitably, variations in sample handling, temperature fluctua-tion, imprecise timing, column degradation and other factors result in systematic errors or biases of the measured abundances between the batches. Numerous methods are available via R packages to assist with batch correction for small molecule om-ics data; however, since these methods were developed by different research teams, the algorithms are available in separate R packages, each with different data input and output formats. We introduce the malbacR package which consolidates eleven common batch effect correction methods for small molecule omics data into one place so users can easily implement and compare: pareto scaling, power scaling, range scaling, ComBat, EigenMS, NOMIS, RUV-random, QC-RLSC, WaveI-CA2.0, TIGER, and SERRF. The malbacR package standardizes data input and output formats across these batch correction methods. The package works in conjunction with the pmartR package, allowing users to seamlessly include batch effect cor-rection in a pmartR workflow without needing any additional data manipulation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Site B - NREL ASSIST (SN11) Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 11) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH), which is a combined data product that uses data from ceilometers at sites A1 and H and scanning lidars from ARM sites C1 and E37. The CBH is weighted inversely proportionally to the distance to the respective site to take into account the spatial variability of clouds (see https://github.com/StefanoWind/ASSIST_analysis/blob/main/awaken_processing/combine_cbh.py). The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at ARM SGP, OK.

17 WIND ENERGY↗

Site G - NREL ASSIST (SN10) Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 10) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH), which is a combined data product that uses data from ceilometers at sites A1 and H and scanning lidars from ARM sites C1 and E37. The CBH is weighted inversely proportionally to the distance to the respective site to take into account the spatial variability of clouds (see https://github.com/StefanoWind/ASSIST_analysis/blob/main/awaken_processing/combine_cbh.py). The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at ARM SGP, OK.

17 WIND ENERGY↗

Site C1a - NREL ASSIST (SN12) Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 12) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH), which is a combined data product that uses data from ceilometers at sites A1 and H and scanning lidars from ARM sites C1 and E37. The CBH is weighted inversely proportionally to the distance to the respective site to take into account the spatial variability of clouds (see https://github.com/StefanoWind/ASSIST_analysis/blob/main/awaken_processing/combine_cbh.py). The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at ARM SGP, OK.

17 WIND ENERGY↗

Data and Code for Understanding Generative AI Content with Embedding Models

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Vargas, Max [Pacific Northwest National Laboratory↗

BEPAM-E Model Code and CABBI Simulation Results for "Repeal of the Clean Power Plan: Social Cost and Distributional Implications"

The dataset consists of results and various input data that are used in the GAMS model for the publication "Repeal of the Clean Power Plan: Social Cost and Distributional Implications". All the data are either excel files or in the .inc format which can be read within GAMS or Notepad. Main data sources include: agriculture, transportation and electricity data. Model details can be found in the paper and the GAMS model package.

carbon abatement↗

Community Solar Program Design Considerations & Modeling Inputs

Designing and modeling a community solar (CS) program is a complex process with numerous variable inputs that are interconnected. Modeling a CS program can be useful to inform the programs design itself while also providing stakeholders of all types with information. Accurate data inputs and assumptions are key to ensuring that a model is informative and as representative of real market conditions as possible. This report is an exploration of community solar program modeling considerations, especially as it relates to data inputs such as capital costs, administrative fees, and subscription size, using a CS program in North Carolina as a case study.

14 SOLAR ENERGY↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

User Manual for MC 2 -3 Gamma Library Generation

This report provides the documentation for the revised procedure to generate the MC 2 -3 gamma library with the latest versions of PreGAMMA and GenGAMMA which have recently been updated to fix the program errors and inconsistent data processing for some isotopes identified in a recent verification and validation study of the MC 2 -3 gamma library. A few associated errors associated with gamma processing were also identified and fixed in the MC 2 -3 code. Addressing the changes in MC 2 -3, the procedure to generate the PMATRX and GAMISO datasets is discussed as well. The input data of PreGAMMA and GenGAMMA are discussed in detail using sample input data. In addition, the structure and formats of the MC 2 -3 gamma library are provided in detail.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Identifying the nature of the QCD transition in heavy-ion collisions with deep learning

In this proceeding, we review our recent work using deep convolutional neural network (CNN) to identify the nature of the QCD transition in a hybrid modeling of heavy-ion collisions. Within this hybrid model, a viscous hydrodynamic model is coupled with a hadronic cascade “after-burner”. As a binary classification setup, we employ two different types of equations of state (EoS) of the hot medium in the hydrodynamic evolution. The resulting final-state pion spectra in the transverse momentum and azimuthal angle plane are fed to the neural network as the input data in order to distinguish different EoS. To probe the effects of the fluctuations in the event-by-event spectra, we explore different scenarios for the input data and make a comparison in a systematic way. We observe a clear hierarchy in the predictive power when the network is fed with the event-by-event, cascade-coarse-grained and event-fine-averaged spectra. The carefully-trained neural network can extract high-level features from pion spectra to identify the nature of the QCD transition in a realistic simulation scenario.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗