Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “feature engineering”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

First indirectly driven liquid-DT filled double shell implosions at the National Ignition Facility

Double shell implosions aim to explore material mixing under fusion conditions in a volume burn geometry using high-Z metal pushers. High-Z pushers are more compressible than low-Z pushers enabling high stagnation pressure which reduces the required implosion speed while maintaining a low pusher adiabat despite strong shock heating. Additionally, the use of a small fuel mass reduces the fuel internal energy required for ignition, thus achieving stable platforms with reasonable fusion output to conduct controlled experiments. These factors make volume burn in a double shell implosion highly promising. Recently, a series of liquid-DT filled, indirectly driven, double shell implosions were conducted at the National Ignition Facility with laser drives reaching up to 1.5 MJ. These experiments achieved a maximum DT neutron yield of 1.67 × 10 14 (yield—479 J), DT ion temperature of 2.6 keV, fuel areal density (ρR) of 0.15 g/cm 2 , and stagnation pressure of 79 Gbar. Over the course of these shots, the DT neutron yield has increased by an order of magnitude largely from improved mitigation of outer shell assembly joint driven instability growth using thicker gold plating at the assembly joint. Further performance improvements are expected by enhancing outer-to-inner shell kinetic energy transfer and refined mitigation of degradations from engineering features such as the outer shell joint, fill tube, and surface roughness.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Unsupervised Learning for Improved Gamma-Ray Spectrometry in Pixelated Cadmium Zinc Telluride (CZT) Detectors

Machine learning has been found to be ubiquitously useful across many industries, presenting an opportunity to improve radiation detection performance using data-driven algorithms. Improved detector resolution can aid in the detection, identification, and quantification of radionuclides. Here, in this work, a novel, data-driven, unsupervised learning approach is developed to improve detector spectral characteristics by learning, and subsequently rejecting, poorly performing regions of the pixelated detector. Feature engineering is used to fit individual characteristic photo peaks to a Doniach lineshape with a linear background model. Then, principal component analysis is used to learn a lower-dimension latent space representation of each photo peak where the pixels are clustered, and subsequently ranked, based on the cluster mean distance to an optimal point. Pixels within the worst cluster(s) are rejected to improve the full-width at half-maximum (FWHM) by 10% to 15% (relative to the bulk detector) at 50% net efficiency when applied to training data obtained from measurements of a 100 μCi 154 Eu source using a H3D M400i pixelated cadmium zinc telluride detector. These results compare well with, but do not outperform, a greedy algorithm that accumulates pixels in order of FWHM from lowest to highest used as a benchmark. In the future, this approach can be extended to include the detector energy and angular response. Finally, the model is applied to newly seen natural and enriched uranium spectra relevant for nuclear safeguards applications.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Photometric identification of compact galaxies, stars, and quasars using multiple neural networks

We present MargNet, a deep learning-based classifier for identifying stars, quasars, and compact galaxies using photometric parameters and images from the Sloan Digital Sky Survey Data Release 16 catalogue. MargNet consists of a combination of convolutional neural network and artificial neural network architectures. Using a carefully curated data set consisting of 240 000 compact objects and an additional 150 000 faint objects, the machine learns classification directly from the data, minimizing the need for human intervention. MargNet is the first classifier focusing exclusively on compact galaxies and performs better than other methods to classify compact galaxies from stars and quasars, even at fainter magnitudes. This model and feature engineering in such deep learning architectures will provide greater success in identifying objects in the ongoing and upcoming surveys, such as Dark Energy Survey and images from the Vera C. Rubin Observatory.

79 ASTRONOMY AND ASTROPHYSICS↗

Extensive Attention Mechanisms in Graph Neural Networks for Materials Discovery

We present our research where attention mechanism is extensively applied to various aspects of graph neural net- works for predicting materials properties. As a result, surrogate models can not only replace costly simulations for materials screening but also formulate hypotheses and insights to guide further design exploration. We predict formation energy of the Materials Project and gas adsorption of crystalline adsorbents, and demonstrate the superior performance of our graph neural networks. Moreover, attention reveals important substructures that the machine learning models deem important for a material to achieve desired target properties. Our model is based solely on standard structural input files containing atomistic descriptions of the adsorbent material candidates. We construct novel methodological extensions to match the prediction accuracy of state-of-the-art models some of which were built with hundreds of features at much higher computational cost. We show that sophisticated neural networks can obviate the need for elaborate feature engineering. Our approach can be more broadly applied to optimize gas capture processes at industrial scale.

Cong, Guojing↗

Tandem Predictions for HPC Jobs

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

HPC↗

Avian Activity Classification Using Recurrent Networks to Fuse Videos with Metadata on Imbalanced Datasets

Activity classification plays a crucial role in various real-life scenarios involving both humans and animals. There is an increasing need for precise activity classification focused on avian-solar interactions, as the usage of solar energy facilities, such as photovoltaic array power stations, has been observed to impact bird species richness, behavior, and activity. However, there has been no work to develop an automated system to monitor and classify these avian-solar interactions. All current methods rely on human observers, which is time and human resources costly and subject to errors related to searcher efficiency. With the recent success of Deep Learning models in activity classification problems, this paper develops a recurrent neural network-based model to automatically classify six avian activities around solar energy facilities. Our proposed model integrates critical feature engineering metadata with video frame data, enabling improved learning and more accurate activity classification. Furthermore, we address the challenge of data imbalance during training and demonstrate the efficacy of our model in detecting and classifying different activities within video tracks. Additionally, we analyze the saliency/backpropagation map of the trained proposed model and validate its decision-making rationale.

Avian activity classification; bidirectional LSTM;↗

ASP (Adaptive Splines for Prediction) [SWR-22-66]

This software package provides functionality to fit piecewise cubic splines to model the relationship between a univariate X (independent variable) and a univariate Y (dependent variable). The output is akin to linear regression but has higher "capacity" in that it can learn more complicated relationships despite leveraging only a single predictor. It thus can learn very reasonable relationships between X and Y with minimal effort. The canonical intended use case is for predicting electric demand (Y) from temperature (X). There can be complicated, non-linear relationships between these two variables, but these can be reasonably approximated by the sum of: 1) an "aggregate" relationship between daily average temperature and daily average demand, and 2) intra-day patterns represented as hourly offsets from the daily average load. The spline functionality is used for learning a robust "aggregate" relationship without the inconvenience of manual feature engineering while still ensuring robust results. The model fitting workhorse is R's built-in smooth.spline. This package provides a suite of options for controlling how smooth.spline fits a spline to the data. In particular, it is designed to return a model with a single critical point (i.e., a single location along the domain/support where the first derivative is zero). This is to ensure that predicted changes in electric demand are always positive as temperatures become more extreme. This helps to prevent overfitting, particularly when the input data has few observations or has significant uncertainty from other sources.

Murphy, Sinnott↗

Risk-informed Predictive Analytics To Achieve Cost-effective Condition-based Monitoring And Maintenance Strategy

The research involves developing risk-informed predictive analytic capabilities to achieve condition-based monitoring and maintenance strategies to reduce overall maintenance costs. The research utilizes data (real-time data, periodic data, and institutional knowledge) related to a particular plant asset from a specific nuclear plant site to develop risk-informed predictive analytic algorithms. The developed algorithms and codes are used to optimize the maintenance strategy and estimate/forecast generation costs based on the state of health of the plant asset. Developed codes specifically include 1. Parameter estimation code based on Bayesian inference 2. Statistical data analysis code 3. Feature engineering code 4. Health classifier code 5. Diagnosis code 6. Prognosis code 7. Hazard code 8. Generation risk code 9. Economic code

Agarwal, Vivek↗

Moltensaltpropnet

MoltenSaltPropnet is a physics-informed machine learning framework that aims to predict the thermophysical properties of molten fluoride and chloride salt mixtures, which are crucial for the design and safety of Generation IV molten salt reactors. The code processes data from the Molten-Salt Thermal Properties Database (MSTDB-TP) and the Janz compendium, converting critically evaluated correlations into fast, differentiable surrogate models for density, viscosity, thermal conductivity, and heat capacity across 448 distinct salt systems. The implementation consists of several key components: 1. Data Curation: The code parses and cleans the raw data, normalizing elemental mole fractions and extracting relevant regression coefficients for various thermophysical properties. 2. Feature Engineering: It generates fixed-length numerical descriptors that encapsulate the composition and temperature, incorporating polynomial interaction terms and dimensionality-reduction techniques to optimize model performance. 3. Coefficient Learning: Four different machine learning architectures are employed: a deep residual network (ResNet), a Kolmogorov–Arnold network (KAN), a sparsity-inducing neural network (SNN), and classical regression models. Each model learns to predict coefficients that define the temperature-dependent correlations for the thermophysical properties. 4. Property Reconstruction: The predicted coefficients are used to compute temperature-dependent property values, ensuring positivity and monotonic trends through a composite loss function that enforces physical constraints. 5. User Interface: An open-source web application enables users to filter the database, train task-specific models, and visualize the results, allowing for rapid exploration of candidate salt mixtures. MoltenSaltPropnet bridges the gap between limited experimental data and high-fidelity reactor simulations, providing a powerful tool for researchers in the field of molten salt reactors and advanced nuclear energy systems.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"

This data package contains the associated data and scripts for Nagamoto, E., Ombadi, M., Ciulla, F. et al. Widespread drought-driven declines in streamflows and water quality in the Upper Colorado River Basin during 1998-2022. Commun Earth Environ 7, 734 (2026). https://doi.org/10.1038/s43247-026-03890-5. This purpose of this study was to investigate the impact of the 21st century drought on water quantity and quality at catchments throughout the Upper Colorado River Basin (UCRB). We used stream flow, water temperature, specific conductance, air temperature, precipitation, and catchment attribute data for over 200 sites in the UCRB, collected from the National Water Information System using Basin3D (Varadharajan, 2023), GAGESII (Falcone, 2010), and the Google Earth Engine. We identified years of severe drought between 1998 and 2022 using the Standardized Precipitation Evaporation Index (SPEI), then calculated the relative change percentage of the stream flow, water temperature, and specific conductance from drought versus non-drought years. We used the attribute information from GAGESII to investigate what physical traits of catchments are associated streamflow vulnerability (greater relative change) or resilience to drought. We used land cover data from the National Land Cover Database (USGS, 2024) to assess any changes to physical attributes that may not be represented in the static attributes information in GAGESII. To increase data availability, we modeled stream temperature using methods from Willard, 2023. While the study period is water years 1998 to 2022, the raw water quantity and quality data extends to 1950 and the meteorological data extends to 1980. The data and code can be downloaded via the UCRB_drought.zip. Within the zip, the files are organized as follows: - INPUTS: Contains all input data used in UCRB_Drought_Workflow.ipynb - OUTPUTS: Contains all intermediate data created from UCRB_Drought_Workflow.ipynb as well as final products including the calculated Standardized Evapotranspiration Index (SPEI) - climatic_variables: The code used to collect meteorologic data from Google Earth Engine - feature_importance: The code used for the catchment attributes analysis - preprocessing: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - pyeto: Code used in UCRB_Drought_Workflow_Preprocessing.ipynb - calculations: Code used in UCRB_Drought_Workflow_Impacts.ipynb - plotting: Code used in UCRB_Drought_Workflow_Impacts.ipynb - README.md - UCRB_Drought_Workflow_Preprocessing.ipynb: The code used to prep raw data for the analysis - UCRB_Drought_Workflow_Impact.ipynb: The code which uses the prepped raw data for analysis, and plots all figures - requirements_ucrb-drought_v2.yml: The requirements file to create a virtual environment and Jupyter Lab kernel to run the code The INPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_RAW" folder contains raw data for streamflow, water temperature, and specific conductance in a ".h5" file. The "NLCD_RAW" folder contains ".csv" files with annual land cover percentages for counties within the UCRB. The "MET_RAW" folder contains a ".csv" file with monthly meteorological data (air temperature and precipitation) for the sites in the UCRB which was obtained from code in the climatic_variables folder. The "GAGESII" folder contains ".csv" files with physical catchment attribute variables for catchments across the country. The "WT_LSTM_data" folder contains ".csv" files with calculated WT (Willard, 2023) and the associated RMSEs. The "Upper_Colorado_River_Basin_Boundary" folder contains geographic data including a shapefile for plotting in the UCRB_Drought_Workflow.ipynb. The "RESERVOIRS_RAW" folder contains ".csv" files for each reservoir in the UCRB with daily reservoir storage. There are also two files in the INPUTS folder that have combined reservoir storage data and reservoir metadata. The OUTPUTS folder is organized into the following major directories and sub-directories. The "RDC_WT_SC_data" folder contains a folder "Water_year" with the associated cleaned data, metadata, and data availability information in ".csv" files, a folder "Median_Relchange" with the relative change comparing drought to non-drought years in ".csv" files, and a folder "Peak95_Min5_Relchange" that has ".csv" files for the relative change in peak (95th %) and minimum (5th %) variables. The "NLCD_data" folder contains the difference in land cover from the beginning to end of the study period and the percentage of the county that is within UCRB bounds can be found in Nagamoto et al (2025)). The "MET_data" folder contains separated monthly air temperature and precipitation data and the calculated PET in ".csv" files. The "SPEI_data" folder contains ".csv" files with calculated SPEI values (one restricted to the study period and the other with information from the entire MET data period). The "Paper_Tables" folder contains two ".csv" files containing site information and data availability and information about the GAGESII trait aggregated categories. The base directory includes the file “flmd.csv” for a list and description of all files and the file “dd.csv” for data dictionaries. Scripts for preprocessing, analysis, and figure generation are located in the associated GitHub repository found at [https://github.com/iNAIADS/drought-impacts/tree/develop/UCRB-drought]. UPDATE 1: Title and code file updated to match submitted manuscript 10-15-2025. UPDATE 2: Code and data files updated to match revised manuscript 3-4-2026. UPDATE 3: Code and data files updated to match revised manuscript 6-7-2026. ** NOTE: DD and FLMD have not been updated yet. UPDATE 4: Added associated Manuscript information and DD and FLMD have been updated. To cite this code, please use the following BibTeX: @misc{nagamoto2025drought, author = {Emily Nagamoto and Fabio Ciulla and Mohammad Ombadi and Jared Willard and Rosemary Carroll and Charuleka Varadharajan}, title = {Dataset: "Widespread Drought-driven Declines in Streamflows and Water quality in the Upper Colorado River Basin (1998-2022)"}, year = {2025}, doi = {10.15485/2551894}, publisher = {ESS-DIVE Repository}, url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2551894} }

54 ENVIRONMENTAL SCIENCES↗

Phase Picking Beyond Local Distances: Where Waveform Filtering Still Matters for Deep Learning Models

Waveform filtering is a standard step in traditional seismic phase picking but often receives little attention in deep learning workflows, where models are typically trained on raw or minimally processed waveforms. Although this strategy performs well for local events, we show that performance can degrade substantially at regional distances. To address this limitation, we introduce two ways to incorporate multiband-filtered waveforms into deep learning phase pickers. The stacking approach concatenates filtered inputs along the channel dimension, while the branching approach processes each frequency band through a dedicated network branch before feature fusion. Both approaches can substantially improve performance across epicentral distances of 0° to 20°, but their effectiveness depends strongly on the selected frequency bands. Tests with multiple filter banks show that filter-bank design should be treated as part of model optimization rather than as a fixed preprocessing choice. Grad-CAM analysis of the branching model indicates that band importance varies among waveform samples and across training realizations, with only a weak overall preference for the 0.25 to 0.5 Hz band. These results show that no single filter band is consistently optimal and demonstrate that explicit feature engineering remains valuable for robust deep learning-based seismic phase picking.

58 GEOSCIENCES↗

Human-Constrained Indicators of Gatekeeping Behavior as a Role in Information Suppression: Finding Invisible Information and the Significant Unsaid

To date, disinformation research has focused largely on the production of false information ignoring the suppression of select information. We term this alternative form of disinformation information suppression. Information suppression occurs when facts are withheld with the intent to mislead. In order to detect information suppression, we focus on understanding the actors who withhold information. In this research, we use knowledge of human behavior to find signatures of different gatekeeping behaviors found in text. Specifically, we build a model to classify the different types of edits on Wikipedia using the added text alone and compare a human-informed feature engineering approach to a featureless algorithm. Being able to computationally distinguish gatekeeping behaviors is a first step towards identifying when information suppression is occurring.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Plan for Developing TRISO Fuel Processing Technologies

This Plan demonstrates the availability of technologies for processing TRISO used nuclear fuel for waste management and actinide recovery purposes. These technologies are judged to be at a very low level of technology readiness and as such they constitute a fertile research area for the DOE-NE’s Office of Materials and Chemical Technologies. Strategies to mature the technologies to a point where they can reasonably be considered in engineering alternatives analyses typically involve laboratory-scale tests using fuel simulant to characterize process streams and demonstrate key engineering features. Several criteria are available to help selecting candidate technologies for further maturation

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Flexible Siting Criteria and Staff Minimization for Micro-Reactors

The economic potential of micro-reactors is vast and underestimated. Commonly-emphasized applications include niche markets such as remote communities, mines and military bases. However, micro-reactors could be used as flexible energy generators also for larger markets, such as mobile and containerized agriculture and manufacturing facilities, district heating, micro-grids for data centers, sea ports, airports and hospitals. The implication is that micro-reactors may have to be deployed also in non-remote locations. Successful implementation of micro-reactors needs a navigable and predictable licensing process, technology-appropriate siting restrictions, risk-informed emergency and safety requirements, and practical operating and maintenance requirements. The primary goal of this project was to develop siting criteria that are tailored to micro-reactors deployable in densely-populated areas, e.g., urban environments. To achieve that goal, we compared the characteristics of the MIT research reactor (MITR) with those of leading micro-reactor concepts (e.g., eVinci, USNC, Aurora), and evaluated whether and how the MITR design basis (e.g., inherent safety features, engineered safety systems, source term, emergency planning and emergency operating procedures) and associated regulations may be applicable to these new micro-reactors as well. What makes MITR a unique analogue in this context is its small power rating (6 MWt) and physical size, mode of operations (24/7 with a somewhat more commercial flavor than typical university reactors), and especially its urban location. Of course significant differences exist, such as mission (power production vs. research) and the reactor design itself. Leveraging the MITR experience, this project was able to generate criteria that will allow micro-reactors to realize their full economic potential as flexible heat and electricity generators for a diverse portfolio of applications in non-remote locations. As such, the outcome of this project might encourage investment in and use of micro-reactors. A second goal of the project was to conceptualize a model of operations for micro-reactors that would minimize the staffing requirements, and thus reduce the cost of electricity and heat generated by these systems. Here too our approach was to systematically review the MITR experience and requirements, as well as survey the innovations in autonomous control technologies and monitoring (e.g., advanced sensors, drones, robotics, AI) that would permit a dramatic reduction in staffing at future micro-reactor installations. The scope of work was expanded after the start date to include also an evaluation of micro-reactor security, using the so-called consequence-based analysis, and the development of a methodology to perform dynamic risk assessment for micro-reactors, using system theory and modeling and simulation.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A Semi-supervised Hybrid Machine Learning Framework for the Qualification of Resistance Spot Welds

• Industries requiring high structural integrity, including automotive, aerospace, and construction, place considerable significance on weld quality classification. • The inspection normally involves human expertise through predefined quality metrics that are subjective, error-prone, and time-intensive • The challenge to classification model development is the scarcity of labeled data and imbalanced distributions in the data that are labeled. • This work develops a new hybrid methodology that achieves clustering using KMeans++ together with supervised classification to overcome these challenges. • The ensemble-based classifiers were identified as optimal, with accuracy enhancements of up to 8% using the pseudo-labeled dataset. • The work provides practical insight into feature engineering and machine learning integration in industrial quality assurance applications.

Rogers, Jeremy K. [Savannah River National Laborat↗

Fingerprinting Interactions between Proteins and Ligands for Facilitating Machine Learning in Drug Discovery

Molecular recognition is fundamental in biology, underpinning intricate processes through specific protein–ligand interactions. This understanding is pivotal in drug discovery, yet traditional experimental methods face limitations in exploring the vast chemical space. Computational approaches, notably quantitative structure–activity/property relationship analysis, have gained prominence. Molecular fingerprints encode molecular structures and serve as property profiles, which are essential in drug discovery. While two-dimensional (2D) fingerprints are commonly used, three-dimensional (3D) structural interaction fingerprints offer enhanced structural features specific to target proteins. Machine learning models trained on interaction fingerprints enable precise binding prediction. Recent focus has shifted to structure-based predictive modeling, with machine-learning scoring functions excelling due to feature engineering guided by key interactions. Notably, 3D interaction fingerprints are gaining ground due to their robustness. Various structural interaction fingerprints have been developed and used in drug discovery, each with unique capabilities. This review recapitulates the developed structural interaction fingerprints and provides two case studies to illustrate the power of interaction fingerprint-driven machine learning. The first elucidates structure–activity relationships in β2 adrenoceptor ligands, demonstrating the ability to differentiate agonists and antagonists. The second employs a retrosynthesis-based pre-trained molecular representation to predict protein–ligand dissociation rates, offering insights into binding kinetics. Despite remarkable progress, challenges persist in interpreting complex machine learning models built on 3D fingerprints, emphasizing the need for strategies to make predictions interpretable. Binding site plasticity and induced fit effects pose additional complexities. Interaction fingerprints are promising but require continued research to harness their full potential.

3D structural interaction fingerprints↗

Unsupervised Detection of SOC Spoofing in OCPP 2.0.1 EV Charging Communication Protocol Using One-Class SVM

The electric vehicles (EVs) market keeps growing globally; thus, it is critical to secure the EV charging communication protocols in order to guarantee reliable and fair charging operations among the customers. The Open Charge Point Protocol (OCPP) 2.0.1 supports the communication between the Electric Vehicle Supply Equipment (EVSE) and Charging Station Management Systems (CSMSs); therefore, it becomes vulnerable to several types of attacks, which aim to jeopardize smart charging, billing, and energy management. Specifically, OCPP 2.0.1 allows the self-reporting of the State of Charge (SOC) values, which makes it vulnerable to spoofing-based cyberattacks, which target manipulating the scheduling priorities, distorting the load forecasts, and extending the charging sessions in an unfair manner. In this paper, we try to address this type of attack by providing a comprehensive analysis of the SOC spoofing attacks and introducing a novel unsupervised detection framework based on the One-Class Support Vector Machine (OCSVM) algorithm. Specifically, two types of attack scenarios are analyzed (i.e., priority manipulation and session extension) by deriving engineered features that capture the nonlinear relationships under normal charging behavior. Detailed simulation-based results are derived by utilizing the DESL-EPFL Level 3 EV charging dataset. Our results demonstrate high F1-score and recall in identifying spoofed SOC values and that the proposed OCSVM model demonstrates superior performance compared to alternative clustering and deep-learning based detectors.

EV charging↗

Data Augmentation for Neutron Spectrum Unfolding with Neural Networks

Neural networks require a large quantity of training spectra and detector responses in order to learn to solve the inverse problem of neutron spectrum unfolding. In addition, due to the under-determined nature of unfolding, non-physical spectra which would not be encountered in usage should not be included in the training set. While physically realistic training spectra are commonly determined experimentally or generated through Monte Carlo simulation, this can become prohibitively expensive when considering the quantity of spectra needed to effectively train an unfolding network. In this paper, we present three algorithms for the generation of large quantities of realistic and physically motivated neutron energy spectra. Using an IAEA compendium of 251 spectra, we compare the unfolding performance of neural networks trained on spectra from these algorithms, when unfolding real-world spectra, to two baselines. We also investigate general methods for evaluating the performance of and optimizing feature engineering algorithms.

McGreivy, James (ORCID:0000000321723411)↗