Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Algorithm testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83

MArVD2: a machine learning enhanced tool to discriminate between archaeal and bacterial viruses in viral datasets

Abstract Our knowledge of viral sequence space has exploded with advancing sequencing technologies and large-scale sampling and analytical efforts. Though archaea are important and abundant prokaryotes in many systems, our knowledge of archaeal viruses outside of extreme environments is limited. This largely stems from the lack of a robust, high-throughput, and systematic way to distinguish between bacterial and archaeal viruses in datasets of curated viruses. Here we upgrade our prior text-based tool (MArVD) via training and testing a random forest machine learning algorithm against a newly curated dataset of archaeal viruses. After optimization, MArVD2 presented a significant improvement over its predecessor in terms of scalability, usability, and flexibility, and will allow user-defined custom training datasets as archaeal virus discovery progresses. Benchmarking showed that a model trained with viral sequences from the hypersaline, marine, and hot spring environments correctly classified 85% of the archaeal viruses with a false detection rate below 2% using a random forest prediction threshold of 80% in a separate benchmarking dataset from the same habitats.

Vik, Dean (ORCID:000000027546899X)↗

Piecewise interaction picture density matrix quantum Monte Carlo

The density matrix quantum Monte Carlo (DMQMC) set of methods stochastically samples the exact N-body density matrix for interacting electrons at finite temperature. We introduce a simple modification to the interaction picture DMQMC (IP-DMQMC) method that overcomes the limitation of only sampling one inverse temperature point at a time, instead allowing for the sampling of a temperature range within a single calculation, thereby reducing the computational cost. At the target inverse temperature, instead of ending the simulation, we incorporate a change of picture away from the interaction picture. The resulting equations of motion have piecewise functions and use the interaction picture in the first phase of a simulation, followed by the application of the Bloch equation once the target inverse temperature is reached. We find that the performance of this method is similar to or better than the DMQMC and IP-DMQMC algorithms in a variety of molecular test systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enhancing Thomson scattering polychromator performance with multi-pass spectral filters

In photon-deficient, noncollective Thomson scattering diagnostics, filter polychromators are typically employed in the spectral analysis of Thomson-scattered signals to achieve acceptable signal-to-noise performance. Currently, the most common polychromator filter configuration employs a set of single-passband optical filters that define individual spectral channels. Here, we introduce a new spectral analysis method for Thomson scattering based on spectral filters with multiple passbands, referred to as Thomson scattering spectral multiplexing. Implementing multi-bandpass spectral filters on polychromators increases the achievable range of electron temperature measurement for a given number of filters employed. In addition, Thomson scattering spectral multiplexing reduces systematic measurement uncertainty, with fewer required spectral channels, thereby decreasing light loss from reduced optical element interactions. A multi-bandpass filter set, optimized by a genetic algorithm, has been successfully installed and tested on the Helically Symmetric eXperiment (HSX), demonstrating the benefits of the Thomson scattering spectral multiplexing method.

Instruments & Instrumentation↗

Achieving Higher Order Accuracy in Space in Hydrodynamic Simulations of Self-Gravitating Gas

Modern astrophysical simulation codes employ a variety of numerical algorithms capable of achieving higher-order accuracy in both space and time. Albeit they succeed in achieving an effective higher spatial resolution and in suppressing the numerical damping of waves, to our knowledge, all current astrophysical simulations invoking self-gravity are limited to second-order accuracy in space. If we can devise an algorithm to evaluate self-gravity with a higher-order spatial accuracy, we can better the evaluation of the gravitational acceleration and gravitational energy release which dictate the evolution of many astrophysical systems. Herein, we present a numerical algorithm for self-gravitating hydrodynamics capable of achieving fourth-order accuracy for a given density distribution on a Cartesian uniform grid. First, we derive the cell-averaged gravitational potential at fourth-order accuracy from the cell-averaged density by solving the Poisson equation. Next, we obtain the cell average of the product of the density and gravitational acceleration, which differs from the cell-averaged density multiplied by the cell-averaged gravitational acceleration. We then show the verification of the algorithm by applying it to critical test problems: (1) maintaining equilibria of self-gravitating slabs, even upon advection, (2) evolving a polytropic sphere with a massive power-law envelope, and (3) conservation of specific entropy during the propagation of a sound wave.

79 ASTRONOMY AND ASTROPHYSICS↗

Coincidence anomaly detection for unsupervised locating of edge localized modes in the DIII-D tokamak dataset

Using supervised learning to train a machine learning model to predict an on-coming edge localized mode (ELM) requires a large number of labeled samples. Creating an appropriate data set from the very large database of discharges at a long-running tokamak, such as DIII-D, would be a very time-consuming process for a human. Considering this need and difficulty, we use coincidence anomaly detection, an unsupervised learning technique, to train an ELM-identifier to identify and label ELMs in the DIII-D discharge database. This ELM-identifier shows, simultaneously, a precision of 0.68 and a recall of 0.63 (AUC is 0.73) on identifying ELMs in example time series pulled from thousands of discharges spanning five years. In a test set of 50 discharges, the algorithm finds over 26 thousand ELM candidates, more than 5 times the existing catalog of ELMs labeled by humans.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

FluxRETAP: a REaction TArget Prioritization genome-scale modeling technique for selecting genetic targets

MOTIVATION: Metabolic engineering is rapidly evolving as a result of new advances in synthetic biology tools and automation platforms that enable high throughput strain construction, as well as the development of machine learning tools (ML) for biology. However, selecting genetic engineering targets that effectively guide the metabolic engineering process is still challenging. ML can provide predictive power for synthetic biology, but current technical limitations prevent the independent use of ML approaches without previous biological knowledge. RESULTS: Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale models for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing the production of a desired metabolite. This method can provide a list of desirable engineering targets that can be combined with current ML pipelines. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production, 50% of targets that experimentally improved taxadiene production in E. coli and ∼60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida, while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets. AVAILABILITY AND IMPLEMENTATION: FluxRETAP is implemented in python and released under the creative commons license. The implementation and code are freely available at: https://github.com/JBEI/FluxRETAP.

Czajka, Jeffrey J↗

QSO photometric redshifts using machine learning and neural networks

ABSTRACT The scientific value of the next generation of large continuum surveys would be greatly increased if the redshifts of the newly detected sources could be rapidly and reliably estimated. Given the observational expense of obtaining spectroscopic redshifts for the large number of new detections expected, there has been substantial recent work on using machine learning techniques to obtain photometric redshifts. Here, we compare the accuracy of the predicted photometric redshifts obtained from deep learning (DL) with the k-nearest neighbour (kNN) and the decision tree regression (DTR) algorithms. We find using a combination of near-infrared, visible, and ultraviolet magnitudes, trained upon a sample of Sloan Digital Sky Survey quasi-stellar objects, that the kNN and DL algorithms produce the best self-validation result with a standard deviation of σΔz = 0.24 (σΔz(norm) = 0.11). Testing on various subsamples, we find that the DL algorithm generally has lower values of σΔz, in addition to exhibiting a better performance in other measures. Our DL method, which uses an easy to implement off-the-shelf algorithm with neither filtering nor removal of outliers, performs similarly to other, more complex, algorithms, resulting in an accuracy of Δz < 0.1 up to z ∼ 2.5. Applying the DL algorithm trained on our 70 000 strong sample to other independent (radio-selected) data sets, we find σΔz ≤ 0.36 (σΔz(norm) ≤ 0.17) over a wide range of radio flux densities. This indicates much potential in using this method to determine photometric redshifts of quasars detected with the Square Kilometre Array.

Curran, S. J.↗

Preliminary Thermoluminescent Dosimeter Glow Curve Analysis with Automated Glow Peak Identification for LiF Mg,Ti

When appropriately analyzed, thermoluminescent dosimeter glow curve analysis allows for improved quantification of thermoluminescent material behavior while flagging abnormalities. The mathematical separation of a glow curve into contributions from energetically unique trap states, or glow curve analysis, may be used to remove undesired effects of signal fading for complex materials. A generalized glow curve analysis software for the separation of glow curves is presented in this paper. Written in C ++ , the software uses the first-order kinetics model with automatic peak identification. The automatic identification of peaks is achieved through a unique peak-finding algorithm. Here, the program was performance tested using experimental glow curve data from LiF:Mg,Ti, and comparative results are presented.

47 OTHER INSTRUMENTATION↗

Attention-Augmented Parametric Kernel Graph Neural Network (APKGNN) for Node Classification

We present a new graph neural network, the Attention-based Parametric-Kernel augmented Graph Neural Network (APKGNN), developed for node classification tasks. Despite extensive work on modeling multi-faceted relationships between connected nodes of a graph, the effect of attention on edge features mapped to relationships has not yet been analyzed through learning representation. This study derives such an attention vector by first calculating node features corresponding to endpoints of an edge and then aggregating these with extracted local intrinsic patches of a given graph to generate augmented local patch vectors. This process uses a parametric kernel based on Gaussian mixture models (GMMs) to embed local neighborhoods of the graph in local patches. The patch vectors then convolve with the above node features to produce an updated node representation. We show that this new learning representation (APKGNN) achieves higher node classification accuracy on tasks - both standard benchmarks (Cora, PubMed, Citeseer) and new experimental short text corpora where nodes correspond to text documents and words. This implementation of the GNN convolution layer outperforms state-of-the-art (SOTA) algorithms, achieving higher training, validation, and test accuracy by a significant margin on three standard benchmark data sets under both SOTA experimental settings and those for new testbeds.

Bose, Avishek↗

Impact of Open Communication Networks on Load Frequency Control with Plug-In Electric Vehicles by Cyber-Physical Dynamic Co-Simulation

With the increasing electrification of the transportation sector to achieve the carbon neutrality objective, despite the challenges of charging electric vehicles (EV), there are also opportunities through smart charging EVs to improve system frequency stability; however, EV control technologies might require nontraditional communication support. This paper investigates the impacts of communication variations of EV on power system load frequency control through a cyber-physical dynamic system (CPDS) co-simulation. Here, the CPDS is built upon our previously developed transmission-and-distribution dynamic co-simulation model with the added communication variation functions (i.e., delay and packet loss). The case studies consider multiple communication variation scenarios when the system experiences an N-1 generation trip contingency. The scenarios include communication delays and packet loss using both homogeneous and heterogeneous assumptions. The outcomes of this work can help improve EV frequency regulation services and provide robust and effective tests for different load frequency control algorithms of the future power systems.

ADVANCED PROPULSION SYSTEMS,ENERGY PLANNING, POLIC↗

Mass Detection for Heavy-Duty Vehicles using Gaussian Belief Propagation

Predicting vehicle mass is critical to accurately estimate energy use and emissions of commercial trucks. However, data from vehicle telematics is often not at sufficient temporal resolution or accuracy for use in model-based detection methods. In this work, a new statistical mass prediction technique is described for heavy-duty vehicles that incorporates the use Gaussian Belief Propagation (GBP) for probabilistic inference. Similar to Bayesian inference models, the GBP model typically requires less labeled training data than other contemporary machine learning techniques. First, a factor graph is constructed, and a set of Gaussian belief nodes with associated means and variances are fitted to the training data. To better handle noisy input data, the GBP mass prediction model utilizes a k-nearest factors (kNF) algorithm for probabilistic inference on unseen testing data. The proposed method is compared with a classical weighted k-nearest neighbors (kNN) regressor. This statistical kNF-GBP model works even with low-quantity, low-quality initial training data, while being capable of realtime mass estimation. Unlike the kNN regressor, the GBP model produces a measure of uncertainty with its predictions. The proposed method is validated using curve-sampled driving data collected from multiple cloud-connected Class 8 regional haul diesel trucks. Both the kNN regressor and the kNF-GBP mass prediction model were able to predict payload mass with coefficients of determination above 0.97 with minimal data preprocessing.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

PMU-Based Decoupled State Estimation for Unsymmetrical Power Systems

Modal decomposition of measurement equations has already been shown to simplify the formulation and resulting computational complexity of three-phase state estimation of systems where all the transmission lines are three-phase and fully transposed. When there are non-transposed and/or mixed-phase lines, modal decomposition can no longer fully decouple the threephase measurement equations. Here, this paper addresses the above shortcoming by proposing a simple yet practical solution based on the commonly used numerical compensation techniques. Thus, it enables application of the powerful decoupling approach to any type of three-phase networks which may contain non-transposed or mixed-phase lines and are fully observable by PMUs. The proposed procedure modifies the measurement set by deriving additive terms that compensate for the neglected unsymmetrical effects. It will be shown that unbalanced systems including nontransposed and mixed-phase elements, can still be transformed into three decoupled subsystems and solved in parallel by the proposed approach. Performance of the proposed algorithm is validated against several IEEE test cases.

42 ENGINEERING↗

A Data-Driven Voltage Control Strategy for Distribution Grids With Distributed Energy Resources

Traditionally, distribution system control approaches have been model-based. The deployment of advanced metering infrastructure has provided electric utilities with the capability of data-driven control with real-time measurements. The shift from model-based to data-driven control represents a significant advancement in the management of distribution systems, offering a more adaptive approach to system control because of the ability to dynamically adapt to changing conditions without the need for system modeling. Here, in this paper, a behavioral data-driven control method is developed to provide voltage regulation to an actual distribution system by controlling the legacy devices and distributed energy resource (DER) assets. The studied distribution system has a load tap changer and three capacitor banks as the legacy devices and photovoltaic systems as the DERs. The performance of the proposed control algorithm is validated using a laboratory test bed setup considering multiple scenarios. The results show that the proposed control achieved 99% voltage regulation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Flux REaction TArget Prioritization (Flux RETAP) v1

Metabolic engineering is evolving rapidly as a result of new advances in synthetic biology and automation, as well as the irruption of machine learning (ML). ML has been shown to provide the predictive power synthetic biology lacked and needed, and to be able to effectively guide the metabolic engineering process. However, current technical limitations prevent the independent application of ML approaches to metabolic engineering without the use of previous biological knowledge in the form of a prioritized list of desirable engineering targets. Here, we present FluxRETAP, a simple and computationally inexpensive method that leverages the prior mechanistic knowledge embedded in genome-scale metabolic models (GSMs) for suggesting targets for genetic overexpression, downregulation or deletion, with the final goal of increasing metabolite production. FluxRETAP captured 100% of reaction targets experimentally verified to improve Escherichia coli isoprenol production in the literature accessible to us, 50% of targets that experimentally improved taxadiene production in E. coli and ~60% of genetic targets from a verified minimal constrained cut-set in Pseudomonas putida while providing additional high priority targets that could be tested. Overall, FluxRETAP is an efficient algorithm for identifying a prioritized list of testable genetic and reaction targets which can also be utilized in ML pipelines.

Czajka, Jeffrey [Battelle Memorial Institute, Paci↗

Evaluation of KDP Estimation Algorithm Performance in Rain Using a Known-Truth Framework

Accurate estimation of specific differential phase ( K DP ) is necessary for rain rate estimation, attenuation correction, and hydrometeor classification algorithms. There are numerous published methods to process polarimetric radar observations of propagation differential phase shift (Φ DP ) and estimate K DP , but the corresponding K DP estimate uncertainty is unquantified. This study provides guidance on how commonly used K DP estimation algorithms perform in various environments. Here, we create numerous synthetic (“true”) K DP profiles, integrate over them to obtain “smoothed” Φ DP , and then add noise typical of S-band operational weather radar measurements. Each algorithm is applied to our noisy Φ DP profiles and compared to the true K DP profile such that the errors and uncertainty are quantified. The synthetic K DP profiles are Gaussian in shape, which allows systematic variations in their magnitude and width to determine how each algorithm performs in smooth, slowly changing K DP profiles, as well as steep profiles. Results demonstrate that algorithm performance is dependent on the Φ DP field received. These results are further supported by an error analysis of each algorithm for two more complicated synthetic K DP profiles. Some K DP algorithms allow users to change various tuning parameters; a subset of these tuning parameters is tested to provide guidance on how changing these parameters impacts algorithm performance. We then provide evidence that our known-truth framework provides insight into algorithm performance in observed data through two case studies.

54 ENVIRONMENTAL SCIENCES↗

Data from: 'Abiotic influences on continuous conifer forest structure across a subalpine watershed'

This package archives the core data used for analysis and inference in 'Abiotic influences on continuous conifer forest structure across a subalpine watershed' (Worsham et al., 2025). All data were collected in the East River, Washington Gulch, Slate River, and Coal Creek watersheds of Colorado. In the paper, we quantified the relative influence of climate, topographic, edaphic, and geologic factors on conifer stand structure and composition, and their functional relationships, at the watershed scale. We used waveform LiDAR data to derive spatially continuous stand structure metrics. We fused these with a species-level classification map to estimate tree species abundance. We applied generalized additive and generalized boosted models to evaluate the covariability of structural and compositional metrics with abiotic variables. The package contains the essential products required for reproducing our analysis and the tables and figures reported in the publication. The products comprise four classes: (1) geospatial data, (2) tabular data used for inferential analysis, (3) tabular data describing analytical results and performance statistics, and (4) a data user guide. (1) includes discretized waveform LiDAR data, locations and attributes of individual tree crowns, sampling locations and domain boundaries, a canopy height model, and raster files of estimated forest structural and compositional metrics at 100 m grid scale. (2) includes all response and explanatory variable values applied in inferential models. Response variables include conifer forest stand density, basal area, 95th percentile height, quadratic mean diameter, and others. Explanatory variables include climatic water deficit, actual evapotranspiration, elevation, heat load, soil available water content, and others. (3) includes results of training and testing several individual tree detection (ITD) algorithms, as well as inferential modeling results. (4) is a PDF user guide for this data package, including detailed descriptions and data dictionaries for all files. The data package root contains 17 assets: 8 compressed tape archive (.tar.gz) files, 5 comma-separated values (.csv) files, 3 Geographic Tagged Image File Format (GeoTIFF) (.tif) files, and 1 Portable Document Format (.pdf) file. The compressed .tar.gz archives contain ESRI shapefiles (.shp) .tif, compressed LASer (.laz), and .csv files. The archives must first be decompressed using the widely distributed command-line software utility TAR. All other files, including constituent files within the .tar.gz archives, can be opened in the open-source R statistical computing environment. Alternatively, .csv files may also be read in any simple text editor software or Microsoft Excel. Geospatial files including .shp and .tif files can also be opened in GIS software, such as QGIS (open-source) or ESRI ArcGIS (proprietary). The .pdf Data User Guide can be read with Adobe Acrobat Reader or other compatible readers.

2018 NEON and 2025 CHESS Campaigns↗

Data-driven learning of nonlocal models: from high-fidelity simulations to constitutive laws

We show that machine learning can improve the accuracy of simulations of stress waves in one-dimensional composite materials. We propose a data-driven technique to learn nonlocal constitutive laws for stress wave propagation models. The method is an optimization-based technique in which the nonlocal kernel function is approximated via Bernstein polynomials. The kernel, including both its functional form and parameters, is derived so that when used in a nonlocal solver, it generates solutions that closely match high-fidelity data. The optimal kernel therefore acts as a homogenized nonlocal continuum model that accurately reproduces wave motion in a smaller-scale, more detailed model that can include multiple materials. We apply this technique to wave propagation within a heterogeneous bar with a periodic microstructure. Several one-dimensional numerical tests illustrate the accuracy of our algorithm. The optimal kernel is demonstrated to reproduce high-fidelity data for a composite material in applications that are substantially different from the problems used as training data.

97 MATHEMATICS AND COMPUTING↗

Remapping of Data Between One-Dimensional Meshes

In this report we present two approaches to data remapping between one-dimensional meshes implemented with the c++ programming language. Our goal was to test the performance of two search algorithms, linear and binary, and verify the accuracy of our implementations of the two methods. We first introduce the concept of data remap and meshing components, as well as their various uses. We then delve into the differences between point-wise and conservative remap, the algorithms used in the implementations, and lastly confirm the implementations work as intended when given various inputs. We expect that, after profiling, the binary search algorithm will be more efficient than the linear algorithm for sorted sets of data, the point-wise remap implementation to accurately approximate the data transfer between two meshes, and the conservative remap implementation to conserve the area underneath the curve of two distinct meshes.

97 MATHEMATICS AND COMPUTING↗