Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Algorithm testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Multiomics Data Collection, Visualization, and Utilization for Guiding Metabolic Engineering

Biology has changed radically in the past two decades, growing from a purely descriptive science into also a design science. The availability of tools that enable the precise modification of cells, as well as the ability to collect large amounts of multimodal data, open the possibility of sophisticated bioengineering to produce fuels, specialty and commodity chemicals, materials, and other renewable bioproducts. However, despite new tools and exponentially increasing data volumes, synthetic biology cannot yet fulfill its true potential due to our inability to predict the behavior of biological systems. Here, we showcase a set of computational tools that, combined, provide the ability to store, visualize, and leverage multiomics data to predict the outcome of bioengineering efforts. We show how to upload, visualize, and output multiomics data, as well as strain information, into online repositories for several isoprenol-producing strain designs. We then use these data to train machine learning algorithms that recommend new strain designs that are correctly predicted to improve isoprenol production by 23%. This demonstration is done by using synthetic data, as provided by a novel library, that can produce credible multiomics data for testing algorithms and computational tools. In short, this paper provides a step-by-step tutorial to leverage these computational tools to improve production in bioengineered strains.

09 BIOMASS FUELS↗

Predicting Biomass Yields of Advanced Switchgrass Cultivars for Bioenergy and Ecosystem Services Using Machine Learning

The production of advanced perennial bioenergy crops within marginal areas of the agricultural landscape is gaining interest due to its potential to sustainably produce feedstocks for biofuels and bioproducts while also improving the sustainability and resilience of commodity crop production. However, predicting the biomass yields of this production system is challenging because marginal areas are often relatively small and spread around agricultural fields and are typically associated with various abiotic conditions that limit crop production. Machine learning (ML) offers a viable solution as a biomass yield prediction tool because it is suited to predicting relationships with complex functional associations. The objectives of this study were to (1) evaluate the accuracy of commonly applied ML algorithms in agricultural applications for predicting the biomass yields of advanced switchgrass cultivars for bioenergy and ecosystem services and (2) determine the most important biomass yield predictors. Datasets on biomass yield, weather, land marginality, soil properties, and agronomic management were generated from three field study sites in two U.S. Midwest states (Illinois and Iowa) over three growing seasons. The ML algorithms evaluated in the study included random forests (RFs), gradient boosting machines (GBMs), artificial neural networks (ANNs), K-neighbors regressor (KNR), AdaBoost regressor (ABR), and partial least squares regression (PLSR). Coefficient of determination (R 2 ) and mean absolute error (MAE) were used to evaluate the predictive accuracy of the tested algorithms. Results showed that the ensemble methods, RF (R 2 = 0.86, MAE = 0.62 Mg/ha), GBM (R 2 = 0.88, MAE = 0.57 Mg/ha), and GBM (R 2 = 0.78, MAE = 0.66 Mg/ha), were the most accurate in predicting biomass yields of the Independence, Liberty, and Shawnee switchgrass cultivars, respectively. This is in agreement with similar studies that apply ML to multi-feature problems where traditional statistical methods are less applicable and datasets used were considered to be relatively small for ANNs. Consistent with previous studies on switchgrass, the most important predictors of biomass yield included average annual temperature, average growing season temperature, sum of the growing season precipitation, field slope, and elevation. This study helps pave the way for applying ML as a management tool for alternative bioenergy landscapes where understanding agronomic and environmental performance of a multifunctional cropping system seasonally and interannually at the sub-field scale is critical.

09 BIOMASS FUELS↗

PV Validation Hub

The Validation Hub will be a clearinghouse for the transfer of novel algorithms and software from the PV research community to industry. Potential algorithms tested in the Hub could include the estimation of various PV loss factors and the detection of various operational issues. The primary function of the Hub will be for developers to submit executable code which will run on hosted data sets. Developers will receive private reports on the accuracy and performance (e.g., run-time) of the submitted algorithms, and public high level summaries will be hosted. These summaries will indicate the organization who submitted the algorithm (e.g., links to GitHub pages, documentation websites, etc.), high-level accuracy metrics, and standardized performance metrics. These results will be stored in a publicly available database, accessible through the Hub, with the ability for users to sort and filter the results. In short, the Hub will be presented to public users as a collection of interactive leaderboards, organized around specific analysis tasks pertinent to the PV data science community. These tasks include things such as the estimation of various PV loss factors and the detection of various operational issues. We will present progress on the development of this hub, including preliminary results of comparative validation of PV data science algorithms and progress towards building the platform itself.

algorithm↗

Evaluating the potential of short-term instrument deployment to improve distributed wind resource assessment

Distributed wind projects, which are connected at the distribution level of an electricity system or in off-grid applications to serve specific or local energy needs, often rely solely on wind resource models to establish wind speed and energy generation expectations. Historically, anemometer loan programs have provided an affordable avenue for more accurate onsite wind resource assessment, and the lowering cost of lidar systems has shown similar advantages for more recent assessments. While a full 12 months of onsite wind measurement is the standard for correcting model-based long-term wind speed estimates for utility-scale wind farms, the time and capital investment involved in gathering onsite measurements must be reconciled with the energy needs and funding opportunities that drive expedient deployment of distributed wind projects. Much literature exists to quantify the performance of correcting long-term wind speed estimates with 1 or more years of observational data, but few studies explore the impacts of correcting with months-long observational periods. This study aims to answer the question of how short you can go in terms of the observational time period needed to make impactful improvements to model-based long-term wind speed estimates. Three algorithms, multivariable linear regression, adaptive regression splines, and regression trees, are evaluated for their skill at correcting long-term wind resource estimates from the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) using months-long periods of observational data from 66 locations across the US. On average, correction with even 1 month of observations provides significant improvement over the baseline ERA5 wind speed estimates and produces median bias magnitudes and relative errors within 0.22 m s −1 and 4 percentage points of the median bias magnitudes and relative errors achieved using the standard 12 months of data for correction. However, in cases when the shortest observational periods (1 to 2 months) used for correction are not well correlated with the overlapping ERA5 reference, the resultant long-term wind speed errors are worse than those produced using ERA5 without correction. Summer months, which are characterized by weaker relative wind speeds and standard deviations for most of the evaluation sites, tend to produce the worst results for long-term correction using months-long observations. The three tested algorithms perform similarly for long-term wind speed bias; however, regression trees perform notably worse than multivariable linear regression and adaptive regression splines in terms of correlation when using 6 months or less of observational data for correction. Translating the analysis to wind energy, median relative errors in the capacity factor are on average within 10 % using 1 month of training. If the observation period used for correction is not well correlated with the reference data, however, misrepresentation of the observed capacity factor can be substantial. The risk associated with poor correlation between the observed and reference datasets decreases with increasing training period length. In the worst-correlation scenarios, the median capacity factor relative errors from using 1, 3, and 6 months are within 47 %, 26 %, and 16 %, respectively.

17 WIND ENERGY↗

HIV drug resistance during antiretroviral therapy scale-up in Uganda, 2012–19: a population-based, longitudinal study

Background With scale-up of antiretroviral therapy (ART) in sub-Saharan Africa, increasing pretreatment HIV drug resistance has been reported; however, the broader effect of ART expansion on population-level resistance patterns remains insufficiently quantified. We aimed to estimate the longitudinal prevalence of drug resistance and resistance-conferring mutations. Methods This study used data collected as part of the Rakai Community Cohort Study (RCCS), an open population-based census and cohort study conducted in southern Uganda. At each survey round, residents aged 15–49 years are invited to participate and receive a structured questionnaire that obtains sociodemographic, behavioural, and health information, including self-reported past and current ART use. Voluntary HIV testing is conducted using a rapid test algorithm and a venous blood sample. People with HIV provide samples for viral load quantification and deep sequencing. We analysed RCCS survey, HIV viral load, and deep sequencing (which was used to predict resistance) data from five survey rounds. The key outcomes were the population prevalence of viraemic people with HIV with non-nucleoside reverse transcriptase inhibitor (NNRTI), nucleoside reverse transcriptase inhibitor (NRTI), protease inhibitor, or multiclass resistance among all participants (regardless of HIV serostatus) in the 2015 and 2017 surveys. Prevalence of class-specific resistance and resistance-conferring substitutions were estimated using robust log-Poisson regression. Findings Between Aug 10, 2011, and Nov 4, 2020, there were 43 361 participants in the RCCS and 7923 (18·27%) people with HIV. Over five survey rounds, 93 622 participant visits occurred, among which 17 460 (18·65%) were from people with HIV. Over the analysis period, the median age of study participants remained similar (28 years [22–35] in 2012 and 29 years [21–38] in 2019). Sufficient data were available to reliably genotype 4072 (90·03%) of 4523 participant visits from 3407 people with HIV for at least one drug. Overall population prevalence of resistance contributed by viraemic pretreatment people with HIV decreased between 2012 and 2017 from 0·56% (95% CI 0·42–0·75) to 0·25% (0·18–0·33) for NNRTI and from 0·24% (0·15–0·37) to 0·05% (0·02–0·10) for NRTI (prevalence ratio 0·44 [0·29–0·68] for NNRTI and 0·21 [0·09–0·47] for NRTI). Between 2012 and 2017, NNRTI resistance among viraemic pretreatment people with HIV increased from 4·86% (3·69–6·42) to 9·61% (7·27–12·7; prevalence ratio 1·98 [1·34–2·91]). The prevalence of NNRTI and NRTI resistance was substantially higher among viraemic treatment-experienced people with HIV (51·49% [46·24–57·34] for NNRTI and 36·46% [30·06–44·22] for NRTI in 2017) than among pretreatment people with HIV. NNRTI and NRTI resistance was predominantly attributable to rtK103N and rtM184V. inT97A was observed at a similar prevalence among viraemic treatment-experienced (9·96% [6·41–15·48]) and viraemic pretreatment (10·56% [8·01–13·93]) people with HIV; no major dolutegravir resistance mutations were observed. Interpretation Despite rising NNRTI resistance among pretreatment people with HIV, overall population prevalence of pretreatment HIV drug-resistant viraemia decreased due to increasing ART uptake and viral suppression. This finding underscores the crucial role of achieving and maintaining high ART coverage in reducing transmission of drug-resistant HIV. The high prevalence of mutations conferring resistance to components of first-line ART regimens among viraemic people with HIV is potentially concerning. Funding National Institutes of Health, Johns Hopkins University Center for AIDS Research, Bill & Melinda Gates Foundation, and the US Centers for Disease Control and Prevention.

59 BASIC BIOLOGICAL SCIENCES↗

Graph interpolating activation improves both natural and robust accuracies in data-efficient deep learning

Improving the accuracy and robustness of deep neural nets (DNNs) and adapting them to small training data are primary tasks in deep learning (DL) research. In this paper, we replace the output activation function of DNNs, typically the data-agnostic softmax function, with a graph Laplacian-based high-dimensional interpolating function which, in the continuum limit, converges to the solution of a Laplace–Beltrami equation on a high-dimensional manifold. Furthermore, we propose end-to-end training and testing algorithms for this new architecture. The proposed DNN with graph interpolating activation integrates the advantages of both deep learning and manifold learning. Compared to the conventional DNNs with the softmax function as output activation, the new framework demonstrates the following major advantages: First, it is better applicable to data-efficient learning in which we train high capacity DNNs without using a large number of training data. Second, it remarkably improves both natural accuracy on the clean images and robust accuracy on the adversarial images crafted by both white-box and black-box adversarial attacks. Third, it is a natural choice for semi-supervised learning. This paper is a significant extension of our earlier work published in NeurIPS, 2018. For reproducibility, the code is available at https://github.com/BaoWangMath/DNN-DataDependentActivation .

Mathematics↗

Suppressing Quantum Circuit Errors Due to System Variability

We present a quantum circuit optimization technique that takes into account the variability in error rates that is inherent across present-day noisy quantum computing platforms. This method can be run after qubit routing or postcompilation and consists of computing isomorphic subgraphs to input circuits and scoring each using heuristic cost functions derived from system calibration data. Using an independent standard algorithmic test suite, we show that it is possible to recover on average nearly 40% of missing fidelity using better qubit selection via efficient to compute cost functions. We demonstrate additional performance gains by considering qubit placement over multiple quantum processors. The overhead from these tools is minimal with respect to other compilation steps, such as qubit routing, as the number of qubits increases. As such, our method can be used to find qubit mappings for problems at the scale of quantum advantage and beyond.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Sandbox for Outer Loop Analysis (SOLA)

SAND2026-18834O Sandbox for Outer Loop Analysis (SOLA) is an object-oriented Matlab library that prototypes outer loop analysis algorithms. It serves as a platform for rapid idea exploration, algorithm testing, and enhancing pedagogy. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

van Bloemen Waanders, Bart [Sandia National Lab. (↗

Data Science and Machine Learning for Genome Security

This report describes research conducted to use data science and machine learning methods to distinguish targeted genome editing versus natural mutation and sequencer machine noise. Genome editing capabilities have been around for more than 20 years, and the efficiencies of these techniques has improved dramatically in the last 5+ years, notably with the rise of CRISPR-Cas technology. Whether or not a specific genome has been the target of an edit is concern for U.S. national security. The research detailed in this report provides first steps to address this concern. A large amount of data is necessary in our research, thus we invested considerable time collecting and processing it. We use an ensemble of decision tree and deep neural network machine learning methods as well as anomaly detection to detect genome edits given either whole exome or genome DNA reads. The edit detection results we obtained with our algorithms tested against samples held out during training of our methods are significantly better than random guessing, achieving high F1 and recall scores as well as with precision overall.

59 BASIC BIOLOGICAL SCIENCES↗

A Novel Machine Learning Approach to Disentangle Multitemperature Regions in Galaxy Clusters

The hot intracluster medium (ICM) surrounding the heart of galaxy clusters is a complex medium that comprises various emitting components. Although previous studies of nearby galaxy clusters, such as the Perseus, the Coma, or the Virgo cluster, have demonstrated the need for multiple thermal components when spectroscopically fitting the ICM’s X-ray emission, no systematic methodology for calculating the number of underlying components currently exists. In turn, underestimating or overestimating the number of components can cause systematic errors in the emission parameter estimations. In this paper, we present a novel approach to determining the number of components using an amalgam of machine learning techniques. Synthetic spectra containing a various number of underlying thermal components were created using well-established tools available from the Chandra X-ray Observatory. The dimensions of the training set was initially reduced using principal component analysis and then categorized based on the number of underlying components using a random forest classifier. Our trained and tested algorithm was subsequently applied to Chandra X-ray observations of the Perseus cluster. Our results demonstrate that machine learning techniques can efficiently and reliably estimate the number of underlying thermal components in the spectra of galaxy clusters, regardless of the thermal model (MEKAL versus APEC). We also confirm that the core of the Perseus cluster contains a mix of differing underlying thermal components. We emphasize that although this methodology was trained and applied on Chandra X-ray observations, it is readily portable to other current (e.g., XMM-Newton, eROSITA) and upcoming (e.g., Athena, Lynx, XRISM) X-ray telescopes. The code is publicly available at https://github.com/XtraAstronomy/Pumpkin.

79 ASTRONOMY AND ASTROPHYSICS↗

Closed Loop Testing of Microphonics Algorithms Using a Cavity Emulator

An analog crystal filter based cavity emulator is modified with reverse biased varactor diodes to provide a tuning range of around 160 Hz. The piezo drive voltage of the resonance controller is used to detune the cavity through the bias voltage. A signal conditioning and summing circuit allows the introduction of microphonics disturbance from a signal source or using real microphonics data from cavity testing. This setup is used in closed loop with a cavity controller and resonance controller to study the effectiveness of resonance control algorithms suitable for superconducting cavities.

43 PARTICLE ACCELERATORS↗

Comprehensive Efficiency Analysis of Current Source Inverter Based on CSI-Type Double Pulse Test and Genetic Algorithm

A current-source inverter (CSI) has a natural output voltage boost feature that can be advantageous for traction applications. Switching frequency is one of the easily-controlled variables that can be adjusted to improve the CSI efficiency at different operating points when the boost function is used. This paper investigates a CSI-based double-pulse test (DPT) measurement that mimics normal operation of the CSI. A loss model of the CSI is developed based on the CSI-type DPT experimental results. The impacts of the switching frequency and voltage boost ratio on the CSI efficiency and output voltage ripple are investigated. Based on the loss model, a genetic algorithm has been introduced that makes it possible to optimize the CSI's switching frequency and modulation index to maximize its efficiency under any desired operating condition.

current-source inverter, genetic algorithm, SPM ma↗

Space Nuclear Power Autonomous Control Algorithm and Control Element Test Bed

Nuclear thermal rockets are currently NASA’s preferred option for use in a manned mission to Mars in the 2040s. The communication delay between an Earth ground station and a spacecraft heading toward Mars can be up to 20 min. Therefore, controlling the nuclear rocket engine would require either a full-time reactor operator on the mission or an autonomous control system for the reactor. The latter idea of making space nuclear reactors fully autonomous has drawn more interest from stakeholders, but such an autonomous control system must be rigorously tested and validated before it is certified for human use. The cost of a full ground test for a space nuclear reactor is tremendous, so a nonnuclear mock reactor test bed was created to test and validate control elements and control algorithms for space nuclear reactors. The test bed consists of control element hardware that inputs physical measurement data into a reactor emulator to produce the reactor’s performance under steady-state, transient, and fault conditions. The control element hardware consists of six full-sized control drums equipped with servo drives and motors and is instrumented with optical encoders, resolvers, and torque sensors for drum movement characterization. In addition to the drums, a two-phase flow loop was designed and built to mimic the valves and turbomachinery associated with the propellant flow through a nuclear thermal rocket engine; components such as pressure sensors, flow meters, thermocouples, and tachometers are instrumented throughout the loop to characterize the fluid flow, valve, and turbomachinery behavior of the system. The data from the physical hardware (e.g., drum position, propellant flow rates) are input to a nuclear reactor simulator to determine the actual nuclear reactor parameters, and the data are sent back to a control algorithm to complete the control loop. The ability to conduct numerous tests of the control systems and autonomous algorithms can help validate the instrumentation and control aspects for a space nuclear reactor for every possible fault situation.

Wilson, Brandon↗

Improving convergence of the Matrix Power Control Algorithm for random vibration testing

Herein, this paper describes modifications to the Matrix Power Control Algorithm (MPCA) to improve convergence for Random Vibration Control (RVC) testing. In particular, this paper presents Multiple-Input Multiple-Output (MIMO) implementations of MPCA in simulation and experiment. An Euler–Bernoulli beam model was simulated with applied base excitations and the Box Assembly with Removable Component (BARC) was used in experiment to validate results. The Bayesian optimization package Dragonfly was used to optimize control parameters. Additionally, a moving-average was employed and optimized to improve the measured response feedback for MPCA, reduce the number of averages needed to be taken between control updates, and further improve convergence. The key results of this paper show that the performance of MPCA can be improved by tuning the control parameters and by applying an optimized moving-average. Furthermore, it is demonstrated that convergence can be achieved within 12 drive-frames, which greatly enhances vibration control capability.

42 ENGINEERING↗

Testing and validating SMDS algorithms implemented in the cloud

Pacific Northwest National Laboratory (PNNL) provided technical assistance to NorthWrite Inc. under the Small Business Vouchers (SBV) Pilot. NorthWrite delivered services to owners of small commercial buildings, using a cloud-based service to monitor, control, and optimize building operations, saving energy and reducing operating costs for the owners while ensuring that occupant comfort needs were consistently met. NorthWrite had a longstanding desire to add a new suite of diagnostic capabilities to their service offering and had been trying, without success, to incorporate several diagnostic algorithms published by PNNL. These algorithms arise from research previously supported by BTO. Through the SBV awarded to NorthWrite, PNNL made available technical knowledge regarding the derivation and application of the following sets of algorithms for use in the NorthWrite Cloud-based service delivery system: • algorithms for monitoring rooftop packaged air conditioners and heat pumps (often referred to as rooftop units or RTUs) and diagnosing faults in these units, • advanced algorithms for automated fault detection and diagnosis of other equipment found in small buildings, and • algorithms for identifying, prioritizing, and assessing the economic effectiveness of implementing energy saving measures in small commercial buildings based on sensed data. The PNNL researchers involved were the original developers of these algorithms, had a unique understanding of the derivation of the algorithms, and had the source data used for this derivation. Further, the PNNL researchers had extensive experience using these algorithms in the laboratory, unique experience applying these algorithms in real-world small commercial buildings and could solve a number of key problems that NorthWrite and other users faced in using these algorithms at scale. In addition, PNNL had recently installed a pair of RTUs in a laboratory setting that were instrumented and connected with data acquisition systems that provided a unique test rig for validating the algorithms before NorthWrite began to deploy their new services in the field.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A set of test problems for nonlinear optimization

The software is a set of test problems for nonlinear optimization algorithms, including subroutines such as linear algebra routines and automatic differentiation algorithms. The test problems come from chemical engineering open literature, and describe optimization tasks related to the design and operation of processes such as carbon capture, Hydrogen production, heat exchange, and distillation.

Parker, Robert↗

Seismic Spatial Gradients and Machine Learning-Based Classifiers for Explosion Monitoring (LDRD 218327)

This final report summarizes the work completed under the Laboratory Directed Research and Development (LDRD) project “Seismic Spatial Gradients as a Machine Learning-Based Classifier for Explosion Monitoring.” The overarching goal of the project was to explore the efficacy of using machine learning-based classification algorithms where the input data are the spatial gradient of the seismic wavefield collected at a single point on the Earth’s surface. The methods that I describe here are in direct contrast to conventional methods of seismic discrimination which typically rely on a spatially extended network of instruments and physics-based wavefield attributes such as, for example, the ratio between $\textit{P}$ and $\textit{S}$ waves. Rather, we use the spatial gradient of the seismic wavefield observed at a single point on the Earth’s surface and data processing approaches inspired by the machine learning community. We tested two algorithms, a neural network and a modified version of principal component analysis termed Spectrally Filtered Principal Component Analysis (SFPCA). To test these algorithms, we first conducted a series of numerical tests using synthetic data and then conducted a small-scale controlled field experiment. The tests using synthetic data showed that both algorithms had high success rates on gradiometric data, even when simulated noise was added to the signal. Furthermore, we found that using seismic spatial gradients increased the performance of our discrimination algorithms when compared to using just the traditional translational motion seismic data. The tests with field data also showed a high degree of discriminative success.

58 GEOSCIENCES↗